Use of a combination of non-error-propagation phasing techniques and allelic balance to improve CNV detection
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- マイオームインコーポレイテッド
- Filing Date
- 2021-10-29
- Publication Date
- 2026-05-12
AI Technical Summary
Current methods for detecting copy number variants (CNVs) are hindered by noise from mixed normal and abnormal tissues, limited dynamic range in sequencing data, and skewed variant allele balance due to uneven amplification, making accurate detection of chromosomal deletions and duplications challenging, particularly in diagnosing diseases like cancer or fetal conditions.
A method involving non-error propagation techniques such as chromosome conformation capture and single cell template strand sequencing is used to determine allelic balance signals, correcting for phase errors by aligning phase sets and combining them with read depth signals to improve ploidy status determination.
This approach enhances the accuracy of CNV detection by reducing false positives and negatives, allowing for precise identification of chromosomal abnormalities and informing treatment decisions, including cancer therapy and prenatal diagnostics.
Smart Images

Figure 00000047_0000 
Figure 00000047_0001 
Figure 00000048_0000
Abstract
Description
[Background technology]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of U.S. Provisional Application No. 63 / 107,464, filed October 30, 2020, which is incorporated herein by reference in its entirety.
[0002] background Copy number variation (CNV) can be an important indicator of disease and disease progression. CNVs have been identified as a major source of structural variation in the genome and typically include both sequence duplications and deletions, ranging in length from 1 kb to 20 Mb. Deletions and duplications of chromosomal segments or entire chromosomes are associated with various conditions, such as disease susceptibility or resistance. However, methods for identifying CNVs remain challenging and are complicated by multiple issues. In some instances, normal and abnormal tissues (containing one or more CNVs) are mixed together, generating noise that precludes the detection of one or more CNVs. Also, available sequencing data may have a limited dynamic range. Furthermore, uneven amplification due to resampling bias can result in skewed variant allele balance.
[0003] Therefore, there is a need for improved methods for more accurately detecting deletions and duplications of chromosome segments or entire chromosomes, including CNVs. Preferably, these methods can be used to more accurately diagnose diseases or increased risks of diseases, such as cancer or CNVs, in fetuses during pregnancy. Summary of the Invention
[0004] overview According to one aspect of the present invention, a method for correcting an allele balance signal for a chromosome segment is disclosed herein. The method includes obtaining a reference genetic code having at least two phase sets, each of which can be at least partially phase-determined. Each phase set contains one or more variants of interest. The method further includes obtaining an allele balance signal for one or more variants of interest from sequencing performed on a sample of genetic material, and obtaining a plurality of reads sequenced using a non-error propagation technique. Each read contains at least one of the one or more variants of interest. The phase alignment of the two phase sets is then determined based on the plurality of reads as being in the same phase or different phases, and the true allele balance signal is determined by confirming, correcting, or providing the phase state of at least one variant of interest based on the determined phase alignment of the two phase sets.
[0005] Non-error propagation techniques may include conformational capture, single-cell template strand sequencing, or chromosome isolation (e.g., via laser capture microdissection or karyotyping). The method may include performing a non-error propagation technique to obtain multiple reads. The method may include sequencing a sample of genetic material to obtain allele balance signals.
[0006] The allele balance signal and the multiple reads can be derived from the same sample of genetic material. The sample can be a body fluid sample (e.g., a blood sample, a saliva sample) or a tissue biopsy sample. The allele balance signal and the multiple reads can be derived from the same cell population. The allele balance signal can be derived from multiple reads derived from extracellular DNA and cellular DNA. The cellular DNA can be derived from cells found in body fluids (e.g., blood or saliva).
[0007] The reference genetic code can be derived from the sequencing used to generate allele balance signal.The reference genetic code can be derived at least in part from the sequencing of the normal tissue of the subject from which the allele balance signal is obtained, from the sequencing of the germline tissue of the subject, or from the sequencing of the genetic material of one or more genetic relatives of the subject.The one or more relatives can be the mother and / or father of the subject.The reference genetic code can be derived at least in part from the germline sequencing of one or more genetic relatives.
[0008] The reference genetic code may be derived, at least in part, from whole genome shotgun sequencing of the subject. The allele balance signal may be derived from whole genome shotgun sequencing. In either case, whole genome shotgun sequencing may be performed on extracellular DNA in a body fluid sample (e.g., a blood sample or a saliva sample). The non-error propagation technique may include single-cell sequencing. The method may further include collecting a sample of genetic material from which the allele balance signal is derived, and / or collecting a sample of genetic material from which multiple reads are derived.
[0009] Correcting allele balance data can include correcting switch errors in at least partially phased reference genetic codes. Allele balance signals can be averaged across multiple binned variants within a region of about 50,000, about 100,000, about 200,000, about 300,000, about 400,000, about 500,000, about 750,000, about 1 million, about 50 million, or about 100 million, at least about 50,000, at least about 100,000, at least about 200,000, at least about 300,000, at least about 400,000, at least about 500,000, at least about 750,000, at least about 1 million, at least about 50 million, or at least about 100 million base pairs, or about 50,000 or less, about 100,000 or less, about 200,000 or less, about 300,000 or less, about 400,000 or less, about 500,000 or less, about 750,000 or less, about 1 million or less, about 50 million or less, or about 100 million or less base pairs. Allele balance can be averaged across one or more haplotype blocks. One or more haplotype blocks can be determined by dilution pool sequencing. The allele balance signal can be derived from the same sequencing used to determine one or more haplotype blocks. The allele balance signal can be filtered for a minimum read depth, such as a minimum read depth of 5, 10, 15, 20 or 25 reads.
[0010] Two phase sets can be adjacent phase sets in reference genetic code.For example, each adjacent phase set can comprise the variant of interest that is not more than about 1,000, about 5,000, about 10,000, about 50,000, about 100,000, about 5 million, about 1 million, about 5 million, about 10 million, about 50 million, about 100 million or about 250 million base pairs away from the variant of interest in the other set.Multiple reads can be filtered for the reads that comprise at least 2, 3, 4 or 5 of the variant of interest from each of the two phase sets.
[0011] The non-error propagation technique may specifically include chromosome conformation capture. The chromosome conformation capture technique may be Hi-C. Determining a phase alignment based on multiple reads may involve determining whether a majority of the reads are matched or mismatched with respect to an estimated phase state alignment between two phase sets, and the estimated phase state alignment between the two phase sets may be based on at least a partial phase state of a reference genetic code. Determining a phase alignment based on multiple reads may include determining or estimating the probability that the amount of match or mismatch observed between the two phase sets from the multiple reads is the result of chance. Optionally, the probability may be a binomial probability, which assumes that the observed fragments are equally likely to be matched or mismatched.
[0012] The method can further include using the corrected allele balance signal to determine a ploidy state for the chromosome segment. For example, determining the ploidy state can be calling copy number variations (CNVs).
[0013] According to another aspect of the present invention, a method for determining the ploidy state of a chromosome segment is disclosed herein. The method includes: obtaining read depth signals for a first set of one or more variants in the chromosome segment; obtaining allele balance signals for a second set of one or more variants in the chromosome segment; and using the read depth signals in combination with the allele balance signals to determine the ploidy state of the chromosome segment.
[0014] Determining the ploidy state of the chromosome segment can include determining whether a CNV is present within the chromosome segment. Obtaining a read depth signal can include obtaining the number of sequencing reads mapped to at least one of the variants in the first set normalized to the total number of reads. The depth signal and / or allele balance signal of the reads can be averaged over a plurality of binned variants within a region of about 50,000, about 100,000, about 200,000, about 300,000, about 400,000, about 500,000, about 750,000, about 1 million, about 50 million, or about 100 million, at least about 50,000, at least about 100,000, at least about 200,000, at least about 300,000, at least about 400,000, at least about 500,000, at least about 750,000, at least about 1 million, at least about 50 million, or at least about 100 million base pairs or less. The depth signal and / or allele balance signal of the reads can be averaged over one or more haplotype blocks. The one or more haplotype blocks may be those determined by dilution pool sequencing. The depth and allele balance signals of the reads may be averaged over the same binned region.
[0015] Using the read depth signal in combination with the allele balance signal may include making a positive or negative determination only if the read depth signal exceeds the read depth threshold and the allele balance signal exceeds the allele balance threshold, or if the read depth signal does not exceed the read depth threshold and the allele balance signal does not exceed the allele balance threshold. Using the read depth signal in combination with the allele balance signal may include integrating the read depth signal and the allele balance signal into a single integrated signal. Integrating the read depth signal and the allele balance signal into a single integrated signal may include multiplying the signals or adding the signals together. The integrated signal can be averaged over multiple binned variants within a region of about 50,000, about 100,000, about 200,000, about 300,000, about 400,000, about 500,000, about 750,000, about 1 million, about 50 million, or about 100 million, at least about 50,000, at least about 100,000, at least about 200,000, at least about 300,000, at least about 400,000, at least about 500,000, at least about 750,000, at least about 1 million, at least about 50 million, or at least about 100 million base pairs or less. The integrated signal can be averaged over one or more haplotype blocks, which may be determined by dilution pool sequencing. The integrated signal can be averaged over multiple bins where the read depth signal and / or allele balance signal are averaged.
[0016] The first set of one or more variants may consist of only one variant. The first set of one or more variants may have at least 2, at least 3, at least 4, at least 5, at least 10, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 200, at least 300, at least 400, at least 500, at least 600, at least 700, at least 800, at least 900, or at least 1,000 variants. The second set of one or more variants may consist of only one variant. The second set of one or more variants may have at least 2, at least 3, at least 4, at least 5, at least 10, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 200, at least 300, at least 400, at least 500, at least 600, at least 700, at least 800, at least 900 or at least 1,000 variants. The first set of one or more variants may be identical to the second set of one or more variants.
[0017] Obtaining a read depth signal and / or obtaining an allele balance signal can include sequencing.The read depth signal and the allele balance signal can be derived from the same sequencing data.The read depth signal and / or the allele balance signal can be filtered for a minimum read depth, such as a minimum read depth of 5, 10, 15, 20 or 25 reads.
[0018] The method may include calculating individual probabilities of correct determination of ploidy state based on depth signals and / or allele balance signals of the reads, or calculating a joint probability of correct determination of ploidy state based on depth signals and allele balance signals of the reads. The probabilities may, for example, measure the probability of one of the following: true positive, false positive, true negative, and false negative. It may be determined that at least one of the following is true: the joint probability of false positives is less than both individual probabilities of false positives, the joint probability of false negatives is less than both individual probabilities of false negatives, the joint probability of true positives is greater than both individual probabilities of true positives, or the joint probability of true negatives is greater than both individual probabilities of true negatives.
[0019] The read depth signal can be offset against a first baseline signal, and / or the allele balance signal can be offset against a second baseline signal. Each baseline signal can be based on the average signal for a second chromosome segment with a known ploidy state. The second chromosome segment can be within the same chromosome as the chromosome segment whose ploidy state is being determined. The read depth signal and / or allele balance signal can be normalized to a measure of noise in the signal. The measure of noise can be the standard deviation or variance of the signal across the chromosome segment whose ploidy state is being determined, across the second chromosome segment with a known ploidy state, across a third chromosome segment with a known ploidy state of interest that is different from the ploidy state of the second chromosome segment, or across the entire chromosome. The variance in the depth signal of the reads and the variance in the allele balance signal can be within 100, 90, 80, 70, 60, 50, 40, 30, 20, 10, 9, 8, 7, 6, 5, 4, 3, 2, 1.9, 1.8, 1.7, 1.6, 1.5, 1.4, 1.3, 1.2, or 1.1 fold of each other. Use of a read depth signal in combination with an allele balance signal may result in a reduction in the false positive and / or false negative rate of at least about 2-fold, at least about 3-fold, at least about 4-fold, at least about 5-fold, at least about 6-fold, at least about 7-fold, at least about 8-fold, at least about 9-fold, at least about 10-fold, at least about 20-fold, at least about 30-fold, at least about 40-fold, at least about 50-fold, at least about 60-fold, at least about 70-fold, at least about 80-fold, at least about 90-fold, at least about 100-fold, at least about 150-fold, at least about 200-fold, at least about 250-fold, or at least about 500-fold, compared to the false positive and / or false negative rates obtained using one or both of the signals individually.
[0020] Using the read depth signal in combination with the allele balance signal can include selecting a read depth threshold and an allele balance threshold. Each signal threshold can be calculated as half the mean value of each signal averaged over multiple variants known to indicate the ploidy state of interest (e.g., aneuploidy). Using the read depth signal in combination with the allele balance signal can include selecting an integrated signal threshold. The integrated signal threshold can be calculated as half the mean value of the integrated signal averaged over multiple variants known to indicate the ploidy state of interest (e.g., aneuploidy).
[0021] The method may result in one or more chromosomal aneuploidies being detected.
[0022] The method can result in the euploidy of all analyzed chromosomes being detected. The method can result in the addition and / or deletion of chromosome segments being detected. The method can result in the CNV being identified.
[0023] Obtaining the allele balance signal may include correcting the original allele balance signal by performing any one of the above-mentioned methods for making such a correction, as described elsewhere herein.
[0024] According to another aspect of the present invention, any of the above-mentioned methods can include obtaining a signal (for example, an allele balance signal or a read depth signal) indicating ploidy state from a sample containing a population of cells with different copy numbers for chromosome segments.Some of the cells in the population of cells can have aneuploidy, while other cells can not.The signal can be derived from a sample containing one or more tumor cells.The sample can further contain non-tumor cells.
[0025] According to another aspect of the present invention, any of the above-described methods may include obtaining a signal indicating ploidy state (e.g., an allele balance signal or a read depth signal) derived from extracellular DNA. The extracellular DNA may be extracellular fetal DNA (cffDNA) or circulating tumor DNA (ctDNA).
[0026] According to another aspect of the present invention, any of the above-described methods can include obtaining a signal indicative of ploidy state (e.g., an allele balance signal or a read depth signal) from an embryo or fetus. The embryo can be an embryo present in vitro, e.g., prior to implantation of the embryo in a uterus.
[0027] According to another aspect of the present invention, a method for detecting chromosomal instability in tumor DNA is disclosed herein.This method comprises determining the ploidy state of one or more chromosomal segments in a sample of genetic material according to any one of the above-mentioned methods for determining the ploidy state.The sample of genetic material is at least partially derived from DNA originating from one or more cells that are known to be tumor cells or suspected to be tumor cells.The identification of the aneuploidy state of one or more chromosomal segments is used to indicate the chromosomal instability of at least some tumor cells.
[0028] The sample may be derived from a subject diagnosed with or suspected of having cancer. The sample may contain circulating tumor DNA. To establish a reference genetic code, sequencing of normal tissue (e.g., germline tissue) or tumor tissue from the subject from which the genetic material was obtained may be used. The method may further include treating one or more cells or subjects from which the genetic material was obtained for cancer based on whether chromosomal instability is demonstrated. If chromosomal instability is demonstrated, the treatment may include administering a poly ADP-ribose polymerase (PARP) inhibitor and / or a platinum-based chemotherapeutic agent to one or more cells or subjects.
[0029] According to another aspect of the present invention, disclosed herein is a method for detecting de novo copy number variation (CNV) in subject.This method comprises determining the ploidy state according to any one of the above-mentioned methods for determining the ploidy state of chromosome segment.The parent of subject is euploid for chromosome segment.By carrying out this method, de novo aneuploid (such as CNV) can be identified in the chromosome segment of subject.
[0030] Determining the ploidy state may involve comparing the ploidy state to a reference genetic code derived from sequencing performed on one or more genetic relatives of the subject.
[0031] The one or more genetic relatives can be the mother and / or father of the subject.Sequencing can be carried out using non-error propagation technology to provide multiple readings according to any one of the above-mentioned methods for providing multiple readings.Sequencing can be carried out on cellular DNA.This method can further comprise determining whether the mother or father of the subject is the cause of aneuploidy.
[0032] The subject may be an embryo. The method may include obtaining a signal indicating ploidy state (e.g., an allele balance signal or a read depth signal) from an embryo biopsy, blastocoelic fluid, or cell culture medium (extracellular DNA in the culture medium). The method may further include selecting an embryo based on the absence or presence of aneuploidy. The embryo may be selected from a plurality of embryos. The selected embryo may be used for in vitro fertilization (IVF), disposed of, or frozen.
[0033] The subject may be a fetus. The method may include obtaining a signal indicating a ploidy state (e.g., an allele balance signal or a read depth signal) derived from extracellular fetal DNA (cffDNA). The method may include treating the fetus and / or the mother based on the identified absence or presence of aneuploidy (e.g., CNV). The treatment may include performing additional tests on the fetus, such as karyotype analysis. The treatment may include terminating the pregnancy. The treatment may include administering prenatal treatment to the fetus for a disease associated with the presence of the detected aneuploidy (e.g., CNV).
[0034] According to another aspect of the present invention, disclosed herein is a method for screening a subject for disease.This method comprises determining whether one or more genetic variants associated with disease exist.The one or more genetic variants include the aneuploidy (for example, CNV) and / or SNP that exist in the same haplotype block as the aneuploidy, which are identified by carrying out any one of the above-mentioned methods for determining ploidy status on one or more other subjects.SNP can be known to be associated with disease.
[0035] CNVs and SNPs may be in linkage disequilibrium. Determining whether one or more genetic variants associated with a disease exist may include sequencing the subject. A portion of the genome containing one or more genetic variants may be targeted for sequencing (e.g., via microarray). The method may include calculating a polygenic risk score (PRS) for the disease based at least in part on one or more genetic variants. The method may further include diagnosing the subject with the disease based at least in part on the presence or absence of one or more genetic variants or based on a PRS based at least in part on one or more genetic variants. The method may include treating the subject based on the presence or absence of one or more genetic variants.
[0036] According to another aspect of the present invention, a method for determining the phase of a germline mosaic variant in a subject is disclosed herein. The method includes obtaining a reference genetic code having at least two phase sets. Each phase set has one or more variants of interest. The reference genetic code can be at least partially phase-determined. The method further includes obtaining a plurality of reads sequenced using a non-error propagation technique. Each read includes at least one of the one or more variants of interest. The phase alignment of the two phase sets is determined based on the plurality of reads as being the same phase or different phases, and a haplotype encompassing a chromosomal segment showing aneuploidy (e.g., CNV) is identified based on the determined phase alignment of the two phase sets.
[0037] The subject may be diagnosed with or suspected of having a genetic disease or condition associated with aneuploidy. The subject may have been diagnosed with or be suspected of having Noonan syndrome or rasopathy. The method may further include screening gametes from the subject for the identified haplotype. The method may further include selecting gametes that do not have the identified haplotype for in vitro fertilization. The method may include screening for the haplotype in embryos during preimplantation genetic testing. The method may include selecting embryos based on the absence or presence of aneuploidy. The embryo may be selected from a plurality of embryos. The method may include using the selected embryo in in vitro fertilization (IVF), disposing of the selected embryo, or freezing the selected embryo. Aneuploidy may be identified by performing any one of the above methods for determining ploidy status. [Brief explanation of the drawings]
[0038] [Figure 1] FIG. 1 shows simulated allele balance data for human chromosome 21, which has an amplification between approximately nucleotide positions 30.2 Mb and 44.3 Mb. [Figure 2] Figure 2 shows the simulated allele balance data when averaged across haplotype blocks. The arrow indicates the approximate location of a switch error in the input phased genotype data that would cause the appearance of monosomy, rather than trisomy, in the chromosome downstream of the actual simulated switch error. [Figure 3] Figure 3 shows simulated allele balance data when averaged over a 300 Kb window of haplotype blocks, illustrated at the bottom of the figure, across the region of the chromosome where aneuploidy is detected. [Figure 4] FIG. 4 shows a summary of the Hi-C data for the genetic sample from which the allele balance data was simulated. [Figure 5] Figure 5 shows the true allele balance signal after the switch errors have been corrected. [Figure 6] Figures 6A-6B illustrate simulated true allele balance signals for a scenario involving a mixture of chromosomes containing normal disomic and abnormal trisomic regions. Figure 6A shows the signal for individual measurements, while Figure 6B shows the signal when averaged over haplotype blocks. [Figure 7] Figure 7 shows schematically the population of disomic measurements and the population of trisomic measurements (shaded) as normal distributions spread across two different signals, X1 and X2, where m1 and m2 refer to the mean measurements for the trisomic population (trisomic region of the chromosome). [Figure 8] Figures 8A-8B show read depth data for regions of a chromosome with simulated amplification. Figure 8A shows the raw depth signal for each indexed location, and Figure 8B illustrates a histogram showing the proportion of measurements for various binned read depths. [Figure 9]Figures 9A-9C show allele balance data for regions of a chromosome with simulated amplification. Figure 9A shows the raw allele balance signal for each indexed position, and Figure 9B illustrates a histogram showing the frequency of measurements for various binned proportions of the A allele. Figure 9C further shows a histogram in which measurements are averaged across 50 adjacent SNPs. [Figure 10] Figure 10 shows the depth signal of the reads across the simulated amplification (trisomy) between positions 30 Mb and 37 Mb, offset against the depth signal of the disomic reads and normalized against the noise (standard deviation) of the depth signal of the trisomic reads. [Figure 11] Figure 11 shows the allele balance signal across the simulated amplification (trisomy) between positions 30 Mb and 37 Mb, offset against the allele balance signal of disomy and normalized against the noise (standard deviation) of the allele balance signal of trisomy. [Figure 12] FIG. 12 shows the integration of additive, counterbalanced, and normalized read depth and allele balance signals. DETAILED DESCRIPTION OF THE INVENTION
[0039] Detailed Description Disclosed herein are methods for improved determination of ploidy state by applying nucleotide sequencing methods that are non-error-propagating in nature to clarify the phase of one or more regions of a genetic code of interest (e.g., a genome of interest), particularly regions that may contain switch errors introduced from previous error-propagation phasing techniques. A phase alignment determined between two or more variants of interest via a non-error-propagation method can be combined with existing phase information for the genetic code of interest. In some instances, the determined phase alignment can be used to correct the phase state of one or more variants of interest that were incorrectly phased (e.g., from a phasing technique that introduced switch errors). In some instances, the determined phase alignment can be used to confirm that the estimated phase state of one or more variants is the true phase state. In some instances, the determined phase alignment can be used to supply missing phase information. Phase state information for a portion of a genetic code of interest that was at least partially determined by a non-error-propagation method can be used to (re)analyze allele balance signals. The true allelic balance signals obtained from using the non-error-propagating phasing method can be used to make improved determinations of ploidy state, such as CNV calling. In certain embodiments, the improved phase state alignment can be used to determine whether an allelic balance signal indicating a shift in allelic balance relative to a reference haplotype corresponds to a deletion or amplification within the genetic code of interest.
[0040] Also disclosed herein is a method for improving the determination of ploidy state by combining allele balance signals with read depth signals.Such signals provide independent information that can improve signal-to-noise ratio and reduce the probability of false positive and / or false negative calls.The combined use can be particularly powerful when the allele balance signals are corrected through a non-error propagation phase determination approach to provide true allele balance signals.
[0041] Phase status and switch errors A switch error occurs when a variant position is imprecisely phased relative to its neighboring variants. As used herein, "variant" can refer to any difference between the sequences of two or more homologous chromosomes, including single nucleotide polymorphisms (SNPs). As used herein, a variant does not imply a sufficiently low frequency in a larger population, unless otherwise indicated by the context. Phasing accuracy can be measured by counting the number of switch errors that occur divided by the number of switch error opportunities, known as the "switch error rate." Switch errors can be classified as long switch errors, point switch errors, or undetermined switch errors. A long switch appears as a large-scale pseudo-recombination event in which there are no other local switches surrounding the long switch (e.g., no other switches within three consecutive heterozygous sites). A point switch is a small-scale switch error that appears as two adjacent switch errors (e.g., two switches within three consecutive heterozygous sites; a pair of switches is counted as one point switch). The remaining switches are considered undetermined (e.g., only two sites within a small phasing block were phased, and therefore the switch error could not be classified as long or point). Because switch errors propagate across larger portions of the genome (e.g., a second switch error in a co-switch restores the nucleotide downstream of the co-switch to its original / proper phase state, so the phase state of loci far downstream from the co-switch is not affected by the co-switch error), long switches are particularly detrimental to genomic analyses that rely on the phase state of loci. Long switch errors can appear as spurious recombination events induced in inferred haplotypes, particularly compared to true haplotypes. A significant limitation of using phase sets has been the presence of long switch errors. These errors directly affect the sensitivity of detecting small (e.g., less than approximately 1 Mb) deletions or amplifications. In contrast to isolated phasing misevents, switch errors can directly affect the relationships of all downstream loci to upstream loci and / or all upstream loci to downstream loci.Regions of the genome with low polymorphism or SNV density are particularly prone to switch errors when phased.
[0042] Population-based phasing approaches, which rely on computational inference of phase from statistical analysis of populations, generally have a higher switching error rate than molecular phasing approaches. However, molecular phasing approaches may be prone to switching errors. For example, many molecular phasing approaches rely on the computational construction of synthetic long reads from short reads, which relies on statistically derived inferences about the alignment of short reads to the genome. For example, haplotype determination based on diluted pool sequencing relies on a low molar concentration of molecules per given compartment to reduce the likelihood that one DNA molecule in the compartment has a sequence that overlaps with another DNA molecule. While such assumptions allow for the acquisition of at least some haplotypes, they may introduce switching errors when performing long-range phasing (e.g., phasing an entire chromosome). To find the most likely phase alignment, several assumptions about the phase alignment of distant variants may be made, which may allow for the introduction of switching errors.
[0043] Phasing approaches that rely directly on determining the proximity of two or more loci in an intact chromosome and phasing one or more variants at these loci relative to each other are generally less prone to switch errors, because the phase alignment is determined by experimental information that directly links one variant to another, and the phase alignment is not based on inferences related to the phase state of more distant variants. Thus, even if a phase determination error occurs using such an approach, the error will not necessarily be propagated to other more distant loci (e.g., downstream loci). Therefore, such "non-error propagation" methods provide an independent phase determination approach compared to population-based phase determination approaches and molecular phase determination approaches, which are prone to switch errors.
[0044] Generally, non-error-propagating approaches and error-propagating approaches are well understood in the art.Examples of non-error-propagating approaches include, but are not limited to, chromosome conformation capture (e.g., Hi-C) for a set of particularly close (e.g., adjacent) phases; single-cell-template strand sequencing; and chromosome sequencing (e.g., obtained by karyotyping or laser capture microdissection).It will be understood that a sequencing technique that can estimate that reads originate from the same chromosome homologue due to the nature of the experimental setting used for sequencing (i.e., a sequencing approach that can experimentally focus or be limited to only one chromosome homologue) is a non-error-propagating approach.Generally, approaches that are prone to error propagation (error propagation) include, but are not limited to, approaches based on parent sperm and / or polar body sequencing; dilution pool sequencing; population reference panel; and long-read sequencing (e.g., nanopore sequencing) when phase determination is not focused on a set of phases within a sufficiently localized region (e.g., within about 50 kb) so that two sets of phases can be captured in a single read.
[0045] According to some aspects of the present invention, a non-error propagation method can be used on a targeted region of DNA to provide accurate phase determination of the targeted region. The phase state information obtained from the non-error propagation method can be combined with the phase state information obtained from the error propagation method. For example, the phase state information obtained from the non-error propagation method can be used to identify and correct switch errors in an estimated phase state alignment (e.g., a phase state obtained from the error propagation method) and / or to confirm the estimated phase state alignment as a true alignment. The phase state information obtained from the non-error propagation method can be used to supply missing phase information in an estimated phase state alignment (e.g., a phase state obtained from the error propagation method).
[0046] Ploidy status The ploidy state of chromosome or chromosome segment can be broadly characterized as euploid (having normal copy number) or aneuploid (having abnormal copy number).To determine the ploidy state of genetic sample, the amount of genetic material present at one or more loci can be used.Aneuploidy can include, for example, unbalanced translocation, uniparental disomy or other global chromosomal abnormalities, including copy number variation (CNV).
[0047] Copy number variation CNVs generally refer to variation in the number of repeats in repeated genomic segments across individual chromosomes. Approximately two-thirds of the entire human genome can be composed of repeats, and 4.8–9.5% of the human genome can be classified as CNVs. CNVs are known to predict disease phenotypes, at least to some extent. CNVs can affect the number of short repeats (e.g., dinucleotide or trinucleotide repeats) or long repeats (e.g., whole-gene repeats) and are generally introduced by duplication or deletion events. CNVs are often assigned to one of two major categories based on the length of the affected sequence. The first category includes copy number variations (CNPs), which are common in the general population and occur with an overall frequency greater than 1%. CNPs are typically small (most less than 10 kb in length) and are often enriched in genes encoding proteins important in drug detoxification and immunity. A subset of these CNPs is highly variable in copy number. As a result, different human chromosomes can have a wide range of copy numbers (e.g., 2, 3, 4, 5, etc.) for a particular set of genes. CNPs associated with immune response genes have recently been linked to susceptibility to complex genetic disorders, including psoriasis, Crohn's disease, and glomerulonephritis.
[0048] The second class of CNVs comprises relatively rare variants, which are much longer than CNPs, ranging in length from hundreds of thousands of base pairs to over a million base pairs.In some cases, these CNVs may occur during the sperm or egg production that gives rise to a particular individual, or may only be inherited within a family for two or three generations.These large and rare structural variants are disproportionately observed in subjects with mental retardation, developmental delay, schizophrenia and autism.Their appearance in such subjects has led to speculation that large and rare CNVs may be more important in neurocognitive diseases than other forms of inherited mutations, including single nucleotide substitutions.
[0049] Gene copy number can be altered in cancer cells. For example, Chr1p duplication is common in breast cancer, and EGFR copy number can be higher than normal in non-small cell lung cancer. Cancer is one of the leading causes of death, and therefore early diagnosis and treatment of cancer is important because it can improve patient outcomes (such as by increasing the probability of remission and the duration of remission). Early diagnosis can also allow patients to receive fewer or less intense treatment options. Many current treatments that destroy cancerous cells also affect normal cells, resulting in various possible side effects such as nausea, vomiting, low blood counts, increased risk of infection, hair loss, and ulcers in the mucous membranes. Therefore, early detection of cancer is desirable because it can reduce the amount and / or number of treatments (such as chemotherapy or radiation) required to eliminate the cancer.
[0050] Copy number variations have also been associated with severe mental and physical disabilities and idiopathic learning disabilities. Noninvasive prenatal testing (NIPT) using extracellular DNA (cfDNA) can be used to detect abnormalities such as fetal trisomies 13, 18, and 21, triploidy, and sex chromosome aneuploidies. Subchromosomal microdeletions, which can also result in severe mental and physical disabilities, are more difficult to detect due to their smaller size. Eight microdeletion syndromes have a combined incidence of more than 1 in 1,000, making them nearly as common as fetal autosomal trisomies. Furthermore, higher copy numbers of CCL3L1 are associated with lower susceptibility to HIV infection, and low copy numbers of FCGR3B (CD16 cell surface immunoglobulin receptor) may increase susceptibility to systemic lupus erythematosus and similar inflammatory autoimmune disorders.
[0051] Determination of ploidy status Various aspects of the present invention include determining or calling the ploidy state (e.g., calling CNV) of a subject, a cell or a population of cells, or other source of genetic material, for either a chromosome or a chromosome segment. As used herein, a chromosome segment can refer to any length or portion of a chromosome sequence that can be characterized as having a copy number, including an entire chromosome. A subject can refer to any organism that has a genome, preferably a diploid genome. Preferably, the subject can be a mammal. According to various aspects, the subject is a human. Determining the ploidy state can include determining the origin of aneuploidy (i.e., determining which chromosome homologue contains aneuploidy). The origin can be identified, for example, as originating from a chromosome inherited from a mother or a father.
[0052] The ploidy state of a chromosome or chromosome segment can be determined with respect to a reference genetic code. The reference genetic code can correspond to the entire genome of a subject, one or more entire chromosomes of a subject, or one or more chromosome segments (on the same or different chromosomes) of a subject. The reference genetic code can be obtained directly or indirectly from a subject whose genetic material is being analyzed according to the methods disclosed herein. For example, the reference genetic code can be derived from sequencing normal genetic material (e.g., normal cells or non-cancerous cells) from the subject. The normal genetic material can be genetic material that is known to be euploid or in which an aneuploidy of known nature has been previously identified. The reference genetic code can be obtained from sequencing the subject's somatic cells and / or germline cells. In some examples, the reference genetic code can be obtained by reconstructing the genetic code from sequencing one or more parents or other genetic relatives of the subject whose genetic material is being analyzed, particularly if the subject is an embryo or fetus, according to methods known in the art. See, for example, International Publication No. 2021 / 067417 to Kumar et al., published April 8, 2021, which is incorporated by reference in its entirety. Constructing a reference genetic code may include sampling somatic and / or germline tissues of one or more genetic relatives. Constructing a reference genetic code may include sampling a subject (e.g., an embryo or fetus), even if only limited genetic information is available. Constructing a reference genetic code may include sequencing cells obtained from the subject. Constructing a reference genetic code may include sequencing extracellular DNA (cfDNA), such as through sampling DNA fragments in the subject's blood, cell culture medium (in the case of an embryo), or the subject's mother's blood (in the case of a fetus). In some embodiments, the subject's genome, or at least the genomes of the subject's normal cells, serves as a reference genetic code to which comparisons can be made to determine the ploidy state (e.g., of abnormal cells such as tumor cells).In some embodiments, the subject's predicted genome (i.e., a genome composed of the particular chromosomes inherited from the subject's parents, absent de novo changes in ploidy state, such as de novo amplification or deletion events) serves as a reference genetic code against which comparisons can be made to determine de novo changes to ploidy state in the subject.
[0053] The reference genetic code may not be phased. Preferably, the reference genetic code is fully phased or at least partially phased. The reference genetic code may be phased by any method known in the art, such as an error propagation phasing approach. For example, the genetic code may be phased by computational techniques involving a reference population panel. The genetic code may be phased by molecular techniques, such as diluted pool sequencing. See, for example, Choi et al., PLoS Genet. 2018 Apr 5; 14(4): e1007308 (doi:10.1371 / journal.pgen.1007308). The genetic code may be phased by sequencing the subject's germline cells and / or one or more genetic relatives (e.g., mother and father) of the subject. See, for example, International Publication No. WO 2021 / 067417 to Kumar et al., published April 8, 2021, which is incorporated by reference in its entirety.
[0054] A haplotype is a continuous phase-determined block of genomic variants specific to any chromosome homolog. According to various aspects, before carrying out the methods of the present invention described herein, haplotype blocks can be pre-constructed so that there is a certainty or at least a sufficiently high confidence of correct phase determination within the haplotype block. For example, haplotype blocks can be constructed from diluted pool sequencing or long-read sequencing, where there is a certainty or a high confidence that there are no switch errors within the haplotype block. Obtaining prior phase state information for the genetic code of interest can include obtaining one or more haplotype blocks. In various embodiments, one or more of the signals described herein can be averaged across the haplotype block or across smaller regions or sections of the haplotype block.
[0055] A non-error-propagating phase decision approach In various embodiments, it may be advantageous to combine a non-error propagation phasing approach with an error propagation phasing approach. Non-error propagation phasing techniques can provide an independent source of information for more traditional error propagation techniques. Error propagation phasing approaches (e.g., population-based phasing and molecular phasing approaches described elsewhere herein) may provide a faster, cheaper, and / or more convenient approach for obtaining large-scale sequence and / or phase state information than non-error propagation approaches. Non-error propagation approaches may provide more accurate phase state information for targeted regions of the genetic code, allowing for better determination of ploidy state (e.g., improving the ability to call CNVs within the targeted region).
[0056] The phase alignment obtained from the non-error propagation technique can be used in a targeted manner. Depending on the method used, targeted phase correction can focus on specific regions of the genetic code, saving resources and enabling more efficient implementation of one or more non-error propagation methods. For example, a specific set of phase states associated with potential switch errors identified from an at least partially phased genome can be used to correct their true set of phase states. The phase alignment can be used to reanalyze the entire phase state alignment of a genome, a chromosome of interest, or a chromosome segment of interest. The phase states can be used to provide missing phase information for specific variants or chromosome segments. The phase alignment can be computationally recalculated using the phase alignment in combination with prior phase state data (e.g., obtained from an error propagation approach). Methods for combining the phase state alignment obtained from the methods described herein with existing phase information are well understood in the art. According to certain aspects of the present invention, the non-error propagation technique can be used in combination with conventional error propagation techniques to provide an improved process for reconstructing the entire genome based on the obtained more accurate phase state information. Non-error propagation techniques may also enable interpretation of the function of variants within a genome.
[0057] As described herein, various phase determination approaches that are understood to be non-error propagating are known in the art. Specific, but non-limiting, examples of such techniques that can be used in a non-error propagating manner are described herein.
[0058] Chromosome three-dimensional structure capture (3C) Chromosome conformation capture (3C) technology is a molecular biology method used to analyze the spatial organization of chromatin within cells. 3C methods generally quantify the number of interactions between genomic loci that are close together in three-dimensional space, including loci that may be separated by many nucleotides in the linear genome sequence (e.g., loci that may be too far apart to be captured together by short-read and / or long-read sequencing). Such interactions can arise, for example, from biological functions such as promoter-enhancer interactions or from random polymer looping, where the undirected physical movement of chromatin causes loci to collide. Interaction frequencies can be analyzed directly, or they can be converted to distances, which can facilitate the reconstruction of the three-dimensional structure. Different 3C-based methods may have different scopes for genome-wide interactions that can be investigated. Deep sequencing of 3C-generated material can be used to generate genome-wide interaction maps.
[0059] In 3C methods, digestion and subsequent religation of DNA in cross-linked chromatin within the cell nucleus allows for the detection of spatial proximity between DNA sequences. Some 3C techniques may be based on high-throughput sequencing. In standard 3C-based protocols, chromatin is typically cross-linked with formaldehyde. The cross-linked chromatin is then fragmented, typically with a restriction enzyme, so that the genome is generally cut approximately every 256 bp or 4096 bp. In situ ligation then ensures preferential ligation between contacting chromatin fragments and cross-linked chromatin fragments. The chromatin is then digested to reverse the cross-links, resulting in linear and / or circular DNA concatemers carrying shuffled genome fragments linked together according to spatial proximity.
[0060] 3C techniques include classical 3C, 4C, 5C, Hi-C, and ChIA-PET methods. Classical 3C, often referred to as the "one-to-one" approach, uses PCR to amplify and quantify specifically targeted ligation junctions. 4C, often referred to as the "one-to-all" approach, is similar to classical 3C, except that a second round of digestion and ligation is performed to generate small DNA circles. Primers designed against specific anchor sequences can then be used in inverse PCR to amplify all contact sequences that formed ligation products with the anchor sequence, although more recent methods can avoid the need for amplification. The contact sequences can then be sequenced by any appropriate means. 5C, often referred to as the "many-to-many" approach, hybridizes and then ligates primers complementary to the fragments of interest to the 3C ligation product to create copies of the junctions of interest to the extent they exist. Universal PCR primers complementary to the tails of the original primers are then used to amplify the ligation products of interest, which can then be sequenced by any appropriate means. Hi-C, often referred to as the "all-to-all" approach, uses restriction enzymes that leave overhangs filled with biotin-labeled nucleotides. After blunt-end ligation, the ligation products are sheared to reduce fragment size, and streptavidin is used to remove the biotin-containing fragments to create an enriched library, which is then sequenced, typically by NGS techniques. Hi-C provides a matrix of pairwise interaction frequencies between fragments across the genome. Resolution can be improved by using a higher restriction site density and / or by increasing sequencing depth, resulting in a higher resolution. 2Pairwise sequencing generally results in a factor of x improvement in resolution. Hi-C, in particular, can improve chromosome-wide phase determination by combining Hi-C and chromatin immunoprecipitation (ChIP). A specific antibody is used to remove ligation junctions bound by the chromatin protein of interest before biotinylating and ligating the fragment ends. Other chromosome conformation capture techniques known in the art include tethered conformation capture (TCC), DNase Hi-C or Micro-C, targeted chromatin capture (T2C), capture Hi-C (Chi-C), HiCap, and Capture-C. Various methods for performing chromosome conformation capture are described, for example, in Denker, et al., Genes Dev. 2016 Jun 15;30(12):1357-82 (doi:10.1101 / gad.281964.116); de Wit, et al., Genes Dev. 2012 Jan 1;26(1):11-24 (doi:10.1101 / gad.179804.111); McCord et al., Mol Cell. 2020 February 20;77(4):688-708) (doi:10.1016 / j.molcel.2019.12.021); or Belton et al., Methods. 2012 Nov;58(3):268-76 (doi:10.1016 / j.ymeth.2012.05.001), each of which is incorporated herein by reference in its entirety.
[0061] Chromosome conformation capture techniques can be used to determine genome phase in a non-error-propagating manner. Due to their inherent spatial proximity, the probability of loci on the same chromosome homolog being linked together is much higher than that of loci on two homologous chromosomes being linked together. Therefore, the overall distribution of ligation fragments generated by 3C techniques can be assumed to favor variants from the same chromosome homolog compared to variants from two or more different homologs. Furthermore, the effect is more dominant the closer the variants or sets of phases are to each other. Therefore, chromosome conformation capture techniques such as Hi-C can be used to align two phases, especially sets of two adjacent phases, without the risk of introducing switching errors.
[0062] The distribution of fragments (ligation products) obtained from the chromosome conformation capture method can be analyzed to determine whether the distribution supports the two phase sets being the same or different phases. The fragments can be filtered to select fragments containing at least one variant from each phase set. The fragments can be grouped into subgroups corresponding to different sets of variants that support the same haplotype call, although each fragment need not contain the same variant. In some embodiments, the fragments can be filtered to only fragments that contain each variant from one or both phase sets. A putative phase or haplotype can be assigned to each phase set so that a putative phase alignment exists. If no prior phase determination has been made, the phase alignment can be assigned randomly. The selected fragments and / or subgroups can be characterized as concordant or discordant with respect to the putative phase alignment. For example, if all of the variants detected within a fragment are derived from the same putative haplotype, the fragment can be considered concordant with the putative phase alignment; otherwise, the fragment can be considered discordant. Given the significantly higher probability of fragments containing variants from the same haplotype or chromosomal homolog, especially for closely spaced variants, the distribution of fragments / subgroups can be expected to be heavily biased toward a predominance of matched or mismatched fragments. A predominance of matched fragments / subgroups suggests that the putative phase alignment is correct, whereas a predominance of mismatched fragments suggests that the putative phase alignment is incorrect. The amount of bias can be quantified by calculating the probability of observing the bias by chance. For example, a binomial probability can be calculated for the probability of observing the measured distribution by chance, with each measurement having a certain probability of being matched or mismatched. The certain probability can be set as a lower bound, with 50% suggesting that the ligation of the phase set is completely random.Alternatively, to account for the higher probability expected from spatial proximity, the certain probability that sets of phases from the same haplotype exist within the same fragment can be set higher (e.g., 60%, 70%, 75%, 80%, 90%, 95%, 99%, 99.9%, etc.). A higher certain probability may be more useful for a smaller number of measurements, while a lower certain probability may be sufficient for a larger number of measurements. If there is high confidence that the observed distribution is not simply the result of chance (e.g., the measurements are statistically significant within a 95% confidence interval), the sets of phases can be accurately aligned based on the chromosome conformation data.
[0063] Single-cell template strand sequencing Single-cell template strand sequencing (Strand-seq) is a single-cell sequencing technique that isolates individual homologs within cells by restricting sequence analysis to the DNA template strand used during DNA replication. This method relies on DNA directionality (distinguished by the 5'-3' orientation of the DNA) by culturing cells in a thymidine analog during a single cell division, allowing nascent DNA strands to be labeled and subsequently selectively removed from analysis. Each single-cell library is multiplexed for storage and sequencing, and the resulting sequence data are aligned and mapped to either the minus or plus strand of a reference genome to assign the template strand status of each chromosome within the cell. See, for example, Porubsky et al., Genome Res. 2016 Nov;26(11):1565-1574 (doi:10.1101 / gr.209841.116); Sanders et al., Nat Protoc. 2017 Jun;12(6):1151-1176 (doi:10.1038 / nprot.2017.029), each of which is incorporated by reference in its entirety. Because sequencing can be limited to a single strand, this technique can be used as a non-error-propagating method as described herein.
[0064] Chromosome isolation Since all sequence readings can be assumed to be from the same homolog, any technique that physically separates one chromosome homolog from another chromosome homolog before sequencing can be considered as a non-error propagation approach to phase determination.For example, the sequencing of chromosomes obtained by karyotype or laser capture microdissection can be used for the non-error propagation technique described herein.For example, see Kang et al., Cytogenet Genome Res.2017;152(4):204-212 (doi:10.1159 / 000481790), the entire contents of which are incorporated herein by reference.
[0065] Sequencing methods Various methods of DNA sequencing are well known in the art and may be used to perform the methods described herein unless otherwise indicated by the context. DNA sequencing may include, for example, Sanger sequencing (chain termination sequencing). DNA sequencing may include the use of next-generation sequencing (NGS) or second-generation sequencing technologies, which are typically characterized by being highly scalable and allowing the entire genome to be sequenced at once. NGS technologies generally allow multiple fragments to be sequenced at once, enabling "massively parallel" sequencing in automated processes. DNA sequencing may include third-generation sequencing technologies (e.g., nanopore sequencing or SMRT sequencing), which generally allow for longer reads to be obtained than can be obtained via second-generation sequencing technologies. Sequencing may include paired-end sequencing, where feasible, in which both ends of a DNA fragment are sequenced, which may improve the ability to align reads to longer sequences. DNA sequencing may include sequencing by synthesis / ligation (e.g., ILLUMINA® sequencing), single molecule real-time (SMRT) sequencing (e.g., PACBIO® sequencing), nanopore sequencing (e.g., OXFORD NANOPORE® sequencing), ion semiconductor sequencing (Ion Torrent sequencing), combinatorial probe anchor synthesis sequencing, pyrosequencing, and the like.
[0066] Shotgun sequencing refers to a method of sequencing random DNA strands from genomes or large genetic samples. DNA is randomly divided into many small segments, which are sequenced (for example, using chain termination) to obtain reads. By performing several rounds of this fragmentation and sequencing, multiple overlapping reads are obtained for the target DNA. A computational algorithm then uses the overlapping ends of different reads to assemble the random segment reads into a continuous sequence. Shotgun sequencing can be used for whole genome sequencing. As described elsewhere herein, any suitable form of sequencing, including those described herein, can be used to identify variants (e.g., SNPs) in a subject, which can then be used as a basis for measuring the genetic signal indicating the ploidy state for the chromosome segment containing the variant. According to one aspect of the present invention, hierarchical sequencing can be used for whole genome sequencing.
[0067] Data collection Genetic material for analysis by the methods described herein can be obtained from a variety of sources, including somatic cells (e.g., white blood cells, cells from tissue biopsies), germ cells (e.g., sperm, eggs, polar bodies), and extracellular DNA. Genetic material can be collected directly from the subject whose genome is being analyzed and / or from the subject's genetic relatives (e.g., mother and / or father). According to various embodiments, genetic signals indicative of ploidy status, such as allele balance signals or read depth signals, can be obtained from extracellular DNA (cfDNA) derived directly from the subject. Extracellular DNA is DNA found outside of cells, for example, freely circulating in the bloodstream or in the cell culture medium of cultured cells, such as embryos grown for in vitro fertilization (IVF).
[0068] Various embodiments of the methods described herein may include obtaining and / or sequencing extracellular DNA. The extracellular DNA may include extracellular fetal DNA (cffDNA). The extracellular DNA may include circulating tumor DNA (ctDNA). Extracellular DNA may provide a relatively abundant source of genetic material that can be obtained from non-invasive or minimally invasive procedures, such as sampling cell culture medium or drawing blood from a subject. The extracellular DNA may provide sufficient genetic information for whole-genome sequencing of the subject from which the extracellular DNA is derived. See, for example, Kitzman et al., Sci Transl Med. 2012 Jun 6;4(137):137ra76 (doi:10.1126 / scitranslmed.3004323). For example, shotgun sequencing of extracellular DNA may be used to sequence one or more chromosomes of a subject. Genetic material from a subject may contain cells with consistent genetic profiles or cells with different genetic profiles (e.g., normal cells and tumor cells). In some examples, a subject's genome can be reconstructed based on sequencing genetic material obtained directly from the subject and sequencing one or more genetic relatives. See, e.g., International Publication No. WO 2021 / 067417 to Kumar et al., published April 8, 2021, which is incorporated by reference in its entirety.
[0069] Extracellular fetal DNA (cffDNA) is fetal DNA that circulates freely in maternal blood. Therefore, cffDNA can be obtained from maternal blood collected, for example, by venipuncture. Analysis of cffDNA is a noninvasive prenatal diagnostic method that can be indicated for pregnant women. cffDNA originates from placental trophoblast cells. When placental microparticles are released into the maternal circulation, fetal DNA is fragmented. cffDNA fragments, approximately 200 bp in length, are significantly smaller than maternal DNA fragments and can be distinguished from them. Approximately 11–13.4% of extracellular DNA in maternal blood is cffDNA, but the amount varies greatly between pregnant women. cffDNA generally becomes detectable after 5–7 weeks of gestation, and its amount increases as pregnancy progresses. The amount of cffDNA in maternal blood rapidly decreases after birth and is generally no longer detectable approximately 2 hours after birth. Analysis of cffDNA may provide an earlier diagnosis of fetal conditions than other techniques. The cffDNA can be analyzed, for example, by massively parallel shotgun sequencing (MPSS), targeted massively parallel sequencing (t-MPS) and SNP assays.
[0070] ctDNA is tumor-derived fragmented DNA in the bloodstream that is not associated with cells. Because ctDNA can reflect the entire tumor genome, its potential clinical utility is gaining momentum. Liquid biopsies, in the form of blood draws, can be taken at various time points throughout a treatment regimen to monitor tumor progression. ctDNA originates directly from the tumor or from circulating tumor cells (CTCs), which are live, intact tumor cells that shed from the primary tumor and enter the bloodstream or lymphatic system. The exact mechanism of ctDNA release remains unclear. Biological processes hypothesized to be involved in ctDNA release include apoptosis and necrosis from dead cells or active release from live tumor cells. Studies in both humans (healthy and cancer patients) and xenograft mice have shown that the size of fragmented cfDNA is primarily 166 bp long, which corresponds to the length of DNA wrapped around a nucleosome plus linker. Fragmentation of this length may indicate apoptotic DNA fragmentation, suggesting that apoptosis may be the primary method of ctDNA release. cfDNA fragmentation is altered in the plasma of cancer patients. In healthy tissues, infiltrating phagocytes are responsible for the clearance of apoptotic or necrotic cellular debris, including cfDNA. While cfDNA is present at low levels in healthy patients, higher levels of ctDNA can be detected in cancer patients as tumor size increases. This likely occurs due to inefficient immune cell infiltration into the tumor site, which reduces the effective clearance of ctDNA from the bloodstream. Comparison of mutations in ctDNA and DNA extracted from the same patient's primary tumor revealed the presence of identical cancer-associated genetic alterations, opening the possibility of analyzing ctDNA to analyze the genetic makeup of tumor cells. Therefore, ctDNA may be used for earlier cancer detection and treatment follow-up monitoring.
[0071] According to various aspects of the present invention, the non-error-propagation phasing techniques described elsewhere herein are performed on cellular DNA (not extracellular DNA) so that intact chromosomes are isolated or effectively isolated to provide accurate phasing (e.g., correct any switch errors). In some embodiments, single-cell sequencing can be performed on one or more cells to obtain the data described herein. The genetic data obtained using non-error-propagation phasing techniques may or may not be sufficient to independently construct a subject's genome or to independently provide a sufficient reference genome. Genetic data obtained from conventional sequencing techniques (e.g., whole-genome shotgun sequencing on extracellular DNA, etc.) combined with error-propagation phasing approaches can be advantageous in providing depth and / or scope of genetic information. Genetic data obtained from non-error-propagation phasing approaches (which may be performed on cellular DNA) can be advantageous in providing more accurate phasing of various phase sets, particularly sets of adjacent or neighboring phases. Thus, using these independent sources of information together can be advantageous.
[0072] According to some aspects of the present invention, cellular DNA sequencing can be performed on blood cells (e.g., white blood cells) or other cells (e.g., cells found in saliva) collected through non-invasive or minimally invasive techniques.Therefore, sequencing of extracellular DNA and cellular DNA can be performed exclusively by non-invasive or minimally invasive procedures such as blood sampling.Extracellular DNA and cellular DNA can be isolated from the same or different samples (e.g., body fluid samples such as blood samples or saliva samples).For example, extracellular DNA can include ctDNA, and cellular DNA can include white blood cell DNA (which should provide normal genetic material except in the case of leukemia).
[0073] According to some aspects of the present invention, sequencing cellular DNA may involve isolating one or more cells from a fetus or embryo according to methods well understood in the art. Such approaches typically require invasive techniques that may pose risks to the embryo or fetus. According to preferred aspects of the present invention, the cellular DNA used for non-error-propagation phasing approaches may be obtained using non-invasive or minimally invasive techniques, such as blood or sperm collection. While non-invasive or minimally invasive techniques for sequencing cellular DNA may not be possible for the subject's own cells in the case of an embryo or fetus, sequencing of cellular DNA may be performed on genetic relatives of the fetus (e.g., the mother and / or father). Because non-error-propagation phasing may only be used to provide an accurate phase state of a set of phases, and not necessarily to independently construct a reference genetic code and / or generate a signal indicating ploidy state, the true phase state of the subject's genome may be inferred from the true phase state of the genomes of one or more genetic relatives who inherited at least some of the same haplotypes as the subject. Thus, the methods described herein may be performed on genetic material obtained by entirely non-invasive or minimally invasive methods, including when the subject is an embryo or fetus.
[0074] Genetic signals indicating ploidy status As used herein, "signal" may refer to one or more measurements that can provide information about the genetic composition of the genetic sample being investigated. The measurements may be raw measurements or processed measurements, such as those derived from mathematical analysis of one or more raw measurements. The signals may be obtained from sequencing data. The signals may be, for example, allele balance signals or read depth signals, as described elsewhere herein. The signals may correspond to values along a continuous or discrete number spectrum. The signals may indicate genetic information at one specific locus. The signals may be averaged from signals measured across multiple loci.
[0075] A locus is a specific, fixed location on a chromosome. A locus identifies the chromosomal location of a specific gene or genetic marker. As used herein, a locus of interest may refer to a locus within the genetic material being analyzed to which one or more measurements can be mapped to derive a signal indicative of the genetic composition of the genetic material. A variant of interest may refer to a locus of interest where there is a difference in the genetic composition at the locus of interest between two or more chromosomal homologs within the genetic material. A SNP may be a variant of interest. As used herein, a "phase set" may refer to a set of one or more adjacent variants of interest whose phase alignment with another phase set can be determined according to the methods described herein. In some examples, a phase set may correspond to a haplotype block or a chromosomal region larger than a haplotype block (e.g., two or more adjacent haplotype blocks). For example, a phase set may include 2, 5, 10, 50, 100, 500, 1,000, 5,000, or more variants. In some examples, a phase set may consist of a single variant. Two aligned phase sets may or may not have the same number of variants of interest. Determining the phase alignment of one phase set with another phase set can include determining that the two phase sets are in phase (i.e., the variants of interest in each phase set belong to the same chromosome homolog) or that the two phase sets are out of phase (i.e., the variants of interest in the first phase set do not belong to the same chromosome homolog as the variants of interest in the second phase set).
[0076] According to some specific aspects, the phase sets can be adjacent phase sets. For example, the first phase set can have a variant of interest that is not more than about 1,000, about 5,000, about 10,000, about 50,000, about 100,000, about 5 million, about 1 million, about 5 million, about 10 million, about 50 million, about 100 million, or about 250 million base pairs away from the variant of interest in the adjacent phase set. The adjacent phase set can be defined to include the variant of interest on either side of a potential switch error. A potential switch error can be identified as a possible switch error between two haplotype blocks. According to some specific aspects, a site where one or more signals suggest a shift between chromosome segments from a euploid segment to an aneuploid segment, or vice versa, can be identified as a potential switch error. According to some specific aspects, a site where one or more signals suggest a change in copy number relative to an adjacent segment can be identified as a potential shift error. According to some particular aspects, sites where one or more signals suggest a shift between chromosomal segments of different aneuploid states (e.g., from trisomy to monosomy, or vice versa) can be identified as potential switch errors.
[0077] Allelic balance (synonymous with allelic balance, allele frequency, or allele frequency) refers to the proportion of reads from a set of sequencing data that cover a variant's location that support that variant. For example, if 100 reads map to a particular variant's locus, 25 of which support that variant and 75 of which do not, the variant would have an allelic balance of 0.25. Heterozygous loci can be filtered for a minimum read depth for inclusion in allelic balance data. The relative proportion of one variant relative to another can indicate differences in locus copy number between different chromosomal homologs in a genetic sample. Comparing the expected copy number based on a reference genetic code with the detected number can indicate, for example, whether an amplification or deletion event occurred for one of the chromosomal homologs (e.g., in all or at least some of the cells from which the genetic sample was derived). Allelic balance signals measured across multiple variants can provide a signal for haplotype or chromosomal balance based on the assignment of alleles to haplotypes or chromosomal homologs. Because allele balance thereby becomes dependent on the phase state of a variant (i.e., whether a relatively high or low proportion of alleles supports a high or low proportion of chromosomal homologs depends on its phase state), the allele balance signal can be altered by phasing errors, such as switch errors. Thus, phase correction can translate directly into allele balance correction, such that a true allele balance signal is obtained from correcting the phase alignment. As used herein, "correcting" a phase alignment or allele balance signal can refer to comparing a phase determination to a prior or other estimated phase determination, or to supplying missing phase information, unless the context dictates otherwise (e.g., "correcting an error"), regardless of whether an incorrect phase has actually been identified and changed.
[0078] Read depth refers to the number of sequencing reads that map to a given locus during one or more sequencing runs. Read depth signal (or depth signal) can be normalized across the total number of reads. Read depth can be expressed in a variety of different ways, including, but not limited to, the absolute number of reads mapped to a particular locus by a sequencing device, or the percentage or proportion of reads mapped to that locus. Thus, for example, in a highly parallel DNA sequencing device, such as ILLUMINA HISEQ®, which generates sequences for 1 million clones, 3,000 sequencing runs of one locus will result in a read depth of 3,000 reads at that locus. The proportion of reads at that locus is 3,000 divided by 1 million total reads, or 0.3% of the total reads. Generally, the greater the read depth at a locus, the more likely the allele balance signal at that locus will be closer to the true allele balance in the original genetic sample. Loci can be filtered for minimum read depth to be included in read depth data.The read depth of specific variant can indicate the relative number of copies of this variant compared with other variants, especially when normalized to the total number of reads.Comparing the relative number of copies of variant with one or more benchmarks, for example, for the known number of copies from reference genetic code, can indicate, for example, whether amplification or deletion event occurs for one of chromosome homologues (for example, in all or at least part of the cells from which genetic sample is derived).
[0079] For example, in addition to any copy number abnormality, noise can be introduced into the signal by many mechanisms, including stochastic events due to sampling, GC bias, and / or uneven distribution of variants across the genome.The signals described herein can generally be averaged across multiple adjacent loci.For example, multiple adjacent loci can include 2, 3, 4, 5, 10, 15, 20, 25, 30, 40, 50, 100, 500, 1,000, 5,000 or more loci.The selection of loci can depend on their density with the region of interest.For example, multiple adjacent loci can include all loci within a region of at least about 50,000, at least about 100,000, at least about 200,000, at least about 300,000, at least about 400,000, at least about 500,000, at least about 750,000, at least about 1 million, at least about 50 million, or at least about 100 million base pairs. The multiple adjacent loci can include all loci within a range of about 50,000 or less, about 100,000 or less, about 200,000 or less, about 300,000 or less, about 400,000 or less, about 500,000 or less, about 750,000 or less, about 1 million or less, about 50 million or less, or about 100 million or less base pairs.The range of adjacent loci can be selected so that the loci are assumed to be located on the same chromosome.Therefore, if there is no aneuploidy for only some of the loci within the selection, the true signal for each locus' allele balance or read depth should be the same.Therefore, averaging across adjacent loci can reduce the noise in the signal described herein.
[0080] Combining allelic balance and read depth According to various aspects of the present invention, allele balance signals and read depth signals can be used in combination to determine ploidy state. Allele balance and read depth can each individually indicate ploidy state determination, as described elsewhere herein. However, because the noise from these signals, i.e., noise in allele balance related to variations in the number of specific DNA molecules sequenced that overlap the interrogated site and noise in read depth related to variations in the total number of DNA molecules sequenced that overlap the interrogated site, are at least somewhat independent, these signals can provide sources of information independent of each other, improving the signal-to-noise ratio and enabling more accurate ploidy state determination. This combination can be particularly useful in scenarios where there are an intermediate number of reads (i.e., enough reads to determine the allele balance at a locus with sufficient fineness, but not so many reads that a read depth signal is apparent). The allele balance signal can be corrected via a non-error-propagating phase determination approach to provide a true allele balance signal, according to the methods described elsewhere herein.
[0081] Signals can be combined and used according to various aspects, as understood in the art. For example, signals can be combined and used together by multivariate logistic regression, log-linear modeling, neural network analysis, n-of-m analysis (aneuploidy is indicated when at least "n" criteria out of a total of "m" criteria are met), decision tree analysis, random forest analysis, rule set, Bayesian method, neural network method, multiplication, addition, etc. Some methods of using signals together can include integrating two signals into a single composite signal through mathematical operations. For example, signals can be multiplied or added together. In various embodiments, one or both signals can be multiplied by a scalar. For example, signals can be normalized to one or more measures of noise, such as the standard deviation or variance measured in the signal (e.g., across multiple chromosomal locations where the signal is measured and / or across multiple runs of analysis).
[0082] For each signal and / or signal combination, one or more threshold levels or values of the signal can be selected as cutoffs to distinguish different copy numbers of loci or chromosome segments. For example, a threshold can be selected to distinguish between loci present in trisomy (three copies of the locus) and loci present in disomy (two copies of the locus), and / or a threshold can be selected to distinguish between loci present in monosomy (one copy of the locus) and loci present in disomy. Signals can be offset or otherwise normalized to signals (e.g., average signal values) for different copy numbers, such as euploid copy numbers. For example, signals can be configured so that a level of 0 indicates a euploid ploidy state, and a sufficient deviation therefrom indicates an aneuploid ploidy state. Different thresholds can be selected to indicate different copy numbers.
[0083] The use of individual signals and / or combined signals can be characterized by the probability that the signal can correctly distinguish two populations with different copy numbers, such as euploid populations and aneuploid populations.Probability can be characterized, for example, as the probability that using a signal threshold correctly identifies which population a variant should be assigned to.Probability can be characterized by the probability of true positive, false positive, true negative and / or false negative.The probability based on individual signals is the individual probability.The probability based on the combined use of two signals is the joint probability.For example, the probability of a true positive aneuploid call is the probability that an aneuploid is correctly identified as an aneuploid based on the criteria for a positive call using two signals combined.As demonstrated elsewhere herein, the combined use of allele balance signals and read depth signals can generally provide a higher joint probability of true positive and / or true negative compared to individual probabilities, and / or a lower joint probability of false positive and / or false negative compared to individual probabilities.
[0084] The ability of a threshold to adequately distinguish two populations (e.g., euploidy versus aneuploidy) can be established using receiver operating characteristic (ROC) analysis, as is known in the art. The area under the ROC curve can provide a measure of the quality of the signal used to distinguish two populations, regardless of the specific threshold. To plot the ROC curve, the true positive rate (TPR) and false positive rate (FPR) are determined as the discrimination threshold is continuously varied. A perfect test to distinguish two populations has an area under the ROC curve of 1.0, while a random test has an area of 0.5. Preferably, the signal provides an ROC curve area greater than 0.5, preferably at least 0.6, more preferably 0.7, even more preferably 0.75, even more preferably at least 0.8, even more preferably at least 0.9, and most preferably at least 0.95.
[0085] A certain threshold can be selected to provide acceptable levels of sensitivity (true positive rate) and specificity (true negative rate).For example, the threshold can be selected so that the false positive rate is approximately equal to the false negative rate.Such a threshold can be assumed to be, for example, half the average signal level for aneuploidy (or a specific aneuploidy state) when offset against the average signal level for euploidy (or non-aneuploidy state).According to some aspects, the threshold can be selected to provide a specificity of greater than 0.5, preferably at least 0.6, more preferably at least 0.7, even more preferably at least 0.8, even more preferably at least 0.9, and most preferably at least 0.95.According to some aspects, the threshold can be selected to provide a sensitivity of greater than 0.5, preferably at least 0.6, more preferably at least 0.7, even more preferably at least 0.8, even more preferably at least 0.9, and most preferably at least 0.95. According to one aspect, the threshold value may be selected to provide an odds ratio different from 1, preferably at least about 2 or more or about 0.5 or less, more preferably at least about 3 or more or about 0.33 or less, even more preferably at least about 4 or more or about 0.25 or less, even more preferably at least about 5 or more or about 0.2 or less, and most preferably at least about 10 or more or about 0.1 or less.
[0086] A specific threshold value can be independently selected from the measurement value of one of the two populations that the threshold value distinguishes.For example, the threshold value for distinguishing aneuploid variants from euploid variants can be set as a specific percentile of the euploid population, such as the 60th percentile, 70th percentile, 80th percentile, 90th percentile, 95th percentile, 99th percentile, etc. (assuming that aneuploid signal should be greater than euploid signal), and this can be established based on the acceptable level of false positives.Alternatively, the threshold value can be set as a specific percentile of the aneuploid population, such as the 1st percentile, 5th percentile, 10th percentile, 20th percentile, 30th percentile, 40th percentile, etc. (assuming that aneuploid signal should be greater than euploid signal), and this can be established based on the acceptable level of false negatives.In some examples, if there is more data available to characterize the euploid population, the euploid signal can be used to establish the threshold value.
[0087] The population described herein can be any group of measurements.Preferably, the population can be a group of measurements obtained from the same sequencing experiment on the same genetic material.By defining the population in this way, noise within the population can be minimized.Such a population can include measurements across different loci that share the same ploidy state.However, the population can also be defined to refer to or include measurements from different sequencing experiments on the same sample of genetic material, different sequencing experiments on different samples of the same genetic material, and / or different sequencing experiments on different genetic material (e.g., different genomes).
[0088] In various embodiments, baseline signal can be established from the same sequencing data that potential aneuploids are to be identified.For example, baseline signal (for example, average signal value) can be established based on the signal measurement value of one or more chromosome segments that are known or confirmed to be euploid.The signals of other segments of chromosomes that are examined to identify potential aneuploids can be offset by this baseline signal, as described elsewhere herein.This can facilitate the comparison of different signal types.
[0089] According to some aspects, populations can be assumed to have normal distribution.Therefore, the characteristics of populations can be calculated from the average signal value for populations and optionally the noise or variance / standard deviation measure within populations.Two populations (for example, euploid populations and aneuploid populations) can be estimated to have roughly the same variance / standard deviation, which can simplify the theoretical characterization of populations, as described elsewhere herein.In particular, when two populations are determined from the same sequencing experiment (for example, for different sections of chromosomes), the noise within each signal can be assumed to be substantially the same.
[0090] According to some embodiments, the allele balance signal and the read depth signal can be obtained from the same sequencing experiment. In other words, reads from a single experiment can be mapped to variants in the reference genetic code, and the relative number of reads mapped to different alleles for the same variant can be used to obtain the allele balance signal, while the total number of reads mapped to a specific variant (optionally normalized to the total number of reads from the experiment) can be used to obtain the read depth signal. In various applications, both signals are obtained from sequencing extracellular DNA, as described elsewhere herein. According to other embodiments, the allele balance signal and the read depth signal can be obtained from different sequencing experiments. Different sequencing experiments can be performed on the same sample of genetic material or different samples of genetic material. When different samples are used, the genetic material can be obtained from the same source (e.g., extracellular DNA) or from different sources (e.g., extracellular DNA vs. cellular DNA or different cell types). In situations where allele balance signals and / or read depth signals are obtained from cellular DNA, the source of genetic material (particular sample and / or cell type) may be the same as or different from that used for any non-error-propagation phase determination, as described elsewhere herein.
[0091] Purpose A variety of potential applications are possible for making ploidy state determinations on samples of genetic material (e.g., on genomes). Some specific, but non-limiting examples of how such determinations can be used to drive subsequent determinations and / or further analysis or treatment are described herein.
[0092] Genetic profiling of tumors with chromosomal instability Genomic instability in tumor cells is often associated with poor patient outcomes and resistance to targeted cancer therapies. The accumulation of genetic and epigenetic lesions in response to environmental exposure to carcinogens and / or random cellular events often leads to the inactivation of tumor suppressor genes, which play critical roles in maintaining the cell cycle, DNA replication, and DNA repair. Loss or inhibition of cellular DNA repair mechanisms often leads to increased mutational burden and genomic instability. CNVs are widespread across many cancer types and can cause the gain of oncogenes and / or the loss of tumor suppressors, which are associated with disease progression and therapeutic response or resistance. Genomic instability is associated with subclonal heterogeneity and is frequently observed in solid tumors across different lesions, within the same tumor, and even within the same solid biopsy site. This tumor cell heterogeneity can complicate therapeutic interventions designed around a single molecular target. Although genome-wide CNV profiles can be used to characterize genome instability, the assessment of genome instability in bulk tumors or biopsies can be complicated due to sample availability and noise resulting from contamination of surrounding tissues or tumor heterogeneity.Tumors with increased genome instability have been shown to respond to certain types of treatment, including, for example, platinum-based chemotherapy and PARP inhibitors.See, for example, Greene et al., PLoS One.2016 Nov 16;11(11):e0165089 (doi:10.1371 / journal.pone.0165089), the entire contents of which are incorporated herein by reference.
[0093] Poly(ADP-ribose) polymerase (PARP), a nuclear enzyme found in nearly all eukaryotic cells, catalyzes the transfer of ADP-ribose units from nicotinamide adenine dinucleotide (NAD+) to nuclear acceptor proteins, resulting in the formation of protein-bound linear and branched homo-ADP-ribose polymers. PARP activation and the resulting formation of poly(ADP-ribose) can be induced by DNA strand breaks following exposure to chemotherapy, ionizing radiation, oxygen free radicals, or nitric oxide (NO). Some forms of cancer are more dependent on PARP than normal cells, making PARP an attractive target for cancer therapy, regardless of the specific cancer indication. Furthermore, PARP is associated with the repair of DNA strand breaks in response to DNA damage caused by radiation or chemotherapy, potentially contributing to the frequent resistance to various types of cancer treatments. Consequently, PARP inhibition may slow intracellular DNA repair and enhance the antitumor effects of cancer treatments. Indeed, in vitro and in vivo data indicate that many PARP inhibitors enhance the effects of cytotoxic drugs such as ionizing radiation or DNA methylating agents. The PARP family of enzymes is widespread, and competitive inhibitors of PARP are known. Approved PARP inhibitors include olaparib (Lynparza®, AstraZeneca); rucaparib (Rubraca®, Clovis Oncology); niraparib (Zejula®, Tesaro); and talazoparib (Talzenna®, Pfizer). Other PARP inhibitors under investigation include veliparib (ABT-888, AbbVie), pamiparib (BGB-290) (BeiGene, Inc.); CEP 9722 (Cephalon); E7016 (Eisai); and 3-aminobenzamide.
[0094] Platinum-based chemotherapy agents (antineoplastic drugs informally referred to as "platins") are coordination complexes of platinum, including cisplatin, oxaliplatin, and carboplatin, as well as several proposed drugs in development. Platinum-based chemotherapy agents cause crosslinking of DNA as single adducts, interstrand crosslinks, intrastrand crosslinks, or DNA-protein crosslinks, which inhibit DNA repair and / or DNA synthesis.
[0095] Other forms of treatment suitable for cancers that show chromosomal instability are understood in the art.Therefore, the method described herein may relate to identifying gene signatures in subjects with cancer that show chromosomal instability and are therefore suitable for a class of therapeutic agents that target genetic mechanisms (for example, inhibit DNA repair so that damaged DNA can be more effectively targeted).These therapeutic agents may be agnostic for certain types of cancer.Therefore, the method described herein can be performed on subjects who have been diagnosed with or suspected of having cancer before or at the same time as specific cancer diagnosis and / or tissue biopsy.Advantageously, the method described herein can be performed based on genetic material collected exclusively from non-invasive or minimally invasive procedures such as blood sampling.The genetic analysis described herein can be performed simultaneously with other routine analysis and / or cancer diagnosis or evaluation based on the same or different biological samples collected at the same time.
[0096] According to certain aspects of the present invention, allele balance signals and / or read depth signals (e.g., used in combination) can be obtained from a sample of genetic material collected from a subject. The signals can be obtained from extracellular DNA containing or suspected of containing ctDNA. The signals can be obtained from cellular DNA, such as tumor tissue. When an allele balance signal is used, the true signal can be determined by correcting the allele balance signal using a non-error-propagation phase determination technique, as described elsewhere herein. The non-error-propagation phase determination technique can be performed on cellular DNA. The cellular DNA can be obtained from blood cells (e.g., white blood cells). According to some aspects in which one or more signals indicating ploidy state are obtained from cellular DNA and non-error-propagation phase determination is performed on cellular DNA, the same source of cellular DNA can be used for both. In some embodiments, the extracellular DNA for obtaining the genetic signal of ploidy state and the cellular DNA for performing non-error-propagation phase determination are obtained from the same biological sample (e.g., a blood sample). To assess the ploidy state of the DNA (e.g., extracellular DNA) being evaluated, a determination of the ploidy state can be made from one or more signals. The determination can be made with respect to a reference genetic code (e.g., a normal cellular genetic code), as described elsewhere herein. The ploidy state can be determined for one or more chromosomal segments. Detection of one or more chromosomal segments exhibiting CNV can be used to identify one or more regions of the genome exhibiting chromosomal instability. Identification of such regions can be used to indicate the presence of tumors susceptible to treatment with therapeutic agents that exploit chromosomal instability, such as treatment with PARP inhibitors and / or platinum-based chemotherapeutic agents. According to some aspects, the determination of the ploidy state is used to treat a subject (e.g., by administering a treatment in vivo). According to some aspects of the present invention, the determination of the ploidy state is used to treat one or more cells in vitro. The one or more cells can include cancer cells. The cells can be cultured (e.g., grown from a tumor biopsy) from a subject with or suspected of having cancer.The cells may include cells from a cancer cell line (e.g., artificially induced to replicate cancer). The cells may include a mixture of normal and cancerous cells.
[0097] De novo or inherited CNV detection The methods described herein can be used to detect variations in ploidy state (e.g., CNV) in a subject. According to some aspects of the present invention, allele balance signals and / or read depth signals (e.g., used in combination) can be obtained from a sample of genetic material collected from a subject. One or more signals can be obtained from extracellular DNA. One or more signals can be obtained from cellular DNA. When an allele balance signal is used, the true signal can be determined by correcting the allele balance signal using a non-error propagation phase determination technique, as described elsewhere herein. The non-error propagation phase determination technique can be performed on cellular DNA. According to some aspects in which one or more signals indicating ploidy state are obtained from cellular DNA and non-error propagation phase determination is performed on cellular DNA, the same source of cellular DNA can be used for both. Cellular DNA can be obtained from blood cells (e.g., white blood cells) or other cells collected by non-invasive or minimally invasive techniques. In some embodiments, the extracellular DNA for obtaining genetic signals of ploidy status and the cellular DNA for performing non-error propagation phase determination are obtained from the same biological sample (e.g., blood sampling). To evaluate the ploidy status of the DNA being evaluated, ploidy status determination can be performed from one or more signals. Allele balance and / or read depth (e.g., used in combination) can be used to identify copy number differences between variants at the same locus, indicating aneuploidy in one of the chromosome homologs.
[0098] The method described herein can be used to detect inherited variation in ploidy state (i.e., variation in ploidy state at one or more loci of one of the chromosomes of interest, where the ploidy state of each chromosome homolog is inherited from the parent) or de novo variation in ploidy state (i.e., a change in the ploidy state of one of the chromosomes of interest, relative to the ploidy state of the corresponding chromosome homolog or haplotype of the parent from which the chromosome homolog or haplotype is inherited). Inherited haplotypes can be used to provide a reference genetic code that can be compared to the ploidy state detected in the subject. If aneuploidy exists in the genetic code of either parent, the aneuploidy can be determined to be inherited. If aneuploidy does not exist in the genetic code of either parent, the aneuploidy can be referred to as de novo variation.
[0099] According to some aspects of the present invention, the origin of the haplotype with aneuploidy state is determined.This determination can be based on, for example, the phase determination of variant and the prior probability of the copy number of mother / father.To confirm the determination, additional sequencing can be carried out on one parent (originating parent) or both parents.For example, whole genome sequencing (for example, shotgun sequencing) can be carried out on (both) parents, which can confirm the corresponding copy number in originating parent.
[0100] According to certain aspects of the present invention, the subject may be an embryo or a fetus. As used herein, "embryo" may refer to a cellular organism produced by sexual reproduction, including a zygote, a morula, and a blastocyst, up to the stage of development at which the embryo becomes a fetus. An embryo may exist in vitro (e.g., for IVF purposes) or in the uterus. As used herein, "fetus" may refer to an unborn child produced by sexual reproduction and present in the uterus, beginning at a developmental stage at which the unborn child can no longer be characterized as an embryo. Thus, a subject may be considered either an embryo or a fetus from the single-cell stage until the fetus is born. In humans, a child is usually considered a fetus at about 8 weeks after conception. The types of genetic material that can be effectively obtained from an embryo or fetus, as well as the techniques and inherent risks involved, are well understood in the art.
[0101] Determining the ploidy state of a fetal embryo (including calling de novo changes) can generally be performed as described elsewhere herein (e.g., for a born child or adult individual). However, de novo detection in a non-born subject can present certain challenges. For example, cellular DNA for performing non-mispropagation phase determination may not be readily available. For example, collecting a bodily fluid sample, such as a blood sample containing circulating blood cells, may be impractical or impossible depending on the stage of development. Furthermore, generally, collecting cellular material from an embryo or fetus may pose a risk to the subject's viability or health (e.g., spontaneous abortion). According to some aspects, cellular DNA can be obtained from a biopsy of the embryo or fetus, as known in the art. In preferred embodiments of performing ploidy state determination on an embryo or fetus, non-mispropagation phase determination can be performed on samples collected from one or more genetic relatives, e.g., the mother and / or father. Cellular DNA can be obtained, for example, from a bodily fluid (e.g., blood) sample or other tissue type obtained from a genetic relative and used to correct the phase state of the reference genetic code, as described elsewhere herein. Extracellular DNA can be collected from genetic relatives as needed. In some embodiments, the reference genetic code can be constructed, at least in part, based on sequencing (e.g., whole-genome shotgun sequencing) of one or more genetic relatives, as known in the art. See, for example, Kitzman et al., Sci Transl Med. 2012 Jun 6;4(137):137ra76 (doi:10.1126 / scitranslmed.3004323). For example, analysis of the genomes of genetic relatives can identify variants for subsequent analysis in the subject. Extracellular DNA from an embryonic or fetal subject can be collected for analysis according to any suitable method known in the art. For example, cffDNA can be collected from the blood of a target fetus or a mother carrying a target embryo until fully developed. Extracellular DNA can be harvested from the blastocoelic fluid of the embryo or from the cell culture medium used to culture the embryo for IVF, as is known in the art.Extracellular DNA from a fetus or embryo can be used, at least in part, to determine a subject's genome (e.g., via whole-genome shotgun sequencing) and / or to establish a reference genetic code for ploidy state calling. See, for example, Kitzman et al., Sci Transl Med. 2012 Jun 6;4(137):137ra76 (doi:10.1126 / scitranslmed.3004323). Sequencing of extracellular DNA can be used, at least in part, to determine the phase of a subject's genome or reference genetic code (e.g., via molecular techniques known in the art). Sequences of one or more genetic relatives and / or population reference panels can be used in combination with sequencing of extracellular DNA to provide an at least partially phased genome (before any correction of phase by non-error-propagating phasing techniques). Extracellular DNA collected from an embryonic or fetal subject can be used to generate allele frequency and / or read depth signals from which ploidy state calls can be made, as described elsewhere herein. The allele frequency signal can be corrected using non-error-propagating phasing techniques performed on the cellular DNA of one or more genetic relatives of the subject.
[0102] Examples of specific associations between aneuploidy (e.g., CNVs or whole chromosome abnormalities) and disease are well known in the art. According to some aspects of the present invention, determining ploidy status can be used to inform decisions regarding IVF. The methods described herein can be performed on a single embryo or on multiple embryos (e.g., multiple embryos candidate for implantation). Determining ploidy status can be used to select one or more embryos for implantation and / or to select one or more embryos for discard / disposal. Determining ploidy status can be used to select one or more embryos for freezing (either when an embryo is selected for possible future implantation, or when an embryo is not the first candidate for implantation but is not desired to be discarded). For example, a disease risk determination can be made for an embryo based at least in part on the detection of the aneuploidy status for a chromosome or chromosome segment (e.g., the identification of CNVs, particularly CNVs with known associations with disease). In some embodiments, embryos without identified aneuploidy (e.g., CNVs) can be selected for implantation or freezing. In some embodiments, embryos may be ranked based entirely or at least in part on the identification of aneuploidies (e.g., by the number of CNVs and / or the presence of specific CNVs). Determination of ploidy state by the methods described herein may be used independently or in combination with existing methods of preimplantation genetic testing (PGT), as is well known in the art.
[0103] According to some aspects of the present invention, determining ploidy status can be used to inform the decision about pregnancy, especially when the subject is a fetus.For example, the decision about whether to continue or terminate pregnancy can be based on determining ploidy status (for example, identifying aneuploidy) in the same manner as the decision about IVF, as described elsewhere herein.The determination of ploidy status by the method described herein can be used independently or in combination with existing methods of prenatal diagnosis, as is well known in the art.
[0104] According to certain aspects of the present invention, the determination of ploidy status can be used to inform further testing and / or diagnostic methods. For example, once aneuploidy is identified, additional PGD or prenatal diagnostic testing can be indicated. In some instances, the additional testing can be specific to one or more diseases associated with the detected aneuploidy. In some instances, particularly when the subject is an embryo or fetus, more invasive procedures can be performed on the subject. For example, a tissue biopsy can be performed directly on the embryo or fetus to sequence cellular DNA or perform other diagnostics on the cellular material. Karyotyping can be performed on the subject. In some embodiments, the additional testing can be performed substantially simultaneously (at approximately the same level of development) with the determination of ploidy status. In some embodiments, the additional testing can be performed on a delayed schedule, allowing for further development to occur (e.g., for development from embryo to fetus and / or after implantation of the embryo via IVF). In some embodiments, additional testing can be performed on a born subject (e.g., an infant or pediatric subject) based on the determination of ploidy status performed when the subject was an embryo and / or fetus.
[0105] According to some aspects of the present invention, the determination of ploidy status can be used to inform treatment decisions for a subject. For example, once aneuploidy is identified, the subject can be treated for a disease or condition associated with aneuploidy. Treatment can include any treatment appropriate for the subject's developmental stage. For example, gene editing can be performed on the embryo, and / or prenatal treatment can be administered to the fetus (or the mother carrying the fetus). In some embodiments, treatment can be performed on a delayed schedule, allowing further development to occur (e.g., for development from embryo to fetus and / or after implantation of the embryo via IVF). In some embodiments, treatment can be administered to a born subject (e.g., an infant or child subject) based on a determination of ploidy status made when the subject was an embryo and / or fetus. Early detection of aneuploidy (e.g., while in utero) can allow for earlier treatment in infants and children, which can result in improved outcomes.
[0106] Disease diagnosis In addition to the diagnosis described elsewhere herein based on the known association of aneuploidy (e.g., CNV) with disease, the methods described herein can be used to identify new associations between aneuploidy and disease.By identifying the same aneuploidy in a population of subjects with a particular disease or predisposition to disease, the association between aneuploidy and disease can be established.
[0107] To clarify the function of SNPs, particularly in relation to disease, one can use the phase determined by non-error-propagation phasing of one or more rare aneuploid variants and identify adjacent SNPs known to be associated with disease (e.g., within the same haplotype block or within a set of two phases determined to be in the same phase alignment by the methods described herein). The rare variant and the identified SNP can be determined to be in linkage disequilibrium. The rare variant can be effectively linked to the identified SNP by increasing the SNP's contribution to disease risk (e.g., in a polygenic risk score (PRS)) compared to other adjacent SNPs (e.g., in linkage disequilibrium with the identified SNP). Thus, linkage of a rare variant to a more common SNP can improve the predictive power of the more common SNP, as it is associated with disease predisposition.
[0108] Once the aneuploid variant associated with disease is identified, sequencing can be carried out in other subjects for diagnostic purposes to determine predisposition to disease.Sequencing can be targeted to capture aneuploid variant.Sequencing can be carried out to target adjacent SNPs, such as the adjacent SNPs that are determined to be in linkage disequilibrium with aneuploid variant (for example, through microarray), as described elsewhere herein.Sequencing can be carried out to target both aneuploid variants (for example, rare variants) and SNPs (for example, common SNPs).
[0109] Diagnosis of disease can be based at least in part on the presence or absence of one or more aneuploid variants, and / or at least in part on one or more SNPs that are determined to be in linkage disequilibrium with one or more aneuploid variants.As is well known in the art, diagnosis can be based on, for example, PRS.Treatment for disease can be informed based on any of the diagnostic methods described herein.For example, a subject can be treated (including preventive treatment) for a disease that the subject has been diagnosed with or has been diagnosed as having or at least having an increased predisposition to developing.Diagnosis and treatment can be performed in combination with other clinical factors and variables, as is understood in the art.
[0110] Determining the phase of germline mosaic variants The methods described herein can be used to identify haplotypes in affected individuals with aneuploid variants. Gametes from affected individuals can be screened for IVF purposes (e.g., to avoid gametes with the identified haplotypes).
[0111] According to one aspect of the present invention, the use of non-error-propagating phasing technology can be applied to determine the phase of germline mosaic variants in affected individuals. Such affected individuals may include, for example, individuals with Noonan syndrome or rasopathy. This phased information can be used to inform decisions regarding IVF, as described elsewhere herein. For example, the phased information can be used to determine which haplotypes should be avoided in subsequent generations using IVF and PGT.
[0112] According to one aspect of the invention, long phased reads can be used to include prediction of rare variants in the genome of an embryo by linking rare variants to common variants (e.g., SNPs) in each of two parents, and then subsequently determining which SNPs were inherited in the embryo and then inferring the inheritance of that rare variant in the embryo. [Example]
[0113] [Example 1] To simulate a chromosomal imbalance (amplification) on human chromosome 21, a dataset of synthetic reads corresponding to specific haplotypes was generated from a phased genome. Briefly, reads from nucleotide positions 30227447 to 44327015 of genetic sample NA12878 were added to data generated using the 10XGENOMICS® synthetic long read approach (CHROMIUM® product) according to the method described in Samadian et al., PLoS Comput Biol. 2018 Mar 28;14(3):e1006080 (doi:10.1371 / journal.pcbi.1006080), which is incorporated herein by reference in its entirety. Input to this software included a phased VCF file containing a phase shift error at approximately 37 Mb and a sequencing file (bam). 200,000 of these reads were then added to a standard set of shotgun reads obtained from the 1000 Genomes repository. Positions predicted to be "0|1" based on the Platinum Genomes variant set for sample NA12878 were assigned the "A" haplotype, and positions predicted to be "1|0" were assigned the "B" haplotype. See, e.g., Eberle et al., Genome Res. 2017 Jan;27(1):157-164 (doi:10.1101 / gr.210500.116), incorporated herein by reference in its entirety. Positions were filtered for depths greater than 5 reads or greater than 20 reads. Each position was assigned the "A" or "B" allele based on the phase of the input phased VCF file. Figure 1 shows the allele balance for heterozygous sites (SNPs) in terms of the proportion of A alleles based on a dataset of synthetic reads for chromosomes.
[0114] As shown in Figure 2, to improve the signal-to-noise ratio of the allele balance signal, consecutive SNPs on the same haplotype determined by dilution pool sequencing were binned, and the allele balance signal was averaged across the binned region. In Figure 3, the allele balance signal was averaged across a 300-Kb window of the haplotype block. As is evident from the averaged allele balance signals in Figures 2 and 3, two different aneuploidies appear to exist: a chromosomal amplification of the A haplotype (specifically, a trisomy from approximately 30 Mb to 37 Mb) followed immediately by a chromosomal deletion of the A haplotype (specifically, a monosomy from approximately 37 Mb to 44 Mb). The haplotype block determined by dilution pool sequencing across the aneuploidy region is shown at the bottom of Figure 3.
[0115] Data from the Hi-C experiment for sample NA12878 was downloaded from staging.4dnucleome.org / filesprocessed / 4DNFIY9YBG6I / . The Hi-C data could be used to identify switch errors in the phased vcf and then correct the allele balance data to accurately call aneuploidies, as described below. Since the reference is hg38, the vcf file was mapped to hg38. The tool "extractHAIRS" from the program HapCut2 was used to generate pieces of evidence supporting various combinations of phase blocks, as described in Edge et al., Genome Res. 2017 May;27(5):801-812 (doi:10.1101 / gr.213462.116), which is incorporated herein by reference in its entirety.
[0116] The Hi-C data were used to evaluate the phase alignment of two phase sets. One phase set was defined as the set of SNPs spanning approximately 30 Mb to 37 Mb, and the second phase set was defined as the remaining SNPs on chromosome 21 from approximately 37 Mb onward. Hi-C fragments containing informative reads (two or more overlapping heterozygous variants) were clustered into sparse subgroups in which variants were self-consistent across the subgroups. As shown in Figure 4, subgroups that at least partially overlapped both phase sets (i.e., subgroups with at least one SNP from each of the two phase sets) were further filtered and evaluated from the Hi-C data. Overlapping subgroups were determined to be either fully concordant (i.e., no discordant haplotype calls, such as "00," "000," or "0000") or discordant (i.e., at least one discordant haplotype call, such as "01," "011," or "0111"). The total number of subgroups, including the distribution of perfectly matched and mismatched fragments, was tabulated. As shown in Figure 4, there were 20 subgroups in total, with 19 mismatches and one match compared to the diluted pool sequence. The number of fragments represents the number of fragment reads within each subgroup, with each fragment containing at least two of the SNPs supporting the haplotype call, but not necessarily each of the SNPs in the subgroup. To assess the distribution of observed concordant and mismatched measurements, we used a binomial distribution to calculate the probability that the observed distribution occurred purely by chance, assuming equal likelihood of obtaining concordant and mismatched measurements. The binomial probability was extremely low, and the probability of a skewed distribution occurring purely by chance was less than 0.01%. Therefore, it was determined that the putative phase alignment between the two phase sets was actually incorrect or misaligned, and therefore the Hi-C measurements overlapping the two phase sets were primarily discordant.Assuming that the phasing of the first phase set (spanning approximately positions 30 Mb to 37 Mb) was correct and that of the second phase set (37 Mb and beyond) was incorrect due to the nature of the switch error introduced between the two phase sets, the phase of the second phase set was reversed, and the true allelic balance signal averaged over a 300 Kb window of the haplotype block was corrected as shown in Figure 5. The true allelic balance signal indicated 14 Mb of aneuploidy spanning approximately positions 30 Mb to 44 Mb, which could theoretically correspond to an amplification of haplotype A or a deletion of haplotype B.
[0117] [Example 2] We replicated the simulated dataset from Example 1, but downsampled reads corresponding to aneuploidy (amplification of haplotype A) on chromosome 21 to approximately 9% of the measured cells, with approximately 91% of the cells showing euploidy across the same chromosome segment. Figure 6A shows the raw allelic balance signal for a 30.3-37 Mb portion of the chromosome relative to the heterozygous locus (SNP). The allelic balance signal across this range has a mean of 0.5232 and a standard deviation of 0.1141. Figure 6B shows the same allelic balance signal averaged over a 300 Kb window of the haplotype block determined by diluted pool sequencing. As is evident from Figure 6B, the allelic balance shift introduced by the 9% aneuploid cells is more readily discernible, with the standard deviation reduced to 0.0258 as a result of binning. Thus, this example demonstrates the ability to call amplification even at low allele fractions.
[0118] [Example 3] In this example, a population of disomy (D) measurements and trisomy (
number
number
number
[0119] The total probability of disomy is equal to the total probability of trisomy (i.e.,
number
number
number
[0120] The probability of a signal X1 corresponding to a false positive (i.e., incorrectly characterizing a disomy as trisomy) was then calculated from the cumulative distribution function using X1 as follows:
number
[0121] A method for making disomy / trisomy calls from using two signals—X1, the read depth signal, and an independent signal X2 (e.g., allele balance signal) together—was computationally simulated according to the calling scheme shown in Table 1 below.
[0122] [Table 1]
[0123] As noted above, the same assumptions were made for the distribution of signal X2 as for the distribution of signal X1. The probabilities of calling a false positive and making no call at all based on using both distributions according to Table 1 are determined as follows in Table 2, where "normcdf" is the normal cumulative distribution function (e.g., as in MATLAB®).
[0124] [Table 2]
[0125] Assuming m1 = 6 and m2 = 6 / sqrt(3), the probability values were calculated as follows: P FPX1 =0.0013;P FPX2 =0.0416; and P FPX1X2 =0.000056.
[0126] [Example 4] Populations of disomy (D) measurements and trisomy
number
number
number
[0127] Again, the total probability of disomy is equal to the total probability of trisomy (i.e.,
number
number
number
number
[0128] The false positive rate was then estimated by integrating the joint probability function as follows:
number
number
number
[0129] X2 was then solved as follows:
number
[0130] Therefore, the false positive rate is
number
[0131] The false positive rate can then be empirically calculated using the following MATLAB® code, where “sum” is the false positive rate for the different signal means m1 and m2: %variables n=2000; m1=6; m2=6 / sqrt(3); lim=20; delta=2*lim / (n-1); x1_vec=[-lim:delta:lim]; x2_vec=[-lim:delta:lim]; sum=0; for x1=x1_vec ind=find(x2_vec>(m1^2+m2^2-2*m1*x1) / (2*m2)); for x2=x2_vec(ind) sum=sum+exp(-0.5*(x1^2+x2^2))*delta^2 / (2*pi); end end sum
[0132] Simulations were performed using the same signal averages as in Example 3, where "sum" corresponds to the probability of observing a false positive in this joint probability scenario, combining signal average m1 with the slightly weaker signal average m2. The probability of a false positive was determined to be P(false positive) = sum = 0.00026, whereas the individual probability (as assessed in Example 3) was determined to be higher: P FPX1 =0.0013 and P FPX2 =0.0416.
[0133] Simulations demonstrate that combining two independent signals, one with a variance three times higher than the other, can reduce the false positive rate by at least five times compared to using either signal alone.
[0134] [Example 5] In a manner similar to Example 1, a synthetic aneuploid mixture of DNA was generated using amplification starting at position 30.3 Mb on chromosome 21. Figure 8A shows the read depth signal for positions 31 Mb to 37 Mb, and Figure 8B illustrates a histogram of binned read depth measurements for positions 31 Mb to 37 Mb. Similarly, Figure 9A shows the allele balance signal for positions 31 Mb to 37 Mb, and Figure 9B illustrates a histogram of binned allele balance measurements for positions 31 Mb to 37 Mb. Figure 9C shows a histogram of binned allele balance measurements, where measurements were averaged across 50 adjacent SNPs.
[0135] The average signal-to-noise was calculated from the aggregated data as described in U.S. Patent No. 8,682,592 to Rabinowitz et al., issued March 25, 2014, which is incorporated herein by reference in its entirety. As described in the theoretical simulations in Examples 3 and 4, the threshold signal value for indicating trisomy was selected to be midway between the average diploid signal and the average triploid signal for both read depth and allelic balance, approximating a scenario in which the probability of calling a false negative equals the probability of calling a false positive, as in Examples 3 and 4, although other thresholds can be selected. The average signal for diploidy was determined by calculating the average measurement across positions 20 Mb to 30.3 Mb, and the average signal for triploidy was determined by calculating the average measurement across positions 30.3 Mb to 37 Mb. Thus, the thresholds were determined to be 31.5 reads per position and 58% A (0.58) for read depth and allelic balance signals, respectively.
[0136] Signal-to-noise plots were generated for the read depth signal and allele balance signal across approximately 2500 measurements / positions of amplification by subtracting the corresponding threshold value from the signal value at each position and then normalizing for the noise level by dividing by the standard deviation measured across the region of amplification. Figure 10 shows the signal-to-noise plot for the read depth signal, and Figure 11 shows the signal-to-noise plot for the allele balance signal. Figure 12 shows the integrated signal resulting from adding together the signal-to-noise values for read depth and allele balance. The mean and standard deviation of the integrated signal shown in Figure 12 were calculated to be 0.4940 and 0.11, respectively.
[0137] Although the present invention has been described and exemplified in sufficient detail to enable those skilled in the art to make and use it, various substitutions, modifications, and improvements will become apparent without departing from the spirit and scope of the invention. The examples provided herein are representative of preferred aspects, are illustrative, and are not intended as limitations on the scope of the invention. Modifications in the examples and other uses will occur to those skilled in the art. These modifications are encompassed within the spirit of the invention and defined by the scope of the claims.
[0138] It will be apparent to those skilled in the art that various substitutions and modifications can be made to the invention disclosed herein without departing from the scope and spirit of the invention. It is understood that various aspects of the invention are combinable unless physically possible or unless the context dictates otherwise.
[0139] All patents and publications mentioned in this specification are indicative of the level of skill of those skilled in the art. All patents and publications are herein incorporated by reference to the same extent as if each individual publication was specifically and individually indicated to be incorporated by reference.
[0140] The present invention illustratively described herein may suitably be practiced in the absence of any element or elements, or any limitation or limitations not specifically disclosed herein. Thus, for example, in each example herein, any of the terms "comprising," "consisting essentially of," and "consisting of" may be replaced with either of the other two terms. The terms and expressions used are used as terms of description, not limitation, and the use of such terms and expressions is not intended to exclude the features shown and described or equivalents thereof, but it is recognized that various modifications are possible within the scope of the claimed invention. Thus, while the present invention has been specifically disclosed by preferred aspects and optional features, it should be understood that modifications and variations of the concepts disclosed herein may be employed by those skilled in the art, and that such modifications and variations are considered to fall within the scope of the invention as defined by the appended claims.
Claims
1. A method for determining the true allele balance signal for a chromosomal segment, To obtain a reference genetic code comprising two or more sets of phases, wherein each set of phases has one or more variants of interest, and the reference genetic code is at least partially phase-determined; Obtaining the allele balance signal for one or more variants of interest from sequencing performed on a sample of genetic material; Obtaining a plurality of sequenced reads using a non-error propagation technique, wherein each read includes at least one of one or more variants of the subject of interest; Based on the aforementioned multiple leads, the phase alignment of the two or more sets of phases is determined to be either the same phase or different phases; Determining the true allele balance signal by correcting the phase state of at least one variant of interest based on the determined phase alignment of the two or more sets of phases, wherein correcting the phase state includes correcting a switch error in the at least partially phase-determined reference genetic code, thereby correcting the allele balance signal; Methods that include...
2. The method according to claim 1, wherein the non-error propagation technique includes capturing chromosome three-dimensional structure, single-cell template strand sequencing, or chromosome isolation.
3. The method according to claim 1 or 2, further comprising using the corrected allele balance signal to determine the ploidy status of a chromosomal segment, wherein determining the ploidy status includes calling a copy number variant (CNV).
4. The method according to claim 3, wherein the non-error propagation technique includes capturing the chromosome three-dimensional structure, and the chromosome three-dimensional structure capture is Hi-C.
5. The method according to claim 4, wherein determining the phase alignment based on the plurality of reads includes determining whether the majority of the reads agree or disagree with respect to an estimated phase state alignment between the two or more sets of phases, the estimated phase state alignment being based on at least partial phase states of the reference genetic code.
6. The method according to claim 5, wherein determining the phase alignment based on the plurality of reads includes determining or estimating the probability that the amount of agreement or mismatch observed between the two or more sets of phases from the plurality of reads is a result of chance.
7. The method according to claim 6, wherein the probability is a binomial probability and it is assumed that the observed fragments have equal probability of being a match or a mismatch.
8. The method according to claim 3, wherein the allele balance signal and the plurality of reads are obtained from a sample of the genetic material.
9. The allele balance signal mentioned above is obtained from the following: (a) Body fluid sample or tissue biopsy sample, (b) The plurality of reads are the same population of cells obtained therefrom, (c) extracellular DNA, (d) Whole genome shotgun sequencing, (e) Sequencing used to determine one or more haplotype blocks, (f) Any combination of (a) to (e), The method according to claim 3.
10. The aforementioned multiple reads are obtained from the following: (a) Body fluid sample, tissue biopsy sample, or population of cells (b) Cellular DNA, or (c) Combinations of (a) to (b), The method according to claim 9.
11. The aforementioned reference genetic code is obtained, at least in part, from the following: (a) Sequence determination used to generate the allele balance signal, (b) sequencing of genetic material from one or more genetically related individuals of the subject from whom the allele balance signal is obtained, and / or (c) Whole genome shotgun sequencing of the subject from which the allele balance signal is obtained, The method according to claim 10.
12. (a) The sequencing used to generate the allele balance signal is from sequencing of normal tissue or germline tissue in the subject from which the allele balance signal is obtained, (b) The sequencing of genetic material from the one or more genetic relatives is derived from the sequencing of germline tissue of the one or more genetic relatives, and / or (c) The one or more genetic relatives are the mother and / or father, The method according to claim 11.
13. This also includes: (a) Performing the non-fault propagation technique to obtain the plurality of leads, (b) Performing the sequencing on a sample of the genetic material in order to obtain the allele balance signal, (c) Collecting a sample of the genetic material, wherein the allele balance signal is obtained from the sample of the genetic material. (d) Collecting a sample of the genetic material, wherein the plurality of reads are obtained from the sample of the genetic material, or (e) Any combination of (a) to (d), The method according to claim 3.
14. The allele balance signal is (a) Averaging across multiple binned variants within a region of at least 50,000 base pairs, (b) Averaging across multiple binned variants within a region of 100 million base pairs or less, (c) Averaged across haplotype blocks, The method according to claim 3.
15. The method according to claim 14, wherein the haplotype block is determined by dilution pool sequencing.
16. The method according to claim 3, wherein the allele balance signal is filtered with respect to a minimum read depth, the minimum read depth being 5, 10, 15, 20, or 25 reads.
17. The method according to claim 3, wherein the two or more sets of phases are adjacent sets of phases in the reference genetic code, and each of the adjacent sets of phases includes a variant of interest not more than 250 million base pairs from the other variant of interest.
18. The method according to claim 3, wherein the plurality of reads are filtered to include reads that include at least two variants of interest from each of the two or more sets of phases.