Methods and compositions for proximity ligation
The method of nucleic acid fragmentation and recombinase-based proximity ligation addresses the challenge of analyzing limited genomic sequences, enabling precise determination of cell-specific mutations and three-dimensional structures using nucleic acid barcodes.
Patent Information
- Application Number
- JP2025041622
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2020-04-23
- Filing Date
- 2025-03-14
- Publication Date
- 2025-07-23
AI Technical Summary
Obtaining high-quality adjacent genomic sequences is difficult, especially when the source material available for sequence analysis is limited, and existing methods for efficiently and accurately analyzing and assembling such data remain a challenge.
A method involving nucleic acid fragmentation, size selection, and proximity ligation of segments using recombinases like integrases or transposases, followed by sequencing and mapping to determine cell-specific mutations and three-dimensional nucleic acid structures, with optional use of nucleic acid barcodes for cell identification.
Enables accurate determination of cell-specific mutations and three-dimensional nucleic acid structures, even with limited sample input, by assigning paired-ends to originating compartments and preserving conformational information.
Smart Images

Figure 2025108429000001_ABST
Abstract
Description
Technical Field
[0001] Cross-reference This application claims the benefit of U.S. Provisional Patent Application No. 62 / 867,463, filed Jun. 27, 2019; U.S. Provisional Patent Application No. 62 / 931,069, filed Nov. 5, 2019; U.S. Provisional Patent Application No. 63 / 011,490, filed Apr. 17, 2020; U.S. Provisional Patent Application No. 62 / 870,297, filed Jul. 3, 2019; and U.S. Provisional Patent Application No. 63 / 014,422, filed Apr. 23, 2020, each of which is hereby incorporated by reference in its entirety.
Background Art
[0002] Obtaining high-quality adjacent genomic sequences is often difficult, especially when the source material available for sequence analysis is limited. Although obtaining raw sequence data has become faster and less costly, appropriate methods for efficiently and accurately analyzing and assembling the data remain a challenge.
Summary of the Invention
[0003] In one aspect, a method of nucleic acid analysis is provided. Optionally, the method comprises: (a) obtaining a stabilized biological sample comprising a nucleic acid molecule complexed with at least one nucleic acid binding protein; (b) contacting the stabilized biological sample with a non-specific endonuclease to cleave the nucleic acid molecule into a plurality of segments; (c) attaching a first segment and a second segment of the plurality of segments at one junction; and (d) subjecting the plurality of segments to size selection to obtain a plurality of selected segments. Optionally, the plurality of selected segments are from about 145 to about 600 bp. Optionally, the plurality of selected segments are from about 100 to about 2500 bp. Optionally, the plurality of selected segments are from about 100 to about 600 bp. Optionally, the plurality of selected segments are from about 600 to about 2500 bp. Optionally, the method further comprises preparing a sequencing library from the plurality of segments prior to step (d). Optionally, the method further comprises subjecting the sequencing library to size selection to obtain a size-selected library. Optionally, the size-selected library is in a size range from about 350 bp to 1000 bp. Optionally, the size selection is performed using gel electrophoresis, capillary electrophoresis, size selection beads, or a gel filtration column. Optionally, the method further comprises analyzing the plurality of selected segments to obtain a QC value. Optionally, the QC value is a chromatin digestion efficiency (CDE) based on the percentage of segments sized between 100 and 2500 bp prior to step (d). Optionally, the method further comprises selecting a sample for further analysis when the CDE value is at least 65%. Optionally, the QC value is a chromatin digestion index (CDI) based on the ratio of the number of mononucleosome-sized segments to the number of dinucleosome-sized segments prior to step (d). Optionally, the method further comprises selecting a sample for further analysis when the CDI value is greater than -1.5 and less than 1.In some cases, the method further comprises binding the plurality of segments to one or more surfaces following the step of contacting the stabilized biological sample with a non-specific endonuclease. In some cases, the one or more surfaces comprise one or more beads. In some cases, the one or more beads are solid phase reversible immobilization (SPRI) beads. In some cases, the stabilized biological sample comprises a stabilized cell lysate. In some cases, the stabilized biological sample comprises stabilized intact cells. In some cases, the stabilized biological sample comprises stabilized intact nuclei. In some cases, step (b) is performed prior to the lysis of intact cells or intact nuclei. In some cases, the method further comprises lysing the cells and / or nuclei in the stabilized biological sample prior to step (c). In some cases, the stabilized biological sample comprises less than 3,000,000 cells. In some cases, the stabilized biological sample comprises less than 1,000,000 cells. In some cases, the stabilized biological sample comprises less than 100,000 cells. In some cases, the stabilized biological sample comprises less than 10 μg of DNA. In some cases, the stabilized biological sample comprises less than 1 μg of DNA. In some cases, the non-specific endonuclease is DNase (DNase). In some cases, the DNase is DNase I. In some cases, the DNase is DNase II. In some cases, the DNase is micrococcal nuclease. In some cases, the DNase is selected from one or more of DNase I, DNase II, and micrococcal nuclease. In some cases, the stabilized biological sample has been treated with a cross-linking agent. In some cases, the cross-linking agent is a chemical fixative. In some cases, the chemical fixative comprises formaldehyde. In some cases, the chemical fixative comprises psoralen. In some cases, the chemical fixative comprises disuccinimidyl glutarate (DSG). In some cases, the chemical fixative comprises ethylene glycol bis(succinimidyl succinate) (EGS).In some cases, the chemical fixative includes disuccinimidyl glutarate (DSG) and ethylene glycol bis(succinimidyl succinate) (EGS). In some cases, the crosslinking agent is ultraviolet light. In some cases, the stabilized biological sample is a crosslinked paraffin-embedded tissue sample. In some cases, the method further includes contacting a selected plurality of segments with an antibody. In some cases, the method further includes performing immunoprecipitation on the plurality of segments. In some cases, immunoprecipitation is performed subsequent to the attaching step. In some cases, the attaching step includes filling in sticky ends using biotinylated nucleotides. In some cases, the attaching step includes filling in sticky ends using non-tagged nucleotides. In some cases, the attaching step includes ligating blunt ends. In some cases, the attaching step includes adding an overhang. In some cases, adding an overhang includes adenylation. In some cases, the attaching step includes contacting at least a first segment and a second segment with at least one bridging oligonucleotide. In some cases, the bridging oligonucleotide is at least 10 bp in length. In some cases, the bridging oligonucleotide is at least 12 bp in length. In some cases, the bridging oligonucleotide is 12 bp in length. In some cases, the bridging oligonucleotide includes a barcode sequence. In some embodiments, the first oligonucleotide includes an affinity tag. In some cases, the affinity tag is biotin. In some cases, the attaching step includes successively contacting at least a first segment and a second segment with a plurality of bridging oligonucleotides. In some cases, the attaching step results in a sample, cell, nucleus, chromosome, or nucleic acid molecule of the stabilized biological sample receiving the unique sequence of the bridging oligonucleotide. In some cases, at least one bridging oligonucleotide is linked to one immunoglobulin-binding protein or one fragment thereof.In some cases, at least one cross-linked oligonucleotide is linked or fused to two or more immunoglobulin-binding proteins or two or more fragments thereof. In some cases, the immunoglobulin-binding protein is selected from Protein A, Protein G, Protein A / G, and Protein L. In some cases, the attaching step includes contacting at least a first segment and a second segment with a barcode. In some cases, the method does not include a shearing step. In some cases, the method further includes (e) obtaining at least some sequences on both sides of the junction to generate a first read pair. In some cases, the method further includes (f) mapping the first read pair to a set of contigs, and (g) determining a path across the set of contigs that represents the order and / or orientation to the genome. Alternatively, or in combination, the method further includes (f) mapping the first read pair to a set of contigs, and (g) determining the presence of a structural variant or a decrease in heterozygosity in a stabilized biological sample from the set of contigs. Alternatively, or in combination, the method further includes (f) mapping the first read pair to a set of contigs, and (g) assigning a phase to the variants in the set of contigs. In some cases, the variant is a human leukocyte antigen (HLA) variant. In some cases, the variant is a killer cell immunoglobulin-like receptor (KIR) variant. Alternatively, or in combination, the method includes (f) mapping the first read pair to a set of contigs, (g) determining the presence of a variant in a stabilized organism from the set of contigs, and (h) performing one or more steps selected from (1) confirming a disease stage, prognosis, or treatment regimen for the stabilized biological sample, (2) selecting a drug based on the presence of the variant, or (3) confirming the efficacy of a drug on the stabilized biological sample. In some cases, the DNase is linked or fused to an immunoglobulin-binding protein or a fragment thereof.In some cases, DNase is linked to two or more immunoglobulin-binding proteins or fragments thereof. In some cases, the immunoglobulin-binding protein is selected from Protein A, Protein G, Protein A / G, and Protein L.
[0004] In another aspect, a method is provided that includes the following steps: (a) obtaining a stabilized biological sample comprising a nucleic acid molecule complexed with at least one nucleic acid-binding protein; (b) contacting the stabilized biological sample with micrococcal nuclease (MNase) to cleave the nucleic acid molecule into a plurality of segments; and (c) attaching a first segment and a second segment of the plurality of segments at one junction. Optionally, the method further includes (d) subjecting the plurality of segments to size selection to obtain a plurality of selected segments. Optionally, the plurality of selected segments are from about 145 to about 600 bp. Optionally, the plurality of selected segments are from about 100 to about 2500 bp. Optionally, the plurality of selected segments are from about 100 to about 600 bp. Optionally, the plurality of selected segments are from about 600 to about 2500 bp. Optionally, the method further includes, prior to step (d), preparing a sequencing library from the plurality of segments. Optionally, the method further includes subjecting the sequencing library to size selection to obtain a size-selected library. Optionally, the size-selected library is in a size range from about 350 bp to 1000 bp. Optionally, the size selection is performed using gel electrophoresis, capillary electrophoresis, size selection beads, or a gel filtration column. Optionally, the method further includes analyzing the plurality of selected segments to obtain a QC value. Optionally, the QC value is a chromatin digestion efficiency (CDE) based on the percentage of segments sized between 100 and 2500 bp prior to step (d). Optionally, the method further includes selecting a sample for further analysis when the CDE value is at least 65%. Optionally, the QC value is a chromatin digestion index (CDI) based on the ratio of the number of mononucleosome-sized segments to the number of dinucleosome-sized segments prior to step (d). Optionally, the method further includes selecting a sample for further analysis when the CDI value is greater than -1.5 and less than 1.In some cases, the method further includes a step of binding the plurality of segments to one or more surfaces following the step of contacting the stabilized biological sample with MNase. In some cases, the one or more surfaces include one or more beads. In some cases, the one or more beads are solid phase reversible immobilization (SPRI) beads. In some cases, the stabilized biological sample includes a stabilized cell lysate. In some cases, the stabilized biological sample includes stabilized intact cells. In some cases, the stabilized biological sample includes stabilized intact nuclei. In some cases, step (b) is performed prior to the lysis of intact cells or intact nuclei. In some cases, the method further includes a step of lysing the cells and / or nuclei in the stabilized biological sample prior to step (c). In some cases, the stabilized biological sample includes less than 3,000,000 cells. In some cases, the stabilized biological sample includes less than 1,000,000 cells. In some cases, the stabilized biological sample includes less than 100,000 cells. In some cases, the stabilized biological sample includes less than 10 μg of DNA. In some cases, the stabilized biological sample includes less than 1 μg of DNA. In some cases, the stabilized biological sample is further treated with DNase. In some cases, the DNase is DNase I. In some cases, the DNase is DNase II. In some cases, the DNase is selected from one or more of DNase I and DNase II. In some cases, the stabilized biological sample is treated with a cross-linking agent. In some cases, the cross-linking agent is a chemical fixative. In some cases, the chemical fixative includes formaldehyde. In some cases, the chemical fixative includes psoralen. In some cases, the chemical fixative includes disuccinimidyl glutarate (DSG). In some cases, the chemical fixative includes ethylene glycol bis(succinimidyl succinate) (EGS). In some cases, the chemical fixative includes disuccinimidyl glutarate (DSG) and ethylene glycol bis(succinimidyl succinate) (EGS). In some cases, the cross-linking agent is ultraviolet light.In some cases, the stabilized biological sample is a cross-linked paraffin-embedded tissue sample. In some cases, the method further comprises contacting a selected plurality of segments with an antibody. In some cases, the method further comprises performing immunoprecipitation on the plurality of segments. In some cases, immunoprecipitation is carried out subsequent to the step of attaching. In some cases, the step of attaching comprises filling in the sticky ends using a biotin-tagged nucleotide. In some cases, the step of attaching comprises filling in the sticky ends using an untagged nucleotide. In some cases, the step of attaching comprises ligating blunt ends. In some cases, the step of attaching comprises adding an overhang. In some cases, adding an overhang comprises adenylation. In some cases, the step of attaching comprises contacting at least a first segment and a second segment with a cross-linking oligonucleotide. In some cases, the cross-linking oligonucleotide is at least 10 bp in length. In some cases, the cross-linking oligonucleotide is at least 12 bp in length. In some cases, the cross-linking oligonucleotide is 12 bp in length. In some cases, the cross-linking oligonucleotide comprises a barcode sequence. In some embodiments, the first oligonucleotide comprises an affinity tag. In some cases, the affinity tag is biotin. In some cases, the step of attaching comprises successively contacting at least a first segment and a second segment with a plurality of cross-linking oligonucleotides. In some cases, the step of attaching results in a sample, cell, nucleus, chromosome, or nucleic acid molecule of the stabilized biological sample receiving the unique sequence of the cross-linking oligonucleotide. In some cases, at least one cross-linking oligonucleotide is linked to one immunoglobulin-binding protein or one fragment thereof. In some cases, at least one cross-linking oligonucleotide is linked to two or more immunoglobulin-binding proteins or two or more fragments thereof. In some cases, the immunoglobulin-binding protein is selected from Protein A, Protein G, Protein A / G, and Protein L.In some cases, the attaching step includes bringing at least the first segment and the second segment into contact with the barcode. In some cases, the method does not include a shearing step. In some cases, the method further includes (e) obtaining at least some sequences on both sides of the junction to generate a first read pair. In some cases, the method further includes (f) mapping the first read pair to a set of contigs, and (g) determining a path across the set of contigs that represents the order and / or orientation to the genome. Alternatively, or in combination, the method further includes (f) mapping the first read pair to a set of contigs, and (g) determining the presence of structural variants or a decrease in heterozygosity in the stabilized biological sample from the set of contigs. Alternatively, or in combination, the method further includes (f) mapping the first read pair to a set of contigs, and (g) assigning phases to the variants in the set of contigs. In some cases, the variant is a human leukocyte antigen (HLA) variant. In some cases, the variant is a killer cell immunoglobulin-like receptor (KIR) variant. Alternatively, or in combination, the method further includes (f) mapping the first read pair to a set of contigs, (g) determining the presence of variants in the stabilized organism from the set of contigs, and (h) performing one or more steps selected from (1) a step of checking the disease stage, prognosis, or treatment regimen for the stabilized biological sample, (2) a step of selecting a drug based on the presence of the variant, or (3) a step of checking the drug efficacy against the stabilized biological sample. In some cases, MNase is linked or fused to an immunoglobulin-binding protein. In some cases, MNase is linked or fused to two or more immunoglobulin-binding proteins or fragments thereof. In some cases, the immunoglobulin-binding protein is selected from protein A, protein G, protein A / G, and protein L.
[0005] In an additional aspect, a nucleic acid library is provided that includes the following: (a) a first cell genomic library component that includes a plurality of first cell genomic fragment pairs, where at least one of the first cell genomic fragment pairs includes two first cell genomic segments that are tethered via a nucleic acid segment that includes a first cell genomic library display tag, the first cell genomic library component; and (b) a second cell genomic library component that includes a plurality of second cell genomic fragment pairs, where at least one of the second cell genomic fragment pairs includes two second cell genomic segments that are tethered via a nucleic acid segment that includes a second cell genomic library display tag, the second cell genomic library component. Optionally, the two first cell genomic segments that are tethered via a nucleic acid segment that includes a first cell genomic library display tag display a first cell genomic structure in a first cell. Optionally, the two second cell genomic segments that are tethered via a nucleic acid segment include a second cell genomic library display tag that displays a second cell genomic structure in a second cell, where the second cell genomic structure is different from the first cell genomic structure. Optionally, the first cell genomic library component is obtained from an isolated eukaryotic nucleus. Optionally, a plurality of the first cell genomic fragment pairs are flanked on both sides by recombinase sites. Optionally, the recombinase site is an integrase integration site. Optionally, the recombinase site is a transposase mosaic end. Optionally, at least one of the recombinase sites of the recombinase site includes an exonuclease-resistant portion. Optionally, the exonuclease-resistant portion includes phosphorothioate. Optionally, the nucleic acid segment that includes a first cell genomic library display tag further includes a recombinase left border and a recombinase right border. Optionally, the recombinase is integrase. Optionally, the recombinase is transposase. Optionally, the nucleic acid segment that includes a first cell genomic library display tag includes an affinity tag. Optionally, the affinity tag includes biotin.In some cases, at least some library members are clone copies. In some cases, the co-occurrence of read pairs mapped to comparable regions of nucleic acid references indicates the distance between the regions in a cell. In some cases, the distance is a relative distance.
[0006] In an additional aspect, a system is provided, the system including a plurality of cell genome aliquots, where at least some of the cell genome aliquots include a genome that combines portions that retain the positional information of genomic components, and the system includes a plurality of recombinase nucleic acid aliquots, where at least some of the recombinase nucleic acid aliquots include an identifiable sequence relative to at least one other aliquot. Optionally, the recombinase is an integrase. Optionally, the recombinase is a transposase. Optionally, at least some of the cell genome aliquots include fragmented genomic molecules. Optionally, at least some of the fragmented genomic molecules include integration site ends. Optionally, the cell genome aliquots include eukaryotic cell genome aliquots. Optionally, the genomic binding portion that retains the positional information of genomic components includes chromatin components. Optionally, the genomic binding portion that retains genomic component positional information includes nucleosomes. Optionally, the plurality of cell genome aliquots include integrase enzymes. Optionally, the plurality of recombinase nucleic acid aliquots include integrase nucleic acid molecules with integrase integration sites. Optionally, the plurality of cell genome aliquots include transposase enzymes. Optionally, the plurality of recombinase nucleic acid aliquots include transposase nucleic acid molecules with transposase mosaic ends. Optionally, at least one of the plurality of integration sites includes an exonuclease-resistant portion. Optionally, at least one of the plurality of mosaic ends includes an exonuclease-resistant portion. Optionally, the exonuclease-resistant portion includes phosphorothioate. Optionally, at least some of the recombinase nucleic acid molecules include a recombinase left border and a recombinase right border. Optionally, the nucleic acid segment including the first cell genome library display tag includes an affinity tag. Optionally, the affinity tag includes biotin. Optionally, the identifiable sequence relative to at least one other aliquot includes a plurality of identical nucleic acid sequences in the aliquot.In some cases, the first aliquot comprises a plurality of nucleic acid molecules having an identifiable arrangement common to at least one other aliquot. In some cases, the plurality of cell genomic aliquots includes single cell genomic aliquots.
[0007] In an additional aspect, a method for assaying chromosomal conformational variations between at least two cells is provided. In some cases, the method comprises obtaining genomic nucleic acids from two cells in which the chromosomal conformational variation is conserved; introducing internal cleavage into the genomic nucleic acids from the two cells; and linking two exposed termini adjacent to the internal cleavage site via one of a plurality of tagged segments, wherein the tag of the first cell is distinguishable from the tag of the second cell. In some cases, the genomic nucleic acids from the two cells are isolated prior to the linking step. In some cases, the method comprises amplifying the nucleic acid molecule resulting from the linking step to create an amplicon comprising chromosomal cleavage junction termini linked by an internal segment comprising a distinguishable sequence such that the chromosomal cleavage junctions from the two cells are distinguishable. In some cases, the method comprises isolating a fragment derived from a first of two cells comprising two previously exposed termini linked by a first tagged segment and a fragment derived from a second of at least two cells comprising two previously exposed termini linked by a second tagged segment. In some cases, the obtained paired-end sequence information comprises at least some first tag information and at least some second tag information. In some cases, the method comprises assigning paired-ends to common proximities in the cell. In some cases, the step of assigning paired-ends to common proximities comprises counting the occurrences of paired-ends that map to two common clusters and relating proximities relative to the occurrences. In some cases, the fragment from the first cell and the fragment from the second cell are isolated in a common volume. In some cases, the fragment from the first cell and the fragment from the second cell are sequenced in a common volume. In some cases, the step of linking two exposed termini adjacent to the internal cleavage site comprises linking two exposed ends that were not adjacent immediately adjacent prior to the step of introducing the internal cleavage.In some cases, the step of linking two exposed terminals that are adjacent to each other at the internal cut site includes linking two exposed terminals that were remote from each other on a common nucleic acid molecule prior to the step of introducing the internal cut. In some cases, the step of linking two exposed terminals that are adjacent to each other at the internal cut site includes linking two exposed ends that were physically adjacent to each other prior to the step of introducing the internal cut. In some cases, at least two cells comprise at least two cell populations.
[0008] In a further aspect, a method is provided that includes obtaining a stabilized sample comprising a nucleic acid molecule complexed with at least one nucleic acid binding protein, cleaving the nucleic acid molecule into a plurality of segments comprising at least a first segment and a second segment, attaching an adapter comprising a first recombinase site to the first segment and the second segment, and contacting the first segment and the second segment with a linker comprising a second recombinase site in the presence of recombinase, thereby generating a linked nucleic acid comprising a first sequence from the first segment, a linker sequence from the linker, and a second sequence from the second segment. Optionally, the recombinase is integrase. Optionally, the recombinase is transposase. Optionally, the method further comprises sequencing at least a portion of the linked nucleic acid. Optionally, the sequencing step comprises sequencing at least a portion of the first sequence and at least a portion of the second sequence. Optionally, the method further comprises mapping at least a portion of the first sequence and at least a portion of the second sequence to a genome. Optionally, the method further comprises performing a three-dimensional genome analysis using information from the sequencing step. Optionally, the stabilized biological sample is a cross-linked sample. Optionally, the step of obtaining a stabilized sample comprises obtaining a sample and stabilizing the sample. Optionally, the step of obtaining a stabilized sample comprises obtaining a pre-stabilized sample. Optionally, the nucleic acid binding protein comprises chromatin or a component thereof. Optionally, the cleaving step comprises enzymatic digestion. Optionally, the enzymatic digestion comprises digestion with one or more restriction enzymes. Optionally, the enzymatic digestion comprises digestion with one or more non-specific nucleases. Optionally, the one or more non-specific nucleases comprise DNase or MNase. Optionally, the step of attaching the first recombinase site comprises ligation. Optionally, the first recombinase site and the second recombinase site comprise integrase sites attP and attB.In some cases, the adapter further includes a sequencing adapter region. In some cases, the sequencing adapter region includes a Y adapter. In some cases, the sequencing adapter region includes a P5 and / or P7 adapter. In some cases, the first recombinase site and the second recombinase site include transposase mosaic ends. In some cases, the linker sequence includes an affinity tag. In some cases, the affinity tag is biotin. In some cases, the linker sequence includes a barcode sequence. In some cases, the barcode sequence indicates the originating compartment. In some cases, the barcode sequence indicates the originating cell. In some cases, the barcode sequence indicates the originating cell population. In some cases, the barcode sequence indicates the originating organism. In some cases, the barcode sequence indicates the originating species.
[0009] In a further aspect, a method is provided that includes obtaining a stabilized biological sample comprising a nucleic acid molecule complexed to at least one nucleic acid binding protein; cleaving the nucleic acid molecule into a plurality of segments comprising at least a first segment and a second segment; attaching the first segment to the second segment, thereby creating a proximally ligated segment; recovering the proximally ligated segment; and sequencing at least a portion of the proximally ligated segment, wherein the sequencing adapter is not attached to the proximally ligated segment after recovery. Optionally, the attaching step is performed by ligating the first segment to the second segment. Optionally, the attaching step is performed using a recombinase. Optionally, the attaching step is performed via a linker. Optionally, the linker comprises an affinity tag. Optionally, the affinity tag is biotin. Optionally, the method further comprises attaching a recombinant adapter comprising recombinase sites to the first segment and the second segment prior to the attaching step of (c). Optionally, the recombinant adapter comprises a sequencing adapter. Optionally, the sequencing adapter comprises a Y adapter. Optionally, the sequencing adapter comprises a P5 and / or P7 adapter.
[0010] A method is provided herein, the method comprising: (a) obtaining a stabilized biological sample comprising a nucleic acid molecule complexed to at least one nucleic acid binding protein; (b) contacting the stabilized biological sample with Nase to cleave the nucleic acid molecule into a plurality of segments; (c) attaching a first segment and a second segment of the plurality of segments at one junction; and (d) subjecting the plurality of segments to size selection to obtain a plurality of selected segments. Optionally, the plurality of selected segments are from about 145 to about 600 bp. Optionally, the plurality of selected segments are from about 100 to about 2500 bp. Optionally, the plurality of selected segments are from about 100 to about 600 bp. Optionally, the plurality of selected segments are from about 600 to about 2500 bp. Optionally, the method further comprises preparing a sequencing library from the plurality of segments prior to step (d). Optionally, the method further comprises subjecting the sequencing library to size selection to obtain a size-selected library. Optionally, the size-selected library is in a size range from about 350 bp to 1000 bp. Optionally, the size selection is performed using gel electrophoresis, capillary electrophoresis, size selection beads, or a gel filtration column. Optionally, the method further comprises analyzing the plurality of selected segments to obtain a QC value. Optionally, the QC value is a chromatin digestion efficiency (CDE) based on the percentage of segments sized between 100 and 2500 bp prior to step (d). Optionally, the method further comprises selecting a sample for further analysis when the CDE value is at least 65%. Optionally, the QC value is a chromatin digestion index (CDI) based on the ratio of the number of mononucleosome-sized segments to the number of dinucleosome-sized segments prior to step (d). Optionally, the method further comprises selecting a sample for further analysis when the CDI value is greater than -1.5 and less than 1. Optionally, the stabilized biological sample comprises a stabilized cell lysate.In some cases, the stabilized biological sample contains stabilized intact cells. In some cases, the stabilized biological sample contains stabilized intact nuclei. In some cases, step (b) is performed prior to the lysis of intact cells or intact nuclei. In some cases, the method further comprises, prior to step (c), lysing the cells and / or nuclei in the stabilized biological sample. In some cases, the stabilized biological sample contains less than 3,000,000 cells. In some cases, the stabilized biological sample contains less than 1,000,000 cells. In some cases, the stabilized biological sample contains less than 100,000 cells. In some cases, the stabilized biological sample contains less than 10 μg of DNA. In some cases, the stabilized biological sample contains less than 1 μg of DNA. In some cases, the DNase is DNase I. In some cases, the DNase is DNase II. In some cases, the DNase is micrococcal nuclease. In some cases, the DNase is selected from one or more of DNase I, DNase II, and micrococcal nuclease. In some cases, the stabilized biological sample has been treated with a cross-linking agent. In some cases, the cross-linking agent is a chemical fixative. In some cases, the chemical fixative contains formaldehyde. In some cases, the chemical fixative contains psoralen. In some cases, the chemical fixative contains dithiobis(succinimidyl glutarate) (DSG). In some cases, the chemical fixative contains ethylene glycol bis(succinimidyl succinate) (EGS). In some cases, the cross-linking agent is ultraviolet light. In some cases, the stabilized biological sample is a cross-linked paraffin-embedded tissue sample. In some cases, the method further comprises contacting a plurality of selected segments with an antibody. In some cases, the attaching step comprises filling in sticky ends and ligating blunt ends using a biotin-tagged nucleotide. In some cases, the attaching step comprises contacting at least a first segment and a second segment with at least one cross-linking oligonucleotide. In some cases, the cross-linking oligonucleotide contains a barcode sequence.In some cases, the attaching step includes continuously contacting at least a first segment and a second segment with a plurality of cross-linked oligonucleotides. In some cases, the attaching step results in a biological sample's cells, nuclei, chromosomes, or stabilized nucleic acid molecules receiving the unique sequence of the cross-linked oligonucleotides. In some cases, the attaching step includes contacting at least a first segment and a second segment with a barcode. In some cases, the method does not include a shearing step. In some cases, the method further includes the step of obtaining at least some sequences on both sides of the junction to generate a first read pair. In some cases, the method further includes the steps of: (f) mapping the first read pair to a set of contigs, and (g) determining a path across the set of contigs that represents the order and / or orientation to the genome. In some cases, the method further includes the steps of: (f) mapping the first read pair to a set of contigs, and (g) determining the presence of structural variants or a decrease in heterozygosity in the stabilized biological sample from the set of contigs. In some cases, the method further includes the steps of: (f) mapping the first read pair to a set of contigs, and (g) assigning a phase to the variants in the set of contigs. In some cases, the variant is a human leukocyte antigen (HLA) variant. In some cases, the variant is a killer cell immunoglobulin-like receptor (KIR) variant. In some cases, the method further includes the steps of: (f) mapping the first read pair to a set of contigs, (g) determining the presence of variants in the stabilized organism from the set of contigs, and (h) performing one or more steps selected from the steps of: (1) confirming the disease stage, prognosis, or treatment regimen for the stabilized biological sample, (2) selecting a drug based on the presence of the variant, or (3) confirming the drug efficacy for the stabilized biological sample.
[0011] In an additional aspect, a method is provided, the method comprising: (a) obtaining a stabilized biological sample comprising a nucleic acid molecule complexed to at least one nucleic acid binding protein; (b) contacting the stabilized biological sample with micrococcal nuclease (MNase) to cleave the nucleic acid molecule into a plurality of segments; and (c) attaching a first segment and a second segment of the plurality of segments at one junction. Optionally, the method herein further comprises, (d) subjecting the plurality of segments to size selection to obtain a plurality of selected segments. Optionally, the plurality of selected segments are from about 145 to about 600 bp. Optionally, the plurality of selected segments are from about 100 to about 2500 bp. Optionally, the plurality of selected segments are from about 100 to about 600 bp. Optionally, the plurality of selected segments are from about 600 to about 2500 bp. Optionally, the method herein further comprises, prior to step (d), preparing a sequencing library from the plurality of segments. Optionally, the method herein further comprises subjecting the sequencing library to size selection to obtain a size-selected library. Optionally, the size-selected library is in a size range from about 350 bp to 1000 bp. Optionally, the size selection is performed using gel electrophoresis, capillary electrophoresis, size selection beads, or a gel filtration column. Optionally, the method further comprises analyzing the plurality of selected segments to obtain a QC value. Optionally, the QC value is a chromatin digestion efficiency (CDE) based on the percentage of segments sized between 100 and 2500 bp prior to step (d). Optionally, the method further comprises selecting a sample for further analysis when the CDE value is at least 65%. Optionally, the QC value is a chromatin digestion index (CDI) based on the ratio of the number of mononucleosome-sized segments to the number of dinucleosome-sized segments prior to step (d). Optionally, the method further comprises selecting a sample for further analysis when the CDI value is greater than -1.5 and less than 1.In some cases, the stabilized biological sample contains a stabilized cell lysate. In some cases, the stabilized biological sample contains stabilized intact cells. In some cases, the stabilized biological sample contains stabilized intact nuclei. In some cases, step (b) is performed prior to the lysis of intact cells or intact nuclei. In some cases, the method herein further comprises lysing cells and / or nuclei in the stabilized biological sample prior to step (c). In some cases, the stabilized biological sample contains less than 3,000,000 cells. In some cases, the stabilized biological sample contains less than 1,000,000 cells. In some cases, the stabilized biological sample contains less than 100,000 cells. In some cases, the stabilized biological sample contains less than 10 μg of DNA. In some cases, the stabilized biological sample contains less than 1 μg of DNA. In some cases, the stabilized biological sample is further treated with DNase. In some cases, the DNase is DNase I. In some cases, the DNase is DNase II. In some cases, the DNase is selected from one or more of DNase I and DNase II. In some cases, the stabilized biological sample is treated with a cross-linking agent. In some cases, the cross-linking agent is a chemical fixative. In some cases, the chemical fixative contains formaldehyde. In some cases, the chemical fixative contains psoralen. In some cases, the chemical fixative contains disuccinimidyl glutarate (DSG). In some cases, the chemical fixative contains ethylene glycol bis(succinimidyl succinate) (EGS). In some cases, the cross-linking agent is ultraviolet light. In some embodiments, the stabilized biological sample is a cross-linked paraffin-embedded tissue sample. In some cases, the method herein comprises contacting a selected plurality of segments with an antibody. In some cases, the attaching step comprises filling sticky ends and ligating blunt ends using biotin-tagged nucleotides.In some cases, the attaching step includes contacting at least a first segment and a second segment with at least one crosslinking oligonucleotide. In some cases, the crosslinking oligonucleotide includes a barcode sequence. In some cases, the attaching step includes successively contacting at least a first segment and a second segment with a plurality of crosslinking oligonucleotides. In some cases, the attaching step results in a cell, nucleus, chromosome, or stabilized nucleic acid molecule of a biological sample receiving the unique sequence of the crosslinking oligonucleotide. In some cases, the attaching step includes contacting at least a first segment and a second segment with a barcode. In some cases, the method does not include a shearing step. In some cases, the method herein further includes (e) obtaining at least some sequences on both sides of a junction to generate a first read pair. In some cases, the method herein further includes (f) mapping the first read pair to a set of contigs, and (g) determining a path across the set of contigs that represents the order and / or orientation to the genome. In some cases, the method herein further includes (f) mapping the first read pair to a set of contigs, and (g) determining the presence of structural variants or a decrease in heterozygosity in a stabilized biological sample from the set of contigs. In some cases, the method herein further includes (f) mapping the first read pair to a set of contigs, and (g) phasing variants in the set of contigs. In some cases, the variant is a human leukocyte antigen (HLA) variant. In some cases, the variant is a killer cell immunoglobulin-like receptor (KIR) variant.In some cases, the method herein further includes: (f) mapping a first read pair to a contiguous set; (g) determining the presence of a variant in a stabilized organism from the set of contigs; and (h) performing one or more steps selected from: (1) confirming a disease stage, prognosis, or treatment regimen for the stabilized biological sample; (2) selecting a drug based on the presence of the variant; or (3) confirming the efficacy of a drug against the stabilized biological sample.
[0012] Incorporation by reference All publications, patents, and patent applications mentioned herein are hereby incorporated by reference to the extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] The patent application file includes at least one drawing created in color. A copy of this patent application with color drawings will be provided by the office upon request and payment of the necessary fees.
[0014] A better understanding of the features and advantages of the present invention will be obtained by referring to the following detailed description of exemplary embodiments in which the principles of the invention are used, and the following accompanying drawings.
[0015]
Figure 1A
Figure 1B
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17
Figure 18
Figure 19A
Figure 19B
Figure 19C
Figure 19D
Figure 20
Figure 21
Figure 22
Figure 23A
Figure 23B
Figure 23C
Figure 23D
Figure 24
Figure 25
Figure 26
Figure 27
DETAILED DESCRIPTION OF THE INVENTION
[0016] In one aspect, compositions, systems, and methods are provided herein for generating very long-range read pairs for nucleic acids related to the determination of genomic sequences containing long-range and structural genomic information and the determination of the physical conformation of nucleic acids in cells, with improved results over methods already disclosed in the art. The methods herein can utilize techniques including, but not limited to, DNase digestion, micrococcal nuclease (MNase) digestion, recombinase treatment, size selection, QC control, whole cell or whole nucleus nuclease digestion, single cell analysis, and low input requirements to achieve optimal results. The methods herein can further include utilizing an immunoglobulin binding protein or fragment thereof to target oligonucleotides, or oligonucleotides and nucleases, to antibody binding sites in a nucleic acid sample. Also provided herein are improved methods of HiChIP, HiChIRP, and methylHiC.
[0017] In another aspect, embodiments are provided herein for nucleic acid conformation assessment, nucleic acid sequence analysis, or nucleic acid phase information determination for single cells or multiple cells, or cell populations.
[0018] In some cases, a nucleic acid sample with a conserved conformation or a conformation that has been reconstructed can be fragmented and distributed to aliquots or compartments to which an aliquot identification sequence segment can be added, such that upon analysis of a paired-end library generated from the sample, the paired end can be assigned to the originating compartment, or cell. Thus, cell-specific mutations in the sequence and / or three-dimensional nucleic acid structure can be determined.
[0019] Nucleic Acid Conformation Assessment Disclosed herein are compositions, systems, and methods for determining the physical conformation of nucleic acids in cells, such as single cells or cell populations, that are distinguishable from the physical conformation of a second cell or population of cells. Through the practice of the disclosure herein, nucleic acid molecules that exhibit three-dimensional nucleic acid relative positions can be generated and optionally provided with tags (e.g., nucleic acid barcodes) for identifying a cell or cell population of common origin for a plurality of the molecules.
[0020] Through the practice of the disclosed methods herein, nucleic acids can be harvested such that all or at least some of the three-dimensional structure in the cell is preserved. Exposed nucleic acid loops of such nucleic acids can be cleaved to expose the ends of internal segments that are randomly reattached such that physically proximal ends are likely to attach to each other (proximity ligation). Thus, by determining which exposed termini become attached to each other, data can be obtained indicative of the physical proximity of nucleic acids with adjacent termini in the native cell structure.
[0021] A related approach is disclosed, for example, in US9434985B2 by Dekker et al., published September 6, 2016, which is hereby incorporated by reference in its entirety.
[0022] Through the practice of the disclosed methods herein, components of the paired-end library are further tagged or given sequence information indicative of the originating cell, such that conformational differences between individual cells of the population are readily discernible for the population of cells, or conformational differences between a first population of cells and a second population of cells are readily discernible even when they are analyzed simultaneously. Tags can include, for example, nucleic acid barcodes. In some cases, a tag can include a junction between two nucleic acid segments that are not adjacent in the genome. Nucleic acid molecules can be generated such that when fully or partially sequenced, at least some genomic sequence sufficient to map the ends of the individual genomes to their genomic loci is obtained in many cases, and such that further tagging or linking sequences sufficient to accurately or prospectively identify the originating cell or cell population are obtained. Thus, sequence information indicating that two regions of the genome are physically proximate to each other is obtained, and on the other hand, information indicating the cell or cell population in which this physical conformation occurs is also obtained, such that the sequence information can be evaluated in the context of other physical conformational information that occurs simultaneously in that cell or cell population.
[0023] Genomic nucleic acids or other nucleic acids in cells can be stabilized, and for eukaryotic cells, nuclei can optionally be isolated according to methods known in the art, such as those incorporated herein or otherwise known methods. For example, in FIGS. 1A and 1B, processed and stabilized tissue samples are illustrated. FIG. 1A illustrates a tissue sample with insufficient processing. FIG. 1B illustrates a fully processed and stabilized tissue sample.
[0024] Nucleic acids consistent with the disclosure herein can include any number of exogenous nucleic acids in a sample, such as prokaryotic primary genomic or plasmid nucleic acids, eukaryotic nuclear, mitochondrial or plastid nucleic acids, or, optionally, cytoplasmic nucleic acids such as rRNA, mRNA, or viral or other pathogen, or other sample exogenous nucleic acids.
[0025] Stabilized nucleic acids can optionally be partitioned such that at least some of the nucleic acids are distributed into individual compartments. Exemplary compartments include wells, droplets in an emulsion, or surface locations (e.g., array spots, beads, etc.), which contain distinct patches of differentially addressed linker molecules as described elsewhere herein. Additional compartments known in the art or available to the skilled artisan are also contemplated and consistent with the methods, compositions, and systems disclosed herein.
[0026] The stabilized nucleic acids can be fragmented to expose internal cleavage for later recombination to obtain nucleic acid structure information about a particular cell. Many fragmentation approaches are known in the art and are consistent with the disclosure herein. The nucleic acids can be fragmented using one or more populations of restriction endonucleases, programmable endonucleases such as CRISPR / Cas molecules linked to guide RNAs, non-specific endonucleases (e.g., DNase), tagmentation, shearing, sonication, heating, or other means. In some cases, the DNase is non-sequence specific. In some cases, the DNase is active against both single-stranded DNA and double-stranded DNA. In some cases, the DNase is specific for double-stranded DNA. In some cases, the DNase is preferential for double-stranded DNA. In some cases, the DNase is specific for single-stranded DNA. In some cases, the DNase is preferential for single-stranded DNA. In some cases, the DNase is DNase I. In some cases, the DNase is DNase II. In some cases, the DNase is selected from one or more of DNase I and DNase II. In some cases, the DNase is micrococcal nuclease. In some cases, the DNase is selected from one or more of DNase I, DNase II, and micrococcal nuclease. Other suitable nucleases are also within the scope of this disclosure.
[0027] In particular, the disclosure of Green et al. in WO2014121091A1, published on August 7, 2014 (subsequently published as US20150363550A1 on December 17, 2015 and issued as US10089437B2 on October 2, 2018), is hereby incorporated by reference in its entirety. Similarly, the disclosure of Fields et al. in WO2016019360A1, published on February 4, 2016 (subsequently published as US20170335369A1 on November 23, 2017), is hereby incorporated by reference in its entirety. Similarly, the disclosure of Green et al. in WO2017147279A1, published on August 31, 2017, is hereby incorporated by reference in its entirety.
[0028] Nucleic acids can be bound to a surface either before or after attachment. Exemplary surfaces include, but are not limited to, beads, arrays, wells. In some cases, the surface is a SPRI surface such as SPRI beads. Binding the nucleic acid to the surface prior to attachment can improve the performance of downstream processes such as reducing ligation or attachment between chromosomes and increasing ligation or attachment within a chromosome.
[0029] Nucleic acids may be immunoprecipitated either before or after attachment. Such methods can include fragmenting chromatin and then contacting the fragments with an antibody that specifically recognizes and binds acetylated histone, specifically H3. Examples of such antibodies include, but are not limited to, anti-acetylated histone H3 available from Upstate Biotechnology, Lake Placid, N.Y. Polynucleotides from the immunoprecipitation can then be collected from the immunoprecipitation. Similar targeted enrichment methods can also be used with compounds including, but not limited to, aptamers, oligonucleotides, or other nucleic acid probes and nucleic acid-guided nucleases (e.g., Cas family enzymes such as Cas9 including catalytically inactive or "dead" nucleases).
[0030] Linking nucleic acids, such as nucleic acids having a barcode, a compartment-specific array, or a compartment-identifying array, can be attached to exposed internal terminals to generate nucleic acid segments having a left genomic segment, often having a compartment-specific or compartment-identifying sequence (e.g., a nucleic acid barcode), a linking region, and a right genomic segment, where the left and right genomic segments map to genomic segments that are physically proximal within the source cell.
[0031] Prior to attachment of the exposed nucleic acid terminal, the terminal can be processed. Such processing can include end polishing or blunt ending. A blunt-ended nucleic acid terminal can be ligated directly, for example, to other blunt-ended exposed nucleic acid termini, adapters, or linkers. Such processing can include generating an overhang, for example, by tailing (e.g., A-tailing or adenylation). In one example, the overhang is 1 nucleotide in size. In one embodiment, the overhang is a single A nucleotide. A tailed exposed nucleic acid terminal can be ligated directly, for example, to other tailed exposed nucleic acid termini or to an adapter or linker. Optionally, blunt ending or tailing can incorporate an affinity-tagged nucleic acid, such as a biotinylated nucleic acid. The affinity tag can be used, for example, in downstream capture or enrichment steps. In other cases, blunt ending or tailing can be performed without incorporating an affinity-tagged nucleic acid (e.g., without a biotinylated nucleic acid). The affinity tag can be subsequently added, if desired, for example, to an adapter or linker (e.g., by crosslinking). In one example, the exposed nucleic acid is end polished, an overhang is generated, and the exposed terminus is attached via a crosslinking oligo.
[0032] Attachment is direct, for example, by ligation.
[0033] Attachment may be by a linker or crosslink, such as ligation of one or more linkers or crosslinking nucleic acids that connect one exposed nucleic acid end to another.
[0034] Attachment may be by use of a capping nucleic acid adapter segment that is consistent with recombinase integration, such as integrase or transposase integration. An adapter having a recombinase site can be added to the exposed nucleic acid termini, and then those termini can be joined, for example, by recombination.
[0035] Taking phiC31 integrase barcode delivery as an example, linkers such as cell discrimination linkers or cell-specific linkers (e.g., nucleic acid barcodes) can be enzymatically added as follows.
[0036] Following exposure of internal nucleic acid termini, the integrase site can be ligated to an exposed nucleic acid end such as an exposed linear chromosomal end, such as an internal end or a terminus with a removed telomere. Exemplary integration sites are nucleic acids containing the attP phiC31 integrase integration site, or the attP integration site, although other integration sites are consistent with the disclosure herein. Ligation results in a population of nucleic acid fragments, at least some of which individually contain cell nucleic acid segments adjacent to each end of the integration site, such as segments containing the attP segment. In various embodiments, one or both of fragmentation and attachment of the integration site occur prior to compartmentalization, or one or both of fragmentation and attachment of the integration site follow compartmentalization.
[0037] Figures 19A-19D show an exemplary overview of an attachment approach based on the integrase of phiC31. In Figure 19A, an overview of the integration of phiC31 into Streptomyces via the integrase is shown. Nucleic acids containing the attP (indicated by the dashed line) and attB (indicated by the solid line) sites are shown, although various embodiments other than attB and attP, and enzyme activities other than the integrase are also contemplated and are consistent with the disclosure herein. In Figure 19B, it is shown that the integrase and associated proteins (indicated by circles) bind to the phage attB site and attP sequences in the bacterial genome and trigger strand exchange. In Figure 19C, the results of the integration event are shown. The integration resolves into a linear nucleic acid lacking attB and attP, but having attL and attR, which are chimeric fragments of the attB and attP portions. The attL and attR sites are 3 bp shorter and different in sequence compared to attB and attP. In Figure 19D, it is shown that circular integration or a linker genome is not required. The integration of linear DNA containing the attB site will cause the cleavage of attP containing the DNA.
[0038] Figure 20 depicts the delivery of integrase sites to the exposed internal termini of stable nucleic acids as contemplated herein. For example, attP can be delivered by adapter ligation onto the exposed internal termini of DNase-digested chromatin (indicated by the cylinder). The nucleic acids can be stabilized, optionally, to protect against contact with binding sites such as nucleosomes, or to protect phase information or three-dimensional physical position.
[0039] Figure 21 shows the generation of linker constructs using integrase sites such as attB sites. For example, a minimal 33 nucleotide attB site is sufficient for integration. The flanking sequences can be replaced using selected sequences such as barcodes or other sequences that specify the nucleic acids of a particular source (e.g., cells, droplets, or other compartments, tissues). Figure 22 demonstrates in vivo ligation by integration of linear attB DNA. The results are linear molecules that were either in phase or had exposed internal termini of nucleic acid segments that were joined in physical proximity on the components of a single library. Library components are restricted by an intact integration site (attP in this case), but internal integration sites are disrupted and replaced by the boundaries of attR and attL, such that primers associated with attP can amplify library fragments. By obtaining the sequences adjacent to the internal termini and mapping them to a set of genomes or contigs, contigs or genomic segments can be assigned to common phases or common three-dimensional positions within the cell.
[0040] Figure 25 shows another example of a recombination-based proximity ligation protocol. Genomic DNA containing cross-linked chromatin is digested, for example, with DNase. The exposed termini are polished and, for example, A-tailed with a single A base overhang. Recombinase sites containing adapters that are compatible with the A-tail, such as attB sites, are ligated to the exposed ends. Linkers with corresponding recombinase sites, such as attP sites, are contacted with the sample, and recombination is carried out using a recombinase enzyme (e.g., phiC31) to achieve proximity ligation. The linker optionally contains an affinity agent such as biotin (b) to enable downstream pull-down or other purification or processing. The cross-linking is reversed, and the proximity-ligated nucleic acids are recovered, for example, containing an attB site up to 40 bp, followed by a genomic DNA region 1 up to 150 bp, followed by an attR site and linker sequence containing an affinity agent up to 90 bp, followed by a genomic DNA region 2 up to 150 bp, followed by an attB site up to 40 bp.
[0041] Figure 26 shows a protocol similar to that shown in Figure 25 and involves exemplary adapter and linker sequences. At the top, non-recombinant gDNA with an EP overhang attB adapter is shown, whose sequence GTGCCAGGGCGTGCCCttGGGCTCCCCGGGCGCGATC has the attB site GCCCTTGGGC, and its complementary sequence CGCGCCCGGGGAGCCCaaGGGCACGCCCTGGCAC has the reverse attB site GCCCAAGGGC. Second from the top, non-recombinant gDNA with an attB adapter is shown, whose sequence is GTGCCAGGGCGTGCCCttGGGCTCCCCGGGCGCGTCCCC, and the complementary sequence is GGGGGACGCGCCCGGGGAGCCCaaGGGCACGCCCTGGCAC. Third from the top, a non-recombinant linker containing an attP site and biotin is shown, whose sequence ggagCCCCAACTGGGGTAACCTttGAGTTCTCTCAGTTGGGGaccatggaga / iBiodT / caCCCCAACTGAGAGAACTCaaAGGTTACCCCAGTTGGGGCACTAC contains the attP site with the sequence ACCTTTGAGT and the linker sequence CATGGAGATC. Fourth from the top, one end of the linker recombined with attB / gDNA is shown, whose sequence is ggagCCCCAACTGGGGTAACCTttGAGTTCTCTCAGTTGGGGaccatggaga / iBiodT / caCCCCAACTGAGAGAACTCaaGGGCACGCCCTGGCAC. At the bottom, both ends of the linker recombined with attB / gDNA are shown, whose sequence GTGCCAGGGCGTGCCCttGAGTTCTCTCAGTTGGGGaccatggaga / iBiodT / caCCCCAACTGAGAGAACTCaaGGGCACGCCCTGGCAC has the attR site with the sequence GCCCTTGAGT and the reverse attR site with the sequence ACTCAAGGGC.
[0042] Figure 23A shows exemplary linker molecules and adapter molecule modifications to facilitate library generation. The linker molecules are given an affinity tag (in this case, biotin, denoted by the circle), while the adapter has an exonuclease resistance modification (in this case, phosphothioation (PS), denoted by the asterisk). The affinity tag facilitates the isolation of the linker molecules regardless of whether they are incorporated into the molecules adjacent to the termini. The exonuclease resistance modification on the linker facilitates selective degradation of nucleic acid molecules to which the linker was not added and linker molecules that were not incorporated into the nucleic acid sample molecules adjacent to the termini. Figure 23B shows the affinity purification (in this case, streptavidin, denoted by the semi-circle arc) of tagged molecules regardless of whether they are incorporated into the ligation sites added to the internal termini. Figure 23C shows that attP-directed amplification is used to selectively amplify affinity isolation molecules that preserve integrase sites such as the attP site (in this case, by targeting sites such as those having primers). The presence of the affinity tag and the attP site indicate molecules in which a successful integration event has occurred. Figure 23D shows an alternative, whereby an exonuclease (denoted by the circular sector or "Pac-Man") is used to remove affinity-labeled molecules lacking the exonuclease resistance modification. The presence of the affinity tag and the exonuclease resistance site indicate molecules in which a successful integration event has occurred.
[0043] Alternatively, transposases such as Tn3, Tn5, Tn7, or sleeping beauty transposase can be used for delivery of the barcode. Following exposure of the internal nucleic acid termini, the mosaic termini can be ligated to an exposed nucleic acid terminus such as an internal terminus or an exposed linear chromosomal terminus such as a terminus from which a telomere has been removed. Exemplary mosaic termini are Tn5 mosaic termini, or nucleic acids comprising Tn5 mosaic termini, although other mosaic termini are consistent with the disclosure herein. Ligation results in a population of nucleic acid fragments, at least some of which individually contain a cellular nucleic acid segment flanking each terminus adjacent to a mosaic terminus such as a Tn5 mosaic terminus.
[0044] The recombinase - adapter molecule can further include sequencing adapter sites, such as P5, P7 sites. Figure 27 (top) shows an exemplary attB adapter with a sequencing Y - adapter, the sequence of the attB adapter is GTGCCAGGGCGTGCCCttGGGCTCCCCGGGCGCG, the sequence of P7 is GATCGGAAGAGCACACGTCTGAACTCCAGTCAC, and the sequence of P5 is ACACTCTTTCCCTACACGACGCTCTTCCGATC. Figure 27 (bottom) shows an overview of non - recombinant gDNA with a recombinase - adapter having a sequencing adapter. The adapter to be sequenced is attached to a portion of the attB site that will remain with the genomic DNA after recombination, thereby enabling post - recombination sequencing without further amplification or adapter ligation. The use of a recombinase - adapter containing a sequencing adapter allows for direct sequencing of proximity ligation products, including but not limited to those shown in Figure 27, where no amplification or separate adapter incorporation step is required. This can suppress biases such as amplification bias in the resulting sequence information.
[0045] In some embodiments, in various embodiments, either or both of fragmentation and attachment of the integration site occur prior to compartmentalization, or either or both of fragmentation and attachment of the integration site follow compartmentalization. FIG. 24 depicts an exemplary system for single cell HiC (or other proximity ligation techniques) using integrase-mediated in vivo ligation. A single cell nucleus is combined with integrase and encapsulated in a first set of compartments. The compartments are, in this case, droplets in an emulsion. The nucleus is subjected to sonication to generate internally exposed ends and preserve local three-dimensional information. Adapters are ligated onto the exposed internal ends. The adapters optionally include exonuclease-resistant ends. In this embodiment, the adapters do not convey information identifying the compartments. In a second set of compartments, linkers with unique molecular identifiers (UMIs), such as compartment discriminator sequences, are encapsulated and optionally subjected to linearization that leads to amplification and cleavage. The first and second sets of compartments are merged at a nearly 1:1 ratio, or under conditions such that nucleic acids from two cells will not result in binding to a single compartment.
[0046] Recombinase sites, such as integrase sites or mosaic ends, can be carried on unmodified single-stranded or double-stranded fragments that will optionally be ligated onto internal nucleic acid ends. Alternatively, for ease of later cleanup of sequencing libraries, integration sites that have some single-stranded or double-stranded fragments internally, such as attP sequences, or mosaic ends, such as Tn3, Tn5, Tn7, or Sleeping Beauty transposase mosaic ends, can include at least one modification, such as a modification that interferes with exonuclease or other nuclease activity. An example is a thiosulfate modification that renders exonuclease digestion of a double-stranded fragment with an integration site added to each end impossible.
[0047] Often, recombinase sites or mosaic ends, such as integration sites, are non-specific, and among them, such integration sites, or attP sequences, or sequences within ends such as Tn3, Tn5, Tn7, or Sleeping Beauty transposase mosaic ends, are not used to specify the cellular source of adjacent nucleic acids. Alternatively, often, after compartmentalization of the nucleic acids, the compartments are provided with adapters having distinguishable, specific, or cell-identifying sequences (e.g., nucleic acid barcodes) adjacent to the integration site or mosaic end, or with distinguishable integration sites or mosaic ends, such that the nucleic acids of the first compartment receive an integration segment or mosaic end having a first cell-identifying segment, and the nucleic acid segment of the second compartment receives an integration segment having a second cell-identifying segment.
[0048] Fragments having recombinase boundaries, such as boundaries containing integrase attP segments, can then be contacted with integration sites, such as the attB phiC31 integration site, in a common solution. In one example, the integration enzyme may include phi31 integrase, the integration boundary may include an attP segment, and the integration site may include an attB integration site. Alternatively, the fragment has a mosaic end boundary, such as a Tn3, Tn5, Tn7, or Sleeping Beauty transposase mosaic end boundary.
[0049] When an attB integration site, or a recombinase site such as Tn3, Tn5, Tn7, or a Sleeping Beauty transposase mosaic end, is located next to a linking segment having a sequence that identifies a compartment or cell, such as something unique to the segment or cell source (e.g., a nucleic acid barcode), the sequence identifies adjacent cellular nucleic acids as having originated from a particular or common cell source or compartment, such that even if they have been enlarged by a fragment of a second compartment prior to or simultaneously with sequencing, a plurality of exposed ends from a common cell joined to a common cell or compartment identification segment can be readily identified as having originated from a common cell.
[0050] When a sequence that identifies a cell is delivered via a fragment adjacent to a recombinase site, integration or translocation is preferably carried out following compartmentalization. Thereby, the nucleic acid content of at least some compartments can be identified by the cell-identifying sequence of its own linker, such that even if the nucleic acids forming a plurality of cell sources have become large for sequencing, an internal terminal pair, and proximity information assigned to the vicinity mapped across and including a set of contigs of a largely or fully sequenced genome, can be assigned to a common cell identified from at least one other cell of the sample, such that differences in the predicted nucleic acid three-dimensional conformation can be established.
[0051] Fragments adjacent to the recombination site variously include left and right boundary fragments (e.g., attB sites, or Tn3, Tn5, Tn7, or Sleeping Beauty transposase mosaic ends) linked by a linker region that optionally includes a compartment (e.g., a nucleic acid barcode) that specifies a cell or sequence. The linker region optionally further includes a moiety to facilitate later isolation. Many affinity tags or modified bases are consistent with the disclosure and this specification. Exemplary moieties facilitate physical or chemical isolation of the linker after integrase treatment or transposase treatment. Any number of affinity tags known to those of skill in the art, such as one or more biotin tags that can facilitate isolation based on avidin or streptavidin, are also consistent with the disclosure herein. Alternatively, any antigen, receptor, or ligand that facilitates isolation without interfering with integrase or transposase activity is suitable for some embodiments herein.
[0052] As described above, some library generation approaches include a cleanup step, such as a step to selectively remove unincorporated reagents. For example, exonuclease treatment is often used to selectively remove unbound linker molecules, genomic fragments lacking an attached integration site, or both. Genomic fragments ligated to an integration site fragment having an exonuclease-resistant modification, such as a thiosulfate backbone, are resistant to exonuclease degradation from their termini, and nucleic acid molecules bounded at both termini by an integration site fragment having an exonuclease-resistant modification, such as a thiosulfate backbone, are resistant to degradation at both termini and can survive exonuclease treatment.
[0053] In or in combination with an interaction, some linker molecules contain a counter-affinity tag on the opposite side of the attP integration site or a recombination site such as Tn3, Tn5, Tn7, or the Sleeping Beauty transposase mosaic end, such that the counter-affinity tag is removed in response to the success of the recombination reaction. In such cases, unwanted reagents can be removed by contacting the counter-affinity tag with a binding partner.
[0054] Integrase activity partially disrupts both integration sites, such as the attB and attP sites, as part of the integration event. Thus, by designing primers to anneal the attP integration site, etc. to the ligated adapter site, at least one linker-spanning clone amplicon can be generated, either alone or in combination with linker-based isolation, such that information identifying the cell or aliquot and information adjacent to the internal termini is amplified and, in some cases, sequencing or other downstream analysis is facilitated.
[0055] Following generation of the library and optionally cleaning up the library, the nucleic acids can be fully or partially sequenced, such that sufficient information is obtained for identifying cells or for cell-specific three-dimensional nucleic acid position assessment. As described above, sequencing is preferably performed such that at least some sequence of the genome is obtained that is sufficient to map the individual genomic ends of library elements to their genomic loci, and further sequences are obtained that are sufficient to link and identify the cells of origin as being exact or promising. Thus, sequence information is obtained that indicates two regions of the genome that are physically proximal to each other, and on the other hand, information indicating the cells in which this physical conformation occurs, such that the sequence information can be evaluated in the context of other physical conformational information that occurs simultaneously in those cells. In many cases, although both approaches are consistent with the disclosure and this specification, this information is obtained by paired-end sequencing rather than by full-length sequencing.
[0056] Compositions and methods for determining the physical conformation of nucleic acids in cells, such as single cells distinguishable from the physical conformation of a second cell, can be implemented on many systems consistent with the disclosure herein. Some systems include, for example, the dispensing of immobilized cell nucleic acid material onto a well plate into an emulsion or a first droplet in a well. These droplets further include recombinase sites, such as integrase sites or mosaic ends optionally modified to be exonuclease resistant, as well as integrase or transposase enzymes and ligase enzymes as described herein. Separately, linker nucleic acid molecules can be constructed for delivery into the first droplet of the emulsion. The linker nucleic acid can optionally be dispensed into a second emulsion or a droplet in a second well, amplified optionally using rolling circle amplification, and processed to generate multiple copies of a given linker molecule for the second emulsion droplet.
[0057] Subsequently, the second emulsion droplet and the first emulsion droplet can be merged in pairs, resulting in the assembly of nucleic acid fragments that ligate integrase or transposase to a linker that is integrase or transposase-compatible, and in many cases, a uniform label is shown for each droplet. However, particularly when data analysis indicates the presence of one or more types of tags in the droplets, droplets with two or more identifiers for a nucleic acid sample may still be able to produce meaningful data.
[0058] Instead of a pair of merges, in some cases, the integrase or transposase-compatible linker can be delivered as a colony of solid particles in a reagent stream that contacts the first emulsion droplet via the droplet heading towards the flow merge, which is incorporated herein by reference in its entirety as described in US20170335369A1 published on November 23, 2017. The linker nucleic acid can optionally be amplified on the solid particles or in a gel. The first emulsion droplet can be merged into the flow, and the second emulsion droplet can be recovered by segmenting or splitting the flow, resulting in a desired ratio of nucleic acid clusters to linker particles, such as 1:1, greater than 1:1, or less than 1:1.
[0059] Alternatively, some systems and methods involve distributing either fixed cellular nucleic acid material, whether amplified or not as described above, into wells of a chip or plate, followed by delivering the linker nucleic acid into the compartments.
[0060] Alternatively, in some cases, the delivery of the linker nucleic acid is not temporarily separated from the compartmentalization. Rather, the linker nucleic acid or factors required for enzyme activity or enzyme activity are isolated until a specific treatment such as heating, electromagnetic activation, or other administration is applied to temporarily activate the enzyme activity to effect covalent bonding of the linker, such as via the linker to the exposed termini of the nucleic acid sample.
[0061] Many integrase enzymes are consistent with the disclosure herein. PhiC31 integrase, such as that commercially available from ThermoFisher, offers many advantages for use in the practice of the methods, operation of the systems, and compositions herein. Some of the advantages of this integrase are as described below. It uses small integration sites (attB / attP). The enzyme itself is a small single polypeptide. Integration is irreversible unless a separate enzyme is used to remove the integration event. The activity is high, and the enzyme is easily engineered to alter activity. Nevertheless, because many integration systems are consistent with the disclosure herein, its use does not require the exclusion of other enzymes. Aspects of the disclosure are described with respect to PhiC31 integrase, but the use of any compatible enzyme is contemplated. Many transposase enzymes are consistent with the disclosure herein. Tn5 transposase, such as that commercially available from ThermoFisher, offers many advantages for use in the practice of the methods, operation of the systems, and compositions herein. Some of the advantages of this transposase are as described below. Tn5 uses a 19bp mosaic end recognition sequence, insertion is nearly unbiased and stable, and Tn5 can be delivered into cells for in vivo transposition or to isolated nucleic acids for in vitro reactions. Nevertheless, because many transposase systems, such as Tn3, Tn7, Sleeping Beauty transposase, are consistent with the disclosure herein, its use does not require the exclusion of other enzymes. Aspects of the disclosure may be described with respect to Tn3, Tn5, Tn7, or Sleeping Beauty transposase, but the use of any compatible enzyme is contemplated.
[0062] Array information obtained from library components is evaluated by many approaches, such as those known in the art, for in vitro proximity ligation, Hi-C, Chicago™, or other three-dimensional conformation analysis contexts in the art. Importantly, cell-specific lead pair frequencies can be obtained, such that the frequency with which terminal adjacent sequences map to a particular region or contig of the genome can be evaluated on a cell-specific basis. That is, cell-specific emergence of promising three-dimensional conformations can be evaluated. In some cases, cell-specific signal strength can also be evaluated in relation to cell-specific distances in the three-dimensional conformation, such that a region of nucleic acid can be concluded to be relatively proximal in one cell compared to a second cell with comparable but “weaker” or more distant proximity. That is, both qualitative and quantitative evaluation of three-dimensional structure is consistent with the disclosure herein. In some cases, proximity of one region to a second region is at least partially evaluated by tallying the number of cluster components of a first cluster that occur simultaneously in paired-end reads having cluster components of a second cluster, particularly in library components that share a common compartment identification sequence, such as a unique compartment tag.
[0063] It is not necessary to create structural information by the composite occurrence of the same terminal adjacent sequences in multiple library components. Rather, in some cases, terminal adjacent sequences (for a common “cluster”) mapped near a second terminal adjacent sequence mapping site can enhance evaluation of three-dimensional conformation when members of both clusters map to non-identical regions of a second cluster on a second region of a nucleic acid reference, such as the genome.
[0064] In some cases, the methods disclosed herein are used to label and / or associate polynucleotides or sequence segments thereof and to utilize the data for various applications. In some cases, the disclosure herein provides methods and computing systems that can generate highly contiguous and accurate human genome assemblies having read pairs of less than about 10,000, about 20,000, about 50,000, about 100,000, about 200,000, about 500,000, about 1 million, about 2 million, about 5 million, about 10 million, about 20 million, about 30 million, about 40 million, about 50 million, about 60 million, about 70 million, about 80 million, about 90 million, about 1 billion, about 2 billion, about 3 billion, about 4 billion, about 5 billion, about 6 billion, about 7 billion, about 8 billion, about 9 billion, about 10 billion. In some cases, the disclosure provides methods for stepwise performing or physically assigning linkage information to about 50%, 60%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or more of the heterozygous variants in the human genome with an accuracy of about 50%, 60%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or more.
[0065] In some embodiments, the compositions and methods described herein enable the study of metagenomes (e.g., those found in the human gut). Thus, it is possible to investigate the partial or entire genomic sequences of some or all of the organisms residing in a given ecological environment. Examples include random sequencing of the microorganisms throughout the gut, those found in certain areas of the skin, and those living in toxic waste sites. The composition of the microbial populations in these environments, as well as the interrelated biochemical aspects encoded by each genome, can be determined using the compositions and methods described herein. The methods described herein can, for example, enable metagenomic studies from complex biological environments containing 2, 3, 4, 5, 6, 7, 8, 9, 10, 12, 15, 20, 25, 30, 40, 50, 60, 70, 80, 90, 100, 125, 150, 175, 200, 250, 300, 400, 500, 600, 700, 800, 900, 1000, 5000, or 10,000 or more organisms and / or organism variants.
[0066] Accordingly, the methods disclosed herein may be applied to a complete human genomic DNA sample, but can also be widely applied to a variety of nucleic acid samples, such as reverse transcribed RNA samples, cell-free circulating DNA samples, cancer tissue samples, crime scene samples, archaeological samples, non-human genomic samples, or environmental samples, such as those containing genetic information from two or more organisms, such as organisms that are not easily cultured under laboratory conditions.
[0067] The high accuracy required for cancer genome sequencing can be achieved using the methods and systems described herein. Incorrect reference genomes can pose difficulties in base calling during antigen administration for sequencing cancer genomes. Heterogeneous samples and small starting materials, such as samples obtained by biopsy, pose further difficulties. Additionally, the detection of large-scale structural variations and / or losses of heterozygosity is often important for distinguishing somatic mutations from errors in base calling, in addition to cancer genome sequencing.
[0068] The systems and methods described herein can generate accurate long sequences from complex samples containing two, three, four, five, six, seven, eight, nine, ten, twelve, fifteen, twenty, or more various genomes. Normal, benign, and / or tumor-derived mixed samples may be analyzed optionally without the need for a normal control. In some embodiments, small samples on the order of 100 ng, or small starting samples on the order of several hundred genome equivalents, are utilized to generate accurate long sequences. The systems and methods described herein may enable the detection of mutations, large-scale structural variations, and rearrangements, and phased variant calls may be obtained over long sequences spanning about 1 kbp, about 2 kbp, about 5 kbp, about 10 kbp, 20 kbp, about 50 kbp, about 100 kbp, about 200 kbp, about 500 kbp, about 1 Mbp, about 2 Mbp, about 5 Mbp, about 10 Mbp, about 20 Mbp, about 50 Mbp, about 100 Mbp, or more nucleotides. For example, phased variant calls may be obtained over long sequences spanning about 1 Mbp or about 2 Mbp.
[0069] In one aspect, the methods disclosed herein are used to assemble multiple contigs derived from a single DNA molecule. In some cases, the method includes generating multiple read pairs from a single DNA molecule cross-linked to a number of nanoparticles, and using the read pairs to assemble the contigs. In some cases, the single DNA molecule is cross-linked outside the cell. In some cases, at least 0.1%, 0.2%, 0.3%, 0.4%, 0.5%, 0.6%, 0.7%, 0.8%, 0.9%, 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, 20%, 25%, 30%, 35%, 40%, 45%, or 50% of the read pairs span a distance greater than 1 kB, 2 kB, 3 kB, 4 kB, 5 kB, 6 kB, 7 kB, 8 kB, 9 kB, 10 kB, 15 kB, 20 kB, 30 kB, 40 kB, 50 kB, 60 kB, 70 kB, 80 kB, 90 kB, 100 kB, 150 kB, 200 kB, 250 kB, 300 kB, 400 kB, 500 kB, 600 kB, 700 kB, 800 kB, 900 kB, or 1 MB on a single DNA molecule. In some cases, at least 0.5%, 0.6%, 0.7%, 0.8%, 0.9%, 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, or 20% of the read pairs span a distance greater than 5 kB, 6 kB, 7 kB, 8 kB, 9 kB, 10 kB, 15 kB, 20 kB, 30 kB, 40 kB, 50 kB, 60 kB, 70 kB, 80 kB, 90 kB, 100 kB, 150 kB, or 200 kB on a single DNA molecule. In further cases, at least 0.5%, 0.6%, 0.7%, 0.8%, 0.9%, 1%, 2%, 3%, 4%, or 5% of the read pairs span a distance greater than 20 kB, 30 kB, 40 kB, 50 kB, 60 kB, 70 kB, 80 kB, 90 kB, or 100 kB on a single DNA molecule. In certain cases, at least 1% or 5% of the read pairs span a distance greater than 50 kB or 100 kB on a single DNA molecule.Optionally, the read pairs are generated within 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 30, 40, 50, or 60 days. In some cases, the read pairs are generated within 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, or 18 days. In further cases, the read pairs are generated within 7, 8, 9, 10, 11, 12, 13, or 14 days. In certain cases, the read pairs are generated within 7 or 14 days.
[0070] The haplotypes determined using the methods and systems described herein may be assigned to computer resources, such as computer resources on a network, such as a cloud system. Short variant calls can be corrected using relevant information stored in computer resources if necessary. Structural variations can be detected based on combined information from short variant calls and information stored in computer resources. Problematic portions of the genome, including but not limited to segmental duplications, regions prone to structural variation, highly variable and medically relevant MHC regions, centromere and telomere regions, and other heterochromatin regions with repetitive regions, low sequence accuracy, high mutation rates, ALU repeats, segmental duplications, or other relevant problematic portions known in the art, can be reassembled for improved accuracy.
[0071] Sample types can be assigned to array information locally or in network-connected computer resources such as the cloud. If the source of the information is known, for example, when the source of the information is from cancer or normal tissue, this source can be assigned to the sample as part of the sample type. Other examples of sample types generally include, but are not limited to, tissue type, sample collection method, presence of infection, type of infection, processing method, sample size, etc. When complete or partial comparative genomic sequences such as normal genomic sequences compared to cancer genomes are available, the differences between the sample data and the comparative genomic sequences can be determined and optionally output.
[0072] The methods of the present disclosure can be used for the analysis of genetic information in genomic regions that can interact with a selected region of interest in addition to the selected region of interest of the genome. The amplification methods as disclosed herein can be used in, but are not limited to, known devices, kits, and methods in the art of gene analysis such as U.S. Pat. Nos. 6,449,562, 6,287,766, 7,361,468, 7,414,117, 6,225,109, and 6,110,709. In some cases, the amplification methods of the present disclosure can be used to amplify target nucleic acids for DNA hybridization studies to determine the presence or absence of polymorphisms. Polymorphisms, or alleles, can be associated with diseases or disorders such as genetic diseases. In some other cases, polymorphisms can be associated with susceptibility to diseases or disorders, for example, polymorphisms are associated with poisoning, degenerative, and age-related diseases, cancer, etc. In other cases, polymorphisms can be associated with useful characteristics such as an increase in coronary artery health, resistance to diseases such as HIV or malaria, or resistance to adult diseases such as osteoporosis, Alzheimer's disease, or dementia.
[0073] The compositions and methods of the present disclosure can be used for diagnostic, prognostic, therapeutic, patient stratification, drug development, treatment selection, and screening purposes. The present disclosure provides the advantage that many different target molecules can be analyzed at once from a single biological molecule sample using the methods of the present disclosure. This enables, for example, various diagnostic tests to be performed on one sample.
[0074] The methods provided herein can greatly advance the field of genomics by overcoming the substantial barriers posed by these repetitive regions, and thereby enable significant progress in many domains of genome analysis. To perform de novo assembly using conventional methods, one must either accept assemblies fragmented into many small scaffolds, devote considerable time and resources to generating large insertion libraries, or use other approaches to generate more contiguous assemblies. Such approaches can include obtaining very deep sequencing coverage, constructing BAC or fosmid libraries, optical mapping, or most likely, some combination of these and other techniques. Due to the extreme resource and time requirements, such approaches are out of reach for most small research institutions and have hampered the study of non-model organisms. The methods described herein can generate very long-range read sets, such that de novo assembly may be achieved in a single sequencing run. This can reduce the assembly cost by orders of magnitude and shorten the time that would otherwise have taken months or years to just a few weeks. In some cases, the methods disclosed herein can generate multiple read sets in less than 14 days, less than 13 days, less than 12 days, less than 11 days, less than 10 days, less than 9 days, less than 8 days, less than 7 days, less than 6 days, less than 5 days, less than 4 days, less than 3 days, less than 2 days, less than 1 day, or within a range spanning any two of the periods specified herein. In some cases, the method can generate multiple read sets in about 10 to 14 days. Genome construction for even the most niche organisms will become routine, phylogenetic flow cytometry analysis will no longer suffer from lack of validation, and projects such as 10k genomes will become feasible.
[0075] The methods described herein are capable of assigning contig information that has been provided, generated, or de novo synthesized into physical linkage groups such as chromosomes or shorter contiguous nucleic acid molecules. Similarly, the methods disclosed herein enable the contigs to be arranged in a linear order along the physical nucleic acid molecules relative to each other. Similarly, the methods disclosed herein enable the contigs to be oriented relative to each other in a linear order along the physical nucleic acid molecules.
[0076] Similarly, the methods disclosed herein can lead to advances in structural and phasing analysis for medical purposes. There is surprising heterogeneity among cancers, among individuals with the same type of cancer, or even within the same tumor. Extracting the cause from the resulting effects requires very high accuracy and throughput at a low cost per sample. In the domain of personalized medicine, one of the gold standards of genomic care is a sequenced genome with all mutants thoroughly characterized and phased, including large and small structural rearrangements and novel mutations. To achieve this with conventional techniques currently requires a similar effort to that required for de novo assembly, which is too expensive and laborious for routine medical treatment. In some cases, the methods disclosed herein can rapidly produce a complete and accurate genome at low cost and thereby provide many highly sought-after capabilities in the study and treatment of human diseases.
[0077] Furthermore, by applying the methods disclosed herein to phasing, it is possible to combine the convenience of statistical approaches with the accuracy of familial genetic analysis, resulting in significant savings in terms of money, labor, and samples compared to using either of these methods alone. Highly desirable de novo variant phasing analysis, which was impossible with prior art, can be readily performed using the methods disclosed herein. This is particularly important because the majority of human variants are rare (minor allele frequency of less than 5%). Phasing information is valuable for population genetics studies that can derive significant benefit from a network of highly connected haplotypes (populations of variants assigned to a single chromosome) compared to unlinked genotypes. Haplotype information can enable higher-resolution investigations of historical changes in population size, migration, and exchange between subpopulations, and can enable tracing specific variants back to particular parents and grandparents. This, in turn, reveals the genetic transmission of variants associated with disease and the interactions between variants when aggregated in a single individual. In further instances, the methods of the disclosure enable the preparation, sequencing, and analysis of libraries of extremely long-range read sets (XLRS) or extremely long-range read pairs (XLRP).
[0078] In some embodiments of the disclosure, a sample of a subject's tissue or DNA is provided and the method returns an assembled genome, an alignment with called variants (including large-scale structural variants), a phased variant call, or any additional analysis. In other embodiments, the methods disclosed herein provide an XLRP library directly for each individual.
[0079] In various embodiments, the methods disclosed herein generate extremely long-range read pairs separated by large distances. The upper limit of this distance can be improved by the ability to collect large-sized DNA samples. In some cases, the read pairs span genomic distances of up to 50, 60, 70, 80, 90, 100, 125, 150, 175, 200, 225, 250, 300, 400, 500, 600, 700, 800, 900, 1000, 1500, 2000, 2500, 3000, 4000, 5000 kbp, or more. In some cases, the read pairs span a genomic distance of up to 500 kbp. In some cases, the read pairs span a genomic distance of up to 2000 kbp. The methods disclosed herein can be integrated and constructed based on standard techniques in molecular biology and are well-suited for increasing efficiency, specificity, and genomic coverage. In some cases, the read pairs are generated within about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, less than 60 days, or less than 90 days. In some cases, the read pairs are generated within less than about 14 days. In further cases, the read pairs are generated within less than about 10 days. In some cases, the methods of the present disclosure provide read pairs having at least about 50%, about 60%, about 70%, about 80%, about 90%, about 95%, about 99%, or about 100% accuracy in accurately ordering and / or orienting a large number of contigs, where about 5%, about 10%, about 15%, about 20%, about 30%, about 40%, about 50%, about 60%, about 70%, about 80%, about 90%, about 95%, about 99%, or more, or about 100% of the read pairs. In some cases, the method provides an accuracy of about 90 - 100% in accurately ordering and / or orienting a large number of contigs.
[0080] In other embodiments, the methods disclosed herein can be used with currently utilized sequencing technologies. In some cases, the methods can be used in combination with well-tested and / or widely deployed sequencing instruments. In further embodiments, the methods disclosed herein can be used with techniques and methodologies derived from currently utilized sequencing technologies.
[0081] The methods disclosed herein can dramatically simplify de novo genome assembly for a wide range of organisms. Using conventional techniques, such assemblies are currently limited by short inserts in economical mate pair libraries. It may be possible to generate read pairs with genomic distances up to 40 - 50 kbp available with fosmids, but these are expensive and laborious and are too short for the longest repetitive regions, including centromeres, which range in size from 300 kbp to 5 Mbp in humans. In some cases, the methods disclosed herein can provide read pairs that can span large distances (e.g., megabases or longer), thereby overcoming problems of scaffold completeness. Thus, generating chromosome-level assemblies can become routine by utilizing the methods disclosed herein. Similarly, obtaining long-range phasing information can provide great additional power in population genetics, phylogenetics, and disease studies. In some cases, the methods disclosed herein enable accurate phasing for multiple individuals, thereby expanding the breadth and depth of the ability to survey genomes with respect to the number of individuals in a population and at deep time scales.
[0082] In the area of personalized medicine, XLRS read sets generated from the methods disclosed herein represent a significant advancement for the personal genome that is accurate, low-cost, phased, and rapidly generated. Conventional methods are insufficient in their ability to phase long-range variants, which hinders the characterization of the phenotypic impact of compound heterozygous genotypes. Additionally, structural variants of substantial interest for genomic diseases are difficult to accurately identify and characterize using prior art because their size is large compared to the reads and read inserts used to study them. Read sets spanning from tens of kilobases to megabases or more help alleviate this difficulty, thereby enabling highly parallelized and individualized analysis of structural variants.
[0083] Fundamental evolutionary and biomedical research can be facilitated by technological advancements in high-throughput sequencing. Generating large amounts of DNA sequence data is now relatively low-cost. However, generating high-quality and highly contiguous genomic sequences with prior art is difficult both theoretically and in practice. Additionally, many organisms, including humans, are diploid, and each of these individuals has two haploid copies of the genome. At sites of heterozygosity (e.g., when the allele given by the mother is different from the allele given by the father), it is difficult to know which set of alleles is derived from which parent (known as haplotype phasing). This information can be extremely important for conducting many evolutionary and biomedical investigations, such as studies of the association between diseases and traits.
[0084] The present disclosure provides a method for genome assembly, the method combining techniques for DNA preparation with tagged sequence reads for high-throughput exploration of short-, medium-, and long-range connections corresponding to sequence reads derived from a single physical nucleic acid molecule that binds to complexes such as chromatin complexes within a given genome. The present disclosure further provides methods of using these connections to assist in genome assembly for phasing of haplotypes and / or for metagenomic studies. While the methods presented herein can be used to determine the assembly of a subject's genome, in some cases it should also be understood that the methods presented herein are used to determine the assembly of portions of a subject's genome, such as the assembly of chromosomes or chromatin of varying lengths of the subject. It should further be understood that in certain cases the methods presented herein are used to determine or direct an assembly of a collection of nucleic acid molecules that are not located on a chromosome. Indeed, any nucleic acid sequencing complicated by the presence of repetitive regions that separate non-repetitive contigs can be facilitated using the methods disclosed herein.
[0085] In further cases, the methods disclosed herein enable accurate and predictive results for genotyping, haplotype phasing, and metagenomics using small amounts of material. In some cases, DNA of less than about 100 picograms (pg), about 200 pg, about 300 pg, about 400 pg, about 500 pg, about 600 pg, about 700 pg, about 800 pg, about 900 pg, about 1.0 nanogram (ng), about 2.0 ng, about 3.0 ng, about 4.0 ng, about 5.0 ng, about 6.0 ng, about 7.0 ng, about 8.0 ng, about 9.0 ng, about 10 ng, about 15 ng, about 20 ng, about 30 ng, about 40 ng, about 50 ng, about 60 ng, about 70 ng, about 80 ng, about 90 ng, about 100 ng, about 200 ng, about 300 ng, about 400 ng, about 500 ng, about 600 ng, about 700 ng, about 800 ng, about 900 ng, about 1.0 microgram (μg), about 1.2 μg, about 1.4 μg, about 1.6 μg, about 1.8 μg, about 2.0 μg, about 2.5 μg, about 3.0 μg, about 3.5 μg, about 4.0 μg, about 4.5 μg, about 5.0 μg, about 6.0 μg, about 7.0 μg, about 8.0 μg, about 9.0 μg, about 10 μg, about 15 μg, about 20 μg, about 30 μg, about 40 μg, about 50 μg, about 60 μg, about 70 μg, about 80 μg, about 90 μg, about 100 μg, about 150 μg, about 200 μg, about 300 μg, about 400 μg, about 500 μg, about 600 μg, about 700 μg, about 800 μg, about 900 μg, or less than about 1000 μg is used with the methods disclosed herein. In some cases, the DNA used in the methods disclosed herein is extracted from about 10,000,000, about 5,000,000, about 4,000,000, about 3,000,000, about 2,000,000, about 1,000,000, about 500,000, about 200,000, about 100,000, about 50,000, about 20,000, about 10,000, about 5,000, about 2,000, about 1,000, about 500, about 200, about 100, about 50, about 20, or less than about 10 cells.
[0086] In a diploid genome, it is often more important to know which allelic variants are physically linked on the same chromosome than to map to homologous positions on chromosome pairs. Mapping alleles or other sequences to the physical chromosomes of a particular polyploid chromosome pair is known as haplotype phasing. Short reads derived from high-throughput sequence data rarely allow direct observation of which allelic variants are linked, especially when allelic variants are separated by distances greater than the longest single read, as is most frequent. Computational estimates of haplotype phasing can be unreliable for long distances. The methods disclosed herein enable determination of which variants are physically linked using allelic variants on read pairs.
[0087] In various cases, the methods and compositions of the disclosure enable haplotype phasing of diploid or polyploid genomes with respect to a large number of allelic variants. The methods described herein thus contribute to determining linked allelic variants based on variant information from labeled sequence segments and / or contigs assembled using them. Examples of allelic variants are known by, but not limited to, 1000genomes, UK10K, HapMap, and other projects for discovering genetic variation in humans. In some cases, the association of a disease with a particular gene is more readily revealed by obtaining haplotype phasing data, for example, by the discovery of unlinked inactivating mutations in both copies of SH3TC(2) that cause Charcot-Marie-Tooth neuropathy (Lupski JR, Reid JG, Gonzaga-Jauregui C, et al. N. Engl. J. Med. 362:1181-91, 2010), and by the discovery of inactivating mutations in both copies of ABCG(5) that cause hypercholesterolemia 9.
[0088] On average, one out of 1,000 sites in humans is heterozygous. In some cases, a single lane of data using high-throughput sequencing methods can generate at least about 150,000,000 reads. In further cases, individual reads are about 100 base pairs in length. Assuming a DNA fragment with an average size of 150 kbp is input and 100 paired-end reads are obtained per fragment, it can be expected to observe 30 heterozygous sites per set, that is, per 100 read pairs. All read pairs containing a heterozygous site within a set are homologous (i.e., molecularly linked) to all other read pairs within the same set. This property, in some cases, enables a greater ability to phase the set, as opposed to specific read pairs. There are approximately 3 billion bases in the human genome, and since 1 in 1,000 is heterozygous, there are approximately 3 million heterozygous sites in the average human genome. For the approximately 45,000,000 read pairs containing heterozygous sites, the average coverage of each heterozygous site to be phased using a single lane of high-throughput sequencing methods is approximately (15X) using a typical high-throughput sequencing machine. Thus, the diploid human genome can be phased with high reliability and completely using one lane of high-throughput sequence data regarding sequence variants from a sample prepared using the methods disclosed herein. In some cases, a lane of data is the data of one set of DNA sequence reads. In further cases, a lane of data is the data of one set of DNA sequence reads from a single run of high-throughput sequencing equipment.
[0089] Because the human genome consists of two sets of homologous chromosomes, understanding the individual's true genetic structure requires depicting the copies or haplotypes of the maternal and paternal genetic material. Obtaining haplotypes in an individual is useful in several ways. For example, haplotypes are clinically useful when predicting the outcome of donor-host matching in organ transplantation. Haplotypes are increasingly being used to detect disease associations. In genes showing compound heterozygosity, haplotypes provide information regarding whether two deleterious variants are located on the same allele (i.e., "in cis" using genetic terminology) or on two different alleles ("in trans"), and this information greatly influences the prediction of whether the inheritance of these variants is harmful, and whether the individual has a single non-functional allele with a functional allele and two harmful variant positions, or whether the individual has two non-functional alleles each with a different defect. Haplotypes of populations provide information about the underlying population structure of interest to both epidemiologists and anthropologists, and have provided knowledge about the history of human evolution. In addition, extensive allelic imbalance in gene expression has been reported, and suggests that genetic or epigenetic differences between allelic phases may contribute to quantitative differences in expression. Understanding haplotype structure will describe the mechanisms of variants contributing to allelic imbalance.
[0090] In certain embodiments, the methods disclosed herein include in vitro techniques for fixing and capturing associations between distant regions of the genome as required for long-range linkage and phasing. Optionally, the methods include constructing and sequencing one or more read sets for delivering very distant read pairs on the genome. In further instances, each read set includes two or more reads labeled with a common barcode, which reads may represent two or more sequence segments from a common polynucleotide. Optionally, the interactions arise primarily from probabilistic associations within a single polynucleotide. Optionally, sequence segments that are close to each other in a polynucleotide interact more frequently and with higher probability, while interactions between distant portions of the molecule are less frequent, such that the genomic distance between sequence segments is inferred. Consequently, there is a systematic relationship between the number of pairs connecting two loci on the input DNA and their proximity.
[0091] In some aspects, the present disclosure provides methods and compositions for generating data to achieve extremely high accuracy in phasing. Compared to conventional methods, the methods described herein can phase a higher percentage of variants. In some cases, phasing is achieved while maintaining a high level of accuracy. In further cases, this phased information is extended over a longer range, e.g., over a length greater than about 200 kbp, about 300 kbp, about 400 kbp, about 500 kbp, about 600 kbp, about 700 kbp, about 800 kbp, about 900 kbp, about 1 Mbp, about 2 Mbp, about 3 Mbp, about 4 Mbp, about 5 Mbp, or greater than about 10 Mbp, or greater than about 10 Mbp up to the full length of the chromosome. In some embodiments, greater than 90% of the heterozygous SNPs for a human sample are phased with an accuracy exceeding 99%, where, for example, fewer than about 250 million reads are used by using just one lane of Illumina HiSeq data. In other cases, about 40%, 50%, 60%, 70%, 80%, 90%, greater than 95%, or greater than about 99% of the heterozygous SNPs for a human sample are phased with a high accuracy of greater than about 70%, 80%, 90%, greater than 95%, or greater than 99%, where, for example, fewer than about 250 million or fewer than about 500 million reads are used by using just one or two lanes of Illumina HiSeq data. In some cases, greater than 95% or greater than 99% of the heterozygous SNPs for a human sample are phased with an accuracy higher than about 95% or 99%, where fewer than about 250 million or fewer than about 500 million reads are used. In further cases, additional variants are captured by increasing the read length up to about 200 bp, 250 bp, 300 bp, 350 bp, 400 bp, 450 bp, 500 bp, 600 bp, 800 bp, 1000 bp, 1500 bp, 2 kbp, 3 kbp, 4 kbp, 5 kbp, 10 kbp, 20 kbp, 50 kbp, or up to about 100 kbp.
[0092] The compositions and methods of the present disclosure can be used for gene expression analysis. The methods described herein distinguish nucleotide sequences. The differences between target nucleotide sequences can be, for example, a single nucleotide difference, a nucleic acid deletion, a nucleic acid insertion, or a rearrangement. Such sequence differences regarding more than one base can also be detected. The processes of the present disclosure can detect infectious diseases, genetic diseases, and cancers. Also, it is useful in environmental monitoring, forensic science, and food science. Examples of genetic analysis that can be performed on nucleic acids include, for example, SNP detection, STR detection, RNA expression analysis, promoter methylation, gene expression, virus detection, virus subtyping, and drug resistance.
[0093] The methods of the present invention can be applied to the analysis of biomolecular samples obtained from or derived from a patient to determine whether an affected cell type is present in the sample, the stage of the disease, the patient's prognosis, the patient's ability to respond to a particular treatment, or the best treatment for the patient. The method can also be applied to identify biomarkers for a particular disease.
[0094] In some embodiments, the methods described herein are used for the diagnosis of diseases. As used herein, the term "diagnose" or "diagnosis" of a disease includes predicting or diagnosing a disease, determining the predisposition of a disease, monitoring the treatment of a disease, diagnosing the treatment response of a disease, or the prognosis of a disease, the progression of a disease, or the response to a particular treatment of a disease. For example, a blood sample can be assayed according to any of the methods described herein to determine the presence and / or amount of markers of a disease or malignant cell type in the sample.
[0095] In some embodiments, the methods and compositions described herein are used for the diagnosis and prognosis of diseases.
[0096] A number of immunological, proliferative, and malignant diseases and disorders are particularly suitable for the methods described herein. Immunological diseases and disorders include allergic diseases and disorders, immunodeficiencies, and autoimmune diseases and disorders. Allergic diseases and disorders include, but are not limited to, allergic rhinitis, allergic conjunctivitis, allergic asthma, atopic eczema, atopic dermatitis, and food allergies. Immunodeficiencies include, but are not limited to, severe combined immunodeficiency (SCID), hypereosinophilic syndrome, chronic granulomatous disease, leukocyte adhesion deficiency I and II, hyper IgE syndrome, Chediak-Higashi, neutrophilia, neutropenia, agranulocytosis, agammaglobulinemia, hyper IgM syndrome, DiGeorge / velo-cardio-facial syndrome, and interferon-gamma-TH1 pathway deficiency. Autoimmune and immunoregulatory disorders include, but are not limited to, rheumatoid arthritis, diabetes, systemic lupus erythematosus, Graves' disease, Graves' ophthalmopathy, Crohn's disease, multiple sclerosis, psoriasis, systemic sclerosis, goiter and lymphomatous goiter (Hashimoto's thyroiditis, struma lymphomatosa), alopecia areata, autoimmune myocarditis, lichen sclerosus, autoimmune uveitis, Addison's disease, atrophic gastritis, myasthenia gravis, idiopathic thrombocytopenic purpura, hemolytic anemia, primary biliary cirrhosis, Wegener's granulomatosis, polyarteritis nodosa, and inflammatory bowel disease, allograft rejection, and tissue destruction due to allergic reactions to infectious bacteria or environmental antigens.
[0097] Proliferative diseases and disorders that can be evaluated by the methods of the present disclosure include, but are not limited to, neonatal hemangioma, secondary progressive multiple sclerosis, chronic progressive myelodysplastic disease, neurofibromatosis, ganglioneuroma, keloid formation, Paget's disease of bone, fibrocystic disease (e.g., of the breast or uterus), sarcoidosis, Peyronie's and Dupuytren's fibrosis, cirrhosis, atherosclerosis, and vascular restenosis.
[0098] Malignant diseases and disorders that can be evaluated by the methods of the present disclosure include both hematological malignancies and solid tumors.
[0099] Hematological malignancies are particularly suitable for the methods of the present disclosure when the sample is a blood sample because such malignancies involve changes in blood - infective cells. Such malignancies include non - Hodgkin lymphoma, Hodgkin lymphoma, non - B - cell lymphoma, and other lymphomas, acute or chronic leukemia, polycythemia, thrombocythemia, multiple myeloma, myelodysplastic syndrome, myeloproliferative disorder, encephalomyelitis, abnormal immune lymphocyte proliferation, and plasma cell disorders.
[0100] Plasma cell diseases that can be evaluated by the methods of the present disclosure include multiple myeloma, amyloidosis, and Waldenström macroglobulinemia.
[0101] Examples of solid tumors include, but are not limited to, colon cancer, breast cancer, lung cancer, prostate cancer, brain tumor, central nervous system tumor, bladder tumor, melanoma, liver cancer, osteosarcoma, and other bone cancers, testicular and ovarian carcinomas, head and neck tumors, and cervical neoplasms.
[0102] Genetic disorders can also be detected by the processes of the present disclosure. This can be performed by prenatal or postnatal screening for chromosomal and genetic abnormalities, or genetic diseases. Examples of detectable genetic diseases include 21 - hydroxylase deficiency, cystic fibrosis, fragile X syndrome, Turner syndrome, Duchenne muscular dystrophy, Down syndrome or other trisomies, heart disease, single - gene diseases, HLA typing, phenylketonuria, sickle - cell anemia, Tay - Sachs disease, thalassemia, Klinefelter syndrome, Huntington's disease, autoimmune diseases, lipidosis, obesity defect, hemophilia, inborn errors of metabolism, and diabetes.
[0103] The methods described herein can be used to diagnose pathogen infections, such as infections by intracellular bacteria and viruses, by determining the presence and / or amount of markers for each of the bacteria or viruses in the sample.
[0104] A variety of infectious diseases can be detected by the processes of the present disclosure. Infectious diseases can be caused by infectious agents of bacteria, viruses, parasites, and fungi. The resistance of various infectious agents to drugs can also be determined using the present disclosure.
[0105] Bacterial infectious agents that can be detected by the present disclosure include Escherichia - coli, Salmonella, Shigella, Klebsiella, Pseudomonas, Listeria - monocytogenes, Mycobacterium - tuberculosis, Mycobacterium - avium - intracellulare, Yersinia, Francisella, Pasteurella, Brucella, Clostridium, Bordetella - pertussis, Bacteroides, Staphylococcus - aureus, Streptococcus - pneumoniae, B - Hemolytic strep., Corynebacteria, Legionella, Mycoplasma, Ureaplasma, Chlamydia, Neisseria - gonorrhoeae, Meningococcus, Haemophilus - influenzae, Enterococcus - faecalis, Proteus - vulgaris, Proteus - mirabilis, Helicobacter - pylori, Treponema - pallidum, Borrelia - burgdorferi, Borrelia - recurrentis, Rickettsia pathogenic microorganisms, Nocardia, and Actinomycetes.
[0106] Fungal infectious agents that can be detected by the present disclosure include Cryptococcus - neoformans, Blastomyces - dermatitidis, Histoplasma - capsulatum, Coccidioides - immitis, Paracoccidioides - brasiliensis, Candida - albicans, Aspergillus - fumigatus, Phycomycetes (Rhizopus), Sporothrix - schenckii, Chromomycosis, and Maduromycosis.
[0107] Viral infectious agents detected by the present disclosure include human immunodeficiency virus, human T-cell lymphocytotrophic virus, hepatitis viruses (e.g., hepatitis B virus and hepatitis C virus), Epstein-Barr virus, cytomegalovirus, human papillomavirus, orthomyxovirus, paramyxovirus, adenovirus, coronavirus, rhabdovirus, poliovirus, togavirus, bunyavirus, arenavirus, rubella virus, and reovirus.
[0108] Parasitic agents that can be detected by the present disclosure include Plasmodium falciparum, Plasmodium malariae, Plasmodium vivax, Plasmodium ovale, Onchoverva volvulus, Leishmania, Trypanosoma species, Schistosoma species, Entamoeba histolytica, Cryptosporidum, Giardia species, Trichimonas species, Balatidium Coli, Wuchereria bancrofti, Toxoplasma species, Enterobius vermicularis, Ascaris lumbricoides, Trichuris trichiura, Dracunculus medinesis, trematodes, Diphyllobothrium latum, Taenia species, Pneumocystis carinii, and Necator americanis.
[0109] The present disclosure is also useful for detecting drug resistance to infectious agents. For example, vancomycin-resistant Enterococcus faecium, methicillin-resistant Staphylococcus aureus, penicillin-resistant Streptococcus pneumoniae, multi-drug resistant Mycobacterium tuberculosis, and AZT-resistant human immunodeficiency virus can all be identified by the present disclosure.
[0110] Thus, the target molecules detected using the compositions and methods of the present disclosure can be either patient markers (such as cancer markers) or markers of foreign infection, such as bacterial or viral markers.
[0111] The compositions and methods of the present disclosure can be used to identify and / or quantify target molecules, the abundance of which indicates a biological state or disease condition, such as a blood marker that is upregulated or downregulated as a result of a disease state, e.g., cancer.
[0112] In some embodiments, the methods and compositions of the present disclosure can be used for cytokine expression. The low sensitivity of the methods described herein is useful for the early detection of cytokines as biomarkers for disease states, diagnosis, or prognosis, such as cancer, and for the identification of subclinical states.
[0113] The various samples from which the target polynucleotide is derived can include multiple samples from the same individual, samples from different individuals, or combinations thereof. In some embodiments, the sample includes multiple polynucleotides from one individual. In some embodiments, the sample includes multiple polynucleotides from two or more individuals. An individual is an organism or a part thereof from which the target polynucleotide is derived, and non-limiting examples thereof include animals, fungi, protists, monera, viruses, mitochondria, and chloroplasts. The polynucleotides of the sample can be isolated from a subject such as a cell sample, a tissue sample, or an organ sample derived therefrom, including, for example, a cultured cell line, a biopsy, a blood sample, or a fluid sample containing cells. The subject is an animal including but not limited to animals such as cows, pigs, mice, rats, chickens, cats, dogs, etc., and is usually a mammal such as a human. The sample can also be artificially derived, such as by chemical synthesis. In some embodiments, the sample contains DNA. In some embodiments, the sample contains genomic DNA. In some embodiments, the sample contains mitochondrial DNA, chloroplast DNA, plasmid DNA, bacterial artificial chromosomes, yeast artificial chromosomes, oligonucleotide tags, or combinations thereof. In some embodiments, the sample contains DNA generated by a primer extension reaction using an appropriate combination of primers and DNA polymerase, including but not limited to polymerase chain reaction (PCR), reverse transcription, and combinations thereof. When the template for the primer extension reaction is RNA, the product of reverse transcription is called complementary DNA (cDNA). Primers useful for the primer extension reaction can include sequences specific for one or more targets, random sequences, partially random sequences, and combinations thereof. Reaction conditions suitable for the primer extension reaction are known in the art. Generally, the polynucleotides of the sample contain polynucleotides in the sample, which may or may not include the target polynucleotide.
[0114] In some embodiments, the nucleic acid template molecule (e.g., DNA or RNA) can be isolated from a biological sample containing various other components such as proteins, lipids, and non-template nucleic acids. The nucleic acid template molecule can be obtained from any cellular material and can be obtained from animals, plants, bacteria, fungi, or other cellular organisms. The biological sample for use in the present disclosure includes viral particles or preparations. The nucleic acid template molecule can be obtained directly from an organism or from a biological sample obtained from an organism, such as blood, urine, cerebrospinal fluid, semen, saliva, sputum, feces, and tissue. Any tissue or body fluid specimen may be used as a source of nucleic acid for the use of the present disclosure. The nucleic acid template molecule can also be isolated from cultured cells such as primary cell cultures or cell lines. The cell or tissue from which the template nucleic acid is obtained can be infected with a virus or other intracellular pathogen. The sample can also be a biological specimen, a cDNA library, viral DNA, or total RNA extracted from genomic DNA. The sample can also be DNA isolated from a source without cellular structure, such as DNA amplified / isolated from a freezer.
[0115] Methods for extracting and purifying nucleic acids are well known in the art. For example, nucleic acids can be purified by organic extraction with phenol, phenol / chloroform / isoamyl alcohol, or similar formulations including TRIzol and TriReagent. Other non-limiting examples of extraction techniques include: (1) organic extraction with ethanol precipitation, with or without the use of an automated nucleic acid extractor, such as the Model 341 DNA Extractor available from Applied Biosystems (Foster City, Calif.), using an organic reagent such as phenol / chloroform (Ausubel et al., 1993); (2) solid-phase adsorption methods (U.S. Patent No. 5,234,809; Walsh et al., 1991); and (3) salt-induced nucleic acid precipitation methods, such as precipitation methods typically referred to as "salting out" methods (Miller et al., (1988)). Another example of nucleic acid isolation and / or purification involves the use of magnetic particles, where the nucleic acid binds specifically or non-specifically to the magnetic particles, and then the beads can be isolated, washed, and the nucleic acid eluted from the beads using a magnet (see, e.g., U.S. Patent No. 5,705,628). In some embodiments, the isolation methods described above may begin with an enzymatic digestion step, such as digestion with proteinase K or other proteases, to remove unwanted proteins from the sample. See, e.g., U.S. Patent No. 7,001,724. If desired, an RNase inhibitor can be added to the lysis buffer. For certain cell or sample types, it may be desirable to add a protein denaturation / digestion step to the protocol. The purification methods can be aimed at isolating DNA, RNA, or both. If both DNA and RNA are isolated together during or after the extraction procedure, additional steps can be utilized to purify one or both separately from the other. For example, purification based on size, sequence, or other physical or chemical properties can also be used to generate sub-fragments of the extracted nucleic acid. In addition to the initial nucleic acid isolation step, nucleic acid purification can be performed after the steps in the methods of the present disclosure, for example, to remove excess or unwanted reagents, reactants, or products.
[0116] The nucleic acid template molecule can be obtained as described in U.S. Patent Application Publication No. 2002 / 0190663 A1, published on October 9, 2003. Usually, nucleic acids can be extracted from biological samples by various techniques such as those described by Maniatis, et al., Molecular Cloning: A Laboratory Manual, Cold Spring Harbor, N.Y., pp. 280-281 (1982). In some cases, the nucleic acid can be first extracted from the biological sample and then cross-linked in vitro. In some cases, native associated proteins (e.g., histones) can be further removed from the nucleic acid.
[0117] In other embodiments, the present disclosure can be readily applied to high molecular weight double-stranded DNA, including, for example, DNA isolated from tissues, cell cultures, body fluids, animal tissues, plants, bacteria, fungi, viruses, etc.
[0118] Hi-C method including size selection Provided herein is a method that includes obtaining a stabilized biological sample comprising a nucleic acid molecule complexed to at least one nucleic acid-binding protein, contacting the stabilized biological sample with a DNase to cleave the nucleic acid molecule into a plurality of segments, attaching a first segment and a second segment of the plurality of segments at one junction, and size selecting the plurality of segments to obtain a plurality of selected segments. Optionally, the plurality of selected segments are from about 145 to about 600 bp. Optionally, the plurality of selected segments are from about 100 to about 2500 bp. Optionally, the plurality of selected segments are from about 100 to about 600 bp. Optionally, the plurality of selected segments are from about 600 to about 2500 bp. Optionally, the plurality of selected segments are between about 100 bp and about 600 bp, between about 100 bp and about 700 bp, between about 100 bp and about 800 bp, between about 100 bp and about 900 bp, between about 100 bp and about 1000 bp, between about 100 bp and about 1100 bp, between about 100 bp and about 1200 bp, between about 100 bp and about 1300 bp, between about 100 bp and about 1400 bp, between about 100 bp and about 1500 bp, between about 100 bp and about 1600 bp, between about 100 bp and about 1700 bp, between about 100 bp and about 1800 bp, between about 100 bp and about 1900 bp, between about 100 bp and about 2000 bp, between about 100 bp and about 2100 bp, between about 100 bp and about 2200 bp, between about 100 bp and about 2300 bp, between about 100 bp and about 2400 bp, or between about 100 bp and about 2500 bp.
[0119] In another aspect of the method involving a size selection step provided herein, the method further includes, prior to the size selection step, preparing a sequencing library from a plurality of segments. In some embodiments, the method further includes subjecting the sequencing library to size selection to obtain a size-selected library. Optionally, the size-selected library is in a size range from about 350 bp to about 1000 bp. Optionally, the size-selected library is in a size range from about 100 bp to 2500 bp, for example, between about 100 bp and about 350 bp, between about 350 bp and about 500 bp, between about 500 bp and about 1000 bp, between about 1000 bp and about 1500 bp to about 2000 bp, between about 2000 bp and about 2500 bp, between about 350 bp and about 1000 bp, between about 350 bp and about 1500 bp, between about 350 bp and about 2000 bp, between about 350 bp and about 2500 bp, between about 500 bp and about 1500 bp, between about 500 bp and about 2000 bp, between about 500 bp and about 2500 bp, between about 1000 bp and about 1500 bp, between about 1000 bp and about 2000 bp, between about 1000 bp and about 2500 bp, between about 1500 bp and about 2000 bp, between about 1500 bp and about 2500 bp, or between about 100 bp and about 700 bp.
[0120] The size selection utilized in the method involving a size selection step provided herein is performed by gel electrophoresis, capillary electrophoresis, size selection beads, gel filtration columns, other suitable methods, or combinations thereof.
[0121] In another aspect, the method with the size selection step provided herein further includes analyzing a plurality of selected segments to obtain one QC value. In some cases, the QC value is selected from chromatin digestion efficiency (CDE) and chromatin digestion index (CDI). CDE is calculated as the proportion of segments having a desired length. For example, in some cases, CDE is calculated as the proportion of segments having a size between 100 and 2500 bp before size selection. In some cases, if the CDE value is at least 65%, the sample is selected for further analysis. In some cases, the sample is selected for further analysis when the CDE value is at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, or at least about 95%. CDI is calculated as the ratio of the number of segments of mononucleosome size to the number of segments of dinucleosome size before size selection. For example, CDI can be calculated as the logarithm of the ratio of fragments having a size of 600 to 2500 bp to fragments having a size of 100 to 600 bp. In some cases, if the CDI value exceeds -1.5 and is less than 1, the sample is selected for further analysis.In some cases, the sample is selected for further analysis when the CDI value is greater than about -2 and less than about 1.5, greater than about -1.9 and less than about 1.5, greater than about -1.8 and less than about 1.5, greater than about -1.7 and less than about 1.5, greater than about -1.6 and less than about 1.5, greater than about -1.5 and less than about 1.5, greater than about -1.4 and less than about 1.5, greater than about -1.3 and less than about 1.5, greater than about -1.2 and less than about 1.5, greater than about -1.1 and less than about 1.5, greater than about -2 and less than about 1.5, greater than about -1 and less than about 1.5, greater than about -0.9 and less than about 1.5, greater than about -0.8 and less than about 1.5, greater than about -0.7 and less than about 1.5, greater than about -0.6 and less than about 1.5, greater than about -0.5 and less than about 1.5, greater than about -2 and less than about 1.4, greater than about -2 and less than about 1.3, greater than about -2 and less than about 1.2, greater than about -2 and less than about 1.1, greater than about -2 and less than about 1, greater than about -2 and less than about 0.9, greater than about -2 and less than about 0.8, greater than about -2 and less than about 0.7, greater than about -2 and less than about 0.6, or greater than about -2 and less than about 0.5.
[0122] In another aspect, the stabilized biological sample used in the methods with a size selection step herein includes a biological sample treated with a stabilizer. In some cases, the stabilized biological sample includes a stabilized cell lysate. Alternatively, the stabilized biological sample includes stabilized intact cells. Alternatively, the stabilized biological sample includes stabilized intact nuclei. In some cases, the step of contacting the stabilized intact cell or intact nucleus sample with DNase is performed prior to lysis of the intact cell or intact nucleus. In some cases, the cells and / or nuclei are lysed prior to attaching the first segment and the second segment of the plurality of segments at the junction.
[0123] In another aspect, the methods herein involving a size selection step are performed on small samples containing only a small number of cells or a small amount of nucleic acid. For example, in some cases, the stabilized biological sample contains less than 3,000,000 cells. In some cases, the stabilized biological sample contains less than 2,000,000 cells. In some cases, the stabilized biological sample contains less than 1,000,000 cells. In some cases, the stabilized biological sample contains less than 500,000 cells. In some cases, the stabilized biological sample contains less than 400,000 cells. In some cases, the stabilized biological sample contains less than 300,000 cells. In some cases, the stabilized biological sample contains less than 200,000 cells. In some cases, the stabilized biological sample contains less than 100,000 cells. In some cases, the stabilized biological sample contains less than 10 μg of DNA. In some cases, the stabilized biological sample contains less than 9 μg of DNA. In some cases, the stabilized biological sample contains less than 8 μg of DNA. In some cases, the stabilized biological sample contains less than 7 μg of DNA. In some cases, the stabilized biological sample contains less than 6 μg of DNA. In some cases, the stabilized biological sample contains less than 5 μg of DNA. In some cases, the stabilized biological sample contains less than 4 μg of DNA. In some cases, the stabilized biological sample contains less than 3 μg of DNA. In some cases, the stabilized biological sample contains less than 2 μg of DNA. In some cases, the stabilized biological sample contains less than 1 μg of DNA. In some cases, the stabilized biological sample contains less than 0.5 μg of DNA.
[0124] In another aspect, the methods involving the size selection step herein can be performed on individual cells or single cells. For example, the methods herein may be performed on cells distributed into individual compartments. Exemplary compartments include, but are not limited to, wells, droplets in an emulsion, or surface locations (e.g., array spots, beads, etc.), and those compartments contain distinct patches of differentially arrayed linker molecules as otherwise described herein. Additional compartments are also contemplated and are consistent with the methods, compositions, and systems disclosed herein.
[0125] In an additional aspect, the stabilized biological sample used in the methods involving the size selection step herein is treated with a nuclease, such as DNase, to create DNA fragments. In some cases, the DNase is non-sequence specific. In some cases, the DNase is active against both single-stranded DNA and double-stranded DNA. In some cases, the DNase is specific for double-stranded DNA. In some cases, the DNase preferentially cleaves double-stranded DNA. In some cases, the DNase is specific for single-stranded DNA. In some cases, the DNase preferentially cleaves single-stranded DNA. In some cases, the DNase is DNase I. In some cases, the DNase is DNase II. In some cases, the DNase is selected from one or more of DNase I and DNase II. In some cases, the DNase is micrococcal nuclease. In some cases, the DNase is selected from one or more of DNase I, DNase II, and micrococcal nuclease. In some cases, the DNase can be bound or fused to an immunoglobulin binding protein or a fragment thereof. The immunoglobulin binding protein can be, for example, protein A, protein G, protein A / G, or protein L. In some embodiments, the DNase can be linked to a fusion protein containing two or more immunoglobulin binding proteins and / or fragments thereof. Other suitable nucleases are also within the scope of this disclosure.
[0126] In an additional aspect, a stabilized biological sample as provided herein for use in a method involving a size selection step is treated with one or more crosslinking agents. In some cases, the crosslinking agent is a chemical fixative. In some cases, the chemical fixative includes formaldehyde having a spacer arm length of about 2.3 - 2.7 angstroms (A). In some cases, the chemical fixative includes a crosslinking agent having a long spacer arm length. For example, the crosslinking agent can have a spacer length of at least about 3A, 4A, 5A, 6A, 7A, 8A, 9A, 10A, 11A, 12A, 13A, 14A, 15A, 16A, 17A, 18A, 19A, or 20A. The chemical fixative can include ethylene glycol bis(succinimidyl succinate) (EGS) and has a spacer arm with a length of about 16.1A. The chemical fixative can include disuccinimidyl glutarate (DSG) and has a spacer arm with a length of about 7.7A. In some cases, the chemical fixative includes formaldehyde and EGS, formaldehyde and DSG, or formaldehyde, EGS, and DSG. When multiple chemical fixatives are utilized, each chemical fixative can be used sequentially, or in other cases, some or all of the multiple chemical fixatives are applied to the sample simultaneously. The use of a crosslinking agent having a long spacer arm can increase the fragments of read pairs having a large (e.g., >1 kb) read pair separation distance. For example, FIG. 7 shows a comparison of the resulting libraries (both DNase digestion and MNase digestion) crosslinked with formaldehyde alone and crosslinked with formaldehyde plus DSG or EGS. DSG has NHS ester reactive groups at both ends and can be reactive with amino groups (e.g., primary amines). DSG is membrane permeable, enabling intracellular crosslinking. DSG can increase crosslinking efficiency in some applications compared to disuccinimidyl suberate (DSS). EGS has NHS ester reactive groups at both ends and can be reactive with amino groups (e.g., primary amines). EGS is membrane permeable, enabling intracellular crosslinking.EGS crosslinking can be reversed, for example, by treating with hydroxylamine at pH 8.5 for 3 to 6 hours. In one example, lactate dehydrogenase retained 60% of its activity after reversible crosslinking with EGS. In some cases, the chemical fixative includes psoralen. In some cases, the crosslinking agent is ultraviolet light. In some embodiments, the stabilized biological sample is a crosslinked paraffin-embedded tissue sample.
[0127] In a further aspect, the methods provided herein involving the size selection step include contacting a plurality of selected segments with an antibody.
[0128] In an additional aspect, the method with the size selection step provided herein includes, at one junction, attaching a first segment and a second segment of a plurality of segments. In some cases, the attaching step includes filling in sticky ends and ligating blunt ends using biotin-tagged nucleotides. In some cases, the attaching step includes contacting at least the first segment and the second segment with a cross-linking oligonucleotide. In some cases, the attaching step includes contacting at least the first segment and the second segment with a barcode. In some embodiments, the cross-linking oligonucleotide herein can be from at least about 5 nucleotides to about 50 nucleotides in length. In some embodiments, the cross-linking oligonucleotide herein can be from at least about 15 nucleotides to about 18 nucleotides in length. In some embodiments, the cross-linking oligonucleotide can be about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, about 20, about 25, about 30, about 35, about 40, about 45, or about 50 nucleotides in length. In some embodiments, the cross-linking oligonucleotide herein may include a barcode. In some embodiments, the cross-linking oligonucleotide can include a plurality of barcodes. In some embodiments, the cross-linking oligonucleotide includes a plurality of cross-linking oligonucleotides connected to each other. In some embodiments, the cross-linking oligonucleotide may be linked or connected to an immunoglobulin-binding protein such as protein A, protein G, protein A / G, or protein L or a fragment thereof. In some cases, the linked cross-linking oligonucleotide can be delivered to the position where an antibody binds in the nucleic acid of the sample.
[0129] The splitting and pooling approach can be utilized to generate cross-linked oligonucleotides with unique barcodes. A population of samples can be divided into multiple groups, and the cross-linked oligonucleotides can be attached to the samples such that the barcode of the cross-linked oligonucleotides is different between groups but the same within a group. The groups of samples can be pooled together again, and this process can be repeated multiple times. Repeating this process results in each sample in the population having a unique set of cross-linked oligonucleotide barcodes, which may enable the analysis of a single sample (e.g., a single cell, a single nucleus, a single chromosome). In one illustrative example, a sample of cross-linked and digested nuclei attached to a solid support of beads is separated across 8 tubes, each containing one of 8 unique members of a first adapter group (first iteration) of double-stranded DNA (dsDNA) adapters that are to be ligated. Each of the 8 adapters may have the same 5’ overhang sequence for ligation to the nucleic acid termini of the cross-linked chromatin assemblies in the nuclei, or otherwise have a unique dsDNA sequence. After the first adapter group is ligated, the nuclei are pooled together again and can be washed to remove the ligation reaction components. The scheme of distributing, ligating, and pooling can be repeated 2 more times (2 iterations). After ligation of the members of each adapter group, the cross-linked chromatin assemblies can be attached to a series of multiple barcodes. In some cases, successive ligation of multiple members of multiple adapter groups (iterations) results in a combination of barcodes. The number of possible combinations of barcodes depends on the number of groups per iteration and the total number of barcode oligonucleotides used. For example, 3 iterations with 8 members can each have 83 possible combinations. In some cases, the combination of barcodes is unique. In some cases, there is redundancy in the combination of barcodes.The total number of barcode combinations can be adjusted by increasing or decreasing the number of groups receiving unique barcodes and / or by increasing or decreasing the number of iterations. If more than one adapter group is used, schemes for distributing, attaching, and pooling can be used for iterative adapter attachment. Optionally, the distribution, attachment, and pooling scheme can be iterated an additional at least 3, 4, 5, 6, 7, 8, 9, or 10 times. Optionally, members of the last adapter group include sequences for subsequent enrichment of adapter-attached DNA, for example, during preparation of a sequencing library by PCR amplification.
[0130] In a further aspect, methods involving a size selection step herein do not include a shearing step (e.g., the nucleic acid is not sheared).
[0131] In a further aspect of methods involving a size selection step herein, the method includes obtaining at least some sequences on both sides of the junction to generate a first read pair. For example, the method can include obtaining sequences of at least about 50 bp, at least about 100 bp, at least about 150 bp, at least about 200 bp, at least about 250 bp, or at least about 300 bp on each side of the junction to generate a first read pair.
[0132] In an additional aspect of methods involving a size selection step herein, the method further includes mapping the first read pair to a set of contigs and determining a path across the set of contigs that represents the order and / or orientation to the genome.
[0133] In a further aspect of methods involving a size selection step herein, the method further includes mapping the first read pair to a set of contigs and determining the presence of structural variants or a decrease in heterozygosity in a stabilized biological sample from the set of contigs.
[0134] In a further aspect of the method involving a size selection step herein, the method includes mapping a first read pair to a set of contiguous contigs and assigning phases to variants in the set of contigs.
[0135] In yet another aspect of the method involving a size selection step herein, the method further includes mapping a first read pair to a set of contigs, determining the presence of variants in the set of contigs from the set of contigs, and performing one or more steps selected from the group consisting of: (1) confirming a disease stage, prognosis, or treatment regimen for a stabilized biological sample; (2) selecting a drug based on the presence of the variant; or (3) confirming the efficacy of a drug against a stabilized biological sample.
[0136] Hi-C method including QC calculation Additionally, a method is provided herein, the method comprising obtaining a stabilized biological sample comprising a nucleic acid molecule complexed to at least one nucleic acid binding protein, contacting the stabilized biological sample with DNase to cleave the nucleic acid molecule into a plurality of segments, attaching a first segment and a second segment of the plurality of segments at one junction, and analyzing the plurality of segments to determine a QC value. Optionally, the QC value is selected from a chromatin digestion efficiency (CDE) and a chromatin digestion index (CDI). The CDE is calculated as the proportion of segments having a desired length. For example, optionally, the CDE is calculated as the proportion of segments having a size between 100 and 2500 bp prior to size selection. Optionally, when the CDE value is at least 65%, the sample is selected for further analysis. In some embodiments, the sample is selected for further analysis when the CDE value is at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, or at least about 95%. The CDI is calculated as the ratio of the number of segments of mononucleosome size to the number of segments of dinucleosome size prior to size selection. For example, the CDI can be calculated as the logarithm of the ratio of fragments having a size of 600 to 2500 bp to fragments having a size of 100 to 600 bp. Optionally, the sample is selected for further analysis when the CDI value is greater than -1.5 and less than 1.In some cases, the sample is selected for further analysis when the CDI value is greater than about -2 and less than about 1.5, greater than about -1.9 and less than about 1.5, greater than about -1.8 and less than about 1.5, greater than about -1.7 and less than about 1.5, greater than about -1.6 and less than about 1.5, greater than about -1.5 and less than about 1.5, greater than about -1.4 and less than about 1.5, greater than about -1.3 and less than about 1.5, greater than about -1.2 and less than about 1.5, greater than about -1.1 and less than about 1.5, greater than about -2 and less than about 1.5, greater than about -1 and less than about 1.5, greater than about -0.9 and less than about 1.5, greater than about -0.8 and less than about 1.5, greater than about -0.7 and less than about 1.5, greater than about -0.6 and less than about 1.5, greater than about -0.5 and less than about 1.5, greater than about -2 and less than about 1.4, greater than about -2 and less than about 1.3, greater than about -2 and less than about 1.2, greater than about -2 and less than about 1.1, greater than about -2 and less than about 1, greater than about -2 and less than about 0.9, greater than about -2 and less than about 0.8, greater than about -2 and less than about 0.7, greater than about -2 and less than about 0.6, or greater than about -2 and less than about 0.5.
[0137] In other embodiments, a method involving the QC quantification step herein may include a step of size selecting a plurality of segments to obtain a plurality of selected segments. Optionally, the plurality of selected segments are from about 145 to about 600 bp. Optionally, the plurality of selected segments are from about 100 to about 2500 bp. Optionally, the plurality of selected segments are from about 100 to about 600 bp. Optionally, the plurality of selected segments are from about 600 to about 2500 bp. Optionally, optionally, the plurality of selected segments are between about 100 and about 600 bp, between about 100 bp and about 700 bp, between about 100 bp and about 800 bp, between about 100 bp and about 900 bp, between about 100 bp and about 1000 bp, between about 100 bp and about 1100 bp, between about 100 bp and about 1200 bp, between about 100 bp and about 1300 bp, between about 100 bp and about 1400 bp, between about 100 bp and about 1500 bp, between about 100 bp and about 1600 bp, between about 100 bp and about 1700 bp, between about 100 bp and about 1800 bp, between about 100 bp and about 1900 bp, between about 100 bp and about 2000 bp, between about 100 bp and about 2100 bp, between about 100 bp and about 2200 bp, between about 100 bp and about 2300 bp, between about 100 bp and about 2400 bp, or between about 100 bp and about 2500 bp.
[0138] In another aspect of the methods provided herein that involve QC quantification steps, the method may further include preparing a sequencing library from a plurality of segments prior to the size selection step. In some embodiments, the method further includes subjecting the sequencing library to size selection to obtain a size-selected library. Optionally, the size-selected library is sized between about 350 bp and about 1000 bp. Optionally, the size-selected library is sized between about 100 bp and about 2500 bp, for example, between about 100 bp and about 350 bp, between about 350 bp and about 500 bp, between about 500 bp and about 1000 bp, between about 1000 bp and about 1500 bp to about 2000 bp, between about 2000 bp and about 2500 bp, between about 350 bp and about 1000 bp, between about 350 bp and about 1500 bp, between about 350 bp and about 2000 bp, between about 350 bp and about 2500 bp, between about 500 bp and about 1500 bp, between about 500 bp and about 2000 bp, between about 500 bp and about 3500 bp, between about 1000 bp and about 1500 bp, between about 1000 bp and about 2000 bp, between about 1000 bp and about 2500 bp, between about 1500 bp and about 2000 bp, between about 1500 bp and about 2500 bp, or between about 100 bp and about 700 bp.
[0139] In the methods involving QC quantification steps herein, the size selection may be performed using gel electrophoresis, capillary electrophoresis, size selection beads, gel filtration columns, or combinations thereof. Other suitable size selection methods are also within the scope of this disclosure.
[0140] In other aspects, the stabilized biological sample used in connection with the QC quantification step herein includes biological material treated with a stabilizer. Optionally, the stabilized biological sample includes a stabilized cell lysate. Alternatively, the stabilized biological sample includes stabilized intact cells. Alternatively, the stabilized biological sample includes stabilized intact nuclei. Optionally, the step of contacting the stabilized intact cell or intact nucleus sample with DNase is performed prior to lysis of the intact cell or intact nucleus. Optionally, the cell and / or nucleus is lysed prior to attaching a first segment and a second segment of a plurality of segments at a junction.
[0141] In other aspects, the methods herein involving QC quantification steps are performed on small samples containing only a small number of cells or a small amount of nucleic acid. In some cases, the stabilized biological sample contains less than 3,000,000 cells. In some cases, the stabilized biological sample contains less than 2,000,000 cells. In some cases, the stabilized biological sample contains less than 1,000,000 cells. In some cases, the stabilized biological sample contains less than 500,000 cells. In some cases, the stabilized biological sample contains less than 400,000 cells. In some cases, the stabilized biological sample contains less than 300,000 cells. In some cases, the stabilized biological sample contains less than 200,000 cells. In some cases, the stabilized biological sample contains less than 100,000 cells. In some cases, the stabilized biological sample contains less than 10 μg of DNA. In some cases, the stabilized biological sample contains less than 9 μg of DNA. In some cases, the stabilized biological sample contains less than 8 μg of DNA. In some cases, the stabilized biological sample contains less than 7 μg of DNA. In some cases, the stabilized biological sample contains less than 6 μg of DNA. In some cases, the stabilized biological sample contains less than 5 μg of DNA. In some cases, the stabilized biological sample contains less than 4 μg of DNA. In some cases, the stabilized biological sample contains less than 3 μg of DNA. In some cases, the stabilized biological sample contains less than 2 μg of DNA. In some cases, the stabilized biological sample contains less than 1 μg of DNA. In some cases, the stabilized biological sample contains less than 0.5 μg of DNA.
[0142] In other aspects, the methods involving the QC quantification steps herein can be performed on individual cells or single cells. For example, the methods herein may be performed on cells that have been distributed into individual compartments. Exemplary compartments include, but are not limited to, wells, droplets in an emulsion, or surface locations (e.g., array spots, beads, etc.), and those compartments contain distinct patches of differentially arrayed linker molecules as described elsewhere herein. Additional compartments are contemplated and are consistent with the methods, compositions, and systems disclosed herein.
[0143] In one aspect, the stabilized biological sample used herein in involving the QC quantification step is treated with a nuclease such as DNase to create DNA fragments. In some cases, the DNase is non-sequence specific. In some cases, the DNase is active against both single-stranded DNA and double-stranded DNA. In some cases, the DNase is specific for double-stranded DNA. In some cases, the DNase preferentially cleaves double-stranded DNA. In some cases, the DNase is specific for single-stranded DNA. In some cases, the DNase preferentially cleaves single-stranded DNA. In some cases, the DNase is DNase I. In some cases, the DNase is DNase II. In some cases, the DNase is selected from one or more of DNase I and DNase II. In some cases, the DNase is micrococcal nuclease. In some cases, the DNase is selected from one or more of DNase I, DNase II, and micrococcal nuclease. In some cases, the DNase is a cross-linked oligonucleotide that can be linked or fused to an immunoglobulin-binding protein such as protein A, protein G, protein A / G, or protein L, or fragments thereof. Other suitable nucleases are also within the scope of this disclosure.
[0144] In an additional aspect, in a method involving a QC quantification step herein, the stabilized biological sample is treated with a crosslinking agent. In some cases, the crosslinking agent is a chemical fixative. In some cases, the chemical fixative includes formaldehyde having a spacer arm length of about 2.3 to 2.7 angstroms (A). In some cases, the chemical fixative includes a crosslinking agent having a long spacer arm length. For example, the crosslinking agent can have a spacer length of at least about 3A, 4A, 5A, 6A, 7A, 8A, 9A, 10A, 11A, 12A, 13A, 14A, 15A, 16A, 17A, 18A, 19A, or 20A. The chemical fixative includes ethylene glycol bis(succinimidyl succinate) (EGS) and has a spacer arm with a length of about 16.1A. The chemical fixative can include disuccinimidyl glutarate (DSG) and has a spacer arm with a length of about 7.7A. In some cases, the chemical fixative includes formaldehyde and EGS, formaldehyde and DSG, or formaldehyde, EGS, and DSG. In some cases where multiple chemical fixatives are utilized, each chemical fixative is used sequentially. In other cases, some or all of the multiple chemical fixatives are applied to the sample simultaneously. The use of a crosslinking agent with a long spacer arm can increase the fragments of read pairs having a large (e.g., >1 kb) read pair separation distance. For example, FIG. 7 shows a comparison for the resulting library (both DNase digestion and MNase digestion) between those crosslinked with formaldehyde alone and those crosslinked with formaldehyde plus DSG or EGS. DSG has NHS ester reactive groups at both ends and can be reactive towards amino groups (e.g., primary amines). DSG is membrane permeable and enables intracellular crosslinking. DSG can increase crosslinking efficiency in some applications compared to disuccinimidyl suberate (DSS). EGS has NHS ester reactive groups at both ends and can be reactive towards amino groups (e.g., primary amines). EGS is membrane permeable and enables intracellular crosslinking.EGS cross-linking can be reversed, for example, by treating with hydroxylamine at pH 8.5 for 3 to 6 hours. In one example, lactate dehydrogenase retained 60% of its activity after reversible cross-linking with EGS. In some cases, the chemical fixative includes psoralen. In some cases, the cross-linking agent is ultraviolet light. In some embodiments, the stabilized biological sample is a cross-linked paraffin-embedded tissue sample.
[0145] In further embodiments, the methods involving the QC quantification steps provided herein include contacting a plurality of selected segments with an antibody.
[0146] In an additional aspect, a method involving the QC quantification step herein includes attaching a first segment and a second segment of a plurality of segments at one junction. Optionally, the attaching step includes filling in sticky ends and ligating blunt ends using biotin-tagged nucleotides. Optionally, the attaching step includes contacting at least the first segment and the second segment with a cross-linking oligonucleotide. Optionally, the attaching step includes contacting at least the first segment and the second segment with a barcode. In some embodiments, the cross-linking oligonucleotide herein can be from at least about 5 nucleotides to about 50 nucleotides in length. In some embodiments, the cross-linking oligonucleotide herein can be from at least about 15 nucleotides to about 18 nucleotides in length. In some embodiments, the cross-linking oligonucleotide is about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, about 20, about 25, about 30, about 35, about 40, about 45, or about 50 nucleotides in length. In some embodiments, the cross-linking oligonucleotide herein can include a barcode. In some embodiments, the cross-linking oligonucleotide can include a plurality of barcodes. In some embodiments, the cross-linking oligonucleotide includes a plurality of cross-linking oligonucleotides connected to each other. In some embodiments, the cross-linking oligonucleotide may be linked or connected to an immunoglobulin-binding protein such as protein A, protein G, protein A / G, or protein L or a fragment thereof. Optionally, the linked cross-linking oligonucleotide can be delivered to the position where an antibody binds in the nucleic acid of the sample.
[0147] In an additional aspect, a method involving the QC quantification step herein does not include a shearing step.
[0148] In a further aspect of the method involving the QC quantification step herein, the method includes obtaining at least some sequences on both sides of the junction to generate a first read pair. For example, the method may include obtaining sequences of at least about 50 bp, at least about 100 bp, at least about 150 bp, at least about 200 bp, at least about 250 bp, or at least about 300 bp on each side of the junction to generate a first read pair.
[0149] In an additional aspect of the method involving the QC quantification step herein, the method further includes mapping the first read pair to a set of contigs and determining a path across the set of contigs that represents the order and / or orientation to the genome.
[0150] In a further aspect of the method involving the QC quantification step herein, the method may include mapping the first read pair to a set of contigs and determining the presence of structural variants or a decrease in heterozygosity in the stabilized biological sample from the set of contigs.
[0151] In an additional aspect of the method involving the QC quantification step herein, the method includes mapping the first read pair to a set of contigs and assigning a phase to the variants in the set of contigs.
[0152] In a further aspect of the method involving the QC quantification step herein, the method includes mapping the first read pair to a set of contigs, determining the presence of variants in the set of contigs from the set of contigs, and performing one or more steps selected from (1) a step of confirming the disease stage, prognosis, or treatment regimen for the stabilized biological sample, (2) a step of selecting a drug based on the presence of the variant, or (3) a step of confirming the drug efficacy for the stabilized biological sample.
[0153] Hi-C method, including digestion of whole cells or whole nuclei Also provided herein is a method that includes obtaining a stabilized biological sample comprising a nucleic acid molecule complexed to at least one nucleic acid-binding protein, contacting the stabilized biological sample with DNase to cleave the nucleic acid molecule into a plurality of segments, and attaching a first segment and a second segment of the plurality of segments at a junction, wherein the stabilized biological sample comprises intact cells and / or whole nuclei. Optionally, the stabilized biological sample comprises stabilized intact cells. Alternatively, or in combination, the stabilized biological sample comprises stabilized intact nuclei. Optionally, the step of contacting the stabilized intact cell or intact nuclear sample with DNase is performed prior to lysis of the intact cell or intact nucleus. Optionally, the cells and / or nuclei are lysed prior to attaching the first segment and the second segment of the plurality of segments at the junction.
[0154] In other aspects, methods herein involving digestion of whole cells or whole nuclei can include subjecting a plurality of segments to size selection to obtain a plurality of selected segments. Optionally, the plurality of selected segments are from about 145 to about 600 bp. Optionally, the plurality of selected segments are from about 100 to about 2500 bp. Optionally, the plurality of selected segments are from about 100 to about 600 bp. Optionally, the plurality of selected segments are from about 600 to about 2500 bp. Optionally, the plurality of selected segments are between about 100 bp and about 600 bp, between about 100 bp and about 700 bp, between about 100 bp and about 800 bp, between about 100 bp and about 900 bp, between about 100 bp and about 1000 bp, between about 100 bp and about 1100 bp, between about 100 bp and about 1200 bp, between about 100 bp and about 1300 bp, between about 100 bp and about 1400 bp, between about 100 bp and about 1500 bp, between about 100 bp and about 1600 bp, between about 100 bp and about 1700 bp, between about 100 bp and about 1800 bp, between about 100 bp and about 1900 bp, between about 100 bp and about 2000 bp, between about 100 bp and about 2100 bp, between about 100 bp and about 2200 bp, between about 100 bp and about 2300 bp, between about 100 bp and about 2400 bp, or between about 100 bp and about 2500 bp.
[0155] In another aspect of the methods provided herein that involve digestion of whole cells or whole nuclei, the method further comprises preparing a sequencing library from a plurality of segments prior to the size selection step. In some embodiments, the method further comprises subjecting the sequencing library to size selection to obtain a size selected library. Optionally, the size selected library is in a size between about 350 bp and about 1000 bp. Optionally, the size selected library is in a size between about 100 bp and about 2500 bp, e.g., between about 100 bp and about 350 bp, between about 350 bp and about 500 bp, between about 500 bp and about 1000 bp, between about 1000 bp and about 1500 bp and about 2000 bp, between about 2000 bp and about 2500 bp, between about 350 bp and about 1000 bp, between about 350 bp and about 1500 bp, between about 350 bp and about 2000 bp, between about 350 bp and about 2500 bp, between about 500 bp and about 1500 bp, between about 500 bp and about 2000 bp, between about 500 bp and about 3500 bp, between about 1000 bp and about 1500 bp, between about 1000 bp and about 2000 bp, between about 1000 bp and about 2500 bp, between about 1500 bp and about 2000 bp, between about 1500 bp and about 2500 bp, or between about 100 bp and about 700 bp.
[0156] The size selection utilized in the methods herein that involve digestion of whole cells or whole nuclei can be performed by gel electrophoresis, capillary electrophoresis, size selection beads, gel filtration columns, or combinations thereof.
[0157] In other aspects, methods herein involving digestion of whole cells or whole nuclei may further include analyzing a plurality of selected segments to obtain one QC value. Optionally, the QC value is selected from chromatin digestion efficiency (CDE) and chromatin digestion index (CDI). CDE is calculated as the proportion of segments having a desired length. For example, optionally, CDE is calculated as the proportion of segments having a size between 100 and 2500 bp prior to size selection. Optionally, when the CDE value is at least 65%, the sample is selected for further analysis. Optionally, the sample is selected for further analysis when the CDE value is at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, or at least about 95%. CDI is calculated as the ratio of the number of segments of mononucleosome size to the number of segments of dinucleosome size prior to size selection. For example, CDI can be calculated as the logarithm of the ratio of fragments having a size of 600 - 2500 bp to fragments having a size of 100 - 600 bp. Optionally, the sample is selected for further analysis when the CDI value exceeds -1.5 and is less than 1.In some cases, the sample is selected for further analysis when the CDI value is greater than approximately -2 and less than approximately 1.5, greater than approximately -1.9 and less than approximately 1.5, greater than approximately -1.8 and less than approximately 1.5, greater than approximately -1.7 and less than approximately 1.5, greater than approximately -1.6 and less than approximately 1.5, greater than approximately -1.5 and less than approximately 1.5, greater than approximately -1.4 and less than approximately 1.5, greater than approximately -1.3 and less than approximately 1.5, greater than approximately -1.2 and less than approximately 1.5, greater than approximately -1.1 and less than approximately 1.5, greater than approximately -2 and less than approximately 1.5, greater than approximately -1 and less than approximately 1.5, greater than approximately -0.9 and less than approximately 1.5, greater than approximately -0.8 and less than approximately 1.5, greater than approximately -0.7 and less than approximately 1.5, greater than approximately -0.6 and less than approximately 1.5, greater than approximately -0.5 and less than approximately 1.5, greater than approximately -2 and less than approximately 1.4, greater than approximately -2 and less than approximately 1.3, greater than approximately -2 and less than approximately 1.2, greater than approximately -2 and less than approximately 1.1, greater than approximately -2 and less than approximately 1, greater than approximately -2 and less than approximately 0.9, greater than approximately -2 and less than approximately 0.8, greater than approximately -2 and less than approximately 0.7, greater than approximately -2 and less than approximately 0.6, or greater than approximately -2 and less than approximately 0.5.
[0158] In other embodiments, the methods herein involving digestion of whole cells or whole nuclei are performed on small samples containing only a few cells or small amounts of nucleic acid. In some cases, the stabilized biological sample contains less than 3,000,000 cells. In some cases, the stabilized biological sample contains less than 2,000,000 cells. In some cases, the stabilized biological sample contains less than 1,000,000 cells. In some cases, the stabilized biological sample contains less than 500,000 cells. In some cases, the stabilized biological sample contains less than 400,000 cells. In some cases, the stabilized biological sample contains less than 300,000 cells. In some cases, the stabilized biological sample contains less than 200,000 cells. In some cases, the stabilized biological sample contains less than 100,000 cells. In some cases, the stabilized biological sample contains less than 10 μg of DNA. In some cases, the stabilized biological sample contains less than 9 μg of DNA. In some cases, the stabilized biological sample contains less than 8 μg of DNA. In some cases, the stabilized biological sample contains less than 7 μg of DNA. In some cases, the stabilized biological sample contains less than 6 μg of DNA. In some cases, the stabilized biological sample contains less than 5 μg of DNA. In some cases, the stabilized biological sample contains less than 4 μg of DNA. In some cases, the stabilized biological sample contains less than 3 μg of DNA. In some cases, the stabilized biological sample contains less than 2 μg of DNA. In some cases, the stabilized biological sample contains less than 1 μg of DNA. In some cases, the stabilized biological sample contains less than 0.5 μg of DNA.
[0159] In other aspects, the methods herein involving digestion of whole cells or whole nuclei can be performed on individual cells or single cells. For example, the methods herein may be performed on cells distributed into individual compartments. Exemplary compartments include, but are not limited to, wells, droplets in an emulsion, or surface locations (e.g., array spots, beads, etc.), which compartments contain distinct patches of differentially arrayed linker molecules as otherwise described herein. Additional compartments are contemplated and are consistent with the methods, compositions, and systems disclosed herein.
[0160] In additional aspects, the stabilized biological samples utilized in the methods herein involving digestion of whole cells or whole nuclei are treated with a nuclease, such as DNase, to create DNA fragments. In some cases, the DNase is non-sequence specific. In some cases, the DNase is active against both single-stranded and double-stranded DNA. In some cases, the DNase is specific for double-stranded DNA. In some cases, the DNase preferentially cleaves double-stranded DNA. In some cases, the DNase is specific for single-stranded DNA. In some cases, the DNase preferentially cleaves single-stranded DNA. In some cases, the DNase is DNase I. In some cases, the DNase is DNase II. In some cases, the DNase is selected from one or more of DNase I and DNase II. In some cases, the DNase is micrococcal nuclease. In some cases, the DNase is selected from one or more of DNase I, DNase II, and micrococcal nuclease. In some cases, the DNase may be linked or fused to an immunoglobulin binding protein, such as protein A, protein G, protein A / G, or protein L, or fragments thereof. Other suitable nucleases are also within the scope of this disclosure.
[0161] In an additional aspect, in the methods utilized herein that involve digestion of whole cells or whole nuclei, the stabilized samples are treated with a cross-linking agent. In some cases, the cross-linking agent is a chemical fixative. In some cases, the chemical fixative includes formaldehyde having a spacer arm length of about 2.3 - 2.7 angstroms (A). In some cases, the chemical fixative includes a cross-linking agent having a long spacer arm length. For example, the cross-linking agent may have a spacer length of at least about 3A, 4A, 5A, 6A, 7A, 8A, 9A, 10A, 11A, 12A, 13A, 14A, 15A, 16A, 17A, 18A, 19A, or 20A. The chemical fixative includes ethylene glycol bis(succinimidyl succinate) (EGS) and has a spacer arm with a length of about 16.1A. The chemical fixative may include disuccinimidyl glutarate (DSG) and has a spacer arm with a length of about 7.7A. In some cases, the chemical fixative includes formaldehyde and EGS, formaldehyde and DSG, or formaldehyde, EGS, and DSG. In some cases where multiple chemical fixatives are utilized, each chemical fixative is used sequentially, and in other cases, some or all of the multiple chemical fixatives are applied to the sample simultaneously. The use of a cross-linking agent having a long spacer arm can increase the fragments of read pairs having a large (e.g., >1 kb) read pair separation distance. For example, FIG. 7 shows a comparison of the resulting libraries (both DNase digestion and MNase digestion) cross-linked with formaldehyde alone and cross-linked with formaldehyde plus DSG or EGS. DSG has NHS ester reactive groups at both ends and can be reactive towards amino groups (e.g., primary amines). DSG is membrane-permeable and enables cross-linking within cells. DSG can increase cross-linking efficiency in some applications compared to disuccinimidyl suberate (DSS). EGS has NHS ester reactive groups at both ends and can be reactive towards amino groups (e.g., primary amines). EGS is membrane-permeable and enables cross-linking within cells.EGS crosslinking can be reversed, for example, by treating with hydroxylamine at pH 8.5 for 3 to 6 hours. In one example, lactate dehydrogenase retained 60% of its activity after reversible crosslinking with EGS. In some cases, the chemical fixative includes psoralen. In some cases, the crosslinking agent is ultraviolet light. In some embodiments, the stabilized biological sample is a crosslinked paraffin-embedded tissue sample.
[0162] In further embodiments, the methods herein involving digestion of whole cells or whole nuclei include contacting a plurality of selected segments with an antibody.
[0163] In an additional aspect, the methods provided herein that involve digestion of whole cells or whole nuclei include attaching a first segment and a second segment of a plurality of segments at one junction. Optionally, the attaching step includes filling in sticky ends and ligating blunt ends using biotin-tagged nucleotides. Optionally, the attaching step includes contacting at least the first segment and the second segment with a crosslinking oligonucleotide. Optionally, the attaching step includes contacting at least the first segment and the second segment with a barcode. In some embodiments, the crosslinking oligonucleotides herein can be from at least about 5 nucleotides to about 50 nucleotides in length. In some embodiments, the crosslinking oligonucleotides herein can be from at least about 15 nucleotides to about 18 nucleotides in length. In some embodiments, the crosslinking oligonucleotide is about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, about 20, about 25, about 30, about 35, about 40, about 45, or about 50 nucleotides in length. In some embodiments, the crosslinking oligonucleotides herein can include a barcode. In some embodiments, the crosslinking oligonucleotide can include a plurality of barcodes. In some embodiments, the crosslinking oligonucleotide includes a plurality of crosslinking oligonucleotides connected to each other. In some embodiments, the crosslinking oligonucleotide may be linked or connected to an immunoglobulin binding protein such as protein A, protein G, protein A / G, or protein L or a fragment thereof. Optionally, the linked crosslinking oligonucleotide can be delivered to the location where an antibody binds in the nucleic acid of the sample.
[0164] The splitting and pooling approach can be utilized to generate cross-linked oligonucleotides with unique barcodes. A population of samples can be divided into multiple groups, and the cross-linked oligonucleotides can be attached to the samples such that the barcode of the cross-linked oligonucleotides differs between groups but is the same within a group. The groups of samples can be pooled together again, and this process can be repeated multiple times. Repeating this process results in each sample in a population having a unique set of cross-linked oligonucleotide barcodes, which may enable the analysis of a single sample (e.g., a single cell, a single nucleus, a single chromosome). In one illustrative example, in one illustrative example, a sample of cross-linked and digested nuclei attached to a solid support of beads is separated across eight tubes, each containing one of eight unique members of a first adapter set (first iteration) of double-stranded DNA (dsDNA) adapters that are to be ligated. Each of the eight adapters may have the same 5' overhang sequence for ligation to the nucleic acid termini of the cross-linked chromatin assemblies in the nuclei, or otherwise have a unique dsDNA sequence. After the first adapter set has been ligated, the nuclei are pooled together again and can be washed to remove the ligation reaction components. The scheme of distributing, ligating, and pooling can be repeated two more times (two iterations). After ligation of the members of each adapter set, the cross-linked chromatin assemblies can be attached to a series of multiple barcodes. In some cases, successive ligation of multiple members of multiple adapter sets (iterations) results in a combination of barcodes. The number of possible combinations of barcodes depends on the number of groups per iteration and the total number of barcode oligonucleotides used. For example, three iterations with eight members can each have 83 possible combinations individually. In some cases, the combination of barcodes is unique. In some cases, there is redundancy in the combination of barcodes.The total number of barcode combinations can be adjusted by increasing or decreasing the number of groups receiving unique barcodes and / or by increasing or decreasing the number of iterations. If more than one adapter group is used, a scheme of distributing, attaching, and pooling can be used for iterative adapter attachment. In some cases, the scheme of distributing, attaching, and pooling can be iterated an additional at least 3, 4, 5, 6, 7, 8, 9, or 10 times. In some cases, the members of the last adapter group include sequences for subsequent enrichment of adapter-attached DNA, for example, during preparation of a sequencing library by PCR amplification.
[0165] In an additional aspect, methods herein involving digestion of whole cells or whole nuclei do not include a shearing step.
[0166] In a further aspect of the methods provided herein involving digestion of whole cells or whole nuclei, the method includes obtaining at least some sequences on both sides of the junction to generate a first read pair. For example, the method can include obtaining sequences of at least about 50 bp, at least about 100 bp, at least about 150 bp, at least about 200 bp, at least about 250 bp, or at least about 300 bp on each side of the junction to generate a first read pair.
[0167] In an additional aspect of the methods provided herein involving digestion of whole cells or whole nuclei, the method further includes mapping the first read pair to a set of contigs and determining a path across the set of contigs that represents the order and / or orientation to the genome.
[0168] In a further aspect of the methods provided herein involving digestion of whole cells or whole nuclei, the method includes mapping the first read pair to a set of contigs and determining the presence of structural variants or a decrease in heterozygosity in a stabilized biological sample from the set of contigs.
[0169] In additional aspects of the methods provided herein that involve digestion of whole cells or whole nuclei, the method includes mapping a first read pair to a set of contigs and assigning a phase to variants in the set of contigs.
[0170] In further aspects of the methods provided herein that involve digestion of whole cells or whole nuclei, the method further includes mapping a first read pair to a set of contigs, determining the presence of variants in the set of contigs from the set of contigs, and performing one or more steps selected from: (1) confirming a disease stage, prognosis, or treatment regimen for a stabilized biological sample; (2) selecting a drug based on the presence of the variant; or (3) confirming the efficacy of a drug against a stabilized biological sample.
[0171] Hi-C method with low nucleic acid input requirements In addition, a method is provided herein, the method comprising obtaining a stabilized biological sample comprising a nucleic acid molecule complexed to at least one nucleic acid binding protein, contacting the stabilized biological sample with DNase to cleave the nucleic acid molecule into a plurality of segments, and attaching a first segment and a second segment of the plurality of segments at a junction, wherein the stabilized biological sample comprises less than 3,000,000 cells or less than 10 μg of DNA. Optionally, the stabilized biological sample comprises less than 3,000,000 cells. Optionally, the stabilized biological sample comprises less than 2,000,000 cells. Optionally, the stabilized biological sample comprises less than 1,000,000 cells. Optionally, the stabilized biological sample comprises less than 500,000 cells. Optionally, the stabilized biological sample comprises less than 400,000 cells. Optionally, the stabilized biological sample comprises less than 300,000 cells. Optionally, the stabilized biological sample comprises less than 200,000 cells. Optionally, the stabilized biological sample comprises less than 100,000 cells. Optionally, the stabilized biological sample comprises less than 10 μg of DNA. Optionally, the stabilized biological sample comprises less than 9 μg of DNA. Optionally, the stabilized biological sample comprises less than 8 μg of DNA. Optionally, the stabilized biological sample comprises less than 7 μg of DNA. Optionally, the stabilized biological sample comprises less than 6 μg of DNA. Optionally, the stabilized biological sample comprises less than 5 μg of DNA. Optionally, the stabilized biological sample comprises less than 4 μg of DNA. Optionally, the stabilized biological sample comprises less than 3 μg of DNA. Optionally, the stabilized biological sample comprises less than 2 μg of DNA. Optionally, the stabilized biological sample comprises less than 1 μg of DNA. Optionally, the stabilized biological sample comprises less than 0.5 μg of DNA.
[0172] In another aspect, the methods having low nucleic acid input requirements herein may be performed on individual cells or a single cell. For example, the methods herein may be performed on cells distributed in individual compartments. Examples of exemplary compartments include, but are not limited to, wells, droplets in an emulsion, or surface positions (such as array spots, beads, etc.) containing individual patches of linker molecules of different sequences as described elsewhere herein. Additional compartments are also contemplated and are consistent with the methods, compositions, and systems disclosed herein.
[0173] In another aspect, the methods having low nucleic acid input requirements herein include subjecting a plurality of segments to size selection to obtain a plurality of selected segments. Optionally, the plurality of selected segments are from about 145 to about 600 bp. Optionally, the plurality of selected segments are from about 100 to about 2500 bp. Optionally, the plurality of selected segments are from about 100 to about 600 bp. Optionally, the plurality of selected segments are from about 600 to about 2500 bp. Optionally, the plurality of selected segments are from about 100 bp to about 600 bp, from about 100 bp to about 700 bp, from about 100 bp to about 800 bp, from about 100 bp to about 900 bp, from about 100 bp to about 1000 bp, from about 100 to about 1100 bp, from about 100 bp to about 1200 bp, from about 100 bp to about 1300 bp, from about 100 bp to about 1400 bp, from about 100 bp to about 1500 bp, from about 100 bp to about 1600 bp, from about 100 bp to about 1700 bp, from about 100 bp to about 1800 bp, from about 100 bp to about 1900 bp, from about 100 bp to about 2000 bp, from about 100 bp to about 2100 bp, from about 100 bp to about 2200 bp, from about 100 bp to about 2300 bp, from about 100 bp to about 2400 bp, or from about 100 bp to about 2500 bp.
[0174] In another aspect of the methods provided herein having low nucleic acid input requirements, the method further includes, prior to the size selection step, preparing a sequencing library from a plurality of segments. In some embodiments, the method further includes subjecting the sequencing library to size selection to obtain a size selected library. Optionally, the size selected library is in the size range of about 350 bp to about 1000 bp. Optionally, the size selected library is in the size range of about 100 bp to about 2500 bp, for example, about 100 bp to about 350 bp, about 350 bp to about 500 bp, about 500 bp to about 1000 bp, about 1000 to about 1500 bp to about 2000 bp, about 2000 to about 2500 bp, about 350 bp to about 1000 bp, about 350 bp to about 1500 bp, about 350 bp to about 2000 bp, about 350 bp to about 2500 bp, about 500 bp to about 1500 bp, about 500 bp to about 2000 bp, about 500 bp to about 3500 bp, about 1000 bp to about 1500 bp, about 1000 bp to about 2000 bp, about 1000 bp to about 2500 bp, about 1500 bp to about 2000 bp, about 1500 bp to about 2500 bp, or about 2000 bp to about 2500 bp.
[0175] Size selection utilized in the methods having low nucleic acid input requirements herein is often performed by gel electrophoresis, capillary electrophoresis, size selection beads, gel filtration columns, or combinations thereof.
[0176] In another aspect, the method with low nucleic acid input requirements herein may further include the step of analyzing a plurality of selected segments to obtain a QC value. In some cases, the QC value is selected from a chromatin digestion efficiency (CDE) and a chromatin digestion index (CDI). CDE is calculated as the proportion of segments having a desired length. For example, in some cases, CDE is calculated as the proportion of segments sized 100-2500 bp before size selection. In some cases, when the CDE value is at least 65%, the sample is selected for further analysis. In some cases, when the CDE value is at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, or at least about 95%, the sample is selected for further analysis. CDI is calculated as the ratio of the number of mononucleosome-sized segments to the number of dinucleosome-sized segments before size selection. For example, CDI may be calculated as the logarithm of the ratio of fragments having a size of 600-2500 bp to fragments having a size of 100-600 bp. In some cases, when the CDI value exceeds -1.5 and is less than 1, the sample is selected for further analysis.In some cases, when the CDI value is greater than about -2 and less than about 1.5, greater than about -1.9 and less than about 1.5, greater than about -1.8 and less than about 1.5, greater than about -1.7 and less than about 1.5, greater than about -1.6 and less than about 1.5, greater than about -1.5 and less than about 1.5, greater than about -1.4 and less than about 1.5, greater than about -1.3 and less than about 1.5, greater than about -1.2 and less than about 1.5, greater than about -1.1 and less than about 1.5, greater than about -2 and less than about 1.5, greater than about -1 and less than about 1.5, greater than about -0.9 and less than about 1.5, greater than about -0.8 and less than about 1.5, greater than about -0.7 and less than about 1.5, greater than about -0.6 and less than about 1.5, greater than about -0.5 and less than about 1.5, greater than about -2 and less than about 1.4, greater than about -2 and less than about 1.3, greater than about -2 and less than about 1.2, greater than about -2 and less than about 1.1, greater than about -2 and less than about 1, greater than about -2 and less than about 0.9, greater than about -2 and less than about 0.8, greater than about -2 and less than about 0.7, greater than about -2 and less than about 0.6, or greater than about -2 and less than about 0.5, the sample is selected for further analysis.
[0177] In another aspect, the stabilized biological sample used in the methods having low nucleic acid input requirements herein comprises a biological material treated with a stabilizer. In some cases, the stabilized biological sample comprises a stabilized cell lysate. Alternatively, the stabilized biological sample comprises stabilized intact cells. Alternatively, the stabilized biological sample comprises stabilized intact nuclei. In some cases, a step of contacting a sample of stabilized intact cells or nuclei with DNase is performed prior to lysis of the intact cells or nuclei. In some cases, the cells and / or nuclei are lysed prior to a step of attaching a first segment and a second segment of a plurality of segments at a junction.
[0178] In an additional aspect, the stabilized biological sample used in the methods having low nucleic acid input requirements herein is treated with a nuclease, such as DNase, to create DNA fragments. In some cases, the DNase is not sequence specific. In some cases, the DNase is active against both single-stranded DNA and double-stranded DNA. In some cases, the DNase is specific for double-stranded DNA. In some cases, the DNase preferentially cleaves double-stranded DNA. In some cases, the DNase is specific for single-stranded DNA. In some cases, the DNase preferentially cleaves single-stranded DNA. In some cases, the DNase is DNase I. In some cases, the DNase is DNase II. In some cases, the DNase is selected from one or more of DNase I and DNase II. In some cases, the DNase is micrococcal nuclease. In some cases, the DNase is selected from one or more of DNase I, DNase II, and micrococcal nuclease. In some cases, the DNase may be linked or fused to an immunoglobulin-binding protein or fragment thereof, such as protein A, protein G, protein A / G, or protein L. Other suitable nucleases are also within the scope of the present disclosure.
[0179] In an additional aspect, the stabilized biological sample used in the methods having low nucleic acid input requirements herein is treated with a crosslinking agent. In some cases, the crosslinking agent is a chemical fixative. In some cases, the chemical fixative includes formaldehyde and has a spacer arm length of about 2.3 - 2.7 angstroms (A). In some cases, the chemical fixative includes a crosslinking agent having a long spacer arm length. For example, the crosslinking agent can have a spacer length of at least about 3A, 4A, 5A, 6A, 7A, 8A, 9A, 10A, 11A, 12A, 13A, 14A, 15A, 16A, 17A, 18A, 19A, or 20A. The chemical fixative can include ethylene glycol bis(succinimidyl succinate) (EGS), which has a spacer arm with a length of about 16.1A. The chemical fixative can include disuccinimidyl glutarate (DSG), which has a spacer arm with a length of about 7.7A. In some cases, the chemical fixative includes formaldehyde and EGS, formaldehyde and DSG, or formaldehyde, EGS, and DSG. In some cases where multiple chemical fixatives are used, each chemical fixative is used sequentially; in other cases, some or all of the multiple chemical fixatives are applied to the sample simultaneously. The use of a crosslinking agent having a long spacer arm can increase the fraction of read pairs having a large (e.g., >1 kb) separation distance. For example, FIG. 7 shows a comparison between a library resulting from crosslinking with formaldehyde only (both digested with DNase and MNase) and a library resulting from crosslinking with formaldehyde and DSG or EGS. DSG has NHS ester reactive groups at both ends and can be reactive towards amino groups (e.g., primary amines). DSG is membrane permeable and enables intracellular crosslinking. DSG can increase crosslinking efficiency compared to disuccinimidyl suberate (DSS) in some applications. EGS has NHS ester reactive groups at both ends and can be reactive towards amino groups (e.g., primary amines). EGS is membrane permeable and enables intracellular crosslinking.For example, EGS crosslinking can be reversed by treating with hydroxylamine at pH 8.5 for 3 to 6 hours; in one example, lactate dehydrogenase retained 60% of its activity after reversible crosslinking with EGS. In some cases, the chemical fixative contains psoralen. In some cases, the crosslinking agent is ultraviolet light. In some cases, the stabilized biological sample is a crosslinked paraffin-embedded tissue sample.
[0180] In a further aspect, the methods provided herein include contacting an antibody with a plurality of selected segments.
[0181] In an additional aspect, the methods provided herein having low nucleic acid input requirements include the step of attaching a first segment and a second segment among a plurality of segments at a junction. In some cases, the attaching step includes filling sticky ends using biotin-tagged nucleotides and ligating blunt ends. In some cases, the attaching step includes contacting at least the first segment and the second segment with a cross-linking oligonucleotide. In some cases, the attaching step includes contacting at least the first segment and the second segment with a barcode. In some embodiments, the cross-linking oligonucleotides herein may be from at least about 5 nucleotides in length to about 50 nucleotides in length. In some embodiments, the cross-linking oligonucleotides herein may be from about 15 to about 18 nucleotides in length. In some embodiments, the cross-linking oligonucleotide may be about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, about 20, about 25, about 30, about 35, about 40, about 45, or about 50 nucleotides in length. In some embodiments, the cross-linking oligonucleotides herein may include a barcode. In some embodiments, the cross-linking oligonucleotide may include a plurality of barcodes. In some embodiments, the cross-linking oligonucleotide includes a plurality of cross-linking oligonucleotides that are connected together. In some embodiments, the cross-linking oligonucleotide may be linked or attached to an immunoglobulin-binding protein or a fragment thereof, such as protein A, protein G, protein A / G, or protein L. In some cases, the linked cross-linking oligonucleotide may be delivered to a position in the sample nucleic acid where an antibody binds.
[0182] The splitting and pooling approach can be used to generate cross-linked oligonucleotides with unique barcodes. The population of samples can be divided into multiple groups, and the cross-linked oligonucleotides can be attached to the samples such that the cross-linked nucleotide barcodes are different between groups but the same within one group. The groups of samples can be pooled together again, and this process can be repeated multiple times. By repeating this process, ultimately each sample within the population will have a unique series of cross-linked oligonucleotide barcodes, enabling the analysis of a single sample (e.g., a single cell, a single nucleus, a single chromosome). In one exemplary example, a sample of cross-linked digested nuclei attached to a solid support of beads is divided across eight tubes, each containing one of eight unique members of a first adapter group (first iteration) that includes a double-stranded DNA (dsDNA) adapter to be ligated. Each of the eight adapters can have the same 5’ overhang sequence for ligation to the nucleic acid termini of the cross-linked chromatin assemblies in the nuclei, but otherwise has a unique dsDNA sequence. After the first adapter group has been ligated, the nuclei can be pooled together again and washed to remove the ligation reaction components. The scheme of partitioning, ligation, and pooling can be repeated two additional times (two iterations). After ligation of the members from each adapter group, the cross-linked chromatin assemblies can be sequentially attached to multiple barcodes. In some cases, sequential ligation (iteration) of multiple members of multiple adapter groups results in a combination of barcodes. The number of possible combinations of barcodes depends on the number of groups per iteration and the total number of barcode oligonucleotides used. For example, three iterations each containing eight members can have 8^3 possible combinations. In some cases, the combination of barcodes is unique. In some cases, the combination of barcodes is redundant.The total number of barcode combinations can be adjusted by increasing or decreasing the number of groups receiving unique barcodes and / or by increasing or decreasing the number of repeats. When more than one adapter group is used, the distribution, attachment, and pooling schemes can be used for repeated adapter attachment. In some cases, the distribution, attachment, and pooling schemes can be additionally repeated at least 3, 4, 5, 6, 7, 8, 9, or 10 times. In some cases, the members of the last adapter group contain sequences for subsequent enrichment of adapter-attached DNA during the preparation of a sequencing library, e.g., through PCR amplification.
[0183] In additional aspects, methods herein having low nucleic acid input requirements do not include a shearing step.
[0184] In further aspects of methods herein having low nucleic acid input requirements, the method includes obtaining at least some sequence on each side of a junction to generate a first read pair. For example, the method may include obtaining at least about 50 bp, at least about 100 bp, at least about 150 bp, at least about 200 bp, at least about 250 bp, or at least about 300 bp of sequence on each side of a junction to generate a first read pair.
[0185] In additional aspects of methods herein having low nucleic acid input requirements, the method includes mapping a first read pair to a set of contigs and determining a path through the set of contigs to the genome that represents order and / or orientation.
[0186] In further aspects of methods herein having low nucleic acid input requirements, the method includes mapping a first read pair to a set of contigs and determining the presence of structural variants or loss of heterozygosity in a stabilized biological sample from the set of contigs.
[0187] In additional aspects of the methods having low nucleic acid input requirements herein, the method includes mapping a first read pair to a set of contigs and assigning a phase to variants in the set of contigs.
[0188] In further aspects of the methods having low nucleic acid input requirements herein, the method includes mapping a first read pair to a set of contigs, determining the presence of variants in the set of contigs from the set of contigs, and performing one or more steps selected from (1) identifying a disease stage, prognosis, or treatment regimen for a stabilized biological sample, (2) selecting a drug based on the presence of the variant, or (3) identifying drug efficacy for a stabilized biological sample.
[0189] Hi-C method using Micrococcus nuclease (MNase) In addition, methods are provided herein, the methods including obtaining a stabilized biological sample comprising a nucleic acid molecule complexed to at least one nucleic acid binding protein, contacting the stabilized biological sample with Micrococcus nuclease (MNase) to cleave the nucleic acid molecule into a plurality of segments, and attaching a first segment and a second segment of the plurality of segments at a junction. The use of MNase in the methods herein may provide information, for example, as to where the DNA binding protein binds to chromatin with a resolution of up to a single base pair, since MNase can cleave all base pairs that are not bound to a DNA binding protein. In addition, the use of MNase digestion may make it possible to create a contact map and topologically associated domains for decoding three-dimensional chromatin structure information. Optionally, MNase may be linked or fused to an immunoglobulin binding protein or fragment thereof, such as protein A, protein G, protein A / G, or protein L.
[0190] For example, the MNase Hi-C method can provide the positions of protein binding or genomic contact interactions at a resolution of about 1 bp, 2 bp, 3 bp, 4 bp, 5 bp, 6 bp, 7 bp, 8 bp, 9 bp, 10 bp, 20 bp, 30 bp, 40 bp, 50 bp, 60 bp, 70 bp, 80 bp, 90 bp, 100 bp, 200 bp, 300 bp, 400 bp, 500 bp, 600 bp, 700 bp, 800 bp, 900 bp, 1000 bp, 2000 bp, 3000 bp, 4000 bp, 5000 bp, 6000 bp, 7000 bp, 8000 bp, 9000 bp, 10 kb, 20 kb, 30 kb, 40 kb, 50 kb, 60 kb, 70 kb, 80 kb, 90 kb, or less than 100 kb, or equal to them. In some cases, protein binding sites, protein footprints, contact interactions, or other features can be mapped within 1000 bp, 900 bp, 800 bp, 700 bp, 600 bp, 500 bp, 400 bp, 300 bp, 200 bp, 190 bp, 180 bp, 170 bp, 160 bp, 150 bp, 140 bp, 130 bp, 120 bp, 110 bp, 100 bp, 90 bp, 80 bp, 70 bp, 60 bp, 50 bp, 40 bp, 30 bp, 20 bp, 10 bp, 9 bp, 8 bp, 7 bp, 6 bp, 5 bp, 4 bp, 3 bp, 2 bp, 1 bp.
[0191] In one aspect, the method comprising the low nucleic acid MNase digestion step herein may further include a step of subjecting a plurality of segments to size selection to obtain a plurality of selected segments. In some cases, the plurality of selected segments can be from about 145 to about 600 bp. In some cases, the plurality of selected segments can be from about 100 to about 2500 bp. In some cases, the plurality of selected segments can be from about 100 to about 600 bp. In some cases, the plurality of selected segments can be from about 600 to about 2500 bp. In some cases, the plurality of selected segments can be from about 100 bp to about 600 bp, from about 100 bp to about 700 bp, from about 100 bp to about 800 bp, from about 100 bp to about 900 bp, from about 100 bp to about 1000 bp, from about 100 to about 1100 bp, from about 100 bp to about 1200 bp, from about 100 bp to about 1300 bp, from about 100 bp to about 1400 bp, from about 100 bp to about 1500 bp, from about 100 bp to about 1600 bp, from about 100 bp to about 1700 bp, from about 100 bp to about 1800 bp, from about 100 bp to about 1900 bp, from about 100 bp to about 2000 bp, from about 100 bp to about 2100 bp, from about 100 bp to about 2200 bp, from about 100 bp to about 2300 bp, from about 100 bp to about 2400 bp, or from about 100 bp to about 2500 bp.
[0192] In another aspect of the method including the MNase digestion step as provided herein, the method may further include the step of preparing a sequencing library from a plurality of segments. In some embodiments, the method may further include the step of subjecting the sequencing library to size selection to obtain a size-selected library. Optionally, the size-selected library may be in the size range of about 350 bp to about 1000 bp. Optionally, the size-selected library may be in the size range of about 100 bp to about 2500 bp, for example, about 100 bp to about 350 bp, about 350 bp to about 500 bp, about 500 bp to about 1000 bp, about 1000 to about 1500 bp, about 2000 to about 2500 bp, about 350 bp to about 1000 bp, about 350 bp to about 1500 bp, about 350 bp to about 2000 bp, about 350 bp to about 2500 bp, about 500 bp to about 1500 bp, about 500 bp to about 2000 bp, about 500 bp to about 3500 bp, about 1000 bp to about 1500 bp, about 1000 bp to about 2000 bp, about 1000 bp to about 2500 bp, about 1500 bp to about 2000 bp, about 1500 bp to about 2500 bp, or about 2000 bp to about 2500 bp.
[0193] In another aspect, the method including the MNase digestion step as provided herein can further include the step of analyzing a plurality of segments to obtain QC values. Optionally, the QC values may be selected from chromatin digestion efficiency (CDE) and chromatin digestion index (CDI). CDE can be calculated as the proportion of segments having a desired length. For example, optionally, CDE can be calculated as the proportion of segments sized 100 bp to 2500 bp before size selection. Optionally, when the CDE value is at least 65%, the sample may be selected for further analysis. Optionally, when the CDE value is at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, or at least about 95%, the sample may be selected for further analysis.
[0194] CDI can be calculated as the ratio of the number of segments of mononucleosome size to the number of segments of dinucleosome size before size selection. For example, CDI may be calculated as the logarithm of the ratio of fragments having a size of 600 - 2500 bp to fragments having a size of 100 - 600 bp. In some cases, when the CDI value exceeds -1.5 and is less than 1, the sample may be selected for further analysis. In some cases, when the CDI value exceeds approximately -2 and is less than approximately 1.5, exceeds approximately -1.9 and is less than approximately 1.5, exceeds approximately -1.8 and is less than approximately 1.5, exceeds approximately -1.7 and is less than approximately 1.5, exceeds approximately -1.6 and is less than approximately 1.5, exceeds approximately -1.5 and is less than approximately 1.5, exceeds approximately -1.4 and is less than approximately 1.5, exceeds approximately -1.3 and is less than approximately 1.5, exceeds approximately -1.2 and is less than approximately 1.5, exceeds approximately -1.1 and is less than approximately 1.5, exceeds approximately -2 and is less than approximately 1.5, exceeds approximately -1 and is less than approximately 1.5, exceeds approximately -0.9 and is less than approximately 1.5, exceeds approximately -0.8 and is less than approximately 1.5, exceeds approximately -0.7 and is less than approximately 1.5, exceeds approximately -0.6 and is less than approximately 1.5, exceeds approximately -0.5 and is less than approximately 1.5, exceeds approximately -2 and is less than approximately 1.4, exceeds approximately -2 and is less than approximately 1.3, exceeds approximately -2 and is less than approximately 1.2, exceeds approximately -2 and is less than approximately 1.1, exceeds approximately -2 and is less than approximately 1, exceeds approximately -2 and is less than approximately 0.9, exceeds approximately -2 and is less than approximately 0.8, exceeds approximately -2 and is less than approximately 0.7, exceeds approximately -2 and is less than approximately 0.6, exceeds approximately -2 and is less than approximately 0.5, the sample may be selected for further analysis.
[0195] In another aspect, a stabilized biological sample used in a method having an MNase digestion step as provided herein may comprise a biological material treated with a stabilizer. Optionally, the stabilized biological sample may comprise a cell lysate. Alternatively, the stabilized biological sample may comprise stabilized intact cells. Alternatively, the stabilized biological sample may comprise stabilized intact nuclei. Optionally, prior to lysis of the intact cells or intact nuclei, a step of contacting a sample of the stabilized intact cells or intact nuclei with MNase may be performed. Optionally, the cells and / or nuclei may be lysed prior to a step of attaching a first segment and a second segment of a plurality of segments at a junction.
[0196] In another aspect, the method comprising the MNase digestion step as provided herein may be performed on a small sample that contains few cells or contains a small amount of nucleic acid. For example, in some cases, the stabilized biological sample may contain fewer than 3,000,000 cells. In some cases, the stabilized biological sample may contain fewer than 2,000,000 cells. In some cases, the stabilized biological sample may contain fewer than 1,000,000 cells. In some cases, the stabilized biological sample may contain fewer than 500,000 cells. In some cases, the stabilized biological sample may contain fewer than 400,000 cells. In some cases, the stabilized biological sample may contain fewer than 300,000 cells. In some cases, the stabilized biological sample may contain fewer than 200,000 cells. In some cases, the stabilized biological sample may contain fewer than 100,000 cells. In some cases, the stabilized biological sample may contain less than 10 μg of DNA. In some cases, the stabilized biological sample may contain less than 9 μg of DNA. In some cases, the stabilized biological sample may contain less than 8 μg of DNA. In some cases, the stabilized biological sample may contain less than 7 μg of DNA. In some cases, the stabilized biological sample may contain less than 6 μg of DNA. In some cases, the stabilized biological sample may contain less than 5 μg of DNA. In some cases, the stabilized biological sample may contain less than 4 μg of DNA. In some cases, the stabilized biological sample may contain less than 3 μg of DNA. In some cases, the stabilized biological sample may contain less than 2 μg of DNA. In some cases, the stabilized biological sample may contain less than 1 μg of DNA. In some cases, the stabilized biological sample may contain less than 0.5 μg of DNA.
[0197] In another aspect, a method comprising the MNase digestion step herein may be performed on individual cells or a single cell. For example, the methods herein may be performed on cells distributed in individual compartments. Examples of exemplary compartments include, but are not limited to, wells, droplets in an emulsion, or surface locations (e.g., array spots, beads, etc.) containing individual patches of linker molecules of different sequences as described elsewhere herein. Additional compartments are contemplated and are consistent with the methods, compositions, and systems disclosed herein.
[0198] In an additional aspect, the stabilized biological sample used in a method comprising the MNase digestion step herein may be further treated with an additional nuclease, such as DNase, to create DNA fragments. In some cases, the DNase may not be sequence-specific. In some cases, the DNase may be active against both single-stranded and double-stranded DNA. In some cases, the DNase may be specific for double-stranded DNA. In some cases, the DNase may preferentially cleave double-stranded DNA. In some cases, the DNase may be specific for single-stranded DNA. In some cases, the DNase may preferentially cleave single-stranded DNA. In some cases, the DNase may be DNase I. In some cases, the DNase may be DNase II. In some cases, the DNase may be selected from one or more of DNase I and DNase II. In some cases, the DNase may be linked or fused to an immunoglobulin-binding protein or fragment thereof, such as protein A, protein G, protein A / G, or protein L. Other suitable nucleases are also within the scope of the present disclosure.
[0199] In an additional aspect, the stabilized biological sample as provided herein for use in a method comprising an MNase digestion step can be treated with a crosslinking agent. Optionally, the crosslinking agent can be a chemical fixative. Optionally, the chemical fixative can include formaldehyde and have a spacer arm length of about 2.3 - 2.7 angstroms (A). Optionally, the chemical fixative can include a crosslinking agent with a long spacer arm length. For example, the crosslinking agent can have a spacer length of at least about 3A, 4A, 5A, 6A, 7A, 8A, 9A, 10A, 11A, 12A, 13A, 14A, 15A, 16A, 17A, 18A, 19A, or 20A. The chemical fixative can include ethylene glycol bis(succinimidyl succinate) (EGS), which has a spacer arm with a length of about 16.1A. The chemical fixative can include disuccinimidyl glutarate (DSG), which has a spacer arm with a length of about 7.7A. Optionally, the chemical fixative can include formaldehyde and EGS, formaldehyde and DSG, or formaldehyde, EGS, and 7DSG. In some cases where multiple chemical fixatives are used, each chemical fixative is used sequentially; in other cases, some or all of the multiple chemical fixatives are applied to the sample simultaneously. The use of a crosslinking agent with a long spacer arm can increase the fraction of read pairs with large (e.g., >1 kb) separation distances. For example, Figure 7 shows a comparison between a library resulting from crosslinking (digested with both DNase and MNase) with formaldehyde only and a library resulting from crosslinking with formaldehyde and DSG or EGS. DSG has NHS ester reactive groups at both ends and can be reactive towards amino groups (e.g., primary amines). DSG is membrane permeable and allows for intracellular crosslinking. DSG can increase crosslinking efficiency compared to disuccinimidyl suberate (DSS) in some applications. EGS has NHS ester reactive groups at both ends and can be reactive towards amino groups (e.g., primary amines). EGS is membrane permeable and allows for intracellular crosslinking.For example, EGS cross-linking can be reversed by treating with hydroxylamine at pH 8.5 for 3 to 6 hours; in one example, lactate dehydrogenase retained 60% of its activity after reversible cross-linking with EGS. Optionally, the chemical fixative may contain psoralen. Optionally, the cross-linking agent may be ultraviolet light. Optionally, the stabilized biological sample may be a cross-linked paraffin-embedded tissue sample.
[0200] In a further aspect, a method provided herein that includes the MNase digestion step may include contacting a plurality of selected segments with an antibody. Optionally, an immunoglobulin-binding protein or a fragment thereof tethered to an oligonucleotide adapter may target an antibody bound to a plurality of selected segments.
[0201] In an additional aspect, a method comprising the MNase digestion step provided herein may include a step of attaching a first segment and a second segment among a plurality of segments at a junction. In some cases, the attaching step may include filling in sticky ends and ligating blunt ends using biotin-tagged nucleotides. In some cases, the attaching step may include contacting at least the first segment and the second segment with a cross-linking oligonucleotide. In some cases, the attaching step may include contacting at least the first segment and the second segment with a barcode. In some embodiments, the cross-linking oligonucleotide herein may be at least about 5 nucleotides to about 50 nucleotides in length. In some embodiments, the cross-linking oligonucleotide herein may be about 15 to about 18 nucleotides in length. In some embodiments, the cross-linking oligonucleotide may be about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, about 20, about 25, about 30, about 35, about 40, about 45, or about 50 nucleotides in length. In some embodiments, the cross-linking oligonucleotide herein may include a barcode.
[0202] In a further aspect of a method comprising the MNase digestion step provided herein, the method may include obtaining at least some sequence on each side of the junction to generate a first read pair. For example, the method may include obtaining at least about 50 bp, at least about 100 bp, at least about 150 bp, at least about 200 bp, at least about 250 bp, or at least about 300 bp of sequence on each side of the junction to generate a first read pair.
[0203] In an additional aspect of a method comprising the MNase digestion step provided herein, the method can include mapping a first read pair to a set of contigs and determining a path through the set of contigs representing order and / or orientation to the genome.
[0204] In a further aspect of the method comprising the MNase digestion step herein, the method may include mapping a first read pair to a set of contigs and determining the presence of structural variants or loss of heterozygosity in a stabilized biological sample from the set of contigs.
[0205] In an additional aspect of the method comprising the MNase digestion step herein, the method may include mapping a first read pair to a set of contigs and assigning a phase to variants in the set of contigs.
[0206] In a further aspect of the method comprising the MNase digestion step herein, the method may include mapping a first read pair to a set of contigs, determining the presence of variants in the set of contigs from the set of contigs, and performing a step selected from one or more of (1) identifying a disease stage, prognosis, or treatment regimen for a stabilized biological sample, (2) selecting a drug based on the presence of the variant, or (3) identifying drug efficacy for a stabilized biological sample.
[0207] Improved methods for HiChIP, HiChIRP, and methylHiC HiChIP is an approach that combines the HiC method with the chromatin immunoprecipitation method, enabling targeted analysis of interactions involving one or more proteins of interest. Nucleic acids ligated in proximity can be prepared, and the targeted regions can be immunoprecipitated for further analysis. HiChIRP, a related approach, uses chromatin isolation by RNA purification (ChIRP) enrichment in combination with the HiC method to enable interrogation of RNAs such as the scaffolding function of long non-coding RNAs (lncRNAs). Methyl-HiC combines methylation analysis with the HiC method to enable simultaneous capture of chromosome conformation and DNA methylome information. Methyl-HiC reveals coordinated DNA methylation states between distal genomic segments that are spatially proximal in the nucleus, depicts the heterogeneity of both chromatin architecture and DNA methylome in a mixed population, and enables simultaneous characterization of chromatin organization and epigenome specific to cell types in complex tissues. These and other methods can be improved by use of the techniques of the present disclosure, including but not limited to size selection steps, surface binding steps (e.g., binding to beads such as SPRI beads), use of crosslinking oligonucleotides to perform proximity ligation, use of recombinases to perform proximity ligation, etc.
[0208] In additional aspects, improved methods for HiChIP, HiChIRP, and methylHiC are provided herein, which methods can include, for example, obtaining a stabilized biological sample comprising nucleic acid molecules complexed with at least one nucleic acid-binding protein by immunoprecipitation of nucleic acids bound to a nucleic acid-binding protein or by immunoprecipitation of methylated nucleic acids; contacting the stabilized biological sample with DNase to cleave the nucleic acid molecules into a plurality of segments; attaching a first segment and a second segment of the plurality of segments at a junction; and size selecting the plurality of segments to obtain a plurality of selected segments. Alternatively, or in combination, the methods herein can include, for example, obtaining a stabilized biological sample comprising nucleic acid molecules complexed with at least one nucleic acid-binding protein by immunoprecipitation of nucleic acids bound to a nucleic acid-binding protein or by immunoprecipitation of methylated nucleic acids; contacting the stabilized biological sample with micrococcal nuclease (MNase) to cleave the nucleic acid molecules into a plurality of segments; and attaching a first segment and a second segment of the plurality of segments at a junction.
[0209] In some aspects of the improved methods for HiChIP, HiChIRP, and methylHiC herein, the stabilized biological sample can comprise intact cells and / or intact nuclei. Optionally, the stabilized biological sample can comprise stabilized intact cells. Alternatively, or in combination, the stabilized biological sample can comprise stabilized intact nuclei. Optionally, a step of contacting a sample of stabilized intact cells or nuclei with DNase can be performed prior to lysis of the intact cells or nuclei. Optionally, the cells and / or nuclei can be lysed prior to the step of attaching a first segment and a second segment of the plurality of segments at a junction.
[0210] In another aspect, the improved methods for HiChIP, HiChIRP, and methylHiC herein can include subjecting a plurality of segments to size selection to obtain a plurality of selected segments. Optionally, the plurality of selected segments can be about 145 to about 600 bp. Optionally, the plurality of selected segments can be about 100 to about 2500 bp. Optionally, the plurality of selected segments can be about 100 to about 600 bp. Optionally, the plurality of selected segments can be about 600 to about 2500 bp. Optionally, the plurality of selected segments can be about 100 bp to about 600 bp, about 100 bp to about 700 bp, about 100 bp to about 800 bp, about 100 bp to about 900 bp, about 100 bp to about 1000 bp, about 100 to about 1100 bp, about 100 bp to about 1200 bp, about 100 bp to about 1300 bp, about 100 bp to about 1400 bp, about 100 bp to about 1500 bp, about 100 bp to about 1600 bp, about 100 bp to about 1700 bp, about 100 bp to about 1800 bp, about 100 bp to about 1900 bp, about 100 bp to about 2000 bp, about 100 bp to about 2100 bp, about 100 bp to about 2200 bp, about 100 bp to about 2300 bp, about 100 bp to about 2400 bp, or about 100 bp to about 2500 bp.
[0211] In another aspect of the method, which includes improved methods for HiChIP, HiChIRP, and methylHiC herein, the method may further include preparing a sequencing library from a plurality of segments prior to the size selection step. In some embodiments, the method may further include subjecting the sequencing library to size selection to obtain a size-selected library. Optionally, the size-selected library may be in the size range of about 350 bp to about 1000 bp. Optionally, the size-selected library may be in the size range of about 100 bp to about 2500 bp, for example, about 100 bp to about 350 bp, about 350 bp to about 500 bp, about 500 bp to about 1000 bp, about 1000 to about 1500 bp, about 2000 to about 2500 bp, about 350 bp to about 1000 bp, about 350 bp to about 1500 bp, about 350 bp to about 2000 bp, about 350 bp to about 2500 bp, about 500 bp to about 1500 bp, about 500 bp to about 2000 bp, about 500 bp to about 3500 bp, about 1000 bp to about 1500 bp, about 1000 bp to about 2000 bp, about 1000 bp to about 2500 bp, about 1500 bp to about 2000 bp, about 1500 bp to about 2500 bp, or about 2000 bp to about 2500 bp.
[0212] The size selection utilized in the method, which includes improved methods for HiChIP, HiChIRP, and methylHiC herein, can be performed by gel electrophoresis, capillary electrophoresis, size selection beads, gel filtration columns, combinations thereof, or any other suitable method.
[0213] In another aspect, methods including improved methods for HiChIP, HiChIRP, and methylHiC may further include the step of further analyzing a plurality of selected segments to obtain QC values. In some cases, the QC value may be selected from a chromatin digestion efficiency (CDE) and a chromatin digestion index (CDI). The CDE can be calculated as the proportion of segments having a desired length. For example, in some cases, the CDE can be calculated as the proportion of segments sized 100-2500 bp before size selection. In some cases, when the CDE value is at least 65%, the sample may be selected for further analysis. In some cases, when the CDE value is at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, or at least about 95%, the sample may be selected for further analysis.
[0214] CDI can be calculated as the ratio of the number of segments of mononucleosome size to the number of segments of dinucleosome size before size selection. For example, CDI may be calculated as the logarithm of the ratio of fragments having a size of 600 - 2500 bp to fragments having a size of 100 - 600 bp. In some cases, when the CDI value exceeds -1.5 and is less than 1, the sample may be selected for further analysis. In some cases, when the CDI value exceeds approximately -2 and is less than approximately 1.5, exceeds approximately -1.9 and is less than approximately 1.5, exceeds approximately -1.8 and is less than approximately 1.5, exceeds approximately -1.7 and is less than approximately 1.5, exceeds approximately -1.6 and is less than approximately 1.5, exceeds approximately -1.5 and is less than approximately 1.5, exceeds approximately -1.4 and is less than approximately 1.5, exceeds approximately -1.3 and is less than approximately 1.5, exceeds approximately -1.2 and is less than approximately 1.5, exceeds approximately -1.1 and is less than approximately 1.5, exceeds approximately -2 and is less than approximately 1.5, exceeds approximately -1 and is less than approximately 1.5, exceeds approximately -0.9 and is less than approximately 1.5, exceeds approximately -0.8 and is less than approximately 1.5, exceeds approximately -0.7 and is less than approximately 1.5, exceeds approximately -0.6 and is less than approximately 1.5, exceeds approximately -0.5 and is less than approximately 1.5, exceeds approximately -2 and is less than approximately 1.4, exceeds approximately -2 and is less than approximately 1.3, exceeds approximately -2 and is less than approximately 1.2, exceeds approximately -2 and is less than approximately 1.1, exceeds approximately -2 and is less than approximately 1, exceeds approximately -2 and is less than approximately 0.9, exceeds approximately -2 and is less than approximately 0.8, exceeds approximately -2 and is less than approximately 0.7, exceeds approximately -2 and is less than approximately 0.6, exceeds approximately -2 and is less than approximately 0.5, the sample may be selected for further analysis.
[0215] In another aspect, the improved methods for HiChIP, HiChIRP, and methyl HiC herein can be performed on small samples that contain little or a small amount of nucleic acid. In some cases, the stabilized biological sample may contain fewer than 3,000,000 cells. In some cases, the stabilized biological sample may contain fewer than 2,000,000 cells. In some cases, the stabilized biological sample may contain fewer than 1,000,000 cells. In some cases, the stabilized biological sample may contain fewer than 500,000 cells. In some cases, the stabilized biological sample may contain fewer than 400,000 cells. In some cases, the stabilized biological sample may contain fewer than 300,000 cells. In some cases, the stabilized biological sample may contain fewer than 200,000 cells. In some cases, the stabilized biological sample may contain fewer than 100,000 cells. In some cases, the stabilized biological sample may contain less than 10 μg of DNA. In some cases, the stabilized biological sample may contain less than 9 μg of DNA. In some cases, the stabilized biological sample may contain less than 8 μg of DNA. In some cases, the stabilized biological sample may contain less than 7 μg of DNA. In some cases, the stabilized biological sample may contain less than 6 μg of DNA. In some cases, the stabilized biological sample may contain less than 5 μg of DNA. In some cases, the stabilized biological sample may contain less than 4 μg of DNA. In some cases, the stabilized biological sample may contain less than 3 μg of DNA. In some cases, the stabilized biological sample may contain less than 2 μg of DNA. In some cases, the stabilized biological sample may contain less than 1 μg of DNA. In some cases, the stabilized biological sample may contain less than 0.5 μg of DNA.
[0216] In another aspect, methods including improved methods for HiChIP, HiChIRP, and methylHiC herein may be performed on individual cells or single cells. For example, the methods herein may be performed on cells distributed in individual compartments. Examples of exemplary compartments include, but are not limited to, wells, droplets in an emulsion, or surface positions (e.g., array spots, beads, etc.) containing individual patches of linker molecules of different sequences as described elsewhere herein. Additional compartments are contemplated and are consistent with the methods, compositions, and systems disclosed herein.
[0217] In an additional aspect, a stabilized biological sample used in a method including improved methods for HiChIP, HiChIRP, and methylHiC herein can be treated with a nuclease such as DNase to create DNA fragments. In some cases, the DNase may not be sequence-specific. In some cases, the DNase may be active against both single-stranded and double-stranded DNA. In some cases, the DNase may be specific for double-stranded DNA. In some cases, the DNase may preferentially cleave double-stranded DNA. In some cases, the DNase may be specific for single-stranded DNA. In some cases, the DNase may preferentially cleave single-stranded DNA. In some cases, the DNase may be DNase I. In some cases, the DNase may be DNase II. In some cases, the DNase may be selected from one or more of DNase I and DNase II. In some cases, the DNase may be micrococcal nuclease. In some cases, the DNase may be selected from one or more of DNase I, DNase II, and micrococcal nuclease. In some cases, the DNase may be linked or fused to an immunoglobulin-binding protein or fragment thereof, such as protein A, protein G, protein A / G, or protein L. Other suitable nucleases are also within the scope of the present disclosure.
[0218] In an additional aspect, a stabilized biological sample used in a method including an improved method for HiChIP, HiChIRP, and methylHiC herein may be treated with a crosslinking agent. In some cases, the crosslinking agent may be a chemical fixative. In some cases, the chemical fixative may include formaldehyde and have a spacer arm length of about 2.3 - 2.7 angstroms (A). In some cases, the chemical fixative may include a crosslinking agent having a long spacer arm length. For example, the crosslinking agent can have a spacer length of at least about 3A, 4A, 5A, 6A, 7A, 8A, 9A, 10A, 11A, 12A, 13A, 14A, 15A, 16A, 17A, 18A, 19A, or 20A. The chemical fixative can include ethylene glycol bis(succinimidyl succinate) (EGS), which has a spacer arm with a length of about 16.1A. The chemical fixative can include disuccinimidyl glutarate (DSG), which has a spacer arm with a length of about 7.7A. In some cases, the chemical fixative includes formaldehyde and EGS, formaldehyde and DSG, or formaldehyde, EGS, and DSG. In some cases where multiple chemical fixatives are used, each chemical fixative is used sequentially; in other cases, some or all of the multiple chemical fixatives are applied to the sample simultaneously. The use of a crosslinking agent with a long spacer arm can increase the fraction of read pairs having a large (e.g., >1kb) separation distance between read pairs. For example, FIG. 7 shows a comparison between a library resulting from crosslinking with formaldehyde only (digested with both DNase and MNase) and a library resulting from crosslinking with formaldehyde and DSG or EGS. DSG has NHS ester reactive groups at both ends and can be reactive towards amino groups (e.g., primary amines). DSG is membrane permeable and enables intracellular crosslinking. DSG can increase crosslinking efficiency compared to disuccinimidyl suberate (DSS) in some applications. EGS has NHS ester reactive groups at both ends and can be reactive towards amino groups (e.g., primary amines). EGS is membrane permeable and enables intracellular crosslinking.For example, EGS cross-linking can be reversed by treating with hydroxylamine at pH 8.5 for 3 to 6 hours; in one example, lactate dehydrogenase retained 60% of its activity after reversible cross-linking with EGS. Optionally, the chemical fixative may contain psoralen. Optionally, the cross-linking agent may be ultraviolet light. Optionally, the stabilized biological sample may be a cross-linked paraffin-embedded tissue sample.
[0219] In additional aspects, methods including improved methods for HiChIP, HiChIRP, and methyl HiC herein may include attaching a first segment and a second segment of a plurality of segments at a junction. Optionally, the attaching step may include filling sticky ends and ligating blunt ends using biotin-tagged nucleotides. Optionally, the attaching step may include contacting at least the first segment and the second segment with a cross-linking oligonucleotide. Optionally, the attaching step may include contacting at least the first segment and the second segment with a barcode. In some embodiments, the cross-linking oligonucleotides herein may be at least about 5 nucleotides to about 50 nucleotides in length. In some embodiments, the cross-linking oligonucleotides herein may be about 15 to about 18 nucleotides in length. In some embodiments, the cross-linking oligonucleotide may be about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, about 20, about 25, about 30, about 35, about 40, about 45, or about 50 nucleotides in length. In some embodiments, the cross-linking oligonucleotides herein may include a barcode.
[0220] In additional aspects, methods including improved methods for HiChIP, HiChIRP, and methyl HiC herein may not include a shearing step.
[0221] In a further aspect of the methods, including the improved methods for HiChIP, HiChIRP, and methylHiC herein, the method may include obtaining at least some sequence on each side of the junction to generate a first read pair. For example, the method may include obtaining a sequence of at least about 50 bp, at least about 100 bp, at least about 150 bp, at least about 200 bp, at least about 250 bp, or at least about 300 bp on each side of the junction to generate a first read pair.
[0222] In an additional aspect of the methods, including the improved methods for HiChIP, HiChIRP, and methylHiC herein, the method may include mapping a first read pair to a set of contigs and determining a path through the set of contigs to the genome that represents the order and / or orientation.
[0223] In a further aspect of the methods, including the improved methods for HiChIP, HiChIRP, and methylHiC herein, the method may include mapping a first read pair to a set of contigs and determining the presence of structural variants or loss of heterozygosity in a stabilized biological sample from the set of contigs.
[0224] In an additional aspect of the methods, including whole cell or whole nucleus digestion herein, the method may include mapping a first read pair to a set of contigs and phasing variants in the set of contigs.
[0225] In a further aspect of the method, including the improved methods for HiChIP, HiChIRP, and methylHiC in this specification, the method comprises mapping a first read pair to a set of contigs, determining the presence of variants in the set of contigs from the set of contigs, and (1) identifying a disease stage, prognosis, or treatment regimen for a stabilized biological sample, (2) selecting a drug based on the presence of the variant, or (3) identifying drug efficacy for a stabilized biological sample, and performing a step selected from one or more of the above.
[0226] Generating long-range read pairs The present disclosure provides methods for generating ultra-long-range read pairs and utilizing that data for all of the above-mentioned pursuits of progress. In some embodiments, the present disclosure provides methods for generating a very close and very accurate human genome assembly with as few as ~300 million read pairs. In other embodiments, the present disclosure provides methods for phasing over 90% of the heterozygous variants in the human genome with an accuracy of over 99%. Further, the range of read pairs generated by the present disclosure can be extended to cover much larger genomic distances. The assembly is generated from a standard shotgun library in addition to the ultra-long-range read pair library. In yet other embodiments, the present disclosure provides software that can utilize both of these sets of sequencing data. The phased variants are generated in a single long-range read pair library, the reads from which are mapped to a reference genome and used to assign the variant to one of the two parental chromosomes of an individual. Finally, the present disclosure provides for the extraction of even larger DNA fragments using known techniques to generate exceptionally long reads.
[0227] The mechanisms by which these repeats interfere with assembly and alignment processes are rather simple and ultimately the result of ambiguity. In the case of large repeat regions, the difficulty can be one of span. If the read or read pair is not long enough to span the repeat region, it may not be possible to connect the regions bounding the repeat elements with confidence. In the case of smaller repeat elements, the problem can primarily be one of placement. When regions are flanked by two repeat elements common to the genome, it becomes difficult, if not impossible, to determine their exact placement because the flanking elements are similar to all the others in their class. In either case, what makes identification, and thus the placement of specific repeats, difficult is the lack of identifying information within the repeats. What is needed is the ability to experimentally establish connections between unique segments bounded or separated by repeat regions.
[0228] The methods of the present disclosure can advance the field of genomics by overcoming the substantial barriers posed by these repetitive regions, and thereby enable significant progress in many domains of genomic analysis. To perform de novo assembly with previous techniques, one must either resort to assemblies fragmented into many small scaffolds to generate a more continuous assembly, or spend a significant amount of time and resources creating large-insert libraries, or using other approaches. Such approaches may involve obtaining very deep sequencing coverage, constructing BAC or fosmid libraries, optical mapping, or some combination of these and / or other techniques. Due to the huge resource and time requirements, such approaches become inaccessible to most small-scale labs, hindering research on non-model organisms. Since the methods described herein can generate very long-range read pairs, de novo assembly can be achieved with a single sequencing run. This will reduce the assembly cost by orders of magnitude and shorten the time required from months or years to weeks. In some cases, the methods disclosed herein can generate multiple read pairs in less than 14 days, less than 13 days, less than 12 days, less than 11 days, less than 10 days, less than 9 days, less than 8 days, less than 7 days, less than 6 days, less than 5 days, less than 4 days, or within the range between any two of the previously defined periods. For example, the method may be able to generate multiple read pairs in about 10 to 14 days. Constructing genomes will become routine even for the most niche organisms, phylogenetic analysis will no longer be hampered by the lack of comparatives, and projects such as Genome10k will be achievable.
[0229] Similarly, structural and phasing analyses for medical purposes continue to be more challenging. There is surprising heterogeneity even among cancers, among individuals with the same type of cancer, or even within the same tumor. Extracting the cause from the resulting effects requires very high accuracy and throughput at low cost per sample. In the area of personalized medicine, one of the absolute criteria for genomic care is a sequenced genome with all variants thoroughly characterized and phased, including large and small structural rearrangements and novel mutations. Achieving this with previous technologies required similar effort to de novo assembly, which is currently too costly and laborious as a routine medical procedure. The disclosed method can rapidly generate a complete and accurate genome at low cost, thereby creating many much-needed capabilities in the study and treatment of human diseases.
[0230] By applying the methods disclosed herein to phasing, the convenience of statistical approaches can be combined with the accuracy of family analysis, providing savings in - money, labor, and samples - over using either method alone. De novo variant phasing is a highly desirable phasing analysis that was prohibited by previous techniques and can be readily performed using the methods disclosed herein. This is particularly important because the majority of human mutations are rare (minor allele frequency < 5%). Phasing information is valuable for population genetic studies that gain significant advantages from networks of highly connected haplotypes (collections of variants assigned to a single chromosome) compared to unlinked genotypes. Haplotype information can enable higher-resolution studies of population size, migration, and historical changes in exchange between subpopulations, and can enable tracing specific variants back to specific parents and grandparents. This, in turn, reveals the genetic transmission of variants associated with disease and the interactions between variants when grouped together in a single individual. The methods of the present disclosure may ultimately enable the preparation, sequencing, and analysis of ultra-long-range read pair (XLRP) libraries.
[0231] In some embodiments of the present disclosure, a tissue or DNA sample from a subject may be provided, and the method may return an assembled genome, an alignment with called variants (including large structural variants), phased variant calls, or any additional analysis. In other embodiments, the methods disclosed herein may provide an XLRP library directly to an individual.
[0232] Ultra-long-range read pair In various embodiments of the present disclosure, the methods disclosed herein can generate ultra-long-range read pairs that are separated over long distances. The upper limit of this distance can be improved by the ability to collect large-sized DNA samples. In some cases, the read pairs can span genomic distances of up to 50, 60, 70, 80, 90, 100, 125, 150, 175, 200, 225, 250, 300, 400, 500, 600, 700, 800, 900, 1000, 1500, 2000, 2500, 3000, 4000, 5000 kbp, or more. In some examples, the read pairs can span genomic distances of up to 500 kbp. In other examples, the read pairs can span genomic distances of up to 2000 kbp. The methods disclosed herein can be integrated and constructed based on standard techniques in molecular biology and are well-suited for increasing efficiency, specificity, and genomic coverage. In some cases, the read pairs can be generated in less than about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 60, or 90 days. In some examples, the read pairs can be generated in less than about 14 days. In further examples, the read pairs can be generated in less than about 10 days. In some cases, the methods of the present disclosure can provide read pairs that exceed about 5%, about 10%, about 15%, about 20%, about 30%, about 40%, about 50%, about 60%, about 70%, about 80%, about 90%, about 95%, about 99%, or about 100%, which have at least about 50%, about 60%, about 70%, about 80%, about 90%, about 95%, about 99%, or about 100% accuracy when correctly ordering and / or orienting multiple contigs. For example, the method can provide about 90 - 100% accuracy when correctly ordering and / or orienting multiple contigs.
[0233] In other embodiments, the methods disclosed herein can be used with currently utilized sequencing technologies. For example, this method can be used in combination with well-tested and / or widely deployed sequencing equipment. In further embodiments, the methods disclosed herein can be used with techniques and methods derived from currently utilized sequencing technologies.
[0234] The methods of the present disclosure dramatically simplify de novo genome assembly for a wide range of organisms. Using previous techniques, such assemblies are currently limited by short inserts in economical mate-pair libraries. It may be possible to generate read pairs at genomic distances up to 40 - 50 kbp accessible with fosmids, but on the other hand, these are expensive and difficult to handle and are too short to span the longest repeat stretches, including those within centromeres, which can range in size from 300 kbp to 5 Mbp in humans. The methods disclosed herein can provide read pairs that can span large distances (e.g., megabases or longer), and thus overcome difficulties regarding the completeness of these scaffolds. Accordingly, generating chromosome-level assemblies can become routine by utilizing the methods of the present disclosure. More laborious means for assembly, which currently take an unthinkable amount of time and money for laboratories and prohibit large genomic catalogs, may become unnecessary and resources may be freed up for more meaningful analysis. Similarly, obtaining long-range phasing information can provide an unthinkable additional power for studies related to population genomics, phylogenetics, and disease. The methods disclosed herein enable a large number of individual accurate phasings, thus expanding the breadth and depth of our ability to investigate genomes at the population and deep-time levels.
[0235] In the field of personalized medicine, the XLRP read pairs generated from the methods disclosed herein represent a significant advancement towards accurate, low-cost, phased, and rapidly generated personalized genomes. Current methods are insufficient in their ability to phase variants over long distances, thereby impeding the characterization of the phenotypic impact of compound heterozygous genotypes. Additionally, substantial structural variants of interest for genomic diseases are large in size compared to the reads and read pair insertions used to study them, making them difficult to accurately identify and characterize with current technologies. Read pairs spanning from tens of kilobases to over megabases can help mitigate this difficulty, thereby enabling highly parallelized personalized analysis of structural displacements.
[0236] Basic evolutionary and biomedical research has been propelled by technological advancements in high-throughput sequencing. Previously, the entirety of genome sequencing and assembly was the provenance of large-scale genome sequencing centers, but now commercially available sequencers are inexpensive enough that most research universities have one or more of these machines. Currently, it is relatively inexpensive to generate large amounts of DNA sequence data. However, generating high-quality and highly contiguous genome sequences remains difficult both theoretically and in practice. Furthermore, since most organisms of interest, including humans, are diploid, each individual has two haploid copies of the genome. At heterozygous sites (e.g., when the allele given by the mother differs from the allele given by the father), it is difficult to know which set of alleles is derived from which parent (known as haplotype phasing). This information can be used to perform many evolutionary and biomedical studies, such as disease and trait-related research.
[0237] In various embodiments, the present disclosure provides methods for genome assembly that combine techniques for DNA preparation using paired-end sequencing for high-throughput discovery of short-, medium-, and long-range connections within a given genome. The present disclosure further provides methods for using these connections to assist in genome assembly for haplotype phasing and / or metagenomic studies. While the methods presented herein can be used to determine the assembly of a subject's genome, the methods presented herein can be used to determine the assembly of a portion of a subject's genome, such as a chromosome, or the assembly of chromatin of a subject of various lengths.
[0238] In some embodiments, the present disclosure provides one or more methods disclosed herein, the method comprising generating a plurality of contigs from sequencing a fragment of target DNA obtained from a subject. A long stretch of target DNA can be fragmented by cutting the DNA with one or more nucleases (such as DNase I, DNase II, micrococcal nuclease). The resulting fragments can be sequenced using a high-throughput sequencing method to obtain a plurality of sequencing reads. Examples of high-throughput sequencing methods that can be used with the methods of the present disclosure include, but are not limited to, the 454 pyrosequencing method developed by Roche Diagnostics, the "cluster" sequencing method developed by Illumina, the SOLiD and ion semiconductor sequencing methods developed by Life Technologies, and the DNA nanoball sequencing method developed by Complete Genomics. Subsequently, overlapping ends of different sequencing reads can be assembled to form contigs. Alternatively, the fragmented target DNA can be cloned into a vector. Subsequently, cells or organisms are transfected with the DNA vector to form a library. After replicating the transfected cells or organisms, the vector is isolated and sequenced to generate a plurality of sequencing reads. Subsequently, overlapping ends of different sequencing reads can be assembled to form contigs.
[0239] Genome assembly, especially using high-throughput sequencing technology, can sometimes cause problems. Often, the assembly consists of thousands or tens of thousands of short contigs. The order and orientation of these contigs are generally unknown, limiting the usefulness of the genome assembly. There are techniques for ordering and orienting these scaffolds, but they are generally expensive, labor-intensive, and often fail to discover very long-range interactions.
[0240] Samples containing target DNA used to generate contigs can be obtained from a subject by any number of means including collecting body fluids (e.g., blood, urine, serum, lymph, saliva, buccal swabs, anal and vaginal secretions, sweat, and semen), taking tissue, or collecting cells / organisms. The resulting samples may be composed of a single type of cell / organism or may be composed of multiple types of cells / organisms. DNA can be extracted and prepared from the subject's samples. For example, the samples may be treated to lyse cells containing polynucleotides using known lysis buffers, sonication techniques, electroporation, etc. The target DNA can be further purified and contaminants such as proteins may be removed by using alcohol extraction, cesium gradients, and / or column chromatography.
[0241] In other embodiments of the disclosure, methods are provided for extracting DNA of very high molecular weight. In some cases, data from an XLRP library can be improved by increasing the fragment size of the input DNA. In some examples, by extracting megabase-sized DNA fragments from cells, it is possible to generate read pairs separated by megabases in the genome. In some cases, the generated read pairs can provide sequence information spanning a span of about 10 kB, about 50 kB, about 100 kB, about 200 kB, about 500 kB, about 1 Mb, about 2 Mb, about 5 Mb, about 10 Mb, greater than about 100 Mb. In some examples, the read pairs can provide sequence information spanning a span greater than about 500 kB. In further examples, the read pairs can provide sequence information spanning a span greater than about 2 Mb. In some cases, very high molecular weight DNA can be extracted by very gentle cell lysis (Teague, B. et al. (2010) Proc. Nat. Acad. Sci. USA 107(24), 10848-53), and agarose plugs (Schwartz, D.C., & Cantor, C.R. (1984) Cell, 37(1), 67-75). In other cases, very high molecular weight DNA can be extracted using commercially available machines capable of purifying DNA molecules up to megabase lengths.
[0242] Investigation of the physical layout of chromosomes In various embodiments, the present disclosure provides one or more methods disclosed herein, the method including the step of investigating the physical layout of chromosomes in a living cell. Examples of techniques for investigating the physical layout of chromosomes through sequencing include techniques of the "C" family such as chromosome conformation capture ("3C"), circular chromosome conformation capture ("4C"), carbon copy chromosome capture ("5C"), and Hi-C based methods, as well as ChIP-based methods such as ChIP-loop, ChIA-PET, and HiChIP. These techniques utilize the fixation of chromatin in living cells to fix the spatial relationships in the nucleus. Through subsequent processing and sequencing of the products, researchers are able to recover an approximate matrix of associations in regions of the genome. In further analysis, these associations can be used to create a three-dimensional geometric map of the chromosomes when the chromosomes are physically arranged in the nucleus of the living body. Such techniques explain the discrete spatial organization of chromosomes in living cells and provide an accurate view of the functional interactions at chromosomal loci. One problem that has plagued these functional studies is the presence of non-specific interactions, associations, which were present in data that was only due to chromosomal proximity. In the present disclosure, these non-specific intra-chromosomal interactions are captured by the methods presented herein to provide valuable information regarding the assembly.
[0243] In some embodiments, the intra-chromosomal interactions have a correlation with chromosomal connectivity. In some cases, the intra-chromosomal data can assist the genome assembly. In some cases, chromatin is reconstructed in vitro. This is because chromatin - specifically histones, the major protein component of chromatin - is important for the fixation underlying the most common "C" family of techniques for detecting the conformation and structure of chromatin through the sequencing of 3C, 4C, 5C, and Hi-C. Chromatin is very non-specific with respect to sequence and usually assembles uniformly across the genome. In some cases, the genomes of species that do not use chromatin can be assembled on the reconstituted chromatin, thereby expanding the horizon for the disclosure to all domains of life.
[0244] Chromatin conformation capture techniques are summarized. Briefly, crosslinks are created between genomic regions that are physically proximal. Crosslinks between DNA molecules, such as genomic DNA, and proteins (such as histones) within chromatin can be accomplished according to appropriate methods described in more detail elsewhere in this specification or known in the art. In some cases, two or more nucleotide sequences can be crosslinked via a protein bound to one or more nucleotide sequences. One approach is to expose chromatin to ultraviolet irradiation (Gilmour et al., Proc. Nat’l. Acad. Sci. USA 81:4275-4279, 1984). Crosslinking of polynucleotide segments can also be performed using other approaches such as chemical or physical (e.g., optical) crosslinking. Suitable chemical crosslinking agents include, but are not limited to, formaldehyde and psoralen (Solomon et al., Proc. Natl. Acad. Sci. USA 82:6470-6474, 1985; Solomon et al., Cell 53:937-947, 1988). For example, crosslinking can be accomplished by adding 2% formaldehyde to a mixture containing DNA molecules and chromatin proteins. Other examples of agents that can be used to crosslink DNA include, but are not limited to, UV light, mitomycin C, nitrogen mustard, melphalan, 1,3-butadiene diepoxide, cis-diamminedichloroplatinum(II), and cyclophosphamide. Preferably, the crosslinking agent forms crosslinks that bridge relatively short distances, such as about 2 Å, thereby selecting for intimate interactions that can be reversed.
[0245] In some embodiments, the DNA molecule may be immunoprecipitated before or after cross-linking. Optionally, the DNA molecule may be fragmented. The fragments may be contacted with a contact partner, such as an antibody that specifically recognizes and binds to acetylated histone, e.g., H3. Examples of such antibodies include, but are not limited to, anti-acetylated histone H3 available from Upstate Biotechnology, Lake Placid, N.Y. The polynucleotide from the immunoprecipitate can then be collected from the immunoprecipitate. Before fragmenting the chromatin, the acetylated histone can cross-link with the adjacent polynucleotide sequence. Thereafter, the mixture is processed to fractionate the polynucleotides in the mixture. The fractionation techniques herein include the use of deoxyribonuclease (DNase) enzymes. DNases suitable for the methods herein include, but are not limited to, DNase I, DNase II, and micrococcal nuclease. The resulting fragments can be of various sizes. The resulting fragments may also include single-stranded overhangs at the 5' or 3' ends.
[0246] In some embodiments, fragments of about 145 bp to about 600 bp can be obtained. Alternatively, fragments of about 100 bp to about 2500 bp, about 100 bp to about 600 bp, or about 600 bp to about 2500 bp can be obtained. The sample can be prepared for sequencing of the cross-linked ligation sequence segments. Optionally, for example, a single short stretch of polynucleotide can be created by ligating two sequence segments cross-linked within the molecule. The sequence information may be obtained from the sample using any suitable sequencing technique described in more detail herein, or other suitable methods such as high-throughput sequencing methods. For example, the ligation product can be subjected to paired-end sequencing to obtain sequence information from each end of the fragment. Pairs of sequence segments can be represented by the obtained sequence information, associating haplotyping information over the linear distance separating the two sequence segments along the polynucleotide.
[0247] One feature of the data generated by Hi-C is that most read pairs are found to be nearly linearly proximal when remapped to the genome. That is, most read pairs are found to be close to each other in the genome. In the resulting dataset, the probability of intrachromosomal contacts is on average much higher than that of interchromosomal contacts, as expected when chromosomes occupy distinct regions. Furthermore, the probability of interaction decays rapidly with linear distance, and loci separated by >200 Mb on the same chromosome are more likely to interact than loci on different chromosomes. This "background" of short- and medium-range intrachromosomal contacts is background noise that is excluded using Hi-C analysis when detecting long-range intrachromosomal contacts, particularly interchromosomal contacts.
[0248] In particular, Hi-C experiments in eukaryotes exhibit two canonical interaction patterns in addition to species-specific and cell-type-specific chromatin interactions. One pattern, distance-dependent decay (DDD), is the general tendency of decay in interaction frequency as a function of genomic distance. The second pattern, the cis-trans ratio (CTR), shows that the interaction frequency between loci located on the same chromosome is significantly higher, even when separated by sequences of dozens of megabases, compared to loci on different chromosomes. These patterns may reflect general macromolecular dynamics, where the likelihood of proximal loci interacting randomly, as well as phenomena such as the formation of chromosomal regions and the tendency of interphase chromosomes to occupy distinct volumes in the nucleus with little mixing, are more likely features of specific nuclear organization. The exact details of these two patterns may vary between species, cell types, and cell states, but they are ubiquitous and prominent. Because these patterns are very powerful and consistent, they are used to evaluate the quality of experiments and are typically normalized from the data to reveal detailed interactions. However, in the methods disclosed herein, genome assemblies can utilize the three-dimensional structure of the genome. The features that make the canonical Hi-C interaction patterns an obstacle to the analysis of specific loop interactions, namely their ubiquity, strength, and consistency, can be used as powerful tools for inferring the genomic positions of contigs.
[0249] In certain implementations, consideration of the physical distance between intrachromosomal read pairs reveals several useful features of data regarding genome assembly. First, short-range interactions are more common than long-range interactions. That is, each read of a read pair is more likely to mate with a region closer in the actual genome than with a region farther away. Second, there are long tails of medium-range and long-range interactions. That is, read pairs convey information about the intrachromosomal sequence at distances of kilobases (kB) or even megabases (Mb). For example, a read pair can provide sequence information spanning a span of about 10 kB, about 50 kB, about 100 kB, about 200 kB, about 500 kB, about 1 Mb, about 2 Mb, about 5 Mb, about 10 Mb, or more than about 100 Mb. These features of the data indicate that regions of the genome that are proximal on the same chromosome are more likely to be physically proximal - which is a predicted result because they are chemically linked to each other through the DNA backbone. Genome-wide chromatin interaction datasets, such as those generated by Hi-C, were hypothesized to provide long-range information regarding the grouping and linear organization of sequences along entire chromosomes.
[0250] The experimental method for Hi-C is simple and relatively low-cost, but current protocols for genome assembly and haplotyping require 3 to 5 million cells, and a fairly large amount of material may not be obtainable, especially from specific human patient samples. In contrast, the methods disclosed herein include methods that enable accurate and predictive results for genotyping assembly, haplotype phasing, and metagenomics with significantly less material from cells. For example, DNA of about 0.1 μg, about 0.2 μg, about 0.3 μg, about 0.4 μg, about 0.5 μg, about 0.6 μg, about 0.7 μg, about 0.8 μg, about 0.9 μg, about 1.0 μg, about 1.2 μg, about 1.4 μg, about 1.6 μg, about 1.8 μg, about 2.0 μg, about 2.5 μg, about 3.0 μg, about 3.5 μg, about 4.0 μg, about 4.5 μg, about 5.0 μg, about 6.0 μg, about 7.0 μg, about 8.0 μg, about 9.0 μg, about 10 μg, about 15 μg, about 20 μg, about 30 μg, about 40 μg, about 50 μg, about 60 μg, about 70 μg, about 80 μg, about 90 μg, about 100 μg, about 150 μg, about 200 μg, about 300 μg, about 400 μg, about 500 μg, about 600 μg, about 700 μg, about 800 μg, about 900 μg, about 1000 μg, about 1200 μg, about 1400 μg, about 1600 μg, about 1800 μg, about 2000 μg, about 2200 μg, about 2400 μg, about 2600 μg, about 2800 μg, about 3000 μg, about 3200 μg, about 3400 μg, about 3600 μg, about 3800 μg, about 4000 μg, about 4200 μg, about 4400 μg, about 4600 μg, about 4800 μg, about 5000 μg, about 5200 μg, about 5400 μg, about 5600 μg, about 5800 μg, about 6000 μg, about 6200 μg, about 6400 μg, about 6600 μg, about 6800 μg, about 7000 μg, about 7200 μg, about 7400 μg, about 7600 μg, about 7800 μg, about 8000 μg, about 8200 μg, about 8400 μg, about 8600 μg, about 8800 μg, about 9000 μg, about 9200 μg, about 9400 μg, about 9600 μg, about 9800 μg, about 10,000 μg can be used in the methods disclosed herein.In some examples, DNA used in the methods disclosed herein can be extracted from cells that are less than about 3,000,000, about 2,500,000, about 2,000,000, about 1,500,000, about 1,000,000, about 500,000, about 100,000, about 50,000, about 10,000, about 5,000, about 1,000, about 500, or about 100.
[0251] Generally, procedures for investigating the physical layout of chromosomes, such as Hi-C based techniques, utilize chromatin formed within a cell / organism, such as chromatin isolated from cultured cells or primary tissue. The present disclosure provides for the use of such techniques with not only chromatin isolated from a cell / organism, but also reconstituted chromatin. Reconstituted chromatin is differentiated from chromatin formed within a cell / organism across various characteristics. First, for many samples, collection of naked DNA samples can be achieved by using a variety of methods ranging from non-invasive to invasive, such as collecting body fluids, swabbing the cheek or rectal area, obtaining epithelial samples, etc. Second, reconstitution of chromatin substantially precludes the formation of inter-chromosomal, and other long-range interactions that generate artifacts for genome assembly and haplotype phasing. In some cases, the sample may have inter-chromosomal or intermolecular cross-links according to the methods and compositions of the present disclosure that are less than, or below, about 20, 15, 12, 11, 10, 9, 8, 7, 6, 5, 4, 3, 2, 1, 0.5, 0.4, 0.3, 0.2, 0.1%. In some examples, the sample may have inter-chromosomal or intermolecular cross-links of less than about 5%. In some examples, the sample may have inter-chromosomal or intermolecular cross-links of less than about 3%. In further examples, the sample may have inter-chromosomal or intermolecular cross-links of less than about 1%. Third, the frequency of sites capable of cross-linking, and thus the frequency of intra-molecular cross-links within a polynucleotide, can be regulated. For example, the DNA to histone ratio may vary, thereby allowing the nucleosome density to be adjusted to a desired value. In some cases, the nucleosome density is reduced below physiological levels. Thus, the distribution of cross-links can be modified to support a longer range of interactions. In some embodiments, subsamples with varying cross-link densities can be adjusted to cover associations over both short and long ranges.For example, the cross-linking conditions can be adjusted such that at least about 1%, about 2%, about 3%, about 4%, about 5%, about 6%, about 7%, about 8%, about 9%, about 10%, about 11%, about 12%, about 13%, about 14%, about 15%, about 16%, about 17%, about 18%, about 19%, about 20%, about 25%, about 30%, about 40%, about 45%, about 50%, about 60%, about 70%, about 80%, about 90%, about 95%, or about 100% of the cross-linking occurs between DNA segments that are at least about 50 kb, about 60 kb, about 70 kb, about 80 kb, about 90 kb, about 100 kb, about 110 kb, about 120 kb, about 130 kb, about 140 kb, about 150 kb, about 160 kb, about 180 kb, about 200 kb, about 250 kb, about 300 kb, about 350 kb, about 400 kb, about 450 kb, or about 500 kb apart on the sample DNA molecule.
[0252] Contact mapping and topology The read pairs generated by the methods of the present disclosure can be used to analyze the three-dimensional structure of the genome, and of the chromosomes and nucleic acid molecules therein. As discussed herein, each read in a read pair can be mapped to different regions in the genome. For a given read pair, it can be inferred that the two different regions in the genome to which they map would have been spatially proximal to each other in order to be ligated together. A contact map can be created for a sample by plotting the read pairs from the sample by the coordinates of both reads in the read pair. An exemplary contact map can be found in FIG. 13, where each point on the contact map represents a read pair plotted according to the mapped positions of that read pair.
[0253] Analysis of contacts across a sample can enable analysis of chromosome and genome structure. Organization into genomic A and B compartments, active and inactive compartments, chromosomal compartments, euchromatin and heterochromatin, topological associated domains (TADs) including TAD subtypes, and other structures can be analyzed at the kilobase or megabase scale. Analysis of contact maps can also enable detection of genomic features such as structural variants, such as rearrangements, translocations, copy number variations, inversions, deletions, and insertions.
[0254] The methods of the present disclosure can provide the positions of protein binding, structural variations, or genomic contact interactions at a resolution of about 1 bp, 2 bp, 3 bp, 4 bp, 5 bp, 6 bp, 7 bp, 8 bp, 9 bp, 10 bp, 20 bp, 30 bp, 40 bp, 50 bp, 60 bp, 70 bp, 80 bp, 90 bp, 100 bp, 200 bp, 300 bp, 400 bp, 500 bp, 600 bp, 700 bp, 800 bp, 900 bp, 1000 bp, 2000 bp, 3000 bp, 4000 bp, 5000 bp, 6000 bp, 7000 bp, 8000 bp, 9000 bp, or less than 10 kb, or equal thereto. In some cases, protein binding sites, protein footprints, contact interactions, or other features can be mapped within 1000 bp, 900 bp, 800 bp, 700 bp, 600 bp, 500 bp, 400 bp, 300 bp, 200 bp, 190 bp, 180 bp, 170 bp, 160 bp, 150 bp, 140 bp, 130 bp, 120 bp, 110 bp, 100 bp, 90 bp, 80 bp, 70 bp, 60 bp, 50 bp, 40 bp, 30 bp, 20 bp, 10 bp, 9 bp, 8 bp, 7 bp, 6 bp, 5 bp, 4 bp, 3 bp, 2 bp, or 1 bp. In one example, the methods of the present disclosure can enable the resolution of sites (e.g., protein binding sites such as CTCF sites) that are within 10,000 bp, 5,000 bp, 2,000 bp, or 1,000 bp of each other on the genome. In some cases, improved resolution or mapping can be achieved by using MNase or other endonucleases that cleave unprotected nucleic acids (e.g., nucleic acids that are not within the footprint of the binding protein), which results in proximity ligation events occurring at the ends of the protected regions (e.g., protein footprints).
[0255] Contig mapping In various embodiments, the present disclosure provides various ways to enable mapping of multiple read pairs to multiple contigs. There are several publicly available computer programs for mapping reads to a contig array. These read mapping program data also provide data explaining how unique a particular read mapping is within the genome. From a population of reads that map uniquely with high confidence within a contig, the distribution of the distance between reads of each read pair can be inferred. For read pairs of reads that map with confidence to different contigs, this mapping data suggests a connection between the two contigs in question. It also suggests the distance between two contigs that is proportional to the distribution of distances learned from the above analysis. Thus, each read pair of reads that map with confidence to different contigs suggests a connection between those two contigs during accurate assembly. The connections inferred from all such mapped read pairs can be summarized in an adjacency matrix, where each contig is represented by both rows and columns. Read pairs that connect contigs are marked with non-zero values in the corresponding rows and columns indicating the contigs to which the reads in the read pair were mapped. Most read pairs will map within a contig, from which the distribution of distances between read pairs can be learned, and from which the adjacency matrix of the contig can be constructed using read pairs that map to different contigs.
[0256] In various embodiments, the present disclosure provides a method that includes constructing an adjacency matrix of contigs using read mapping data from read pair data. In some embodiments, the adjacency matrix uses a weighting scheme for read pairs that incorporates a tendency for short-range interactions over long-range interactions. Read pairs that span short distances are typically more common than those that span longer distances. A function that describes the probability of a particular distance can be fit using read pair data that maps to a single contig to learn this distribution. Thus, one important feature of read pairs that map to different contigs is the position on the contig where they map. In the case of read pairs that both map near one end of a contig, the estimated distance between these contigs can be short, and thus the distance between the joined reads can be small. Since shorter distances between read pairs are more common than longer distances, this configuration provides stronger evidence that these two contigs are adjacent than read mapping far from the ends of the contig. Thus, the connections in the adjacency matrix are further weighted by the distance of the reads to the ends of the contig. In further embodiments, the adjacency matrix can be further rescaled to down-weight a large number of contacts on several contigs that represent uninteresting regions of the genome. These regions of the genome can be identified by having a high rate of read mapping to them and are more likely to contain spurious read mappings that can give false information to the assembly. In yet further embodiments, this scaling can be directed by searching for one or more conserved binding sites for one or more agents that regulate chromatin scaffold interactions, such as the transcription repressor CTCF, an endocrine receptor, cohesin, or a histone that is commonly modified.
[0257] In some embodiments, the present disclosure provides one or more methods as disclosed herein, the method comprising analyzing an adjacency matrix, thereby determining a path through contigs that represent their order and / or orientation on the genome. In other embodiments, the path through the contigs can be selected such that each contig is visited exactly once. In further embodiments, the path through the contigs is selected to maximize the sum of the weights of the bridges visited by the path through the adjacency matrix. In this way, most likely, the contig connections are proposed for an accurate assembly. In still further embodiments, the path through the contigs can be selected such that each contig is visited exactly once and the weights at the ends of the adjacency matrix are maximized.
[0258] Haplotype phasing In a diploid genome, it is often important to know which allelic variants are linked on the same chromosome. This is known as haplotype phasing. In short reads from high-throughput sequence data, it is rarely possible to directly observe which allelic variants are linked. Computational inference of haplotype phasing can be unreliable over long distances. The present disclosure provides one or more methods that enable determination of which allelic variants are linked using allelic variants on read pairs. In some cases, phasing by the methods of the present disclosure is performed without imputation.
[0259] In various embodiments, the methods and compositions of the present disclosure enable haplotype phasing of diploid or polyploid genomes with respect to multiple allelic variants. Accordingly, the methods described herein can provide determination of linked allelic variants that are linked based on variant information from read pairs and / or contigs assembled using the same. Examples of allelic variants include, but are not limited to, those well-known by the 1000 Genomes Project, UK10K, HapMap, and other projects for discovering genetic variation among humans. The association of a particular gene with a disease can be more readily revealed by having haplotype phasing data, as demonstrated, for example, by the identification of unlinked inactivating mutations in both copies of SH3TC2 leading to Charcot-Marie-Tooth neuropathy (Lupski JR, Reid JG, Gonzaga-Jauregui C et al. N. Engl. J. Med. 362:1181-91, 2010), and the identification of unlinked inactivating mutations in both copies of ABCG5 leading to hypercholesterolemia 9 (Rios J, Stein E, Shendure J et al. Hum. Mol. Genet. 19:4313-18, 2010).
[0260] Humans, on average, are heterozygous at one site in every 1,000. In some cases, data from a single lane using high-throughput sequencing methods can generate at least about 150,000,000 read pairs. The read pairs can be about 100 base pairs in length. From these parameters, it is estimated that one-tenth of all reads from a human sample cover heterozygous sites. Thus, on average, one-hundredth of all read pairs from a human sample are estimated to cover pairs of heterozygous sites. Thus, about 1,500,000 read pairs (one-hundredth of 150,000,000) provide phasing data using a single lane. The human genome has about 3 billion bases, and since one in 1,000 is heterozygous, there are about 3 million heterozygous sites in the average human genome. Using about 1,500,000 read pairs representing pairs of heterozygous sites, the average coverage of each heterozygous site phased using a single lane of high-throughput sequencing methods is about (1X) using a typical high-throughput sequencing machine. Thus, the diploid human genome can be reliably and completely phased in one lane of high-throughput sequence data related to sequence variants from a sample prepared using the methods disclosed herein. In some examples, a lane of data can be a set of DNA sequence read data. In further examples, a lane of data can be a set of DNA sequence read data from a single run of a high-throughput sequencing instrument.
[0261] Since the human genome consists of two homologous sets of chromosomes, to understand the true genetic makeup of an individual, it is necessary to depict the maternal and paternal copies or haplotypes of the genetic material. Obtaining an individual's haplotypes is useful in several respects. First, haplotypes are useful clinically in predicting the outcome of donor-host matching in organ transplantation and are increasingly being used as a means to detect disease associations. Second, in genes showing compound heterozygosity, haplotypes provide information about whether two deleterious variants are located on the same allele, which greatly influences the prediction of whether the inheritance of these variants is harmful. Third, haplotypes from groups of individuals provide information about population structure and the history of human evolution. Finally, the recently described extensive allelic imbalance in gene expression suggests that genetic or epigenetic differences between alleles can contribute to quantitative differences in expression. Understanding haplotype structure depicts the mechanism of variants contributing to allelic imbalance.
[0262] In certain embodiments, the methods disclosed herein include in vitro techniques for fixing and capturing associations between distant regions of the genome, such as those required for long-range linkage and phasing. Optionally, the method includes constructing and sequencing an XLRP library to provide genomically very distant read pairs. Optionally, the interactions primarily result from random associations within a single DNA fragment. In some instances, the genomic distance between segments can be inferred because segments that are close to each other in a DNA molecule interact more frequently and with higher probability, while interactions between distant parts of the molecule are less frequent. Consequently, there is a systematic relationship between the number of pairs connecting two loci and their proximity on the input DNA. The present disclosure can generate read pairs that span the largest DNA fragments upon extraction. The maximum length of the input DNA for this library is 150 kbp, which is the longest meaningful read pair observed from the sequencing data. This suggests that the method can further link genomically distant loci when larger input DNA fragments are provided. Applying improved assembly software tools that are specially adapted to process the type of data generated by the method can enable a complete genome assembly.
[0263] Very high phasing accuracy can be achieved with the data generated using the methods and compositions of the present disclosure. In comparison with previous methods, the methods described herein can phase a higher percentage of variants. Phasing can be achieved while maintaining a high level of accuracy. The techniques herein can enable phasing with an accuracy of greater than about 70%, 80%, 90%, 95%, 96%, 97%, 98%, 99%, 99.9%, 99.99%, or 99.999%. The techniques herein can enable accurate phasing with a sequencing depth of less than 500x, less than 450x, less than 400x, less than 350x, less than 300x, less than 250x, less than 200x, less than 150x, less than 100x, or less than 50x. This phasing information can be extended over long distances, for example, greater than about 200 kbp, about 300 kbp, about 400 kbp, about 500 kbp, about 600 kbp, about 700 kbp, about 800 kbp, about 900 kbp, about 1 Mbp, about 2 Mbp, about 3 Mbp, about 4 Mbp, about 5 Mbp, or greater than about 10 Mbp. In some embodiments, greater than 90% of the heterozygous SNPs in a human sample can be phased with an accuracy of greater than 99% using less than about 250 million reads or read pairs, for example, by using only 1 lane of Illumina HiSeq data. In other cases, greater than about 40%, 50%, 60%, 70%, 80%, 90%, 95%, or 99% of the heterozygous SNPs for a human sample can be phased with an accuracy of greater than about 70%, 80%, 90%, 95%, 96%, 97%, 98%, 99%, 99.9%, 99.99%, or 99.999% using less than about 250 million or less than about 500 million reads and read pairs, for example, by using only 1 or 2 lanes of Illumina HiSeq data. For example, greater than 95% or 99% of the heterozygous SNPs in a human sample can be phased with an accuracy of greater than 95% or 99% using about 250 million or about 500 million reads.In further cases, additional variants can be captured by increasing the read length to about 200 bp, 250 bp, 300 bp, 350 bp, 400 bp, 450 bp, 500 bp, 600 bp, 800 bp, 1000 bp, 1500 bp, 2 kbp, 3 kbp, 4 kbp, 5 kbp, 10 kbp, 20 kbp, 50 kbp, or 100 kbp.
[0264] In other embodiments of the present disclosure, data from the XLRP library can be used to confirm the phasing ability of long-range read pairs. The accuracy of these results is equivalent to the state-of-the-art previously available, but is significantly extended to even longer distances. The current sample preparation protocol for a particular sequencing method recognizes variants located within the read length of the phasing target site, e.g., within 150 bp. In one example, 44% of the 1,703,909 heterozygous SNPs present were phased with an accuracy exceeding 99% from an XLRP library constructed for the assembly benchmark sample NA12878. In some cases, this ratio can be extended to almost all variable sites by a judicious choice of enzyme or digestion conditions.
[0265] Haplotype phasing can involve phasing the human leukocyte antigen (HLA) region (e.g., HLA-A, B, and C of class I; HLA-DRB1 / 3 / 4 / 5, HLA-DQA1, HLA-DQB1, HLA-DPA1, HLA-DPB1 of class II). The HLA region of the genome is densely polymorphic and can be difficult to sequence or phase with standard sequencing approaches. The techniques of the present disclosure can provide improved sequencing and improved phasing of the HLA region of the genome. Using the techniques of the present disclosure, the HLA region of the genome can be accurately phased as part of the phasing of a larger region (e.g., a chromosome arm, a chromosome, the entire genome) or by itself (e.g., by target enrichment such as hybrid capture). In one example, the HLA region itself was accurately phased at a sequencing depth of about 300x. These techniques can provide advantages over conventional approaches for HLA analysis, such as long-range PCR, which can involve complex protocols and a number of separate reactions. As further discussed herein, samples can be multiplexed for sequencing analysis, for example, by including sample identification barcodes in cross-linking oligonucleotides or elsewhere, and by demultiplexing sequence information based on the barcodes. In one example, multiple samples are subjected to proximity ligation, barcoded with sample identification barcodes (e.g., in cross-linking oligonucleotides), the HLA region is targeted (e.g., by hybrid capture), and multiplex sequencing is performed to enable phasing of HLA for multiple samples. In some cases, phasing of the HLA region is performed without imputation.
[0266] Haplotype phasing can include phasing the killer cell immunoglobulin-like receptor (KIR) region. The KIR region of the genome is highly homologous and structurally dynamic for transposon-mediated recombination, and sequencing or phasing using standard sequencing approaches can be difficult. The techniques of the present disclosure can provide improved sequencing of the KIR region of the genome and improved phasing accuracy. Using the techniques of the present disclosure, the KIR region of the genome can be accurately phased as part of the phasing of a larger region (e.g., a chromosome arm, a chromosome, the entire genome) or by itself (e.g., by target enrichment such as hybrid capture). These techniques can provide advantages over conventional approaches for HLA analysis, such as long-range PCR, which can involve complex protocols and a number of separate reactions. As further discussed herein, samples can be multiplexed for sequencing analysis by, for example, including sample identification barcodes in cross-linked oligonucleotides or elsewhere and demultiplexing sequence information based on the barcodes. In one example, multiple samples are subjected to proximity ligation, barcoded with sample identification barcodes (e.g., in cross-linked oligonucleotides), the KIR region is targeted (e.g., by hybrid capture), and multiplex sequencing is performed to enable phasing of the KIR region for multiple samples. At least about 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, or more genes and / or pseudogenes can be phased. In some cases, phasing of the KIR region is performed without imputation.
[0267] Metagenomic analysis In some embodiments, the compositions and methods described herein enable metagenomic investigations, such as those found in the human gut. Thus, it is possible to investigate the partial or entire genomic sequences of some or all of the organisms inhabiting a given ecological environment. Examples include random sequencing of all gut microbiota, microbiota identified in specific areas of the skin, and microbiota inhabiting toxic waste sites. The composition of the microbial populations in these environments can be determined using the compositions and methods described herein, as well as the interrelated biochemical aspects encoded by their respective genomes. The methods described herein enable metagenomic studies from complex biological environments, including, for example, those containing 2, 3, 4, 5, 6, 7, 8, 9, 10, 12, 15, 20, 25, 30, 40, 50, 60, 70, 80, 90, 100, 125, 150, 175, 200, 250, 300, 400, 500, 600, 700, 800, 900, 1000, 5000, or more than 10,000 organisms and / or variants of organisms.
[0268] The high accuracy required by cancer genome sequencing can be achieved using the methods and systems described herein. Incorrect reference genomes can pose difficulties in basecalling when sequencing cancer genomes. Heterogeneous samples and small starting materials, such as samples obtained by biopsy, pose further difficulties. Additionally, the detection of large-scale structural variations and / or losses of heterozygosity is generally important for cancer genome sequencing, as well as the ability to distinguish somatic mutations from errors in basecalling.
[0269] Improved sequencing accuracy The systems and methods described herein can generate accurate long sequences from complex samples containing two, three, four, five, six, seven, eight, nine, ten, twelve, fifteen, twenty, or more various genomes. Normal, benign, and / or tumor-derived mixed samples may be analyzed optionally without the need for normal controls. In some embodiments, small starting samples as little as 100 ng or even a few hundred genome equivalents are utilized to generate accurate long sequences. The systems and methods described herein may enable the detection of large-scale structural variations and rearrangements, and phased variant calls may be obtained over long sequences spanning about 1 kbp, about 2 kbp, about 5 kbp, about 10 kbp, 20 kbp, about 50 kbp, about 100 kbp, about 200 kbp, about 500 kbp, about 1 Mbp, about 2 Mbp, about 5 Mbp, about 10 Mbp, about 20 Mbp, about 50 Mbp, about 100 Mbp, or more nucleotides. For example, phased variant calls may be obtained over long sequences spanning about 1 Mbp or about 2 Mbp.
[0270] Haplotypes determined using the methods and systems described herein may be assigned to computer resources, such as computer resources on a network, such as a cloud system. Short variant calls can be corrected using relevant information stored in computer resources if necessary. Structural variations can be detected based on combined information from short variant calls and information stored in computer resources. Other heterochromatic regions including, but not limited to, segmental duplications, regions prone to structural variations, highly variable and medically relevant MHC regions, centromere and telomere regions, and regions with repetitive regions, as well as problematic parts of the genome such as low sequence accuracy, high mutation rates, ALU repeats, segmental duplications, or other relevant problematic parts known in the art, can be reassembled for improved accuracy.
[0271] The sample type can be assigned to array information, either locally or in network-connected computer resources such as the cloud. If the source of the information is known, for example, when the source of the information is from cancer or normal tissue, this source ...
Claims
1. (a) obtaining a stabilized biological sample comprising a nucleic acid molecule complexed with at least one nucleic acid binding protein; (b) contacting the stabilized biological sample with a non-specific endonuclease to cleave the nucleic acid molecule into a plurality of segments; (c) attaching a first segment and a second segment of the plurality of segments at one junction; and (d) subjecting the plurality of segments to size selection to obtain a plurality of selected segments.
2. The method according to claim 1, wherein the plurality of selected segments are about 145 to about 600 bp.
3. The method according to claim 1, wherein the plurality of selected segments are about 100 to about 2500 bp.
4. The method according to claim 1, wherein the plurality of selected segments are about 100 to about 600 bp.
5. The method according to claim 1, wherein the plurality of selected segments are about 600 to about 2500 bp.
6. The method according to claim 1, further comprising adjusting a sequencing library from the plurality of segments prior to step (d).
7. The method according to claim 6, further comprising subjecting the sequencing library to size selection to obtain a size-selected library.
8. The method according to claim 7, wherein the size-selected library has a size from about 350 bp to about 1000 bp.
9. The method according to any one of claims 1 to 8, wherein the size selection is performed using gel electrophoresis, capillary electrophoresis, size selection beads, or a gel filtration column.
10. The method according to any one of claims 1 to 9, further comprising analyzing the plurality of selected segments to obtain a QC value.
11. The method according to claim 10, wherein the QC value is a chromatin digestion efficiency (CDE) based on the ratio of segments sized 100 bp to 2500 bp before step (d).
12. The method according to claim 11, further comprising selecting a sample for further analysis when the CDE value is at least 65%.
13. The method according to claim 10, wherein the QC value is a chromatin digestion index (CDI) based on the ratio of the number of segments of mononucleosome size to the number of segments of dinucleosome size before step (d).
14. The method according to claim 13, further comprising the step of selecting a sample for further analysis when the CDI value is greater than -1.5 and less than 1.
15. The method according to claim 1, further comprising the step of binding a plurality of segments to one or more surfaces following the step of contacting the stabilized biological sample with a non-specific endonuclease.
16. The method according to claim 15, wherein the one or more surfaces comprise one or more beads.
17. The method according to claim 16, wherein the one or more beads are solid phase reversible immobilization (SPRI) beads.
18. The method according to any one of claims 1 to 14, wherein the stabilized biological sample comprises a stabilized cell lysate.
19. The method according to any one of claims 1 to 14, wherein the stabilized biological sample comprises stabilized intact cells.
20. The method according to any one of claims 1 to 14, wherein the stabilized biological sample comprises stabilized intact nuclei.
21. The method according to claim 19 or 20, wherein step (b) is performed prior to the lysis of intact cells or intact nuclei.
22. The method according to claim 1, further comprising the step of lysing cells and / or nuclei in the stabilized biological sample prior to step (c).
23. The method according to any one of claims 1 to 20, wherein the stabilized biological sample comprises less than 3,000,000 cells.
24. The method according to any one of claims 1 to 23, wherein the stabilized biological sample comprises less than 1,000,000 cells.
25. The method according to any one of claims 1 to 24, wherein the stabilized biological sample comprises less than 100,000 cells.
26. The method according to any one of claims 1 to 25, wherein the stabilized biological sample comprises less than 10 μg of DNA.
27. The method according to any one of claims 1 to 26, wherein the stabilized biological sample contains less than 1 μg of DNA.
28. The method according to any one of claims 1 to 27, wherein the non-specific endonuclease is DNase.
29. The method according to claim 28, wherein the DNase is DNase I.
30. The method according to claim 28, wherein the DNase is DNase II.
31. The method according to claim 28, wherein the DNase is micrococcal nuclease.
32. The method according to claim 28, wherein the DNase is selected from one or more of DNase I, DNase II, and micrococcal nuclease.
33. The method according to any one of claims 1 to 32, wherein the stabilized biological sample has been treated with a cross-linking agent.
34. The method according to claim 33, wherein the cross-linking agent is a chemical fixative.
35. The method according to claim 34, wherein the chemical fixative contains formaldehyde.
36. The method according to claim 34, wherein the chemical fixative contains psoralen.
37. The method according to claim 34, wherein the chemical fixative contains dithiobis(succinimidyl glutarate) (DSG).
38. The method according to claim 34, wherein the chemical fixative contains ethylene glycol bis(succinimidyl succinate) (EGS).
39. The method according to claim 34, wherein the chemical fixative contains dithiobis(succinimidyl glutarate) (DSG) and ethylene glycol bis(succinimidyl succinate) (EGS).
40. The method according to claim 33, wherein the cross-linking agent is ultraviolet light.
41. The method according to any one of claims 1 to 40, wherein the stabilized biological sample is a cross-linked paraffin-embedded tissue sample.
42. The method according to any one of claims 1 to 41, further comprising the step of contacting a plurality of selected segments with an antibody.
43. The method according to claim 1, further comprising the step of performing immunoprecipitation on a plurality of segments.
44. The method according to claim 43, characterized in that immunoprecipitation is carried out after the step of attaching.
45. The method according to any one of claims 1 to 42, wherein the step of attaching comprises filling the sticky ends using biotin-tagged nucleotides.
46. The method according to any one of claims 1 to 42, wherein the step of attaching comprises filling the sticky ends using untagged nucleotides.
47. The method according to any one of claims 1 to 42, wherein the step of attaching comprises ligating blunt ends.
48. The method according to any one of claims 1 to 42, wherein the step of attaching comprises adding an overhang.
49. The method according to claim 48, wherein adding an overhang comprises adenylation.
50. The method according to any one of claims 1 to 45, wherein the step of attaching comprises contacting at least a first segment and a second segment with at least one cross-linking oligonucleotide.
51. The method according to claim 50, characterized in that the cross-linking oligonucleotide has a length of at least 10 bp.
52. The method according to claim 50, characterized in that the cross-linking oligonucleotide has a length of at least 12 bp.
53. The method according to claim 50, characterized in that the cross-linking oligonucleotide has a length of 12 bp.
54. The method according to claim 50, characterized in that the cross-linking oligonucleotide comprises a barcode sequence.
55. The method according to claim 50, characterized in that the cross-linking oligonucleotide comprises an affinity tag.
56. The method according to claim 55, characterized in that the affinity tag is biotin.
57. The method according to claim 50, wherein the step of attaching comprises successively contacting at least a first segment and a second segment with a plurality of cross-linking oligonucleotides.
58. The method according to claim 55, wherein the step of attaching results in a sample, cell, nucleus, chromosome, or nucleic acid molecule of a stabilized biological sample receiving the unique sequence of the cross-linking oligonucleotide.
59. The method according to any one of claims 50 to 56, characterized in that at least one crosslinked oligonucleotide is linked to one immunoglobulin-binding protein or a fragment thereof.
60. The method according to any one of claims 50 to 57, characterized in that at least one crosslinked oligonucleotide is linked or fused to two or more immunoglobulin-binding proteins or fragments thereof.
61. The method according to claim 57 or 58, characterized in that the immunoglobulin-binding protein is selected from protein A, protein G, protein A / G, and protein L.
62. The method according to any one of claims 1 to 45, wherein the step of attaching comprises contacting at least a first segment and a second segment with a barcode.
63. The method according to any one of claims 1 to 60, characterized in that the method does not include a shearing step.
64. The method according to claim 1, further comprising obtaining at least some sequences on both sides of the junction to generate a first read pair.
65. (f) mapping the first read pair to a set of contigs, and (g) determining a path across the set of contigs that represents the order and / or orientation to the genome, the method according to claim 62.
66. (f) mapping the first read pair to a set of contigs, and (g) determining the presence of structural variants or a decrease in heterozygosity in the stabilized biological sample from the set of contigs, the method according to claim 62.
67. (f) mapping the first read pair to a set of contigs, and (g) assigning phases to the variants in the set of contigs, the method according to claim 62.
68. The method according to claim 65, characterized in that the variant is a human leukocyte antigen (HLA) variant.
69. The method according to claim 65, characterized in that the variant is a killer cell immunoglobulin-like receptor (KIR) variant.
70. (f) mapping the first read pair to a set of contigs, (g) further comprising determining the presence of variants in the set of contigs, and (h)(1) a step of checking the disease stage, prognosis, or treatment policy for the stabilized biological sample; (2) a step of selecting a drug based on the presence of the variant; or (3) a step of checking the drug efficacy for the stabilized biological sample, the method according to claim 62, further comprising performing one or more steps selected therefrom.
71. The method according to any one of claims 1 to 68, wherein the DNase is linked or fused to an immunoglobulin binding protein or a fragment thereof.
72. The method according to any one of claims 1 to 69, wherein the DNase is linked to two or more immunoglobulin binding proteins or fragments thereof.
73. The method according to claim 69 or 70, wherein the immunoglobulin binding protein is selected from protein A, protein G, protein A / G, and protein L.
Citation Information
Patent Citations
Methods for genome assembly, haplotype phasing, and target-independent nucleic acid detection
JP2019500009A