Methods and compositions for preparing an array determination library
By stabilizing nucleic acid samples, cleaving them with a transposase, ligating, and circularizing the segments, the method addresses the challenge of obtaining high-quality genomic sequences, achieving improved accuracy and efficiency in genomic sequence determination and three-dimensional structure analysis.
Patent Information
- Application Number
- JP2024566469
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-03-14
- Filing Date
- 2023-05-10
- Publication Date
- 2025-05-30
AI Technical Summary
Obtaining high-quality continuous genomic sequences is challenging, especially with limited source materials, as existing methods struggle with efficiently and accurately analyzing and assembling raw sequence data.
The method involves obtaining a stabilized sample of nucleic acid molecules complexed with nucleic acid binding proteins, cleaving these molecules into segments using a transposase, ligating these segments to form a ligated nucleic acid, and then circularizing and sequencing it to map to a genome.
This method enables the generation of extremely long-range read pairs, improving the determination of genomic sequences, structural genomic information, and the physical three-dimensional structure of nucleic acids in cells, with enhanced accuracy and efficiency compared to previous methods.
Smart Images

Figure 2025516619000001_ABST
Abstract
Description
Technical Field
[0001] Cross - Reference to Related Applications This application claims the benefit of U.S. Provisional Patent Application No. 63 / 340,734, filed May 11, 2022, and U.S. Provisional Patent Application No. 63 / 490,192, filed Mar. 14, 2023, each of which is hereby incorporated by reference in its entirety.
Background Art
[0002] Obtaining high - quality continuous genomic sequences is often difficult, especially when the source materials available for sequence analysis are limited. Although raw sequence data has become available faster and at lower cost, suitable methods for efficiently and accurately analyzing and assembling the data remain an issue.
Summary of the Invention
[0003] In one aspect, methods of nucleic acid processing are provided herein. Optionally, the method includes obtaining a stabilized sample comprising a nucleic acid molecule complexed with at least one nucleic acid binding protein. Optionally, the method includes cleaving the nucleic acid molecule into a plurality of segments comprising at least a first segment and a second segment, wherein the cleavage is effected by a transposase. Optionally, the method includes producing a ligated nucleic acid comprising a first sequence from the first segment and a second sequence from the second segment by ligating the first segment to the second segment. Optionally, the transposase is a Tn5 transposase. Optionally, the method further includes circularizing the ligated nucleic acid by ligating the 5' end of the ligated nucleic acid to the 3' end of the ligated nucleic acid. Optionally, the method further includes sequencing at least a portion of the ligated nucleic acid. Optionally, the sequencing step includes sequencing at least a portion of the first sequence and at least a portion of the second sequence. Optionally, the method includes mapping at least a portion of the first sequence and at least a portion of the second sequence to a genome. Optionally, the method further includes performing a three-dimensional genome analysis using information from the sequencing step. Optionally, the stabilized sample is a cross-linked sample. Optionally, the step of obtaining a stabilized sample includes obtaining a sample and stabilizing the sample. Optionally, the step of obtaining a stabilized sample includes obtaining a pre-stabilized sample. Optionally, the nucleic acid binding protein includes chromatin or a component thereof. Optionally, a linker sequence is ligated between the first segment and the second segment. Optionally, the linker sequence includes a barcode sequence. Optionally, the barcode sequence indicates a partition of origin. Optionally, the barcode sequence indicates a cell of origin. Optionally, the barcode sequence indicates a cell population of origin. Optionally, the barcode sequence indicates an organism of origin. Optionally, the cleavage occurs in open and closed chromatin compartments.In some cases, at least 10% of the cleavage occurs in the closed chromatin compartment. In some cases, at least 20% of the cleavage occurs in the closed chromatin compartment. In some cases, at least 30% of the cleavage occurs in the closed chromatin compartment. In some cases, the stabilized biological sample contains 50,000 or fewer cells. In some cases, the stabilized biological sample contains at least 10,000 cells. In some cases, the stabilized biological sample contains stabilized nuclei. In some cases, the stabilized biological sample contains 50,000 or fewer nuclei. In some cases, the stabilized biological sample contains at least 10,000 nuclei. In some cases, the bound nucleic acid does not contain an affinity tag. In some cases, the bound nucleic acid does not contain biotin. In some cases, the circularized bound nucleic acid does not contain an affinity tag. In some cases, the circularized bound nucleic acid does not contain biotin. In some cases, the bound nucleic acid and / or the circularized bound nucleic acid are isolated without using an affinity tag. In some cases, the bound nucleic acid and / or the circularized bound nucleic acid are isolated without using streptavidin.
[0004] Incorporation by reference All publications, patents, and patent applications mentioned herein are hereby incorporated by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference.
Brief Description of the Drawings
[0005] An understanding of the features and advantages of the present invention will be obtained by reference to the following detailed description, which describes exemplary embodiments in which the principles of the present invention are used, and the accompanying drawings.
[0006]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9A
Figure 9B
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
MODE FOR CARRYING OUT THE INVENTION
[0007] In one aspect, provided herein are compositions, systems, and methods for generating extremely long-range read pairs of nucleic acids related to the determination of genomic sequences including long-range and structural genomic information, the determination of the physical three-dimensional structure of nucleic acids in cells, and having results improved over several other methods. The methods herein can utilize techniques including, but not limited to, transposase fragmentation of cross-linked nucleic acids and ligation-based binding of transposase-fragmented nucleic acids.
[0008] Transposase-fragmented chromatin In one aspect, provided herein are methods of nucleic acid processing. Such methods can include obtaining a stabilized sample comprising a nucleic acid molecule complexed with at least one nucleic acid binding protein, and cleaving the nucleic acid molecule into a plurality of segments including at least a first segment and a second segment, wherein the cleavage is effected by a transposase. The methods herein can further include ligating the first segment and the second segment, thereby creating a ligated nucleic acid. Optionally, the ligated nucleic acid is further ligated to produce a circularized ligated nucleic acid.
[0009] In various aspects of the methods herein, the step of cleaving the stabilized nucleic acid is effected using a transposase. Optionally, the cleavage is effected in permeabilized cells. Optionally, the cleavage is effected in permeabilized nuclei. Optionally, the transposase is Tn5, Tn3, Tn7, sleeping beauty transposase, or a combination thereof. Optionally, the transposase is Tn5 transposase.
[0010] In various aspects of the methods herein, cleavage occurs in both open chromatin and closed chromatin. In some cases, at least about 10% to at least about 50% of the cleavage occurs in closed chromatin. For example, at least about 10%, at least about 15%, at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, or more of the cleavage occurs in closed chromatin. In some embodiments, closed chromatin is transcriptionally inactive and binds to one or more nucleosomes or other chromatin proteins. In some embodiments, open chromatin is transcriptionally active and binds little to nucleosomes or other chromatin proteins.
[0011] In various aspects of the methods herein, a linker containing a recombinase site can be contacted with a nucleic acid cleaved in the presence of a recombinase, and the recombinase site comprises two recombinase sites oriented as direct repeats. In some cases, the presence of the recombinase sites oriented as direct repeats can prevent the resulting product from forming a stable hairpin structure. In some cases, the bound nucleic acid does not form a hairpin loop. In some cases, the resulting product is more readily sequenced than a product in which the recombinase sites are oriented as inverted repeats.
[0012] In various aspects of the methods herein, a first segment and a second segment of a cleaved nucleic acid are contacted with a linker containing an integrase site in the presence of a recombinase. In some cases, the recombinase is an integrase. In some cases, the integrase is PhiC31 integrase, Bxb1 integrase, or a combination thereof.
[0013] In various aspects, methods of processing nucleic acids are provided in which the joined nucleic acids are circularized. Optionally, the ends are removed to expose the first and second segments prior to ligation. In some embodiments, the circularized product is amplified using PCR to generate a sequencing library. An example of this method is illustrated in FIGS. 6 and 7.
[0014] In an exemplary embodiment, the sample can be prepared and cross-linked prior to subjecting it to in situ tagmentation that fragments the chromatin and leaves mosaic ends at each end of the fragmented chromatin. The tagmented chromatin can then be ligated using adapters and ligase. The ends can be removed and the cross-linking reversed. The nucleic acids can be captured and the resulting fragments are then circularized via ligation, resulting in a circular nucleic acid having two genomic DNA fragments linked to the mosaic ends of each end all joined together by the adapter. Genomic DNA for analysis can be amplified using adapter PCR and purification / size selection suitable for sequencing analysis (FIGS. 6 and 7).
[0015] Circularization-based approaches can offer several advantages. As discussed above, when the adapter site flanks the genomic sequence of interest, circularization can generate nucleic acid molecules, enabling facile production of nucleic acid molecules for sequencing a higher percentage of the genomic sequence (e.g., by elimination of linker sequences). Additionally, circularization-based approaches can obviate the need for affinity tag enrichment approaches. Existing proximity ligation approaches generally use affinity tag enrichment (e.g., incorporating biotinylated nucleic acids into proximity ligation sites, which can then be enriched with surface-bound streptavidin) to ensure that the ultimately sequenced nucleic acid is representative of proximity ligation events rather than, for example, typical genomic DNA, or alternatively, as presented herein, circularization may require nucleic acids of at least a certain length, so circularization can be performed to enrich nucleic acids that have undergone proximity ligation. For example, nucleic acids of less than about 250 base pairs, such as mononucleosome-sized fragments that did not ligate to a partner during proximity ligation, may not circularize. In some cases, enrichment of circularized molecules can be performed by cleanup, bead binding, or size selection, among other means. In other cases, enrichment of circularized molecules is not necessary, and instead primer-based amplification (e.g., adapter PCR) generates amplification products suitable for sequencing only from circularized molecules.
[0016] In various aspects of the methods herein, the linker includes a mosaic end, a sequencing adapter, and an attB sequence. Alternatively, the linker includes a mosaic end and a sequencing adapter, and the attB sequence is added to the transposase product prior to recombination, for example, using a ligase.
[0017] In certain aspects of the methods herein, the method can further include sequencing at least a portion of the nucleic acid ligated via any suitable method, such as the methods provided herein. Optionally, the sequencing step can include sequencing at least a portion of a first sequence and at least a portion of a second sequence. In certain cases, the method can further include mapping at least a portion of the first sequence and at least a portion of the second sequence to a genome. In various cases, the method can further include performing a three-dimensional genome analysis using information from the sequencing.
[0018] In various aspects of the methods herein, the stabilized sample can be a cross-linked sample. Optionally, the stabilized sample can be a cross-linked cell. Optionally, the stabilized sample can be a cross-linked nucleus. Optionally, the stabilized sample can be a cross-linked chromatin. Optionally, the step of obtaining a stabilized sample can include obtaining a sample and stabilizing the sample. Optionally, the step of obtaining a stabilized sample can include obtaining a pre-stabilized sample. Optionally, the nucleic acid-binding protein can include chromatin or a component thereof.
[0019] In various aspects of the methods herein, the recombinase sites can include attP and attB integrase sites. In some cases, the first recombinase site can be different from the second recombinase site. In some cases, the first recombinase site can be an attP or attB integrase site. In some cases, the second recombinase site can be an attP or attB integrase site. In various cases, the first recombinase site is an attP integrase site and the second recombinase site is an attB integrase site. In various cases, the first recombinase site is an attB integrase site and the second recombinase site is an attP integrase site. In some cases, the first recombinase site and the second recombinase site can include transposase mosaic ends.
[0020] In various aspects of the methods herein, the linker can include additional sequences. In some cases, the linker sequence can include a barcode sequence. In some cases, the barcode sequence can indicate the partition of origin. In some cases, the barcode sequence can indicate the cell of origin. In some cases, the barcode sequence can indicate the cell population of origin. In some cases, the barcode sequence can indicate the organism of origin. In some cases, the barcode sequence can indicate the species of origin. In some cases, the linker can include an adapter. In some cases, the adapter can include a P5 sequence. In some cases, the adapter can include a P7 sequence.
[0021] In various aspects of the methods herein, the method can be completed in less than 1 day. In some cases, the method can be completed in less than 8 hours. In some cases, the method can be completed in less than 6 hours. In some cases, the method can be completed in 4 hours or less. In some cases, the method can be completed in 4 - 6 hours. In some cases, the method can be completed in 4 - 8 hours. In some cases, the method can be completed in 3 - 4 hours.
[0022] In various aspects of the methods herein, the methods may require a very low input of sample material. In some cases, the stabilized sample can contain 50,000 cells or fewer. In some cases, the sample can contain 40,000 cells or fewer. In some cases, the sample can contain 30,000 cells or fewer. In some cases, the sample can contain 20,000 cells or fewer. In some cases, the sample can contain at least 10,000 cells. In some cases, the sample can contain at least 20,000 cells. In some cases, the sample can contain at least 30,000 cells. In some cases, the sample can contain at least 40,000 cells. In some cases, the sample can contain from about 10,000 cells to about 50,000 cells. In some cases, the sample can contain from about 20,000 cells to about 50,000 cells. In some cases, the sample can contain from about 30,000 cells to about 50,000 cells. In some cases, the sample can contain from about 40,000 cells to about 50,000 cells. In some cases, the sample can contain from about 10,000 cells to about 40,000 cells. In some cases, the sample can contain from about 10,000 cells to about 30,000 cells. In some cases, the sample can contain from about 10,000 cells to about 20,000 cells. In some cases, the sample can contain from about 20,000 cells to about 50,000 cells. In some cases, the sample can contain from about 20,000 cells to about 40,000 cells. In some cases, the sample can contain from about 20,000 cells to about 30,000 cells. In some cases, the sample can contain from about 30,000 cells to about 50,000 cells. In some cases, the sample can contain from about 30,000 cells to about 40,000 cells.
[0023] In various aspects of the methods herein, the stabilized sample may contain nuclei. In some cases, the stabilized sample may contain 50,000 nuclei or fewer. In some cases, the sample may contain 40,000 nuclei or fewer. In some cases, the sample may contain 30,000 nuclei or fewer. In some cases, the sample may contain 20,000 nuclei or fewer. In some cases, the sample may contain at least 10,000 nuclei. In some cases, the sample may contain at least 20,000 nuclei. In some cases, the sample may contain at least 30,000 nuclei. In some cases, the sample may contain at least 40,000 nuclei. In some cases, the sample may contain from about 10,000 nuclei to about 50,000 nuclei. In some cases, the sample may contain from about 20,000 nuclei to about 50,000 nuclei. In some cases, the sample may contain from about 30,000 nuclei to about 50,000 nuclei. In some cases, the sample may contain from about 40,000 nuclei to about 50,000 nuclei. In some cases, the sample may contain from about 10,000 nuclei to about 40,000 nuclei. In some cases, the sample may contain from about 10,000 nuclei to about 30,000 nuclei. In some cases, the sample may contain from about 10,000 nuclei to about 20,000 nuclei. In some cases, the sample may contain from about 20,000 nuclei to about 50,000 nuclei. In some cases, the sample may contain from about 20,000 nuclei to about 40,000 nuclei. In some cases, the sample may contain from about 20,000 nuclei to about 30,000 nuclei. In some cases, the sample may contain from about 30,000 nuclei to about 50,000 nuclei. In some cases, the sample may contain from about 30,000 nuclei to about 40,000 nuclei.
[0024] Recombinase sites oriented as direct repeats In another aspect, provided is a method of nucleic acid processing, comprising the step of obtaining a stabilized sample comprising a nucleic acid molecule complexed with at least one nucleic acid binding protein. Next, the method can comprise the steps of cleaving the nucleic acid molecule into a plurality of segments comprising at least a first segment and a second segment, and ligating a first recombinase site to the first segment and the second segment. Then, the method comprises the step of contacting the first segment and the second segment with a linker comprising a second recombinase site in the presence of a recombinase, whereby a proximally ligated nucleic acid comprising a first sequence from the first segment, a linker sequence from the linker, and a second sequence from the second segment is generated, and the second recombinase site comprises two recombinase sites oriented as direct repeats. Optionally, the stabilized sample may not be sonicated.
[0025] In various aspects of the methods herein, the step of cleaving the stabilized nucleic acid may be performed using a transposase. Optionally, the cleavage may be performed in permeabilized cells. Optionally, the cleavage may be performed in permeabilized nuclei. Optionally, the transposase may be Tn5, Tn3, Tn7, sleeping beauty transposase, or a combination thereof. Optionally, the transposase may be Tn5 transposase.
[0026] In various aspects of the methods herein, the linker comprising the recombinase site may be contacted with the cleaved nucleic acid in the presence of a recombinase, and the recombinase site comprises two recombinase sites oriented as direct repeats. Optionally, the presence of the recombinase sites oriented as direct repeats can prevent the resulting product from forming a stable hairpin structure. Optionally, the proximally ligated nucleic acid does not form a hairpin loop. Optionally, the resulting product may be more readily sequenced than a product in which the recombinase sites are oriented as inverted repeats.
[0027] In various aspects of the methods herein, the first and second segments of the cleaved nucleic acid may be contacted with a linker containing an integrase site in the presence of a recombinase. Optionally, the recombinase may be an integrase. Optionally, the integrase may be PhiC31 integrase, Bxb1 integrase, or a combination thereof.
[0028] In various aspects of the methods herein, the method can further include sequencing at least a portion of the proximally ligated nucleic acid via any suitable method, such as the methods provided herein. Optionally, the sequencing step may include sequencing at least a portion of the first sequence and at least a portion of the second sequence. Optionally, the method further includes mapping at least a portion of the first sequence and at least a portion of the second sequence to the genome. Optionally, the method further includes performing a three-dimensional genome analysis using information from the sequencing step.
[0029] In various aspects of the methods herein, the stabilized sample may be a cross-linked sample. Optionally, the stabilized sample may be a cross-linked cell. Optionally, the stabilized sample may be a cross-linked nucleus. Optionally, the stabilized sample may be a cross-linked chromatin. Optionally, the step of obtaining a stabilized sample includes obtaining a sample and stabilizing the sample. Optionally, the step of obtaining a stabilized sample includes obtaining a pre-stabilized sample. Optionally, the nucleic acid-binding protein includes chromatin or a component thereof.
[0030] In various aspects of the methods herein, the recombinase sites can include attP and attB integrase sites. In some cases, the first recombinase site can be different from the second recombinase site. In some cases, the first recombinase site can be an attP or attB integrase site. In some cases, the second recombinase site can be an attP or attB integrase site. In various cases, the first recombinase site is an attP integrase site and the second recombinase site is an attB integrase site. In various cases, the first recombinase site is an attB integrase site and the second recombinase site is an attP integrase site. In some cases, the first recombinase site and the second recombinase site can include transposase mosaic ends.
[0031] In various aspects of the methods herein, the linker can include additional sequences. In some cases, the linker sequence can include a barcode sequence. In some cases, the barcode sequence can indicate the partition of origin. In some cases, the barcode sequence can indicate the cell of origin. In some cases, the barcode sequence can indicate the cell population of origin. In some cases, the barcode sequence can indicate the organism of origin. In some cases, the barcode sequence can indicate the species of origin. In some cases, the linker can include an adapter. In some cases, the adapter can include a P5 sequence. In some cases, the adapter can include a P7 sequence.
[0032] In various aspects of the methods herein, the method can be completed in less than 1 day. In some cases, the method can be completed in less than 8 hours. In some cases, the method can be completed in less than 6 hours. In some cases, the method can be completed in 4 hours or less. In some cases, the method can be completed in 4 to 6 hours. In some cases, the method can be completed in 4 to 8 hours. In some cases, the method can be completed in 3 to 4 hours.
[0033] In various aspects of the methods herein, the methods can require very low inputs of sample material. In some cases, the stabilized sample can contain 50,000 cells or fewer. In some cases, the sample can contain 40,000 cells or fewer. In some cases, the sample can contain 30,000 cells or fewer. In some cases, the sample can contain 20,000 cells or fewer. In some cases, the sample can contain at least 10,000 cells. In some cases, the sample can contain at least 20,000 cells. In some cases, the sample can contain at least 30,000 cells. In some cases, the sample can contain at least about 40,000 cells. In some cases, the sample can contain from about 10,000 cells to about 50,000 cells. In some cases, the sample can contain from about 20,000 cells to about 50,000 cells. In some cases, the sample can contain from about 30,000 cells to about 50,000 cells. In some cases, the sample can contain from about 40,000 cells to about 50,000 cells. In some cases, the sample can contain from about 10,000 cells to about 40,000 cells. In some cases, the sample can contain from about 10,000 cells to about 30,000 cells. In some cases, the sample can contain from about 10,000 cells to about 20,000 cells. In some cases, the sample can contain from about 20,000 cells to about 50,000 cells. In some cases, the sample can contain from about 20,000 cells to about 40,000 cells. In some cases, the sample can contain from about 20,000 cells to about 30,000 cells. In some cases, the sample can contain from about 30,000 cells to about 50,000 cells. In some cases, the sample can contain from about 30,000 cells to about 40,000 cells.
[0034] In various aspects of the methods herein, the stabilized sample may contain nuclei. In some cases, the stabilized sample may contain 50,000 nuclei or fewer. In some cases, the sample may contain 40,000 nuclei or fewer. In some cases, the sample may contain 30,000 nuclei or fewer. In some cases, the sample may contain 20,000 nuclei or fewer. In some cases, the sample may contain at least 10,000 nuclei. In some cases, the sample may contain at least 20,000 nuclei. In some cases, the sample may contain at least 30,000 nuclei. In some cases, the sample may contain at least 40,000 nuclei. In some cases, the sample can contain from about 10,000 nuclei to about 50,000 nuclei. In some cases, the sample can contain from about 20,000 nuclei to about 50,000 nuclei. In some cases, the sample can contain from about 30,000 nuclei to about 50,000 nuclei. In some cases, the sample can contain from about 40,000 nuclei to about 50,000 nuclei. In some cases, the sample can contain from about 10,000 nuclei to about 40,000 nuclei. In some cases, the sample can contain from about 10,000 nuclei to about 30,000 nuclei. In some cases, the sample can contain from about 10,000 nuclei to about 20,000 nuclei. In some cases, the sample can contain from about 20,000 nuclei to about 50,000 nuclei. In some cases, the sample can contain from about 20,000 nuclei to about 40,000 nuclei. In some cases, the sample can contain from about 20,000 nuclei to about 30,000 nuclei. In some cases, the sample can contain from about 30,000 nuclei to about 50,000 nuclei. In some cases, the sample can contain from about 30,000 nuclei to about 40,000 nuclei.
[0035] Proximity ligation for making concatemers Compositions, systems, and methods are provided herein that enable concatemer formation using proximity ligation. For example, a biological sample, such as a stabilized biological sample having a nucleic acid molecule complexed with a nucleic acid-binding protein, can be contacted with a dendrimer to form a complex. In another example, the biological sample can be stabilized by contacting it with a dendrimer to form a complex. Next, the nucleic acid molecule can be cleaved into a plurality of segments, for example, at least a first segment and a second segment. Thereafter, the plurality of segments can be joined at a plurality of junctions, for example, the first segment and the second segment can be joined at one junction.
[0036] In certain aspects of the methods herein, a biological sample, such as a stabilized biological sample having a nucleic acid molecule complexed with a nucleic acid-binding protein and a dendrimer. Optionally, the dendrimer is conjugated to psoralen. Optionally, the dendrimer is conjugated to azido-Peg4-N-hydroxysuccinimide (NHS) ester. Optionally, the NHS ester of the azido-Peg4-NHS ester reacts with a primary amine on the dendrimer to yield a dendrimer having a reactive azido group. Optionally, carboxylated beads (e.g., magnetic beads) are prepared by conjugating dibenzocyclooctyne-amine (DBCO)-Peg4-amine building blocks using 1-ethyl-3-(3-dimethylaminopropyl)carbodiimide (EDC) / sulfo-NHS chemistry. These prepared beads can be used, for example, to isolate the dendrimer by magnetic separation prior to proximity ligation.
[0037] In some cases, the dendrimer is modified with a compound or contacts a compound. For example, in some cases, the dendrimer is modified with psoralen. In some cases, the psoralen includes N-hydroxysuccinimide (NHS) ester-conjugated psoralen. In some cases, the dendrimer includes a polyamidoamine (PAMAM) dendrimer. In some cases, the dendrimer is modified with a crosslinking agent, such as chloromethine, cyclophosphamide, chlorambucil, uramustine, melphalan, bendamustine, bis(2-chloroethyl)ethylamine, bis(2-chloroethyl)methylamine, tris(2-chloroethyl)amine, isofamide, carmustine, lomustine, streptozocin, busulfan, cisplatin, carboplatin, cicycloplatin, eptaplatin, lobaplatin, miltiplatin, nedaplatin, oxaliplatin, picoplatin, satraplatin, triplatin tetranitrate, procarbazine, altretamine, dacarbazine, mitozolomide, temozolomide, mitomycin C, nitrous acid, formaldehyde, acetylaldehyde, doxorubicin, daunorubicin, epirubicin, or idarubicin. In some cases, the dendrimer is modified with an intercalator, an antibiotic, or a minor groove binder.
[0038] The methods herein can include a step of separating a compound from the dendrimer. For example, a compound such as psoralen can be separated from the dendrimer using heat. In some cases, a compound such as psoralen is separated from the dendrimer using alkaline conditions or high pH. Alternatively, a compound such as psoralen is separated from the dendrimer using heat and alkaline conditions. The compound (e.g., psoralen) can further be separated from the dendrimer using UV radiation.
[0039] Any suitable dendrimer can be used in the methods of the present specification. The molecular weight of the dendrimer can be from about 5 kilodaltons (kDa) to about 125 kDa. In some cases, the molecular weight of the dendrimer is from 6 kDa to 8 kDa. In some cases, the molecular weight of the dendrimer is from 25 kDa to 35 kDa. In some cases, the molecular weight of the dendrimer is from 110 kDa to 125 kDa. In some cases, the dendrimer contains 32 to 512 reactive groups. In some cases, the dendrimer contains about 32 reactive groups. In some cases, the dendrimer contains about 128 reactive groups. In some cases, the dendrimer contains about 512 reactive groups. In some cases, the dendrimer is a Gen3 dendrimer. In some cases, the dendrimer is a Gen5 dendrimer. In some cases, the dendrimer is a Gen7 dendrimer.
[0040] The method of the present specification can combine at least a part of segments into a concatemer. For example, to form a concatemer, at least 2 segments, at least 3 segments, at least 4 segments, at least 5 segments, at least 6 segments, at least 7 segments, at least 8 segments, at least 9 segments, at least 10 segments, or more can be combined. In some cases, oligonucleotides are ligated between each segment. In some cases, the oligonucleotide is a crosslinking oligonucleotide. In some cases, the oligonucleotide is an adapter oligonucleotide. In some cases, the oligonucleotide is a punctuation oligonucleotide. In some cases, the crosslinking oligonucleotide, adapter oligonucleotide, and / or punctuation oligonucleotide contain a barcode sequence. In some cases, the crosslinking oligonucleotide, adapter oligonucleotide, and / or punctuation oligonucleotide are modified with a dibenzo-cyclooctyne (DBCO) moiety. In some cases, the DBCO moiety facilitates copper free click chemistry. In some cases, multiple oligonucleotides are ligated continuously between each segment. The ligation can result in a stabilized biological sample of a sample, cell, nucleus, chromosome, or nucleic acid molecule that receives the unique sequence of the oligonucleotide (e.g., crosslinking oligonucleotide).
[0041] In some cases, after contacting a dendrimer with a stabilized biological sample to form a complex, the complex is photoactivated, for example, by exposing the complex to UV radiation having a wavelength of about 360 nm, thereby generating a crosslinked complex. In some cases, the crosslinking is reversable and leaves no adducts on the nucleic acid.
[0042] The methods herein can further include subjecting the plurality of segments to size selection to obtain a plurality of selected segments. The size selection herein can include any suitable range of segment sizes.
[0043] Cleavage in the methods provided herein can be performed using any suitable method, e.g., by using a nuclease or deoxyribonuclease (DNase). In some cases, the DNase can include DNase I, DNase II, micrococcal nuclease, restriction endonuclease, or combinations thereof.
[0044] The stabilized biological samples of the methods herein can be stabilized by treatment with a stabilizer or a cross-linking reagent. In some cases, the cross-linking agent is a chemical fixative such as formaldehyde, psoralen, disuccinimidyl glutarate (DSG), ethylene glycol bis(succinimidyl succinate) (EGS), ultraviolet light, or a combination thereof. In some cases, the cross-linking agent includes chloromethine, cyclophosphamide, chlorambucil, uramustine, melphalan, bendamustine, bis(2-chloroethyl)ethylamine, bis(2-chloroethyl)methylamine, tris(2-chloroethyl)amine, isofamide, carmustine, lomustine, streptozocin, busulfan, cisplatin, carboplatin, cicycloplatin, eptaplatin, lobaplatin, miltiplatin, nedaplatin, oxaliplatin, picoplatin, satraplatin, triplatin tetranitrate, procarbazine, altretamine, dacarbazine, mitozolomide, temozolomide, mitomycin C, nitrous acid, formaldehyde, acetylaldehyde, doxorubicin, daunorubicin, epirubicin, or idarubicin. In some cases, the cross-linking agent includes an intercalator, an antibiotic, or a minor groove binder. The stabilized biological sample can be a cross-linked paraffin-embedded tissue sample. In some cases, the stabilized biological sample includes stabilized intact cells or stabilized intact nuclei. In some cases, the method includes a step of lysing the cells and / or nuclei in the stabilized biological sample. The cutting step of the methods herein can be performed prior to the lysis of the intact cells or intact nuclei.
[0045] The methods of this specification can be performed on stabilized biological samples containing a small number of cells. For example, in some cases, the stabilized biological sample contains less than 3,000,000 cells. The stabilized biological sample can contain less than about 1,000,000 cells, less than about 500,000 cells, less than about 400,000 cells, less than about 300,000 cells, less than about 200,000 cells, less than about 100,000 cells, or less.
[0046] In aspects of the methods of this specification, the method can further include obtaining at least some sequences on each side of a junction to generate a first read pair. In addition, the method can further include mapping the first read pair to a set of contigs and determining a path through the set of contigs that represents the order and / or orientation relative to the genome. Alternatively, or in combination, the method can include mapping the first read pair to a set of contigs and determining the presence of a structural variant or loss of heterozygosity in the stabilized biological sample from the set of contigs. Alternatively, or in combination, the method can include mapping the first read pair to a set of contigs and assigning phases to variants in the set of contigs. Alternatively, or in combination, the method can include mapping the first read pair to a set of contigs, determining the presence of variants in the set of contigs from the set of contigs, and performing one or more steps selected from identifying a disease stage, prognosis, or course of treatment for the stabilized biological sample, selecting a drug based on the presence of the variant, or identifying the drug efficacy for the stabilized biological sample.
[0047] In aspects of the methods herein, proximity ligation can be performed using click chemistry, including copper-free click chemistry, such as using DBCO-modified crosslinking oligonucleotides joined between each segment of the concatemer. The concatemer can then be joined, for example, via a dendrimer. To enrich the ligated molecules, the features of the crosslinking oligonucleotides can be targeted. In one example, the DBCO-containing oligonucleotide can be reacted with an azide-biotin moiety that can be isolated with a streptavidin substrate such as beads. In another example, the DBCO-containing oligonucleotide can be reacted with an azide-modified NHS-S-S-dPEG4-biotin containing a disulfide bond, and an azide-PEG3 amine can be used to add an azide to the NHS-S-S-dPEG4-biotin, and this disulfide bond can be reduced using DTT and heating, for example, by heating at 70 °C for about 10 minutes, to isolate nucleic acids for library preparation.
[0048] In aspects of the methods herein, the dendrimer with which the nucleic acid fragments are in contact can be separated or isolated from the remaining nucleic acids in the sample prior to proximity ligation of the nucleic acid fragments. This step can ensure that the concatemer formed by proximity ligation contains fragments that were in contact with the same dendrimer. This can mean that all segments of a given concatemer were in proximity to each other in the original stabilized sample. Thus, such an approach can provide not just pairwise information about which nucleic acid regions were in proximity to which other regions, but much more complex proximity information, for example, information that three, four, five, six, seven, eight, nine, ten, or more nucleic acid regions were all in proximity to each other.
[0049] In some cases, by separating or isolating the dendrimer in contact with the nucleic acid fragment, barcoding or tagging of these fragments can be enabled instead of proximity ligation. The fragments associated with a given dendrimer can be barcoded or tagged, for example, in droplets or wells. After sequencing, the sequences can be associated based on their barcodes, and the proximity information can be derived based on the barcodes rather than being present in the same concatemer as described above. This proximity information can be used as discussed herein. In one example, the dendrimer complexes with the nucleic acids in the sample to stabilize them, then the nucleic acids are fragmented, and then the dendrimer is isolated with the nucleic acid fragments with which it is complexed and encapsulated in droplets. The nucleic acids in the droplets are labeled with droplet-specific barcodes or labels, then the nucleic acids are sequenced, and the barcode or label information is used to associate the fragments that were in proximity to each other in the sample.
[0050] Long non-coding RNA analysis Methods for analyzing long non-coding RNA binding sites are provided herein. In some cases, such methods include obtaining a stabilized biological sample comprising a DNA molecule complexed with at least one nucleic acid binding protein and at least one non-coding RNA. Next, the method can include contacting the DNA molecule with Tn5 transposase and an oligonucleotide comprising mosaic ends and a detectable label, thereby fragmenting the DNA molecule and ligating the oligonucleotide to the ends of the fragmented DNA molecule. The fragments can be contacted with T4 RNA ligase, thereby ligating the non-coding RNA to the oligonucleotide and reversing the cross-linking. Then, double-stranded DNA fragments can be generated by extending the ligated RNA with reverse transcriptase. The double-stranded DNA fragments can then be contacted with an endonuclease conjugated to an agent that binds to the detectable label, thereby digesting the DNA near the detectable label. Thereafter, sequencing adapters can be ligated to generate a sequencing library. In some cases, the oligonucleotide is adenylated at one end to facilitate ligation to the non-coding RNA. In some cases, the oligonucleotide further comprises a barcode. In some cases, the stabilized biological sample is contacted with RNase H prior to transposase treatment.
[0051] In aspects of the methods herein, in some cases, the non-coding RNA is a long non-coding RNA. In some cases, the non-coding RNA is an enhancer RNA. In some cases, the non-coding RNA is a miRNA. In some cases, the non-coding RNA is a Y RNA. In some cases, the non-coding RNA is an RNase P. In some cases, the non-coding RNA is a piRNA. In some cases, the non-coding RNA is an Xist.
[0052] In aspects of the methods herein, optionally, the detectable label comprises a modified nucleotide capable of click chemistry reaction. Optionally, the detectable label comprises biotin. Optionally, the agent comprises an antibody, protein A, protein G, or streptavidin. Optionally, the DNA bound to the non-coding RNA is concentrated prior to further analysis.
[0053] In aspects of the methods herein, an endonuclease is used to cleave exogenous sample DNA prior to analysis. Optionally, the endonuclease comprises DNaseI, DNaseII, micrococcal nuclease, restriction endonuclease, or a combination thereof.
[0054] In aspects of the methods herein, the sequences are obtained from double-stranded DNA fragments containing non-coding RNA. Any suitable sequencing method further comprising the methods described herein can be used.
[0055] A variety of suitable stabilized biological samples are contemplated for use in the methods herein. The stabilized biological samples, as specifically described elsewhere herein, are cross-linked using a cross-linking agent, such as a fixative, or UV light. For example, optionally, the stabilized biological sample is a cross-linked paraffin-embedded tissue sample. Optionally, the stabilized biological sample comprises a stabilized cell lysate. Optionally, the stabilized biological sample comprises stabilized intact cells. Optionally, the stabilized biological sample comprises stabilized intact nuclei.
[0056] Evaluation of nucleic acid three-dimensional structure Compositions, systems, and methods are disclosed herein related to determining the physical three-dimensional structure of nucleic acids in a cell, such as a single cell or cell population, distinguishable from the physical three-dimensional structure of a second cell or cell population. Through the practice of the disclosure herein, nucleic acid molecules indicative of three-dimensional nucleic acid relative positions are generated and, optionally, tags (e.g., nucleic acid barcodes) are provided to identify a common origin cell or population for a plurality of molecules.
[0057] Through the implementation of the methods disclosed herein, nucleic acids can be obtained such that all or at least some of their three-dimensional arrangements in a cell are preserved. Cleaving the exposed nucleic acid loops of such nucleic acids exposes internal segment ends that are randomly recombined with each other, such that physically proximal exposed ends are more likely to bind to each other (proximity ligation). Thus, by determining which exposed ends become bound to each other, it is possible to obtain useful data regarding the physical proximity of end-adjacent nucleic acids in the native cellular arrangement.
[0058] A related approach is disclosed, for example, in US9434985B2, published September 6, 2016, to Dekker et al., which is hereby incorporated by reference in its entirety.
[0059] Through the implementation of the methods disclosed herein, paired-end library components can be further tagged with sequence information indicative of cell origin, or otherwise provided, such that differences in three-dimensional structure between individual cells of a population are readily distinguishable for the population of cells, or such that differences in three-dimensional structure between a first population of cells and a second population of cells are readily distinguishable even if they are analyzed simultaneously. Tags can include, for example, nucleic acid barcodes. In some cases, a tag can include a junction between two nucleic acid segments that are not adjacent in the genome. Nucleic acid molecules can be generated such that when fully or partially sequenced, at least some genomic sequence sufficient to map each genomic end to its genomic locus is obtained, and further, tagging sequences or linked sequences sufficient to identify the cell or cell population of accurate or likely origin are obtained. Thus, useful sequence information for two regions of the genome that are physically proximal to each other can be obtained, along with useful information about the cell or cell population in which this physical three-dimensional structure occurs, and this can be evaluated in the context of other physical three-dimensional structure information that occurs simultaneously in that cell or cell population.
[0060] It is possible to stabilize genomic nucleic acids or other nucleic acids in cells, and for eukaryotic cells, the nucleus is optionally isolated according to known methods incorporated herein or otherwise known methods.
[0061] Nucleic acids consistent with the disclosure herein can include any number of cellular nucleic acids, such as prokaryotic primary genomic or plasmid nucleic acids, eukaryotic nuclear, mitochondrial or plastid nucleic acids, or, optionally, cytoplasmic nucleic acids in a sample, such as rRNA, mRNA, or exogenous nucleic acids, such as viruses or other pathogens or other exogenous nucleic acids of the sample.
[0062] The stabilized nucleic acids can, optionally, be distributed such that at least some of the nucleic acids are distributed into individual partitions. Exemplary partitions include wells, droplets in an emulsion, or surface locations (e.g., array spots, beads, etc.) that include discrete patches of differentially addressable linker molecules as described elsewhere herein. Additional partitions are contemplated and are consistent with the methods, compositions, and systems disclosed herein.
[0063] Stabilized nucleic acids can be fragmented to expose internal cleavage sites for later recombination in order to obtain nucleic acid placement information for a particular cell. Many fragmentation approaches are known and are consistent with the disclosure herein. Nucleic acids can be fragmented using one or a plurality of populations of restriction endonucleases, programmable endonucleases such as CRISPR / Cas molecules bound to guide RNA, non-specific endonucleases (e.g., DNase), tagmentation, shearing, sonication, heating, or other mechanisms. In some cases, DNase is non-sequence specific. In some cases, DNase is active against both single-stranded DNA and double-stranded DNA. In some cases, DNase is specific for double-stranded DNA. In some cases, DNase is preferred over double-stranded DNA. In some cases, DNase is specific for single-stranded DNA. In some cases, DNase is preferred over single-stranded DNA. In some cases, DNase is DNaseI. In some cases, DNase is DNaseII. In some cases, DNase is selected from one or more of DNaseI and DNaseII. In some cases, DNase is micrococcal nuclease. In some cases, DNase is selected from one or more of DNaseI, DNaseII, and micrococcal nuclease. Other suitable nucleases are also within the scope of the present disclosure.
[0064] In particular, the disclosure of W02014121091A1, published on August 7, 2014, by Green et al. (subsequently published as US20150363550A1 on December 17, 2015, and as US10089437B2 on October 2, 2018), is hereby incorporated by reference in its entirety. Similarly, the disclosure of W02016019360A1, published on February 4, 2016, by Fields et al. (subsequently published as US20170335369A1 on November 23, 2017), is hereby incorporated by reference in its entirety. Similarly, the disclosure of WO2017147279A1, published on August 31, 2017, by Green et al., is hereby incorporated by reference in its entirety.
[0065] The nucleic acid can bind to the surface either before or after binding. Exemplary surfaces include, but are not limited to, beads, arrays, and wells. In some cases, the surface is a solid-phase reversible immobilization (SPRI) surface such as SPRI beads. By binding the nucleic acid to the surface before binding, the performance of downstream processes can be improved, for example, by reducing interchromosomal ligation or binding and increasing intrachromosomal ligation or binding.
[0066] The nucleic acid may be immunoprecipitated either before or after binding. Such methods can include a step of fragmenting chromatin and a step of contacting the fragments with an antibody that specifically recognizes and binds to acetylated histone, particularly H3. Examples of such antibodies include, but are not limited to, anti-acetylated histone H3 available from Upstate Biotechnology, Lake Placid, NY. The polynucleotide from the immunoprecipitate can then be collected from the immunoprecipitate. Similar targeted enrichment methods can also be employed using target-specific compounds including, but not limited to, aptamers, oligonucleotides, or other nucleic acid probes, and nucleic acid-guided nucleases (e.g., Cas family enzymes such as Cas9 including catalytically inactive or "dead" nucleases).
[0067] A concatenated nucleic acid, e.g., a concatenated nucleic acid having a barcode, a partition-specific sequence, or a partition identification sequence, can be bound to the exposed internal ends to generate a nucleic acid segment having a left genomic segment, often a concatenated region having a partition-specific sequence or a partition identification sequence (e.g., a nucleic acid barcode), and a right genomic segment, where the left and right genomic segments map to genomic segments that are physically proximal in the source cell.
[0068] Before ligating the exposed nucleic acid ends, the ends can be processed. Such processing can include end polishing or blunt ending. Nucleic acid ends with blunt ends exposed can be ligated, for example, directly to other nucleic acid ends with blunt ends exposed, or to an adapter or linker. Such processing can include, for example, generating overhangs by tailing (e.g., A-tailing or adenylation). In one example, the overhang is 1 nucleotide in size. In one example, the overhang is a single A nucleotide. Nucleic acid ends with tailed ends exposed can be ligated, for example, directly to other nucleic acid ends with tailed ends exposed, or to an adapter or linker. In some cases, blunt ending or tailing can incorporate affinity-tagged nucleic acids such as biotinylated nucleic acids. The affinity tag can be used, for example, in downstream capture or enrichment steps. In other cases, blunt ending or tailing can be performed without incorporating affinity-tagged nucleic acids (e.g., without biotinylated nucleic acids). Subsequently, the affinity tag can be added, if desired, to an adapter or linker (e.g., by crosslinking). In one example, the exposed nucleic acid is blunt-ended at the ends, overhangs are generated, and the exposed ends are ligated via a crosslinking oligo.
[0069] The ligation can be direct, for example, via ligation.
[0070] The ligation can be via a linker or crosslink, for example, by ligation of one or more linker or crosslink nucleic acids that bind one exposed nucleic acid end to another nucleic acid end.
[0071] The ligation can be via the use of a capping nucleic acid adapter segment, for example, that matches the incorporation of a recombinase such as integrase or transposase. An adapter having a recombinase site can be added to the exposed nucleic acid ends, and these ends can then be ligated, for example, by recombination.
[0072] As an example, for phiC31 integrase barcode delivery, linkers such as cell identification linkers or cell-specific linkers (e.g., nucleic acid barcodes) can be enzymatically added as follows.
[0073] After exposure of the internal nucleic acid ends, the integrase site can be ligated to the exposed nucleic acid ends, such as internal ends or exposed linear chromosomal ends, e.g., those with telomeres removed. Exemplary integration sites are nucleic acids containing the attP phiC31 integrase integration site, or the attP integration site, although other integration sites are consistent with the disclosure herein. Ligation results in a population of nucleic acid fragments, at least some of which individually contain cell nucleic acid segments bounded at each end by an integration site, such as a segment containing the attP segment. In various embodiments, one or both of fragmentation and attachment of the integration site occur before partitioning, or one or both of fragmentation and attachment of the integration site occur after partitioning.
[0074] Alternatively, transposases such as Tn3, Tn5, Tn7, or sleeping beauty transposase can be used for barcode delivery. After exposure of the internal nucleic acid ends, the mosaic ends can be ligated to the exposed nucleic acid ends, such as internal ends or exposed linear chromosomal ends, e.g., those with telomeres removed. Exemplary mosaic ends are nucleic acids containing the Tn5 mosaic end, or the Tn5 mosaic end, although other mosaic ends are consistent with the disclosure herein. Ligation results in a population of nucleic acid fragments, at least some of which individually contain cell nucleic acid segments bounded at each end by a mosaic end, such as the Tn5 mosaic end.
[0075] In various embodiments, one or both of fragmentation and joining of mosaic ends occur before partitioning, or one or both of fragmentation and joining of mosaic ends occur after partitioning. In an exemplary system for single cell HiC (or other proximity ligation techniques), integrase-mediated in-aggregate ligation is used. Single cell nuclei are encapsulated in a first set of partitions in combination with integrase. The partitions are, in this case, droplets in an emulsion. By subjecting the nuclei to strand breakage, internally exposed ends are generated and local three-dimensional information is preserved. Adapters are ligated to the exposed internal ends. The adapters optionally include exonuclease-resistant ends. In this embodiment, the adapters do not convey partition identification information. In a second set of partitions, linkers having partition identification sequences such as unique molecular identifiers (UMIs) are encapsulated and optionally subjected to amplification and cleavage-directed linearization. The first and second sets of partitions are combined in a ratio of about 1:1 or under conditions such that nucleic acids from two cells are unlikely to be combined into a single resulting partition.
[0076] Recombinase sites, such as integrase sites or mosaic ends, can optionally be carried on unmodified single-stranded or double-stranded fragments that are ligated onto internal nucleic acid ends. Alternatively, some single-stranded or double-stranded fragments having attP sequences or mosaic ends, such as integration sites for Tn3, Tn5, Tn7, or sleeping beauty transposase mosaic ends, can include at least one modification, such as a modification that interferes with exonuclease or other nuclease activity, to facilitate subsequent cleanup of the sequencing library. An example is a thiophosphate modification to preclude exonuclease degradation of fragments with double-stranded fragments having integration sites added to each end.
[0077] Often, recombinase sites such as integration sites or mosaic ends are non-specific in that the sequences of such integration sites or mosaic ends, e.g., the attP sequence or the Tn3, Tn5, Tn7, or sleeping beauty transposase mosaic ends, are not used to specify the cellular source of the adjacent nucleic acid. Alternatively, often after nucleic acid partitioning, an adapter having a separate, specific sequence or cell identification sequence (e.g., a nucleic acid barcode) adjacent to the integration site or mosaic end is provided to the partition, or a separate integration site or mosaic end can be provided, such that the nucleic acid of the first partition receives an integration segment or mosaic end having a first identification segment, and the nucleic acid segment of the second partition receives an integration segment having a second identification segment.
[0078] A fragment having a recombinase boundary such as a boundary containing an integrase attP segment can then be contacted in a common solution with an integration site such as the attB phiC31 integration site. For example, the integrating enzyme can comprise phi31 integrase, the integration boundary can comprise an attP segment, and the integration site can comprise an attB integration site. Alternatively, the fragment has a mosaic end boundary such as Tn3, Tn5, Tn7, or a sleeping beauty transposase mosaic end boundary.
[0079] When an attB integration site, or a recombinase site such as a Tn3, Tn5, Tn7, or sleeping beauty transposase mosaic end, is adjacent to a contiguous segment having a sequence that identifies a partition or cell, such as one specific to a segment or cell source (e.g., a nucleic acid barcode), the sequence identifies adjacent cell nucleic acids as arising from a particular or common cell source or partition, and thus, multiple exposed ends from a common cell joined by a common cell identification segment or partition identification segment can be readily identified as arising from a common cell even if they are bulked with fragments of a second partition prior to or simultaneously with sequencing.
[0080] When a cell identification sequence is delivered via a recombinase site boundary fragment, integration or translocation is preferably performed after partitioning. The nucleic acid content of at least some partitions can thereby be identified by the cell identification sequence of its linker, and thus, even after nucleic acids from multiple cell sources are bulked for sequencing, the internal end pairs and the proximity information assigned to the vicinity where they map in a contig set that contains a mostly or fully sequenced genome can be assigned to a common cell that is distinguished from at least one other cell of the sample, thereby enabling the establishment of differences in the predicted nucleic acid three-dimensional structures.
[0081] Recombinant site boundary fragments variously include a left boundary fragment and a right boundary fragment (e.g., attB site or Tn3, Tn5, Tn7, or sleeping beauty transposase mosaic end) joined by a linker region that optionally includes a cell or partition designation sequence (e.g., nucleic acid barcode). The linker region optionally further includes a moiety to facilitate subsequent isolation. Many affinity tags or modified bases are consistent with the disclosure herein. Exemplary moieties facilitate physical or chemical isolation of the linker after integrase or transposase treatment. Any number of affinity tags, such as one or more biotin tags that can facilitate avidin or streptavidin-based isolation, are consistent with the disclosure herein. Alternatively, any antigen, receptor, or ligand that facilitates isolation without interfering with integrase or transposase activity is suitable for some embodiments herein.
[0082] As described above, some library generation approaches include cleanup steps such as selectively removing unincorporated reagents. For example, exonuclease treatment is often used to selectively remove unbound linker molecules, genomic fragments that do not bind to any integration sites, or both unbound linker molecules and genomic fragments that do not bind to any integration sites. Genomic fragments ligated to integration site fragments having exonuclease-resistant modifications such as a thiophosphate backbone are resistant to exonuclease digestion from their ends, and nucleic acid molecules joined at both ends by integration site fragments having exonuclease-resistant modifications such as a thiophosphate backbone are resistant to digestion at both ends and can survive exonuclease treatment.
[0083] Alternatively or in combination, some linker molecules contain inverse affinity tags on the opposite side of a recombination site, such as an attP integration site, or the mosaic ends of Tn3, Tn5, Tn7, or sleeping beauty transposase, and the inverse affinity tags are removed following the success of the recombination reaction. In such cases, unwanted reagents can be removed by contacting them with the binding partner of the inverse affinity tag.
[0084] Integrase activity partially disrupts both integration sites, such as attB and attP sites, as part of the integration event. Thus, by designing primers that anneal to the ligated adapter sites, such as attP integration sites, alone or in combination with linker-based isolation, cell or aliquot identification information and internal end-adjacent information can be amplified, and in some cases, a cloned amplicon spanning at least one linker can be generated to facilitate sequencing or other downstream analysis.
[0085] After library generation and optionally library cleanup, the nucleic acids can be fully or partially sequenced to obtain sufficient information for cell identification or cell-specific three-dimensional nucleic acid position assessment. As described above, sequencing is preferably performed to obtain at least some genomic sequences sufficient to map each genomic end of the library components to their genomic loci, and further to obtain sufficient linked sequences to identify the exact or possible starting cells. Thus, useful sequence information for two regions of the genome that are physically close to each other can be obtained, along with useful information about the cell in which this physical three-dimensional structure occurs, and it can be evaluated in the context of other physical three-dimensional structure information that occurs simultaneously in that cell. In many cases, this information is obtained by paired-end sequencing rather than full-length sequencing, but both approaches and other approaches are consistent with the disclosure herein.
[0086] Compositions and methods for determining the physical three-dimensional structure of nucleic acids in cells, such as single cells, distinguishable from the physical three-dimensional structure on a second cell, can be implemented on many systems consistent with the disclosure herein. Some systems involve the distribution of fixed cell nucleic acid material in the first droplet of an emulsion or in a well, e.g., on a well plate. These droplets optionally contain recombinase sites such as integrase sites or mosaic ends modified to be exonuclease resistant as described herein, as well as integrase or transposase enzymes and ligase enzymes. Separately, linker nucleic acid molecules can be configured for delivery to the first droplet of an emulsion. The linker nucleic acid is optionally distributed to droplets of a second emulsion or a second well and can be optionally amplified, e.g., using rolling circle amplification, to generate multiple copies of a given linker molecule / emulsion droplet and processed to do so.
[0087] Subsequently, the second emulsion droplet and the first emulsion droplet can be paired to assemble integrase or transposase ligation nucleic acid fragments using an integrase or transposase compatible linker, often showing uniform labeling per droplet. However, droplets having two or more identifiers per nucleic acid sample can still potentially yield meaningful data, especially when data analysis indicates the presence of multiple types of tags in the droplet.
[0088] As an alternative to the conjugate, in some cases, the integrase or transposase compatibility linker can be delivered as a colony of solid particles in a reagent stream that contacts the first emulsion droplet via coalescence from a droplet to a stream, such as those described in US20170335369A1, published November 23, 2017, which is hereby incorporated by reference in its entirety. The linker nucleic acid can be optionally amplified on the solid particles or in a gel. The first emulsion droplet can be coalesced into the stream, and the second emulsion droplet can be recovered by segmenting or partitioning the stream such that a desired ratio of nucleic acid cluster to linker particles, such as 1:1, greater than 1:1, or less than 1:1, is obtained.
[0089] Alternatively, some systems and methods include dispensing fixed cell nucleic acid material into wells of a chip or plate and then delivering the linker nucleic acid, unamplified or amplified as described above, to the partition.
[0090] Alternatively, in some cases, the delivery of the linker nucleic acid is not temporally separated from the partitioning. Rather, the linker nucleic acid or the enzyme activity or factors required for the enzyme activity are isolated until a specific treatment, such as heat, electromagnetic activation, or other administration, that temporarily activates the enzyme activity that results in a covalent bond to the exposed end of the nucleic acid sample of the linker, such as via the linker.
[0091] Many integrase enzymes are consistent with the disclosure herein. PhiC31 integrase, such as that commercially available from ThermoFisher, exhibits many advantages for the practice of the methods, operation of the systems, and use in the compositions herein. Some of the advantages of this integrase are as follows. It uses small integration sites (attB / attP). The enzyme itself is a small single polypeptide. Integration is irreversible without using a separate enzyme to excise the integration event. The activity is high and the enzyme is easily manipulated to change the activity. Nevertheless, its use is not required to the exclusion of other enzymes, as many integration systems are consistent with the disclosure herein. Aspects of the disclosure may be described with respect to PhiC31 integrase, but the use of any suitable enzyme is contemplated.
[0092] Many transposase enzymes are consistent with the disclosure herein. Tn5 transposase, such as that commercially available from Lucigen, exhibits many advantages for the practice of the methods, operation of the systems, and use in the compositions herein. Some of the advantages of this transposase are as follows. Tn5 uses a 19bp mosaic end recognition sequence and the insertions are mostly unbiased and stable. Tn5 can be delivered to cells for in vivo transposition or to isolated nucleic acids for in vitro reactions. Nevertheless, its use is not required to the exclusion of other enzymes, as many transposase systems, such as Tn3, Tn7, or sleeping beauty transposase, are consistent with the disclosure herein. Aspects of the disclosure may be described with respect to Tn3, Tn5, Tn7, or sleeping beauty transposase, but the use of any suitable enzyme is contemplated.
[0093] The array information obtained from library components is evaluated by many approaches such as those in the context of Hi-C, Chicago (registered trademark) in vitro proximity ligation, or other three-dimensional structure analysis. Importantly, the frequency of cell-specific read pairs can be obtained such that the frequency of end-adjacent array mapping to a particular region of the genome or a particular contig can be evaluated cell-specifically. That is, the cell-specific occurrence of possible three-dimensional structures can be evaluated. In some cases, the cell-specific intensity of a signal correlated with the cell-specific distance in the three-dimensional structure can also be evaluated, such that it can be concluded that a particular region of the nucleic acid is relatively close to a second cell in which they are equivalent but "weak" or more distantly proximate, while there is no signal indicating proximity among a third cell. That is, both qualitative and quantitative evaluations of the three-dimensional structure are consistent with the disclosure herein. In some cases, the proximity of one region to a second region is evaluated, at least in part, by counting the number of cluster components of a first cluster that occur simultaneously in paired-end reads that pair with cluster components of a second cluster, particularly in library components that share a common partition identification sequence such as a unique partition tag.
[0094] The configuration information need not be generated through multiple occurrences of the same end-adjacent array in multiple library components. Rather, in some cases, end-adjacent arrays that map near (to a common "cluster") a second end-adjacent array mapping site can enhance three-dimensional structure evaluation when both members of the cluster map to non-identical regions of a second cluster on a second region of a nucleic acid reference such as the genome.
[0095] In some cases, the methods disclosed herein are used to label and / or associate polynucleotides or sequence segments thereof and utilize that data for various applications. In some cases, the present disclosure provides methods for generating highly contiguous and accurate human genome assemblies having read pairs of less than about 10,000, less than about 20,000, less than about 50,000, less than about 100,000, less than about 200,000, less than about 500,000, less than about 1 million, less than about 2 million, less than about 5 million, less than about 10 million, less than about 20 million, less than about 30 million, less than about 40 million, less than about 50 million, less than about 60 million, less than about 70 million, less than about 80 million, less than about 90 million, less than about 100 million, less than about 200 million, less than about 300 million, less than about 400 million, less than about 500 million, less than about 600 million, less than about 700 million, less than about 800 million, less than about 900 million, or less than about 1 billion. In some cases, the present disclosure provides methods for phasing or assigning physical linkage information to about 50%, about 60%, about 70%, about 75%, about 80%, about 85%, about 90%, about 91%, about 92%, about 93%, about 94%, about 95%, about 96%, about 97%, about 98%, about 99%, or more of the heterozygous variants in the human genome with an accuracy of about 50%, about 60%, about 70%, about 75%, about 80%, about 85%, about 90%, about 91%, about 92%, about 93%, about 94%, about 95%, about 96%, about 97%, about 98%, about 99%, or more.
[0096] In some embodiments, the compositions and methods described herein enable the investigation of metagenomes, e.g., those found in the human gut. Thus, it is possible to investigate the partial or complete genomic sequences of some or all of the organisms inhabiting a given ecological environment. Examples include random sequencing of all gut microbes, microbes found in specific regions of the skin, and microbes surviving at toxic waste sites. The composition of microbial populations in these environments can be determined using the compositions and methods described herein, as well as the interrelated biochemical aspects encoded by their respective genomes. The methods described herein can enable metagenomic studies from complex biological environments, e.g., biological environments containing 2, 3, 4, 5, 6, 7, 8, 9, 10, 12, 15, 20, 25, 30, 40, 50, 60, 70, 80, 90, 100, 125, 150, 175, 200, 250, 300, 400, 500, 600, 700, 800, 900, 1000, 5000, 10000, or more organisms and / or variants of organisms.
[0097] Thus, the methods disclosed herein can be applied to intact human genomic DNA samples, but also to a wide variety of nucleic acid samples such as reverse transcribed RNA samples, cell-free DNA samples, cancer tissue samples, crime scene samples, archaeal samples, non-human genomic samples, or environmental samples such as environmental samples containing genetic information from more than one organism, e.g., organisms that are not easily cultured under laboratory conditions.
[0098] The high accuracy required for cancer genome sequencing can be achieved using the methods and systems described herein. Incorrect reference genomes can make basecalling difficult when sequencing cancer genomes. Heterogeneous samples and small starting materials, e.g., samples obtained by biopsy, pose further difficulties. Furthermore, the detection of large-scale structural variants and / or loss of heterozygosity is often essential not only for cancer genome sequencing, but also for the ability to distinguish somatic variants from basecalling errors.
[0099] The systems and methods described herein can generate accurate long reads from complex samples containing two, three, four, five, six, seven, eight, nine, ten, twelve, fifteen, twenty, or more diverse genomes. Normal, benign, and / or tumor-derived mixed samples can optionally be analyzed without the need for normal controls. In some embodiments, small starting samples of 100 ng or a few hundred genome equivalents are utilized to generate accurate long reads. The systems and methods described herein may enable the detection of large-scale structural variants and rearrangements. Phased variant calls can be obtained over long reads spanning about 1 kbp, about 2 kbp, about 5 kbp, about 10 kbp, about 20 kbp, about 50 kbp, about 100 kbp, about 200 kbp, about 500 kbp, about 1 Mbp, about 2 Mbp, about 5 Mbp, about 10 Mbp, about 20 Mbp, about 50 Mbp, or about 100 Mbp, or more nucleotides. For example, phased variant calls can be obtained over long reads spanning about 1 Mbp or about 2 Mbp.
[0100] In certain embodiments, the methods disclosed herein are used to assemble multiple contigs derived from a single DNA molecule. In some cases, the method includes generating multiple read pairs from a single DNA molecule crosslinked to multiple nanoparticles and assembling the contigs using the read pairs. In certain cases, the single DNA molecule is crosslinked outside the cell. In some cases, at least 0.1%, 0.2%, 0.3%, 0.4%, 0.5%, 0.6%, 0.7%, 0.8%, 0.9%, 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, 20%, 25%, 30%, 35%, 40%, 45%, or 50% of the read pairs span a distance greater than 1 kB, 2 kB, 3 kB, 4 kB, 5 kB, 6 kB, 7 kB, 8 kB, 9 kB, 10 kB, 15 kB, 20 kB, 30 kB, 40 kB, 50 kB, 60 kB, 70 kB, 80 kB, 90 kB, 100 kB, 150 kB, 200 kB, 250 kB, 300 kB, 400 kB, 500 kB, 600 kB, 700 kB, 800 kB, 900 kB, or 1 MB on the single DNA molecule. In certain cases, at least 0.5%, 0.6%, 0.7%, 0.8%, 0.9%, 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, or 20% of the read pairs span a distance greater than 5 kB, 6 kB, 7 kB, 8 kB, 9 kB, 10 kB, 15 kB, 20 kB, 30 kB, 40 kB, 50 kB, 60 kB, 70 kB, 80 kB, 90 kB, 100 kB, 150 kB, or 200 kB on the single DNA molecule. In further cases, at least 0.5%, 0.6%, 0.7%, 0.8%, 0.9%, 1%, 2%, 3%, 4%, or 5% of the read pairs span a distance greater than 20 kB, 30 kB, 40 kB, 50 kB, 60 kB, 70 kB, 80 kB, 90 kB, or 100 kB on the single DNA molecule. In certain cases, at least 1% or 5% of the read pairs span a distance greater than 50 kB or 100 kB on the single DNA molecule.In some cases, the read pairs are generated within 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 30, 40, 50, or 60 days. In certain cases, the read pairs are generated within 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, or 18 days. In further cases, the read pairs are generated within 7, 8, 9, 10, 11, 12, 13, or 14 days. In certain cases, the read pairs are generated within 7 or 14 days.
[0101] The haplotypes determined using the methods and systems described herein may be assigned to computer resources, such as computer resources on a network, e.g., a cloud system. Short variant calls can be corrected using relevant information stored in computer resources if necessary. Structural variants can be detected based on combined information from short variant calls and information stored in computer resources. Problematic parts of the genome, such as segmental duplications, regions prone to structural variation, highly variable and medically relevant MHC regions, centromere and telomere regions, as well as repetitive regions, low sequence accuracy, high variant rate, ALU repeats, other heterochromatin regions with segmental duplications, or other relevant problematic parts in the art can be reassembled for improved accuracy.
[0102] The type of sample can be assigned to the array information locally or in network-connected computer resources such as the cloud. If the source of the information is known, for example, if the source of the information is from cancer or normal tissue, this source can be assigned to the sample as part of the type of sample. Examples of other sample types typically include, but are not limited to, tissue type, sample collection method, presence of infection, type of infection, processing method, sample size, etc. If a complete or partial comparative genomic sequence, such as a normal genome in comparison to a cancer genome, is available, the difference between the sample data and the comparative genomic sequence can be determined and optionally output.
[0103] The methods of the present disclosure can be used to analyze genetic information of genomic regions that can interact with a selected region of interest in addition to the selected region of interest of the genome. Amplification methods as disclosed herein can be used, but are not limited to, those found in U.S. Patent Nos. 6,449,562, 6,287,766, 7,361,468, 7,414,117, 6,225,109, and 6,110,709, for devices, kits, and methods for genetic analysis. In some cases, the amplification methods of the present disclosure can be used to amplify target nucleic acids for DNA hybridization studies to determine the presence or absence of polymorphisms. Polymorphisms or alleles can be associated with diseases or disorders such as genetic diseases. In some other cases, polymorphisms can be associated with susceptibility to a disease or disorder, for example, polymorphisms can be associated with poisoning, degenerative and age-related diseases, cancer, etc. In other cases, polymorphisms can be associated with useful traits such as increased coronary artery health, resistance to diseases such as HIV or malaria, or resistance to adult diseases such as osteoporosis, Alzheimer's disease, or dementia.
[0104] The compositions and methods of the present disclosure can be used for diagnostic, prognostic, therapeutic, patient stratification, drug development, treatment selection, and screening purposes. The present disclosure provides the advantage that many different target molecules can be analyzed from a single biological sample at one time using the methods of the present disclosure. This enables, for example, performing various diagnostic tests on one sample.
[0105] The methods provided herein can significantly advance the field of genomics by overcoming the substantial barriers posed by these repetitive regions, thereby enabling important advances in many areas of genomic analysis. To perform de novo assembly using prior art, one must either prepare an assembly fragmented into many small scaffolds, or use other approaches for generating large insertion libraries or more contiguous assemblies, which requires significant time and resources. Such approaches may include obtaining very deep sequencing coverage, constructing BAC or fosmid libraries, optical mapping, or perhaps some combination of these and other techniques. Due to the severe requirements for resources and time, such approaches have not penetrated most small-scale laboratories, hindering research on non-model organisms. The methods described herein can generate very long-range read sets, such that de novo assembly can be achieved with the execution of a single sequencing run. This reduces the assembly cost by orders of magnitude and shortens the time required from months or years to weeks. In some cases, the methods disclosed herein can enable the generation of multiple read sets in less than 14 days, less than 13 days, less than 12 days, less than 11 days, less than 10 days, less than 9 days, less than 8 days, less than 7 days, less than 6 days, less than 5 days, less than 4 days, less than 3 days, less than 2 days, less than 1 day, or within the range between any two of the aforementioned specified periods. In some cases, the method can enable the generation of multiple read sets in about 10 to 14 days. Constructing genomes becomes routine even for most niches of organisms, phylogeny analysis is not troubled by lack of comparison, and projects such as Genome 10k can be realized.
[0106] The methods described herein enable the assignment of previously provided contig information, previously generated contig information, or de novo synthesized contig information to physical linkage groups such as chromosomes or shorter contiguous nucleic acid molecules. Similarly, the methods disclosed herein enable the contigs to be arranged relative to one another in a linear order along a physical nucleic acid molecule. Similarly, the methods disclosed herein enable the contigs to be oriented relative to one another in a linear order along a physical nucleic acid molecule.
[0107] Similarly, the methods disclosed herein can provide advances in structural and phasing analysis for medical purposes. There is a surprising heterogeneity even within cancers, individuals with the same type of cancer, or even within the same tumor. Extracting the cause from the resulting effects requires very high accuracy and throughput at low cost per sample. In the area of personalized medicine, one of the absolute benchmarks for genomic care is a sequenced genome with all variants thoroughly characterized and phased, including large and small structural rearrangements and novel mutations. Achieving this with previous technologies requires similar effort to that required for de novo assembly, which is now too expensive and requires conventional medical procedures. In some cases, the methods disclosed herein can rapidly produce a low-cost, complete, and accurate genome, thereby providing many highly sought-after capabilities in the study and treatment of human diseases.
[0108] Furthermore, applying the methods disclosed herein to phasing can combine the convenience of statistical approaches with the accuracy of family analysis, resulting in greater savings - cost, labor, and samples - than using either method alone. A novel variant phasing analysis, a highly desirable phasing analysis that was prohibitively expensive with previous techniques, can be readily implemented using the methods disclosed herein. This is particularly important because the majority of human variants are rare (minor allele frequency of less than 5%). Phasing information is valuable for population genetic studies that obtain significant advantages from highly linked haplotype networks (sets of variants assigned to a single chromosome) compared to unlinked genotypes. Haplotype information can enable higher resolution studies of the history of population size, migration, and exchange between subpopulations, and enable tracking specific variants back to specific parents and grandparents. This, in turn, reveals the genetic transmission of disease - related variants and the interactions between variants when grouped in a single individual. In further instances, the methods of the disclosure enable the preparation, sequencing, and analysis of extremely long - range read sets (XLRS) or extremely long - range read pairs (XLRP) libraries.
[0109] In some embodiments of the disclosure, a tissue or DNA sample from a subject is provided and the method returns to the assembled genome, alignment with called variants (including large structural variants), phased variant calls, or any additional analysis. In other embodiments, the methods disclosed herein directly provide an XLRP library for an individual.
[0110] In various embodiments, the methods disclosed herein generate extremely long read pairs that are separated by large distances. The upper limit of this distance can be improved by the ability to collect large-sized DNA samples. In some cases, the read pairs span genomic distances of up to 50, 60, 70, 80, 90, 100, 125, 150, 175, 200, 225, 250, 300, 400, 500, 600, 700, 800, 900, 1000, 1500, 2000, 2500, 3000, 4000, 5000 kbp, or more. In some cases, the read pairs span genomic distances of up to 500 kbp. In other cases, the read pairs span genomic distances of up to 2000 kbp. The methods disclosed herein can be incorporated and constructed based on standard techniques in molecular biology and are well-suited for increased efficiency, specificity, and genomic coverage. In some cases, the read pairs are generated in less than about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 60, or 90 days. In some cases, the read pairs are generated in less than about 14 days. In further cases, the read pairs are generated in less than about 10 days. In some cases, the methods of the present disclosure provide at least about 50%, about 60%, about 70%, about 80%, about 90%, about 95%, about 99%, or about 100% accuracy in the correct ordering and / or orientation of multiple contigs for about 5%, about 10%, about 15%, about 20%, about 30%, about 40%, about 50%, about 60%, about 70%, about 80%, about 90%, about 95%, about 99%, or more than about 100% of the read pairs. In some cases, the method provides about 90 - 100% accuracy in the correct ordering and / or orientation of multiple contigs.
[0111] In other embodiments, the methods disclosed herein are used in conjunction with currently utilized sequencing technologies. In some cases, this method is used in combination with well-tested and / or widely deployed sequencing equipment. In further embodiments, the methods disclosed herein are used with techniques and approaches derived from currently utilized sequencing technologies.
[0112] The methods disclosed herein can dramatically simplify de novo genome assembly for a wide variety of organisms. Using previous techniques, such assemblies are currently limited by short inserts in economical mate-pair libraries. It may be possible to generate read pairs at genomic distances of up to 40-50 kbp accessible with fosmids, but these are expensive, difficult to handle, and too short to span the longest repetitive stretches, including those within centromeres, which range in size from 300 kbp to 5 Mbp in humans. In some cases, the methods disclosed herein can provide read pairs that can span long distances (e.g., megabases and above), thereby overcoming these scaffold completeness issues. Thus, the generation of chromosome-level assemblies can be routine by utilizing the methods disclosed herein. Similarly, the acquisition of long-range phasing information can provide additional substantial power to population genomic studies, phylogenetic studies, and disease studies. In certain cases, the methods disclosed herein enable accurate phasing for multiple individuals, thus expanding the breadth and depth of our ability to probe genomes at the population and deep time levels.
[0113] In the field of personalized medicine, the XLRS read sets generated from the methods disclosed herein represent a significant advancement to accurate, low-cost, phased, and rapidly produced personal genomes. Previous methods have insufficient ability to phase variants at long distances, thereby hindering the characterization of the phenotypic impact of compound heterozygous genotypes. Furthermore, structural variants of substantial interest in genomic diseases are large in size compared to the reads and read inserts used to study them, making it difficult to accurately identify and characterize them using previous techniques. Read sets spanning tens of kilobases to megabases and above can help alleviate this difficulty, thereby enabling highly parallel and individualized analysis of structural polymorphisms.
[0114] Basic evolutionary and biomedical research can be advanced by technological progress in high-throughput sequencing. Currently, it is relatively inexpensive to generate large amounts of DNA sequence data. However, it is difficult, both theoretically and practically, to produce high-quality and highly contiguous genomic sequences using previous technologies. Furthermore, many organisms, including humans, are diploid, and each individual has two haploid copies of the genome. At heterozygous sites (e.g., where the allele given by the mother is different from the allele given by the father), it is difficult to know which set of alleles came from which parent (known as haplotype phasing). This information can be extremely important for conducting some evolutionary and biomedical research, such as studies of disease and trait associations.
[0115] The present disclosure provides methods for genomic assembly that combine tagged sequence reads and techniques for DNA preparation for high-throughput discovery of short, medium, and long-range interactions corresponding to sequence reads from a single physical nucleic acid molecule bound to a complex such as a chromatin complex within a given genome. The present disclosure further provides methods for using these interactions to assist in genomic assembly, for haplotype phasing, and / or for metagenomic studies. The methods presented herein can be used to determine the assembly of a genome of interest, but in certain cases, it should also be understood that the methods presented herein can be used to determine the assembly of a portion of a genome of interest, such as a chromosome, or the assembly of chromatin of various lengths of the object. In certain cases, it should also be understood that the methods presented herein can be used to determine or direct the assembly of non-chromosomal nucleic acid molecules. Indeed, any nucleic acid whose sequencing is complicated by the presence of repetitive regions that separate non-repetitive contigs can be facilitated using the methods disclosed herein.
[0116] In further cases, the methods disclosed herein enable accurate and predictive results for genotyping assembly, haplotype phasing, and metagenomics using small amounts of material. In some cases, less than about 100 picograms (pg), less than about 200 pg, less than about 300 pg, less than about 400 pg, less than about 500 pg, less than about 600 pg, less than about 700 pg, less than about 800 pg, less than about 900 pg, less than about 1.0 nanogram (ng), less than about 2.0 ng, less than about 3.0 ng, less than about 4.0 ng, less than about 5.0 ng, less than about 6.0 ng, less than about 7.0 ng, less than about 8.0 ng, less than about 9.0 ng, less than about 10 ng, less than about 15 ng, less than about 20 ng, less than about 30 ng, less than about 40 ng, less than about 50 ng, less than about 60 ng, less than about 70 ng, less than about 80 ng, less than about 90 ng, less than about 100 ng, less than about 200 ng, less than about 300 ng, less than about 400 ng, less than about 500 ng, less than about 600 ng, less than about 700 ng, less than about 800 ng, less than about 900 ng, less than about 1.0 microgram (μg), less than about 1.2 μg, less than about 1.4 μg, less than about 1.6 μg, less than about 1.8 μg, less than about 2.0 μg, less than about 2.5 μg, less than about 3.0 μg, less than about 3.5 μg, less than about 4.0 μg, less than about 4.5 μg, less than about 5.0 μg, less than about 6.0 μg, less than about 7.0 μg, less than about 8.0 μg, less than about 9.0 μg, less than about 10 μg, less than about 15 μg, less than about 20 μg, less than about 30 μg, less than about 40 μg, less than about 50 μg, less than about 60 μg, less than about 70 μg, less than about 80 μg, less than about 90 μg, less than about 100 μg, less than about 150 μg, less than about 200 μg, less than about 300 μg, less than about 400 μg, less than about 500 μg, less than about 600 μg, less than about 700 μg, less than about 800 μg, less than about 900 μg, or less than about 1000 μg of DNA is used with the methods disclosed herein.In some cases, the DNA used in the methods disclosed herein is extracted from fewer than about 10,000,000, fewer than about 5,000,000, fewer than about 4,000,000, fewer than about 3,000,000, fewer than about 2,000,000, fewer than about 1,000,000, fewer than about 500,000, fewer than about 200,000, fewer than about 100,000, fewer than about 50,000, fewer than about 20,000, fewer than about 10,000, fewer than about 5,000, fewer than about 2,000, fewer than about 1,000, fewer than about 500, fewer than about 200, fewer than about 100, fewer than about 50, fewer than about 20, or fewer than about 10 cells.
[0117] In a diploid genome, it is often important to know which allelic variants are physically associated on the same chromosome rather than mapping to homologous positions on chromosome pairs. Mapping an allele or other sequence to a specific physical chromosome of a diploid chromosome pair is known as haplotype phasing. Short reads from high-throughput sequence data, while most frequent, rarely allow direct observation of which allelic variants are associated when the allelic variants are separated by a distance greater than the longest single read. Computer-based inference of haplotype phasing can be uncertain at long distances. The methods disclosed herein enable determination of which allelic variants are physically associated using allelic variants on read pairs.
[0118] In various cases, the methods and compositions of the present disclosure enable haplotype phasing of diploid or polyploid genomes with respect to multiple allelic variants. Accordingly, the methods described herein provide determination of associated allelic variants based on variant information from labeled sequence segments and / or assembled contigs that use it. Examples of allelic variants include, but are not limited to, those known from the 1000 Genomes, UK10K, HapMap, and other projects to discover genetic variation among humans. In some cases, for example, the discovery of unlinked inactivating mutations in both copies of SH3TC2 that cause Charcot-Marie-Tooth neuropathy (Lupski JR, Reid JG, Gonzaga-Jauregui C, et al. N. Engl. J. Med. 362:1181-91, 2010), and unlinked inactivating mutations in both copies of ABCG5 that cause hypercholesterolemia 9 (Rios J, Stein E, Shendure J, et al. Hum. Mol. Genet. 19:4313-18, 2010), as demonstrated by having haplotype phasing data, more readily reveals the association of diseases to specific genes.
[0119] Humans are heterozygous at an average of one site per 1,000. In some cases, a single lane of data using high-throughput sequencing generates at least about 150,000,000 reads. In further cases, individual reads are about 100 base pairs in length. The inventors assume that the input DNA fragments are of an average size of 150 kbp and, when obtaining 100 paired-end reads per fragment, expect to observe 30 heterozygous sites per set, i.e., per 100 read pairs. All read pairs containing a heterozygous site within a set are homologous (i.e., molecularly linked) to all other read pairs within the same set. This property, in some cases, enables a greater force of phasing within the set, as opposed to a single pair of reads. With about 3 billion bases in the human genome and 1 in 1,000 being heterozygous, there are about 3 million heterozygous sites in the average human genome. For about 45,000,000 read pairs containing heterozygous sites, the average coverage of each heterozygous site phased using a single lane of high-throughput sequencing data is about (15-fold (15X)) using a typical high-throughput sequencer. Thus, the diploid human genome can be reliably and fully phased with one lane of high-throughput sequence data related to sequence variants from a sample prepared using the methods disclosed herein. In some cases, a lane of data is a set of DNA sequence read data. In further cases, a lane of data is a set of DNA sequence read data from one run of a high-throughput sequencer.
[0120] Since the human genome consists of two homologous sets of chromosomes, understanding an individual's true genetic makeup requires characterization of the maternal and paternal copies, or haplotypes, of the genetic material. Obtaining haplotypes in an individual is useful in several ways. For example, haplotypes are clinically useful in predicting the outcome of donor-host compatibility in organ transplantation. Haplotypes are increasingly being used to detect disease associations. In genes showing compound heterozygosity, haplotypes provide information regarding whether two deleterious variants are located on the same allele (i.e., in "cis," using genetic terms), or on two different alleles ("trans"), and greatly influence prediction of whether the inheritance of these variants will be deleterious, affecting conclusions as to whether an individual has a single non-functional allele with a functional allele and two deleterious variant positions, or whether the individual has two non-functional alleles each with a different defect. Haplotypes from populations are of interest to both epidemiologists and anthropologists, providing information regarding population structure and being useful in the history of human evolution. In addition, extensive allelic imbalance in gene expression has been reported, suggesting that genetic or epigenetic differences between allelic phases may contribute to quantitative differences in expression. Understanding haplotype structure will describe the mechanisms of variants contributing to allelic imbalance.
[0121] In certain embodiments, the methods disclosed herein include in vitro techniques for fixing and capturing associations between distal regions of the genome required for long-range ligation and phasing. Optionally, the method includes constructing and sequencing one or more read sets to deliver very genomically distal read pairs. In further cases, each read set includes two or more reads labeled with a common barcode, which may represent two or more sequence segments from a common polynucleotide. Optionally, the interactions arise primarily from random associations within a single polynucleotide. Optionally, sequence segments that are proximal to each other in the polynucleotide interact more frequently and with a higher probability, while interactions between distal portions of the molecule are less frequent, so the genomic distance between sequence segments is inferred. Thus, there is a systematic relationship between the number of pairs that join two loci and their proximity on the input DNA.
[0122] In some aspects, the present disclosure provides methods and compositions for generating data to achieve extremely high phasing accuracy. Compared to previous methods, the methods described herein can phase a higher percentage of variants. In some cases, phasing is achieved while maintaining a high level of accuracy. In further cases, this phasing information is extended over long distances, for example, beyond about 200 kbp, about 300 kbp, about 400 kbp, about 500 kbp, about 600 kbp, about 700 kbp, about 800 kbp, about 900 kbp, about 1 Mbp, about 2 Mbp, about 3 Mbp, about 4 Mbp, about 5 Mbp, or about 10 Mbp, or up to and including the full length of the chromosome, beyond about 10 Mbp. In some embodiments, over 90% of the heterozygous SNPs in a human sample are phased with an accuracy of over 99% using fewer than about 250 million reads, for example, using only 1 lane of Illumina HiSeq data. In other cases, over about 40%, 50%, 60%, 70%, 80%, 90%, 95%, or 99% of the heterozygous SNPs for a human sample are phased with an accuracy of over about 70%, 80%, 90%, 95%, or 99% using fewer than about 250 million or fewer than about 500 million reads, for example, using only 1 or 2 lanes of Illumina HiSeq data. In some cases, over 95% or 99% of the heterozygous SNPs for a human sample are phased with an accuracy of over 95% or 99% using fewer than about 250 million or fewer than about 500 million reads. In further cases, additional variants are captured by increasing the read length to about 200 bp, 250 bp, 300 bp, 350 bp, 400 bp, 450 bp, 500 bp, 600 bp, 800 bp, 1000 bp, 1500 bp, 2 kbp, 3 kbp, 4 kbp, 5 kbp, 10 kbp, 20 kbp, 50 kbp, or 100 kbp.
[0123] The compositions and methods of the present disclosure can be used for gene expression analysis. The methods described herein distinguish nucleotide sequences. The differences between target nucleotide sequences can be, for example, a single nucleotide difference, a nucleic acid deletion, a nucleic acid insertion, or a rearrangement. Such sequence differences containing more than one base can also be detected. The processes of the present disclosure can detect infectious diseases, genetic diseases, and cancers. Further, the above processes are also useful in environmental monitoring, forensic science, and food science. Examples of genetic analysis that can be performed on nucleic acids include, for example, SNP detection, STR detection, RNA expression analysis, promoter methylation, gene expression, virus detection, virus subtype classification, and drug resistance.
[0124] The method can be applied to the analysis of a biomolecular sample obtained from or derived from a patient to determine whether an affected cell type is present in the sample, the stage of the disease, the patient's prognosis, the patient's ability to respond to a particular treatment, or the best treatment for the patient. The method can also be applied to identify biomarkers for a particular disease.
[0125] In some embodiments, the methods described herein are used for the diagnosis of diseases. As used herein, the term "diagnosing" or "diagnosis" of a disease includes predicting or diagnosing a disease, determining the predisposition of a disease, monitoring the treatment of a disease, diagnosing the treatment response, or prognosis, progression, or response to a particular treatment of a disease. For example, a blood sample can be assayed according to any of the methods described herein to determine the presence and / or amount of markers of a disease or malignant cell type in the sample, thereby diagnosing or staging a disease or cancer.
[0126] In some embodiments, the methods and compositions described herein are used for the diagnosis and prognosis of diseases.
[0127] A number of immunological, proliferative, and malignant diseases and disorders are particularly suitable for the methods described herein. Immune diseases and disorders include allergic diseases and disorders, immune dysfunction, and autoimmune diseases and disorders. Allergic diseases and disorders include, but are not limited to, allergic rhinitis, allergic conjunctivitis, allergic asthma, atopic eczema, atopic dermatitis, and food allergies. Immunodeficiencies include, but are not limited to, severe combined immunodeficiency (SCID), eosinophilia syndrome, chronic granulomatous disease, leukocyte adhesion deficiency I and II, hyper IgE syndrome, Chediak-Higashi, neutrophilia, neutropenia, agammaglobulinemia, hyper IgM syndrome, DiGeorge / velo-cardio-facial syndrome, and interferon-gamma-TH1 pathway deficiency. Autoimmune and immunoregulatory disorders include, but are not limited to, rheumatoid arthritis, diabetes, systemic lupus erythematosus, Graves' disease, Graves' ophthalmopathy, Crohn's disease, multiple sclerosis, psoriasis, systemic sclerosis, goiter and lymphomatous goiter (Hashimoto's thyroiditis, lymphadenoid goiter), alopecia areata, autoimmune myocarditis, lichen sclerosus, autoimmune uveitis, Addison's disease, atrophic gastritis, myasthenia gravis, idiopathic thrombocytopenic purpura, hemolytic anemia, primary biliary cirrhosis, Wegener's granulomatosis, polyarteritis nodosa, and inflammatory bowel disease, allograft rejection, and tissue destruction due to allergic reactions to infectious bacteria or environmental antigens.
[0128] Proliferative diseases and disorders that can be evaluated by the methods of the present disclosure include, but are not limited to, neonatal hemangiomas, secondary progressive multiple sclerosis, chronic progressive myelodysplastic diseases, neurofibromatosis, ganglioneuromas, keloid formation, Paget's disease of bone, fibrocystic disease (e.g., of the breast or uterus), sarcoidosis, Peyronie's and Dupuytren's fibrosis, cirrhosis, atherosclerosis, and vascular restenosis.
[0129] Malignant diseases and disorders that can be evaluated by the methods of the present disclosure include both hematological malignancies and solid tumors.
[0130] Hematological malignancies are particularly suitable for the methods of the present disclosure when the sample is a blood sample, as they involve changes in blood-derived cells. Such malignancies include non-Hodgkin lymphoma, Hodgkin lymphoma, non-B cell lymphoma, and other lymphomas, acute or chronic leukemia, polycythemia, thrombocythemia, multiple myeloma, myelodysplastic syndromes, myeloproliferative disorders, myelofibroses, abnormal immune lymphocyte proliferation, and plasma cell disorders.
[0131] Plasma cell diseases that can be evaluated by the methods of the present disclosure include multiple myeloma, amyloidosis, and Waldenström macroglobulinemia.
[0132] Examples of solid tumors include, but are not limited to, colon cancer, breast cancer, lung cancer, prostate cancer, brain tumors, central nervous system tumors, bladder tumors, melanoma, liver cancer, osteosarcoma, and other bone cancers, testicular and ovarian carcinomas, head and neck tumors, and cervical neoplasms.
[0133] Genetic diseases can also be detected by the processes of the present disclosure. This can be performed by screening for chromosomal and genetic abnormalities, or prenatal or postnatal screening for genetic diseases. Examples of detectable genetic diseases include 21-hydroxylase deficiency, cystic fibrosis, fragile X syndrome, Turner syndrome, Duchenne muscular dystrophy, Down syndrome or other trisomies, heart diseases, single gene diseases, HLA typing, phenylketonuria, sickle cell anemia, Tay-Sachs disease, thalassemia, Klinefelter syndrome, Huntington's disease, autoimmune diseases, lipidosis, obesity defect, hemophilia, inborn errors of metabolism, and diabetes.
[0134] The methods described herein can be used to diagnose pathogen infections, such as intracellular bacterial and viral infections, by determining the presence and / or amount of markers for each of the bacteria or viruses in the sample.
[0135] A variety of infectious diseases can be detected by the processes of the present disclosure. Infectious diseases can be caused by infectious agents of bacteria, viruses, parasites, and fungi. The resistance of various infectious agents to drugs can also be determined using the present disclosure.
[0136] Bacterial infectious agents that can be detected by the present disclosure include Escherichia - coli, Salmonella, Shigella, Klebsiella, Pseudomonas, Listeria - monocytogenes, Mycobacterium - tuberculosis, Mycobacterium - avium - intracellulare, Yersinia, Francisella, Pasteurella, Brucella, Clostridium, Bordetella - pertussis, Bacteroides, Staphylococcus - aureus, Streptococcus - pneumoniae, B - Hemolytic strep., Corynebacteria, Legionella, Mycoplasma, Ureaplasma, Chlamydia, Neisseria gonorrhoeae, Neisseria meningitidis, Haemophilus influenzae, Enterococcus - faecalis, Proteus vulgaris, Proteus mirabilis, Helicobacter pylori, Treponema pallidum, Borrelia burgdorferi, Borrelia recurrentis, Rickettsia pathogens, Nocardia, and Actinomycetes.
[0137] Fungal infectious agents that can be detected by the present disclosure include Cryptococcus - neoformans, Blastomyces - dermatitidis, Histoplasma - capsulatum, Coccidioides - immitis, Paracoccidioides - brasiliensis, Candida - albicans, Aspergillus fumigatus, Zygomycetes (Rhizopus), Sporothrix - schenckii, Chromomycosis, and Maduromycosis.
[0138] Viral infectious agents detected by the present disclosure include human immunodeficiency virus, human T-cell lymphocytotrophic virus, hepatitis viruses (e.g., hepatitis B virus and hepatitis C virus), Epstein-Barr virus, cytomegalovirus, human papillomavirus, orthomyxovirus, paramyxovirus, adenovirus, coronavirus, rhabdovirus, poliovirus, togavirus, bunyavirus, arenavirus, rubella virus, and reovirus.
[0139] Parasitic agents that can be detected by the present disclosure include Plasmodium falciparum, Plasmodium malariae, Plasmodium vivax, Plasmodium ovale, Onchocerca volvulus, Leishmania, Trypanosoma species, Schistosoma species, Entamoeba histolytica, Cryptosporidium, Giardia species, Trichomonas species, Balantidium coli, Wuchereria bancrofti, Toxoplasma species, Enterobius vermicularis, Ascaris lumbricoides, Trichuris trichiura, Dracunculus medinensis, trematodes, Diphyllobothrium latum, Taenia species, Pneumocystis carinii, and Necator americanis.
[0140] The present disclosure is also useful for detecting drug resistance to infectious agents. For example, vancomycin-resistant Enterococcus faecium, methicillin-resistant Staphylococcus aureus, penicillin-resistant Streptococcus pneumoniae, multidrug-resistant Mycobacterium tuberculosis, and AZT-resistant human immunodeficiency virus can all be identified by the present disclosure.
[0141] Therefore, the target molecules detected using the compositions and methods of the present disclosure can be either a patient marker (such as a cancer marker) or a marker of infection by foreign substances such as bacterial or viral markers.
[0142] The compositions and methods of the present disclosure can be used to identify and / or quantify a target molecule, the abundance of which indicates a blood marker that is upregulated or downregulated as a result of a biological state or disease condition, such as a disease state.
[0143] In some embodiments, the methods and compositions of the present disclosure can be used for cytokine expression. The low sensitivity of the methods described herein is useful for the early detection of cytokines as biomarkers for disease states, diagnosis, or prognosis, such as cancer, and for the identification of subclinical states.
[0144] The various samples from which the target polynucleotide is derived can include multiple samples from the same individual, samples from different individuals, or combinations thereof. In some embodiments, the sample comprises multiple polynucleotides from one individual. In some embodiments, the sample comprises multiple polynucleotides from two or more individuals. An individual is any organism or part thereof from which the target polynucleotide can be derived, non-limiting examples of which include plants, animals, fungi, protists, Monera, viruses, mitochondria, and chloroplasts. Sample polynucleotides can be isolated from a subject, which can be, but is not limited to, an animal including animals such as cows, pigs, mice, rats, chickens, cats, dogs, etc., and is usually a mammal such as a human, and can be, for example, cell samples, tissue samples, or organ samples derived therefrom, such as cultured cell lines, biopsies, blood samples, or fluid samples containing cells. The sample can also be obtained artificially, such as by chemical synthesis. In some embodiments, the sample contains DNA. In some embodiments, the sample contains genomic DNA. In some embodiments, the sample contains mitochondrial DNA, chloroplast DNA, plasmid DNA, bacterial artificial chromosomes, yeast artificial chromosomes, oligonucleotide tags, or combinations thereof. In some embodiments, the sample contains DNA generated by a primer extension reaction using an appropriate combination of primers and DNA polymerase, including, but not limited to, polymerase chain reaction (PCR), reverse transcription, and combinations thereof. When the template for the primer extension reaction is RNA, the product of reverse transcription is called complementary DNA (cDNA). Primers useful for the primer extension reaction can include sequences specific for one or more targets, random sequences, partially random sequences, and combinations thereof. Reaction conditions suitable for the primer extension reaction are known in the art. Generally, the polynucleotides of the sample include any polynucleotides present in the sample, which may or may not include the target polynucleotide.
[0145] In some embodiments, nucleic acid template molecules (e.g., DNA or RNA) are isolated from biological samples containing various other components such as proteins, lipids, and non-template nucleic acids. The nucleic acid template molecules can be obtained from any cellular material and can be obtained from animals, plants, bacteria, fungi, or other cellular organisms. Biological samples for use in the present disclosure include viral particles or preparations. Nucleic acid template molecules can be obtained directly from an organism or from a biological sample obtained from an organism, such as blood, urine, cerebrospinal fluid, semen, saliva, sputum, feces, and tissue. Any tissue or body fluid specimen may be used as a source of the nucleic acids used in the present disclosure. Nucleic acid template molecules can also be isolated from cultured cells such as primary cell cultures or cell lines. The cells or tissues from which the template nucleic acid is obtained can be infected with a virus or other intracellular pathogen. The sample can also be a biological specimen, a cDNA library, viral DNA, or total RNA extracted from genomic DNA. The sample can also be DNA isolated from a source without cellular structure, such as DNA amplified / isolated from a freezer.
[0146] Methods for the extraction and purification of nucleic acids are known. For example, nucleic acids can be purified by organic extraction with phenol, phenol / chloroform / isopentyl alcohol, or similar formulations including TRIzol and TriReagent. Other non-limiting examples of extraction techniques include (1) organic extraction with or without the use of an automated nucleic acid extractor, such as the Model 341 DNA Extractor available from Applied Biosystems (Foster City, Calif.), using an organic reagent such as phenol / chloroform (Ausubel et al., 1993), followed by ethanol precipitation, (2) solid-phase adsorption methods (U.S. Patent No. 5,234,809, Walsh et al., 1991), and (3) salt-induced nucleic acid precipitation methods, such as precipitation methods typically referred to as "salting out" methods (Miller et al., (1988)). Another example of nucleic acid isolation and / or purification involves the use of magnetic particles to which the nucleic acid can bind specifically or non-specifically, followed by isolation of the beads using a magnet, washing of the nucleic acid, and elution of the nucleic acid from the beads (see, for example, U.S. Patent No. 5,705,628). In some embodiments, the isolation methods described above may be initiated by an enzymatic digestion step, such as digestion with proteinase K or other similar proteases, which helps to remove unwanted proteins from the sample. See, for example, U.S. Patent No. 7,001,724. An RNase inhibitor can be added to the lysis buffer if desired. For certain cell or sample types, it may be desirable to add a protein denaturation / digestion step to the protocol. The purification method can be aimed at isolating DNA, RNA, or both. If both DNA and RNA are isolated together during or after the extraction procedure, additional steps can be utilized to purify one or both separately from the other. For example, purification based on size, sequence, or other physical or chemical properties can also be used to generate sub-fractions of the extracted nucleic acids. In addition to the initial nucleic acid isolation step, nucleic acid purification can be performed after the steps in the methods of the present disclosure, for example, to remove excess or unwanted reagents, reactants, or products.
[0147] The nucleic acid template molecule can be obtained as described in U.S. Patent Application Publication No. US2002 / 0190663 A1, published on October 9, 2003. Generally, nucleic acids can be extracted from biological samples by various techniques such as those described in Maniatis, et al., Molecular Cloning: A Laboratory Manual, Cold Spring Harbor, N.Y., pp. 280-281 (1982). In some cases, the nucleic acid can first be extracted from the biological sample and then cross-linked in vitro. In some cases, native associated proteins (e.g., histones) can be further removed from the nucleic acid.
[0148] In other embodiments, the present disclosure can be readily applied to any high molecular weight double-stranded DNA containing DNA isolated from, for example, tissues, cell cultures, body fluids, animal tissues, plants, bacteria, fungi, viruses, and the like.
[0149] Hi-C method including size selection
[0150] A method is provided herein that includes the steps of obtaining a stabilized biological sample comprising a nucleic acid molecule complexed with at least one nucleic acid-binding protein, contacting the stabilized biological sample with DNase to cleave the nucleic acid molecule into a plurality of segments, joining a first segment and a second segment of the plurality of segments at a junction, and subjecting the plurality of segments to size selection to obtain a plurality of selected segments. Optionally, the plurality of selected segments are from about 145 to about 600 bp. Optionally, the plurality of selected segments are from about 100 to about 2500 bp. Optionally, the plurality of selected segments are from about 100 to about 600 bp. Optionally, the plurality of selected segments are from about 600 to about 2500 bp. Optionally, the plurality of selected segments are from about 100 bp to about 600 bp, from about 100 bp to about 700 bp, from about 100 bp to about 800 bp, from about 100 bp to about 900 bp, from about 100 bp to about 1000 bp, from about 100 bp to about 1100 bp, from about 100 bp to about 1200 bp, from about 100 bp to about 1300 bp, from about 100 bp to about 1400 bp, from about 100 bp to about 1500 bp, from about 100 bp to about 1600 bp, from about 100 bp to about 1700 bp, from about 100 bp to about 1800 bp, from about 100 bp to about 1900 bp, from about 100 bp to about 2000 bp, from about 100 bp to about 2100 bp, from about 100 bp to about 2200 bp, from about 100 bp to about 2300 bp, from about 100 bp to about 2400 bp, or from about 100 bp to about 2500 bp.
[0151] In another aspect of the method comprising a size selection step provided herein, the method further comprises, prior to the size selection step, preparing a sequencing library from a plurality of segments. In some embodiments, the method further comprises subjecting the sequencing library to size selection to obtain a size selected library. Optionally, the size selected library is in the size range of about 350 bp to about 1000 bp. Optionally, the size selected library is in the size range of about 100 bp to about 2500 bp, such as about 100 bp to about 350 bp, about 350 bp to about 500 bp, about 500 bp to about 1000 bp, about 1000 bp to about 1500 bp and about 2000 bp, about 2000 bp to about 2500 bp, about 350 bp to about 1000 bp, about 350 bp to about 1500 bp, about 350 bp to about 2000 bp, about 350 bp to about 2500 bp, about 500 bp to about 1500 bp, about 500 bp to about 2000 bp, about 500 bp to about 2500 bp, about 1000 bp to about 1500 bp, about 1000 bp to about 2000 bp, about 1000 bp to about 2500 bp, about 1500 bp to about 2000 bp, about 1500 bp to about 2500 bp, or about 2000 bp to about 2500 bp.
[0152] The size selection utilized in the method comprising a size selection step provided herein can be performed using gel electrophoresis, capillary electrophoresis, size selection beads, gel filtration columns, other suitable methods, or combinations thereof.
[0153] In another aspect, a method that includes the size selection step provided herein can further include analyzing a plurality of selected segments to obtain a QC value. In some cases, the QC value is selected from a chromatin digestion efficiency (CDE) and a chromatin digestion index (GDI). CDE is calculated as the percentage of segments having a desired length. For example, in some cases, CDE is calculated as the percentage of segments sized 100 - 2500 bp before size selection. In some cases, a sample is selected for further analysis when the CDE value is at least 65%. In some cases, a sample is selected for further analysis when the CDE value is at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, or at least about 95%. CDI is calculated as the ratio of the number of segments of mononucleosome size to the number of segments of dinucleosome size before size selection. For example, CDI can be calculated as the logarithm of the ratio of fragments having a size of 600 - 2500 bp to fragments having a size of 100 - 600 bp. In some cases, a sample is selected for further analysis when the CDI value is greater than -1.5 and less than 1.In some cases, when the CDI value is greater than about -2 and less than about 1.5, greater than about -1.9 and less than about 1.5, greater than about -1.8 and less than about 1.5, greater than about -1.7 and less than about 1.5, greater than about -1.6 and less than about 1.5, greater than about -1.5 and less than about 1.5, greater than about -1.4 and less than about 1.5, greater than about -1.3 and less than about 1.5, greater than about -1.2 and less than about 1.5, greater than about -1.1 and less than about 1.5, greater than about -2 and less than about 1.5, greater than about -1 and less than about 1.5, greater than about -0.9 and less than about 1.5, greater than about -0.8 and less than about 1.5, greater than about -0.7 and less than about 1.5, greater than about -0.6 and less than about 1.5, greater than about -0.5 and less than about 1.5, greater than about -2 and less than about 1.4, greater than about -2 and less than about 1.3, greater than about -2 and less than about 1.2, greater than about -2 and less than about 1.1, greater than about -2 and less than about 1, greater than about -2 and less than about 0.9, greater than about -2 and less than about 0.8, greater than about -2 and less than about 0.7, greater than about -2 and less than about 0.6, or greater than about -2 and less than about 0.5, the sample is selected for further analysis.
[0154] In another aspect, the stabilized biological sample used in the method including the size selection step herein comprises biological material treated with a stabilizer. In some cases, the stabilized biological sample comprises a stabilized cell lysate. Alternatively, the stabilized biological sample comprises stabilized intact cells. Alternatively, the stabilized biological sample comprises stabilized intact nuclei. In some cases, the step of contacting the stabilized intact cells or intact nucleus sample with DNase is performed prior to lysis of the intact cells or intact nuclei. In some cases, the cells and / or nuclei are lysed prior to joining a first segment and a second segment of a plurality of segments at a junction.
[0155] In another aspect, the method comprising the size selection step herein is performed on a small sample containing a small number of cells or a small amount of nucleic acid. For example, in some cases, the stabilized biological sample contains less than 3,000,000 cells. In some cases, the stabilized biological sample contains less than 2,000,000 cells. In some cases, the stabilized biological sample contains less than 1,000,000 cells. In some cases, the stabilized biological sample contains less than 500,000 cells. In some cases, the stabilized biological sample contains less than 400,000 cells. In some cases, the stabilized biological sample contains less than 300,000 cells. In some cases, the stabilized biological sample contains less than 200,000 cells. In some cases, the stabilized biological sample contains less than 100,000 cells. In some cases, the stabilized biological sample contains less than 50,000 cells. In some cases, the stabilized biological sample contains less than 40,000 cells. In some cases, the stabilized biological sample contains less than 30,000 cells. In some cases, the stabilized biological sample contains less than 20,000 cells. In some cases, the stabilized biological sample contains less than 10,000 cells. In some cases, the stabilized biological sample can contain about 10,000 cells. In some cases, the stabilized biological sample contains less than 10 μg of DNA. In some cases, the stabilized biological sample contains less than 9 μg of DNA. In some cases, the stabilized biological sample contains less than 8 μg of DNA. In some cases, the stabilized biological sample contains less than 7 μg of DNA. In some cases, the stabilized biological sample contains less than 6 μg of DNA. In some cases, the stabilized biological sample contains less than 5 μg of DNA. In some cases, the stabilized biological sample contains less than 4 μg of DNA. In some cases, the stabilized biological sample contains less than 3 μg of DNA. In some cases, the stabilized biological sample contains less than 2 μg of DNA. In some cases, the stabilized biological sample contains less than 1 μg of DNA. In some cases, the stabilized biological sample contains less than 0.5 μg of DNA.
[0156] In another aspect, the methods including the size selection step herein can be performed on individual cells or a single cell. For example, the methods herein can be performed on cells dispensed into individual partitions. Exemplary partitions include, but are not limited to, wells, droplets in an emulsion, or surface locations (e.g., array spots, beads, etc.) that include discrete patches of differentially arrayed linker molecules as described elsewhere herein. Additional partitions are contemplated and are consistent with the methods, compositions, and systems disclosed herein.
[0157] In a further aspect, the stabilized biological sample used in the methods including the size selection step herein is treated with a nuclease such as DNase to generate DNA fragments. In some cases, the DNase is non-sequence specific. In some cases, the DNase is active against both single-stranded DNA and double-stranded DNA. In some cases, the DNase is specific for double-stranded DNA. In some cases, the DNase preferentially cleaves double-stranded DNA. In some cases, the DNase is specific for single-stranded DNA. In some cases, the DNase preferentially cleaves single-stranded DNA. In some cases, the DNase is DNaseI. In some cases, the DNase is DNaseII. In some cases, the DNase is selected from one or more of DNaseI and DNaseII. In some cases, the DNase is micrococcal nuclease. In some cases, the DNase is selected from one or more of DNaseI, DNaseII, and micrococcal nuclease. In some cases, the DNase can be bound or fused to an immunoglobulin-binding protein or a fragment thereof. The immunoglobulin-binding protein can be, for example, protein A, protein G, protein A / G, or protein L. In some embodiments, the DNase can be bound to a fusion protein comprising two or more immunoglobulin-binding proteins and / or fragments thereof. Other suitable nucleases are within the scope of the present disclosure.
[0158] In a further aspect, the stabilized biological sample provided herein for use in a method comprising a size selection step is treated with one or more crosslinking agents. Optionally, the crosslinking agent is a chemical fixative. Optionally, the chemical fixative comprises formaldehyde having a spacer arm length of about 2.3 - 2.7 angstroms (A). Optionally, the chemical fixative comprises a crosslinking agent having a long spacer arm length. For example, the crosslinking agent can have a spacer length of at least about 3A, 4A, 5A, 6A, 7A, 8A, 9A, 10A, 11A, 12A, 13A, 14A, 15A, 16A, 17A, 18A, 19A, or 20A. The chemical fixative can comprise ethylene glycol bis(succinimidyl succinate) (EGS) having a spacer arm with a length of about 16.1A. The chemical fixative can comprise disuccinimidyl glutarate (DSG) having a spacer arm with a length of about 7.7A. Optionally, the chemical fixative comprises formaldehyde and EGS, formaldehyde and DSG, or formaldehyde, EGS, and DSG. Optionally, when multiple chemical fixatives are used, each chemical fixative is used sequentially, and in other cases, some or all of the multiple chemical fixatives are applied to the sample simultaneously. The use of a crosslinking agent having a long spacer arm can increase the proportion of read pairs having a large (e.g., >1kb) read pair separation distance. DSG is membrane permeable and enables intracellular crosslinking. DSG can enhance crosslinking efficiency compared to disuccinimidyl suberate (DSS) in some applications. EGS has NHS ester reactive groups at both ends and can be reactive towards amino groups (e.g., primary amines). EGS is membrane permeable and enables intracellular crosslinking. EGS crosslinking can be reversed, for example, by treatment with hydroxylamine at pH 8.5 for 3 - 6 hours. For example, lactate dehydrogenase retained 60% of its activity after reversible crosslinking with EGS. Optionally, the chemical fixative comprises psoralen.In some cases, the crosslinking agent is ultraviolet light, chloromethine, cyclophosphamide, chlorambucil, uracil mustard, melphalan, bendamustine, bis(2-chloroethyl)ethylamine, bis(2-chloroethyl)methylamine, tris(2-chloroethyl)amine, isofamide, carmustine, lomustine, streptozocin, busulfan, cisplatin, carboplatin, cicycloplatin, eptaplatin, lobaplatin, miltiplatin, nedaplatin, oxaliplatin, picoplatin, satraplatin, triplatin tetranitrate, procarbazine, altretamine, dacarbazine, mitozolomide, temozolomide, mitomycin C, nitrous acid, formaldehyde, acetylaldehyde, doxorubicin, daunorubicin, epirubicin, or idarubicin. In some cases, the crosslinking agent includes an intercalator, an antibiotic, or a minor groove binder. In some cases, the stabilized biological sample is a crosslinked paraffin-embedded tissue sample.
[0159] In a further aspect, a method comprising the size selection step provided herein includes contacting a plurality of selected segments with an antibody.
[0160] In a further aspect, a method comprising the size selection step provided herein includes the step of joining a first segment and a second segment of a plurality of segments at a junction. In some cases, the joining step includes filling in sticky ends using biotin-tagged nucleotides and ligating blunt ends. In some cases, the joining step includes contacting at least the first segment and the second segment with a crosslinking oligonucleotide. In some cases, the joining step includes contacting at least the first segment and the second segment with a barcode. In some embodiments, the crosslinking oligonucleotides herein can be at least about 5 nucleotides to about 50 nucleotides in length. In some embodiments, the crosslinking oligonucleotides herein can be about 15 to about 18 nucleotides in length. In some embodiments, the crosslinking oligonucleotide can be about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, about 20, about 25, about 30, about 35, about 40, about 45, or about 50 nucleotides in length. In some embodiments, the crosslinking oligonucleotides herein can include a barcode. In some embodiments, the crosslinking oligonucleotide can include a plurality of barcodes. In some embodiments, the crosslinking oligonucleotide includes a plurality of crosslinking oligonucleotides bound to each other. In some embodiments, the crosslinking oligonucleotide can be coupled or linked to an immunoglobulin-binding protein or a fragment thereof, such as protein A, protein G, protein A / G, or protein L. In some cases, the coupled crosslinking oligonucleotide can be delivered to a position in a sample nucleic acid to which an antibody binds.
[0161] Using a partitioning and pooling approach, cross-linked oligonucleotides with unique barcodes can be generated. A population of samples can be divided into multiple groups, and the cross-linked oligonucleotides can be ligated to the samples such that the cross-linked oligonucleotide barcodes differ between groups but are the same within a group. The groups of samples can be pooled again, and this process can be repeated multiple times. Repeating this process will ultimately result in each sample in the population having a unique set of cross-linked oligonucleotide barcodes, enabling the analysis of a single sample (e.g., single cell, single nucleus, single chromosome). In one exemplary example, a sample of cross-linked digested nuclei bound to a solid support of beads is divided into eight tubes, each containing one of eight unique members of a first adapter group (first iteration) that includes a double-stranded DNA (dsDNA) adapter to be ligated. Each of the eight adapters can have the same 5' overhang sequence for ligation to the nucleic acid termini of cross-linked chromatin aggregates in the nucleus, but otherwise has a unique dsDNA sequence. After ligating the first adapter group, the nuclei can be pooled back and washed to remove the ligation reaction components. The partitioning, ligation, and pooling scheme can be repeated two more times (two iterations). After ligation of the members from each adapter group, the cross-linked chromatin aggregates can be successively ligated to multiple barcodes. In some cases, successive ligation (iteration) of multiple members of multiple adapter groups results in a combination of barcodes. The number of possible combinations of barcodes depends on the number of groups per iteration and the total number of barcode oligonucleotides used. For example, three iterations each containing eight members can have 8^3 possible combinations. In some cases, the combination of barcodes is unique. In some cases, the combination of barcodes is redundant. The total number of barcode combinations can be adjusted by increasing or decreasing the number of groups receiving unique barcodes and / or by increasing or decreasing the number of iterations.When an adapter group greater than 1 is used, a scheme of distribution, ligation, and pooling can be used for iterative adapter ligation. In some cases, the scheme of distribution, ligation, and pooling can be further repeated at least 3, 4, 5, 6, 7, 8, 9, or 10 times. In some cases, the members of the last adapter group include sequences for subsequent enrichment of the DNA ligated to the adapter, for example, during preparation of a sequencing library by PCR amplification.
[0162] In a further aspect, a method comprising a size selection step herein does not include a shearing step (e.g., the nucleic acid is not sheared).
[0163] In a further aspect of a method comprising a size selection step herein, the method includes obtaining at least some sequences on each side of the junction to generate a first read pair. For example, the method can include obtaining sequences of at least about 50 bp, at least about 100 bp, at least about 150 bp, at least about 200 bp, at least about 250 bp, or at least about 300 bp on each side of the junction to generate a first read pair.
[0164] In a further aspect of a method comprising a size selection step herein, the method includes mapping a first read pair to a set of contigs and determining a path through the set of contigs that represents the order and / or orientation relative to the genome.
[0165] In a further aspect of a method comprising a size selection step herein, the method includes mapping a first read pair to a set of contigs and determining the presence of a structural variant or loss of heterozygosity in a stabilized biological sample from the set of contigs.
[0166] In a further aspect of a method comprising a size selection step herein, the method includes mapping a first read pair to a set of contigs and assigning phases to variants in the set of contigs.
[0167] In a further aspect of the method including the size selection step in this specification, the method includes mapping a first read pair to a contiguous set, determining the presence of variants in the contiguous set from the contiguous set, and performing a step selected from one or more of (1) identifying a disease stage, prognosis, or treatment course for a stabilized biological sample, (2) selecting a drug based on the presence of the variant, or (3) identifying the drug efficacy for the stabilized biological sample.
[0168] Hi-C method including QC calculation Further, a method is provided herein that includes obtaining a stabilized biological sample comprising a nucleic acid molecule complexed with at least one nucleic acid-binding protein, contacting the stabilized biological sample with a DNase to cleave the nucleic acid molecule into a plurality of segments, ligating a first segment and a second segment of the plurality of segments at a junction, and analyzing the plurality of segments to determine a QC value. Optionally, the QC value is selected from a chromatin digestion efficiency (CDE) and a chromatin digestion index (GDI). The CDE is calculated as the proportion of segments having a desired length. For example, optionally, the CDE is calculated as the proportion of segments sized 100-2500 bp prior to size selection. Optionally, a sample is selected for further analysis when the CDE value is at least 65%. Optionally, a sample is selected for further analysis when the CDE value is at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, or at least about 95%. The CDI is calculated as the ratio of the number of segments of mononucleosome size to the number of segments of dinucleosome size prior to size selection. For example, the CDI can be calculated as the logarithm of the ratio of fragments having a size of 600-2500 bp to fragments having a size of 100-600 bp. Optionally, a sample is selected for further analysis when the CDI value is greater than -1.5 and less than 1.In some cases, when the CDI value is greater than about -2 and less than about 1.5, greater than about -1.9 and less than about 1.5, greater than about -1.8 and less than about 1.5, greater than about -1.7 and less than about 1.5, greater than about -1.6 and less than about 1.5, greater than about -1.5 and less than about 1.5, greater than about -1.4 and less than about 1.5, greater than about -1.3 and less than about 1.5, greater than about -1.2 and less than about 1.5, greater than about -1.1 and less than about 1.5, greater than about -2 and less than about 1.5, greater than about -1 and less than about 1.5, greater than about -0.9 and less than about 1.5, greater than about -0.8 and less than about 1.5, greater than about -0.7 and less than about 1.5, greater than about -0.6 and less than about 1.5, greater than about -0.5 and less than about 1.5, greater than about -2 and less than about 1.4, greater than about -2 and less than about 1.3, greater than about -2 and less than about 1.2, greater than about -2 and less than about 1.1, greater than about -2 and less than about 1, greater than about -2 and less than about 0.9, greater than about -2 and less than about 0.8, greater than about -2 and less than about 0.7, greater than about -2 and less than about 0.6, or greater than about -2 and less than about 0.5, the sample is selected for further analysis.
[0169] In another aspect, a method that includes the QC determination step of the present specification may include a step of subjecting a plurality of segments to size selection to obtain a plurality of selected segments. In some cases, the plurality of selected segments are from about 145 to about 600 bp. In some cases, the plurality of selected segments are from about 100 to about 2500 bp. In some cases, the plurality of selected segments are from about 100 to about 600 bp. In some cases, the plurality of selected segments are from about 600 to about 2500 bp. In some cases, the plurality of selected segments are from about 100 bp to about 600 bp, from about 100 bp to about 700 bp, from about 100 bp to about 800 bp, from about 100 bp to about 900 bp, from about 100 bp to about 1000 bp, from about 100 bp to about 1100 bp, from about 100 bp to about 1200 bp, from about 100 bp to about 1300 bp, from about 100 bp to about 1400 bp, from about 100 bp to about 1500 bp, from about 100 bp to about 1600 bp, from about 100 bp to about 1700 bp, from about 100 bp to about 1800 bp, from about 100 bp to about 1900 bp, from about 100 bp to about 2000 bp, from about 100 bp to about 2100 bp, from about 100 bp to about 2200 bp, from about 100 bp to about 2300 bp, from about 100 bp to about 2400 bp, or from about 100 bp to about 2500 bp.
[0170] In another aspect of the method including the QC determination step provided herein, the method can further include preparing a sequencing library from a plurality of segments prior to the size selection step. In some embodiments, the method further includes subjecting the sequencing library to size selection to obtain a size selected library. Optionally, the size selected library is in the size range of about 350 bp to about 1000 bp. Optionally, the size selected library is in the size range of about 100 bp to about 2500 bp, such as about 100 bp to about 350 bp, about 350 bp to about 500 bp, about 500 bp to about 1000 bp, about 1000 bp to about 1500 bp and about 2000 bp, about 2000 bp to about 2500 bp, about 350 bp to about 1000 bp, about 350 bp to about 1500 bp, about 350 bp to about 2000 bp, about 350 bp to about 2500 bp, about 500 bp to about 1500 bp, about 500 bp to about 2000 bp, about 500 bp to about 2500 bp, about 1000 bp to about 1500 bp, about 1000 bp to about 2000 bp, about 1000 bp to about 2500 bp, about 1500 bp to about 2000 bp, about 1500 bp to about 2500 bp, or about 2000 bp to about 2500 bp.
[0171] The size selection utilized in the method including the QC determination step herein can be performed using gel electrophoresis, capillary electrophoresis, size selection beads, gel filtration columns, or combinations thereof. Other suitable methods of size selection are also within the scope of the present disclosure.
[0172] In another aspect, the stabilized biological sample used in the QC determination process of the present specification comprises a biological material treated with a stabilizer. In some cases, the stabilized biological sample comprises a stabilized cell lysate. Alternatively, the stabilized biological sample comprises stabilized intact cells. Alternatively, the stabilized biological sample comprises stabilized intact nuclei. In some cases, the step of contacting the stabilized intact cells or intact nucleus sample with DNase is performed prior to lysis of the intact cells or intact nuclei. In some cases, the cells and / or nuclei are lysed prior to joining a first segment and a second segment of a plurality of segments at a junction.
[0173] In another aspect, the methods including the QC determination step herein are performed on a small sample containing a small number of cells or a small amount of nucleic acid. In some cases, the stabilized biological sample contains less than 3,000,000 cells. In some cases, the stabilized biological sample contains less than 2,000,000 cells. In some cases, the stabilized biological sample contains less than 1,000,000 cells. In some cases, the stabilized biological sample contains less than 500,000 cells. In some cases, the stabilized biological sample contains less than 400,000 cells. In some cases, the stabilized biological sample contains less than 300,000 cells. In some cases, the stabilized biological sample contains less than 200,000 cells. In some cases, the stabilized biological sample contains less than 100,000 cells. In some cases, the stabilized biological sample contains less than 50,000 cells. In some cases, the stabilized biological sample contains less than 40,000 cells. In some cases, the stabilized biological sample contains less than 30,000 cells. In some cases, the stabilized biological sample contains less than 20,000 cells. In some cases, the stabilized biological sample contains less than 10,000 cells. In some cases, the stabilized biological sample contains about 10,000 cells. In some cases, the stabilized biological sample contains less than 10 μg of DNA. In some cases, the stabilized biological sample contains less than 9 μg of DNA. In some cases, the stabilized biological sample contains less than 8 μg of DNA. In some cases, the stabilized biological sample contains less than 7 μg of DNA. In some cases, the stabilized biological sample contains less than 6 μg of DNA. In some cases, the stabilized biological sample contains less than 5 μg of DNA. In some cases, the stabilized biological sample contains less than 4 μg of DNA. In some cases, the stabilized biological sample contains less than 3 μg of DNA. In some cases, the stabilized biological sample contains less than 2 μg of DNA. In some cases, the stabilized biological sample contains less than 1 μg of DNA. In some cases, the stabilized biological sample contains less than 0.5 μg of DNA.
[0174] In another aspect, the methods described herein that include the QC determination step can be performed on individual cells or a single cell. For example, the methods described herein can be performed on cells distributed into individual partitions. Exemplary partitions include, but are not limited to, wells, droplets in an emulsion, or surface locations (e.g., array spots, beads, etc.) that include discrete patches of differentially addressable linker molecules as described elsewhere herein. Additional partitions are contemplated and are consistent with the methods, compositions, and systems disclosed herein.
[0175] In a further aspect, the stabilized biological sample used in the methods described herein that include the QC determination step is treated with a nuclease, such as DNase, to generate DNA fragments. In some cases, the DNase is non-sequence specific. In some cases, the DNase is active against both single-stranded DNA and double-stranded DNA. In some cases, the DNase is specific for double-stranded DNA. In some cases, the DNase preferentially cleaves double-stranded DNA. In some cases, the DNase is specific for single-stranded DNA. In some cases, the DNase preferentially cleaves single-stranded DNA. In some cases, the DNase is DNaseI. In some cases, the DNase is DNaseII. In some cases, the DNase is selected from one or more of DNaseI and DNaseII. In some cases, the DNase is micrococcal nuclease. In some cases, the DNase is selected from one or more of DNaseI, DNaseII, and micrococcal nuclease. In some cases, the DNase can be bound or fused to an immunoglobulin-binding protein or fragment thereof, such as protein A, protein G, protein A / G, or protein L. Other suitable nucleases are within the scope of the present disclosure.
[0176] In a further aspect, the stabilized biological sample used in the methods including the QC determination step herein is treated with a crosslinking agent. In some cases, the crosslinking agent is a chemical fixative. In some cases, the chemical fixative includes formaldehyde having a spacer arm length of about 2.3 to 2.7 angstroms (A). In some cases, the chemical fixative includes a crosslinking agent having a long spacer arm length. For example, the crosslinking agent can have a spacer length of at least about 3A, 4A, 5A, 6A, 7A, 8A, 9A, 10A, 11A, 12A, 13A, 14A, 15A, 16A, 17A, 18A, 19A, or 20A. The chemical fixative can include ethylene glycol bis(succinimidyl succinate) (EGS) having a spacer arm with a length of about 16.1A. The chemical fixative can include disuccinimidyl glutarate (DSG) having a spacer arm with a length of about 7.7A. In some cases, the chemical fixative includes formaldehyde and EGS, formaldehyde and DSG, or formaldehyde, EGS, and DSG. When multiple chemical fixatives are used, in some cases, each chemical fixative is used sequentially, and in other cases, some or all of the multiple chemical fixatives are applied to the sample simultaneously. The use of a crosslinking agent having a long spacer arm can increase the proportion of read pairs having a large (e.g., >1 kb) read pair separation distance. DSG is membrane-permeable and enables intracellular crosslinking. DSG can enhance crosslinking efficiency compared to disuccinimidyl suberate (DSS) in some applications. EGS has NHS ester reactive groups at both ends and can be reactive towards amino groups (e.g., primary amines). EGS is membrane-permeable and enables intracellular crosslinking. EGS crosslinking can be reversed, for example, by treating with hydroxylamine at pH 8.5 for 3 to 6 hours. For example, lactate dehydrogenase retained 60% of its activity after reversible crosslinking with EGS. In some cases, the chemical fixative includes psoralen.In some cases, the crosslinking agent is ultraviolet light, chloromethine, cyclophosphamide, chlorambucil, uracil mustard, melphalan, bendamustine, bis(2-chloroethyl)ethylamine, bis(2-chloroethyl)methylamine, tris(2-chloroethyl)amine, isofamide, carmustine, lomustine, streptozocin, busulfan, cisplatin, carboplatin, cicycloplatin, eptaplatin, lobaplatin, milphlatine, nedaplatin, oxaliplatin, picoplatin, satraplatin, triplatin tetranitrate, procarbazine, altretamine, dacarbazine, mitozolomide, temozolomide, mitomycin C, nitrous acid, formaldehyde, acetylaldehyde, doxorubicin, daunorubicin, epirubicin, or idarubicin. In some cases, the crosslinking agent includes an intercalator, an antibiotic, or a minor groove binder. In some cases, the stabilized biological sample is a crosslinked paraffin-embedded tissue sample.
[0177] In a further aspect, the method comprising the QC determination step provided herein includes contacting a plurality of selected segments with an antibody.
[0178] In a further aspect, a method comprising the QC determination step provided herein includes a step of joining a first segment and a second segment of a plurality of segments at a junction. In some cases, the joining step includes filling in sticky ends using biotin-tagged nucleotides and ligating blunt ends. In some cases, the joining step includes contacting at least the first segment and the second segment with a cross-linking oligonucleotide. In some cases, the joining step includes contacting at least the first segment and the second segment with a barcode. In some embodiments, the cross-linking oligonucleotides herein can be at least about 5 nucleotides to about 50 nucleotides in length. In some embodiments, the cross-linking oligonucleotides herein can be about 15 to about 18 nucleotides in length. In some embodiments, the cross-linking oligonucleotide can be about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, about 20, about 25, about 30, about 35, about 40, about 45, or about 50 nucleotides in length. In some embodiments, the cross-linking oligonucleotides herein can include a barcode. In some embodiments, the cross-linking oligonucleotide can include a plurality of barcodes. In some embodiments, the cross-linking oligonucleotide includes a plurality of cross-linking oligonucleotides linked to each other. In some embodiments, the cross-linking oligonucleotide can be coupled or linked to an immunoglobulin-binding protein such as protein A, protein G, protein A / G, or protein L or a fragment thereof. In some cases, the coupled cross-linking oligonucleotide can be delivered to a position in the sample nucleic acid to which the antibody binds.
[0179] In a further aspect, a method comprising the QC determination step provided herein does not include a shearing step.
[0180] In a further aspect of the method including the QC determination step herein, the method includes obtaining at least some sequences on each side of the junction to generate a first read pair. For example, the method may include obtaining sequences of at least about 50 bp, at least about 100 bp, at least about 150 bp, at least about 200 bp, at least about 250 bp, or at least about 300 bp on each side of the junction to generate a first read pair.
[0181] In a further aspect of the method including the QC determination step herein, the method includes mapping the first read pair to a set of contigs and determining a path through the set of contigs that represents the order and / or orientation relative to the genome.
[0182] In a further aspect of the method including the QC determination step herein, the method may include mapping the first read pair to a set of contigs and determining the presence of a structural variant or loss of heterozygosity in the stabilized biological sample from the set of contigs.
[0183] In a further aspect of the method including the QC determination step herein, the method includes mapping the first read pair to a set of contigs and assigning phases to variants in the set of contigs.
[0184] In a further aspect of the method including the QC determination step herein, the method includes mapping the first read pair to a set of contigs, determining the presence of variants in the set of contigs, and performing a step selected from one or more of (1) identifying a disease stage, prognosis, or treatment course for the stabilized biological sample, (2) selecting a drug based on the presence of the variant, or (3) identifying the drug efficacy for the stabilized biological sample.
[0185] Hi-C method including whole cell or whole nucleus digestion A method is further provided herein that includes the steps of obtaining a stabilized biological sample comprising a nucleic acid molecule complexed with at least one nucleic acid binding protein, contacting the stabilized biological sample with DNase to cleave the nucleic acid molecule into a plurality of segments, and ligating a first segment and a second segment of the plurality of segments at a junction, wherein the stabilized biological sample comprises intact cells and / or intact nuclei. Optionally, the stabilized biological sample comprises stabilized intact cells. Alternatively, or in combination, the stabilized biological sample comprises stabilized intact nuclei. Optionally, the step of contacting the stabilized intact cell or intact nucleus sample with DNase is performed prior to lysis of the intact cell or intact nucleus. Optionally, the cells and / or nuclei are lysed prior to ligating the first segment and the second segment of the plurality of segments at the junction.
[0186] In another aspect, a method that includes digestion of whole cells or whole nuclei herein can include the step of size selecting a plurality of segments to obtain a plurality of selected segments. Optionally, the plurality of selected segments are from about 145 to about 600 bp. Optionally, the plurality of selected segments are from about 100 to about 2500 bp. Optionally, the plurality of selected segments are from about 100 to about 600 bp. Optionally, the plurality of selected segments are from about 600 to about 2500 bp. Optionally, the plurality of selected segments are from about 100 bp to about 600 bp, from about 100 bp to about 700 bp, from about 100 bp to about 800 bp, from about 100 bp to about 900 bp, from about 100 bp to about 1000 bp, from about 100 bp to about 1100 bp, from about 100 bp to about 1200 bp, from about 100 bp to about 1300 bp, from about 100 bp to about 1400 bp, from about 100 bp to about 1500 bp, from about 100 bp to about 1600 bp, from about 100 bp to about 1700 bp, from about 100 bp to about 1800 bp, from about 100 bp to about 1900 bp, from about 100 bp to about 2000 bp, from about 100 bp to about 2100 bp, from about 100 bp to about 2200 bp, from about 100 bp to about 2300 bp, from about 100 bp to about 2400 bp, or from about 100 bp to about 2500 bp.
[0187] In another aspect of the methods provided herein that include digestion of whole cells or whole nuclei, the method further includes preparing a sequencing library from a plurality of segments prior to the size selection step. In some embodiments, the method further includes subjecting the sequencing library to size selection to obtain a size-selected library. Optionally, the size-selected library is on the order of about 350 bp to about 1000 bp. Optionally, the size-selected library is on the order of about 100 bp to about 2500 bp, such as about 100 bp to about 350 bp, about 350 bp to about 500 bp, about 500 bp to about 1000 bp, about 1000 bp to about 1500 bp and about 2000 bp, about 2000 bp to about 2500 bp, about 350 bp to about 1000 bp, about 350 bp to about 1500 bp, about 350 bp to about 2000 bp, about 350 bp to about 2500 bp, about 500 bp to about 1500 bp, about 500 bp to about 2000 bp, about 500 bp to about 3500 bp, about 1000 bp to about 1500 bp, about 1000 bp to about 2000 bp, about 1000 bp to about 2500 bp, about 1500 bp to about 2000 bp, about 1500 bp to about 2500 bp, or about 2000 bp to about 2500 bp.
[0188] The size selection utilized in the methods provided herein that include digestion of whole cells or whole nuclei can be performed using gel electrophoresis, capillary electrophoresis, size selection beads, gel filtration columns, or combinations thereof.
[0189] In another aspect, methods involving digestion of whole cells or whole nuclei herein may further include the step of further analyzing a plurality of selected segments to obtain QC values. In some cases, the QC values are selected from a chromatin digestion efficiency (CDE) and a chromatin digestion index (GDI). CDE is calculated as the percentage of segments having a desired length. For example, in some cases, CDE is calculated as the percentage of segments sized 100-2500 bp prior to size selection. In some cases, a sample is selected for further analysis when the CDE value is at least 65%. In some cases, a sample is selected for further analysis when the CDE value is at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, or at least about 95%. CDI is calculated as the ratio of the number of segments of mononucleosome size to the number of segments of dinucleosome size prior to size selection. For example, CDI can be calculated as the logarithm of the ratio of fragments having a size of 600-2500 bp to fragments having a size of 100-600 bp. In some cases, a sample is selected for further analysis when the CDI value is greater than -1.5 and less than 1.In some cases, when the CDI value is greater than about -2 and less than about 1.5, greater than about -1.9 and less than about 1.5, greater than about -1.8 and less than about 1.5, greater than about -1.7 and less than about 1.5, greater than about -1.6 and less than about 1.5, greater than about -1.5 and less than about 1.5, greater than about -1.4 and less than about 1.5, greater than about -1.3 and less than about 1.5, greater than about -1.2 and less than about 1.5, greater than about -1.1 and less than about 1.5, greater than about -2 and less than about 1.5, greater than about -1 and less than about 1.5, greater than about -0.9 and less than about 1.5, greater than about -0.8 and less than about 1.5, greater than about -0.7 and less than about 1.5, greater than about -0.6 and less than about 1.5, greater than about -0.5 and less than about 1.5, greater than about -2 and less than about 1.4, greater than about -2 and less than about 1.3, greater than about -2 and less than about 1.2, greater than about -2 and less than about 1.1, greater than about -2 and less than about 1, greater than about -2 and less than about 0.9, greater than about -2 and less than about 0.8, greater than about -2 and less than about 0.7, greater than about -2 and less than about 0.6, or greater than about -2 and less than about 0.5, the sample is selected for further analysis.
[0190] In another aspect, methods involving digestion of whole cells or whole nuclei herein are performed on small samples containing a small number of cells or a small amount of nucleic acid. In some cases, the stabilized biological sample contains fewer than 3,000,000 cells. In some cases, the stabilized biological sample contains fewer than 2,000,000 cells. In some cases, the stabilized biological sample contains fewer than 1,000,000 cells. In some cases, the stabilized biological sample contains fewer than 500,000 cells. In some cases, the stabilized biological sample contains fewer than 400,000 cells. In some cases, the stabilized biological sample contains fewer than 300,000 cells. In some cases, the stabilized biological sample contains fewer than 200,000 cells. In some cases, the stabilized biological sample contains fewer than 100,000 cells. In some cases, the stabilized biological sample contains fewer than 50,000 cells. In some cases, the stabilized biological sample contains fewer than 40,000 cells. In some cases, the stabilized biological sample contains fewer than 30,000 cells. In some cases, the stabilized biological sample contains fewer than 20,000 cells. In some cases, the stabilized biological sample contains fewer than 10,000 cells. In some cases, the stabilized biological sample contains about 10,000 cells. In some cases, the stabilized biological sample contains less than 10 μg of DNA. In some cases, the stabilized biological sample contains less than 9 μg of DNA. In some cases, the stabilized biological sample contains less than 8 μg of DNA. In some cases, the stabilized biological sample contains less than 7 μg of DNA. In some cases, the stabilized biological sample contains less than 6 μg of DNA. In some cases, the stabilized biological sample contains less than 5 μg of DNA. In some cases, the stabilized biological sample contains less than 4 μg of DNA. In some cases, the stabilized biological sample contains less than 3 μg of DNA. In some cases, the stabilized biological sample contains less than 2 μg of DNA. In some cases, the stabilized biological sample contains less than 1 μg of DNA. In some cases, the stabilized biological sample contains less than 0.5 μg of DNA.
[0191] In another aspect, methods involving digestion of whole cells or whole nuclei of this specification can be performed on individual or single cells. For example, the methods of this specification can be performed on cells dispensed into individual partitions. Exemplary partitions include, but are not limited to, wells, droplets in an emulsion, or surface locations (e.g., array spots, beads, etc.) containing discrete patches of differentially addressable linker molecules as described elsewhere in this specification. Additional partitions are contemplated and are consistent with the methods, compositions, and systems disclosed herein.
[0192] In a further aspect, the stabilized biological sample used in the methods involving digestion of whole cells or whole nuclei of this specification is treated with a nuclease, such as DNase, to create DNA fragments. In some cases, the DNase is non-sequence specific. In some cases, the DNase is active against both single-stranded DNA and double-stranded DNA. In some cases, the DNase is specific for double-stranded DNA. In some cases, the DNase preferentially cleaves double-stranded DNA. In some cases, the DNase is specific for single-stranded DNA. In some cases, the DNase preferentially cleaves single-stranded DNA. In some cases, the DNase is DNaseI. In some cases, the DNase is DNaseII. In some cases, the DNase is selected from one or more of DNaseI and DNaseII. In some cases, the DNase is micrococcal nuclease. In some cases, the DNase is selected from one or more of DNaseI, DNaseII, and micrococcal nuclease. In some cases, the DNase can be bound or fused to an immunoglobulin-binding protein or fragment thereof, such as protein A, protein G, protein A / G, or protein L. Other suitable nucleases are within the scope of the present disclosure.
[0193] In a further aspect, the stabilized biological sample used in a method involving digestion of whole cells or whole nuclei of the present specification is treated with a crosslinking agent. In some cases, the crosslinking agent is a chemical fixative. In some cases, the chemical fixative includes formaldehyde having a spacer arm length of about 2.3 to 2.7 angstroms (A). In some cases, the chemical fixative includes a crosslinking agent having a long spacer arm length. For example, the crosslinking agent can have a spacer length of at least about 3A, 4A, 5A, 6A, 7A, 8A, 9A, 10A, 11A, 12A, 13A, 14A, 15A, 16A, 17A, 18A, 19A, or 20A. The chemical fixative can include ethylene glycol bis(succinimidyl succinate) (EGS) having a spacer arm with a length of about 16.1A. The chemical fixative can include disuccinimidyl glutarate (DSG) having a spacer arm with a length of about 7.7A. In some cases, the chemical fixative includes formaldehyde and EGS, formaldehyde and DSG, or formaldehyde, EGS, and DSG. When multiple chemical fixatives are used, in some cases, each chemical fixative is used sequentially, and in other cases, some or all of the multiple chemical fixatives are applied to the sample simultaneously. The use of a crosslinking agent having a long spacer arm can increase the proportion of read pairs having a large (e.g., >1 kb) read pair separation distance. DSG is membrane-permeable and enables intracellular crosslinking. DSG can enhance crosslinking efficiency compared to disuccinimidyl suberate (DSS) in some applications. EGS has NHS ester reactive groups at both ends and can be reactive towards amino groups (e.g., primary amines). EGS is membrane-permeable and enables intracellular crosslinking. EGS crosslinking can be reversed, for example, by treatment with hydroxylamine at pH 8.5 for 3 to 6 hours. For example, lactate dehydrogenase retained 60% of its activity after reversible crosslinking with EGS. In some cases, the chemical fixative includes psoralen.In some cases, the crosslinking agent is ultraviolet light, chloromethine, cyclophosphamide, chlorambucil, uracil mustard, melphalan, bendamustine, bis(2-chloroethyl)ethylamine, bis(2-chloroethyl)methylamine, tris(2-chloroethyl)amine, isofamide, carmustine, lomustine, streptozocin, busulfan, cisplatin, carboplatin, cicycloplatin, eptaplatin, lobaplatin, miltiplatin, nedaplatin, oxaliplatin, picoplatin, satraplatin, triplatin tetranitrate, procarbazine, altretamine, dacarbazine, mitozolomide, temozolomide, mitomycin C, nitrous acid, formaldehyde, acetylaldehyde, doxorubicin, daunorubicin, epirubicin, or idarubicin. In some cases, the crosslinking agent includes an intercalator, an antibiotic, or a minor groove binder. In some cases, the stabilized biological sample is a crosslinked paraffin-embedded tissue sample.
[0194] In a further aspect, a method that includes digestion of whole cells or whole nuclei herein includes contacting a plurality of selected segments with an antibody.
[0195] In a further aspect, methods involving digestion of whole cells or whole nuclei herein include the step of joining a first segment and a second segment of a plurality of segments at a junction. Optionally, the joining step includes filling in sticky ends using biotin-tagged nucleotides and ligating blunt ends. Optionally, the joining step includes contacting at least the first segment and the second segment with a crosslinking oligonucleotide. Optionally, the joining step includes contacting at least the first segment and the second segment with a barcode. In some embodiments, the crosslinking oligonucleotides herein can be at least about 5 nucleotides in length to about 50 nucleotides in length. In some embodiments, the crosslinking oligonucleotides herein can be about 15 to about 18 nucleotides in length. In some embodiments, the crosslinking oligonucleotide can be about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, about 20, about 25, about 30, about 35, about 40, about 45, or about 50 nucleotides in length. In some embodiments, the crosslinking oligonucleotides herein can include a barcode. In some embodiments, the crosslinking oligonucleotide can include a plurality of barcodes. In some embodiments, the crosslinking oligonucleotide includes a plurality of crosslinking oligonucleotides bound to each other. In some embodiments, the crosslinking oligonucleotide can be coupled or linked to an immunoglobulin-binding protein such as protein A, protein G, protein A / G, or protein L or a fragment thereof. Optionally, the coupled crosslinking oligonucleotide can be delivered to a position in a sample nucleic acid to which an antibody binds.
[0196] Using a partitioning and pooling approach, cross-linked oligonucleotides with unique barcodes can be generated. A population of samples can be divided into multiple groups, and the cross-linked oligonucleotides can be ligated to the samples such that the cross-linked oligonucleotide barcodes differ between groups but are the same within a group. The groups of samples can be pooled again, and this process can be repeated multiple times. Repeating this process ultimately results in each sample in the population having a unique series of cross-linked oligonucleotide barcodes, enabling the analysis of a single sample (e.g., single cell, single nucleus, single chromosome). In one exemplary example, a sample of cross-linked digested nuclei bound to a solid support of beads is divided into eight tubes, each containing one of eight unique members of a first adapter group (first iteration) that includes a double-stranded DNA (dsDNA) adapter to be ligated. Each of the eight adapters can have the same 5’ overhang sequence for ligation to the nucleic acid termini of cross-linked chromatin aggregates in the nuclei, but otherwise has a unique dsDNA sequence. After ligating the first adapter group, the nuclei can be pooled back and washed to remove the ligation reaction components. The partitioning, ligation, and pooling scheme can be repeated two more times (two iterations). After ligation of members from each adapter group, cross-linked chromatin aggregates can be successively ligated to multiple barcodes. In some cases, successive ligation (iteration) of multiple members of multiple adapter groups results in a combination of barcodes. The number of possible combinations of barcodes depends on the number of groups per iteration and the total number of barcode oligonucleotides used. For example, three iterations each containing eight members can have 83 possible combinations. In some cases, the combination of barcodes is unique. In some cases, the combination of barcodes is redundant. The total number of combinations of barcodes can be adjusted by increasing or decreasing the number of groups receiving unique barcodes and / or by increasing or decreasing the number of iterations.When an adapter set greater than 1 is used, distribution, ligation, and pooling schemes can be used for iterative adapter ligation. Optionally, the distribution, ligation, and pooling schemes can be further repeated at least 3, 4, 5, 6, 7, 8, 9, or 10 times. Optionally, the members of the last adapter set include sequences for subsequent enrichment of the DNA ligated to the adapter, for example, during preparation of a sequencing library by PCR amplification.
[0197] In a further aspect, methods that include digestion of whole cells or whole nuclei herein do not include a shearing step.
[0198] In a further aspect of methods that include digestion of whole cells or whole nuclei herein, the method includes obtaining at least some sequences on each side of the junction to generate a first read pair. For example, the method can include obtaining sequences of at least about 50 bp, at least about 100 bp, at least about 150 bp, at least about 200 bp, at least about 250 bp, or at least about 300 bp on each side of the junction to generate a first read pair.
[0199] In a further aspect of methods that include digestion of whole cells or whole nuclei herein, the method includes mapping a first read pair to a set of contigs and determining a path through the set of contigs that represents the order and / or orientation relative to the genome.
[0200] In a further aspect of methods that include digestion of whole cells or whole nuclei herein, the method includes mapping a first read pair to a set of contigs and determining the presence of a structural variant or loss of heterozygosity in a stabilized biological sample from the set of contigs.
[0201] In a further aspect of methods that include digestion of whole cells or whole nuclei herein, the method includes mapping a first read pair to a set of contigs and assigning a phase to the variants in the set of contigs.
[0202] In a further aspect of the method involving digestion of whole cells or whole nuclei of the present specification, the method comprises mapping a first read pair to a set of contigs, determining the presence of variants in the set of contigs from the set of contigs, and (1) identifying a disease stage, prognosis, or course of treatment for a stabilized biological sample, (2) selecting a drug based on the presence of the variant, or (3) identifying the drug efficacy for a stabilized biological sample. The method includes performing a step selected from one or more of the above.
[0203] Hi-C method with low nucleic acid input requirements A step of obtaining a stabilized biological sample comprising a nucleic acid molecule complexed with at least one nucleic acid-binding protein, a step of contacting the stabilized biological sample with DNase to cleave the nucleic acid molecule into a plurality of segments, and a step of ligating a first segment and a second segment of the plurality of segments at a junction are further provided herein, and the stabilized biological sample contains less than 3,000,000 cells or less than 10 μg of DNA. In some cases, the stabilized biological sample contains less than 3,000,000 cells. In some cases, the stabilized biological sample contains less than 2,000,000 cells. In some cases, the stabilized biological sample contains less than 1,000,000 cells. In some cases, the stabilized biological sample contains less than 500,000 cells. In some cases, the stabilized biological sample contains less than 400,000 cells. In some cases, the stabilized biological sample contains less than 300,000 cells. In some cases, the stabilized biological sample contains less than 200,000 cells. In some cases, the stabilized biological sample contains less than 100,000 cells. In some cases, the stabilized biological sample contains less than 50,000 cells. In some cases, the stabilized biological sample contains less than 40,000 cells. In some cases, the stabilized biological sample contains less than 30,000 cells. In some cases, the stabilized biological sample contains less than 20,000 cells. In some cases, the stabilized biological sample contains less than 10,000 cells. In some cases, the stabilized biological sample contains about 10,000 cells. In some cases, the sample contains at least 10,000 cells. In some cases, the sample contains at least 20,000 cells. In some cases, the sample contains at least 30,000 cells. In some cases, the sample contains at least 40,000 cells. In some cases, the sample contains from about 10,000 cells to about 50,000 cells. In some cases, the sample contains from about 20,000 cells to about 50,000 cells. In some cases, the sample contains from about 30,000 cells to about 50,000 cells. In some cases, the sample contains from about 40,000 cells to about 50,000 cells.In some cases, the sample contains approximately 10,000 cells to approximately 40,000 cells. In some cases, the sample contains approximately 10,000 cells to approximately 30,000 cells. In some cases, the sample contains approximately 10,000 cells to approximately 20,000 cells. In some cases, the sample contains approximately 20,000 cells to approximately 50,000 cells. In some cases, the sample contains approximately 20,000 cells to approximately 40,000 cells. In some cases, the sample contains approximately 20,000 cells to approximately 30,000 cells. In some cases, the sample contains approximately 30,000 cells to approximately 50,000 cells. In some cases, the sample contains approximately 30,000 cells to approximately 40,000 cells. In some cases, the stabilized biological sample contains less than 10 μg of DNA. In some cases, the stabilized biological sample contains less than 9 μg of DNA. In some cases, the stabilized biological sample contains less than 8 μg of DNA. In some cases, the stabilized biological sample contains less than 7 μg of DNA. In some cases, the stabilized biological sample contains less than 6 μg of DNA. In some cases, the stabilized biological sample contains less than 5 μg of DNA. In some cases, the stabilized biological sample contains less than 4 μg of DNA. In some cases, the stabilized biological sample contains less than 3 μg of DNA. In some cases, the stabilized biological sample contains less than 2 μg of DNA. In some cases, the stabilized biological sample contains less than 1 μg of DNA. In some cases, the stabilized biological sample contains less than 0.5 μg of DNA.
[0204] In various aspects of the methods herein, the stabilized sample may contain nuclei. In some cases, the stabilized biological sample contains 50,000 nuclei or fewer. In some cases, the sample contains 40,000 nuclei or fewer. In some cases, the sample contains 30,000 nuclei or fewer. In some cases, the sample contains 20,000 nuclei or fewer. In some cases, the sample contains at least 10,000 nuclei. In some cases, the sample contains at least 20,000 nuclei. In some cases, the sample contains at least 30,000 nuclei. In some cases, the sample contains at least 40,000 nuclei. In some cases, the sample contains from about 10,000 nuclei to about 50,000 nuclei. In some cases, the sample contains from about 20,000 nuclei to about 50,000 nuclei. In some cases, the sample contains from about 30,000 nuclei to about 50,000 nuclei. In some cases, the sample contains from about 40,000 nuclei to about 50,000 nuclei. In some cases, the sample contains from about 10,000 nuclei to about 40,000 nuclei. In some cases, the sample contains from about 10,000 nuclei to about 30,000 nuclei. In some cases, the sample contains from about 10,000 nuclei to about 20,000 nuclei. In some cases, the sample contains from about 20,000 nuclei to about 50,000 nuclei. In some cases, the sample contains from about 20,000 nuclei to about 40,000 nuclei. In some cases, the sample contains from about 20,000 nuclei to about 30,000 nuclei. In some cases, the sample contains from about 30,000 nuclei to about 50,000 nuclei. In some cases, the sample contains from about 30,000 nuclei to about 40,000 nuclei.
[0205] In another aspect, the methods herein with low nucleic acid input requirements can be performed on individual or single cells. For example, the methods herein can be performed on cells distributed into individual partitions. Exemplary partitions include, but are not limited to, wells, droplets in an emulsion, or surface locations (e.g., array spots, beads, etc.) containing discrete patches of differentially addressable linker molecules as described elsewhere herein. Additional partitions are contemplated and are consistent with the methods, compositions, and systems disclosed herein.
[0206] In another aspect, the methods having low nucleic acid input requirements herein may include subjecting a plurality of segments to size selection to obtain a plurality of selected segments. Optionally, the plurality of selected segments are from about 145 to about 600 bp. Optionally, the plurality of selected segments are from about 100 to about 2500 bp. Optionally, the plurality of selected segments are from about 100 to about 600 bp. Optionally, the plurality of selected segments are from about 600 to about 2500 bp. Optionally, the plurality of selected segments are from about 100 bp to about 600 bp, from about 100 bp to about 700 bp, from about 100 bp to about 800 bp, from about 100 bp to about 900 bp, from about 100 bp to about 1000 bp, from about 100 bp to about 1100 bp, from about 100 bp to about 1200 bp, from about 100 bp to about 1300 bp, from about 100 bp to about 1400 bp, from about 100 bp to about 1500 bp, from about 100 bp to about 1600 bp, from about 100 bp to about 1700 bp, from about 100 bp to about 1800 bp, from about 100 bp to about 1900 bp, from about 100 bp to about 2000 bp, from about 100 bp to about 2100 bp, from about 100 bp to about 2200 bp, from about 100 bp to about 2300 bp, from about 100 bp to about 2400 bp, or from about 100 bp to about 2500 bp.
[0207] In another aspect of the methods provided herein with low nucleic acid input requirements, the method further includes preparing a sequencing library from a plurality of segments prior to the size selection step. In some embodiments, the method further includes subjecting the sequencing library to size selection to obtain a size selected library. Optionally, the size selected library is about 350 bp to about 1000 bp in size. Optionally, the size selected library is about 100 bp to about 2500 bp in size, such as about 100 bp to about 350 bp, about 350 bp to about 500 bp, about 500 bp to about 1000 bp, about 1000 bp to about 1500 bp and about 2000 bp, about 2000 bp to about 2500 bp, about 350 bp to about 1000 bp, about 350 bp to about 1500 bp, about 350 bp to about 2000 bp, about 350 bp to about 2500 bp, about 500 bp to about 1500 bp, about 500 bp to about 2000 bp, about 500 bp to about 3500 bp, about 1000 bp to about 1500 bp, about 1000 bp to about 2000 bp, about 1000 bp to about 2500 bp, about 1500 bp to about 2000 bp, about 1500 bp to about 2500 bp, or about 2000 bp to about 2500 bp.
[0208] The size selection utilized in the methods with low nucleic acid input requirements herein can be performed using gel electrophoresis, capillary electrophoresis, size selection beads, gel filtration columns, or combinations thereof.
[0209] In another aspect, the methods having low nucleic acid input requirements herein may further include the step of analyzing a plurality of selected segments to obtain a QC value. Optionally, the QC value is selected from a chromatin digestion efficiency (CDE) and a chromatin digestion index (GDI). CDE is calculated as the percentage of segments having a desired length. For example, optionally, CDE is calculated as the percentage of segments sized 100 - 2500 bp before size selection. Optionally, a sample is selected for further analysis when the CDE value is at least 65%. Optionally, a sample is selected for further analysis when the CDE value is at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, or at least about 95%. CDI is calculated as the ratio of the number of segments of mononucleosome size to the number of segments of dinucleosome size before size selection. For example, CDI can be calculated as the logarithm of the ratio of fragments having a size of 600 - 2500 bp to fragments having a size of 100 - 600 bp. Optionally, a sample is selected for further analysis when the CDI value is greater than -1.5 and less than 1.In some cases, when the CDI value is greater than about -2 and less than about 1.5, greater than about -1.9 and less than about 1.5, greater than about -1.8 and less than about 1.5, greater than about -1.7 and less than about 1.5, greater than about -1.6 and less than about 1.5, greater than about -1.5 and less than about 1.5, greater than about -1.4 and less than about 1.5, greater than about -1.3 and less than about 1.5, greater than about -1.2 and less than about 1.5, greater than about -1.1 and less than about 1.5, greater than about -2 and less than about 1.5, greater than about -1 and less than about 1.5, greater than about -0.9 and less than about 1.5, greater than about -0.8 and less than about 1.5, greater than about -0.7 and less than about 1.5, greater than about -0.6 and less than about 1.5, greater than about -0.5 and less than about 1.5, greater than about -2 and less than about 1.4, greater than about -2 and less than about 1.3, greater than about -2 and less than about 1.2, greater than about -2 and less than about 1.1, greater than about -2 and less than about 1, greater than about -2 and less than about 0.9, greater than about -2 and less than about 0.8, greater than about -2 and less than about 0.7, greater than about -2 and less than about 0.6, or greater than about -2 and less than about 0.5, the sample is selected for further analysis.
[0210] In another aspect, a stabilized biological sample used in a method having low nucleic acid input requirements herein comprises a biological material treated with a stabilizer. In some cases, the stabilized biological sample comprises a stabilized cell lysate. Alternatively, the stabilized biological sample comprises stabilized intact cells. Alternatively, the stabilized biological sample comprises stabilized intact nuclei. In some cases, the step of contacting the stabilized intact cells or intact nucleus sample with DNase is performed prior to lysis of the intact cells or intact nuclei. In some cases, the cells and / or nuclei are lysed prior to joining a first segment and a second segment of a plurality of segments at a junction.
[0211] In a further aspect, the stabilized biological sample used in the methods having low nucleic acid input requirements herein is treated with a nuclease, such as DNase, to generate DNA fragments. In some cases, the DNase is non-sequence specific. In some cases, the DNase is active against both single-stranded and double-stranded DNA. In some cases, the DNase is specific for double-stranded DNA. In some cases, the DNase preferentially cleaves double-stranded DNA. In some cases, the DNase is specific for single-stranded DNA. In some cases, the DNase preferentially cleaves single-stranded DNA. In some cases, the DNase is DNaseI. In some cases, the DNase is DNaseII. In some cases, the DNase is selected from one or more of DNaseI and DNaseII. In some cases, the DNase is micrococcal nuclease. In some cases, the DNase is selected from one or more of DNaseI, DNaseII, and micrococcal nuclease. In some cases, the DNase can be bound or fused to an immunoglobulin-binding protein or fragment thereof, such as protein A, protein G, protein A / G, or protein L. Other suitable nucleases are also within the scope of the disclosure.
[0212] In a further aspect, the stabilized biological sample used in the methods having low nucleic acid input requirements herein is treated with a crosslinking agent. In some cases, the crosslinking agent is a chemical fixative. In some cases, the chemical fixative includes formaldehyde having a spacer arm length of about 2.3 to 2.7 angstroms (A). In some cases, the chemical fixative includes a crosslinking agent having a long spacer arm length. For example, the crosslinking agent can have a spacer length of at least about 3A, 4A, 5A, 6A, 7A, 8A, 9A, 10A, 11A, 12A, 13A, 14A, 15A, 16A, 17A, 18A, 19A, or 20A. The chemical fixative can include ethylene glycol bis(succinimidyl succinate) (EGS) having a spacer arm with a length of about 16.1A. The chemical fixative can include disuccinimidyl glutarate (DSG) having a spacer arm with a length of about 7.7A. In some cases, the chemical fixative includes formaldehyde and EGS, formaldehyde and DSG, or formaldehyde, EGS, and DSG. In some cases where multiple chemical fixatives are used, each chemical fixative is used sequentially, and in other cases, some or all of the multiple chemical fixatives are applied to the sample simultaneously. The use of a crosslinking agent having a long spacer arm can increase the proportion of read pairs having a large (e.g., >1 kb) read pair separation distance. DSG is membrane-permeable and enables intracellular crosslinking. DSG can enhance crosslinking efficiency compared to disuccinimidyl suberate (DSS) in some applications. EGS has NHS ester reactive groups at both ends and can be reactive towards amino groups (e.g., primary amines). EGS is membrane-permeable and enables intracellular crosslinking. EGS crosslinking can be reversed, for example, by treatment with hydroxylamine at pH 8.5 for 3 to 6 hours. For example, lactate dehydrogenase retained 60% of its activity after reversible crosslinking with EGS. In some cases, the chemical fixative includes psoralen.In some cases, the crosslinking agent is ultraviolet light, chloromethine, cyclophosphamide, chlorambucil, uracil mustard, melphalan, bendamustine, bis(2-chloroethyl)ethylamine, bis(2-chloroethyl)methylamine, tris(2-chloroethyl)amine, isofamide, carmustine, lomustine, streptozocin, busulfan, cisplatin, carboplatin, cicycloplatin, eptaplatin, lobaplatin, milphlatine, nedaplatin, oxaliplatin, picoplatin, satraplatin, triplatin tetranitrate, procarbazine, altretamine, dacarbazine, mitozolomide, temozolomide, mitomycin C, nitrous acid, formaldehyde, acetylaldehyde, doxorubicin, daunorubicin, epirubicin, or idarubicin. In some cases, the crosslinking agent includes an intercalator, an antibiotic, or a minor groove binder. In some cases, the stabilized biological sample is a crosslinked paraffin-embedded tissue sample.
[0213] In a further aspect, the methods provided herein include contacting a plurality of selected segments with an antibody.
[0214] In a further aspect, the methods provided herein with low nucleic acid input requirements include the step of joining a first segment and a second segment of a plurality of segments at a junction. Optionally, the joining step includes filling in sticky ends using biotin-tagged nucleotides and ligating blunt ends. Optionally, the joining step includes contacting at least the first segment and the second segment with a crosslinking oligonucleotide. Optionally, the joining step includes contacting at least the first segment and the second segment with a barcode. In some embodiments, the crosslinking oligonucleotides herein can be at least about 5 nucleotides to about 50 nucleotides in length. In some embodiments, the crosslinking oligonucleotides herein can be about 15 to about 18 nucleotides in length. In some embodiments, the crosslinking oligonucleotide can be about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, about 20, about 25, about 30, about 35, about 40, about 45, or about 50 nucleotides in length. In some embodiments, the crosslinking oligonucleotides herein can include a barcode. In some embodiments, the crosslinking oligonucleotide can include a plurality of barcodes. In some embodiments, the crosslinking oligonucleotide includes a plurality of crosslinking oligonucleotides bound to each other. In some embodiments, the crosslinking oligonucleotide can be coupled or linked to an immunoglobulin-binding protein such as protein A, protein G, protein A / G, or protein L or a fragment thereof. Optionally, the coupled crosslinking oligonucleotide can be delivered to a position in the sample nucleic acid to which the antibody binds.
[0215] Cross - linking oligonucleotides with unique barcodes can be generated using a partitioning and pooling approach. A population of samples can be divided into multiple groups, and the cross - linking oligonucleotides can be ligated to the samples such that the cross - linking oligonucleotide barcodes are different between groups but the same within a group. The groups of samples can be pooled again, and this process can be repeated multiple times. Repeating this process will ultimately result in each sample in the population having a unique set of cross - linking oligonucleotide barcodes, enabling the analysis of a single sample (e.g., single cell, single nucleus, single chromosome). In one exemplary example, a sample of cross - linked digested nuclei bound to a solid support of beads is divided into eight tubes, each containing one of eight unique members of a first adapter group (first iteration) that includes a double - stranded DNA (dsDNA) adapter to be ligated. Each of the eight adapters can have the same 5’ overhang sequence for ligation to the nucleic acid termini of cross - linked chromatin aggregates in the nucleus, but otherwise has a unique dsDNA sequence. After ligating the first adapter group, the nuclei can be pooled back, washed to remove ligation reaction components. The partitioning, ligation, and pooling scheme can be repeated two more times (two iterations). After ligation of members from each adapter group, cross - linked chromatin aggregates can be successively ligated to multiple barcodes. In some cases, successive ligation (iteration) of multiple members of multiple adapter groups results in a combination of barcodes. The number of possible combinations of barcodes depends on the number of groups per iteration and the total number of barcode oligonucleotides used. For example, three iterations each containing eight members can have 8^3 possible combinations. In some cases, the combination of barcodes is unique. In some cases, the combination of barcodes is redundant. The total number of barcode combinations can be adjusted by increasing or decreasing the number of groups receiving unique barcodes and / or by increasing or decreasing the number of iterations.When an adapter group greater than 1 is used, for iterative adapter ligation, partitioning, ligation, and pooling schemes can be used. Optionally, the partitioning, ligation, and pooling schemes can be further repeated at least 3, 4, 5, 6, 7, 8, 9, or 10 times. Optionally, the members of the last adapter group contain sequences for subsequent enrichment of the DNA ligated to the adapter, for example, during preparation of a sequencing library by PCR amplification.
[0216] In a further aspect, the methods having the low nucleic acid input requirements herein do not include a shearing step.
[0217] In a further aspect of the methods having the low nucleic acid input requirements herein, the method includes obtaining at least some sequences on each side of the junction to generate a first read pair. For example, the method can include obtaining sequences of at least about 50 bp, at least about 100 bp, at least about 150 bp, at least about 200 bp, at least about 250 bp, or at least about 300 bp on each side of the junction to generate a first read pair.
[0218] In a further aspect of the methods having the low nucleic acid input requirements herein, the method includes mapping a first read pair to a set of contigs and determining a path through the set of contigs that represents the order and / or orientation relative to the genome.
[0219] In a further aspect of the methods having the low nucleic acid input requirements herein, the method can include mapping a first read pair to a set of contigs and determining the presence of a structural variant or loss of heterozygosity in the stabilized biological sample from the set of contigs.
[0220] In a further aspect of the methods having the low nucleic acid input requirements herein, the method includes mapping a first read pair to a set of contigs and assigning phases to variants in the set of contigs.
[0221] In a further aspect of the methods having low nucleic acid input requirements of the present specification, the method comprises mapping a first read pair to a set of contigs, determining the presence of variants in the set of contigs from the set of contigs, and (1) identifying a disease stage, prognosis, or treatment course for a stabilized biological sample, (2) selecting a drug based on the presence of the variant, or (3) performing a step selected from one or more of identifying the drug efficacy for the stabilized biological sample.
[0222] Hi-C method using micrococcal nuclease (MNase) Provided herein is a method that may include obtaining a stabilized biological sample comprising a nucleic acid molecule complexed with at least one nucleic acid-binding protein, contacting the stabilized biological sample with micrococcal nuclease (MNase) to cleave the nucleic acid molecule into a plurality of segments, and ligating a first segment and a second segment of the plurality of segments at the junction. The use of MNase in the methods herein can provide specific information regarding where the DNA-binding protein is bound to chromatin with up to single-base pair resolution because, for example, MNase can cleave all base pairs that are not bound to a DNA-binding protein. Further, the use of MNase digestion can enable the generation of contact maps and topologically associated domains for decoding three-dimensional chromatin structure information. Optionally, MNase can be bound or fused to an immunoglobulin-binding protein or fragment thereof, such as protein A, protein G, protein A / G, or protein L.
[0223] For example, the MNase Hi-C method can provide the positions of protein binding or genomic contact interactions at a resolution of about 1 bp, 2 bp, 3 bp, 4 bp, 5 bp, 6 bp, 7 bp, 8 bp, 9 bp, 10 bp, 20 bp, 30 bp, 40 bp, 50 bp, 60 bp, 70 bp, 80 bp, 90 bp, 100 bp, 200 bp, 300 bp, 400 bp, 500 bp, 600 bp, 700 bp, 800 bp, 900 bp, 1000 bp, 2000 bp, 3000 bp, 4000 bp, 5000 bp, 6000 bp, 7000 bp, 8000 bp, 9000 bp, 10 kb, 20 kb, 30 kb, 40 kb, 50 kb, 60 kb, 70 kb, 80 kb, 90 kb, or 100 kb or less. In some cases, protein binding sites, protein footprints, contact interactions, or other features can be mapped within 1000 bp, within 900 bp, within 800 bp, within 700 bp, within 600 bp, within 500 bp, within 400 bp, within 300 bp, within 200 bp, within 190 bp, within 180 bp, within 170 bp, within 160 bp, within 150 bp, within 140 bp, within 130 bp, within 120 bp, within 110 bp, within 100 bp, within 90 bp, within 80 bp, within 70 bp, within 60 bp, within 50 bp, within 40 bp, within 30 bp, within 20 bp, within 10 bp, within 9 bp, within 8 bp, within 7 bp, within 6 bp, within 5 bp, within 4 bp, within 3 bp, within 2 bp, or within 1 bp.
[0224] In certain embodiments, a method that includes an MNase digestion step may include a step of size selecting a plurality of segments to obtain the plurality of selected segments. Optionally, the plurality of selected segments can be from about 145 to about 600 bp. Optionally, the plurality of selected segments can be from about 100 to about 2500 bp. Optionally, the plurality of selected segments can be from about 100 to about 600 bp. Optionally, the plurality of selected segments can be from about 600 to about 2500 bp. Optionally, the plurality of selected segments can be from about 100 bp to about 600 bp, from about 100 bp to about 700 bp, from about 100 bp to about 800 bp, from about 100 bp to about 900 bp, from about 100 bp to about 1000 bp, from about 100 bp to about 1100 bp, from about 100 bp to about 1200 bp, from about 100 bp to about 1300 bp, from about 100 bp to about 1400 bp, from about 100 bp to about 1500 bp, from about 100 bp to about 1600 bp, from about 100 bp to about 1700 bp, from about 100 bp to about 1800 bp, from about 100 bp to about 1900 bp, from about 100 bp to about 2000 bp, from about 100 bp to about 2100 bp, from about 100 bp to about 2200 bp, from about 100 bp to about 2300 bp, from about 100 bp to about 2400 bp, or from about 100 bp to about 2500 bp.
[0225] In another aspect of the method comprising the MNase digestion step provided herein, the method may further comprise preparing a sequencing library from a plurality of segments. In some embodiments, the method may further comprise subjecting the sequencing library to size selection to obtain a size-selected library. Optionally, the size-selected library may be in the size range of about 350 bp to about 1000 bp. Optionally, the size-selected library may be in the size range of about 100 bp to about 2500 bp, such as about 100 bp to about 350 bp, about 350 bp to about 500 bp, about 500 bp to about 1000 bp, about 1000 bp to about 1500 bp, about 2000 bp to about 2500 bp, about 350 bp to about 1000 bp, about 350 bp to about 1500 bp, about 350 bp to about 2000 bp, about 350 bp to about 2500 bp, about 500 bp to about 1500 bp, about 500 bp to about 2000 bp, about 500 bp to about 2500 bp, about 1000 bp to about 1500 bp, about 1000 bp to about 2000 bp, about 1000 bp to about 2500 bp, about 1500 bp to about 2000 bp, about 1500 bp to about 2500 bp, or about 2000 bp to about 2500 bp.
[0226] In another aspect, the method comprising the MNase digestion step provided herein can further comprise analyzing a plurality of segments to obtain QC values. Optionally, the QC values can be selected from chromatin digestion efficiency (CDE) and chromatin digestion index (GDI). CDE can be calculated as the percentage of segments having a desired length. For example, optionally, CDE can be calculated as the percentage of segments in the size range of 100 bp to 2500 bp before size selection. Optionally, a sample is selected for further analysis when the CDE value is at least 65%. Optionally, a sample can be selected for further analysis when the CDE value is at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, or at least about 95%.
[0227] CDI can be calculated as the ratio of the number of segments of mononucleosome size to the number of segments of dinucleosome size before size selection. For example, CDI can be calculated as the logarithm of the ratio of fragments having a size of 600 to 2500 bp to fragments having a size of 100 to 600 bp. In some cases, the sample is selected for further analysis if the CDI value is greater than -1.5 and less than 1. In some cases, the CDI value is greater than about -2 and less than about 1.5, greater than about -1.9 and less than about 1.5, greater than about -1.8 and less than about 1.5, greater than about -1.7 and less than about 1.5, greater than about -1.6 and less than about 1.5, greater than about -1.5 and less than about 1.5, greater than about -1.4 and less than about 1.5, greater than about -1.3 and less than about 1.5, greater than about -1.2 and less than about 1.5, greater than about -1.1 and less than about 1.5, greater than about -2 and less than about 1.5, greater than about -1 and less than about 1.5, greater than about -0.9 and less than about 1.5, greater than about -0.8 and less than about 1.5, greater than about -0.7 and less than about 1.5, greater than about -0.6 and less than about 1.5, greater than about -0.5 and less than about 1.5, greater than about -2 and less than about 1.4, greater than about -2 and less than about 1.3, greater than about -2 and less than about 1.2, greater than about -2 and less than about 1.1, greater than about -2 and less than about 1, greater than about -2 and less than about 0.9, greater than about -2 and less than about 0.8, greater than about -2 and less than about 0.7, greater than about -2 and less than about 0.6, or greater than about -2 and less than about 0.5, and the sample is selected for further analysis.
[0228] In another aspect, the stabilized biological sample used in the methods provided herein that include an MNase digestion step comprises biological material that has been treated with a stabilizer. Optionally, the stabilized biological sample can include a stabilized cell lysate. Alternatively, the stabilized biological sample can include stabilized intact cells. Alternatively, the stabilized biological sample can include stabilized intact nuclei. Optionally, the step of contacting the stabilized intact cells or intact nuclei sample with MNase can be performed prior to lysis of the intact cells or intact nuclei. Optionally, the cells and / or nuclei can be lysed prior to joining a first segment and a second segment of a plurality of segments at a junction.
[0229] In another aspect, the methods of the present disclosure that include the MNase digestion step can be performed on a small sample containing a small number of cells or a small amount of nucleic acid. For example, in some cases, the stabilized biological sample can contain less than 3,000,000 cells. In some cases, the stabilized biological sample can contain less than 2,000,000 cells. In some cases, the stabilized biological sample can contain less than 1,000,000 cells. In some cases, the stabilized biological sample can contain less than 500,000 cells. In some cases, the stabilized biological sample can contain less than 400,000 cells. In some cases, the stabilized biological sample can contain less than 300,000 cells. In some cases, the stabilized biological sample can contain less than 200,000 cells. In some cases, the stabilized biological sample can contain less than 100,000 cells. In some cases, the stabilized biological sample can contain less than 50,000 cells. In some cases, the stabilized biological sample can contain less than 40,000 cells. In some cases, the stabilized biological sample can contain less than 30,000 cells. In some cases, the stabilized biological sample can contain less than 20,000 cells. In some cases, the stabilized biological sample can contain less than 10,000 cells. In some cases, the stabilized biological sample can contain about 10,000 cells. In some cases, the stabilized biological sample can contain less than 10 μg of DNA. In some cases, the stabilized biological sample can contain less than 9 μg of DNA. In some cases, the stabilized biological sample can contain less than 8 μg of DNA. In some cases, the stabilized biological sample can contain less than 7 μg of DNA. In some cases, the stabilized biological sample can contain less than 6 μg of DNA. In some cases, the stabilized biological sample can contain less than 5 μg of DNA. In some cases, the stabilized biological sample can contain less than 4 μg of DNA. In some cases, the stabilized biological sample can contain less than 3 μg of DNA. In some cases, the stabilized biological sample can contain less than 2 μg of DNA. In some cases, the stabilized biological sample can contain less than 1 μg of DNA. In some cases, the stabilized biological sample can contain less than 0.5 μg of DNA.
[0230] In another aspect, the methods of the present disclosure that include the MNase digestion step may be performed on individual cells or a single cell. For example, the methods of the present disclosure can be performed on cells distributed in individual partitions. Exemplary partitions include, but are not limited to, wells, droplets in an emulsion, or surface locations (e.g., array spots, beads, etc.) that include discrete patches of differentially addressable linker molecules as described elsewhere in the present disclosure. Additional partitions are contemplated and are consistent with the methods, compositions, and systems disclosed herein.
[0231] In a further aspect, the stabilized biological sample used in the methods of the present disclosure that include the MNase digestion step can be further treated with an additional nuclease, such as DNase, to generate DNA fragments. In some cases, the DNase can be non-sequence specific. In some cases, the DNase can be active against both single-stranded and double-stranded DNA. In some cases, the DNase can be specific for double-stranded DNA. In some cases, the DNase can preferentially cleave double-stranded DNA. In some cases, the DNase can be specific for single-stranded DNA. In some cases, the DNase can preferentially cleave single-stranded DNA. In some cases, the DNase can be DNaseI. In some cases, the DNase can be DNaseII. In some cases, the DNase can be selected from one or more of DNaseI and DNaseII. In some cases, the DNase can be bound or fused to an immunoglobulin-binding protein or fragment thereof, such as Protein A, Protein G, Protein A / G, or Protein L. Other suitable nucleases are within the scope of the present disclosure.
[0232] In a further aspect, the stabilized biological sample provided herein for use in a method comprising an MNase digestion step can be treated with a crosslinking agent. Optionally, the crosslinking agent can be a chemical fixative. Optionally, the chemical fixative contains formaldehyde having a spacer arm length of about 2.3 to 2.7 angstroms (A). Optionally, the chemical fixative contains a crosslinking agent having a long spacer arm length. For example, the crosslinking agent can have a spacer length of at least about 3A, 4A, 5A, 6A, 7A, 8A, 9A, 10A, 11A, 12A, 13A, 14A, 15A, 16A, 17A, 18A, 19A, or 20A. The chemical fixative can contain ethylene glycol bis(succinimidyl succinate) (EGS) having a spacer arm with a length of about 16.1A. The chemical fixative can contain disuccinimidyl glutarate (DSG) having a spacer arm with a length of about 7.7A. Optionally, the chemical fixative contains formaldehyde and EGS, formaldehyde and DSG, or formaldehyde, EGS, and DSG. Optionally, when multiple chemical fixatives are used, each chemical fixative is used sequentially, and in other cases, some or all of the multiple chemical fixatives are applied to the sample simultaneously. The use of a crosslinking agent with a long spacer arm can increase the proportion of read pairs having a large (e.g., >1 kb) read pair separation distance. DSG is membrane-permeable and enables intracellular crosslinking. DSG can enhance crosslinking efficiency compared to disuccinimidyl suberate (DSS) in some applications. EGS has NHS ester reactive groups at both ends and can be reactive towards amino groups (e.g., primary amines). EGS is membrane-permeable and enables intracellular crosslinking. EGS crosslinking can be reversed, for example, by treatment with hydroxylamine at pH 8.5 for 3 to 6 hours. For example, lactate dehydrogenase retained 60% of its activity after reversible crosslinking with EGS. Optionally, the chemical fixative can contain psoralen.In some cases, the cross-linking agent can be ultraviolet light, chloromethane, cyclophosphamide, chlorambucil, uracil mustard, melphalan, bendamustine, bis(2-chloroethyl)ethylamine, bis(2-chloroethyl)methylamine, tris(2-chloroethyl)amine, isofamide, carmustine, lomustine, streptozocin, busulfan, cisplatin, carboplatin, cicycloplatin, eptaplatin, lobaplatin, miltiplatin, nedaplatin, oxaliplatin, picoplatin, satraplatin, triplatin tetranitrate, procarbazine, altretamine, dacarbazine, mitozolomide, temozolomide, mitomycin C, nitrous acid, formaldehyde, acetylaldehyde, doxorubicin, daunorubicin, epirubicin, or idarubicin. In some cases, the cross-linking agent includes an intercalator, an antibiotic, or a minor groove binder. In some cases, the stabilized biological sample can be a cross-linked paraffin-embedded tissue sample.
[0233] In a further aspect, a method comprising the MNase digestion step provided herein can include contacting a plurality of selected segments with an antibody. In some cases, an immunoglobulin-binding protein or a fragment thereof tethered to an oligonucleotide adapter can be tethered to an antibody bound to a plurality of selected segments.
[0234] In a further aspect, a method comprising the MNase digestion step provided herein includes a step of joining a first segment and a second segment of a plurality of segments at a junction. Optionally, the joining step may include filling in sticky ends using biotin-tagged nucleotides and ligating blunt ends. Optionally, the joining step may include contacting at least the first segment and the second segment with a crosslinking oligonucleotide. Optionally, the joining step may include contacting at least the first segment and the second segment with a barcode. In some embodiments, the crosslinking oligonucleotide herein can be at least about 5 nucleotides to about 50 nucleotides in length. In some embodiments, the crosslinking oligonucleotide herein can be about 15 to about 18 nucleotides in length. In some embodiments, the crosslinking oligonucleotide can be about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, about 20, about 25, about 30, about 35, about 40, about 45, or about 50 nucleotides in length. In some embodiments, the crosslinking oligonucleotide herein can include a barcode.
[0235] In a further aspect of the method comprising the MNase digestion step herein, the method can include obtaining at least some sequences on each side of the junction to generate a first read pair. For example, the method can include obtaining sequences of at least about 50 bp, at least about 100 bp, at least about 150 bp, at least about 200 bp, at least about 250 bp, or at least about 300 bp on each side of the junction to generate a first read pair.
[0236] In a further aspect of the method comprising the MNase digestion step herein, the method can include mapping a first read pair to a set of contigs and determining a path through the set of contigs that represents the order and / or orientation relative to the genome.
[0237] In a further aspect of the method comprising the MNase digestion step of the present specification, the method can include mapping a first read pair to a set of contigs and determining the presence of a structural variant or loss of heterozygosity in the stabilized biological sample from the set of contigs.
[0238] In a further aspect of the method comprising the MNase digestion step herein, the method includes mapping a first read pair to a set of contigs and assigning phases to variants in the set of contigs.
[0239] In a further aspect of the method comprising the MNase digestion step of the present specification, the method includes mapping a first read pair to a set of contigs and performing a step selected from one or more of determining the presence of variants in the set of contigs and (1) identifying a disease stage, prognosis, or treatment course for the stabilized biological sample, (2) selecting a drug based on the presence of the variant, or (3) identifying drug efficacy for the stabilized biological sample from the set of contigs.
[0240] Improved methods for HiChIP, HiChIRP, and methylHiC HiChIP is an approach that combines the HiC method with the chromatin immunoprecipitation method, enabling targeted analysis of interactions involving one or more proteins of interest. Proximity-ligated nucleic acids can be prepared and immunoprecipitated for further analysis of the target region. The related approach, HiChIRP, uses chromatin isolation by RNA purification (ChIRP) enrichment in combination with the HiC method to enable investigation of RNAs such as the scaffolding function of long non-coding RNAs (lncRNAs). Methyl-HiC combines methylation analysis with the HiC method, enabling simultaneous capture of chromosomal conformation and DNA methylome information. Methyl-HiC reveals the coordinated DNA methylation status between distal genomic segments that are spatially proximal in the nucleus, depicts the heterogeneity of both chromatin structure and DNA methylome in a mixed population, and enables simultaneous characterization of cell-type-specific chromatin organization and epigenome in complex tissues. These and other methods can be improved by the use of the techniques of the present disclosure, including but not limited to size selection steps, surface binding steps (e.g., binding to beads such as SPRI beads), use of crosslinking oligonucleotides for performing proximity ligation, use of recombinases for performing proximity ligation, etc.
[0241] In a further aspect, for example, a step of obtaining a stabilized biological sample comprising a nucleic acid molecule complexed with at least one nucleic acid binding protein by immunoprecipitation of a nucleic acid bound to a nucleic acid binding protein or by immunoprecipitation of a methylated nucleic acid; a step of contacting the stabilized biological sample with DNase to cleave the nucleic acid molecule into a plurality of segments; a step of ligating a first segment and a second segment of the plurality of segments at a junction; and a step of subjecting the plurality of segments to size selection to obtain a plurality of selected segments can be included. An improved method for HiChIP, HiChIRP, and methylHiC is provided herein. Alternatively, or in combination, the methods herein include, for example, a step of obtaining a stabilized biological sample comprising a nucleic acid molecule complexed with at least one nucleic acid binding protein by immunoprecipitation of a nucleic acid bound to a nucleic acid binding protein or by immunoprecipitation of a methylated nucleic acid; a step of contacting the stabilized biological sample with micrococcal nuclease (MNase) to cleave the nucleic acid molecule into a plurality of segments; and a step of ligating a first segment and a second segment of the plurality of segments at a junction.
[0242] In some aspects of the improved methods for HiChIP, HiChIRP, and methylHiC herein, the stabilized biological sample can comprise intact cells and / or intact nuclei. Optionally, the stabilized biological sample can comprise stabilized intact cells. Alternatively, or in combination, the stabilized biological sample can comprise stabilized intact nuclei. Optionally, the step of contacting the stabilized intact cell or intact nucleus sample with DNase can be performed prior to lysis of the intact cell or intact nucleus. Optionally, the cells and / or nuclei can be lysed prior to ligating a first segment and a second segment of the plurality of segments at a junction.
[0243] In other aspects, methods including the improved methods of HiChIP, HiChIRP, and methylHiC of the present specification can include a step of subjecting a plurality of segments to size selection to obtain a plurality of selected segments. Optionally, the plurality of selected segments can be from about 145 to about 600 bp. Optionally, the plurality of selected segments can be from about 100 to about 2500 bp. Optionally, the plurality of selected segments can be from about 100 to about 600 bp. Optionally, the plurality of selected segments can be from about 600 to about 2500 bp. Optionally, the plurality of selected segments can be from about 100 bp to about 600 bp, from about 100 bp to about 700 bp, from about 100 bp to about 800 bp, from about 100 bp to about 900 bp, from about 100 bp to about 1000 bp, from about 100 bp to about 1100 bp, from about 100 bp to about 1200 bp, from about 100 bp to about 1300 bp, from about 100 bp to about 1400 bp, from about 100 bp to about 1500 bp, from about 100 bp to about 1600 bp, from about 100 bp to about 1700 bp, from about 100 bp to about 1800 bp, from about 100 bp to about 1900 bp, from about 100 bp to about 2000 bp, from about 100 bp to about 2100 bp, from about 100 bp to about 2200 bp, from about 100 bp to about 2300 bp, from about 100 bp to about 2400 bp, or from about 100 bp to about 2500 bp.
[0244] In another aspect of the methods including the improved methods of HiChIP, HiChIRP, and methylHiC of the present specification, the method may further include the step of preparing a sequencing library from a plurality of segments before the size selection step. In some embodiments, the method may further include the step of subjecting the sequencing library to size selection to obtain a size-selected library. Optionally, the size-selected library may be in the size range of about 350 bp to about 1000 bp. Optionally, the size-selected library may be in the size range of about 100 bp to about 2500 bp, such as about 100 bp to about 350 bp, about 350 bp to about 500 bp, about 500 bp to about 1000 bp, about 1000 bp to about 1500 bp, about 2000 bp to about 2500 bp, about 350 bp to about 1000 bp, about 350 bp to about 1500 bp, about 350 bp to about 2000 bp, about 350 bp to about 2500 bp, about 500 bp to about 1500 bp, about 500 bp to about 2000 bp, about 500 bp to about 2500 bp, about 1000 bp to about 1500 bp, about 1000 bp to about 2000 bp, about 1000 bp to about 2500 bp, about 1500 bp to about 2000 bp, about 1500 bp to about 2500 bp, or about 2000 bp to about 2500 bp.
[0245] The size selection utilized in the methods including the improved methods for HiChIP, HiChIRP, and methylHiC of the present specification can be performed using gel electrophoresis, capillary electrophoresis, size selection beads, gel filtration columns, combinations thereof, or any other suitable method.
[0246] In another aspect, methods including the improved methods for HiChIP, HiChIRP, and methylHiC herein may further include the step of further analyzing a plurality of selected segments to obtain QC values. In some cases, the QC values may be selected from chromatin digestion efficiency (CDE) and chromatin digestion index (GDI). CDE can be calculated as the percentage of segments having a desired length. For example, in some cases, CDE can be calculated as the percentage of segments sized 100 - 2500 bp before size selection. In some cases, a sample is selected for further analysis when the CDE value is at least 65%. In some cases, a sample may be selected for further analysis when the CDE value is at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, or at least about 95%.
[0247] CDI can be calculated as the ratio of the number of segments of mononucleosome size to the number of segments of dinucleosome size before size selection. For example, CDI can be calculated as the logarithm of the ratio of fragments having a size of 600 - 2500 bp to fragments having a size of 100 - 600 bp. In some cases, the sample is selected for further analysis when the CDI value is greater than -1.5 and less than 1. In some cases, the sample is selected for further analysis when the CDI value is greater than approximately -2 and less than approximately 1.5, greater than approximately -1.9 and less than approximately 1.5, greater than approximately -1.8 and less than approximately 1.5, greater than approximately -1.7 and less than approximately 1.5, greater than approximately -1.6 and less than approximately 1.5, greater than approximately -1.5 and less than approximately 1.5, greater than approximately -1.4 and less than approximately 1.5, greater than approximately -1.3 and less than approximately 1.5, greater than approximately -1.2 and less than approximately 1.5, greater than approximately -1.1 and less than approximately 1.5, greater than approximately -2 and less than approximately 1.5, greater than approximately -1 and less than approximately 1.5, greater than approximately -0.9 and less than approximately 1.5, greater than approximately -0.8 and less than approximately 1.5, greater than approximately -0.7 and less than approximately 1.5, greater than approximately -0.6 and less than approximately 1.5, greater than approximately -0.5 and less than approximately 1.5, greater than approximately -2 and less than approximately 1.4, greater than approximately -2 and less than approximately 1.3, greater than approximately -2 and less than approximately 1.2, greater than approximately -2 and less than approximately 1.1, greater than approximately -2 and less than approximately 1, greater than approximately -2 and less than approximately 0.9, greater than approximately -2 and less than approximately 0.8, greater than approximately -2 and less than approximately 0.7, greater than approximately -2 and less than approximately 0.6, or greater than approximately -2 and less than approximately 0.5.
[0248] In other aspects, methods including the improved methods for HiChIP, HiChIRP, and methylHiC herein can be performed on small samples containing a small number of cells or a small amount of nucleic acid. In some cases, the stabilized biological sample can contain fewer than 3,000,000 cells. In some cases, the stabilized biological sample can contain fewer than 2,000,000 cells. In some cases, the stabilized biological sample can contain fewer than 1,000,000 cells. In some cases, the stabilized biological sample can contain fewer than 500,000 cells. In some cases, the stabilized biological sample can contain fewer than 400,000 cells. In some cases, the stabilized biological sample can contain fewer than 300,000 cells. In some cases, the stabilized biological sample can contain fewer than 200,000 cells. In some cases, the stabilized biological sample can contain fewer than 100,000 cells. In some cases, the stabilized biological sample can contain fewer than 50,000 cells. In some cases, the stabilized biological sample can contain fewer than 40,000 cells. In some cases, the stabilized biological sample can contain fewer than 30,000 cells. In some cases, the stabilized biological sample can contain fewer than 20,000 cells. In some cases, the stabilized biological sample can contain fewer than 10,000 cells. In some cases, the stabilized biological sample can contain about 10,000 cells. In some cases, the stabilized biological sample can contain less than 10 μg of DNA. In some cases, the stabilized biological sample can contain less than 9 μg of DNA. In some cases, the stabilized biological sample can contain less than 8 μg of DNA. In some cases, the stabilized biological sample can contain less than 7 μg of DNA. In some cases, the stabilized biological sample can contain less than 6 μg of DNA. In some cases, the stabilized biological sample can contain less than 5 μg of DNA. In some cases, the stabilized biological sample can contain less than 4 μg of DNA. In some cases, the stabilized biological sample can contain less than 3 μg of DNA. In some cases, the stabilized biological sample can contain less than 2 μg of DNA. In some cases, the stabilized biological sample can contain less than 1 μg of DNA. In some cases, the stabilized biological sample can contain less than 0.5 μg of DNA.
[0249] In other embodiments, the methods, including the improved methods for HiChIP, HiChIRP, and methylHiC herein, may be performed on individual cells or single cells. For example, the methods herein can be performed on cells distributed into individual partitions. Exemplary partitions include, but are not limited to, wells, droplets in an emulsion, or surface locations (e.g., array spots, beads, etc.) that include discrete patches of differentially addressable linker molecules as described elsewhere herein. Additional partitions are contemplated and are consistent with the methods, compositions, and systems disclosed herein.
[0250] In further embodiments, the methods, including the improved methods for HiChIP, HiChIRP, and methylHiC herein, can be treated with a nuclease, such as DNase, to generate DNA fragments. In some cases, the DNase can be non-sequence specific. In some cases, the DNase can be active against both single-stranded DNA and double-stranded DNA. In some cases, the DNase can be specific for double-stranded DNA. In some cases, the DNase can preferentially cleave double-stranded DNA. In some cases, the DNase can be specific for single-stranded DNA. In some cases, the DNase can preferentially cleave single-stranded DNA. In some cases, the DNase can be DNaseI. In some cases, the DNase can be DNaseII. In some cases, the DNase can be selected from one or more of DNaseI and DNaseII. In some cases, the DNase can be micrococcal nuclease. In some cases, the DNase can be selected from one or more of DNaseI, DNaseII, and micrococcal nuclease. In some cases, the DNase can be bound or fused to an immunoglobulin-binding protein or fragment thereof, such as protein A, protein G, protein A / G, or protein L. Other suitable nucleases are within the scope of the present disclosure.
[0251] In a further aspect, methods including the improved methods for HiChIP, HiChIRP, and methylHiC herein can be treated with a crosslinking agent. In some cases, the crosslinking agent can be a chemical fixative. In some cases, the chemical fixative can include formaldehyde having a spacer arm length of about 2.3 to 2.7 angstroms (A). In some cases, the chemical fixative includes a crosslinking agent having a long spacer arm length. For example, the crosslinking agent can have a spacer length of at least about 3A, 4A, 5A, 6A, 7A, 8A, 9A, 10A, 11A, 12A, 13A, 14A, 15A, 16A, 17A, 18A, 19A, or 20A. The chemical fixative can include ethylene glycol bis(succinimidyl succinate) (EGS) having a spacer arm with a length of about 16.1A. The chemical fixative can include disuccinimidyl glutarate (DSG) having a spacer arm with a length of about 7.7A. In some cases, the chemical fixative includes formaldehyde and EGS, formaldehyde and DSG, or formaldehyde, EGS, and DSG. When multiple chemical fixatives are used, in some cases, each chemical fixative is used sequentially, and in other cases, some or all of the multiple chemical fixatives are applied to the sample simultaneously. The use of a crosslinking agent having a long spacer arm can increase the proportion of read pairs having a large (e.g., >1 kb) read pair separation distance. DSG is membrane permeable and enables intracellular crosslinking. DSG can enhance crosslinking efficiency compared to disuccinimidyl suberate (DSS) in some applications. EGS has NHS ester reactive groups at both ends and can be reactive towards amino groups (e.g., primary amines). EGS is membrane permeable and enables intracellular crosslinking. EGS crosslinking can be reversed, for example, by treatment with hydroxylamine at pH 8.5 for 3 to 6 hours. For example, lactate dehydrogenase retained 60% of its activity after reversible crosslinking with EGS. In some cases, the chemical fixative can include psoralen.In some cases, the crosslinking agent can be ultraviolet light, chloromethine, cyclophosphamide, chlorambucil, uracil mustard, melphalan, bendamustine, bis(2-chloroethyl)ethylamine, bis(2-chloroethyl)methylamine, tris(2-chloroethyl)amine, isofamide, carmustine, lomustine, streptozocin, busulfan, cisplatin, carboplatin, cicycloplatin, eptaplatin, lobaplatin, miltiplatin, nedaplatin, oxaliplatin, picoplatin, satraplatin, triplatin tetranitrate, procarbazine, altretamine, dacarbazine, mitozolomide, temozolomide, mitomycin C, nitrous acid, formaldehyde, acetylaldehyde, doxorubicin, daunorubicin, epirubicin, or idarubicin. In some cases, the crosslinking agent includes an intercalator, an antibiotic, or a minor groove binder. In some cases, the stabilized biological sample can be a crosslinked paraffin-embedded tissue sample.
[0252] In a further aspect, methods including the improved methods for HiChIP, HiChIRP, and methylHiC herein may include the step of joining a first segment and a second segment of a plurality of segments at a junction. In some cases, the joining step can include filling in sticky ends using biotin-tagged nucleotides and ligating blunt ends. In some cases, the joining step can include contacting at least the first segment and the second segment with a crosslinking oligonucleotide. In some cases, the joining step can include contacting at least the first segment and the second segment with a barcode. In some embodiments, the crosslinking oligonucleotides herein can be at least about 5 nucleotides in length to about 50 nucleotides in length. In some embodiments, the crosslinking oligonucleotides herein can be about 15 to about 18 nucleotides in length. In some embodiments, the crosslinking oligonucleotide can be about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, about 20, about 25, about 30, about 35, about 40, about 45, or about 50 nucleotides in length. In some embodiments, the crosslinking oligonucleotides herein can include a barcode.
[0253] In a further aspect, methods including the improved methods for HiChIP, HiChIRP, and methylHiC herein do not include a shearing step.
[0254] In a further aspect of methods including the improved methods for HiChIP, HiChIRP, and methylHiC herein, the method includes obtaining at least some sequences on each side of the junction to generate a first read pair. For example, the method can include obtaining sequences of at least about 50 bp, at least about 100 bp, at least about 150 bp, at least about 200 bp, at least about 250 bp, or at least about 300 bp on each side of the junction to generate a first read pair.
[0255] In a further aspect of the methods including the improved methods for HiChIP, HiChIRP, and methylHiC of the present specification, the method includes mapping a first read pair to a set of contigs and determining a path through the set of contigs that represents the order and / or orientation relative to the genome.
[0256] In a further aspect of the methods including the improved methods for HiChIP, HiChIRP, and methylHiC of the present specification, the method may include mapping a first read pair to a set of contigs and determining the presence of a structural variant or loss of heterozygosity in a stabilized biological sample from the set of contigs.
[0257] In a further aspect of the methods including whole cell or whole nucleus digestion of the present specification, the method may include mapping a first read pair to a set of contigs and assigning a phase to a variant in the set of contigs.
[0258] In a further aspect of the methods including the improved methods for HiChIP, HiChIRP, and methylHiC of the present specification, the method includes mapping a first read pair to a set of contigs, determining the presence of a variant in the set of contigs, and performing a step selected from one or more of (1) identifying a disease stage, prognosis, or course of treatment for a stabilized biological sample, (2) selecting a drug based on the presence of the variant, or (3) identifying the drug efficacy for a stabilized biological sample.
[0259] Generation of long-range read pairs The present disclosure provides methods for generating extremely long-range read pairs and using that data for all of the aforementioned advancements in tracking. In some embodiments, the present disclosure provides methods for generating a highly contiguous and accurate human genome assembly having only about 300 million read pairs. In other embodiments, the present disclosure provides methods for phasing over 90% of the heterozygous variants in the human genome with greater than 99% accuracy. Further, the range of read pairs generated by the present disclosure can be extended to span much larger genomic distances. The assembly is made from a standard shotgun library in addition to an extremely long-range read pair library. In still other embodiments, the present disclosure provides software that can utilize both of these sets of sequencing data. The phased variants are generated using a single long-range read pair library, and the reads from there are mapped to a reference genome and then used to assign variants to one of the two parental chromosomes of an individual. Finally, the present disclosure results in the extraction of even larger DNA fragments using known techniques to generate exceptionally long reads.
[0260] The mechanisms by which these repeats interfere with the assembly and alignment processes are fairly straightforward and ultimately result in ambiguity. In the case of large repeat regions, the problem can be one of span. If a read or read pair is not long enough to span the repeat region, it may not be possible to join the regions flanking the repeat element in a reliable manner. In the case of small repeat elements, the problem can primarily be one of placement. If a region is flanked by two repeat elements that are common in the genome, determining its exact placement becomes difficult, if not impossible, because the flanking elements are similar to all of the others in their class. In both cases, what makes identification, and thus placement of a particular repeat sequence, difficult is the lack of differentiating information within the repeat sequence. What is needed is the ability to experimentally establish the connections between unique segments surrounded or separated by repeat regions.
[0261] The methods of the present disclosure can advance the field of genomics by overcoming the substantial barriers posed by these repetitive regions, thereby enabling significant progress in many areas of genomic analysis. To perform de novo assembly using prior art, one must either prepare assemblies fragmented into many small scaffolds or use other approaches for generating large insertion libraries or more contiguous assemblies, which requires substantial time and resources. Such approaches can include obtaining very deep sequencing coverage, constructing BAC or fosmid libraries, optical mapping, or some combination of these and / or other techniques. Due to the severe requirements for resources and time, such approaches have not penetrated most small-scale laboratories and have hindered research on non-model organisms. The methods described herein can generate very long-range read pairs, so that de novo assembly can be achieved with a single sequencing run. This reduces the assembly cost by orders of magnitude and shortens the time required from months or years to weeks. In some cases, the methods disclosed herein can generate multiple read pairs in less than 14 days, less than 13 days, less than 12 days, less than 11 days, less than 10 days, less than 9 days, less than 8 days, less than 7 days, less than 6 days, less than 5 days, less than 4 days, or within the range between any two of the aforementioned specified periods. For example, the method can enable generating multiple read pairs in about 10 to 14 days. Constructing genomes can become routine even for most niches of organisms, phylogenetic analysis can be free from the lack of comparison, and projects such as Genome 10k can be realized.
[0262] Similarly, there remain challenges in structural and phasing analysis for medical purposes. There is surprising heterogeneity among cancers, among individuals with the same type of cancer, or even within the same tumor. Extracting the causative from the resulting effects requires very high accuracy and throughput at low cost per sample. In the area of personalized medicine, one of the absolute criteria for genomic care is a sequenced genome with all variants thoroughly characterized and phased, including large and small structural rearrangements and novel mutations. Achieving this with previous technologies requires a similar effort to that required for de novo assembly, which is now too expensive and requires conventional medical procedures. The disclosed method can rapidly produce a low-cost, complete, and accurate genome, thereby providing many highly sought-after capabilities in the study and treatment of human diseases.
[0263] Applying the methods disclosed herein to phasing can combine the convenience of statistical approaches with the accuracy of family analysis, resulting in greater savings - in cost, labor, and samples - than using either method alone. Novel variant phasing analysis, a highly desirable phasing analysis that was prohibitive with previous technologies, can be readily performed using the methods disclosed herein. This is particularly important because the majority of human variants are rare (minor allele frequency of less than 5%). Phasing information is valuable for population genetic studies that gain significant advantages from highly linked haplotype networks (sets of variants assigned to a single chromosome) compared to unlinked genotypes. Haplotype information can enable higher-resolution studies of the history of changes in population size, migration, and exchange among subpopulations and can enable tracking specific variants back to specific parents and grandparents. This, in turn, reveals the genetic transmission of disease-related variants and the interactions among variants when grouped in a single individual. The methods of the present disclosure can ultimately enable the preparation, sequencing, and analysis of extremely long-range read pair (XLRP) libraries.
[0264] In some embodiments of the present disclosure, a tissue or DNA sample from a subject may be provided, and the method can return an assembled genome, an alignment with called variants (including large structural variants), a phased variant call, or any additional analysis. In other embodiments, the methods disclosed herein can directly provide an XLRP library for an individual.
[0265] Very long-range read pairs In various embodiments of the present disclosure, the methods disclosed herein can generate extremely long-range read pairs that are separated by a far distance. The upper limit of this distance can be improved by the ability to collect large-sized DNA samples. In some cases, the read pairs can span genomic distances of up to 50, 60, 70, 80, 90, 100, 125, 150, 175, 200, 225, 250, 300, 400, 500, 600, 700, 800, 900, 1000, 1500, 2000, 2500, 3000, 4000, 5000 kbp, or more. In some examples, the read pairs can span a genomic distance of up to 500 kbp. In other examples, the read pairs can span a genomic distance of up to 2000 kbp. The methods disclosed herein can be incorporated and constructed based on standard techniques in molecular biology and are well-suited for increasing efficiency, specificity, and genomic coverage. In some cases, the read pairs can be generated in less than about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 60, or 90 days. In some examples, the read pairs can be generated in less than about 14 days. In further examples, the read pairs can be generated in less than about 10 days. In some cases, the methods of the present disclosure can provide at least about 50%, about 60%, about 70%, about 80%, about 90%, about 95%, about 99%, or about 100% accuracy in the correct ordering and / or orientation of multiple contigs for about 5%, about 10%, about 15%, about 20%, about 30%, about 40%, about 50%, about 60%, about 70%, about 80%, about 90%, about 95%, about 99%, or more than about 100% of the read pairs. For example, the method can provide an accuracy of about 90 - 100% in the correct ordering and / or orientation of multiple contigs.
[0266] In other embodiments, the methods disclosed herein can be used with currently utilized sequencing technologies. For example, the methods can be used in combination with well-tested and / or widely deployed sequencing equipment. In further embodiments, the methods disclosed herein can be used with techniques and approaches derived from currently utilized sequencing technologies.
[0267] The methods of the present disclosure can dramatically simplify de novo genome assembly for a wide variety of organisms. Using previous techniques, such assemblies are currently limited by short inserts in economical mate pair libraries. It may be possible to generate read pairs at genomic distances up to 40-50 kbp accessible with fosmids, but these are expensive, difficult to handle, and - in humans - too short to span the longest repeat stretches, which can range in size from -300 kbp to 5 Mbp and include those within centromeres. The methods disclosed herein provide read pairs that can span long distances (e.g., megabases and above), thereby overcoming these scaffold integrity challenges. Thus, the production of chromosome-level assemblies can become routine by utilizing the methods of the present disclosure. More labor-intensive means for assembly - which currently require an unthinkable amount of time and money in research laboratories and impede broad genomic cataloging - may become unnecessary, freeing up resources for more meaningful analyses. Similarly, the acquisition of long-range phasing information can provide additional immense power for population genomic studies, phylogenetic studies, and disease studies. The methods disclosed herein enable accurate phasing for multiple individuals, thus expanding the breadth and depth of our ability to probe genomes at the population and deep time levels.
[0268] In the field of personalized medicine, the XLRP read pairs generated from the methods disclosed herein represent a significant advancement to accurate, low-cost, phased, and rapidly produced individual genomes. Current methods are insufficient in their ability to phase variants over long distances, thereby preventing the characterization of the phenotypic impact of compound heterozygous genotypes. Further, structural variants of substantial interest in genomic diseases are large in size compared to the reads and read pair inserts used to study them, making them difficult to accurately identify and characterize using current techniques. Read pairs spanning tens of kilobases to megabases and above can help alleviate this difficulty, thereby enabling highly parallel and individualized analysis of structural polymorphisms.
[0269] Basic evolutionary and biomedical research has been propelled by technological advances in high-throughput sequencing. Whole genome sequencing and assembly are used as the source of large genome sequencing centers, but commercial sequencers are now inexpensive enough that most research universities have one or more of these machines. Currently, it is relatively inexpensive to generate large amounts of DNA sequence data. However, it remains difficult both theoretically and practically to produce high-quality and highly contiguous genomic sequences using current techniques. Further, since most organisms, including humans, that one would aim to analyze are diploid, each individual has two haploid copies of the genome. At heterozygous sites (e.g., where the allele given by the mother is different from the allele given by the father), it is difficult to know which set of alleles came from which parent (known as haplotype phasing). This information can be used to conduct some evolutionary and biomedical research such as studies of disease and trait associations.
[0270] In various embodiments, the present disclosure provides methods for genome assembly that combine techniques for DNA preparation with paired-end sequencing for high-throughput discovery of short, medium, and long junctions within a given genome. The present disclosure further provides methods of using these junctions to assist in genome assembly, for haplotype phasing, and / or for metagenomic studies. The methods presented herein can be used to determine the assembly of a target genome, but it should also be understood that the methods presented herein can be used to determine the assembly of portions of a target genome, such as chromosomes, or the assembly of target chromatin of various lengths.
[0271] In some embodiments, the present disclosure provides one or more methods disclosed herein that include generating a plurality of contigs from sequencing fragments of target DNA obtained from a subject. A long stretch of target DNA can be fragmented by cleaving the DNA with one or more nucleases (e.g., DNaseI, DNaseII, Micrococcus nuclease, etc.). To obtain a plurality of sequencing reads, the resulting fragments can be sequenced using a high-throughput sequencing method. Examples of high-throughput sequencing methods that can be used with the methods of the present disclosure include, but are not limited to, the 454 pyrosequencing method developed by Roche Diagnostics, the "cluster" sequencing method developed by Illumina, the SOLiD and Ion semiconductor sequencing methods developed by Life Technologies, and the DNA nanoball sequencing method developed by Complete Genomics. Next, overlapping ends of different sequencing reads can be assembled to form contigs. Alternatively, the fragmented target DNA can be cloned into a vector. A DNA vector is then transfected into a cell or organism to form a library. After replicating the transfected cell or organism, the vector is isolated and sequenced to generate a plurality of sequencing reads. Thereafter, overlapping ends of different sequencing reads can be assembled to form contigs.
[0272] Genome assembly, particularly by high-throughput sequencing technologies, can be problematic. Often, the assembly consists of thousands or tens of thousands of short contigs. The order and orientation of these contigs are generally unknown, limiting the usefulness of the genome assembly. There are techniques for ordering and orienting these scaffolds, but they are generally expensive, require a lot of manual labor, and often fail to discover very long-range interactions.
[0273] A sample containing the target DNA used to generate contigs can be obtained from a subject by any number of means, including collecting a body fluid (e.g., blood, urine, serum, lymph, saliva, oral swabs, anal and vaginal secretions, sweat, and semen), collecting a tissue, or collecting cells / organisms. The resulting sample may be composed of a single type of cell / organism or may be composed of multiple types of cells / organisms. The DNA can be extracted and prepared from the subject's sample. For example, the sample can be treated to lyse cells containing polynucleotides using known lysis buffers, sonication techniques, electroporation, etc. The target DNA can be further purified to remove contaminants such as proteins by using alcohol extraction, cesium gradients, and / or column chromatography.
[0274] In other embodiments of the present disclosure, methods for extracting very high molecular weight DNA are provided. In some cases, data from an XLRP library can be improved by increasing the fragment size of the input DNA. In some examples, extraction of megabase-sized fragments of DNA from cells can generate read pairs separated by megabases in the genome. In some cases, the generated read pairs can provide sequence information spanning a span of about 10 kB, about 50 kB, about 100 kB, about 200 kB, about 500 kB, about 1 Mb, about 2 Mb, about 5 Mb, about 10 Mb, or greater than about 100 Mb. In some examples, the read pairs can provide sequence information spanning a span of greater than about 500 kB. In further examples, the read pairs can provide sequence information spanning a span of greater than about 2 Mb. In some cases, very high molecular weight DNA can be extracted by very gentle cell lysis (Teague, B. et al. (2010) Proc. Nat. Acad. Sci. USA 107(24), 10848-53) and agarose plugs (Schwartz, D. C., & Cantor, C. R. (1984) Cell, 37(1), 67-75). In other cases, very high molecular weight DNA can be extracted using commercially available machines that can purify DNA molecules up to megabase in length.
[0275] Probing the physical layout of chromosomes In various embodiments, the present disclosure provides one or more methods disclosed herein that include the step of probing the physical layout of chromosomes within a living cell. Examples of techniques for investigating the physical layout of chromosomes through sequencing include techniques of the "C" family such as chromosome conformation capture ("3C"), circular chromosome conformation capture ("4C"), carbon copy chromosome capture ("5C"), and Hi-C based methods, as well as ChIP-based methods such as ChIP-loop, ChIP-PET, and HiChIP. These techniques utilize the fixation of chromatin in living cells to solidify the spatial relationships within the nucleus. Subsequent processing and sequencing of the products enable the researcher to recover a matrix of proximity relationships between genomic regions. Further analysis can use these relationships to generate a three-dimensional geometric map of the chromosomes when they are physically arranged in the living nucleus. Such techniques explain the discrete spatial organization of chromosomes in living cells and provide an accurate view of the functional interactions between chromosomal loci. One problem that has hampered these functional studies has been the presence of non-specific interactions, the associations present in the data that do not arise from more than chromosomal proximity. In the present disclosure, these non-specific intra-chromosomal interactions are captured by the methods presented herein in order to provide valuable information for assembly.
[0276] In some embodiments, intrachromosomal interactions are associated with chromosomal connectivity. In some cases, intrachromosomal data can assist in genome assembly. In some cases, chromatin is reconstituted in vitro. This can be advantageous because chromatin, specifically the histone which is the major protein component of chromatin, is important for immobilization under the most common "C" family of techniques for detecting the conformation and structure of chromatin through sequencing: 3C, 4C, 5C, and Hi-C. Chromatin is very non-specific with respect to sequence and typically assembles uniformly across the genome. In some cases, genomes of species that do not use chromatin can be assembled on reconstituted chromatin, thereby extending the horizons of the present disclosure to all domains of life.
[0277] Summarize the chromatin three-dimensional structure capture technology. Briefly, cross-links are formed between genomic regions that are physically close. Cross-linking of DNA molecules in chromatin, such as proteins (such as histones) to genomic DNA, can be achieved according to appropriate methods that are described in more detail elsewhere in this specification or are otherwise known. In some cases, two or more nucleotide sequences can be cross-linked via proteins bound to one or more nucleotide sequences. One approach is to expose chromatin to ultraviolet irradiation (Gilmour et al., Proc. Nat’l. Acad. Sci. USA 81:4275-4279, 1984). Cross-linking of polynucleotide segments can also be carried out using other approaches such as chemical or physical (e.g., optical) cross-linking. Suitable chemical cross-linking agents include, but are not limited to, formaldehyde and psoralen (Solomon et al., Proc. Nat’l. Acad.Sci.USA 82:6470-6474, 1985; Solomon et al., Cell 53:937-947, 1988). For example, cross-linking can be performed by adding 2% formaldehyde to a mixture containing DNA molecules and chromatin proteins. Other examples of agents that can be used to cross-link DNA include, but are not limited to, UV light, mitomycin C, nitrogen mustard, melphalan, 1,3-butadiene diepoxide, cis-diamminedichloroplatinum(II), and cyclophosphamide. Appropriately, the cross-linking agent forms cross-links that cross-link relatively short distances (e.g., about 2 Å), thereby selecting for intimate interactions that can be reversed.
[0278] In some embodiments, the DNA molecule can be immunoprecipitated before or after crosslinking. Optionally, the DNA molecule can be fragmented. The fragments can be contacted with a binding partner such as an antibody that specifically recognizes and binds to acetylated histone, e.g., H3. Examples of such antibodies include, but are not limited to, anti-acetylated histone H3 available from Upstate Biotechnology, Lake Placid, NY. The polynucleotide from the immunoprecipitate can then be collected from the immunoprecipitate. Before fragmenting the chromatin, the acetylated histone can be crosslinked to adjacent polynucleotide sequences. The mixture is then processed to fractionate the polynucleotides in the mixture. Fractionation techniques herein include the use of deoxyribonuclease (DNase) enzymes. DNases suitable for the methods herein include, but are not limited to, DNaseI, DNaseII, and Micrococcus nuclease. The resulting fragments can be of different sizes. The resulting fragments can further include single-stranded overhangs at the 5' or 3' ends.
[0279] In some embodiments, fragments of about 145 bp to about 600 bp can be obtained. Alternatively, fragments of about 100 bp to about 2500 bp, about 100 bp to about 600 bp, or about 600 to about 2500 can be obtained. The sample can be prepared for sequencing of the crosslinked binding sequence segments. Optionally, a single short stretch of polynucleotide can be created, e.g., by ligating two sequence segments crosslinked within the molecule. Sequence information can be obtained from the sample using any suitable sequencing technique described in more detail elsewhere herein, or other suitable methods such as high-throughput sequencing methods. For example, the ligation product can be subjected to paired-end sequencing to obtain sequence information from each end of the fragment. Pairs of sequence segments can be represented in the resulting sequence information by associating haplotyping information over the linear distance separating the two sequence segments along the polynucleotide.
[0280] One feature of the data generated by Hi-C is that most read pairs are found to be linearly proximal when mapped back to the genome. That is, most read pairs are found to be proximal to each other in the genome. In the resulting dataset, the probability of intrachromosomal contacts is much higher on average than the probability of interchromosomal contacts, as expected if chromosomes occupy separate regions. Furthermore, the probability of interaction decays rapidly with linear distance, but loci even >200 Mb apart on the same chromosome are more likely to interact than loci on different chromosomes. This "background" of short- and medium-range intrachromosomal contacts is background noise to be removed using Hi-C analysis in the detection of long-range intrachromosomal and especially interchromosomal contacts.
[0281] Specifically, Hi-C experiments in eukaryotes have shown two canonical interaction patterns in addition to species- and cell type-specific chromatin interactions. One pattern, distance-dependent decay (DDD), is the general trend of the decay of interaction frequency according to genomic distance. The second pattern, the cis-trans ratio (CTR), is the interaction frequency between loci located on the same chromosome, which is significantly higher compared to loci on different chromosomes, even when separated by sequences of tens of megabases. These patterns may reflect general polymer dynamics in which proximal loci have a higher probability of interacting randomly, as well as specific nuclear organization features such as the formation of chromosomal regions and the phenomenon of interphase chromosomes that tend to occupy separate volumes in the nucleus with little mixing. The exact details of these two patterns can vary between species, cell types, and cell conditions, but they are ubiquitous and prominent. Because these patterns are very strong and consistent, they are used to evaluate the quality of experiments and are usually normalized from the data to reveal detailed interactions. However, in the methods disclosed herein, the genome assembly can utilize the three-dimensional structure of the genome. The features that make the canonical Hi-C interaction patterns an obstacle to the analysis of specific loop interactions, i.e., their ubiquity, strength, and consistency, can be used as powerful tools for inferring the genomic positions of contigs.
[0282] In certain embodiments, investigation of the physical distance between intrachromosomal read pairs reveals several useful features of data regarding genome assembly. First, shorter distance interactions are more common than longer distance interactions. That is, each read of a read pair is more likely to associate with regions closer in the actual genome than with regions far apart. Second, there are long tails of intermediate and long distance interactions. That is, read pairs retain information regarding intrachromosomal placement at distances of kilobases (kb) or even megabases (Mb). For example, read pairs can provide sequence information spanning spans of about 10 kb, about 50 kb, about 100 kb, about 200 kb, about 500 kb, about 1 Mb, about 2 Mb, about 5 Mb, about 10 Mb, or over about 100 Mb. These features of the data simply indicate that regions of the genome that are close on the same chromosome are likely to be physically closer, which is an expected result since they are chemically bonded to each other via the DNA backbone. Genome-wide chromatin interaction data sets such as those generated by Hi-C were hypothesized to provide long-range information regarding the grouping and linear organization of sequences along entire chromosomes.
[0283] The Hi-C experimental method is simple and relatively low-cost, but current protocols for genome assembly and haplotyping require a relatively large amount of material, particularly 3 to 5 million cells from specific human patient samples, i.e., an amount that may not be feasible to obtain. In contrast, the methods disclosed herein include methods that enable accurate and predictive results for genotype assembly, haplotype phasing, and metagenomics using significantly less cell-derived material. For example, DNA of less than about 0.1 μg, about 0.2 μg, about 0.3 μg, about 0.4 μg, about 0.5 μg, about 0.6 μg, about 0.7 μg, about 0.8 μg, about 0.9 μg, about 1.0 μg, about 1.2 μg, about 1.4 μg, about 1.6 μg, about 1.8 μg, about 2.0 μg, about 2.5 μg, about 3.0 μg, about 3.5 μg, about 4.0 μg, about 4.5 μg, about 5.0 μg, about 6.0 μg, about 7.0 μg, about 8.0 μg, about 9.0 μg, about 10 μg, about 15 μg, about 20 μg, about 30 μg, about 40 μg, about 50 μg, about 60 μg, about 70 μg, about 80 μg, about 90 μg, about 100 μg, about 150 μg, about 200 μg, about 300 μg, about 400 μg, about 500 μg, about 600 μg, about 700 μg, about 800 μg, about 900 μg, about 1000 μg, about 1200 μg, about 1400 μg, about 1600 μg, about 1800 μg, about 2000 μg, about 2200 μg, about 2400 μg, about 2600 μg, about 2800 μg, about 3000 μg, about 3200 μg, about 3400 μg, about 3600 μg, about 3800 μg, about 4000 μg, about 4200 μg, about 4400 μg, about 4600 μg, about 4800 μg, about 5000 μg, about 5200 μg, about 5400 μg, about 5600 μg, about 5800 μg, about 6000 μg, about 6200 μg, about 6400 μg, about 6600 μg, about 6800 μg, about 7000 μg, about 7200 μg, about 7400 μg, about 7600 μg, about 7800 μg, about 8000 μg, about 8200 μg, about 8400 μg, about 8600 μg, about 8800 μg, about 9000 μg, about 9200 μg, about 9400 μg, about 9600 μg, about 9800 μg, or about 10,000 μg can be used with the methods disclosed herein.In some examples, the DNA used in the methods disclosed herein can be extracted from fewer than about 3,000,000, about 2,500,000, about 2,000,000, about 1,500,000, about 1,000,000, about 500,000, about 100,000, about 50,000, about 10,000, about 5,000, about 1,000, about 500, or about 100 cells.
[0284] Generally, procedures for probing the physical layout of chromosomes, such as Hi-C based techniques, utilize chromatin formed within a cell / organism, such as chromatin isolated from cultured cells or primary tissue. The present disclosure provides the use of such techniques with reconstituted chromatin, in addition to chromatin isolated from a cell / organism. Reconstituted chromatin is distinguishable from chromatin formed within a cell / organism across various characteristics. First, for many samples, collection of a naked DNA sample can be achieved by using a variety of methods ranging from non-invasive methods, such as collecting body fluids, swabbing the cheek or rectal area, obtaining epithelial samples, to invasive methods. Second, reconstitution of chromatin substantially prevents the formation of interchromosomal and other long-distance interactions that generate artifacts for genome assembly and haplotype phasing. In some cases, the sample may have less than about 20, 15, 12, 11, 10, 9, 8, 7, 6, 5, 4, 3, 2, 1, 0.5, 0.4, 0.3, 0.2, 0.1% or less of the interchromosomal or intermolecular cross-links according to the methods and compositions of the present disclosure. In some examples, the sample may have less than about 5% interchromosomal or intermolecular cross-links. In some examples, the sample may have less than about 3% interchromosomal or intermolecular cross-links. In further examples, the sample may have less than about 1% interchromosomal or intermolecular cross-links. Third, the frequency of sites capable of cross-linking, and thus the frequency of intramolecular cross-links within a polynucleotide, can be adjusted. For example, the DNA to histone ratio may vary, and by doing so, the nucleosome density can be adjusted to a desired value. In some cases, the nucleosome density is reduced below physiological levels. Thus, the distribution of cross-links can be altered to support more long-distance interactions. In some embodiments, sub-samples with varying cross-link densities can be prepared to cover both short-distance and long-distance associations.For example, the cross-linking conditions can be adjusted such that at least about 1%, about 2%, about 3%, about 4%, about 5%, about 6%, about 7%, about 8%, about 9%, about 10%, about 11%, about 12%, about 13%, about 14%, about 15%, about 16%, about 17%, about 18%, about 19%, about 20%, about 25%, about 30%, about 40%, about 45%, about 50%, about 60%, about 70%, about 80%, about 90%, about 95%, or about 100% of the cross-linking occurs between DNA segments that are at least about 50 kb, about 60 kb, about 70 kb, about 80 kb, about 90 kb, about 100 kb, about 110 kb, about 120 kb, about 130 kb, about 140 kb, about 150 kb, about 160 kb, about 180 kb, about 200 kb, about 250 kb, about 300 kb, about 350 kb, about 400 kb, about 450 kb, or about 500 kb apart on the sample DNA molecule.
[0285] Contact Mapping and Topology Using the read pairs generated by the methods of the present disclosure, the three-dimensional structure of the genome and the chromosomes and nucleic acid molecules therein can be analyzed. As discussed herein, each read in a read pair can be mapped to different regions in the genome. For a given read pair, it can be inferred that the two different regions in the genome to which they map are spatially proximate to each other such that they can be ligated together. By plotting the read pairs from the sample according to the coordinates of both reads in the read pair, a contact map for the sample can be generated.
[0286] Analysis of contacts across the sample can enable analysis of chromosomal and genomic structure. Organization into genomic A and B compartments, active and inactive compartments, chromosomal compartments, euchromatin and heterochromatin, topologically associating domains (TADs) including TAD subtypes, and other structures can be analyzed at scales on the order of kilobase or megabase scale. Analysis of the contact map can enable detection of genomic features such as structural variants such as rearrangements, translocations, copy number polymorphisms, inversions, deletions, and insertions.
[0287] The methods of the present disclosure can provide the positions of protein binding, structural changes, or genomic contact interactions at a resolution of about 1 bp, 2 bp, 3 bp, 4 bp, 5 bp, 6 bp, 7 bp, 8 bp, 9 bp, 10 bp, 20 bp, 30 bp, 40 bp, 50 bp, 60 bp, 70 bp, 80 bp, 90 bp, 100 bp, 200 bp, 300 bp, 400 bp, 500 bp, 600 bp, 700 bp, 800 bp, 900 bp, 1000 bp, 2000 bp, 3000 bp, 4000 bp, 5000 bp, 6000 bp, 7000 bp, 8000 bp, 9000 bp, 10 kb, 20 kb, 30 kb, 40 kb, 50 kb, 60 kb, 70 kb, 80 kb, 90 kb, or 100 kb or less. In some cases, protein binding sites, protein footprints, contact interactions, or other features can be mapped within 1000 bp, 900 bp, 800 bp, 700 bp, 600 bp, 500 bp, 400 bp, 300 bp, 200 bp, 190 bp, 180 bp, 170 bp, 160 bp, 150 bp, 140 bp, 130 bp, 120 bp, 110 bp, 100 bp, 90 bp, 80 bp, 70 bp, 60 bp, 50 bp, 40 bp, 30 bp, 20 bp, 10 bp, 9 bp, 8 bp, 7 bp, 6 bp, 5 bp, 4 bp, 3 bp, 2 bp, or 1 bp. In one example, the methods of the present disclosure can enable the resolution of sites (e.g., protein binding sites such as CTCF sites) that are within 10,000 bp, 5,000 bp, 2,000 bp, or 1,000 bp of each other on the genome. In some cases, improved resolution or mapping can be achieved by using MNase or other endonucleases that digest unprotected nucleic acids (e.g., nucleic acids that are not within the footprint of the binding protein), thereby resulting in proximity ligation events that occur at the edges of the protected regions (e.g., protein footprints).
[0288] Contig mapping In various embodiments, the present disclosure provides various ways to enable mapping of a plurality of read pairs to a plurality of contigs. There are several publicly available computer programs for mapping reads to a contig array. This read mapping program data further provides data describing how specific read mappings are unique within the genome. From the population of uniquely mapped reads, it is possible to infer the distribution of distances between reads in each read pair with high confidence within a contig. For read pairs where reads are mapped to different contigs in a reliable manner, this mapping data implies a connection between the two contigs in question. This means a distance between two contigs that is proportional to the distribution of distances learned from the above analysis. Thus, each read pair where the reads are mapped to different contigs implies a connection between these two contigs in the correct assembly. The connections inferred from all such mapped read pairs can be summarized in an adjacency matrix where each contig is represented by both rows and columns. Read pairs that connect contigs are marked as non-zero values in the corresponding rows and columns, indicating the contigs to which the reads in the read pair are mapped. Most read pairs are mapped within a contig, from which the distribution of distances between read pairs can be learned, and from which an adjacency matrix of contigs can be constructed using read pairs that are mapped to different contigs.
[0289] In various embodiments, the present disclosure provides a method that includes constructing an adjacency matrix of contigs using read mapping data from read pair data. In some embodiments, the adjacency matrix uses a weighting scheme for read pairs that incorporates the tendency of short-range interactions relative to long-range interactions. Read pairs spanning short distances are generally more common than read pairs spanning long distances. A function describing the probability of a particular distance can be fit using read pair data mapped to a single contig to learn this distribution. Thus, one important feature of read pairs mapping to different contigs is their position on the contig to which they map. In the case of read pairs both mapping near one end of a contig, the inferred distance between these contigs can be short, and thus the distance between the joined reads is small. Since shorter distances between read pairs are more common than long distances, this configuration provides stronger evidence that these two contigs are adjacent relative to reads mapping far from the ends of the contig. Thus, the connections in the adjacency matrix are further weighted by the distance of the reads to the ends of the contig. In further embodiments, the adjacency matrix can be further rescaled to downweight the weights of a large number of contact points on some contigs representing uninteresting regions of the genome. These regions of the genome can be identified by having a high proportion of read mapping to them and are likely to include spurious read mappings that could misidentify the assembly a priori. In yet further embodiments, this scaling can be directed by searching for one or more conserved binding sites for one or more agents that regulate chromatin scaffold interactions, such as the transcription repressor CTCF, nuclear receptors, cohesin, or covalently modified histones.
[0290] In some embodiments, the present disclosure provides one or more methods disclosed herein that include analyzing an adjacency matrix to determine a path through contigs that represent an order and / or orientation with respect to a genome. In other embodiments, the path through the contigs can be selected to pass through each contig exactly once. In further embodiments, the path through the contigs is selected to maximize the sum of the weights of the edges through which the path through the adjacency matrix passes. In this way, the most likely contig connections for a correct assembly are proposed. In still further embodiments, the path through the contigs can be selected to pass through each contig exactly once and to maximize the edge weighting of the adjacency matrix.
[0291] Haplotype phasing In a diploid genome, it is often important to know which allelic variants are linked on the same chromosome. This is known as haplotype phasing. Short reads from high-throughput sequence data rarely allow direct observation of which allelic variants are linked. Computerized inference of haplotype phasing can be uncertain at long distances. The present disclosure provides one or more methods that enable determination of which allelic variants are linked using allelic variants on read pairs. In some cases, phasing by the methods of the present disclosure is performed without complementation.
[0292] In various embodiments, the methods and compositions of the present disclosure enable haplotype phasing of diploid or polyploid genomes with respect to multiple allelic variants. Accordingly, the methods described herein can provide for the determination of associated allelic variants based on variant information from read pairs and / or assembled contigs that use it. Examples of allelic variants include, but are not limited to, those known from the 1000 Genomes, UK10K, HapMap, and other projects to discover genetic variation among humans. In some cases, for example, the discovery of unlinked inactivating mutations in both copies of SH3TC2 that cause Charcot-Marie-Tooth neuropathy (Lupski JR, Reid JG, Gonzaga-Jauregui C, et al. N. Engl. J. Med. 362:1181-91, 2010), and unlinked inactivating mutations in both copies of ABCG5 that cause hypercholesterolemia 9 (Rios J, Stein E, Shendure J, et al. Hum. Mol. Genet. 19:4313-18, 2010) demonstrate that having haplotype phasing data can more readily reveal the association of a disease to a particular gene.
[0293] Humans are heterozygous at an average of one site per 1,000. In some cases, a single lane of data using high-throughput sequencing can generate at least about 150,000,000 read pairs. The read pairs can be about 100 base pairs in length. From these parameters, it is estimated that one-tenth of the total reads from a human sample cover heterozygous sites. Thus, on average one-hundredth of the total read pairs from a human sample are estimated to cover pairs of heterozygous sites. Thus, about 1,500,000 read pairs (one percent of 150,000,000) provide phasing data using a single lane. With about 3 billion bases in the human genome and one out of 1,000 being heterozygous, there are about 3 million heterozygous sites in the average human genome. With about 1,500,000 read pairs representing pairs of heterozygous sites, the average coverage of each heterozygous site phased using a single lane of high-throughput sequencing is about (1X) using a typical high-throughput sequencer. Thus, a diploid human genome can be reliably and completely phased with one lane of high-throughput sequence data related to sequence variants from samples prepared using the methods disclosed herein. In some examples, a lane of data can be a set of DNA sequence read data. In further examples, a lane of data can be a set of DNA sequence read data from one run of a high-throughput sequencer.
[0294] Since the human genome consists of two homologous sets of chromosomes, understanding an individual's true genetic makeup requires characterization of the maternal and paternal copies or haplotypes of the genetic material. Obtaining haplotypes in an individual is useful in several ways. First, haplotypes are clinically useful in predicting the outcome of donor-host compatibility in organ transplantation and are increasingly used as a means of detecting disease associations. Second, in genes that exhibit compound heterozygosity, haplotypes provide information regarding whether two deleterious variants are located on the same allele and greatly influence prediction of whether the inheritance of these variants is harmful. Third, haplotypes from populations provide information regarding the population structure and evolutionary history of a race. Finally, extensive allelic imbalance in recently described gene expression suggests that genetic or epigenetic differences between alleles may contribute to quantitative differences in expression. Understanding haplotype structure will describe the mechanisms of variants that contribute to allelic imbalance.
[0295] In certain embodiments, the methods disclosed herein include in vitro techniques for fixing and capturing associations between distal regions of the genome required for long-range ligation and phasing. Optionally, the method includes constructing and sequencing an XLRP library to deliver very genomically distant read pairs. Optionally, the interactions arise primarily from random associations within a single DNA fragment. In some examples, sequence segments that are proximal to each other in a DNA molecule interact more frequently and with higher probability, while interactions between distant parts of the molecule are less frequent, allowing the genomic distance between segments to be inferred. Thus, there is a systematic relationship between the number of pairs joining two loci and their proximity on the input DNA. The present disclosure can generate read pairs that can span the largest DNA fragments in the extraction. The input DNA for this library has a maximum length of 150 kbp, which is the longest meaningful read pair observed from the sequencing data. This suggests that the method can still join more genomically distant loci when larger input DNA fragments are provided. A complete genome assembly may be possible by applying improved assembly software tools specifically adapted to handle the types of data generated by the method.
[0296] Data generated using the methods and compositions of the present disclosure can achieve extremely high phasing accuracy. Compared to previous methods, the methods described herein can phase a higher percentage of variants. Phasing can be achieved while maintaining a high level of accuracy. The techniques herein can enable phasing with accuracies exceeding about 70%, 80%, 90%, 95%, 96%, 97%, 98%, 99%, 99.9%, 99.99%, or 99.999%. The techniques herein can enable accurate phasing at about 500× sequencing depth, 450× sequencing depth, 400× sequencing depth, 350× sequencing depth, 300× sequencing depth, 250× sequencing depth, 200× sequencing depth, 150× sequencing depth, 100× sequencing depth, or less than 50× sequencing depth. This phasing information can extend over longer distances, e.g., about 200 kbp, about 300 kbp, about 400 kbp, about 500 kbp, about 600 kbp, about 700 kbp, about 800 kbp, about 900 kbp, about 1 Mbp, about 2 Mbp, about 3 Mbp, about 4 Mbp, about 5 Mbp, or more than about 10 Mbp. In some embodiments, more than 90% of the heterozygous SNPs for a human sample can be phased with an accuracy exceeding 99% using less than about 250 million reads or read pairs, e.g., by using only 1 lane of Illumina HiSeq data. In other cases, more than about 40%, 50%, 60%, 70%, 80%, 90%, 95%, or 99% of the heterozygous SNPs for a human sample can be phased with an accuracy exceeding about 70%, 80%, 90%, 95%, 96%, 97%, 98%, 99%, 99.9%, 99.99%, or 99.999% using less than about 250 million or less than about 500 million reads or read pairs, e.g., by using only 1 or 2 lanes of Illumina HiSeq data. For example, more than 95% or 99% of the heterozygous SNPs for a human sample can be phased with an accuracy exceeding 95% or 99% using less than about 250 million or less than about 500 million reads.In further cases, additional variants can be captured by increasing the read length to about 200 bp, 250 bp, 300 bp, 350 bp, 400 bp, 450 bp, 500 bp, 600 bp, 800 bp, 1000 bp, 1500 bp, 2 kbp, 3 kbp, 4 kbp, 5 kbp, 10 kbp, 20 kbp, 50 kbp, or 100 kbp.
[0297] In other embodiments of the present disclosure, data from the XLRP library can be used to confirm the phasing ability of long-range read pairs. The accuracy of these results is equivalent to the best available technology previously, but extends significantly further to longer distances. The current sample preparation protocol for a particular sequencing method recognizes variants located within a read length of the target site for phasing, e.g., within 150 bp. In one example, 44% of the 1,703,909 heterozygous SNPs present were phased with an accuracy exceeding 99% from the XLRP library constructed for the assembly benchmark sample NA12878. In some cases, this percentage can be extended to almost all variable sites using judicious selection of enzymes or digestion conditions.
[0298] Haplotyping can include haplotyping of the human leukocyte antigen (HLA) region (e.g., class I HLA-A, B, and C, class II HLA-DRB1 / 3 / 4 / 5, HLA-DQA1, HLA-DQB1, HLA-DPA1, HLA-DPB1). FIG. 4 shows an exemplary haplotyped HLA genotype. The HLA region of the genome is a dense polymorphism and can be difficult to sequence or haplotype with standard sequencing approaches. The techniques of the present disclosure can provide improved sequencing and haplotyping accuracy of the HLA region of the genome. Using the techniques of the present disclosure, the HLA region of the genome can be accurately haplotyped as part of the haplotyping of a larger region (e.g., a chromosomal arm, a chromosome, the entire genome) or itself (e.g., by target enrichment such as hybrid capture). In one example, the HLA region itself was accurately haplotyped at a sequencing depth of about 300×. These techniques can provide advantages over conventional approaches for HLA analysis such as long-range PCR, which can include complex protocols and many separate reactions. As further discussed herein, for example, samples can be multiplexed for sequencing analysis by including sample identification barcodes in cross-linking oligonucleotides or elsewhere and demultiplexing sequence information based on the barcodes. In one example, multiple samples are subjected to proximity ligation, barcoded with a sample identification barcode (e.g., in a cross-linking oligonucleotide), the HLA region is targeted (e.g., by hybrid capture), multiplexed sequencing is performed, and haplotyping of the HLA region for multiple samples is enabled. In some cases, haplotyping of the HLA region is performed without complementation.
[0299] Haplotyping can include phasing of the killer cell immunoglobulin-like receptor (KIR) region. The KIR region of the genome is highly homologous and structurally dynamic due to transposon-mediated recombination, and can be difficult to sequence or phase using standard sequencing approaches. The techniques of the present disclosure can provide improved sequencing and phasing accuracy of the KIR region of the genome. Using the techniques of the present disclosure, the KIR region of the genome can be accurately phased as part of the phasing of a larger region (e.g., a chromosomal arm, a chromosome, the entire genome) or of itself (e.g., by target enrichment such as hybrid capture). These techniques can provide advantages over conventional approaches for HLA analysis such as long-range PCR, which can involve complex protocols and many separate reactions. As further discussed herein, for example, samples can be multiplexed for sequencing analysis by including sample identification barcodes in crosslinking oligonucleotides or elsewhere and demultiplexing sequence information based on the barcodes. For example, multiple samples can be subjected to proximity ligation, barcoded with a sample identification barcode (e.g., in a crosslinking oligonucleotide), the KIR region targeted (e.g., by hybrid capture), multiplex sequencing performed, and phasing of the KIR region for multiple samples enabled. At least about 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, or more genes and / or pseudogenes can be phased. In some cases, phasing of the KIR region is performed without complementation.
[0300] Metagenomic analysis In some embodiments, the compositions and methods described herein enable the investigation of metagenomes, such as those found in the human gut. Thus, it is possible to investigate the partial or complete genomic sequences of some or all of the organisms inhabiting a given ecological environment. Examples include random sequencing of all gut microbiota, microbiota found in specific regions of the skin, and microbiota surviving in toxic waste sites. The composition of the microbial populations in these environments can be determined using the compositions and methods described herein, as well as the associated biochemical aspects encoded by their respective genomes. The methods described herein can enable metagenomic studies from complex biological environments, such as biological environments containing 2, 3, 4, 5, 6, 7, 8, 9, 10, 12, 15, 20, 25, 30, 40, 50, 60, 70, 80, 90, 100, 125, 150, 175, 200, 250, 300, 400, 500, 600, 700, 800, 900, 1000, 5000, 10000, or more organisms and / or variants of organisms.
[0301] The high accuracy required for cancer genome sequencing can be achieved using the methods and systems described herein. Incorrect reference genomes can make basecalling difficult when sequencing cancer genomes. Heterogeneous samples and small starting materials, such as samples obtained by biopsy, pose further difficulties. Furthermore, the detection of large-scale structural variants and / or loss of heterozygosity is often essential not only for cancer genome sequencing but also for the ability to distinguish somatic variants from basecalling errors.
[0302] Improved sequencing accuracy The systems and methods described herein can generate accurate long sequences from complex samples containing two, three, four, five, six, seven, eight, nine, ten, twelve, fifteen, twenty, or more diverse genomes. Normal, benign, and / or tumor-derived mixed samples may optionally be analyzed without the need for normal controls. In some embodiments, small starting samples of 100 ng or a few hundred genome equivalents are utilized to generate accurate long sequences. The systems and methods described herein may enable the detection of large-scale structural variants and rearrangements. Phased variant calls can be obtained over long sequences spanning about 1 kbp, about 2 kbp, about 5 kbp, about 10 kbp, about 20 kbp, about 50 kbp, about 100 kbp, about 200 kbp, about 500 kbp, about 1 Mbp, about 2 Mbp, about 5 Mbp, about 10 Mbp, about 20 Mbp, about 50 Mbp, or about 100 Mbp, or more nucleotides. For example, phased variant calls can be obtained over long sequences spanning about 1 Mbp or about 2 Mbp.
[0303] Haplotypes determined using the methods and systems described herein may be assigned to computer resources, such as computer resources on a network, e.g., a cloud system. Short variant calls can be corrected, if necessary, using related information stored in computer resources. Structural variants can be detected based on combined information from short variant calls and information stored in computer resources. Problematic portions of the genome, e.g., segmental duplications, regions prone to structural polymorphisms, highly variable and medically relevant MHC regions, centromere and telomere regions, and, without limitation, repetitive regions, regions of low sequence accuracy, high variant rates, ALU repeats, other heterochromatic regions with segmental duplications, or other relevant problematic portions in the art, can be reassembled for improved accuracy.
[0304] The type of sample can be assigned to the array information locally or in network-connected computer resources such as the cloud. If the source of the information is known, for example, if the source of the information is from cancer or normal tissue, this source can be assigned to the sample as part of the type of sample. Examples of other sample types usually include, but are not limited to, tissue type, sample collection method, presence of infection, type of infection, treatment method, sample size, etc. If a complete or partial comparative genomic sequence, such as a normal genome in comparison to a cancer genome, is available, the difference between the sample data and the comparative genomic sequence can be determined and optionally output.
[0305] Clinical use The methods of the present disclosure can be used for the analysis of genetic information of genomic regions that can interact with a selected region of interest in addition to the selected region of interest of the genome. Amplification methods as disclosed herein can be used in devices, kits, and methods for genetic analysis, such as those found in U.S. Patent Nos. 6,449,562, 6,287,766, 7,361,468, 7,414,117, 6,225,109, and 6,110,709, among others. In some cases, the amplification methods of the present disclosure can be used to amplify target nucleic acids for DNA hybridization studies to determine the presence or absence of polymorphisms. Polymorphisms or alleles can be associated with diseases or disorders such as genetic diseases. In other cases, polymorphisms can be associated with susceptibility to diseases or disorders, for example, polymorphisms can be associated with poisoning, degenerative and age-related diseases, cancer, etc. In other cases, polymorphisms can be associated with useful characteristics such as an increase in coronary artery health, resistance to diseases such as HIV or malaria, or resistance to adult diseases such as osteoporosis, Alzheimer's disease, or dementia.
[0306] The compositions and methods of the present disclosure can be used for diagnostic, prognostic, therapeutic, patient stratification, drug development, treatment selection, and screening purposes. The present disclosure provides the advantage that many different target molecules can be analyzed from a single biological sample at once using the methods of the present disclosure. This enables, for example, performing various diagnostic tests on a single sample.
[0307] The compositions and methods of the present disclosure can be used in genomics. The methods described herein can quickly provide highly desirable answers for this application. The methods and compositions described herein can be used in the process of finding biomarkers that can be used for diagnosis or prognosis and as indicators of health and disease. The methods and compositions described herein can be used for screening drugs, for example, for drug development, treatment selection, determination of treatment effectiveness, and / or identification of targets for pharmaceutical development. The ability to test gene expression during a screening assay for a drug is very important because proteins are the final gene products in the body. In some embodiments, the methods and compositions described herein simultaneously measure both protein and gene expression that provide the most information regarding the particular screening being performed.
[0308] The compositions and methods of the present disclosure can be used for gene expression analysis. The methods described herein distinguish nucleotide sequences. The differences between target nucleotide sequences can be, for example, a single nucleotide difference, a nucleic acid deletion, a nucleic acid insertion, or a rearrangement. Such sequence differences containing more than one base can also be detected. The processes of the present disclosure can detect infectious diseases, genetic diseases, and cancers. Furthermore, the above processes are also useful in environmental monitoring, forensic science, and food science. Examples of gene analysis that can be performed on nucleic acids include, for example, SNP detection, STR detection, RNA expression analysis, promoter methylation, gene expression, virus detection, virus subtype classification, and drug resistance.
[0309] This method can be applied to the analysis of a biological molecule sample obtained from or derived from a patient to determine whether an affected cell type is present in the sample, the stage of the disease, the patient's prognosis, the patient's ability to respond to a particular treatment, or the best treatment for the patient. This method can also be applied to identify biomarkers for a particular disease.
[0310] In some embodiments, the methods described herein are used for the diagnosis of a disease. As used herein, the terms "diagnose" or "diagnosis" of a disease include predicting or diagnosing a disease, determining a predisposition to a disease, monitoring the treatment of a disease, diagnosing a treatment response, or prognosis, progression, or response to a particular treatment of a disease. For example, a blood sample can be assayed according to any of the methods described herein to determine the presence and / or amount of markers of a disease or malignant cell type in the sample, thereby diagnosing or staging the disease or cancer.
[0311] In some embodiments, the methods and compositions described herein are used for the diagnosis and prognosis of a disease.
[0312] A number of immunological, proliferative, and malignant diseases and disorders are particularly suitable for the methods described herein. Immune diseases and disorders include allergic diseases and disorders, immune dysfunction, and autoimmune diseases and disorders. Allergic diseases and disorders include, but are not limited to, allergic rhinitis, allergic conjunctivitis, allergic asthma, atopic eczema, atopic dermatitis, and food allergies. Immunodeficiencies include, but are not limited to, severe combined immunodeficiency (SCID), hypereosinophilic syndrome, chronic granulomatous disease, leukocyte adhesion deficiency I and II, hyper IgE syndrome, Chediak-Higashi, neutrophilia, neutropenia, agammaglobulinemia, hyper IgM syndrome, DiGeorge / velo-cardio-facial syndrome, and interferon-gamma-TH1 pathway deficiency. Autoimmune and immunoregulatory disorders include, but are not limited to, rheumatoid arthritis, diabetes, systemic lupus erythematosus, Graves' disease, Graves' ophthalmopathy, Crohn's disease, multiple sclerosis, psoriasis, systemic sclerosis, goiter and lymphomatous goiter (Hashimoto's thyroiditis, lymphadenoid goiter), alopecia areata, autoimmune myocarditis, lichen sclerosus, autoimmune uveitis, Addison's disease, atrophic gastritis, myasthenia gravis, idiopathic thrombocytopenic purpura, hemolytic anemia, primary biliary cirrhosis, Wegener's granulomatosis, polyarteritis nodosa, and inflammatory bowel disease, allograft rejection, and tissue destruction due to allergic reactions to infectious bacteria or environmental antigens.
[0313] Proliferative diseases and disorders that can be evaluated by the methods of the present disclosure include, but are not limited to, neonatal hemangioma, secondary progressive multiple sclerosis, chronic progressive myelodysplastic disease, neurofibromatosis, ganglioneuroma, keloid formation, Paget's disease of bone, fibrocystic disease (e.g., of the breast or uterus), sarcoidosis, Peyronie's and Dupuytren's fibrosis, cirrhosis, atherosclerosis, and vascular restenosis.
[0314] Malignant diseases and disorders that can be evaluated by the methods of the present disclosure include both hematological malignancies and solid tumors.
[0315] Hematological malignancies are particularly suitable for the methods of the present disclosure when the sample is a blood sample, as they involve changes in blood-derived cells. Such malignancies include non-Hodgkin lymphoma, Hodgkin lymphoma, non-B cell lymphoma, and other lymphomas, acute or chronic leukemia, polycythemia, thrombocythemia, multiple myeloma, myelodysplastic syndromes, myeloproliferative disorders, myelofibroses, atypical lymphoproliferation, and plasma cell disorders.
[0316] Plasma cell diseases that can be evaluated by the methods of the present disclosure include multiple myeloma, amyloidosis, and Waldenström macroglobulinemia.
[0317] Examples of solid tumors include, but are not limited to, colon cancer, breast cancer, lung cancer, prostate cancer, brain tumors, central nervous system tumors, bladder tumors, melanoma, liver cancer, osteosarcoma, and other bone cancers, testicular and ovarian car...
Claims
1. A method for nucleic acid processing, comprising: (a) obtaining a stabilized sample comprising a nucleic acid molecule complexed with at least one nucleic acid binding protein; (b) cleaving the nucleic acid molecule into a plurality of segments comprising at least a first segment and a second segment, wherein the cleavage is effected by a transposase; (c) producing a proximally ligated nucleic acid comprising a first sequence from the first segment and a second sequence from the second segment by ligating the first segment to the second segment. A method comprising the above steps.
2. The method according to claim 1, wherein the transposase is Tn5 transposase.
3. The method according to claim 1, further comprising circularizing the proximally ligated nucleic acid by ligating the 5' end of the proximally ligated nucleic acid to the 3' end thereof, thereby producing a circularized proximally ligated nucleic acid.
4. The method according to claim 1, further comprising sequencing at least a part of the proximally ligated nucleic acid.
5. The method according to claim 4, wherein the sequencing step comprises sequencing at least a part of the first sequence and at least a part of the second sequence.
6. The method according to claim 5, further comprising mapping at least a part of the first sequence and at least a part of the second sequence to a genome.
7. The method according to claim 4, further comprising performing a three-dimensional genome analysis using information from the sequencing step.
8. The method according to claim 1, wherein the stabilized sample is a cross-linked sample.
9. The method according to claim 1, wherein the step of obtaining the stabilized sample comprises obtaining a sample and stabilizing the sample.
10. The method according to claim 1, wherein the step of obtaining the stabilized sample comprises obtaining a pre-stabilized sample.
11. The method according to claim 1, wherein the nucleic acid binding protein comprises chromatin or a component thereof.
12. The method according to claim 1, wherein a linker sequence is ligated between the first segment and the second segment.
13. The method according to claim 12, wherein the linker sequence comprises a barcode sequence.
14. The method according to claim 13, wherein the barcode array indicates a partition of origin. **Claim 15** The method according to claim 13, wherein the barcode array indicates a cell of origin. **Claim 16** The method according to claim 13, wherein the barcode array indicates a cell population of origin. **Claim 17** The method according to claim 13, wherein the barcode array indicates an organism of origin. **Claim 18** The method according to claim 1, wherein the cleavage occurs in open and closed chromatin compartments. **Claim 19** The method according to claim 18, wherein at least 10% of the cleavage occurs in the closed chromatin compartment. **Claim 20** The method according to claim 18, wherein at least 20% of the cleavage occurs in the closed chromatin compartment. **Claim 21** The method according to claim 18, wherein at least 30% of the cleavage occurs in the closed chromatin compartment. **Claim 22** The method according to claim 1, wherein the stabilized sample contains 50,000 cells or less. **Claim 23** The method according to claim 22, wherein the stabilized sample contains at least 10,000 cells. **Claim 24** The method according to claim 1, wherein the stabilized sample contains stabilized nuclei. **Claim 25** The method according to claim 24, wherein the stabilized sample contains 50,000 nuclei or less. **Claim 26** The method according to claim 1, wherein the proximity-ligated nucleic acid does not contain an affinity tag. **Claim 27** The method according to claim 1, wherein the circularized proximity-ligated nucleic acid does not contain an affinity tag. **Claim 28** The method according to claim 3, wherein the circularized proximity-ligated nucleic acid has a length of more than 250 base pairs. **Claim 29** The method according to claim 3, wherein the circularizing step does not circularize a nucleic acid having a length of less than 250 base pairs. **Claim 30** The method according to claim 27 or claim 28, wherein the proximity-ligated nucleic acid and / or the circularized proximity-ligated nucleic acid is isolated without using an affinity tag.