Methods and compositions for nucleic acid analysis

By separating nucleic acid samples into discrete partitions and using bar coding and cross-linking techniques, the problem of structural and molecular background loss in nucleic acid sequencing in existing technologies has been solved, achieving efficient preservation of the three-dimensional spatial structure and molecular information of nucleic acids and providing more comprehensive sequence analysis.

CN115369161BActive Publication Date: 2026-03-0610X GENOMICS INC
View PDF 13 Cites 0 Cited by

Patent Information

Application Number
CN202211197364.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2015-12-04
Filing Date
2016-12-02
Publication Date
2026-03-06
Estimated Expiration
2036-12-02

AI Technical Summary

Technical Problem

Existing multinucleotide sequencing methods cannot effectively preserve the three-dimensional spatial structure and molecular background of the original nucleic acid molecules, resulting in a lack of position specificity of sequence information and loss of structural information.

Method used

By separating nucleic acid samples into discrete partitions and using bar coding and cross-linking techniques in conjunction with sequencing reactions, the three-dimensional structural and molecular backgrounds of nucleic acids are preserved. This includes using tag libraries and bead separation techniques to ensure that spatially close nucleic acid sequences are assigned to the same partitions, enabling short-read and long-read sequencing.

Benefits of technology

This technology preserves the structural and molecular background of nucleic acids during sequencing, providing information about genomic locus interactions and chromosome conformation, and improving the position specificity and structural fidelity of sequence information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure HDA0003870919450000011
    Figure HDA0003870919450000011
  • Figure HDA0003870919450000021
    Figure HDA0003870919450000021
  • Figure HDA0003870919450000031
    Figure HDA0003870919450000031
Patent Text Reader

Abstract

The present invention relates to methods, compositions, and systems for analyzing sequence information while preserving the structural and molecular background of the sequence information.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-reference to related applications

[0002] This application claims the benefit of U.S. Provisional Application No. 62 / 263,532, filed December 4, 2015, which is incorporated herein by reference in its entirety for all purposes. Background of the Invention

[0003] Polynucleotide sequencing is increasingly used in medical applications such as genetic screening and genotyping of tumors. Many polynucleotide sequencing methods rely on sample processing techniques for the original sample, including the random fragmentation of polynucleotides. These processing techniques can provide advantages in throughput and efficiency, but the resulting sequence information from these processed samples may lack important contextual information about the position of specific sequences within the broader linear (two-dimensional) sequence of the original nucleic acid molecule containing those sequences. The structural background of the original sample in three-dimensional space is also lost due to many sample processing and sequencing techniques. Therefore, sequencing techniques that preserve the structural and molecular background of the identified nucleic acid sequences are needed. Invention Overview

[0004] Therefore, the present invention provides methods, systems, and compositions for providing sequence information that preserves the molecular and structural background of the original nucleic acid molecule.

[0005] In some aspects, this disclosure provides a method for analyzing nucleic acids while maintaining a structural background. The method includes the steps of: (a) providing a sample containing nucleic acids, wherein the nucleic acids comprise a three-dimensional structure; (b) separating portions of the sample into discrete partitions such that portions of the three-dimensional structure of the nucleic acids are also separated into discrete partitions; and (c) obtaining sequence information from the nucleic acids, thereby analyzing the nucleic acids while maintaining a structural background.

[0006] In some implementations, the sequence information from obtaining step (c) includes identifying nucleic acids that are spatially close to each other.

[0007] In a further embodiment, the sequence information obtained from obtaining step (c) includes identifying nucleic acids that are spatially close to each other.

[0008] In a further embodiment, step (c) provides information about intrachromosomal and / or interchromosomal interactions between genomic loci.

[0009] In a further implementation, step (c) provides information about chromosome conformation.

[0010] In a further implementation, prior to the separation step (b), at least some three-dimensional structures are processed to connect different parts of nucleic acids that are close to each other within the three-dimensional structures.

[0011] In any implementation, the nucleic acid is not isolated from the sample prior to separation step (b).

[0012] In any implementation, prior to obtaining step (c), nucleic acids within discrete partitions are barcoded to form multiple barcoded segments, wherein each segment within a given discrete partition contains a common barcode, such that the barcode identifies the nucleic acid from the given partition.

[0013] In a further embodiment, step (c) includes a sequencing reaction selected from the group consisting of short read length sequencing reactions and long read length sequencing reactions.

[0014] In some aspects, this disclosure provides a method for analyzing nucleic acids while maintaining structural background, the method comprising the steps of: (a) forming ligated nucleic acids within a sample such that spatially adjacent nucleic acid segments are linked; (b) processing the ligated nucleic acids to produce a plurality of ligation products, wherein the ligation products contain portions of spatially adjacent nucleic acid segments; (c) depositing the plurality of ligation products into discrete partitions; (d) barcoding the ligation products within the discrete partitions to form a plurality of barcoded fragments, wherein each fragment within a given discrete partition contains a common barcode, thereby associating each fragment with the ligated nucleic acid from which it is derived; and (e) obtaining sequence information from the plurality of barcoded fragments to analyze nucleic acids from the sample while maintaining structural background.

[0015] In some aspects, this disclosure provides a method for analyzing nucleic acids while maintaining structural background, the method comprising the steps of: (a) forming ligated nucleic acids within a sample such that spatially adjacent nucleic acid segments are linked; (b) depositing the ligated nucleic acids into discrete partitions; (c) processing the ligated nucleic acids to produce a plurality of ligation products, wherein the ligation products contain portions of spatially adjacent nucleic acid segments; (d) barcoding the ligation products within the discrete partitions to form a plurality of barcoded fragments, wherein each fragment within a given discrete partition contains a common barcode, thereby associating each fragment with the ligated nucleic acid from which it is obtained; and (e) obtaining sequence information from the plurality of barcoded fragments to analyze nucleic acids from the sample while maintaining structural background.

[0016] In some aspects, this disclosure provides a method for analyzing nucleic acids while maintaining structural background, the method comprising the steps of: (a) cross-linking nucleic acids within a sample to form cross-linked nucleic acids, wherein the cross-linking forms covalent bonds between spatially adjacent nucleic acid fragments; (b) depositing the cross-linked nucleic acids into discrete partitions; (c) processing the cross-linked nucleic acids to generate a plurality of ligation products, wherein the ligation products contain portions of the spatially adjacent nucleic acid segments; and (d) obtaining sequence information from the plurality of ligation products to analyze nucleic acids from the sample while maintaining structural background.

[0017] In any implementation, the sample is a formalin-fixed paraffin sample.

[0018] In any embodiment, the discrete partitions comprise beads. In a further embodiment, the beads are gel beads.

[0019] In any implementation, the sample includes a tumor sample.

[0020] In any implementation, the sample comprises a mixture of tumor and normal cells.

[0021] In any implementation, the sample comprises a nuclear matrix.

[0022] In any implementation, the nucleic acid includes RNA.

[0023] In any implementation, the amount of nucleic acid in the sample is less than 5 ng / ml, 10 ng / ml, 15 ng / ml, 20 ng / ml, 25 ng / ml, 30 ng / ml, 35 ng / ml, 40 ng / ml, 45 ng / ml, or 50 ng / ml.

[0024] In some aspects, the present invention provides a method for analyzing nucleic acids while maintaining structural background, wherein the method includes the following steps: (a) providing a sample containing nucleic acids; (b) applying a tag library to the sample such that different geographic regions of the sample receive different tags or tags of different concentrations; (c) separating portions of the sample into discrete partitions such that portions of the tag library and portions of the nucleic acids are also separated into discrete partitions; (d) obtaining sequence information from the nucleic acids; and (e) identifying tags or tag concentrations in the discrete partitions, thereby analyzing nucleic acids while maintaining structural background. Brief description of the attached diagram

[0025] Figure 1 A schematic diagram of the molecular and structural backgrounds for the methods described herein is provided.

[0026] Figure 2 A schematic diagram of the method described in this article is provided.

[0027] Figure 3 This illustrates a typical workflow for determining sequence information using the methods and compositions disclosed herein.

[0028] Figure 4 A schematic diagram is provided of a method for combining nucleic acid samples with beads and distributing the nucleic acid and beads into discrete droplets.

[0029] Figure 5 A schematic diagram is provided for a method of barcoding and amplifying chromosomal nucleic acid fragments.

[0030] Figure 6 This diagram illustrates the use of barcoding nucleic acid fragments in assigning sequence data to their original source nucleic acid molecules.

[0031] Figure 7 A schematic diagram of an exemplary sample preparation method is provided. Invention Details

[0032] Unless otherwise indicated, this invention can be practiced using conventional techniques and descriptions within the art, including those of organic chemistry, polymer technology, molecular biology (including recombinant techniques), cell biology, biochemistry, and immunology. These conventional techniques include polymer array synthesis, hybridization, ligation, phage display, and label-based detection of hybridization. Specific descriptions of suitable techniques can be found in the examples below. However, other equivalent conventional procedures can certainly be used. The routine techniques and descriptions described can be found in standard laboratory manuals such as the following: Genome Analysis: A Laboratory Manual Series (Volumes I-IV), Using Antibodies: A Laboratory Manual, Cells: A Laboratory Manual, PCR Primer: A Laboratory Manual, and Molecular Cloning: A Laboratory Manual (all from Cold Spring Harbor Laboratory Press), Stryer, L. (1995) Biochemistry (4th edition), Freeman, New York, Gait, “Oligonucleotide Synthesis: A Practical Approach” 1984, IRL Press, London, Nelson and Cox (2000), Lehninger, Principles of Biochemistry 3rd edition, WH Freeman Pub., New York, NY, and Berg et al., (2002) Biochemistry, 5th edition, WH Freeman Pub., New York, NY, all of which are incorporated herein by reference in their entirety for all purposes.

[0033] It should be noted that, unless the context explicitly specifies otherwise, the singular forms “a,” “an,” and “the” used herein and in the appended claims include multiple indicators. Thus, for example, a reference to “a polymerase” refers to a reagent or a mixture of said reagents, and a reference to “the method” includes references to equivalent steps and methods known to those skilled in the art, and so on.

[0034] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. All disclosures mentioned herein are incorporated by reference for the purpose of describing and disclosing apparatuses, compositions, formulations, and methodologies described in the publications and that may be used in conjunction with the invention described herein.

[0035] If a range of values ​​is provided, it should be understood that, unless the context clearly specifies otherwise, this invention includes all interpolated values ​​(up to one-tenth of the lower limit unit) between the upper and lower limits of the range, and any other stated or interpolated values ​​within the stated range. This invention also includes the upper and lower limits of the smaller ranges that may be independently included in these smaller ranges, and any specific excluded limits belonging to the stated range. When the range includes one or two limits, the range excluding one or both of those included limits is also included in this invention.

[0036] In the following description, numerous specific details are set forth in order to provide a more detailed understanding of the invention. However, it will be apparent to those skilled in the art that the invention may be practiced without one or more of these specific details. In other instances, well-known features and procedures familiar to those skilled in the art have not been described in order to avoid obscuring the invention.

[0037] As used herein, the term "comprising" is intended to mean that a composition and method include the stated elements, but does not exclude other elements. "Constitutes substantially of..." when used to define a composition and method should mean excluding other elements that are of any significant importance to the composition or method. "Constitutes of..." should mean excluding trace amounts of other components beyond the claimed composition and substantial method steps. Embodiments defined by these transitional terms are within the scope of this invention. Therefore, it is intended that the method and composition may include additional steps and components (comprising) or optionally include insignificant steps and components (consistently of...) or optionally include only the method steps or components (consistent of...).

[0038] All numerical designations such as pH, temperature, time, concentration, and molecular weight (including range) are approximate values, varying in increments of 0.1 (+) or (-). Although not always explicitly stated, it should be understood that all numerical designations are preceded by the term "about". In addition to the small increments of "X" such as "X+0.1" or "X-0.1", the term "about" also includes the exact value "X". Although not always explicitly stated, it should also be understood that the reagents described herein are exemplary only, and equivalents of said reagents are known in the art.

[0039] I. Overview

[0040] This disclosure provides methods, compositions, and systems for characterizing genetic materials. Generally, the methods, compositions, and systems described herein provide ways to analyze components of a sample while preserving information about the original structure and molecular background of those components in the sample. Although much of the discussion herein pertains to the analysis of nucleic acids, it should be understood that the methods and systems discussed herein can be applied to other components of a sample, including proteins and other molecules.

[0041] Deoxyribonucleic acid (DNA) is a linear molecule, and therefore the genome is often described and evaluated using linear dimensions. However, chromosomes are not rigid, and the spatial distance between two genomic loci does not always correspond to their distance along the linear sequence of the genome. In three-dimensional space, regions separated by many megabases can be directly adjacent. From a regulatory perspective, understanding long-range interactions between genomic loci can be useful. For example, gene enhancers, silencers, and insulators can function over vast genomic distances. The ability to preserve the structural and molecular context of sequence reads provides the means to understand such long-range interactions.

[0042] As used herein, “preserving structural background” means that multiple sequence reads or portions of sequence reads can be attributed to the original three-dimensional relative positions of those reads within a sample. In other words, a sequence read can be associated with the relative positions of adjacent nucleic acids (and in some cases, related proteins) within the sample. This spatial information can be obtained using the methods discussed herein, even if those adjacent nucleic acids are not physically located within the linear sequence of a single original nucleic acid molecule. References Figure 1 The schematic diagram shows that in sample (101), sequences (104) and (105) lie within the linear sequences of two distinct original nucleic acid molecules (102 and (103), respectively), but are spatially close to each other within the sample. The methods and compositions described herein provide the ability to preserve information about the structural background of sequence reads and thus assign reads from sequences (104) and (105) to their relative spatial proximity within the original sample to the original nucleic acid molecules (102) and (103) from which those sequence reads were obtained.

[0043] The methods and compositions discussed herein also provide sequence information that preserves molecular background. As used herein, “preserving molecular background” means that multiple sequence reads or multiple portions of sequence reads can be attributed to a single original molecule of nucleic acid. Although this single nucleic acid molecule can be of any length, in a preferred aspect, it will be a relatively long molecule, thereby allowing preservation of a long range of molecular background. Specifically, the individual original molecule is preferably substantially longer than the typical short read sequence length, for example, longer than 200 bases, and typically at least 1,000 bases or longer, 5,000 bases or longer, 10,000 bases or longer, 20,000 bases or longer, 30,000 bases or longer, 40,000 bases or longer, 50,000 bases or longer, 60,000 bases or longer, 70,000 bases or longer, 80,000 bases or longer, 90,000 bases or longer, or 100,000 bases or longer, and in some cases up to 1 megabase or longer.

[0044] Typically, the methods described herein involve analyzing nucleic acids while maintaining structural and molecular backgrounds. Such analyses include methods in which a sample containing nucleic acids, wherein the nucleic acids have a three-dimensional structure, is provided. Parts of the sample are separated into discrete partitions, such that portions of the three-dimensional structure of the nucleic acids are also separated into discrete partitions—nucleic acid sequences spatially close to each other tend to be separated into the same partitions, thus preserving this three-dimensional information of spatial proximity even when subsequent sequence reads originate from sequences that were not originally on the same individual original nucleic acid molecule. See again Figure 1 If sample 101, containing nucleic acid molecules 102, 103, and 106, is separated into discrete partitions, such that subsets of the sample are allocated to different discrete partitions, then due to the physical distance between nucleic acid molecules 106 and 102 and 103, nucleic acid molecules 102 and 103 are more likely to be placed in the same partition than nucleic acid molecule 106. Therefore, nucleic acid molecules within the same discrete partition are molecules that are spatially close to each other in the original sample. The sequence information obtained from nucleic acids within discrete partitions thus provides a way to analyze nucleic acids, for example, through nucleic acid sequencing, and assigns those sequence reads back to the structural context of the original nucleic acid molecule.

[0045] In a further example, structural background (also referred to herein as “geographical background”) can be maintained by encoding the geographical location of the sample using tags (such as barcode oligonucleotides). In some cases, this may include injecting a viral library encoding a set of barcode-coded sequences (such as mRNA sequences) into the sample. The barcodes are transmitted through the sample via active processes or by diffusion. When the sample is subsequently further processed according to methods described herein and known in the art, the barcodes can be correlated with structural locations to identify nucleic acid sequences within the sample from the same geographical location. In instances where barcodes are distributed throughout the sample via active processes, sequences with the same barcodes can be geographically linked and / or linked via the same processes. As will be understood, the system of using tags to encode structural background can be used alone or in combination with methods described herein that utilize discrete partitions to further preserve both structural and molecular backgrounds. In instances using tags for encoding spatial locations and barcodes for identifying molecules isolated into the same discrete partitions, the sample is essentially tagged or “double-barcoded,” where one set of barcodes is used to identify the spatial location and the other set is partition-specific. In the example described, both sets of barcodes can be used to provide information to preserve the structural and molecular background of the sequence reads generated from the sample.

[0046] In some instances, sequence information obtained from nucleic acids provides information about intrachromosomal and / or interchromosomal interactions between genomic loci. In further instances, sequence information includes information about chromosome conformation.

[0047] In a further example, before being separated into discrete partitions, nucleic acids in a sample can be processed to connect different regions of their three-dimensional structure, such that sequence regions that are close to each other within those three-dimensional structures are attached to each other. Thus, separating the sample into discrete partitions will separate these connected regions into the same partition, thereby further ensuring the preservation of the structural background of any sequence reads from those nucleic acids.

[0048] In some cases, nucleic acid ligation can be accomplished using any method known in the art for cross-linking spatially proximal molecules. The cross-linking agents may include, but are not limited to, alkylating agents, cisplatin, nitrous oxide, psoralen, aldehydes, acrolein, glyoxal, osmium tetroxide, carbodiimide, mercuric chloride, zinc salts, picric acid, potassium dichromate, ethanol, methanol, acetone, acetic acid, etc. In specific instances, nucleic acids are ligated using schemes designed for analyzing the three-dimensional structure of the genome, such as the “Hi-C” scheme described, for example, by Dekker et al., “Capturing chromosome conformation” Science 295:1306-1311 (2002) and Berkum et al., J.Vis.Exp.(39),e1869,doi:10.3791 / 1869 (2010), each of which is incorporated herein by reference, in its entirety and particularly, for all teachings relating to the ligation of nucleic acid molecules. Such schemes typically involve generating molecular libraries by cross-linking samples to ligate closely spatially proximal genomic loci. In a further embodiment, the intervening DNA loops between crosslinks are digested, and then the intra-sequence regions are decrosslinked to add to the library. The digestion and decrosslinking steps can occur before the step of separating the sample into discrete partitions, or they can occur within the partitions after the separation step.

[0049] In a further instance, nucleic acids may undergo a labeling or barcoding step that provides a common barcode for all nucleic acids within a partition. As will be understood, this barcoding can occur with or without the nucleic acid linking / crosslinking steps discussed above. The barcoding techniques disclosed herein provide a unique ability to provide separate structural and molecular contexts for genomic regions—that is, to provide a broader or even longer inferred context across multiple sample nucleic acid molecules and / or for a specific chromosome by assigning certain sequence reads to individual sample nucleic acid molecules and by variant coordination assembly. As used herein, the term “genomic region” or “region” refers to any defined length of genome and / or chromosome. For example, a genomic region can refer to associations (i.e., interactions) between more than one chromosome. A genomic region can also encompass an entire chromosome or a portion of a chromosome. Furthermore, a genomic region can include specific nucleic acid sequences on a chromosome (i.e., open reading frames and / or regulatory genes) or non-coding regions between genes.

[0050] Using barcoding provides the additional advantage of facilitating the differentiation between minority and majority components of the total nucleic acid population extracted from a sample, such as for the detection and characterization of circulating tumor DNA in the bloodstream, and also reduces or eliminates amplification bias during optional amplification steps. Furthermore, implementation in a microfluidic manner provides the ability to operate with extremely small sample volumes and low DNA input, as well as the ability to rapidly process large numbers of sample partitions (droplets) to facilitate whole-genome labeling.

[0051] In addition to providing the ability to obtain sequence information from the whole or selected regions of the genome, the methods and systems described herein can also provide other characterizations of genomic material, including but not limited to the identification of haplotype phasing, structural variations, and copy number variations, as described in USSN 14 / 316,383, 14 / 316,398, 14 / 316,416, 14 / 316,431, 14 / 316,447, and 14 / 316,463, all of which are incorporated herein by reference in their entirety for all purposes, and particularly for all written descriptions, figures, and working examples relating to the characterization of genomic material.

[0052] Typically, the method of the present invention includes, as follows: Figure 2 The steps described are illustrated, providing schematic diagrams of the method of the invention discussed in further detail herein. As will be understood, Figure 2 The methods outlined herein are exemplary implementations that can be changed or modified as needed, as described herein. Figure 2As shown, the method described herein may include an optional step 201, in which sample nucleic acids are processed to link nucleic acids that are spatially close to each other. With or without this initial processing step (201), the method described herein will in most instances include a step (202) in which a sample containing nucleic acids is distributed. Typically, each partition containing nucleic acids from the genomic region of interest will undergo a process to generate a barcoded fragment (203). Those fragments may then be pooled (204) prior to sequencing (205). The sequence reads from (205) can be attributed to the original structural and molecular background (206) typically attributed to the partition-specific barcode (203). Each partition may include more than one nucleic acid in some instances, and in some cases will contain hundreds of nucleic acid molecules. The barcoded fragments of step 203 can be generated using any method known in the art – in some instances, oligonucleotides are included with the sample in different partitions. The oligonucleotides may contain random sequences designed to randomly primate multiple different regions of the sample, or they may contain specific primer sequences targeted upstream of a target region of the sample to primate it. In further examples, these oligonucleotides also contain barcode sequences, such that the replication process also barcodes the resulting replicated fragments of the original sample nucleic acid. Particularly elegant methods for using these barcode oligonucleotides in amplification and barcoding of samples are described in detail in USSN 14 / 316,383; 14 / 316,398; 14 / 316,416; 14 / 316,431; 14 / 316,447; and 14 / 316,463, each of which is incorporated herein by reference, in its entirety and particularly for all purposes, of the teachings relating to barcoding and amplification of oligonucleotides. Also included in the partitions are extension reaction reagents such as DNA polymerase, nucleosides triphosphates, and cofactors (e.g., Mg). 2+ or Mn 2+(etc.) The sample is then used as a template to extend primer sequences to generate complementary fragments of the template strand annealed with the primers, and said complementary fragments comprise oligonucleotides and their associated barcode sequences. Annealing and extending multiple primers to different portions of the sample can produce large aggregates of overlapping complementary fragments of the sample, each having its own barcode sequence indicating the partition from which it was generated. In some cases, these complementary fragments themselves can serve as templates initiated by oligonucleotides present in the partitions to generate complement, which again comprises barcode sequences. In a further example, the replication process is constructed such that when the first complement is replicated, it generates two complementary sequences at or near its ends to allow the formation of hairpin or partial hairpin structures, which reduces the molecule's ability to become the basis for generating additional iterative copies. The advantage of the methods and systems described herein is that attaching partition- or sample-specific barcodes to copied fragments preserves the original molecular background of the sequencing fragments, thereby allowing them to be attributed to their original partitions and therefore to their original sample nucleic acid molecules.

[0053] Typically, a sample is combined with a set of oligonucleotide tags that are releasably attached to the beads prior to the dispensing step. Methods for bar-encoded nucleic acids are known in the art and are described herein. In some instances, methods are utilized, such as those described in Amini et al., 2014, Nature Genetics, Advance Online Publication, which are incorporated herein by reference, in whole and in particular, all teachings relating to the attachment of barcodes or other oligonucleotide tags to nucleic acids for all purposes. Methods for processing nucleic acids and sequencing nucleic acids according to the methods and systems described in this application are also described in further detail in USSN 14 / 316,383; 14 / 316,398; 14 / 316,416; 14 / 316,431; 14 / 316,447; and 14 / 316,463, which are incorporated herein by reference, in whole and in particular, all written descriptions, figures, and working examples relating to the processing of nucleic acids and the sequencing and other characterization of genomic material for all purposes.

[0054] In addition to the workflows described above, methods including chip-based and solution-based capture methods can be used to enrich, isolate, or separate (i.e., “pull down”) targeted genomic regions for further analysis, particularly sequencing. These methods utilize probes complementary to the genomic region of interest or to regions near or adjacent to it. For example, in hybridization (or chip-based) capture, a microarray containing capture probes (typically single-stranded oligonucleotides) with sequences that together cover the region of interest is immobilized on a surface. Genomic DNA is fragmented and can be further processed, such as end repair, to produce blunt ends and / or to add additional features such as universal priming sequences. These fragments hybridize with probes on the microarray. Unhybridized fragments are washed away, and desired fragments are eluted or otherwise processed on the surface for sequencing or other analysis, thus enriching the population of fragments remaining on the surface containing fragments of the targeted region of interest (e.g., regions containing sequences complementary to those sequences contained in the capture probes). The enriched population of fragments can be further amplified using any amplification techniques known in the art. Exemplary methods for the targeted pull-down enrichment method are described in USSN 14 / 927,297, filed October 29, 2015, which is incorporated herein by reference, in its entirety and particularly for all teachings relating to the targeted pull-down enrichment method and sequencing method, including all written descriptions, figures, and embodiments. Populations of targeted genomic regions can be further enriched prior to the aforementioned pull-down method by methods that increase the coverage of those targeted regions. This increased coverage can be achieved, for example, using targeted amplification methods, including those described, for example, in USSN 62 / 119,996, filed February 24, 2015, which is incorporated herein by reference, in its entirety and particularly for all teachings relating to targeted coverage of nucleic acid molecules.

[0055] In specific cases, the methods described herein include the step of selectively amplifying selected regions of the genome prior to sequencing. This amplification, typically performed using methods known in the art (including, but not limited to, PCR amplification), provides at least 1X, 10X, 20X, 50X, 100X, 200X, 500X, 1000X, 1500X, 2000X, 5000X, or 10000X coverage of the selected regions of the genome, thereby providing a sufficient amount of nucleic acid to allow de novo sequencing of those selected regions. In a further embodiment, the amplification provides at least 1X-20X, 50X-100X, 200X-1000X, 1500X-5000X, 5000X-10,000X, 1000X-10000X, 1500X-9000X, 2000X-8000X, 2500X-7000X, 3000X-6500X, 3500X-6000X, and 4000X-5500X coverage of selected regions of the genome.

[0056] Amplification is typically performed by extending primers complementary to sequences in or near selected regions of the genome. In some cases, primer libraries designed to tile across the region of interest are used—in other words, primer libraries designed to amplify regions at specific distances along selected regions of the genome. In some cases, selective amplification utilizes primers complementary to 10, 15, 20, 25, 50, 100, 200, 250, 500, 750, 1000, or 10000 bases along selected regions of the genome. In further instances, primer-tiled libraries are designed to capture mixtures of distances, which can be random mixtures of distances or intelligently designed such that specific portions or percentages of selected regions are amplified by different primer pairs. Further information on targeted coverage of the genome used according to the methods described herein is provided, for example, in USSN 62 / 146,834, filed April 13, 2015, which is incorporated herein by reference, for all purposes, in whole and in particular, all teachings relating to targeted coverage of the genome.

[0057] Typically, the methods and systems described herein provide nucleic acids for analyses such as sequencing. Sequencing information is obtained using methods with the advantages of extremely low sequencing error rates and high throughput from short-read sequencing technologies. As mentioned above, nucleic acid sequencing is generally performed in a manner that preserves both the structural and molecular backgrounds of sequence reads or portions of sequence reads. This means that multiple sequence reads or portions of sequence reads can be assigned to their spatial locations (structural background) relative to other nucleic acids in the original sample, as well as to their locations along the linear sequence of a single original molecule of nucleic acid (molecular background). Although this single nucleic acid molecule can be of any length, it will preferably be a relatively long molecule, thus allowing for the preservation of a long range of molecular backgrounds. Specifically, the individual original molecule is preferably substantially longer than the typical short read sequence length, for example, longer than 200 bases, and typically at least 1,000 bases or longer, 5,000 bases or longer, 10,000 bases or longer, 20,000 bases or longer, 30,000 bases or longer, 40,000 bases or longer, 50,000 bases or longer, 60,000 bases or longer, 70,000 bases or longer, 80,000 bases or longer, 90,000 bases or longer, or 100,000 bases or longer, and in some cases up to 1 megabase or longer.

[0058] As noted above, the methods and systems described herein provide a separate molecular background for short reads of longer nucleic acids. As used herein, a separate molecular background refers to a sequence background that extends beyond a specific read, for example, relating to adjacent or proximal sequences not included within the read itself, and thus generally ensuring that they are not wholly or partially included in the short reads used for paired reads, such as reads of about 150 or about 300 bases. In a particularly preferred aspect, the methods and systems provide a long-range sequence background for short reads. This long-range background includes the relationship or association of a given read with reads that are spaced more than 1 kb, 5 kb, 10 kb, 15 kb, 20 kb, 30 kb, 40 kb, 50 kb, 60 kb, 70 kb, 80 kb, 90 kb, or even 100 kb or longer. As will be understood, by providing a long range of individual molecular backgrounds, it is also possible to derive phase information for variants within those individual molecular backgrounds; for example, variants of a particular long molecule are typically phase-separated by definition.

[0059] By providing a longer range of individual molecular backgrounds, the methods and systems of the present invention also provide much longer inferred molecular backgrounds (also referred to herein as “long virtual single-molecule reads”). Sequence backgrounds as described herein can include mapping or providing concatenations of different (typically on the kilobase scale) ranges of fragments across a complete genome sequence. These methods include mapping short sequence reads to contigs of individual longer molecules or linking molecules, and long-range sequencing of the bulk of longer individual molecules, such as those with consecutive determined sequences, where such determined sequences are longer than 1 kb, longer than 5 kb, longer than 10 kb, longer than 15 kb, longer than 20 kb, longer than 30 kb, longer than 40 kb, longer than 50 kb, longer than 60 kb, longer than 70 kb, longer than 80 kb, longer than 90 kb, or even longer than 100 kb. Similar to sequence background, assigning short sequences to longer nucleic acids (e.g., a single long nucleic acid molecule or a collection of linked nucleic acid molecules or contigs) can include mapping the short sequences against the longer nucleic acid segments to provide a high level of sequence background and providing the assembled sequence from the short sequences through these longer nucleic acids.

[0060] Furthermore, while a long-range sequence background associated with a long individual molecule can be utilized, having such a long-range sequence background also allows for the inference of even longer sequence backgrounds. For example, by providing the aforementioned long-range molecular background, overlapping variant portions, such as phase variants, translocation sequences, etc., within long sequences from different original molecules can be identified, thereby allowing for inferred bonding between those molecules. Such inferred bonding or molecular backgrounds are referred to herein as “inferred overlap groups.” In some cases, when discussed in the context of phase sequences, inferred overlap groups can represent typical phase sequences, for example, where phase overlap groups substantially longer than those of individual original molecules can be inferred from overlapping phase variants. These phase overlap groups are referred to herein as “phase blocks.”

[0061] By starting with a longer single-molecule read (e.g., the “long virtual single-molecule read” discussed above), inferred contigs or phase blocks longer than those that could be obtained using short-read sequencing techniques or other phase sequencing methods can be derived. See, for example, published U.S. Patent Application No. 2013-0157870. In particular, using the methods and systems described herein, inferred contig or phase block lengths having an N50 of at least about 10 kb, at least about 20 kb, or at least about 50 kb can be obtained (wherein the sum of block lengths greater than the N50 value is 50% of the sum of all block lengths). In a more preferred aspect, inferred contig or phase block lengths having an N50 of at least about 100 kb, at least about 150 kb, at least about 200 kb, and in many cases at least about 250 kb, at least about 300 kb, at least about 350 kb, at least about 400 kb, and in some cases at least about 500 kb or longer are obtained. In other cases, maximum block lengths exceeding 200kb, 300kb, 400kb, 500kb, 1Mb, or even 2Mb can be obtained.

[0062] In one respect, and in conjunction with any methods described above and subsequently herein, the methods and systems described herein provide for the compartmentalization, deposition, or partitioning of sample nucleic acids or fragments thereof into discrete compartments or partitions (which are interchangeably referred to herein as partitions), wherein each partition maintains its own contents separate from the contents of other partitions. Unique identifiers such as barcodes may be delivered before, subsequently, or simultaneously to the partitions containing the compartmentalized or partitioned sample nucleic acids to allow for the subsequent attribution of features such as nucleic acid sequence information to the sample nucleic acids contained within a particular compartment, and particularly to relatively long segments of consecutive sample nucleic acids that can be originally deposited into the partition. This subsequent attribution further allows attribution to the original structural context of those sample nucleic acids in the original sample, since nucleic acids that are close to each other in the three dimensions of the original sample are more likely to be deposited into the same partition. Thus, attribution of sequence reads to partitions (and the nucleic acids contained within those partitions) provides not only a molecular context regarding the linear position of the original nucleic acid molecule from which the sequence reads are obtained, but also a structural context for identifying sequence reads of nucleic acids that are spatially close to each other in the three-dimensional context of the original sample.

[0063] The sample nucleic acids used in the methods described herein typically represent multiple overlapping portions of the overall sample to be analyzed, such as entire chromosomes, exomes, or other large genomic regions. These sample nucleic acids can include whole genomes, individual chromosomes, exomes, amplicon sequences, or any of the various nucleic acids of interest. Sample nucleic acids are typically partitioned such that they exist as relatively long fragments or segments of continuous nucleic acid molecules within the partition. These fragments of sample nucleic acids can typically be longer than 1 kb, 5 kb, 10 kb, 15 kb, 20 kb, 30 kb, 40 kb, 50 kb, 60 kb, 70 kb, 80 kb, 90 kb, or even 100 kb, allowing for such a long range of structural and molecular backgrounds.

[0064] Sample nucleic acids are typically also allocated at a certain level, such that a given partition has a very low probability of containing two overlapping fragments of a genomic locus. This is typically accomplished by providing sample nucleic acids at low input volumes and / or concentrations during the allocation process. As a result, in a preferred embodiment, a given partition may comprise many long but non-overlapping fragments of the starting sample nucleic acids. Sample nucleic acids in different partitions are then associated with unique identifiers, wherein for any given partition, the nucleic acids contained therein have the same unique identifier, but different partitions may comprise different unique identifiers. Furthermore, since the allocation step configures sample components into very small partitions or droplets, it should be understood that, in order to achieve the desired configuration as described above, there is no need for extensive sample dilution, as would be required in higher-capacity processes such as in the wells of tubes or multi-well plates. Additionally, due to the high level of barcode diversity employed in the system described herein, diverse barcodes can be configured among a large number of genomic equivalents, as provided above. Specifically, the previously described multi-well plate methods (see, for example, U.S. Publication Applications Nos. 2013-0079231 and 2013-0157870) typically operate with only a few hundred different barcode sequences and employ limiting dilution processes of their samples to be able to assign barcodes to different cells / nucleic acids. Therefore, they will typically operate with far fewer than 100 cells, which will generally provide a genome:(barcode type) ratio of approximately 1:10 and certainly far greater than 1:100. On the other hand, the system described herein can operate at a genome:(barcode type) ratio of approximately 1:50 or lower, 1:100 or lower, 1:1000 or lower, or even smaller due to the high level of barcode diversity, such as more than 10,000, 100,000, 500,000, etc., while also allowing the loading of a higher number of genomes (e.g., approximately more than 100 genomes per assay, more than 500 genomes per assay, 1000 genomes per assay, or even more), while still providing a significantly improved barcode diversity per genome.

[0065] In a further example, the oligonucleotides included along with the sample portions divided into discrete partitions may comprise at least a first region and a second region. The first region may be a barcode region, which may be substantially the same barcode sequence among the oligonucleotides within a given partition, but may be, and in most cases, different barcode sequences between different partitions. The second region may be an N-mer (random N-mer or N-mer designed to target a specific sequence) that can be used to initiate nucleic acids within the sample within a partition. In some cases, where the N-mer is designed to target a specific sequence, it may be designed to target a specific chromosome (e.g., chromosome 1, 13, 18, or 21) or a region of the chromosome, such as an exon or other targeting region. In some cases, the N-mer may be designed to target a specific gene or gene region, such as a gene or region associated with a disease or condition (e.g., cancer). Within a partition, an amplification reaction may be performed using the second N-mer to initiate the nucleic acid sample at different locations along the length of the nucleic acid. As a result of the amplification, each partition may contain amplified products of nucleic acids attached to the same or nearly identical barcodes and may represent overlapping smaller fragments of nucleic acids in each partition. Barcodes can serve as markers indicating sets of nucleic acids originating from the same region and therefore potentially from the same strand of nucleic acid. After amplification, the nucleic acids can be pooled, sequenced, and aligned using sequencing algorithms. Since shorter sequence reads can be aligned with their associated barcode sequences and assigned to a single long fragment of the sample nucleic acid, all identified variants on that sequence can be assigned to a single original fragment and a single original chromosome. Furthermore, the chromosome contribution can be further characterized by aligning multiple co-localized variants across multiple long fragments. Therefore, conclusions regarding the phasing of specific gene variants can then be drawn, such as through analysis across long-range genomic sequences—for example, the identification of sequence information for segments of poorly characterized regions across the genome. This information can also be used to identify haplotypes, which are typically a specific group of gene variants residing on the same or different nucleic acid strands. Copy number variations can also be identified in this manner.

[0066] The described methods and systems offer significant advantages over current nucleic acid sequencing technologies and their associated sample preparation methods. Whole-sample preparation and sequencing methods tend to primarily identify and characterize the majority of components in a sample and are not designed to identify and characterize the small percentage of components that constitute the total DNA in the extracted sample, such as genetic material from poorly characterized or highly polymorphic regions of the genome contributed by a single chromosome, material from one or a few cells, or fragmented tumor cell DNA molecules circulating in the bloodstream. The methods described herein include methods for selectively amplifying genetic material from these minority components, and the ability to preserve the molecular background of this genetic material further provides genetic characterization of these components. The described methods and systems also offer significant advantages for detecting populations present within larger samples. Therefore, they are particularly useful for assessing haplotype and copy number variations—the methods disclosed herein can also be used to provide sequence information for genomic regions that are poorly characterized or poorly represented in a population of nucleic acid targets due to biases introduced during sample preparation.

[0067] The barcoding technique disclosed herein provides a unique ability to provide a separate molecular background for a given set of genetic markers, that is, to assign a given set of genetic markers (as opposed to a single marker) to a separate sample nucleic acid molecule, and through variant coordination assembly, to provide a broader or even longer range of inferred separate structural and molecular backgrounds across multiple sample nucleic acid molecules and / or to a specific chromosome. These genetic markers may include specific genetic loci, such as variants, like SNPs, or they may include short sequences. Furthermore, barcoding provides the added advantage of facilitating the differentiation of minority and majority components of the total nucleic acid population extracted from a sample, for example, for the detection and characterization of circulating tumor DNA in the bloodstream, and also reduces or eliminates amplification bias during optional amplification steps. Moreover, implementation in a microfluidic manner provides the ability to operate with extremely small sample volumes and low DNA input levels, as well as the ability to rapidly process large numbers of sample partitions (droplets) to facilitate whole-genome labeling.

[0068] As previously stated, the advantage of the methods and systems described herein lies in their ability to achieve the desired results using ubiquitous short-read sequencing technologies. These technologies are readily available and widely distributed within the research community using well-characterized and efficient protocols and reagent systems. These short-read sequencing technologies include those available from, for example, Illumina, Inc. (GAIIx, NextSeq, MiSeq, HiSeq, X10), Thermo-Fisher's Ion Torrent division (Ion Proton and IonPGM), pyrosequencing methods, and others.

[0069] Of particular advantage is that the methods and systems described herein utilize these short-read sequencing technologies and operate with their associated low error rates and high throughput. Specifically, the methods and systems described herein achieve the desired individual molecular read lengths or backgrounds as described above, but with individual sequencing reads shorter than 1000 bp, shorter than 500 bp, shorter than 300 bp, shorter than 200 bp, shorter than 150 bp, or even shorter (excluding pair extensions); and with sequencing error rates of less than 5%, less than 1%, less than 0.5%, less than 0.1%, less than 0.05%, less than 0.01%, less than 0.005%, or even less than 0.001% for said individual molecular read lengths.

[0070] II. Workflow Overview

[0071] In one exemplary aspect, the methods and systems described in this disclosure provide for depositing or distributing samples into discrete partitions, wherein each partition maintains its own contents separate from the contents in other partitions. As discussed further in detail herein, samples may include patient-derived samples, such as cell or tissue samples, which may contain nucleic acids and, in some cases, related proteins. In a specific aspect, samples used in the methods described herein include formalin-fixed paraffin-embedded (FFPE) cell and tissue samples, and any other sample types in which the sample is at high risk of degradation.

[0072] As used herein, partitioning refers to a variety of different forms of vessels or containers, such as pores, tubes, micro- or nanopores, through-holes, etc. However, in a preferred aspect, partitioning can flow within a fluid stream. These containers can consist of, for example, microcapsules or microvesicles having an external barrier surrounding an internal fluid center or core, or they can be porous matrices capable of entraining and / or retaining material within their matrix. However, in a preferred aspect, these partitions can comprise droplets of aqueous fluid within a non-aqueous continuous phase such as an oil phase. Various different containers are described, for example, in U.S. Patent Application No. 13 / 966,150, filed August 13, 2013. Similarly, emulsion systems for generating stable droplets in non-aqueous or oil continuous phases are described in detail, for example, in U.S. Patent Application No. 2010-0105112. In some cases, microfluidic channel networks are particularly suitable for generating partitioning as described herein. Examples of the microfluidic devices include those detailed in Provisional U.S. Patent Application No. 61 / 977,804, filed April 4, 2014, the entire disclosure of which is incorporated herein by reference for all purposes. Alternative mechanisms may also be used to dispense individual cells, including porous membranes through which an aqueous mixture of cells is extruded into a non-aqueous fluid. Such systems are typically available from, for example, Nanomi, Inc.

[0073] In the case of droplets in an emulsion, sample material is typically dispensed into discrete partitions by allowing an aqueous sample-containing stream to flow into a junction, into which a non-aqueous stream of the dispensing fluid, such as fluorinated oil, also flows, causing aqueous droplets containing sample material to form within the dispensing fluid. As described below, partitions such as droplets typically also include co-partitioned barcode oligonucleotides. The relative amount of sample material within any particular partition can be adjusted by controlling a variety of different parameters of the system, including, for example, the concentration of the sample in the aqueous stream, the flow rate of the aqueous stream and / or the non-aqueous stream, etc. The partitions described herein are typically characterized by extremely small volumes. For example, in the case of droplet-based partitioning, the droplets can have a total volume of less than 1000 pL, less than 900 pL, less than 800 pL, less than 700 pL, less than 600 pL, less than 500 pL, less than 400 pL, less than 300 pL, less than 200 pL, less than 100 pL, less than 50 pL, less than 20 pL, less than 10 pL, or even less than 1 pL. In the case of co-partitioning with beads, as will be understood, the sample fluid volume within the partition can be less than 90%, less than 80%, less than 70%, less than 60%, less than 50%, less than 40%, less than 30%, less than 20%, or even less than 10% of the aforementioned volumes. In some cases, using low reaction volume partitioning is particularly advantageous when reacting with very small amounts of starting reagents, such as input nucleic acids. Methods and systems for analyzing samples with low-input nucleic acids are proposed in U.S. Provisional Patent Application No. 62 / 017,580 (Attorney’s Case No. 43487-727.101), filed June 26, 2014, the entire disclosure of which is incorporated herein by reference.

[0074] In cases involving samples that have undergone degradation and / or contain low concentrations of the component of interest, the samples may be further processed prior to dispensing or within partitions to further release nucleic acids and / or any associated proteins for further analysis. For example, nucleic acids contained in FFPE samples are typically extracted using methods known in the art. To isolate longer nucleic acid molecules, the samples may also be processed by adding an organic catalyst to remove formaldehyde adducts (see, for example, Karmakar et al., (2015), Nature Chemistry, DOI:10.1038 / NCHEM.2307, which is incorporated herein by reference, in whole and in particular, all teachings relating to the handling and processing of FFPE samples).

[0075] Once a sample is introduced into its respective partition, the nucleic acids within that partition can be amplified to increase the amount of nucleic acid available for subsequent applications (such as sequencing methods described herein and known in the art). In some embodiments, this amplification is performed using primer libraries targeting different portions of the genome sequence, such that the resulting amplified products represent sequences from sub-parts of the original nucleic acid molecule. In embodiments focusing on selected genomic regions, the amplification may include one or more rounds of selective amplification such that the region of interest is present at a higher proportion compared to other regions of the genome (although, as will be understood, those other regions of the genome may also be amplified, but to a lesser extent, as they are not of interest for de novo coverage). In some embodiments, the amplification provides at least 1X, 2X, 5X, 10X, 20X, 30X, 40X, or 50X coverage of the entire or selected regions of the genome. In a further implementation, all nucleic acids within the partition are amplified, but selected genomic regions are amplified in a targeted manner such that at least 1-5, 2-10, 3-15, 4-20, 5-25, 6-30, 7-35, 8-40, 9-45, or 10-50 times more amplicones are generated from these selected genomic regions compared to other parts of the genome.

[0076] Simultaneously or subsequently with the amplification described above, nucleic acids (or fragments thereof) within a partition are provided with unique identifiers, allowing them to be attributed to their respective origins when characterizing those nucleic acids. Therefore, sample nucleic acids are typically co-assigned with unique identifiers (e.g., barcode sequences). In a particularly preferred aspect, the unique identifiers are provided in the form of oligonucleotides containing nucleic acid barcode sequences that can be attached to those samples. Oligonucleotides are partitioned such that the nucleic acid barcode sequences contained within them are identical between oligonucleotides in a given partition, but between different partitions, the oligonucleotides may and preferably have different barcode sequences. In an exemplary aspect, only one nucleic acid barcode sequence will be associated with a given partition, although in some cases, two or more different barcode sequences may exist.

[0077] Nucleic acid barcode sequences typically comprise 6 to 20 or more nucleotides within the sequence of oligonucleotides. These nucleotides can be completely continuous, i.e., within a single segment of adjacent nucleotides, or they can be separated into two or more separate subsequences by one or more nucleotides. Typically, the length of the separated subsequences can be from about 4 to about 16 nucleotides.

[0078] Co-assigned oligonucleotides often also contain other functional sequences that can be used to process the assigned nucleic acids. These sequences include, for example, targeted or random / universal amplification primer sequences for amplifying genomic DNA from individual nucleic acids within the partition, while attaching associated barcode sequences, sequencing primers, hybridization or probe sequences, such as those for identifying the presence of sequences or for pulling down barcodes of nucleic acids, or any of many other potential functional sequences. Furthermore, the co-assignment of oligonucleotides and associated barcodes and other functional sequences, along with the sample material, is described, for example, in USSN 14 / 175,935; 14 / 316,383; 14 / 316,398; 14 / 316,416; 14 / 316,431; 14 / 316,447; and 14 / 316,463, which are incorporated herein by reference, in their entirety and particularly for all written descriptions, figures, and working examples relating to the processing of nucleic acids and the sequencing and other characterization of genomic material, for all purposes.

[0079] In summary, in one exemplary method, beads are provided, each bead comprising a large number of the aforementioned oligonucleotides releasably attached to the bead. All oligonucleotides attached to a particular bead may comprise the same nucleic acid barcode sequence, but the number of diverse barcode sequences may span a representative bead population. Typically, the bead population can provide a diverse barcode sequence library comprising at least 1,000 different barcode sequences, at least 10,000 different barcode sequences, at least 100,000 different barcode sequences, or in some cases, at least 1,000,000 different barcode sequences. Additionally, each bead typically provides a large number of attached oligonucleotide molecules. Specifically, the number of oligonucleotide molecules, including the barcode sequence on individual beads, can be at least about 10,000 oligonucleotides, at least 100,000 oligonucleotide molecules, at least 1,000,000 oligonucleotide molecules, at least 100,000,000 oligonucleotide molecules, and in some cases at least 1 billion oligonucleotide molecules.

[0080] Oligonucleotides can be released from beads when specific stimuli are applied. In some cases, the stimulus can be light stimulation, such as by breaking light-labile bonds, which can release oligonucleotides. In other cases, thermal stimulation can be used, where an increase in temperature in the bead environment can lead to bond cleavage or other releases of oligonucleotides from the beads. In still others, chemical stimulation can be used to break the bonds between the oligonucleotides and the beads, or other methods can induce the release of oligonucleotides from the beads.

[0081] According to the methods and systems described herein, beads including attached oligonucleotides can be co-dispensed with individual samples, such that a single bead and a single sample are contained within a single partition. In some cases, where single-bead partitioning is desired, it may be necessary to control the relative flow rate of the fluid so that, on average, the partition contains less than 1 bead / partition, ensuring that those occupied partitions are predominantly single-occupied. Similarly, it may be desirable to control the flow rate to provide a higher percentage of partition occupancy, e.g., allowing only a small percentage of unoccupied partitions. In a preferred aspect, flow and channel structure are controlled to ensure a desired number of single-occupied partitions, said number being less than a certain level of unoccupied partitions and less than a certain level of multi-occupied partitions.

[0082] Figure 3 The illustration depicts a specific exemplary method for barcoding and subsequent sequencing of nucleic acids in a sample. First, a sample containing nucleic acids, 300, can be obtained from a source, and a set of barcoded beads, 310, can also be obtained. The beads are preferably linked to oligonucleotides containing one or more barcoded sequences, as well as primers such as random N-mers or other primers. Preferably, the barcoded sequence can be released from the barcoded beads, for example, by breaking the bonds between the barcode and the beads, or by degrading the beads below to release the barcode, or a combination of both. For example, in some preferred aspects, the barcoded beads can be degraded or dissolved by a reagent such as a reducing agent to release the barcoded sequence. In this example, a small sample containing nucleic acids 305, barcoded beads 315, and optionally other reagents such as a reducing agent 320 is combined and dispensed. For example, the dispensing may include introducing the components into a droplet generation system, such as a microfluidic device 325. With the aid of a microfluidic device 325, a water-in-oil emulsion 330 can be formed, wherein the emulsion contains aqueous droplets comprising sample nucleic acids 305, a reducing agent 320, and barcode beads 315. The reducing agent can dissolve or degrade the barcode beads, thereby releasing barcode-containing oligonucleotides and random N-mers 335 from the beads within the droplets. The random N-mers can then trigger different regions of the sample nucleic acid, amplifying to produce amplified copies of the sample, each copy being labeled with a barcode sequence 340. Preferably, each droplet contains a set of oligonucleotides containing the same barcode sequence and different random N-mer sequences. Subsequently, the emulsion is broken 345 and additional sequences (e.g., sequences that facilitate a particular sequencing method, additional barcodes, etc.) 350 can be added via, for example, an amplification method (e.g., PCR). Sequencing 355 can then be performed, and algorithms are applied to interpret the sequencing data 360. Sequencing algorithms are typically capable, for example, of analyzing the barcodes to compare sequencing reads and / or identifying the sample to which a specific sequence read belongs. Furthermore, and as described in this paper, these algorithms can also be used to assign copied sequences to their original molecular background.

[0083] As will be understood, samples can be amplified according to any of the methods described herein to provide coverage of the whole genome or selected regions of the genome before or simultaneously with barcoding sequence labeling 340. For implementations where targeted coverage is desired, targeted amplification typically results in a larger population of amplicon sequences representing nucleic acids (or portions thereof) in partitions containing those selected regions of the genome compared to amplicon sequences from other regions of the genome. As a result, a larger number of amplified copies 340 containing the barcoded sequence will be present within partitions of the selected regions of the genome compared to other regions of the genome. In implementations where whole-genome amplification is desired, amplification can be performed using primer libraries designed to minimize amplification bias and provide robust horizontal coverage across the entire genome.

[0084] As noted above, while single occupancy may be the most desirable state, it should be understood that multiple occupancy partitions or unoccupied partitions are often possible. Examples of microfluidic channel structures used for co-distributing samples containing barcode oligonucleotides and beads are provided in... Figure 4 The diagram is schematically illustrated. As shown, fluidly communicating channel sections 402, 404, 406, 408, and 410 are provided at channel junction 412. An aqueous stream containing a single sample 414 flows through channel section 402 to channel junction 412. As described elsewhere herein, these samples can be suspended in the aqueous fluid prior to the dispensing process.

[0085] Simultaneously, an aqueous stream containing barcode-carrying beads 416 flows through channel section 404 to channel junction 412. A non-aqueous dispensing fluid is introduced into channel junction 412 from each side channel 406 and 408, and the combined stream flows into outlet channel 410. Within channel junction 412, the two combined aqueous streams from channel sections 402 and 404 are combined and dispensed into droplets 418 comprising the co-dispensed sample 414 and beads 416. As previously noted, by controlling the flow characteristics of each fluid combined in channel junction 412, and by controlling the geometry of channel junction 412, the combination and dispensing can be optimized to achieve the desired occupancy level of beads, sample, or both within the resulting partition 418.

[0086] As will be understood, many other reagents may be co-distributed with the sample and beads, including, for example, chemical stimuli, nucleic acid extension, transcription and / or amplification reagents such as polymerases, reverse transcriptases, nucleoside triphosphates or NTP analogs, primer sequences and additional cofactors, such as divalent metal ions for the reaction, ligation reaction reagents such as ligases and ligation sequences, dyes, labels or other labeling reagents. Primer sequences may include random primer sequences targeting selected regions of the amplified genome or targeted PCR primers or combinations thereof.

[0087] Once co-distributed, the oligonucleotides placed on the beads can be used for barcoding and amplification of the distributed sample. Particularly elegant methods for using these barcoded oligonucleotides in amplifying and barcoding samples are described in detail in USSN 14 / 175,935; 14 / 316,383; 14 / 316,398; 14 / 316,416; 14 / 316,431; 14 / 316,447; and 14 / 316,463, the entire disclosure of which is incorporated herein by reference. In short, on one hand, the oligonucleotides present on the beads are co-distributed with the sample and released from their beads into the partition containing the sample. The oligonucleotides typically include a primer sequence at their 5' end along with the barcoded sequence. The primer sequences can be random or structured. Random primer sequences are generally designed to randomly prime many different regions of the sample. Structured primer sequences can include a range of different structures, including defined sequences upstream of a specific target region of a sample, and primers having some partially defined structure. These primers include, but are not limited to, primers containing a percentage of specific bases (e.g., a percentage of GC N-mers), primers containing partially or fully degenerate sequences, and / or primers containing partially random and partially structured sequences as described herein. As will be understood, any one or more of the above-described types of random and structured primers can be included in oligonucleotides in any combination.

[0088] Once released, the primer portion of the oligonucleotide can be annealed to a complementary region of the sample. Extension reaction reagents, such as DNA polymerase, nucleoside triphosphate, and cofactors (e.g., Mg2+ or Mn2+), co-distributed with the sample and beads, are then used to extend the primer sequences using the sample as a template to produce fragments complementary to the template strand annealed by the primers. These complementary fragments comprise the oligonucleotide and its associated barcode sequence. Annealing and extending multiple primers to different portions of the sample can produce large aggregates of overlapping complementary fragments of the sample, each having its own barcode sequence indicating the partition from which it was generated. In some cases, these complementary fragments themselves can serve as templates initiated by the oligonucleotide present in the partitions to generate complement, which again comprises the barcode sequence. In some cases, the replication process is constructed such that when the first complement is replicated, it generates two complementary sequences at or near its ends to allow the formation of hairpin or partial hairpin structures, which reduces the molecule's ability to become the basis for generating additional iterative copies. Figure 5 The diagram shows one example of it.

[0089] As shown in the figure, an oligonucleotide including a barcode sequence is co-distributed with sample nucleic acid 504 in droplets 502, such as an emulsion. As noted elsewhere herein, oligonucleotide 508 can be provided on beads 506 co-distributed with sample nucleic acid 504, as shown in page A, and the oligonucleotide is preferably released from beads 506. In addition to one or more functional sequences such as sequences 510, 514, and 516, oligonucleotide 508 also includes a barcode sequence 512. For example, oligonucleotide 508 is shown as containing barcode sequence 512 as well as sequence 510, which can serve as an attachment or fixation sequence for a given sequencing system, such as the P5 sequence used for attachment in flow cells of an Illumina HiSeq or Miseq system. As shown, the oligonucleotide also includes a primer sequence 516, which can include random or targeted N-mers for initiating partial replication of sample nucleic acid 504. Oligonucleotide 508 also includes sequence 514, which can provide a sequencing initiation region, such as a "read 1" or R1 initiation region, for initiating polymerase-mediated template-directed sequencing via a synthesis reaction in a sequencing system. In many cases, barcode sequence 512, fixed sequence 510, and R1 sequence 514 can be common to all oligonucleotides attached to a given bead. Primer sequence 516 can be different for random N-mer primers, or for certain targeted applications, it can be common to oligonucleotides on a given bead.

[0090] Based on the presence of primer sequence 516, the oligonucleotide can initiate the sample nucleic acid as shown in page B, which allows for the extension of oligonucleotides 508 and 508a using polymerase and other extension reagents also co-distributed with bead 506 and sample nucleic acid 504. As shown in page C, after oligonucleotide extension, for random N-mer primers, annealing is performed to multiple different regions of sample nucleic acid 504; producing multiple overlapping complements or fragments of the nucleic acid, such as fragments 518 and 520. Although including sequence portions complementary to parts of the sample nucleic acid, such as sequences 522 and 524, these constructs are generally referred to herein as fragments of sample nucleic acid 504 containing an attached barcode sequence. As should be understood, a replicated portion of the template sequence as described above is generally referred to herein as a “fraction” of that template sequence. However, nevertheless, the term “fraction” encompasses any representation of a portion of the original nucleic acid sequence, such as a template or sample nucleic acid, including those produced by other mechanisms that provide portions of the template sequence, such as actual fragmentation of a given sequence molecule, achieved, for example, by enzymatic, chemical, or mechanical fragmentation. However, in a preferred embodiment, the template or sample nucleic acid sequence fragment will represent a copy of the potential sequence or its complement.

[0091] The barcoded nucleic acid fragments can then be characterized, for example, by sequence analysis, or they can be further amplified in the process, as shown in panel D. For example, additional oligonucleotides, such as oligonucleotide 508b, are also released from bead 506 and can trigger fragments 518 and 520. Specifically, again, based on the presence of the random N-mer primer 516b in oligonucleotide 508b (which in many cases will differ from other random N-mers in a given partition, such as primer sequence 516), the oligonucleotide is annealed with fragment 518 and extended to produce complement 526 containing at least a portion of fragment 518 containing sequence 528, which contains a copy of a portion of the sample nucleic acid sequence. Oligonucleotide 508b continues to extend until it has been replicated through the oligonucleotide portion 508 of fragment 518. As noted elsewhere herein, and as illustrated in page D, the oligonucleotide can be configured to terminate replication by polymerase at a desired point, for example, after replication by sequences 516 and 514 of the oligonucleotide 508 included in fragment 518. As described herein, this can be accomplished by various methods, including, for example, incorporating different nucleotides and / or nucleotide analogs that cannot be processed by the polymerase used. For example, this could include including uracil-containing nucleotides within sequence region 512 to prevent non-uracil-resistant polymerases from stopping replication of that region. As a result, fragment 526 is produced comprising, at one end, the full-length oligonucleotide 508b, including barcode sequence 512, attachment sequence 510, R1 primer region 514, and random N-mer sequence 516b. At the other end of the sequence will be complement 516' of the random N-mer of the first oligonucleotide 508 and complement of all or part of the R1 sequence, as shown in sequence 514'. Then, the R1 sequence 514 and its complement 514' can hybridize together to form a partial hairpin structure 528. As will be understood, because the random N-mers differ in different oligonucleotides, it is expected that these sequences and their complements will not participate in hairpin formation; for example, it is expected that the sequence 516', which is the complement of the random N-mer 516, will not be complementary to the random N-mer sequence 516b. This is not the case for other applications, such as targeting primers, where the N-mers are common to the oligonucleotides within a given partition. By forming these partial hairpin structures, it allows the removal of first-order repeats of the sample sequence from further replication, for example, preventing iterative copying. Partial hairpin structures also provide useful structures for subsequent processing of fragments, such as fragment 526.

[0092] Fragments from multiple different partitions can then be aggregated for sequencing on a high-throughput sequencer as described in this article. Because each fragment is encoded according to its original partition, the sequence of that fragment can be attributed back to its origin based on the presence of a barcode. This is in Figure 6The diagram is illustrated schematically. As shown in one example, nucleic acid 604, derived from a first source 600 (e.g., a single chromosome, nucleic acid chain, etc.), and nucleic acid 606, derived from a different chromosome 602 or nucleic acid chain, are each partitioned together with their own set of barcode oligonucleotides as described above.

[0093] Within each partition, each nucleic acid 604 and 606 is then processed to individually provide overlapping sets of second fragments of the first fragment, such as second fragment sets 608 and 610. This processing also provides second fragments with barcode sequences that are identical for each second fragment derived from a particular first fragment. As shown, the barcode sequence for second fragment set 608 is represented by "1", while the barcode sequence for fragment set 610 is represented by "2". A diverse barcode library can be used to differentiately barcode a large number of different fragment sets. However, each second fragment set from a different first fragment does not necessarily need to be barcoded with a different barcode sequence. In fact, in many cases, multiple different first fragments can be processed simultaneously to include the same barcode sequence. A diverse barcode library is described in detail elsewhere in this paper.

[0094] Barcoded fragments, such as those from fragment sets 608 and 610, can then be aggregated for sequencing using sequences obtained, for example, through synthesis techniques from Illumina or the Ion Torrent division of Thermo Fisher, Inc. Once sequenced, reads from aggregated fragment 612 can be assigned to their respective fragment sets, such as aggregated reads 614 and 616, at least in part based on the included barcodes, and optionally and preferably in part based on the sequence of the fragment itself. Furthermore, reads can be assigned to the structural context of the relative positions of the nucleic acids from which those reads were obtained with respect to other closely spatially proximate nucleic acid molecules within the original sample. The associated reads from each fragment set are then assembled to provide an assembled sequence for each sample fragment, such as sequences 618 and 620, which can then be further assigned back to their respective original chromosomes or source nucleic acid molecules (600 and 602). Methods and systems for assembling genome sequences are described, for example, in U.S. Patent Application No. 14 / 752,773, filed June 26, 2015, the entire disclosure of which is incorporated herein by reference in its entirety and particularly for all teachings relating to genome sequence assembly.

[0095] III. Methods and compositions for preserving structural background

[0096] This disclosure provides methods, compositions, and systems for characterizing genetic material. Generally, the methods, compositions, and systems described herein provide methods for analyzing components of a sample while preserving information about the structure and molecular background of those components as they are in the sample. In other words, the descriptions herein typically relate to the spatial detection of nucleic acids in samples, including tissue samples that have been or will be fixed using methods known in the art, such as formalin-fixed paraffin-embedded samples. As will be understood, any method described in this section can be combined with any method described above in the sections entitled “Overview” and “Workflow Overview”, as well as the nucleic acid sequencing methods described in subsequent sections of this specification.

[0097] Typically, the methods disclosed herein involve identifying and / or analyzing nucleic acids in a sample, including the sample's genome, particularly the whole genome. The methods described herein provide the ability to quantitatively or qualitatively analyze the distribution, location, or expression of nucleic acid sequences (including genomic sequences) in a sample, while preserving the spatial context within the sample. The methods disclosed herein offer advantages over conventional methods for geocoding nucleic acids in samples because information about the structural context is preserved in high-throughput processing methods without the need to identify specific molecular targets (such as specific genes or other nucleic acid sequences) prior to processing the sample for sequence reads. Furthermore, small amounts of nucleic acids are required, which is particularly advantageous in samples such as FFPE samples, in which the input nucleic acids, especially DNA, are often fragmented or present at low concentrations.

[0098] Although most of the discussion here concerns nucleic acid analysis, it should be understood that the methods and systems discussed in this article can be applied to other components of the sample, including proteins and other molecules.

[0099] As discussed above, maintaining structural context (also referred to herein as maintaining geographic context and encoding geography) means using methods that allow the acquisition of multiple sequence reads or sequences that can be attributed to the original three-dimensional relative positions of those reads within a sample. In other words, a sequence read can be associated with its relative position within the sample relative to adjacent nucleic acids (and in some cases, related proteins) in that sample. This spatial information can be obtained even if those adjacent nucleic acids are not physically located within the linear sequence of a single original nucleic acid molecule.

[0100] Generally, the method described herein involves providing analysis of samples containing nucleic acids, wherein the nucleic acids have a three-dimensional structure. Parts of the sample are separated into discrete partitions, such that portions of the three-dimensional structure of the nucleic acids are also separated into discrete partitions—nucleic acid sequences spatially close to each other tend to be separated into the same partitions, thus preserving the three-dimensional information of spatial proximity even when subsequent sequence reads originate from sequences that were not originally on the same individual original nucleic acid molecule. Reference Figure 1 If sample 101, containing nucleic acid molecules 102, 103, and 106, is separated into discrete partitions, such that subsets of the sample are allocated to different discrete partitions, then, due to the physical distance between nucleic acid molecules 106 and 102 and 103, nucleic acid molecules 102 and 103 are more likely to be placed in the same partition than nucleic acid molecule 106. Therefore, nucleic acid molecules within the same discrete partition are those molecules that are spatially close to each other in the original sample. The sequence information obtained from nucleic acids within discrete partitions thus provides a way to analyze nucleic acids, for example, through nucleic acid sequencing analysis, and to assign those sequence reads back to the structural background of the original nucleic acid molecule.

[0101] In some instances, a tagged library is applied to a sample to spatially or geocode the sample. In some embodiments, the tag is an oligonucleotide tag (which may include “oligonucleotide barcodes” and “DNA barcodes”), but as will be understood, any type of tag that can be added to a sample can be used, including but not limited to particles, beads, dyes, molecular inverted probes (MIPs), etc. The tagged library can be applied to the sample by simple diffusion or through active processes, such as cellular processes in tissue culture or cell culture samples. Cellular transport processes include, but are not limited to, osmosis, diffusion facilitated by the involvement of cell transport proteins, passive transport, and active transport by the involvement of cell transport proteins and energy input from molecules such as ATP. Typically, tagging is applied so that different spatial / geographical locations within the sample receive different tags and / or different concentrations of tags. Any further processing of the sample and analysis of nucleic acids within the sample can be attributed to a specific spatial context by identifying the tags. For example, refer to Figure 1 Adding a tagged library to sample 101 will produce nucleic acids 102 and 103, which are spatially close to each other with different portions or concentrations of the tagged library, unlike nucleic acid 106. Any further processing of the sample according to the workflow described herein will then produce nucleic acids 102 and 103 associated with the same portion / concentration of the tagged library, and therefore the identification of these tags will indicate that nucleic acids 102 and 103 are spatially close to each other in the original sample 101. The identification of nucleic acid 106 with tags of different portions / concentrations will show that nucleic acid 106 is located in a different spatial position in the original sample than nucleic acids 102 and 103.

[0102] In a further example, partition-specific barcoding is employed, allowing any obtained sequence read to be attributed back to the partition where the original nucleic acid molecule was located. As discussed above, associating sequence reads with specific partitions identifies nucleic acid molecules that are spatially close to each other in the original sample's geographic location. For example... Figure 2 Further use of the workflows illustrated in the diagrams also provides information about the molecular background of sequence reads, allowing individual sequence reads to be attributed to the individual nucleic acid molecules from which they originated.

[0103] To enable sample labeling, samples can be processed using any method known in the art to allow the application of exogenous molecules such as oligonucleotide tags or other tags. For example, in embodiments using FFPE samples, a tag can be applied to the sample by heating it to allow the tag to embed within the sample, after which the sample can be cooled and further processed according to any of the methods described herein, including sorting into discrete partitions and further analysis to identify nucleic acid sequences in the sample and tags that are also spatially close to those sequence reads, thereby preserving the structural background of those sequence reads. Other sample processing methods include tissue processing methods that remove the extracellular matrix and / or other structural barriers while preserving molecular and protein elements. These methods include, in some non-limiting instances, the use of the CLARITY method and other tissue clearing and labeling methods, including those described, for example, in the following literature: Tomer et al., Vol. 9, No. 7, 2014, Nature Protocols; Kebschull et al., Neuron, Vol. 91, No. 5, September 7, 2016, pp. 975-987; Chung, K. et al., Structural and molecular interrogation of intrinsic biological systems. Nature 497, 332-337 (2013); Susaki, E.A. et al., Whole-brain imaging with single-cell resolution using chemical cocktails and computational analysis. Cell 157, 726-739 (2014); and Lee et al., ACT-PRESTO: Rapid and consistent tissue clearing and labeling method for 3-dimensional (3D) imaging, Scientific Reports, 2016 / 01 / 11 / online; Volume 6, page 18631, each of which is incorporated herein by reference, in whole and in particular, for all purposes, any teachings relating to the processing of samples for structural and molecular inquiry methods.

[0104] In some embodiments, the methods described herein are used in combination with imaging techniques to identify the spatial location of tags within samples, particularly those immobilized on a glass slide, such as FFPE samples. The imaging techniques can allow sequence reads to be correlated with specific locations on the slide, which allows for correlation with other pathological / imaging studies that can be performed on those samples. For example, imaging techniques can be used to provide preliminary identification of a pathology. Sequencing techniques described herein that further provide sequence reads while maintaining structural context can be combined with the imaging analysis to correlate sequence reads with structural context to confirm or provide additional information regarding the preliminary identification of a pathology. Furthermore, imaging techniques can be used in combination with tags having optical properties, such that specific tags are associated with specific regions of the imaged sample. Sequence reads associated with those identified tags can then be further correlated with regions of the imaged sample by their position relative to these tags. However, it should be understood that the methods described herein are independent of any such imaging techniques, and the ability to preserve structural context does not depend on using imaging techniques to determine the spatial information of nucleic acids in a sample.

[0105] In one exemplary aspect, an oligonucleotide gradient is generated within the sample to provide a coordinate system that can be decoded through subsequent processing via sequencing. Such a gradient would allow cells and / or nucleic acids in the sample to be labeled with oligonucleotides or oligonucleotide concentrations, which can be mapped to physical locations within the original sample. This coordinate system can be developed by diffusing an oligonucleotide library into the sample and / or by injecting oligonucleotides into specific regions of the sample. When diffusion is used, standard calculations of diffusion kinetics will provide the correlation between the concentration of the oligonucleotide tag and its spatial location within the original sample. Therefore, any other nucleic acids identified with an oligonucleotide tag of that concentration can also be correlated with a specific geographic region of the sample.

[0106] In a further exemplary embodiment, the method includes a process for analyzing nucleic acids while maintaining structural context, wherein a tag library is applied to the sample such that different geographic regions of the sample receive different tags. The sample portions now containing their original nucleic acids and the added tags are then separated into discrete partitions such that portions of the tag library and nucleic acid portions of the sample that are geographically close to each other are ultimately in the same discrete partition. Sequencing processes, such as those described in detail herein, are used to provide sequence reads of the nucleic acids in the discrete partitions. Tags may also be identified before, after, or simultaneously with those sequencing processes. The correlation between sequence reads and a specific tag (or the tag concentration in embodiments using a concentration gradient of tags) thus helps to provide spatial context for the sequence reads. As discussed above, embodiments in which tags for spatial encoding are used in combination with partition-specific barcodes further provide structural and molecular context for the sequence reads.

[0107] IV. Application Methods and Systems for Nucleic Acid Sequencing

[0108] The methods, compositions, and systems described herein are particularly suitable for nucleic acid sequencing technologies. The sequencing technologies may include any techniques known in the art, including short-read and long-read sequencing technologies. In some aspects, the methods, compositions, and systems described herein are used for high-accuracy short-read sequencing technologies.

[0109] Generally speaking, the methods and systems described in this paper utilize the advantages of extremely low sequencing error rates and high throughput of short-read sequencing technologies to perform genome sequencing. As previously mentioned, the advantage of the methods and systems described in this paper lies in their ability to obtain the desired results using ubiquitous short-read sequencing technologies. These technologies are readily available and widely distributed in the research community with well-characterized and efficient protocols and reagent systems. These short-read sequencing technologies include those available from, for example, Illumina, Inc. (GAIIx, NextSeq, MiSeq, HiSeq, X10), Thermo-Fisher's IonTorrent division (Ion Proton and Ion PGM), pyrosequencing methods, and others.

[0110] Of particular advantage is that the methods and systems described herein utilize these short-read sequencing technologies and operate with their associated low error rates. Specifically, the methods and systems described herein achieve the desired individual molecular read lengths or backgrounds as described above, but with individual sequencing reads shorter than 1000 bp, shorter than 500 bp, shorter than 300 bp, shorter than 200 bp, shorter than 150 bp, or even shorter (excluding mating pair extensions); and with sequencing error rates of less than 5%, less than 1%, less than 0.5%, less than 0.1%, less than 0.05%, less than 0.01%, less than 0.005%, or even less than 0.001% for said individual molecular read lengths.

[0111] Methods for processing and sequencing nucleic acids according to the methods and systems described in this application are also described in further detail in USSN 14 / 316,383; 14 / 316,398; 14 / 316,416; 14 / 316,431; 14 / 316,447; and 14 / 316,463, which are incorporated herein by reference, in their entirety and particularly with respect to all written descriptions, figures, and working examples relating to the processing of nucleic acids and the sequencing and other characterization of genomic material, for all purposes.

[0112] In some embodiments, the methods and systems described herein for obtaining sequence information while preserving structural and molecular background are used for whole-genome sequencing. In some embodiments, the methods described herein are used for sequencing targeted regions of the genome. In further embodiments, the sequencing methods described herein include a combination of deep coverage of selected regions with lower-level connective reads spanning a longer range across the genome. As will be understood, this combination of de novo sequencing and resequencing provides an efficient method for sequencing the entire genome and / or a large portion of the genome. Targeted coverage of poorly characterized and / or highly polymorphic regions further provides the amount of nucleic acid material required for de novo sequence assembly, while connective genome sequencing on other regions of the genome maintains high-throughput sequencing of the remainder of the genome. The methods and compositions described herein are suitable for allowing this combination of de novo sequencing and connective read sequencing because the same sequencing platform can be used for both types of coverage. The population of nucleic acids and / or nucleic acid fragments sequenced according to the methods described herein can contain sequences from both the genomic regions used for de novo sequencing and the genomic regions used for resequencing.

[0113] In specific instances, the methods described herein include the step of amplifying all or selected regions of the genome prior to sequencing. Such amplification, typically performed using methods known in the art (including, but not limited to, PCR amplification), provides at least 1X, 2X, 3X, 4X, 5X, 6X, 7X, 8X, 9X, 10X, 11X, 12X, 13X, 14X, 15X, 16X, 17X, 18X, 19X, or 20X coverage of all or selected regions of the genome. In further embodiments, the amplification provides at least 1X–30X, 2X–25X, 3X–20X, 4X–15X, or 5X–10X coverage of all or selected regions of the genome.

[0114] Amplification is typically performed by extending primers complementary to sequences within or near a selected region of the genome to cover the entire genome and / or selected target regions. In some cases, primer libraries designed to tile across the genome of interest are used—in other words, primer libraries designed to amplify regions at specific distances along the genome, whether this is across a selected region or across the entire genome. In some cases, selective amplification utilizes primers complementary to selected regions along the genome at 10, 15, 20, 25, 50, 100, 200, 250, 500, 750, 1000, or 10000 base pairs. In further instances, primer-tiled libraries are designed to capture a mixture of distances, which can be a random mixture of distances or intelligently designed such that specific portions or percentages of the selected region are amplified by different primer pairs. In a further embodiment, the primer pairs are designed such that each pair amplifies approximately 1-5%, 2-10%, 3-15%, 4-20%, 5-25%, 6-30%, 7-35%, 8-40%, 9-45%, or 10-50% of any consecutive region of a selected portion of the genome.

[0115] In some embodiments and according to any of the above descriptions, amplification occurs across genomic regions of at least 3 megabase pairs (Mb) in length. In a further embodiment, selected regions of the genome are selectively amplified according to any of the methods described herein, and said selected regions are at least 3.5, 4, 4.5, 5, 5.5, 6, 6.5, 7, 7.5, 8, 8.5, 9, 9.5, or 10 Mb in length. In yet another embodiment, the selected regions of the genome are about 2–20, 3–18, 4–16, 5–14, 6–12, or 7–10 Mb in length. Amplification can occur across these regions using a single primer pair complementary to sequences at or near the ends of these regions. In other embodiments, amplification is performed using a library of primer pairs that are laid out across the length of the region, thereby amplifying regular segments, random segments, or combinations of different segment distances along the region according to the coverage described above.

[0116] In some implementations, primers used for selectively amplifying selected regions of the genome contain uracil, thereby not amplifying the primers themselves.

[0117] Regardless of the sequencing platform used, and generally speaking, and according to any of the methods described herein, nucleic acid sequencing is typically performed in a manner that preserves the structural and molecular background of the sequence reads or portions thereof. This means that multiple sequence reads or portions thereof can be assigned to their relative spatial locations (structural background) within the original sample of other nucleic acids and / or their locations within the linear sequence of a single original molecule of nucleic acid (molecular background).

[0118] As will be understood, although a single primitive molecule of nucleic acid can be of any length in a variety of ways, in a preferred respect, it will be a relatively long molecule, thus allowing for the preservation of a long range of molecular background. Specifically, a single primitive molecule is preferably substantially longer than the typical short read sequence length, for example, longer than 200 bases, and typically at least 1,000 bases or longer, 5,000 bases or longer, 10,000 bases or longer, 20,000 bases or longer, 30,000 bases or longer, 40,000 bases or longer, 50,000 bases or longer, 60,000 bases or longer, 70,000 bases or longer, 80,000 bases or longer, 90,000 bases or longer, or 100,000 bases or longer, and in some cases 1 megabase or longer.

[0119] Typically, the method of the present invention includes, as follows: Figure 2 The illustrated steps provide a schematic diagram of the method of the invention, which is discussed in further detail herein. As will be understood, Figure 2 The methods outlined herein are exemplary implementations that can be changed or modified as needed and as described herein.

[0120] like Figure 2 As shown, the method described herein will include a sample allocation step (202) in most instances. Prior to this allocation step, an optional step (201) may be present, in which nucleic acids in the sample are joined to attach sequence regions that are spatially close to each other. Generally, each partition containing nucleic acids from the genomic region of interest will undergo some form of fragmentation, and fragments specific to the partitions containing them, typically by barcoding, will generally retain the original molecular background of the fragments (203). Each partition may include more than one nucleic acid in some instances, and in some cases will contain hundreds of nucleic acid molecules—in the case of multiple nucleic acids within a partition, any particular locus of the genome will typically be represented by a single nucleic acid prior to barcoding. As discussed above, the barcoded fragments of step 203 can be generated using any method known in the art—in some instances, oligonucleotides are used for samples within different partitions. The oligonucleotides may contain random sequences designed to randomly provoke multiple different regions of the sample, or they may contain specific primer sequences targeted upstream of a target region of the sample to provoke it. In further examples, these oligonucleotides also contain barcode sequences, allowing the replication process to also barcode the resulting replicated fragments of the original sample nucleic acid. Extension reaction reagents such as DNA polymerase, nucleosides triphosphates, and cofactors (e.g., Mg) are also included in the partitioning. 2+ or Mn 2+(etc.) The sample is then used as a template to extend the primer sequence to generate complementary fragments of the template strand annealed with the primers, and said complementary fragments comprise oligonucleotides and their associated barcode sequences. Annealing and extending multiple primers to different portions of the sample can produce large aggregates of overlapping complementary fragments of the sample, each having its own barcode sequence indicating the partition from which it was generated. In some cases, these complementary fragments themselves can serve as templates initiated by oligonucleotides present in the partitions to generate complement, which again comprises barcode sequences. In a further example, the replication process is constructed such that when the first complement is replicated, it generates two complementary sequences at or near its ends to allow the formation of hairpin or partial hairpin structures, which reduces the molecule's ability to become the basis for generating additional iterative copies.

[0121] Back Figure 2 The method illustrated herein allows for the optional aggregation of barcoded fragments (204) once partition-specific barcodes are attached to the copied fragments. The aggregated fragments are then sequenced (205), and their sequences are assigned to their original molecular background (206), thereby identifying the target region of interest and linking it to the original molecular background. The advantage of the methods and systems described herein is that attaching partition- or sample-specific barcodes to the copied fragments before enriching fragments for the target genomic region preserves the original molecular background of those target regions, allowing them to be assigned to their original partitions and thus to their original sample nucleic acid molecules.

[0122] In addition to the workflows described above, methods including chip-based and solution-based capture methods can be used to further enrich, isolate, or separate (i.e., “pull down”) target genomic regions for further analysis, particularly sequencing. These methods utilize probes complementary to the genomic region of interest or to regions near or adjacent to it. For example, in hybridization (or chip-based) capture, a microarray containing capture probes (typically single-stranded oligonucleotides) with sequences that together cover the region of interest is immobilized on a surface. Genomic DNA is fragmented and can be further processed, such as end repair, to produce blunt ends and / or to add additional features such as universal priming sequences. These fragments hybridize with probes on the microarray. Unhybridized fragments are washed away, and desired fragments are eluted or otherwise processed on the surface for sequencing or other analysis, thus enriching the residual fragment population on the surface containing fragments of the target region of interest (e.g., regions containing sequences complementary to those sequences contained in the capture probes). The enriched fragment population can be further amplified using any amplification techniques known in the art. Exemplary methods for such targeted pull-down enrichment methods are described in USSN 62 / 072,164, filed October 29, 2014, which is incorporated herein by reference, in its entirety and in particular, for all purposes, all teachings relating to targeted pull-down enrichment methods and sequencing methods, including all written descriptions, figures and embodiments.

[0123] In some instances, instead of whole-genome sequencing, the aim is to focus on selected regions of the genome. The methods described herein are particularly suitable for such analyses because the ability to target these subsets is an advantageous feature of these methods, even when genomic subsets are at large linear distances but potentially very close together in the three-dimensional background of the original sample. In some aspects, methods for covering selected regions of the genome include methods in which discrete partitions containing nucleic acid molecules and / or fragments from those selected regions are themselves classified for further processing. As will be understood, this classification of discrete partitions can be performed in any combination with other selective amplification and / or targeted pull-down methods for the genomic regions of interest described herein, particularly in any combination with the steps of the above-described workflow.

[0124] Generally, methods for classifying discrete partitions include the following steps: partitions containing at least a portion of one or more selected parts of the genome are separated from partitions that do not contain any sequences from those parts of the genome. These methods include the step of providing a population within discrete partitions containing sequences from one or more selected parts of the genome that are rich in sequences containing fragments of at least a portion of those parts of the genome. This enrichment is typically achieved by directed PCR amplification of fragments used within discrete partitions containing at least a portion of one or more selected parts of the genome. This directed PCR amplification thus produces amplicons containing at least a portion of one or more selected parts of the genome. In some embodiments, these amplicons are attached to a detectable tag, which in some non-limiting embodiments may include a fluorescent molecule. Generally, this attachment occurs such that only those amplicons generated from fragments containing one or more selected parts of the genome are attached to the detectable tag. In some embodiments, the attachment of the detectable tag occurs during selective amplification of one or more selected parts of the genome. The detectable tag in further embodiments may include, but is not limited to, fluorescent tags, electrochemical tags, magnetic beads, and nanoparticles. This attachment of the detectable tag can be accomplished using methods known in the art. In a further embodiment, discrete partitions containing at least a portion of one or more selected parts of the genome are sorted based on signals emitted from detectable markers of amplicones connected to these partitions.

[0125] In a further embodiment, the step of classifying discrete partitions containing selected portions of the genome from those discrete partitions lacking said sequences includes the following steps: (a) providing starting genomic material; (b) distributing individual nucleic acid molecules from the starting genomic material into the discrete partitions such that each discrete partition contains a first individual nucleic acid molecule; (c) providing a population of sequences rich in fragments containing at least a portion of one or more selected portions of the genome within at least some of the discrete partitions; (d) attaching a common barcode sequence to the fragments within each discrete partition such that each fragment is assigned to the discrete partition containing it; (e) separating discrete partitions containing fragments containing at least a portion of one or more selected portions of the genome from discrete partitions lacking fragments containing one or more selected portions of the genome; and (f) obtaining sequence information from the fragments containing at least a portion of one or more selected portions of the genome to sequence one or more targeted portions of the genomic sample while preserving the molecular background. As will be understood, step (a) of such a method may include more than one individual nucleic acid molecule.

[0126] In a further embodiment and according to any of the above embodiments, discrete partitions are combined and fragments are aggregated before obtaining sequence information from the fragments. In a further embodiment, the step of obtaining sequence information from the fragments is performed in a manner that maintains the structural and molecular background of the fragment sequence, such that identification also includes identifying fragments derived from nucleic acids located in close physical proximity within the original sample and / or located on the same first individual nucleic acid molecule. In a further embodiment, this acquisition of sequence information includes sequencing reactions selected from the group consisting of short-read sequencing reactions and long-read sequencing reactions. In yet another embodiment, the sequencing reaction is a short-read high-precision sequencing reaction.

[0127] In a further embodiment and according to any of the above embodiments, the discrete partitions are contained in the liquid of the emulsion. In a further embodiment, the barcoded fragments within the discrete partitions represent approximately 1X-10X coverage of one or more selected portions of the genome. In a further embodiment, the barcoded fragments within the discrete partitions represent approximately 2X-5X coverage of one or more selected portions of the genome. In yet another embodiment, the barcoded fragments of the amplicon within the discrete partitions represent at least 1X coverage of one or more selected portions of the genome. In a further embodiment, the barcoded fragments within the discrete partitions represent at least 2X or 5X coverage of one or more selected portions of the genome.

[0128] In addition to providing the ability to obtain sequence information from selected regions of the genome, the methods and systems described herein can also provide other characterizations of genomic materials, including but not limited to haplotype phasing, identification of structural variations, and identification of copy number variations, as detailed in USSN 14 / 316,383, 14 / 316,398, 14 / 316,416, 14 / 316,431, 14 / 316,447, and 14 / 316,463, which are incorporated herein by reference in their entirety for all purposes. Furthermore, all written descriptions, figures, and working examples relating to the characterization of genomic materials are incorporated herein by reference in their entirety and in particular for all purposes.

[0129] In one respect, and in conjunction with any methods described above and subsequently herein, the methods and systems described herein provide for the compartmentalization, deposition, or partitioning of sample nucleic acids or fragments thereof into discrete compartments or partitions (which are interchangeably referred to herein as partitions), wherein each partition maintains its own contents separate from the contents of other partitions. Unique identifiers such as barcodes may be delivered before, subsequently, or simultaneously to the partitions containing the compartmentalized or partitioned sample nucleic acids to allow for the subsequent attribution of features such as nucleic acid sequence information to the sample nucleic acids contained within a specific compartment, and particularly to relatively long segments of continuous sample nucleic acids that can be originally deposited into the partitions.

[0130] The sample nucleic acids used in the methods described herein typically represent multiple overlapping portions of the overall sample to be analyzed, such as entire chromosomes, exomes, or other large genomic regions. These sample nucleic acids can include whole genomes, individual chromosomes, exomes, amplicon sequences, or any of the various nucleic acids of interest. Sample nucleic acids are typically partitioned such that they exist as relatively long fragments or segments of continuous nucleic acid molecules within the partition. These fragments of sample nucleic acids can typically be longer than 1 kb, 5 kb, 10 kb, 15 kb, 20 kb, 30 kb, 40 kb, 50 kb, 60 kb, 70 kb, 80 kb, 90 kb, or even 100 kb, allowing for such a long range of molecular backgrounds.

[0131] Sample nucleic acids are typically also allocated at a certain level, thus a given partition has a very low probability of including two overlapping fragments of the starting sample nucleic acid. This is typically accomplished by providing the sample nucleic acids at low input volumes and / or concentrations during the allocation process. As a result, in a preferred embodiment, a given partition may include many long but non-overlapping fragments of the starting sample nucleic acid. Sample nucleic acids in different partitions are then associated with unique identifiers, wherein for any given partition, the nucleic acids contained therein have the same unique identifier, but different partitions may include different unique identifiers. Furthermore, since the allocation step configures the sample components into very small partitions or droplets, it should be understood that, in order to achieve the desired configuration as described above, there is no need for extensive dilution of the sample, as would be required in higher-capacity processes such as in the wells of tubes or multi-well plates. Additionally, due to the high level of barcode diversity employed in the system described herein, diverse barcodes can be configured among a large number of genomic equivalents, as provided above. Specifically, the previously described multi-well plate methods (see, for example, U.S. Publication Applications Nos. 2013-0079231 and 2013-0157870) typically operate with only a few hundred different barcode sequences and employ limiting dilution processes of their samples to be able to assign barcodes to different cells / nucleic acids. Therefore, they will typically operate with far fewer than 100 cells, which will generally provide a genome:(barcode type) ratio of approximately 1:10 and certainly far greater than 1:100. On the other hand, the system described herein can operate at a genome:(barcode type) ratio of approximately 1:50 or lower, 1:100 or lower, 1:1000 or lower, or even smaller, due to the high level of barcode diversity, such as more than 10,000, 100,000, 600,000, 700,000, etc., for a variety of barcode types. It also allows for loading a higher number of genomes (e.g., approximately more than 100 genomes per assay, more than 500 genomes per assay, 1000 genomes per assay, or even more), while still providing a significantly improved barcode diversity per genome.

[0132] Typically, the sample is combined with a set of oligonucleotide tags releasably attached to the beads prior to the dispensing step. Methods for barcoding nucleic acids are known in the art and are described herein. In some instances, methods such as those described in Amini et al. (2014, Nature Genetics, Advance Online Publication) are utilized, which are incorporated herein by reference, in their entirety and particularly, for all purposes, all teachings relating to attaching barcodes or other oligonucleotide tags to nucleic acids. In further instances, the oligonucleotides may comprise at least a first region and a second region. The first region may be a barcode region, which may be substantially the same barcode sequence among oligonucleotides within a given partition, but may be, and in most cases, different barcode sequences between different partitions. The second region may be an N-mer (random N-mer or N-mer designed to target a specific sequence) that can be used to initiate nucleic acids within the sample within a partition. In some cases where the N-mer is designed to target a specific sequence, it may be designed to target a specific chromosome (e.g., chromosome 1, 13, 18, or 21) or a region of the chromosome, such as an exon or other targeting region. As discussed in this paper, N-mers can also be designed for selected regions of the genome that tend to be poorly characterized or highly polymorphic or divergent relative to a reference sequence. In some cases, N-mers can be designed to target specific genes or gene regions, such as genes or regions associated with diseases or conditions (e.g., cancer). Within a partition, an amplification reaction can be performed using a second N-mer to initiate nucleic acid samples at different locations along the length of the nucleic acid. As a result of amplification, each partition may contain amplified products of nucleic acids attached to the same or nearly identical barcodes and representing overlapping smaller fragments of nucleic acids in each partition. The barcodes can serve as markers indicating sets of nucleic acids originating from the same partition and therefore potentially from the same strand of nucleic acid. After amplification, the nucleic acids can be pooled, sequenced, and aligned using sequencing algorithms. Since shorter sequence reads can be aligned with their associated barcode sequences and assigned to a single long fragment of the sample nucleic acid, all identified variants on that sequence can be assigned to a single original fragment and a single original chromosome. Furthermore, the chromosome contribution can be further characterized by aligning multiple colocalized variants across multiple long fragments. Therefore, conclusions regarding the phasing of specific gene variants can then be drawn, such as through analysis across long genomic sequences—for example, the identification of sequence information from poorly characterized regions across the genome. This information can also be used to identify haplotypes, which are typically a specific group of gene variants residing on the same or different nucleic acid strands. Copy number variations can also be identified in this manner.

[0133] The methods and systems described herein offer significant advantages over current nucleic acid sequencing technologies and their associated sample preparation methods. Whole-sample preparation and sequencing methods tend to primarily identify and characterize the majority of components in a sample and are not designed to identify and characterize the small percentage of components that constitute the total DNA in the extracted sample, such as genetic material contributed by a single chromosome from poorly characterized or highly polymorphic regions of the genome, material from one or a few cells, or fragmented tumor cell DNA molecules circulating in the bloodstream. The methods described herein include methods for selectively amplifying genetic material from these minority components, and the ability to preserve the molecular background of this genetic material further provides genetic characterization of these components. The methods and systems described herein also offer significant advantages for detecting populations present within larger samples. Thus, they are particularly useful for assessing haplotype and copy number variations—the methods disclosed herein can also be used to provide sequence information of sequences spatially close to each other within the three-dimensional space of the original sample or to obtain sequence information of the original nucleic acid molecules from which those sequences are derived.

[0134] The barcoding technology disclosed herein provides a unique ability to provide individual structural and molecular backgrounds for sequences and regions of the genome. These regions of the genome can include a given set of genetic markers, i.e., assigning a given set of genetic markers (as opposed to a single marker) to individual sample nucleic acid molecules, and through variant coordination assembly, to provide a broader or even longer range of inferred individual molecular backgrounds across multiple sample nucleic acid molecules and / or to a specific chromosome. These genetic markers can include specific genetic loci, such as variants, like SNPs, or they can include short sequences. Furthermore, the use of barcoding provides the added advantage of facilitating the differentiation of minority and majority components of the total nucleic acid population extracted from a sample, for example, for the detection and characterization of circulating tumor DNA in the bloodstream, and also reduces or eliminates amplification bias during optional amplification steps. Moreover, implementation in a microfluidic manner provides the ability to work with extremely small sample volumes and low input DNA amounts, as well as the ability to rapidly process large numbers of sample partitions (droplets) to facilitate whole-genome labeling.

[0135] As noted above, the methods and systems described herein provide separate structural and molecular backgrounds for short reads of longer nucleic acids. As used herein, structural background refers to the position of the sequence within the original nucleic acid molecule in three-dimensional space of the original sample. As discussed above, although genomes are generally considered linear, chromosomes are not rigid, and the spatial distance between two genomic loci is not necessarily related to their distance along the genome—megabase-separated genomic regions along a linear sequence can be directly close to each other in three-dimensional space. By preserving information about the original spatial proximity of the sequence reads, the methods and compositions described herein provide a way to assign sequence reads to long-range genomic interactions.

[0136] Similarly, the methods described herein can provide sequence background beyond a specific sequence read, such as that associated with adjacent or proximal sequences not included in the read itself, and thus generally will not be wholly or partially included in short reads such as reads of about 150 or about 300 bases used for paired reads. In a particularly preferred aspect, the methods and systems provide a long-range sequence background for short reads. This long-range background includes the relationship or association of a given sequence read with reads that are spaced apart by distances greater than 1 kb, 5 kb, 10 kb, 15 kb, 20 kb, 30 kb, 40 kb, 50 kb, 60 kb, 70 kb, 80 kb, 90 kb, or even 100 kb or longer. By providing a longer range of individual molecular backgrounds, the methods and systems of the present invention also provide a much longer inferred molecular background. Sequence background as described herein can include, for example, lower-resolution background from contigs of short sequence reads mapped to individual longer molecules or linked molecules, and higher-resolution sequence background from long-range sequencing of a large portion of a longer single molecule having a continuous defined sequence, such as a single molecule, wherein the defined sequence is longer than 1 kb, longer than 5 kb, longer than 10 kb, longer than 15 kb, longer than 20 kb, longer than 30 kb, longer than 40 kb, longer than 50 kb, longer than 60 kb, longer than 70 kb, longer than 80 kb, longer than 90 kb, or even longer than 100 kb. Similar to sequence background, attributing short sequences to longer nucleic acids (e.g., individual long nucleic acid molecules or sets of linked nucleic acid molecules or contigs) can include mapping the short sequences against the longer nucleic acid segments to provide a high level of sequence background and providing the assembled sequence from the short sequences through these longer nucleic acids.

[0137] The methods, compositions, and systems described herein allow for the characterization of long-range interactions across the genome, as well as the characterization of related proteins and other molecules within a sample. Like higher-level protein organization, the bending and folding of DNA and chromatin produce functionally important structures at various scales. On a small scale, DNA is known to often wrap around proteins such as histones to produce structures called nucleosomes. These nucleosomes package into larger “chromatin filaments,” and the packaging patterns are already involved in cellular processes such as transcription. Functional structures also exist on a much larger scale: regions separated by multi-megabase-long linear sequences of the genome can be directly adjacent in three-dimensional space. These long-range interactions between genomic loci can act on functional properties: for example, gene enhancers, silencers, and insulators can all act across vast genomic distances, and their primary modes of action can involve direct physical association with target genes, non-coding RNAs, and / or regulatory elements. Long-range interactions are not limited to elements located in the cis configuration, i.e., along the same chromosome, but can also occur between genomic loci located in the trans configuration, i.e., on different chromosomes. The presence of long-range interactions complicates efforts to understand pathways regulating cellular processes because interacting regulatory elements can be located at large genomic distances from target genes, or even on another chromosome. In the case of oncogenes and other disease-related genes, identifying long-range gene regulators could be of significant use in identifying genomic variants responsible for disease states and the processes that cause them. Therefore, the ability of the methods described herein to preserve both structural and molecular background provides a way to identify long-range genomic interactions and characterize any relevant proteins.

[0138] The methods described herein are particularly useful for characterizing nucleic acids from FFPE tissue samples, including historical FFPE tissue samples. FFPE samples often present challenges for nucleic acid characterization because nucleic acids are frequently fragmented or otherwise degraded, limiting the amount of information that can be obtained using conventional methods. The structural and molecular background information preserved in the methods described herein provides a unique opportunity for these samples, as this background information can provide characterization of long-range genomic interactions, even for degraded samples, since long-range information can be obtained using short-read sequencing techniques. Applications of FFPE nucleic acid characterization include comparing sequences from one or more historical samples with sequences from samples from subjects, such as cancer patients, to provide diagnostic or prognostic information. For example, the status of one or more molecular markers in historical samples can be correlated with one or more treatment outcomes, and the correlation between treatment outcomes and the status of one or more molecular markers in historical samples can be used to predict treatment outcomes for subjects, such as cancer patients. These predictions can serve as a basis for determining whether to recommend drug treatment options to subjects.

[0139] V. Sample

[0140] As will be understood, the methods and systems discussed herein can be used to obtain sequence information from any type of genomic material. This genomic material can be obtained from samples taken from patients. Exemplary samples and types of genomic materials used in the methods and systems discussed herein include, but are not limited to, polynucleotides, nucleic acids, oligonucleotides, cell-free nucleic acids, circulating tumor cells (CTCs), nucleic acid fragments, nucleotides, DNA, RNA, peptide polynucleotides, complementary DNA (cDNA), double-stranded DNA (dsDNA), single-stranded DNA (ssDNA), plasmid DNA, coplasmal DNA, chromosomal DNA, genomic DNA (gDNA), viral DNA, bacterial DNA, mtDNA (mitochondrial DNA), ribosomal RNA, cell-free DNA, cell-free fetal DNA (cffDNA), mRNA, rRNA, tRNA, nRNA, siRNA, snRNA, snoRNA, scaRNA, microRNA, dsRNA, viral RNA, and so on. In short, the samples used can vary depending on specific processing requirements.

[0141] In certain aspects, samples used in this invention include formalin-fixed paraffin-embedded (FFPE) cell and tissue samples, and any other sample type in which the sample has a high risk of degradation. Other types of fixed samples include, but are not limited to, samples fixed using the following: acrolein, glyoxal, osmium tetroxide, carbodiimide, mercuric chloride, zinc salts, picric acid, potassium dichromate, ethanol, methanol, acetone, and / or acetic acid.

[0142] In a further embodiment, the sample used in the methods and systems described herein includes a nuclear matrix. "Nuclear matrix" refers to any composition comprising nucleic acids and proteins. Nucleic acids can be organized into chromosomes, where proteins (i.e., histones, for example) can be associated with chromosomes that have regulatory functions.

[0143] The methods and systems provided herein are particularly applicable to nucleic acid sequencing applications, wherein the starting nucleic acid (e.g., DNA, mRNA, etc.) – or the starting target nucleic acid – is present in small amounts, or wherein it is present in a relatively low proportion of the total nucleic acid in the sample for analysis of the target nucleic acid. In one aspect, this disclosure provides methods for analyzing nucleic acids wherein the input nucleic acid molecule is present in an amount less than 50 nanograms (ng). In a further embodiment, the input amount of the nucleic acid molecule is less than 40 ng. In some embodiments, the amount is less than 20 ng. In some embodiments, the amount is less than 10 ng. In some embodiments, the amount is less than 5 ng. In some embodiments, the amount is less than 1 ng. In some embodiments, the amount is less than 0.1 ng. Methods for isolating and analyzing nucleic acids with small initial input amounts are further described, for example, in USSN 14 / 752,602, filed June 26, 2015, which is incorporated herein by reference, in its entirety and particularly for the isolation and characterization of nucleic acids obtained from samples in which small amounts of nucleic acid are present.

[0144] As will be understood, samples may be processed at any point during the description of the methods described herein using methods known in the art. For example, samples may be processed before assignment or after they have been assigned to discrete partitions.

[0145] In some embodiments, the sample is processed to ensure the retention of longer nucleic acid chains. In embodiments using FFPE samples, the sample may be processed to remove formaldehyde adducts, thereby increasing nucleic acid yield. Such processing methods may, in a non-limiting example, include the use of water-soluble organic catalysts to accelerate the inversion of formaldehyde adducts from RNA and DNA bases, as described in Karmakar et al., (2015), Nature Chemistry, DOI:10.1038 / NCHEM.2307, which is incorporated herein by reference in its entirety and particularly with respect to all teachings relating to the handling and processing of FFPE samples.

[0146] Any substance containing nucleic acids can be a source of the sample. The substance can be a fluid, such as a biological fluid. Fluid substances can include, but are not limited to, blood, umbilical cord blood, saliva, urine, sweat, serum, semen, vaginal fluid, gastric and digestive juices, cerebrospinal fluid, placental fluid, cavity fluid, eye discharge, serum, breast milk, lymph, or combinations thereof. The substance can be a solid, such as biological tissue. The substance can include normal healthy tissue, diseased tissue, or a mixture of healthy and diseased tissue. In some cases, the substance can contain a tumor. The tumor can be benign (non-cancerous) or malignant (cancerous). Non-limiting examples of tumors may include: fibrosarcoma, myxosarcoma, liposarcoma, chondrosarcoma, osteosarcoma, chordoma, angiosarcoma, endothelial sarcoma, lymphangiosarcoma, lymphangioendothelial sarcoma, synovium, mesothelioma, Ewing's sarcoma, leiomyosarcoma, rhabdomyosarcoma, gastrointestinal cancers, colon cancer, pancreatic cancer, breast cancer, genitourinary cancers, ovarian cancer, prostate cancer, squamous cell carcinoma, basal cell carcinoma, adenocarcinoma, sweat gland carcinoma, sebaceous gland carcinoma, papillary carcinoma, papillary adenocarcinoma, cystic adenocarcinoma, medullary carcinoma. Bronchial carcinoma, renal cell carcinoma, cholangiocarcinoma, choriocarcinoma, seminoma, embryonal carcinoma, Wilms' tumor, cervical cancer, endocrine system cancer, testicular tumor, lung cancer, small cell lung cancer, non-small cell lung cancer, bladder cancer, epithelial carcinoma, glioma, astrocytoma, medulloblastoma, craniopharyngioma, ependymoma, pineal tumor, hemangioblastoma, acoustic neuroma, oligodendroglioma, meningioma, melanoma, neuroblastoma, retinoblastoma, or combinations thereof. The substance can be associated with various types of organs. Non-limiting examples of organs may include the brain, liver, lung, kidney, prostate, ovary, spleen, lymph nodes (including tonsils), thyroid gland, pancreas, heart, skeletal muscle, intestine, larynx, esophagus, stomach, or combinations thereof. In some cases, the substance may include a variety of cells, including but not limited to: eukaryotic cells, prokaryotic cells, fungal cells, heart cells, lung cells, kidney cells, hepatocytes, pancreatic cells, germ cells, stem cells, induced pluripotent stem cells, gastrointestinal cells, blood cells, cancer cells, bacterial cells, bacterial cells isolated from human microbiome samples, etc. In some cases, the substance may contain cellular contents, such as, for example, the contents of a single cell or the contents of multiple cells. Methods and systems for analyzing individual cells are provided, for example, in USSN 14 / 752,641, filed June 26, 2015, the entire disclosure of which is incorporated herein by reference.

[0147] Samples can be obtained from a variety of subjects. Subjects can be live or deceased. Examples of subjects may include, but are not limited to, humans, mammals, non-human mammals, rodents, amphibians, reptiles, canines, felines, bovines, equines, goats, sheep, hens, mice, rabbits, insects, slugs, microorganisms, bacteria, parasites, or fish. In some cases, subjects may be patients with a disease or condition, suspected of having a disease or condition, or at risk of developing a disease or condition. In some cases, subjects may be pregnant women. In some cases, subjects may be normally healthy pregnant women. In some cases, subjects may be pregnant women who may be at risk of carrying a baby with certain birth defects.

[0148] Samples can be obtained from the subject by any means known in the art. For example, samples can be obtained from the subject by entering the circulatory system (e.g., via a syringe or other device into a vein or artery), collecting secreted biological samples (e.g., saliva, sputum, urine, feces, etc.), surgery (e.g., biopsy), collecting biological samples (e.g., intraoperative samples, postoperative samples, etc.), swabbing (e.g., oral swabs, oropharyngeal swabs), or pipetting.

[0149] VI. Implementation Plan

[0150] In some aspects, this disclosure provides methods for analyzing nucleic acids while maintaining a structural background. Such methods include the steps of: (a) providing a sample containing nucleic acids, wherein the nucleic acids comprise a three-dimensional structure; (b) separating portions of the sample into discrete partitions such that portions of the three-dimensional structure of the nucleic acids are also separated into the discrete partitions; and (c) obtaining sequence information from the nucleic acids, thereby analyzing the nucleic acids while maintaining a structural background.

[0151] In some implementations, the sequence information from obtaining step (c) includes identifying nucleic acids that are spatially close to each other.

[0152] In any implementation, obtaining step (c) provides information about intrachromosomal and / or interchromosomal interactions between genomic loci.

[0153] In any implementation, obtaining step (c) provides information about chromosome conformation.

[0154] In any implementation, prior to separation step (b), at least some of the three-dimensional structures are processed to connect different portions of nucleic acids that are close to each other within the three-dimensional structure.

[0155] In any implementation, the sample is a formalin-fixed paraffin sample.

[0156] In any implementation, the nucleic acid is not isolated from the sample prior to separation step (b).

[0157] In any implementation, the discrete partitions contain beads.

[0158] In any implementation, the beads are gel beads.

[0159] In any implementation, prior to obtaining step (c), nucleic acids within discrete partitions are barcoded to form multiple barcoded segments, wherein each segment within a given discrete partition contains a common barcode, such that the barcode identifies the nucleic acid from the given partition.

[0160] In any implementation, obtaining step (c) includes a sequencing reaction selected from the group consisting of short read length sequencing reactions and long read length sequencing reactions.

[0161] In any implementation, the sample comprises a tumor sample.

[0162] In any implementation, the sample comprises a mixture of tumor and normal cells.

[0163] In any implementation, the sample comprises a nuclear matrix.

[0164] In any implementation, the nucleic acid includes RNA.

[0165] In any implementation, the amount of nucleic acid in the sample is less than 5 ng / ml, 10 ng / ml, 15 ng / ml, 20 ng / ml, 25 ng / ml, 30 ng / ml, 35 ng / ml, 40 ng / ml, 45 ng / ml, or 50 ng / ml.

[0166] In some aspects, this disclosure provides a method for analyzing nucleic acids while maintaining structural background, the method comprising the steps of: (a) forming ligated nucleic acids within a sample such that spatially adjacent nucleic acid segments are linked; (b) processing the ligated nucleic acids to produce a plurality of ligation products, wherein the ligation products contain portions of spatially adjacent nucleic acid segments; (c) depositing the plurality of ligation products into discrete partitions; (d) barcoding the ligation products within the discrete partitions to form a plurality of barcoded fragments, wherein each fragment within a given discrete partition contains a common barcode, thereby associating each fragment with the ligated nucleic acid from which it is derived; and (e) obtaining sequence information from the plurality of barcoded fragments to analyze nucleic acids from the sample while maintaining structural background.

[0167] In a further embodiment, processing step (b) includes blunt-end ligation under conditions favorable to intramolecular ligation, such that spatially adjacent nucleic acid segments are ligated within the same molecule.

[0168] In any implementation, conditions that facilitate intramolecular ligation include diluting the sample to reduce the concentration of nucleic acids to below 10 ng / μL.

[0169] In any implementation, the nucleic acid is not isolated from the sample prior to step (a).

[0170] In any implementation, prior to step (a), nucleic acid immunoprecipitation is performed, such that the associated DNA-binding proteins remain bound to the nucleic acid.

[0171] In any implementation, the partition contains beads.

[0172] In any implementation, the beads are gel beads.

[0173] In any implementation, the sample includes a tumor sample.

[0174] In any implementation, the sample comprises a mixture of tumor cells and normal cells.

[0175] In any implementation, the processing step includes reversing the connection after the formation of the connection product.

[0176] In any implementation, obtaining step (e) provides information about intrachromosomal and / or interchromosomal interactions between genomic loci.

[0177] In any implementation, obtaining step (e) provides information about chromosome conformation.

[0178] In any implementation, chromosome conformation is associated with disease state.

[0179] In any implementation, the processing step produces a ligation product containing nucleic acids that were initially closely spaced in the sample.

[0180] In any implementation, obtaining step (e) includes a sequencing reaction selected from the group consisting of short read length sequencing reactions and long read length sequencing reactions.

[0181] In any implementation, the sequencing reaction is a short-read, high-precision sequencing reaction.

[0182] In any implementation, the formation step (a) includes cross-linking the nucleic acids in the sample.

[0183] In any implementation, the formation step (a) generates covalent bonds between spatially adjacent nucleic acid segments.

[0184] In some aspects, this disclosure provides a method for analyzing nucleic acids while maintaining structural background, the method comprising the steps of: (a) forming ligated nucleic acids within a sample such that spatially adjacent nucleic acid segments are linked; (b) depositing the ligated nucleic acids into discrete partitions; (c) processing the ligated nucleic acids to produce a plurality of ligation products, wherein the ligation products contain portions of spatially adjacent nucleic acid segments; (d) barcoding the ligation products within the discrete partitions to form a plurality of barcoded fragments, wherein each fragment within a given discrete partition contains a common barcode, thereby associating each fragment with the ligated nucleic acid from which it is obtained; and (e) obtaining sequence information from the plurality of barcoded fragments to analyze nucleic acids from the sample while maintaining structural background.

[0185] In a further embodiment, processing step (c) includes blunt-end ligation under conditions favorable to intramolecular ligation, such that spatially adjacent nucleic acid segments are ligated within the same molecule.

[0186] In any implementation, the sample is a formalin-fixed paraffin sample.

[0187] In any implementation, the sample includes a nuclear matrix.

[0188] In any implementation, the nucleic acid includes RNA.

[0189] In any implementation, the nucleic acid is not isolated from the sample prior to step (a).

[0190] In any implementation, prior to step (a), nucleic acid immunoprecipitation is performed, such that the associated DNA-binding proteins remain bound to the nucleic acid.

[0191] In any implementation, the partition contains beads.

[0192] In any implementation, the beads are gel beads.

[0193] In any implementation, the sample includes a tumor sample.

[0194] In any implementation, the sample comprises a mixture of tumor and normal cells.

[0195] In any implementation, processing step (c) produces a ligation product containing nucleic acids that were initially closely spaced in the sample.

[0196] In any implementation, obtaining step (e) provides information about intrachromosomal and / or interchromosomal interactions between genomic loci.

[0197] In any implementation, obtaining step (e) includes a sequencing reaction selected from the group consisting of short read length sequencing reactions and long read length sequencing reactions.

[0198] In any implementation, the sequencing reaction is a short-read, high-precision sequencing reaction.

[0199] In some aspects, this disclosure provides a method for analyzing nucleic acids while maintaining structural background, the method comprising the steps of: (a) cross-linking nucleic acids within a sample to form cross-linked nucleic acids, wherein the cross-linking forms covalent bonds between spatially adjacent nucleic acid fragments; (b) depositing the cross-linked nucleic acids into discrete partitions; (c) processing the cross-linked nucleic acids to generate a plurality of ligation products, wherein the ligation products contain portions of the spatially adjacent nucleic acid segments; and (d) obtaining sequence information from the plurality of ligation products to analyze nucleic acids from the sample while maintaining structural background.

[0200] In a further embodiment, processing step (b) includes blunt-end ligation under conditions favorable to intramolecular ligation, such that spatially adjacent nucleic acid segments are ligated within the same molecule.

[0201] In any implementation, the sample is a formalin-fixed paraffin sample.

[0202] In any implementation, the sample includes a nuclear matrix.

[0203] In any implementation, the nucleic acid includes RNA.

[0204] In any implementation, the nucleic acid is not isolated from the sample prior to the crosslinking step (a).

[0205] In any implementation, the amount of nucleic acid in the sample is less than 5 ng / ml, 10 ng / ml, 15 ng / ml, 20 ng / ml, 25 ng / ml, 30 ng / ml, 35 ng / ml, 40 ng / ml, 45 ng / ml, or 50 ng / ml.

[0206] In any implementation, prior to crosslinking step (a), nucleic acid is immunoprecipitated so that the associated DNA-binding protein remains bound to the nucleic acid.

[0207] In any implementation, the connection product is associated with a barcode prior to obtaining step (d).

[0208] In any implementation, the connection products within the same partition receive a common barcode, such that the barcode identifies the connection product from a given partition.

[0209] In any implementation, step (d) includes a sequencing reaction selected from the group consisting of short read length sequencing reactions and long read length sequencing reactions.

[0210] While preferred embodiments of the invention have been shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. Numerous variations, modifications, and substitutions will now occur to those skilled in the art without departing from the invention. It should be understood that various alternatives to the embodiments of the invention described herein can be used to practice the invention. It is intended that the foregoing claims define the scope of the invention and therefore cover methods and structures within the scope of these claims, as well as their equivalents. Example

[0211] Example 1: Sample Preparation

[0212] Modify the sample preparation method to provide long DNA molecules from FFPE samples. Figure 7 The illustration shows an exemplary workflow in which modified instructions are used to prepare FFPE samples for whole-genome sequencing (WGS) and whole-exome sequencing (WES). For example, after DNA extraction, the standard thermal cycling protocol is modified at 701 to move the 98°C denaturation step from the end of each cycle to the beginning. Additionally, a 70°C hold is added at the end of each cycle for 2 minutes.

[0213] During post-cycle cleaning 702 and WES library preparation and target enrichment steps 704 and 705, 1.8X solid phase reversible immobilization (SPRI) beads, exceeding the normal protocol, were used.

[0214] Another modification includes changing the conditions during the shearing step 703, where instead of a standard ultrasonic generator with a peak incident power of 50, an ultrasonic generator with a peak incident power of approximately 450 is used.

[0215] Another modification that can be used in certain situations is to first process the FFPE sample with an organic catalyst to remove the formaldehyde adduct, as described, for example, in Karmakar et al., (2015), Nature Chemistry, DOI:10.1038 / NCHEM.2307. Such protocols involve adding 5 mM of the organic catalyst to the sample in 30 mM pH 7 Tris buffer to achieve adduct reversal. Effective organic catalysts include, but are not limited to, water-soluble bifunctional catalysts, such as the o-aminobenzoate and aminophenyl phosphate catalysts described by Karmakar et al. Adduct reversal has the effect of increasing the yield of nucleic acids generated from the sample.

[0216] Example 2: Bar coding of FFPE samples

[0217] FFPE samples (which may include FFPE samples on a slide) can be labeled with DNA barcodes applied in a spatially defined pattern, such as those used in DNA microarray printing. The DNA barcode (hereinafter referred to as barcode-1) is elongated so that it does not diffuse in subsequent steps or be covalently applied to the FFPE sample. To enable barcoding of DNA for embedding within the FFPE slide, the sample is heated, and then the barcode is added. The barcode is typically a barcode library, providing different barcodes for different portions of the slide. Barcodes can also be added at different concentrations in different portions of the slide to aid in geocoding—in which case the barcode library may contain the same or different barcodes. After barcoding, the slide is cooled and then typically divided into portions by cutting, typically using methods such as laser microdissection, mechanical / acoustic means, etc. Fluoresceins or quantum dots (Qdots) can also be used instead of barcodes; however, barcoding allows for the large-scale, parallel, random encapsulation of sample portions while preserving local spatial information (e.g., tumors relative to normal cells).

[0218] The sample portion containing the barcode can then be placed into a sequencing system, including droplet-based systems such as the 10X Genomics Chromium. TM The system allows each droplet to encapsulate a single barcode portion.

[0219] Dewaxing of samples can be carried out by heating in a droplet. Paraffin is immiscible in water but soluble in some oils, so it can be easily removed from the droplet after being placed on a heating plate. Xylene can also be used in the liquid-liquid extraction process to dewax a portion of the sample and prepare its nucleic acid contents for further processing.

[0220] Other steps include decrosslinking the methylene bridges in the deparaffinized sample. For this step, specialized chemical methods can be used to remove the crosslinks, thereby allowing the contained nucleic acids to be used for any subsequent processing, including the nucleic acid barcoding, amplification, and library preparation steps discussed herein (see, for example...). Figure 2 Please note that spatially bar-coding DNA is also encapsulated within the droplet. The second bar-coding step for individual nucleic acids will be used to analyze the nucleic acids. and Barcodes (used for spatially coded samples) are used for barcoding. Sequence reads can then be pieced together to provide information, which can then be compared to the original spatial location within the sample and thus correlated with pathological data.

[0221] In an alternative version of this spatial coding workflow, the decrosslinking step is first performed within the droplet, followed by the attachment of nucleic acids (including genomic DNA and spatially encoded barcodes) from the sample to the particle or otherwise separation from the sample. The nucleic acids are then re-encapsulated and the barcoding and sequencing workflow described herein is performed, including... Figure 2 The method of drawing.

[0222] This specification provides a complete description of methods, systems, and / or structures, and their uses, in relation to embodiments of the technology currently described. While various aspects of the technology have been described above with a degree of specificity or by reference to one or more unique aspects, those skilled in the art can make numerous changes to the disclosed aspects without departing from the spirit or scope of this technology. Because many aspects can arise without departing from the spirit and scope of the technology currently described, appropriate scope exists in the appended claims below. Other aspects are therefore covered. Furthermore, it should be understood that any operation may be performed in any order unless otherwise expressly asserted or the language itself requires a particular order. All things contained in the foregoing description and shown in the accompanying drawings are intended to be illustrative of specific aspects only and not limited to the illustrated embodiments. Unless otherwise clearly apparent or explicitly stated from the context, any concentration values ​​provided herein are generally given as mixture values ​​or percentages, without regard to any transformations that occur when or after the addition of a particular component of the mixture. All disclosed references and patent documents referenced herein, if not expressly incorporated herein, are incorporated herein by reference in their entirety for all purposes. As defined in the following claims, changes in detail or structure may be made without departing from the essential elements of the technology.

Claims

1. A method of analyzing a plurality of nucleic acids in a formalin-fixed paraffin-embedded (FFPE) tissue sample, the method comprising: (a) providing a FFPE tissue sample, the FFPE tissue sample comprising a plurality of nucleic acids preserved in spatial locations of the FFPE tissue sample; (b) applying a plurality of geographic tags to the FFPE tissue sample, wherein the applying comprises applying different geographic tags or different concentrations of geographic tags to different geographic regions of the FFPE tissue sample, wherein a geographic tag comprises an oligonucleotide tag, a particle, or a dye; (c) after (b), partitioning the FFPE tissue sample comprising the applied plurality of geographic tags into discrete partitions, wherein a partition of the discrete partitions comprises a portion of nucleic acids from the plurality of nucleic acids and a geographic tag from the plurality of geographic tags, wherein each partition comprises a plurality of partition-specific tags; (d) obtaining sequencing information from the nucleic acids in the discrete partitions; and (e) identifying a signature of the geographic tags in the discrete partitions, wherein a signature of geographic tags provides the partition with information relating to the original relative spatial location of the portion of nucleic acids in the FFPE tissue sample, thereby analyzing the plurality of nucleic acids.

2. The method of claim 1, wherein a plurality of tagged nucleic acid fragments is generated in each discrete partition prior to the obtaining in (d), wherein each tagged nucleic acid fragment in each discrete partition comprises a fragment copy of the portion of nucleic acids in the partition and a partition-specific tag.

3. The method of claim 1, wherein the partition-specific tags are oligonucleotide barcodes.

4. The method of claim 2, wherein the tagged nucleic acid fragments in a discrete partition each comprise the same partition-specific tag.

5. The method of claim 2, wherein the sequence information obtained in (d) is obtained by sequencing the plurality of tagged nucleic acid fragments in each partition.

6. The method of claim 1, wherein each partition further comprises tagged particles comprising at least a portion of the plurality of partition-specific tags, wherein the partition-specific tags are releasably attached to the particles.

7. The method of claim 6, wherein the partition-specific tags each comprise an oligonucleotide barcode tag and a random n-mer oligonucleotide.

8. The method of claim 6, wherein the partition-specific tags in a discrete partition are the same and the partition-specific tags between discrete partitions are different.

9. The method of claim 7, wherein the partition-specific tags in each discrete partition comprise the same oligonucleotide barcode and different random n-mer oligonucleotides.

10. The method of claim 6, wherein the particles are beads.

11. The method of claim 10, wherein the beads are gel beads.

12. The method of claim 1, wherein the geographic tags are applied to the sample in step (b) such that different regions of the sample receive different concentrations of geographic tags.

13. The method of claim 12, wherein the signature identified in (e) is a concentration of tags in each discrete partition.

14. The method of claim 1, wherein in step (b) the geographic tags are applied to the sample such that different regions of the sample receive different geographic tags.

15. The method of claim 14, wherein the features identified in (e) are the sequences of the geographic tags in each discrete partition.

16. The method of claim 1, wherein the geographic tags comprise oligonucleotide barcodes.

17. The method of claim 1, wherein prior to (d) being obtained, the tissue sample is deparaffinized.

18. The method of claim 1, wherein the nucleic acids comprise RNA.

19. The method of claim 1, wherein the amount of nucleic acids in each discrete partition is less than 10 ng / ml.

20. The method of claim 1, wherein the amount of nucleic acids in each discrete partition is less than 1 ng / ml.

21. The method of claim 1, wherein the tissue sample is a tumor sample.

22. The method of claim 1, wherein the geographic tags are attached to particles.

23. The method of claim 1, wherein obtaining (d) comprises a sequencing reaction selected from the group consisting of a short read length sequencing reaction and a long read length sequencing reaction.

24. The method of claim 23, wherein the sequencing reaction is a short read high accuracy sequencing reaction.

25. The method of claim 1, further comprising sorting the discrete partitions according to the presence of one or more selected portions of the genome prior to (d) being obtained.

26. A system configured to analyze a plurality of nucleic acids in a formalin fixed paraffin embedded (FFPE) tissue sample, the analysis comprising: (a) providing a FFPE tissue sample, the FFPE tissue sample comprising a plurality of nucleic acids preserved in spatial locations in the FFPE tissue sample; (b) applying a plurality of geographic tags to the FFPE tissue sample, wherein the applying comprises applying different geographic tags or different concentrations of geographic tags to different geographic regions of the FFPE tissue sample, wherein the geographic tags comprise oligonucleotide tags, particles, or dyes; (c) after (b), partitioning the FFPE tissue sample comprising the applied plurality of geographic tags into discrete partitions, wherein a partition of the discrete partitions comprises a portion of the plurality of nucleic acids and a geographic tag from the plurality of geographic tags, wherein each partition comprises a plurality of partition specific tags; (d) obtaining sequencing information from the nucleic acids in the discrete partitions; and (e) identifying features of the geographic tags in the discrete partitions, wherein a feature of a geographic tag provides the partition with information relating to the original relative spatial location of the portion of nucleic acids in the FFPE tissue sample, thereby analyzing the plurality of nucleic acids.

27. The system of claim 26, wherein prior to (d) being obtained, a plurality of tagged nucleic acid fragments are generated in each discrete partition, wherein each tagged nucleic acid fragment in each discrete partition comprises a fragment copy of the portion of nucleic acids in the partition and a partition specific tag.

28. The system of claim 27, wherein the partition specific tags are oligonucleotide barcodes.

29. The system of claim 27, wherein the tagged nucleic acid fragments in a discrete partition each comprise the same partition-specific tag.

30. The system of claim 27, wherein the sequence information obtained in (d) is obtained by sequencing a plurality of tagged nucleic acid fragments in each partition.

31. The system of claim 26, wherein each partition further comprises tagged particles comprising at least a portion of the plurality of partition-specific tags, wherein the partition- specific tags are releasably attached to the particles.

32. The system of claim 31, wherein the partition-specific tags each comprise an oligonucleotide barcode tag and a random n-mer oligonucleotide.

33. The system of claim 31, wherein the partition-specific tags in a discrete partition are the same and the partition-specific tags between discrete partitions are different.

34. The system of claim 32, wherein the partition-specific tags in each discrete partition comprise the same oligonucleotide barcode and different random n-mer oligonucleotides.

35. The system of claim 31, wherein the particles are beads.

36. The system of claim 35, wherein the beads are gel beads.

37. The system of claim 26, wherein the geographical tags are applied to the sample in step (b) such that different regions of the sample receive different concentrations of the geographical tags.

38. The system of claim 37, wherein the feature identified in (e) is the concentration of the tag in each discrete partition.

39. The system of claim 26, wherein the geographical tags are applied to the sample in step (b) such that different regions of the sample receive different geographical tags.

40. The system of claim 39, wherein the feature identified in (e) is the sequence of the geographical tag in each discrete partition.

41. The system of claim 26, wherein the geographical tags comprise oligonucleotide barcodes.

42. The system of claim 26, wherein the tissue sample is deparaffinized prior to (d) is obtained.

43. The system of claim 26, wherein the nucleic acids comprise RNA.

44. The system of claim 26, wherein the amount of nucleic acids in each discrete partition is less than 10 ng / ml.

45. The system of claim 26, wherein the amount of nucleic acids in each discrete partition is less than 1 ng / ml.

46. The system of claim 26, wherein the tissue sample is a tumor sample.

47. The system of claim 26, wherein the geographical tags are attached to particles.

48. The system of claim 26, wherein obtaining (d) comprises a sequencing reaction selected from the group consisting of a short read length sequencing reaction and a long read length sequencing reaction.

49. The system of claim 48, wherein the sequencing reaction is a short read high accuracy sequencing reaction.

50. The system of claim 26, further comprising sorting the discrete partitions according to the presence of one or more selected portions of the genome prior to (d) is obtained.

Citation Information

Patent Citations

  • Fluorocarbon emulsion stabilizing surfactants

    US20100105112A1

  • Methods for obtaining a sequence

    US20130079231A1

  • Methods for obtaining a sequence

    US20130157870A1

  • Capsule array devices and methods of use

    US20140155295A1

  • Compositions and methods for sample processing

    US20140378322A1