Lipid-modified oligonucleotides and methods of using the same

Lipid-modified oligonucleotides enhance multiplexing capacity and reduce costs in single-cell RNA sequencing by embedding into cell membranes, addressing limitations in current methods and improving dataset quality.

JP2025166026APending Publication Date: 2025-11-05RGT UNIV OF CALIFORNIA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025128610
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2019-07-08
Filing Date
2025-07-31
Publication Date
2025-11-05

AI Technical Summary

Technical Problem

Current single-cell RNA sequencing methods face limitations in multiplexing capacity, efficiency, and cost due to the use of molecular barcodes attached after cell isolation, which are constrained by physical dimensions of microfluidic devices, leading to high technical noise and batch effects.

Method used

The use of lipid-modified oligonucleotides for labeling cells, allowing sample-specific information incorporation, enhancing multiplexing capacity and reducing technical noise by embedding into cell membranes, thereby improving sample throughput and reducing costs.

Benefits of technology

Enhanced sample multiplexing and reduced technical noise in single-cell RNA sequencing, making datasets more informative and cost-effective.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025166026000001_ABST
    Figure 2025166026000001_ABST
Patent Text Reader

Abstract

To provide methods and applications of single-cell barcoding and methods of nucleotide sequencing using compositions comprising lipid-modified oligonucleotides.SOLUTION: A composition comprising: (a) a first lipid-conjugated DNA oligonucleotide comprising a first lipid moiety, a first hybridization region, and a first primer region; (b) a second lipid-conjugated DNA oligonucleotide comprising a second hybridization region and a second lipid moiety, wherein the second hybridization region is a reverse complement of the first hybridization region; and (c) a third DNA oligonucleotide comprising a second primer region, a barcode region, and a capture sequence, wherein the second primer region is a reverse complement of the first primer region.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Provisional Application No. 62 / 871,702, filed July 8, 2019, which is incorporated by reference in its entirety.

[0002] Technical Field The present disclosure relates generally to methods and applications of single-cell barcoding and methods for sequencing nucleotides using compositions comprising lipid-modified oligonucleotides. [Background technology]

[0003] background Single-cell RNA sequencing has become a powerful tool for mapping transcriptional changes in cells. A major advantage of this technique is its ability to interrogate cellular diversity within a sample. All single-cell RNA sequencing protocols share a common initial step: RNA transcribed from cells is converted into cDNA. The next step is amplification by methods such as PCR and in vitro transcription (IVT). Subsequent steps, once sequencing is complete, allow for quantification of gene product expression levels. Isolation and barcoding of RNA from single cells are the first and crucial limiting steps in single-cell RNA-seq.

[0004] Recently, large-scale screening (genetic perturbation using shRNA or CRISPR) combined with single-cell RNA sequencing has been performed to understand complex biological phenomena. These have the obvious advantage of easy sample identification / barcoding via shRNA or CRISPR gRNA sequencing. While genetic perturbation techniques can be directly combined with barcoding (i.e., by adding a polyA+ barcode to the gRNA or shRNA itself), chemical / drug / patient screening without genetic manipulation has not been performed using barcoded methods that can be "read out" via scRNA-seq. See, for example, Adamson et al., Cell 167(7):1867-82 (2016); Aarts et al., Genes Dev. 31(20):2085-98 (2017); and Jaitin et al., Cell 167(7):1883-96 (2016) (Non-Patent Documents 1-3).

[0005] In single-cell RNA sequencing assays, multiplexing is traditionally achieved by attaching molecular barcodes to cDNA fragments or beads, depending on the application. This is done after cell isolation using either droplet microfluidics or microwells. To enable labeling of cDNA derived from cells isolated in individual droplets (or microwells), beads can be used with real-time PCR (RT) primers that also contain barcodes (or, in some cases, have barcodes on the beads). RT and barcoding can therefore occur in individual droplets or wells. Current methods have at least the following drawbacks: high cost, low efficiency, and low multiplexing capacity. The current sample multiplexing capacity of commercially available droplet microfluidics systems for single-cell RNA sequencing is limited to eight due to the number of separate channels used for cell emulsion and co-encapsulation with mRNA capture beads. [Prior art documents] [Non-patent literature]

[0006] [Non-Patent Document 1] Adamson et al.,Cell 167(7):1867-82(2016) [Non-patent document 2] Aarts et al.,Genes Dev.31(20):2085-98(2017) [Non-patent document 3] Jaitin et al.,Cell 167(7):1883-96(2016) Summary of the Invention

[0007] The present disclosure relates to compositions comprising oligonucleotides specifically designed to label cells, and compositions of pooled cells labeled with distinct exogenous oligonucleotide barcodes corresponding to different sample preparations (e.g., patients, perturbations, replicates of a single experiment, etc.). By incorporating sample-specific information in the form of lipid-modified exogenous oligonucleotides, sample throughput levels are no longer limited by the physical dimensions of the microfluidic device. Enhanced sample multiplexing reduces the cost of single-cell RNA sequencing, limits technical noise arising from batch effects, and makes single-cell transcriptome datasets more informative.

[0008] The present disclosure relates to compositions and methods for barcoding single cells and for RNA sequencing analysis using lipid-modified oligonucleotides on tissue sections taken from a subject's sample. In some embodiments, the present disclosure provides compositions comprising a lipid-conjugated DNA oligonucleotide comprising a lipid moiety, a barcode region, and a capture sequence. In some embodiments, the present disclosure provides compositions comprising: (a) a first lipid-conjugated DNA oligonucleotide comprising a lipid moiety and a first primer region, and (b) a second DNA oligonucleotide comprising a second primer region, a barcode region, and a capture sequence, wherein the second primer region is the reverse complement of the first primer region. In some embodiments, the present disclosure provides a composition comprising: (a) a first lipid-conjugated DNA oligonucleotide comprising a first lipid moiety, a first hybridization region, and a first primer region; (b) a second lipid-conjugated DNA oligonucleotide comprising a second hybridization region and a second lipid moiety, wherein the second hybridization region is the reverse complement of the first hybridization region; and (c) a third DNA oligonucleotide comprising a second primer region, a barcode region, and a capture sequence, wherein the second primer region is the reverse complement of the first primer region.

[0009] In some embodiments, the present disclosure provides methods for labeling a cell sample, isolating endogenous DNA from a cell sample, or sequencing a nucleic acid sequence from a cell sample, the methods comprising: (a) exposing the cell sample to one or more anchor lipid-modified oligonucleotides disclosed herein or any of the compositions of the present disclosure for a time sufficient for the anchor lipid-modified oligonucleotides to embed themselves into the cell membrane of the cell; (b) exposing the cell sample to one or more labeled oligonucleotides complementary to the anchor lipid-modified oligonucleotides, such that the anchor lipid-modified oligonucleotides embed themselves into the cell membrane of the cell; (c) exposing one or more of the labeled oligonucleotides to one or more of the anchor lipid-modified oligonucleotides for a period of time sufficient to form complementary strands of nucleic acid; (c) ligating one or more of the labeled oligonucleotides to one or more of the anchor lipid-modified oligonucleotides; and, optionally, (d) detecting the presence of one or more of the labeled oligonucleotides by detecting one or more unique nucleotide sequences corresponding to one or more of the labeled oligonucleotides; and / or (e) isolating the cells based on the presence of one or more of the labeled oligonucleotides, wherein the presence of one or more of the labeled oligonucleotides is determined by detecting one or more unique nucleotide sequences corresponding to one or more of the labeled oligonucleotides. In some embodiments, the cell sample used in the disclosed methods is a three-dimensional preparation of cells derived from a tissue sample of interest, and has a thickness of about 0.1 microns to about 99 microns. In some embodiments, the cell membrane is the plasma membrane separating the cytoplasm from the outside of the cell, or the cell membrane is the nuclear membrane.

[0010] In some embodiments, the present disclosure provides methods for sequencing nucleic acids from one or more cells of a sample from a subject, comprising: (a) dividing the one or more cells into containers in a three-dimensional preparation or slice having a thickness of about 0.1 microns to about 99 microns corresponding to a region of the sample; (b) labeling the one or more cells corresponding to the region of the sample with an oligonucleotide disclosed herein; (c) isolating the nucleic acids from the one or more cells; and (d) sequencing the nucleic acids from the one or more cells. In some embodiments, the disclosed methods further comprise (e) compiling sequence information from each of the one or more cells. In some embodiments, the disclosed methods further comprise correlating the nucleic acid sequence information and / or expression profile to the spatial location of the one or more cells within the sample or the subject.

[0011] In some embodiments, the compiling step in the disclosed method comprises creating an expression profile for each of the one or more cells corresponding to the region of the sample. In some embodiments, the labeling step in the disclosed method comprises: (w) exposing the one or more cells to one or more anchor lipid-modified oligonucleotides for a time sufficient for the anchor lipid-modified oligonucleotides to embed themselves in the cell membrane of the cell, and / or (x) exposing the one or more cells to one or more labeled oligonucleotides that are complementary to the anchor lipid-modified oligonucleotides for a time sufficient for the anchor lipid-modified oligonucleotides to form complementary strands of nucleic acid with one or more of the labeled oligonucleotides, and / or (y) ligating one or more of the labeled oligonucleotides to one or more of the anchor lipid-modified oligonucleotides, and optionally (z) detecting the presence of one or more of the labeled oligonucleotides by detecting one or more unique nucleotide sequences corresponding to one or more of the labeled oligonucleotides. In some embodiments, the step of isolating the cell in the method of the present disclosure comprises isolating the cell based on the presence of one or more labeled oligonucleotides, and the presence of one or more labeled oligonucleotides is determined by detecting one or more unique nucleotide sequences corresponding to one or more of the labeled oligonucleotides.In some embodiments, the anchor lipid-modified oligonucleotides used in the method of the present disclosure include: (a) a first lipid-conjugated DNA oligonucleotide comprising a first lipid moiety, a first hybridization region and a first primer region; (b) a second lipid-conjugated DNA oligonucleotide comprising a second hybridization region and a second lipid moiety, wherein the second hybridization region is the reverse complement of the first hybridization region; and (c) a third DNA oligonucleotide comprising a second primer region, a barcode region and a capture sequence, wherein the second primer region is the reverse complement of the first primer region.

[0012] In some embodiments, the present disclosure provides methods for spatially localizing a pattern of nucleic acid expression within a sample or tissue of a subject, the method comprising: (a) dividing one or more cells from the sample or tissue corresponding to a region of the sample into one of a plurality of containers; (b) exposing the one or more cells corresponding to the region of the sample with known oligonucleotides as disclosed herein for a time sufficient to allow incorporation of the known oligonucleotides into the one or more cells, each oligonucleotide being unique to and corresponding to one of the plurality of containers to which the one or more cells are exposed; (c) isolating nucleic acids from the one or more cells in response to the known oligonucleotides; (d) quantifying expression of nucleic acids from the one or more cells and / or sequencing the nucleic acids; (e) normalizing the expression of the nucleic acids in an expression profile; and (f) analyzing the expression profile of the one or more cells relative to the known oligonucleotides disclosed herein. , correlating with the spatial location of cells relative to the tissue or within the tissue, wherein the known oligonucleotides disclosed herein in each of the plurality of containers are independently selected from one or a combination of the following: (x) a first lipid-conjugated DNA oligonucleotide comprising a first lipid moiety, a first hybridization region, and a first primer region, and / or (y) a second lipid-conjugated DNA oligonucleotide comprising a second hybridization region and a second lipid moiety, wherein the second hybridization region is a reverse complement of the first hybridization region, and / or (z) a third DNA oligonucleotide comprising a second primer region, a barcode region, and a capture sequence, wherein the second primer region is a reverse complement of the first primer region, and wherein the dividing step comprises placing a thin slice or three-dimensional preparation of the sample having a thickness of about 0.1 to about 99 microns into one of the plurality of containers. In some embodiments, the disclosed method further comprises (g) subjecting the cells to flow cytometry.In some embodiments, the step of isolating the cells in the methods of the present disclosure comprises subjecting the cells to flow cytometry.

[0013] In some embodiments, the present disclosure provides methods for sequencing nucleic acids from one or more cells of a sample from a subject, comprising: (a) dividing the one or more cells corresponding to a region of the sample into containers; (b) labeling the one or more cells corresponding to a region of the sample with an oligonucleotide disclosed herein; (c) isolating the nucleic acids from the one or more cells; and (d) sequencing the nucleic acids from the one or more cells, wherein the dividing step comprises placing a thin slice or three-dimensional preparation of the sample having a thickness of about 0.1 to about 99 microns into one of a plurality of containers. In some embodiments, the disclosed methods further comprise (e) compiling sequence information from each of the one or more cells. In some embodiments, the disclosed methods further comprise correlating the nucleic acid sequence information and / or expression profile to the spatial location of the one or more cells within the sample or the subject.

[0014] In some embodiments, the compiling step of the disclosed method comprises creating an expression profile of each of the one or more cells corresponding to the region of the sample. In some embodiments, the labeling step of the disclosed method comprises: (w) exposing the one or more cells to one or more anchor lipid-modified oligonucleotides for a time sufficient for the anchor lipid-modified oligonucleotides to embed themselves into the cell membrane of the cells, and / or (x) exposing the one or more cells to one or more labeled oligonucleotides complementary to the anchor lipid-modified oligonucleotides for a time sufficient for the anchor lipid-modified oligonucleotides to form complementary strands of nucleic acid with one or more of the labeled oligonucleotides, and / or (y) ligating one or more of the labeled oligonucleotides to one or more of the anchor lipid-modified oligonucleotides, and optionally (z) detecting the presence of one or more of the labeled oligonucleotides by detecting one or more unique nucleotide sequences corresponding to one or more of the labeled oligonucleotides. In some embodiments, the step of isolating the nucleic acid of the one or more cells of the disclosed method comprises isolating the one or more cells based on the presence of one or more labeled oligonucleotides, wherein the presence of the one or more labeled oligonucleotides is determined by detection of one or more unique nucleotide sequences corresponding to one or more of the labeled oligonucleotides.

[0015] In some embodiments, the present disclosure provides a method for identifying spatial expression patterns of nucleic acids within a tissue of a subject, comprising: (a) dividing one or more cells from a sample corresponding to a region of the tissue into one of a plurality of containers; (b) exposing the one or more cells with a lipid-conjugated DNA oligonucleotide comprising a lipid moiety, a barcode region, and a capture sequence for a time sufficient to allow the lipid moiety to embed itself within the cellular membrane of the one or more cells, wherein the barcode region of the lipid-conjugated DNA oligonucleotide is unique to each one of the plurality of containers to which the one or more cells are exposed; (c) sequencing nucleic acids captured by the capture sequence of the lipid-conjugated DNA oligonucleotide in the one or more cells; and (d) correlating the sequenced nucleic acids from the one or more cells to the spatial location of the one or more cells within and / or relative to the tissue according to the barcode region contained in each of the sequenced nucleic acids. In some embodiments, the lipid-conjugated DNA oligonucleotide used in the disclosed method comprises: (i) a first lipid-conjugated DNA oligonucleotide comprising the lipid moiety and a first primer region; and (ii) a second DNA oligonucleotide comprising a second primer region, the barcode region, and the capture sequence, wherein the second primer region is the reverse complement of the first primer region. In some embodiments, the first lipid-conjugated DNA oligonucleotide further comprises a first hybridization region. In some embodiments, the disclosed method further comprises exposing the one or more cells to a second lipid-conjugated DNA oligonucleotide prior to the sequencing step, the second lipid-conjugated DNA oligonucleotide comprising a second hybridization region and a second lipid moiety, and the second hybridization region is the reverse complement of the first hybridization region. In some embodiments, the disclosed method further comprises (e) subjecting the cells to flow cytometry.In some embodiments, the dividing step of the disclosed method comprises placing a tissue slice or a three-dimensional preparation of the sample having a thickness of about 0.1 to about 99 microns into one of a plurality of containers. In some embodiments, one or more containers used in the disclosed method are multi-well. In some embodiments, the sample used in the disclosed method is a tissue slice.

[0016] In some embodiments, the first lipid-conjugated DNA oligonucleotide of the present disclosure, including the compositions and methods of the present disclosure, comprises, in a 5' to 3' direction, the first lipid moiety, the first hybridization region, and the first primer region. In other embodiments, the first lipid-conjugated DNA oligonucleotide of the present disclosure, including the compositions and methods of the present disclosure, comprises, in a 3' to 5' direction, the first lipid moiety, the first hybridization region, and the first primer region.

[0017] In some embodiments, the second lipid-conjugated DNA oligonucleotide of the present disclosure, including the compositions and methods of the present disclosure, comprises, in a 5' to 3' direction, the second hybridization region and the second lipid moiety. In some embodiments, the second lipid-conjugated DNA oligonucleotide of the present disclosure, including the compositions and methods of the present disclosure, comprises, in a 3' to 5' direction, the second hybridization region and the second lipid moiety.

[0018] In some embodiments, the third DNA oligonucleotide of the present disclosure, including the compositions and methods of the present disclosure, comprises, from 5' to 3', the second primer region, the barcode region, and the capture sequence. In some embodiments, the third DNA oligonucleotide of the present disclosure, including the compositions and methods of the present disclosure, comprises, from 3' to 5', the second primer region, the barcode region, and the capture sequence.

[0019] In some embodiments, the first lipid moiety used in a first lipid-conjugated DNA oligonucleotide of this disclosure, including the compositions and methods of this disclosure, comprises a fatty acid having from about 12 to about 28 carbons. In some embodiments, the second lipid moiety used in a second lipid-conjugated DNA oligonucleotide of this disclosure, including the compositions and methods of this disclosure, comprises a fatty acid having from about 12 to about 28 carbons. In some embodiments, both the first lipid moiety and the second lipid moiety used in a lipid-conjugated DNA oligonucleotide of this disclosure, including the compositions and methods of this disclosure, comprise a fatty acid having from about 12 to about 28 carbons.

[0020] In some embodiments, the first lipid moiety used in the first lipid-conjugated DNA oligonucleotides of the present disclosure, including the compositions and methods of the present disclosure, is a compound of Formula I: TIFF2025166026000002.tif40128 or a physiologically acceptable salt thereof, In the formula, n 1 is 5 to 25, and n 2 is 1-25, X is selected from the group consisting of NH, CH, O, and CH—R, and R is a C12-C28 monoglyceride, alkenyl, alkyl, aryl, or aralkyl. In some embodiments, the first lipid moiety comprises a lipid selected from lignoceric acid and cholesterol. In some embodiments, the cholesterol is cholesterol-triethylene glycol (TEG).

[0021] In some embodiments, the second lipid moiety used in the second lipid-conjugated DNA oligonucleotides of the present disclosure, including the compositions and methods of the present disclosure, is a compound of Formula II: TIFF2025166026000003.tif38128 or a physiologically acceptable salt thereof, In the formula, n 1 is 5 to 25, and n 2is 0-24, X is selected from the group consisting of NH, CH, O, and CH—R, and R is a C12-C28 monoglyceride, alkenyl, alkyl, aryl, or aralkyl. In some embodiments, the second lipid moiety comprises a lipid selected from palmitic acid and cholesterol. In some embodiments, the cholesterol is cholesterol-triethylene glycol (TEG).

[0022] In some embodiments, the capture sequence used in the third DNA oligonucleotide of the present disclosure, including the compositions and methods of the present disclosure, is a polyadenylation region.

[0023] In some embodiments, a first lipid-conjugated DNA oligonucleotide of the present disclosure, including the compositions and methods thereof, comprises a nucleic acid sequence having at least about 70% sequence identity to the nucleic acid sequence of SEQ ID NO: 1 (GTAACGATCCAGCTGTCACTTGGAATTCTCGGGTGCCAAGG). In some embodiments, a second lipid-conjugated DNA oligonucleotide of the present disclosure, including the compositions and methods thereof, comprises a nucleic acid sequence having at least about 70% sequence identity to the nucleic acid sequence of SEQ ID NO: 2 (AGTGACAGCTGGATCGTTAC). In some embodiments, a third DNA oligonucleotide of the present disclosure, including the compositions and methods thereof, comprises a nucleic acid sequence having at least 70% sequence identity to the nucleic acid sequence of SEQ ID NO: 3 (CCTTGGCACCCGAGAATTCCANNNNNNAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAA).

[0024] In some embodiments, the first lipid-conjugated DNA oligonucleotide of this disclosure, including the compositions and methods of this disclosure, is bound to a solid support. In some embodiments, the second lipid-conjugated DNA oligonucleotide of this disclosure, including the compositions and methods of this disclosure, is bound to a solid support. In some embodiments, the third lipid-conjugated DNA oligonucleotide of this disclosure, including the compositions and methods of this disclosure, is bound to a solid support. In some embodiments, one or more of the first lipid-conjugated DNA oligonucleotide, the second lipid-conjugated DNA oligonucleotide, and the third lipid-conjugated DNA oligonucleotide are bound to a solid support. In some embodiments, the solid support is a bead. In some embodiments, the capture sequence used in the third DNA oligonucleotide of this disclosure, including the compositions and methods of this disclosure, includes a polyadenylation region, and the bead includes a poly(T) region that hybridizes to the polyadenylation region of the third DNA oligonucleotide. [Brief explanation of the drawings]

[0025] [Figure 1] Schematic and example results of an experiment on adult intestine used as a representative example of the barcode classification workflow are shown. Note that these data were obtained from tissue samples greater than 1 cm thick and are only representative of the data that can be collected using the prophetic methods shown in the examples. [Figure 2] We present results from an experiment on the developing intestine, which is used as a representative example of the barcode classification workflow. Note that these data were obtained from tissue samples greater than 1 cm thick and are only representative of the data that can be collected using the prophetic methods shown in the Examples. DETAILED DESCRIPTION OF THE INVENTION

[0026] Detailed Description of the Embodiments Before describing exemplary embodiments, it is to be understood that this disclosure is not limited to particular embodiments described, as such may, of course, vary. It is also to be understood that the terminology used herein is for the purpose of describing particular embodiments only, and is not intended to be limiting, since the scope of the present disclosure will be limited only by the appended claims.

[0027] Where a range of values ​​is given, it is understood that each intervening value between the upper and lower limits of that range, to the tenth of the unit of the lower limit, unless the context clearly dictates otherwise, is also specifically disclosed. Each smaller range between any stated value or intervening value in a stated range and any other stated value or intervening value in that stated range is encompassed within this disclosure. The upper and lower limits of these smaller ranges may independently be included or excluded within the range, and each range where one, neither, or both limits are included in the smaller range is also encompassed within this disclosure, depending on any specifically excluded limits in the stated range. When a stated range includes one or both of its limits, ranges excluding either or both of those included limits are also included within this disclosure.

[0028] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs. Although any methods and materials similar or equivalent to those described herein can also be used in the practice or testing of embodiments of the present disclosure, some possible exemplary methods and materials may be described herein. Any and all publications mentioned herein are incorporated by reference in their entirety to disclose and describe the methods and / or materials in connection with which the publications are cited. To the extent there is a conflict, it is understood that the present disclosure supersedes any disclosure of the incorporated publication.

[0029] It is further noted that the claims may be drafted to exclude any optional element. Accordingly, this statement is intended to serve as a predicate for use of exclusive terminology such as "solely," "only," and the like in conjunction with the recitation of claim elements or for use of a "negative" limitation.

[0030] The publications discussed herein are provided solely for their disclosure prior to the filing date of the present application. Nothing herein should be construed as an admission that the present disclosure is not entitled to antedate such publications. Further, the dates of publication provided may be different from the actual publication dates, which may need to be independently confirmed. To the extent that such publications may provide definitions of terms that contradict express or implied definitions in the present disclosure, the definitions in the present disclosure shall control.

[0031] As will be apparent to one of skill in the art after reading this disclosure, each of the individual embodiments described and illustrated herein has distinct components and features which may be readily separated from or combined with the features of any of the other several embodiments without departing from the scope or spirit of the disclosure. Any recited method can be carried out in the order of events recited or in any other order which is logically possible.

[0032] definition Before describing the compositions and methods of the present invention, it is to be understood that this disclosure is not limited to the particular molecules, compositions, methods, or protocols described, as these may vary. It is also to be understood that the terminology used in this description is for the purpose of describing particular types or embodiments only, and is not intended to limit the scope of the disclosure, which is limited only by the appended claims. It is to be understood that these embodiments are not limited to the particular methods, protocols, cell lines, vectors, and reagents described, as these may vary. It is also to be understood that the terminology used herein is for the purpose of describing particular embodiments only, and is not intended to limit the scope of the embodiments of the present invention or the claims.

[0033] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art. Although any methods and materials similar or equivalent to those described herein can be used in practicing or testing embodiments of the present disclosure, the methods, devices, and materials are described herein. All publications mentioned herein are incorporated by reference. Nothing herein should be construed as an admission that the present disclosure is not entitled to antedate such disclosure by virtue of prior disclosure.

[0034] As used herein in the specification and claims, the indefinite articles "a" and "an" should be understood to mean "at least one," unless clearly indicated to the contrary.

[0035] As used herein in the specification and claims, the phrase "and / or" should be understood to mean "either or both" of the elements so conjunctivated, i.e., elements that are present conjunctively in some cases and disjunctively in other cases. Other elements other than those specifically identified by the "and / or" clause may optionally be present, whether related to the elements specifically identified or not, unless clearly indicated to the contrary. Thus, as a non-limiting example, when used in conjunction with open-ended language such as "comprising," a reference to "A and / or B" can refer, in one embodiment, to A without B (optionally including elements other than B); in another embodiment, to B without A (optionally including elements other than A); in yet another embodiment, to both A and B (optionally including other elements); etc.

[0036] As used herein in the specification and claims, "or" should be understood to have the same meaning as "and / or" as defined above. For example, when separating items in a list, "or" or "and / or" shall be construed as inclusive, i.e., at least one of some of the elements or list of elements, but including more than one, and optionally, additional items not listed. Only terms clearly indicated to the contrary, e.g., "only one of" or "exactly one of," or, when used in the claims, "consisting of," shall refer to the inclusion of exactly one element of some of the elements or list of elements. In general, as used herein, the term "or" shall be construed as indicating exclusive alternatives (i.e., "one or the other, but not both") only when preceded by terms of exclusivity, i.e., "any of," "one of," "only one of," or "exactly one of." "Consisting essentially of," when used in the claims, shall have its ordinary meaning as used in the field of patent law.

[0037] As used herein, the term "about" when referring to a measurable value, e.g., an amount, a time period, etc., is meant to encompass variations of ±20%, ±10%, ±5%, ±1%, or ±0.1% from the specified value, as such variations are appropriate for performing the methods of the present disclosure.

[0038] As used herein, the phrase "an integer between X and Y" means any integer, inclusive of the endpoints. That is, when a range is disclosed, each integer within the range, including the endpoints, is disclosed. For example, the phrase "an integer between X and Y" discloses 1, 2, 3, 4, or 5, as well as the range 1 to 5.

[0039] As used herein, the term "endogenous" refers to a material that originates within an organism. An "endogenous" nucleic acid therefore refers to a nucleic acid that originates or is produced within an organism, tissue, or cell.

[0040] As used herein, the term "normalizing" or "normalized" refers to the expression level of a nucleic acid relative to the average expression level of one or a set of reference nucleic acids, which are based on minimal variation across tissues or cells.

[0041] The term "labeled oligonucleotide" refers to an oligonucleotide that comprises a sequence complementary to at least a portion of an anchor lipid-modified oligonucleotide, as defined herein, such that the anchor lipid-modified oligonucleotide forms a complementary strand of nucleic acid with the labeled oligonucleotide. In some embodiments, the labeled oligonucleotide comprises a primer region, as defined herein. In some embodiments, the labeled oligonucleotide comprises a barcode region, as defined elsewhere herein. In some embodiments, the labeled oligonucleotide comprises a capture sequence, as defined herein.

[0042] The terms "lipid-modified oligonucleotide," "lipid DNA," "hydrophobic anchor oligonucleotide," and similar terms should be interpreted broadly to include any oligonucleotide or polynucleotide that is attached in any manner to a hydrophobic, lipophilic, or amphiphilic region that can insert into a membrane, regardless of whether the "lipid-modified oligonucleotide," "lipid DNA," "hydrophobic anchor oligonucleotide," or a portion thereof, actually inserts into the membrane.

[0043] The term "membrane" or any similar term is used broadly and generically herein to refer to any lipid-containing membrane, cell membrane, nuclear membrane, monolayer, bilayer, vesicle, liposome, lipid bilayer, etc., and the present disclosure is not intended to be limited to any particular membrane.

[0044] As used herein, the terms "subject," "individual," or "patient," used interchangeably, mean any animal, including a mammal, e.g., a mouse, rat, other rodent, rabbit, dog, cat, pig, cow, sheep, horse, or primate, e.g., a human.

[0045] As used herein, the term "kit" refers to a set of components provided in the context of a system for diagnosing a subject for a disease or infection based on the presence, absence, and / or quantity of nucleotide sequences expressed from a sample or cell, and / or the sequencing of nucleotides and / or the isolation of nucleotide sequences and / or the spatial location of nucleotide sequences expressed in a sample or cell. Such systems may include, for example, systems for storing, identifying, or delivering genes (e.g., oligonucleotides, enzyme-encoding oligonucleotides, extracellular matrix components, etc., in appropriate containers) and / or supporting materials (e.g., buffers, media, cells, instructions for performing the assay, etc.) expressed in one or more cells from one location to another. For example, in some embodiments, a kit includes one or more containers (e.g., boxes) containing key reaction reagents and / or supporting materials. As used herein, the term "fragment kit" refers to a diagnostic assay comprising two or more separate containers, each containing a subportion of the overall kit components. The containers can be delivered to the intended recipient together or separately. For example, a first container may contain a solid support or a polystyrene plate for use in a cell culture assay, while a second container may contain cells, such as control cells. As another example, the kit may include a first container containing a solid support, such as a chip or slide, with one or more ligands having affinity for one or more of the biomarkers disclosed herein, and a second container containing one or more reagents necessary for detecting and / or quantifying the amount of lipid-modified oligonucleotides in a sample. The term "fragmentation kit" is intended to include, but is not limited to, a kit containing an analyte-specific reagent (ASR) as defined by Section 520(e) of the Federal Food, Drug, and Cosmetic Act.Any delivery system that includes two or more separate containers, each containing a subportion of the total kit components, is included in the term "fragmented kit." In contrast, an "integrated kit" refers to a delivery system that contains all components in a single container (e.g., a single box containing each of the desired components). The term "kit" includes both fragmented kits and integrated kits.

[0046] As used herein, the term "animal" includes, but is not limited to, humans and non-human vertebrates, such as wild animals, rodents, e.g., rats, ferrets, and domestic animals, and livestock, e.g., dogs, cats, horses, pigs, cows, sheep, and goats. In some embodiments, the animal is a mammal. In some embodiments, the animal is a human. In some embodiments, the animal is a non-human mammal.

[0047] As used herein, the term "mammal" refers to any animal within the class Mammalia, for example, a rodent (i.e., a mouse, rat, or guinea pig), monkey, cat, dog, cow, horse, pig, or human. In some embodiments, the mammal is a human. In some embodiments, the mammal refers to any non-human mammal. The present disclosure relates to any of the methods or compositions disclosed herein where a sample is taken from a mammal or a non-human mammal. The present disclosure relates to any of the methods or compositions disclosed herein where a sample is taken from a human.

[0048] As used herein, the phrase "in need thereof" means that an animal or mammal has been identified or is suspected of being in need of a particular method or treatment, as indicated based on the presence, absence, and / or amount of a biomarker. In some embodiments, the identification can be by any diagnostic or observational method. The animal or mammal may be in need of any of the methods and treatments described herein. In some embodiments, the animal or mammal is in, or will be moved to, an environment where a particular disorder or condition is prevalent or likely to occur.

[0049] The specific use of the terms "nucleic acid," "oligonucleotide," and "polynucleotide" should not be considered limiting in any way and may be used interchangeably herein. "Oligonucleotide" is used when the relevant nucleic acid molecule typically contains fewer than about 100 bases. "Polynucleotide" is used when the relevant nucleic acid molecule typically contains more than about 100 bases. Both terms are used to refer to DNA, RNA, modified or synthetic DNA or RNA (including, but not limited to, nucleic acids containing synthetic and naturally occurring base analogs, dideoxy or other sugars, thiols, or other non-natural or natural polymer backbones), or other nucleobase-containing polymers capable of hybridizing to DNA and / or RNA. Thus, these terms should not be construed to define or limit the length of the nucleic acids referred to and used herein, nor should these terms be used to limit the nature of the polymer backbone to which the nucleobases are attached.

[0050] Polynucleotides of the present disclosure may be single-stranded, double-stranded, or triple-stranded, or may include combinations of these conformations. Generally, polynucleotides contain phosphodiester linkages, but in some cases, include nucleic acid analogs that may have alternative backbones, including, for example, phosphoramide, phosphorothioate, phosphorodithioate, O-methylphosphoramidite linkages, and peptide nucleic acid backbones and linkages, as outlined below. Other analog nucleic acids include morpholinos, locked nucleic acids (LNAs), and those with positive, non-ionic, and non-ribose backbones. Nucleic acids containing one or more carbocyclic sugars are also included within the definition of nucleic acids. These modifications of the ribose phosphate backbone may be made to facilitate the addition of additional moieties, such as labels, or to increase the stability and half-life of such molecules in physiological environments.

[0051] The term "nucleic acid sequence" or "polynucleotide sequence" refers to a contiguous string of nucleotide bases, and in certain contexts also refers to the specific arrangement of the nucleotide bases relative to one another as they appear in a polynucleotide.

[0052] As used herein, the terms "comprising" (and any form of comprising, e.g., "comprise," "comprises," and "comprised"), "having" (and any form of having, e.g., "have" and "has"), "including" (and any form of comprising, e.g., "includes" and "include"), or "containing" (and any form of comprising, "contains" and "contain") are inclusive or open-ended and do not exclude additional, unrecited elements or method steps.

[0053] As used herein, the term "fluorogenic probe" refers to any molecule (e.g., dye, peptide, or fluorescent marker) that emits light of a known and / or detectable wavelength upon exposure to light of a known wavelength. In some embodiments, the fluorogenic probe is a substrate or peptide with a known cleavage site recognizable by any of the enzymes expressed by one or more animals or unicellular organisms. In some embodiments, the fluorogenic probe is attached to any one or more of the oligonucleotide sequences disclosed herein. In some embodiments, the attachment of the fluorogenic probe to the oligonucleotide disclosed herein creates a chimeric molecule capable of emitting fluorescence(s) upon exposure of the substrate to the enzyme and light of the known wavelength, and exposure to the enzyme creates a reaction product that is quantifiable in the presence of a fluorometer or spectrophotometer. In some embodiments, the fluorogenic probe is fully quenched upon exposure to light of a known wavelength, after which the substrate is enzymatically cleaved and the fluorogenic probe emits light of a known wavelength, the intensity of which can be quantified by absorbance measurements or intensity levels in the presence of a fluorometer, and optionally after cleavage of the probe from the oligonucleotide to which it is attached. In some embodiments, the fluorogenic probe is a coumarin-based dye or a rhodamine-based dye, and has a measurable or quantifiable fluorescence emission spectrum in the presence of or upon exposure to light of a predetermined wavelength. In some embodiments, the fluorogenic probe comprises rhodamine. In some embodiments, the fluorogenic probe comprises rhodamine-100. Coumarin-based fluorogenic probes are known in the art, for example, in U.S. Pat. Nos. 7,625,758 and 7,863,048, which are incorporated herein by reference in their entireties. In some embodiments, the fluorogenic probe is a component of, covalently bound to, non-covalently bound to, or intercalated with one or more substrates for any of the enzymes disclosed herein. In some embodiments, the fluorogenic probe is selected from ACC or AMC. In some embodiments, the fluorogenic probe is a fluorescein molecule.In some embodiments, the fluorogenic probe is capable of emitting a resonant wave that is detectable and / or quantifiable by a fluorometer after exposure to one or more enzymes that catalyze one or more cleavages of the lipid-modified oligonucleotides disclosed herein.

[0054] As used herein, the term "score" refers to a single value or range of values ​​that can be used as a component of a predictive model for the diagnosis, prognosis, or clinical treatment plan of a subject, and the single value is calculated by combining and / or normalizing raw data values ​​with or against a control value based on the characteristic or metric measured by the system. In some embodiments, the score is calculated through an interpretation function or algorithm. In some embodiments, the subject is suspected of having the expression of a gene that promotes or contributes to the likelihood of causing a pathological condition, or whose expression is correlated with the presence of a pathogen. The calculation of the score can be achieved using a known algorithm that can be implemented in a computer program in the device used to sequence or analyze the sample. For example, when using a BioRAD ddSEQ single-cell isolator, the algorithms required to detect, quantify, or visualize probes and / or probe-labeled barcodes are described in BioRad publication 7139, entitled "Implementing the Drop-Seq Protocol on Bio-Rad's ddSEQ Single-Cell Isolator," the contents of which are incorporated by reference in their entirety, and can be found at https: / / www.bio-rad.com / en-us / product / ddseq-single-cell-isolator?ID=OKNWBSE8Z. In some embodiments, the methods disclosed herein include the substeps of detecting the presence, absence, or amount of a given barcode nucleotide by calculating the amount of probe in a control sample, calculating the amount of probe in a cell sample, and normalizing the signal obtained from the cell sample by subtracting the signal obtained from the control sample from the signal obtained from the cell sample.

[0055] To facilitate detection of the lipid-modified oligonucleotides disclosed herein, for example, a detectable substance may be pre-applied to a surface, such as a plate, well, bead, or other solid support, including one or more reaction vessels. In some embodiments, the sample may be pre-mixed with a diluent or reagent and then applied to the surface. The detectable substance may function as a lipid oligonucleotide that can be detected visually or by an instrument. Generally, any substance capable of generating a signal that can be detected visually or by an instrument may be used as a detection probe. Suitable detectable substances may include, for example, luminescent compounds (e.g., fluorescent, phosphorescent, etc.), radioactive compounds, optical compounds (e.g., colored dyes or metallic substances, such as gold), liposomes or other vesicles containing signal-generating substances, enzymes and / or substrates, etc. Other suitable detectable substances may be described in U.S. Patent No. 5,670,381 to Jou et al. and U.S. Patent No. 5,252,459 to Tarcha et al., both of which are incorporated herein by reference in their entireties for all purposes. If the detectable substance is colored, the ideal electromagnetic radiation is light of the complementary dominant wavelength. For example, a blue detection probe strongly absorbs red light. In some embodiments, the lipid-modified oligonucleotide comprises a probe. In some embodiments, the detectable probe comprises or consists of a light-emitting compound that generates an optically detectable signal corresponding to the level or amount of lipid oligonucleotide in a sample. For example, suitable fluorescent molecules may include, but are not limited to, fluorescein, europium chelate, phycobiliprotein, phycoerythrin, phycocyanin, allophycocyanin, o-phthalaldehyde, fluorescamine, rhodamine, and derivatives and analogs thereof. Another suitable fluorescent compound is a semiconductor nanocrystal, commonly referred to as a "quantum dot." For example, such a nanocrystal may comprise a core of the formula CdX, where X is Se, Te, S, etc. The nanocrystal may also be passivated with an overlying shell of the formula YZ, where Y is Cd or Zn and Z is S or Se.Other examples of suitable semiconductor nanocrystals are also described in U.S. Patent No. 6,261,779 to Barbera-Guillem, et al. and U.S. Patent No. 6,585,939 to Dapprich, which are incorporated herein by reference in their entireties for all purposes.

[0056] Furthermore, suitable phosphorescent compounds may include metal complexes of one or more metals, such as ruthenium, osmium, rhenium, iridium, rhodium, platinum, indium, palladium, molybdenum, technetium, copper, iron, chromium, tungsten, zinc, etc. Particularly preferred are ruthenium, rhenium, osmium, platinum, and palladium. The metal complexes may include one or more ligands that promote the dissolution of the complexes in aqueous or non-aqueous environments. For example, some suitable ligands include, but are not limited to, pyridine, pyrazine, isonicotinamide, imidazole, bipyridine, terpyridine, phenanthroline, dipyridophenazine, porphyrin, porphine, and their derivatives. Such ligands can be substituted with, for example, alkyl, substituted alkyl, aryl, substituted aryl, aralkyl, substituted aralkyl, carboxylate, carboxaldehyde, carboxamide, cyano, amino, hydroxy, imino, hydroxycarbonyl, aminocarbonyl, amidine, guanidinium, ureido, sulfur-containing groups, phosphorus-containing groups, and carboxylic acid esters of N-hydroxy-succinimide.

[0057] Porphyrin and porphine metal complexes have pyrrole groups linked by methylene bridges to form a ring structure, chelating the metal in the internal cavity. Many of these molecules exhibit strong phosphorescence at room temperature in an oxygen-free environment in a suitable solvent (e.g., water). Some suitable porphyrin complexes capable of exhibiting phosphorescence include, but are not limited to, platinum(II) coproporphyrin-I and III, palladium(II) coproporphyrin, ruthenium coproporphyrin, zinc(II)-coproporphyrin-I, their derivatives, and the like. Similarly, some suitable porphine complexes capable of exhibiting phosphorescence include, but are not limited to, platinum(II) tetramesofluorophenylporphine and palladium(II) tetramesofluorophenylporphine. Still other suitable porphyrin and / or porphine complexes are described in U.S. Patent No. 4,614,723 to Schmidt et al., U.S. Patent No. 5,464,741 to Hendrix, U.S. Patent No. 5,518,883 to Soini, U.S. Patent No. 5,922,537 to Ewart et al., U.S. Patent No. 6,004,530 to Sagner et al., and U.S. Patent No. 6,582,930 to Ponomarev et al., which are incorporated herein by reference in their entireties for all purposes.

[0058] As used herein, "sequence identity" is determined using a standalone executable BLAST engine program (bl2seq) for blasting two sequences. This can be searched using default parameters from the National Center for Biotechnology Information (NCBI) ftp site (Tatusova and Madden, FEMS Microbiol Lett., 1999, 174, 247-250, incorporated herein by reference in its entirety). The use of the term "homologous to" is synonymous with the measured "sequence identity."

[0059] As used herein, the term "sample" generally refers to a limited amount of something, and is intended to be similar to and symbolic of a larger amount of that thing. In this disclosure, a sample is a collection, swab, scraping, abrasion, biopsy, excised tissue, or surgical resection undergoing testing for an assay or method disclosed herein. In some embodiments, the sample is taken from a patient or subject suspected of containing hyperproliferative cells. In some embodiments, a sample suspected of containing an infection is compared to a "control sample" known to be free of one or more cells. In some embodiments, a sample suspected of containing pathogen cells is compared to a control sample known to be free of pathogen cells. In some embodiments, a sample suspected of containing hyperproliferative cells is compared to a control sample known to be free of hyperproliferative cells. In some embodiments, the sample is a scraping of a surrounding area or location, such as a laboratory bench or medical device. The present disclosure contemplates using any one or more of the methods disclosed herein to identify, detect, and / or quantify the amount of potentially harmful gene expression or the amount of harmful pathogens or cells in a particular item or location based on the expression of harmful genes or nucleic acid sequences.

[0060] The terms "complementary" or "complementarity" refer to polynucleotides (i.e., a sequence of nucleotides) related by the base-pairing rules; for example, the sequence "5'-AGT-3'" is complementary to the sequence "5'-ACT-3'." Complementarity can be "partial," where only some of the nucleobases match according to the base-pairing rules, or it can be "complete" or "total" complementarity between the nucleic acids. The degree of complementarity between nucleic acid strands can have significant effects on the efficiency and strength of hybridization between nucleic acid strands under defined conditions. This is particularly important for methods that rely on binding between nucleobases.

[0061] Any probe disclosed herein may be an antibody. As used herein, the term "antibody" refers to a polypeptide or group of polypeptides that are composed of at least one binding domain formed by folding of the polypeptide chain and have a three-dimensional binding space with an internal surface shape and charge distribution complementary to the antigenic determinant characteristics of the antigen. Antibodies typically have a tetrameric form and comprise two identical pairs of polypeptide chains, each pair having one "light" chain and one "heavy" chain. The variable regions of each light / heavy chain pair form the antibody binding site. As used herein, a "targeted binding agent" refers to an antibody or binding fragment thereof that preferentially binds to a target site. In one embodiment, the targeted binding agent is specific for only one target site. In other embodiments, the targeted binding agent is specific for multiple target sites. In one embodiment, the targeted binding agent may be a monoclonal antibody, and the target site may be an epitope or antigen on the surface of a cell that contains one or more of the modified oligonucleotides disclosed herein. "Binding fragments" of antibodies can be produced by recombinant DNA techniques or by enzymatic or chemical cleavage of intact antibodies. Binding fragments include Fab, Fab', F(ab')2, Fv, and single-chain antibodies. An antibody other than a "bispecific" or "bifunctional" antibody is understood to have each of its binding sites identical. An antibody substantially inhibits adhesion of a receptor to a counterreceptor if the excess antibody reduces the amount of receptor bound to the counterreceptor by at least about 20%, 40%, 60%, or 80%, more usually by more than about 85% (as measured in in vitro competitive binding assays). Antibodies can be oligoclonal, polyclonal, monoclonal, chimeric, CDR-grafted, multispecific, bispecific, catalytic, chimeric, humanized, fully human, anti-idiotypic, and antibodies that can be labeled in soluble or bound form, as well as fragments, variants, or derivatives thereof, alone or in combination with other amino acid sequences provided by known techniques. The antibody may be derived from any species. The term antibody also includes binding fragments of the antibodies of the invention.Exemplary fragments include Fv, Fab, Fab', single-chain antibodies (svFC), dimeric variable regions (diabodies), and disulfide-stabilized variable regions (dsFv). As discussed herein, minor variations in the amino acid sequence of an antibody or immunoglobulin molecule are contemplated as encompassed by the present invention, provided that the variations maintain at least 75%, more preferably at least 80%, 90%, 95%, and most preferably 99% sequence identity to the antibody or immunoglobulin molecule described herein. Specifically, conservative amino acid substitutions are contemplated. Conservative substitutions are those made within a family of amino acids that have related side chains. Genetically encoded amino acids are generally classified into the following families: (1) acidic = aspartic acid, glutamic acid, (2) basic = lysine, arginine, histidine, (3) nonpolar = alanine, valine, leucine, isoleucine, proline, phenylalanine, methionine, tryptophan, and (4) uncharged polar = glycine, asparagine, glutamine, cysteine, serine, threonine, tyrosine. More preferred families are as follows: serine and threonine are the aliphatic hydroxy family, asparagine and glutamine are the amide-containing family, alanine, valine, leucine, and isoleucine are the aliphatic family, and phenylalanine, tryptophan, and tyrosine are the aromatic family. For example, it is reasonable to predict that a single substitution of leucine with isoleucine or valine, a single substitution of aspartic acid with glutamic acid, a single substitution of threonine with serine, or similar substitution of an amino acid with a structurally related amino acid will not significantly affect the binding function or properties of the resulting molecule, particularly if the substitution does not involve an amino acid within a framework region. Whether an amino acid change results in a functional peptide can be readily determined by assaying the specific activity of the polypeptide derivative. Assays are described in detail herein. Fragments or analogs of antibody or immunoglobulin molecules can be readily prepared by those skilled in the art. Preferred amino and carboxy termini of fragments or analogs occur near the boundaries of functional domains.Structural and functional domains can be identified by comparing the nucleotide and / or amino acid sequence data to public or proprietary sequence databases. Preferably, computer-based comparison methods are used to identify sequence motifs or domains of predicted protein conformation that occur in other proteins of known structure and / or function. Methods for identifying protein sequences that fold into known three-dimensional structures are known. See, for example, Bowie et al., Science 253:164 (1991), incorporated by reference in its entirety. Antibodies can be fragmented using conventional techniques, and the fragments can be screened for utility in the same manner as described above for whole antibodies. For example, F(ab')2 fragments can be generated by treating an antibody with pepsin. The resulting F(ab')2 fragment can be treated to reduce disulfide bridges to generate Fab' fragments.

[0062] The present disclosure also contemplates the use of one or more chimeric antibody derivatives, i.e., antibody molecules that combine a non-human animal variable region with a human constant region. Chimeric antibody molecules can include, for example, an antigen-binding domain derived from an antibody of a mouse, rat, or other species and a human constant region. Various approaches for producing chimeric antibodies have been described and can be used to generate chimeric antibodies containing immunoglobulin variable regions that recognize selected antigens on the surface of differentiated or tumor cells. See, e.g., Morrison et al., 1985; Proc. Natl. Acad. Sci. USA 81, 6851; Takeda et al., 1985, Nature 314:452; Cabilly et al., U.S. Pat. No. 4,816,567; Boss et al., U.S. Pat. No. 4,816,397; Tanaguchi et al., European Patent Publication Nos. EP 171496, 0173494, and GB 2177096B. In any of the methods of the present disclosure, the method can include exposing a known substrate to any one or more of the enzymes listed in Table 1, followed by exposing any antibody having affinity for any of the reaction products created by cleavage of the substrate.

[0063] Chemical conjugation is based on the use of homo- and heterobifunctional reagents and E-amino or hinge-region thiol groups. Homobifunctional reagents, such as 5,5'-dithiobis(2-nitrobenzoic acid) (DNTB), create disulfide bonds between two Fab fragments, and o-phenylenedimaleimide (O-PDM) creates thioether bonds between two Fab fragments (Brenner et al., 1985; Glennie et al., 1987). Heterobifunctional reagents, such as N-succinimidyl-3-(2-pyridyldithio)propionate (SPDP), link exposed amino groups of antibodies and Fab fragments, regardless of class or isotype (Van Dijk et al., 1989).

[0064] The assay device of the present disclosure may be used in a variety of formats to determine the presence or absence of lipid-modified oligonucleotides or nucleic acid sequences, or functional fragments thereof, in samples or cells isolated from a subject. For example, a "sandwich" format typically involves mixing a test sample with a lipid-modified nucleic acid sequence conjugated to a specific binding member (e.g., an antibody) for the analyte to form a complex between the analyte and the conjugated probe. These complexes are then contacted with a receptive material (e.g., an antibody) immobilized within a detection zone. Binding occurs between the analyte / probe conjugate complex and the immobilized receptive material, thereby localizing a detectable "sandwich" complex to indicate the presence of the analyte or antigen in any one of the cells disclosed herein. This technique may be used to obtain quantitative or semi-quantitative results. Some examples of such sandwich-type assays are described in U.S. Patent No. 4,168,146.

[0065] As used herein, the term "hybridization" is used in reference to the pairing of complementary nucleic acids. Hybridization and the strength of hybridization (i.e., the strength of the association between nucleic acids) depend on the degree of complementarity between the nucleic acids, the stringency of the conditions involved, and the T of the hybrid formed. m The hybridization temperature is affected by factors such as the thermal melting temperature T of the hybridization conditions necessary to obtain specific hybridization of the target site to the target nucleic acid. The "hybridization" method involves annealing one nucleic acid to another complementary nucleic acid, i.e., a nucleic acid having a complementary nucleotide sequence. Hybridization is performed under conditions that allow specific hybridization. The length, secondary structure, and GC content of the complementary sequence determine the thermal melting temperature T of the hybridization conditions necessary to obtain specific hybridization of the target site to the target nucleic acid. mHybridization can be performed under stringent conditions. The phrase "stringent hybridization conditions" refers to conditions under which a probe will hybridize to its target subsequence, usually in a complex nucleic acid mixture, but will not hybridize to other sequences at detectable or significant levels. Stringent conditions are sequence-dependent and vary depending on the environment. Stringent conditions include conditions in which the salt concentration is less than about 1.0 M sodium ion, e.g., less than about 0.01 M (or other salts), including sodium ion concentrations of about 0.001 M to about 1.0 M, pH is about 6 to about 8, and temperature is about 20°C to about 65°C. Stringent conditions can also be achieved with the addition of destabilizing agents, such as, but not limited to, formamide.

[0066] The oligonucleotide sequences, nucleic acid sequences, or other drugs of the present disclosure can be administered as, among others, pharmaceutically acceptable salts, esters, or amides.The term "salt" refers to inorganic and organic salts of the compounds of the present disclosure.The salts can be prepared in situ during the final isolation and purification process of the compound, or can be prepared independently by reacting the purified compound in its free base or free acid form with an appropriate organic or inorganic base or organic or inorganic acid and isolating the formed salt.Representative salts include hydrobromide, hydrochloride, sulfate, bisulfate, nitrate, acetate, oxalate, palmitate, stearate, laurate, borate, benzoate, lactate, phosphate, tosylate, citrate, maleate, fumarate, succinate, tartrate, naphthylate, mesylate, glucoheptonate, lactobionate, and laurylsulfonate. The salts may include alkali metal and alkaline earth metal cations, such as sodium, lithium, potassium, calcium, magnesium, etc., as well as non-toxic ammonium, quaternary ammonium, and amine cations, including, but not limited to, ammonium, tetramethylammonium, tetraethylammonium, methylamine, dimethylamine, trimethylamine, triethylamine, ethylamine, etc. See, e.g., S.M. Berge, et al., "Pharmaceutical Salts," J. Pharm. Sci., 66:1-19 (1977). In some embodiments, the compositions disclosed herein include one or more salts of the oligonucleotide sequences disclosed herein.

[0067] "Thermal melting point", "melting temperature" or "T m The term "temperature" as used herein refers to the temperature (under defined ionic strength, pH, and nucleic acid concentration) at which 50% of the probes complementary to a target hybridize to the target sequence at equilibrium (because the target sequence is present in excess, T m (At equilibrium, 50% of the probes are occupied.) d" is used to define the temperature at which at least half of the probe dissociates from its perfectly matched target nucleic acid.

[0068] The formation of a double-stranded molecule with fully formed hydrogen bonds between corresponding nucleotides is called a "match" or "perfect match," while a double-stranded molecule with a single or several mismatched nucleotide pairs is called a "mismatch." Any combination of single-stranded RNA or DNA molecules can form a double-stranded molecule (DNA:DNA, DNA:RNA, RNA:DNA, or RNA:RNA) under appropriate experimental conditions. Similarly, synthetic analogs can form double-stranded molecules with each other or with RNA and DNA under appropriate conditions.

[0069] The phrase "selectively (or specifically) hybridize" refers to the ability of a molecule to bind, duplex, or hybridize under stringent hybridization conditions only to a particular nucleotide sequence when present in a complex mixture (e.g., total cellular DNA or RNA or library DNA or RNA). Those skilled in the art will readily recognize that alternative hybridization and wash conditions can be used to provide conditions of similar stringency, and will recognize that a combination of parameters is far more useful than the measurement of any single parameter.

[0070] "Binding" refers to a sequence-specific, non-covalent interaction between macromolecules (e.g., between a protein and a nucleic acid). Not all components of a binding interaction need be sequence-specific (e.g., contacts with phosphate residues in a DNA backbone), as long as the interaction as a whole is sequence-specific. Such interactions generally have a dissociation constant (K d )10 -6 M -1 "Affinity" refers to the strength of binding, and an increase in binding affinity is d correlates with a decrease in

[0071] The term "substantially similar," as used in the context of nucleic acid or amino acid sequence identity, means at least about 50%, at least about 51%, at least about 52%, at least about 53%, at least about 54%, at least about 55%, at least about 56%, at least about 57%, at least about 58%, at least about 59%, at least about 60%, at least about 61%, at least about 62%, at least about 63%, at least about 64%, at least about 65%, at least about 66%, at least about 67%, at least about 68%, at least about 69%, at least about 70%, at least about 71%, at least about 72%, at least about 73%, at least about 74%, at least about 75%, at least about 76%, at least about 77%, at least about 78%, at least about 79%, at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, at least about 100%, at least about 101%, at least about 102%, at least about 103%, at least about 104%, at least about 105%, at least about 106%, at least about 107%, at least about 108%, at least about 109%, at least about 110%, at least about 111%, at least about 112%, at least about 113%, at least about 114%, at least about 115%, at least about 116%, at "SEQ ID NO: 1" refers to two or more sequences that share about 74%, at least about 75%, at least about 76%, at least about 77%, at least about 78%, at least about 79%, at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% sequence identity.

[0072] As used herein, "% sequence identity" is determined using the EMBOSS pairwise alignment algorithm tool available from The European Bioinformatics Institute (EMBL-EBI), part of the European Molecular Biology Laboratory (EMBL), which can be accessed at the website found by placing "www" before "ebi.ac.uk / Tools / emboss / align / ". This tool utilizes the Needleman-Wunsch global alignment algorithm (Needleman, S. B. and Wunsch, C. D. (1970) J. Mol. Biol. 48, 443-453; Kruskal, J. B. (1983) An overview of sequence comparison In D. Sankoff and B. Kruskal, (ed.), Time warps, string edits and macromolecules: the theory and practice of sequence comparison, pp. 1-44, Addison Wesley). Default settings are used, including gap open: 10.0 and gap extend 0.5. The default matrix "Blosum62" is used for amino acid sequences, and the default matrix "DNAfull" is used for nucleic acid sequences.

[0073] The term "operably linked" refers to the juxtaposition of two or more components (e.g., sequence elements) where the components are arranged so as to permit both components to function normally and to allow the potential for at least one of the components to mediate a function that affects at least one of the other components.

[0074] As used herein, "barcode" refers to a sequence, tag, or combination of tags associated with a polynucleotide, and its identity (e.g., of the DNA sequence of the tag) can be used to distinguish polynucleotides in a sample. In certain embodiments, the barcode on a polynucleotide is used to identify the origin from which the polynucleotide originates. For example, a nucleic acid sample can be a pool of polynucleotides from different origins (e.g., polynucleotides from different individuals, different tissues, or cells, or polynucleotides isolated at different times), and each polynucleotide from a different origin is tagged with a unique barcode. Thus, the barcode provides a correlation between the polynucleotide and its origin. In certain embodiments, the barcode is used to uniquely tag each individual polynucleotide in a sample. By determining the number of unique barcodes in a sample, the number of individual polynucleotides present in the sample (or the number of original polynucleotides from which the manipulated polynucleotide sample was obtained, see, for example, U.S. Patent No. 7,537,897, incorporated herein by reference in its entirety) can be determined. Barcodes are approximately 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, The barcodes can range from 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, or about 100 nucleotide bases in length or more and may comprise multiple subunits, with each different barcode having a different identity and / or subunit order.Exemplary nucleic acid tags utilized as barcodes are described in U.S. Patent No. 7,544,473 and U.S. Patent No. 7,393,665, both of which are incorporated herein by reference in their entireties for their description of nucleic acid tags and their use in identifying oligonucleotides. In certain embodiments, the set of barcodes used to tag multiple samples is determined by a particular common characteristic (e.g., T), as the methods described herein can accommodate a wide variety of unique sets of barcodes. mIt is emphasized here that barcodes need only be unique within a given experiment. Thus, the same barcode may be used to tag different samples processed in different experiments. Furthermore, in a particular experiment, a user may use the same barcode to tag different subsets of samples within the same experiment. For example, all samples obtained from individuals with a particular phenotype may be tagged with the same barcode; e.g., all samples obtained from control (or wild-type) subjects may be tagged with a first barcode, while subjects with a pathological condition may be tagged with a second barcode (different from the first barcode). As another example, it may be desirable to tag different samples from the same source with different barcodes (e.g., samples obtained over time or from different sites within a tissue). Furthermore, barcodes can be generated by a variety of different methods, for example, combinatorial tagging approaches in which one barcode is attached by ligation and a second barcode is attached by primer extension. In some embodiments, multiple unique barcodes can be attached to the same sample, increasing its uniqueness relative to other samples. Alternatively, one barcode can represent a class of sample (e.g., a well plate), and a second or third barcode can represent a specific well within that plate. In some embodiments, a sample may be tagged with multiple barcodes by hybridizing multiple barcode oligonucleotides to lipid-modified or hydrophobic anchor oligonucleotides, or a sample may be labeled with multiple barcode lipid-modified or hydrophobic anchor oligonucleotides. In some embodiments, individual cells can be barcoded via split-pool labeling, generating a unique barcode profile distinct from all other cells in the pool. Thus, barcodes can be designed and implemented in a variety of different ways to track polynucleotide fragments throughout processing and analysis, and no limitation in this regard is intended.

[0075] "Polymerase chain reaction" or "PCR" refers to a reaction for in vitro amplification of a specific DNA sequence by simultaneous primer extension of complementary strands of DNA. In other words, PCR is a reaction for creating multiple copies or replicas of a target nucleic acid flanked by primer sites. Such a reaction involves one or more repetitions of the following steps: (i) denaturation of the target nucleic acid, (ii) annealing of the primers to the primer sites, and (iii) extension of the primers by a nucleic acid polymerase in the presence of nucleoside triphosphates. Typically, the reaction is cycled in a thermal cycler at different temperatures, each optimized for each step. The specific temperatures, the duration of each step, and the rate of change between steps depend on many factors well known to those skilled in the art. Examples are found in McPherson et al., editors, PCR: A Practical Approach and PCR2: A Practical Approach (IRL Press, Oxford, 1991 and 1995, respectively). For example, in conventional PCR using Taq DNA polymerase, double-stranded target nucleic acids are denatured at temperatures above 90°C, primers are annealed at temperatures between 50 and 75°C, and primers are extended at temperatures between 72 and 78°C. The term "PCR" encompasses derivative forms of the reaction, including, but not limited to, RT-PCR, real-time PCR, nested PCR, quantitative PCR, multiplex PCR, and the like. Reaction volumes range from hundreds of nanoliters, e.g., 200 nL, to hundreds of μL, e.g., 200 μL. "Reverse transcription PCR," or "RT-PCR," refers to PCR preceded by a reverse transcription reaction that converts target RNA into complementary single-stranded DNA, followed by amplification of the DNA. See, for example, Tecott et al., U.S. Patent No. 5,168,038, incorporated herein by reference. "Real-time PCR" refers to PCR in which the amount of reaction product, i.e., amplicon, is measured as the reaction progresses. There are many forms of real-time PCR, which differ primarily in the detection chemistry used to measure the reaction product.For example, Gelfand et al., U.S. Pat. No. 5,210,015 ("TAQMAN®"); Wittwer et al., U.S. Pat. Nos. 6,174,670 and 6,569,627 (intercalating dyes); Tyagi et al., U.S. Pat. No. 5,925,517 (molecular beacons). These patents are incorporated herein by reference. Real-time PCR detection chemistry is reviewed in Mackay et al., Nucleic Acids Research, 30:1292-1305 (2002), which is also incorporated herein by reference. "Nested PCR" refers to a two-stage PCR in which the amplicon from the first PCR is subjected to a second PCR using a new set of primers, at least one of which binds to the interior of the first amplicon. As used herein, "first primer" in reference to a nested amplification reaction refers to the primer used to generate the first amplicon, and "second primer" refers to one or more primers used to generate the second, i.e., nested, amplicon. "Multiplex PCR" refers to PCR in which multiple target sequences (or a single target sequence and one or more reference sequences) are run simultaneously in the same reaction mixture. For example, Bernard et al., Anal. Biochem. 273:221-28 (1999) (two-color real-time PCR). Typically, a different primer set is used for each sequence to be amplified.

[0076] The term "hyperproliferative cell" refers to a cancerous, precancerous, hyperplastic, or senescent cell, meaning a cell that is unable to undergo mitosis normally. In some embodiments, the hyperproliferative cell is a tumor cell. In some embodiments, the hyperproliferative cell contains a dysfunctional cell cycle that renders the cell deficient in apoptosis or metabolically unstable, such that the cell proliferates faster than a metabolically stable cell of the same type.

[0077] As used herein, "expression" refers to the process by which a polynucleotide is transcribed from a DNA template (e.g., into mRNA or other RNA transcript) and / or the process by which a transcribed mRNA is subsequently translated into a peptide, polypeptide, or protein. The transcript and the encoded polypeptide are sometimes collectively referred to as the "gene product." If the polynucleotide is derived from genomic DNA, expression may include splicing of the mRNA in a eukaryotic cell.

[0078] The term "functional fragment" refers to any portion of a polypeptide or nucleic acid sequence to which a full-length polypeptide or nucleic acid, respectively, is related, and is of sufficient length and structure to confer at least a similar or substantially similar biological effect as the full-length polypeptide or nucleic acid from which the fragment is based. In some embodiments, a functional fragment is a portion of a full-length or wild-type nucleic acid sequence encoding any one of the nucleic acid sequences disclosed herein, which portion encodes a polypeptide of a particular length and / or structure less than the full length, but which still encodes a domain that is biologically functional in relation to the full-length or wild-type protein. In some embodiments, the functional fragment may have less, approximately the same, or more biological activity than the wild-type or full-length polypeptide sequence from which the fragment is based. In some embodiments, the functional fragment is derived from a sequence of an organism, e.g., a human. In such embodiments, the functional fragment may retain 99%, 98%, 97%, 96%, 95%, 94%, 93%, 92%, 91%, or 90% sequence identity to the wild-type human sequence from which it is derived, hi some embodiments, the functional fragment may retain 85%, 80%, 75%, 70%, 65%, or 60% sequence identity to the wild-type sequence or nucleotide oligo portion from which it is derived.

[0079] Lipid-modified oligonucleotides The present disclosure relates to a composition and method of use for a cell barcoding method that uses a recently developed set of specific lipid-conjugated or hydrophobic-anchored oligonucleotides to efficiently label single cells derived from different patients or test conditions. Oligonucleotide barcodes (designed with PCR handles, unique identifiers, and capture sequences) are subsequently introduced into the cells, and a subset of the cells can be processed for preparation of RNA sequencing libraries in a droplet microfluidics system.

[0080] Stable embedding of lipid-modified oligonucleotides into the plasma membrane of cells via a two-component system has been disclosed in Selden et al., J. Am. Chem. Soc. 134:765-68 (2012), Weber et al., BioMacromolecules 15:4621-26 (2014), and published U.S. patent application Ser. No. 2017 / 0305955, each of which is incorporated herein by reference in its entirety.

[0081] A general and non-limiting example of the lipid-modified oligonucleotide disclosed herein is shown in the claims.This lipid-modified oligonucleotide comprises three oligonucleotides: a first oligonucleotide comprising, in 5' to 3' direction, a first lipid moiety, a first hybridization region and a first primer region; a second oligonucleotide comprising, in 5' to 3' direction, a second hybridization region and a second lipid moiety, wherein the second hybridization region is the reverse complement of the first hybridization region; and a third oligonucleotide comprising, in 5' to 3' direction, a second primer region, a barcode region and a capture sequence, wherein the second primer region is the reverse complement of the first primer region.

[0082] The present disclosure also relates to microfluidics and labeled nucleic acids. For example, certain aspects generally relate to systems and methods for labeling nucleic acids in microfluidic droplets or other compartments, e.g., derived from cells. In one set of embodiments, particles can be prepared to include oligonucleotides that can be used to identify target nucleic acids bound to the particle's surface. The oligonucleotides can include a "barcode" or unique sequence that can be used to distinguish nucleic acids in one droplet from nucleic acids in another, e.g., even after the nucleic acids have been pooled together or removed from the droplet. Certain embodiments of the present invention generally relate to systems and methods for attaching additional or arbitrary sequences, e.g., recognition sequences that can be used to selectively identify or amplify desired sequences suspected to be present in the droplet, to nucleic acids in microfluidic droplets or other compartments. Such systems can be useful, for example, for selective amplification in various applications, e.g., high-throughput sequencing applications.

[0083] Some aspects of the present disclosure generally relate to systems and methods for encapsulating or encapsulating nucleic acids in lipid-modified or hydrophobic anchor oligonucleotides within individual spots on microfluidic droplets or other suitable compartments, such as microwells in a microwell plate, a slide, or other surface. In some cases, the nucleic acid and the oligonucleotide can be ligated or attached together. The nucleic acid may originate from lysed cells, organelles, or other substances within the droplet. The oligonucleotides within a droplet may be distinguishable from oligonucleotides in other droplets, for example, within multiple droplets or a population of droplets. For example, the oligonucleotides may contain one or more unique sequences or "barcodes" that differ among various droplets. Thus, nucleic acids within each droplet can be uniquely identified by identifying the barcode associated with the nucleic acid. This can be important, for example, for sequencing or other analysis, when the droplets are "disrupted" or ruptured and nucleic acids from different droplets are subsequently mixed or pooled together.

[0084] The present disclosure relates to cells comprising one or more lipid-modified oligonucleotides, wherein the lipid-modified oligonucleotides comprise a lipid moiety region and, optionally, a capture region. In some embodiments, the cells are hyperproliferative cells, transformed cells from a cell line, or primary cells isolated from a subject or patient.

[0085] The present disclosure also relates to a system comprising a plate, dish or other solid support containing one or more containers with samples, biopsies, cells or tissues immobilized or absorbed to the surface of the container and exposing one or more barcodes, each container containing a unique barcode or barcodes corresponding to the location of expression of the cells and / or nucleic acids in the sample, biopsy, cell or tissue.

[0086] Lipid subregion In some embodiments, the lipid moiety region comprises an alkyl chain and an alkenyl, alkyl, aryl, or aralkyl chain, the alkenyl, alkyl, aryl, or aralkyl chain being about 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 20, 21, 22 It may contain 5, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, or 100 carbon atoms or more.In some embodiments, the alkyl chain is about 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, or 100 carbon atoms or more, The alkyl, aryl, or aralkyl chain may be about 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55 , 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, or 100 carbon atoms or more. In some embodiments, the chains share the same number of carbon atoms. In other embodiments, one chain has about 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 fewer carbon atoms than the other chain.The lipid moiety may comprise two or more alkenyl, aryl, or aralkyl chains, each chain being 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, 122, 123, 124, 125, 1 2, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, or 100 carbon atoms or more.

[0087] In some embodiments, the lipid moiety region can comprise one or more unsaturated carbon bonds. In some embodiments, the unsaturated bonds are all contained in the same chain. In still other embodiments, the unsaturated bonds can be contained in two or more chains.

[0088] In certain embodiments, the lipid moiety region comprises a dialkylphosphoglyceride, and the polynucleotide is conjugated to the dialkylphosphoglyceride. In some embodiments, each chain of the dialkylphosphoglyceride has the same number of carbon atoms as the other chain. In other embodiments, the number of carbon atoms is different between the two alkyl chains of the dialkylphosphoglyceride. In some embodiments, each chain has 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 11 , 92, 93, 94, 95, 96, 97, 98, 99, or 100 carbon atoms or more. In some embodiments, each chain has about 12 carbon atoms, or about 14 carbon atoms, or about 16 carbon atoms, or about 18 carbon atoms, or about 20 carbon atoms, or about 22 carbon atoms. In some embodiments, at least one chain has about 12 carbon atoms, or about 14 carbon atoms, or about 16 carbon atoms, or about 18 carbon atoms, or about 20 carbon atoms, or about 22 carbon atoms.

[0089] The lipid moiety region may comprise a monoalkylamide, and the polynucleotide may be conjugated to the monoalkylamide. In some embodiments, the monoalkylamide chain is about 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, The monoalkylamide chain may have 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, or about 100 carbon atoms or more. In some embodiments, the monoalkylamide chain has about 12 carbon atoms, or about 14 carbon atoms, about 16 carbon atoms, about 18 carbon atoms, about 20 carbon atoms, or about 22 carbon atoms. In certain embodiments, the monoalkylamide comprises about 16 or 18 carbon atoms.

[0090] In another embodiment, the lipid moiety and the polynucleotide are linked by a compound containing a phosphate group. In another embodiment, the lipid moiety and the polynucleotide are linked by a compound containing a urea group. In yet another embodiment, the lipid moiety and the polynucleotide are linked by a compound containing a sulfonyl group. In another embodiment, the lipid moiety and the polynucleotide are linked by a compound containing a sulfonamide, ether, thioether, carbamate, or carbonate group.

[0091] In still other embodiments, the lipid moiety region can comprise a sterol group.In some embodiments, the sterol group can be natural or synthetic, and can be derived from a sterol compound that has (or is modified to have) a functional group that can be used for binding to polynucleotide.For example, sterols derived from biological sources are usually found as free sterol alcohols, acylated (sterol esters), alkylated (steryl alkyl ethers), sulfated (cholesterol sulfate), or linked to glycoside moieties (steryl glycosides), which can be acylated themselves (acylated sterol glycosides) (see, for example, Fahy et al., J. Lipid Res. 46: 839-61 (2005)).This reference is incorporated in its entirety). Examples include (1) sterols available from animal sources, i.e., referred to herein as "zoosterols," e.g., zoosterol cholesterol and certain steroid hormones, and (2) sterols available from plant, fungal, and marine sources, i.e., referred to herein as "phytosterols," e.g., the plant sterols campesterol, sitosterol, stigmasterol, and ergosterol. These sterols generally have at least one free hydroxyl group, usually at the 3-position of ring A, another position, or a combination thereof, or can be modified to incorporate suitable hydroxyl or other functional groups, as needed.

[0092] Particularly important sterols are simple sterols that have an inherent functional group for attachment to the polynucleotide, particularly those in which the inherent functional group is a hydroxyl, particularly simple sterol alcohols that have a hydroxyl group located at the 3-position of the A ring (e.g., cholesterol, β-sitosterol, stigmasterol, campesterol, and brassicasterol, ergosterol, etc., and derivatives thereof).

[0093] Cholesterol is particularly important in certain embodiments for inclusion in the lipid moiety region.Representative sterols of the cholesterol class of interest (including substituted cholesterol) include, for example: (1) natural and synthetic sterols, such as cholesterol (ovine), cholesterol (plant-derived), desmosterol, stigmasterol, β-sitosterol, thiocholesterol, 3-cholesteryl acrylate; (2) A-ring substituted oxysterols, such as cholestanol and cholestenone; (3) B-ring substituted oxysterols, such as 7-ketocholesterol, 5α,6α-epoxycholesterol, 5β,6β-epoxycholesterol, and 7-dehydrocholesterol; (4) D-ring substituted oxysterols, such as 25-ketocholesterol and 15- Ketocholestane, (5) side-chain substituted oxysterols such as 25-hydroxycholesterol, 27-hydroxycholesterol, 24(R / S)-hydroxycholesterol, 24(R / S),25-epoxycholesterol, and 24(S),25-epoxycholesterol, (6) lanosterols such as 24-dihydrolanosterol and lanosterol, (7) fluorinated sterols such as F7-cholesterol, F7-5α,6α-epoxycholesterol, F7-5β,6β-epoxycholesterol, and F7-7-ketocholesterol, (8) fluorinated cholesterols such as 25-NBD cholesterol, dehydroergosterol, and cholesterol trienes. These compounds may include deuterated and non-deuterated forms and are commercially available, for example, from Avanti Polar Lipids, Inc.

[0094] In certain embodiments, the lipid moiety can comprise saturated or unsaturated, straight or branched, substituted or unsubstituted aliphatic chain.Particularly important is the saturated or unsaturated, straight or branched, substituted or unsubstituted hydrocarbon chain that has 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39 or 40 carbons.

[0095] Further embodiments may include elements based on or derivable from various lipids, such as fatty acids, glycerolipids, glycerophospholipids, sphingolipids, prenol lipids, polyprenol lipids, and glycolipids, e.g., from the lipids described in Fahy et al., J. Lipid Res. 46:839-61 (2005).

[0096] The "anchor" lipid-modified or hydrophobic anchor oligonucleotide (e.g., a lipid-modified oligonucleotide comprising, in a 5' to 3' or 3' to 5' direction, a first lipid moiety, a first hybridization region, and a first primer region) and the "co-anchor" lipid-modified or hydrophobic anchor oligonucleotide (e.g., a lipid-modified oligonucleotide comprising, in a 5' to 3' or 3' to 5' direction, a second hybridization region and a second lipid moiety, wherein the second hybridization region is the reverse complement of the first hybridization region) may comprise the same lipid moiety or different lipid moieties (e.g., different carbon chain lengths, different compositions, or different modifications). In some embodiments, the "anchor" lipid-modified or hydrophobic anchor oligonucleotide comprises a lipid moiety that contains the same number of carbons as the lipid moiety of the "co-anchor" lipid-modified or hydrophobic anchor oligonucleotide. In some embodiments, the "anchor" lipid-modified or hydrophobic anchor oligonucleotide comprises a lipid moiety that contains about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, or 12 more carbons than the lipid moiety of the "co-anchor" lipid-modified or hydrophobic anchor oligonucleotide. In some embodiments, the "co-anchor" lipid-modified or hydrophobic anchor oligonucleotide comprises a lipid moiety that contains about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, or 12 more carbons than the lipid moiety of the "anchor" lipid-modified or hydrophobic anchor oligonucleotide. In some embodiments, the "anchor" lipid-modified or hydrophobic anchor oligonucleotide comprises a lipid moiety that contains about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, or 12 more carbons than the lipid moiety of the "co-anchor" lipid-modified or hydrophobic anchor oligonucleotide. In some embodiments, only the anchor lipid-modified or hydrophobic anchor oligonucleotide is used without the corresponding co-anchor lipid-modified or hydrophobic anchor oligonucleotide.

[0097] In some embodiments, the lipid moiety (i.e., the lipid moiety in an embodiment of only an anchor lipid-modified or hydrophobic anchor oligonucleotide, or either the first or second lipid moiety in an embodiment of both an anchor lipid-modified or hydrophobic anchor oligonucleotide and a co-anchor lipid-modified or hydrophobic anchor oligonucleotide) is a compound of Formula I: TIFF2025166026000004.tif40128 or a physiologically acceptable salt thereof, In the formula, n 1 is 5 to 25, and n 2 is 1 to 25, X is selected from the group consisting of NH, CH2, O, and CH-R, and R is a C12 to C28 monoglyceride, alkenyl, alkyl, aryl, or aralkyl.

[0098] In some embodiments, the lipid moiety (i.e., the lipid moiety in an embodiment of only an anchor lipid-modified or hydrophobic anchor oligonucleotide, or either the first or second lipid moiety in an embodiment of both an anchor lipid-modified or hydrophobic anchor oligonucleotide and a co-anchor lipid-modified or hydrophobic anchor oligonucleotide) is a compound of Formula II: TIFF2025166026000005.tif44128 or a physiologically acceptable salt thereof, In the formula, n 1 is 5 to 25, and n 2 is 0 to 24, X is selected from the group consisting of NH, CH2, O, and CH-R, and R is a C12 to C28 monoglyceride, alkenyl, alkyl, aryl, or aralkyl.

[0099] In some embodiments, the lipid moiety, the first lipid moiety, the second lipid moiety, or both lipid moieties comprise a compound of formula III: TIFF2025166026000006.tif35153 In some embodiments, the "anchor" lipid-modified or hydrophobic anchor oligonucleotide comprises a sterol moiety and the "co-anchor" lipid-modified or hydrophobic anchor oligonucleotide comprises a lipid moiety. In some embodiments, the "co-anchor" lipid-modified or hydrophobic anchor oligonucleotide comprises a sterol moiety and the "anchor" lipid-modified or hydrophobic anchor oligonucleotide comprises a lipid moiety. In some embodiments, both the "anchor" lipid-modified or hydrophobic anchor oligonucleotide and the "co-anchor" lipid-modified or hydrophobic anchor oligonucleotide comprise sterol moieties.

[0100] Hybridization region The anchor lipid-modified or hydrophobic anchor oligonucleotide and the co-anchor lipid-modified or hydrophobic anchor oligonucleotide comprise complementary hybridization regions. The anchor lipid-modified or hydrophobic anchor oligonucleotide may be any of the following: , 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, or 100 or more nucleotide bases, which oligonucleotide may be DNA or RNA, and may be modified or synthetic DNA or RNA.

[0101] The co-anchor lipid-modified or hydrophobic anchor oligonucleotide may be 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 105, 106, 107, 108, 109, 110, 111, 120, 121, 122, 123, 124, 125, 126, 127, 128, 129, 130, 131, 132, 133, 134, 135, 136, 137, 1 The second hybridization region comprises a lipid moiety operably linked (e.g., covalently linked) to an oligonucleotide of 4, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, or 100 or more nucleotide bases, which oligonucleotide may be DNA or RNA, and which may be modified or synthetic DNA or RNA. In some embodiments, the second hybridization region can be the same type of nucleic acid as the first hybridization region (e.g., if the first hybridization region is DNA, then the second hybridization region is DNA), or the second hybridization region can be a different type of nucleic acid than the first hybridization region (e.g., if the first hybridization region is DNA, then the second hybridization region can be RNA, modified or synthetic DNA, or modified or synthetic RNA).

[0102] The second hybridization region is the reverse complement of the first hybridization region. In some embodiments, the complementarity can be perfect complementarity (i.e., the second hybridization region is the same length as the first hybridization region, and each base of the second hybridization region is perfect complement to that base pair of the first hybridization region). In some embodiments, the first hybridization region is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, , 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, or 80 or more additional bases. In some embodiments, the second hybridization region is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, 122 , 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, or 80 or more additional bases.In some embodiments, the first hybridization region is at least about 50%, at least about 51%, at least about 52%, at least about 53%, at least about 54%, at least about 55%, at least about 56%, at least about 57%, at least about 58%, at least about 59%, at least about 60%, at least about 61%, at least about 62%, at least about 63%, at least about 64%, at least about 65%, at least about 66%, at least about 67%, at least about 68%, at least about 69%, at least about 70%, at least about 71%, at least about 72%, at least about The sequence identity of the fragments of the present invention may be at least about 73%, at least about 74%, at least about 75%, at least about 76%, at least about 77%, at least about 78%, at least about 79%, at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100%.

[0103] In some embodiments, the first hybridization region has a hybridization sequence of at least about 50%, at least about 51%, at least about 52%, at least about 53%, at least about 54%, at least about 55%, at least about 56%, at least about 57%, at least about 58%, at least about 59%, at least about 60%, at least about 61%, at least about 62%, at least about 63%, at least about 64%, at least about 65%, at least about 66%, at least about 67%, at least about 68%, at least about 69%, at least about 70%, at least about 71%, at least about 72%, at least about 73%, at least about 74%, at least about 75%, at least about 76%, at least about 77%, at least about 78%, at least about 79%, at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, at least about 100%, at least about 101%, at least about 102%, at least about 103%, at least about 104%, at least about 105%, at least about 106%, at least about 107%, at least about 108%, at least about 109%, at least about 110%, at least about 111%, at least about 112%, at least about 113%, at least about 114%, at least about 115%, at least about 116%, at least about 117%, at least 2%, at least about 73%, at least about 74%, at least about 75%, at least about 76%, at least about 77%, at least about 78%, at least about 79%, at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% sequence identity.

[0104] In some embodiments, the second hybridization region has a hybridization ratio of at least about 50%, at least about 51%, at least about 52%, at least about 53%, at least about 54%, at least about 55%, at least about 56%, at least about 57%, at least about 58%, at least about 59%, at least about 60%, at least about 61%, at least about 62%, at least about 63%, at least about 64%, at least about 65%, at least about 66%, at least about 67%, at least about 68%, at least about 69%, at least about 70%, at least about 71%, at least about 72%, at least about 73%, at least about 74%, at least about 75%, at least about 76%, at least about 77%, at least about 78%, at least about 79%, at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, at least about 100%, at least about 101%, at least about 102%, at least about 103%, at least about 104%, at least about 105%, at least about 106%, at least about 107%, at least about 108%, at least about 109%, at least about 110%, at least about 111%, at least about 112%, at least about 113%, at least about 114%, at least about 115%, at least about 116%, at least about 117%, at least 2%, at least about 73%, at least about 74%, at least about 75%, at least about 76%, at least about 77%, at least about 78%, at least about 79%, at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% sequence identity.

[0105] Primer region The anchor lipid-modified or hydrophobic anchor oligonucleotide and the barcode oligonucleotide comprise complementary primer regions. The anchor lipid-modified or hydrophobic anchor oligonucleotide comprises a lipid moiety operably linked (e.g., covalently linked) to a first hybridization region, which is operably linked (e.g., covalently linked) to a first primer region comprising an oligonucleotide of 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 or more nucleotide bases. The oligonucleotide may be DNA or RNA, and may be modified or synthetic DNA or modified or synthetic RNA.

[0106] The barcode oligonucleotide is operably linked (e.g., covalently linked) to a barcode region (described below), which in turn comprises a second primer region operably linked (e.g., covalently linked) to a capture sequence (described below), the second primer region comprising an oligonucleotide of 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 or more nucleotide bases. The oligonucleotide may be DNA or RNA, and may be modified or synthetic DNA or modified or synthetic RNA. In some embodiments, the second primer region is the same type of nucleic acid as the first primer region (e.g., if the first primer region is DNA, then the second primer region is DNA), or the second primer region can be a different type of nucleic acid than the first primer region (e.g., if the first primer region is DNA, then the second primer region can be RNA, modified or synthetic DNA, or modified or synthetic RNA).

[0107] The second primer region is a reverse complement of the first primer region. In some embodiments, the complementarity can be perfect complementarity (i.e., the second primer region is the same length as the first primer region, and each base of the second primer region is a perfect complement to that base pair of the first primer region). In some embodiments, the first primer region comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 or more additional bases compared to the second primer region. In some embodiments, the second primer region comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 or more additional bases compared to the first primer region. In some embodiments, the first primer region is at least about 50%, at least about 51%, at least about 52%, at least about 53%, at least about 54%, at least about 55%, at least about 56%, at least about 57%, at least about 58%, at least about 59%, at least about 60%, at least about 61%, at least about 62%, at least about 63%, at least about 64%, at least about 65%, at least about 66%, at least about 67%, at least about 68%, at least about 69%, at least about 70%, at least about 71%, at least about 72%, at least about 73%, at least about 74%, at least about 75%, at least about 76%, at least about 77%, at least about 78%, at least about 79%, at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, at least about 100%, at least about 101%, at least about 102%, at least about 103%, at least about 104%, at least about 105%, at least about 106%, at least about 107%, at least about 108%, at least about 109%, at least about 110%, at least about 111%, at least about 112%, at least about 113%, at least about 114%, at least about 115%, at least about 116%, at least about 117%, at least about 118%, %, at least about 74%, at least about 75%, at least about 76%, at least about 77%, at least about 78%, at least about 79%, at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% sequence identity.

[0108] In some embodiments, the first primer region has a sequence similar to SEQ ID NO: 5 (TGGAATTCTCGGGTGCCAAGG), at least about 50%, at least about 51%, at least about 52%, at least about 53%, at least about 54%, at least about 55%, at least about 56%, at least about 57%, at least about 58%, at least about 59%, at least about 60%, at least about 61%, at least about 62%, at least about 63%, at least about 64%, at least about 65%, at least about 66%, at least about 67%, at least about 68%, at least about 69%, at least about 70%, at least about 71%, at least about 72%, at least about 73%, at least about 74%, at least about 75%, at least about 76%, at least about 77%, at least about 78%, at least about 79%, at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% sequence identity.

[0109] In some embodiments, the second primer region has a sequence similar to SEQ ID NO: 6 (CCTTGGCACCCGAGAATTCCA) at least about 50%, at least about 51%, at least about 52%, at least about 53%, at least about 54%, at least about 55%, at least about 56%, at least about 57%, at least about 58%, at least about 59%, at least about 60%, at least about 61%, at least about 62%, at least about 63%, at least about 64%, at least about 65%, at least about 66%, at least about 67%, at least about 68%, at least about 69%, at least about 70%, at least about 71%, at least about 72%, at least about 73%, at least about 74%, at least about 75%, at least about 76%, at least about 77%, at least about 78%, at least about 79%, at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% sequence identity.

[0110] Barcode Area The barcode oligonucleotide comprises a second primer region operably linked (e.g., covalently linked) to a barcode region, which in turn is operably linked (e.g., covalently linked) to a capture sequence (described below), and the barcode region is 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37 , 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, or 100 or more nucleotide bases. The oligonucleotides may be DNA or RNA, and may be modified or synthetic DNA or RNA. Methods for designing sets of barcode sequences are shown, for example, in U.S. Patent No. 6,235,475, the contents of which are incorporated herein by reference in their entireties. Attachment of barcode sequences to nucleic acid templates is shown in U.S. Publication Nos. 2008 / 0081330 and 2011 / 0301042, the contents of each of which are incorporated herein by reference in their entireties. Methods for designing sets of barcode sequences and other methods for combining barcode sequences are set forth in U.S. Patent Nos. 6,138,077, 6,352,828, 5,636,400, 6,172,214, 6,235,475, 7,393,665, 7,544,473, 5,846,719, 5,695,934, 5,604,097, 6,150,516, RE39,793, 7,537,897, 6,172,218, and 5,863,722, the contents of each of which are incorporated herein by reference in their entirety.Barcodes for sequencing and copy number estimation are described in U.S. Publication No. 2016 / 0046986, which is incorporated by reference in its entirety.

[0111] Barcodes can be completely random or can be designed with a predetermined sequence. They can have random or semi-random regions and other constant regions. The barcodes can also contain other regions, such as priming sites, adapters, or other complementary regions that facilitate further processing and analysis. The identity of a particular barcode itself relative to its second primer region can be created prior to any binding or capture of an anchor lipid modification or hydrophobic anchor oligonucleotide by hybridization of the first and second primer regions. Thus, in some embodiments, a database of all barcodes is created and, in some embodiments, stored on a computer storage medium.

[0112] In some embodiments, the last nucleotide of the barcode region is not identical to the first nucleotide of the capture sequence, e.g., if the capture sequence is a polyadenylation tail, the last nucleotide of the barcode region is in some embodiments not an adenine.

[0113] The barcode allows for tagging or tracking of cells or membranes containing the lipid-modified or hydrophobic anchor oligonucleotide, providing the opportunity for subsequent identification and origin of a particular cell or membrane. Assigning a barcode to an individual oligonucleotide or subgroup of oligonucleotides may allow for the assignment of a unique identity to an individual sequence, fragment of a sequence, or cell. This may allow for data to be obtained from individual samples, not limited to sample averages.

[0114] In some embodiments, the oligonucleotides may share a common barcode and therefore may subsequently be identified as originating from the same target cell. Multiple cells (of the same or different types) can be identified using multiple barcodes, each barcode identifying a particular cell type or multiple cells of a particular cell type.

[0115] A single cell or membrane may contain multiple lipid-modified or hydrophobic-anchored oligonucleotides, each with a different first primer region. Thus, such cells or membranes can be isolated and / or identified via different barcode oligonucleotides, each with a second primer region complementary to a different first primer region, and a single barcode per barcode oligonucleotide, or different barcodes for some or all of the barcode oligonucleotides.

[0116] Capture sequence In some embodiments, the barcode oligonucleotide comprises a second primer region operably linked (e.g., covalently linked) to the barcode region, which in turn is operably linked (e.g., covalently linked) to a capture sequence (described below), wherein the capture sequence is 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, The oligonucleotides include oligonucleotides of 5, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, or 100 or more nucleotide bases, which may be DNA or RNA, and which may be modified or synthetic DNA or RNA.

[0117] In some embodiments, the capture sequence is a polyadenylated tail ("poly(A) tail"), i.e., the entire capture sequence consists of adenine bases. In some embodiments, the capture sequence is at least about 50%, at least about 51%, at least about 52%, at least about 53%, at least about 54%, at least about 55%, at least about 56%, at least about 57%, at least about 58%, at least about 59%, at least about 60%, at least about 61%, at least about 62%, at least about 63%, at least about 64%, at least about 65%, at least about 66%, at least about 67%, at least about 68%, at least about 69%, at least about 70%, at least about 71%, at least about 72%, at least about 73%, or at least about 74% of the poly(A) tail. In some embodiments, the capture sequence has a sequence identity of at least about 74%, at least about 75%, at least about 76%, at least about 77%, at least about 78%, at least about 79%, at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100%. In some embodiments, the capture sequence has the sequence of SEQ ID NO: 7 (AAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAA).

[0118] In some embodiments, the capture sequence is a polythymine tail ("poly(T) tail"), i.e., the entire capture sequence consists of thymine bases. In some embodiments, the capture sequence consists of at least about 50%, at least about 51%, at least about 52%, at least about 53%, at least about 54%, at least about 55%, at least about 56%, at least about 57%, at least about 58%, at least about 59%, at least about 60%, at least about 61%, at least about 62%, at least about 63%, at least about 64%, at least about 65%, at least about 66%, at least about 67%, at least about 68%, at least about 69%, at least about 70%, at least about 71%, at least about 72%, at least about 73%, at least about 74%, at least about 75%, at least about 76%, at least about 77%, at least about 78%, at least about 79%, at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, at least about 100%, at least about 101%, at least about 102%, at least about 103%, at least about 104%, at least about 105%, at least about 106%, at least about 107%, at least about 108%, at least about 109%, at least about 110%, at least about 111%, at least about 112%, at least about 113%, at The sequence identity of the fragments of the present invention may be at least about 74%, at least about 75%, at least about 76%, at least about 77%, at least about 78%, at least about 79%, at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100%.

[0119] In some embodiments, the capture sequence is a polyuracil tail ("poly(U) tail"), i.e., the entire capture sequence consists of uracil bases. In some embodiments, the capture sequence is at least about 50%, at least about 51%, at least about 52%, at least about 53%, at least about 54%, at least about 55%, at least about 56%, at least about 57%, at least about 58%, at least about 59%, at least about 60%, at least about 61%, at least about 62%, at least about 63%, at least about 64%, at least about 65%, at least about 66%, at least about 67%, at least about 68%, at least about 69%, at least about 70%, at least about 71%, at least about 72%, at least about 73%, or at least about 74% of the poly(U) tail. The sequence identity of the fragments of the present invention may be at least about 74%, at least about 75%, at least about 76%, at least about 77%, at least about 78%, at least about 79%, at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100%.

[0120] In some embodiments, the capture sequence is a variant of a poly(A), poly(T), or poly(U) tail. Such variants include bases other than pure poly(A), poly(T), or poly(U) tails. For example, variants include AAUAAA, AUUAAA, and the like. Capture sequences can include poly(A) variants such as AACAAG, AACAAA, AAUAAU, AAUAAG, UAUAAA, AGUAAA, AAUACA, CAUAAA, AAUAUA, GAUAAA, AAUGAA, AAGAAA, ACUAAA, AAUAGA, AAUAAU, AUUACA, AUUAUA, AACAAG, or AAUAAG, each variant optionally containing additional nucleotide bases, e.g., 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31 , 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, or 100 additional nucleotide bases, wherein a majority, and in some embodiments all, of the additional nucleotide bases are adenine. Variant poly(T) and poly(U) capture sequences can be constructed similarly.

[0121] In some embodiments, instead of a capture sequence, one member of a coupling pair (e.g., antibody / antigen, receptor / ligand, or avidin-biotin pair, such as those described in U.S. Patent Application Publication No. 2006 / 0252077, incorporated herein by reference) may be linked to each fragment so that it can be captured on a surface coated with the other member of the coupling pair. Following capture, the sequence may be analyzed, for example, by single-molecule detection / sequencing, such as that described in U.S. Patent No. 7,283,337, incorporated herein by reference.

[0122] Methods for the synthesis of lipid-modified or hydrophobic anchor oligonucleotides Oligonucleotides can be synthesized using protocols known in the art, such as those described in Caruthers et al., Meth. Enzymol. 211:3 (1992), WO99 / 54459, Wincott et al., Nucleic Acids Res. 23:2677 (1995), Wincott et al., Meth. Mol. Bio. 74:59 (1997), Brennan et al., Biotechnol. Bioeng. 61:33 (1998), and U.S. Patent No. 6,001,311. All of these references are incorporated herein by reference. Oligonucleotide synthesis uses conventional nucleic acid protecting and coupling groups, such as dimethoxytrityl at the 5' end and phosphoramidite at the 3' end.

[0123] The oligonucleotides disclosed herein include natural, synthetic, or modified oligonucleotides. Modified nucleic acids have one or more modifications, such as base modifications or backbone modifications, that confer new or enhanced functionality (e.g., improved stability) to the nucleic acid. A nucleoside can be a base-sugar combination in which the base moiety is a heterocyclic base. Heterocyclic bases include purines and pyrimidines. A nucleotide is a nucleoside that further contains a phosphate group covalently linked to the sugar moiety of the nucleoside. In the case of nucleosides containing a pentofuranosyl sugar, the phosphate group can be linked to the 2', 3', or 5' hydroxyl moiety of the sugar. In the formation of an oligonucleotide, the phosphate groups covalently link adjacent nucleosides to each other to form a linear polymeric compound. In some cases, the respective ends of this linear polymeric compound can be further joined to form a circular compound. Furthermore, linear compounds can fold to yield fully or partially double-stranded compounds due to the complementarity of their internal nucleotide bases. Within oligonucleotides, the phosphate groups can be considered to form the internucleoside backbone of the oligonucleotide. The linkage or backbone of RNA and DNA can be a 3' to 5' phosphodiester bond.

[0124] Examples of suitable nucleic acids containing modifications include nucleic acids with modified backbones or non-natural internucleoside linkages. Nucleic acids with modified backbones include those that retain a phosphorus atom in the backbone and those that do not have a phosphorus atom in the backbone. Suitable modified oligonucleotide backbones containing a phosphorus atom include, for example, phosphorothioates, chiral phosphorothioates, phosphorodithioates, phosphotriesters, aminoalkylphosphotriesters, methyl and other alkyl phosphonates, including 3'-alkylene phosphonates, 5'-alkylene phosphonates, and chiral phosphonates, phosphinates, phosphoramidates, including 3'-amino phosphoramidates and aminoalkyl phosphoramidates, phosphorodiamidates, thionophosphoramidates, thionoalkylphosphonates, thionoalkylphosphotriesters, selenophosphates and boranophosphates having normal 3'-5' linkages, their 2'-5' linked analogs, and those having reversed polarity in which one or more internucleotide linkages are 3'-3', 5'-5', or 2'-2' linkages. Suitable oligonucleotides with inverted polarity contain a single 3'-3' linkage at the 3'-most internucleotide linkage, i.e., a single inverted nucleoside residue that may be basic (either absent a nucleobase or having a hydroxyl group instead).Various salts (e.g., potassium or sodium), mixed salts, and free acid forms are also included.

[0125] In some embodiments, the nucleic acids of interest have one or more phosphorothioate and / or heteroatom internucleoside linkages, particularly --CH2--NH--O--CH2--, --CH2--N(CH3)--O--CH2-- (known as a methylene(methyleneimino) or MMI backbone), --CH2--O--N(CH3)--CH2--, --CH2--N(CH3)--N(CH3)--CH2--, and --O--N(CH3)--CH2--CH2-- (the natural phosphodiester internucleotide linkage is represented by --O--P(=O)(OH)--O--CH2--). MMI-type internucleoside linkages are disclosed in the above-referenced U.S. Patent No. 5,489,677. Suitable amide internucleoside linkages are disclosed in US Pat. No. 5,602,240.

[0126] Also suitable are nucleic acids with morpholino backbone structures, such as those described in U.S. Patent No. 5,034,506.For example, in some embodiments, the nucleic acid of interest comprises a six-membered morpholino ring instead of a ribose ring.In some of these embodiments, phosphodiamidate or other non-phosphodiester internucleoside linkages replace phosphodiester linkages.

[0127] Suitable modified polynucleotide backbones that do not contain phosphorus atoms have backbones formed by short chain alkyl or cycloalkyl internucleoside linkages, mixed heteroatom and alkyl or cycloalkyl internucleoside linkages, or one or more short chain heteroatom or heterocyclic internucleoside linkages, including those with morpholino linkages (formed in part from the sugar portion of the nucleoside), siloxane backbones, sulfide, sulfoxide and sulfone backbones, formacetyl and thioformacetyl backbones, methyleneformacetyl and thioformacetyl backbones, riboacetyl backbones, alkene-containing backbones, sulfamate salt backbones, methyleneimino and methylenehydrazino backbones, sulfonate and sulfonamide backbones, amide backbones, and other backbones with mixed N, O, S, and CH2 moieties.

[0128] Also included are nucleic acid mimetics. The term "mimetics" as applied to polynucleotides encompasses polynucleotides in which only the furanose ring, or both the furanose ring and the internucleotide linkage, are replaced with non-furanose groups; replacement of only the furanose ring is also referred to as a sugar surrogate. The heterocyclic base moiety or modified heterocyclic base moiety is maintained for hybridization with an appropriate target nucleic acid. One such nucleic acid, i.e., a polynucleotide mimic that has been shown to have excellent hybridization properties, is called a peptide nucleic acid (PNA). In PNA, the sugar backbone of a polynucleotide is replaced with an amide-containing backbone, particularly an aminoethylglycine backbone. The nucleotides are retained and are directly or indirectly linked to the aza nitrogen atoms of the amide portion of the backbone.

[0129] One polynucleotide mimic with excellent hybridization properties is peptide nucleic acid (PNA). The backbone of PNA compounds is two or more linked aminoethylglycine units, which gives PNA an amide-containing backbone. The heterocyclic base moiety is directly or indirectly bound to the aza nitrogen atom of the amide portion of the backbone. Representative U.S. patents describing the preparation of PNA compounds include, but are not limited to, U.S. Patent Nos. 5,539,082, 5,714,331, and 5,719,262.

[0130] Another class of suitable polynucleotide mimics is based on linked morpholino units (morpholino nucleic acids) having heterocyclic bases attached to a morpholino ring. Several linking groups capable of linking the morpholino monomer units of morpholino nucleic acids have been reported. One class of linking groups was selected to obtain nonionic oligomeric compounds. The nonionic morpholino-based oligomeric compounds are less likely to have undesired interactions with cellular proteins. Morpholino-based polynucleotides are nonionic mimics of oligonucleotides that are less likely to have undesired interactions with cellular proteins (Braasch et al., Biochemistry, 41(14):4503-10(2002)). Morpholino-based polynucleotides are disclosed in U.S. Patent No. 5,034,506. Various compounds of the morpholino class of polynucleotides have been prepared with a variety of different linking groups connecting the monomer subunits.

[0131] Another suitable class of polynucleotide mimics is called cyclohexenyl nucleic acids (CeNA). The furanose ring normally present in DNA / RNA molecules is replaced with a cyclohexenyl ring. CeNA DMT-protected phosphoramidite monomers have been prepared and used in the synthesis of oligomeric compounds following classical phosphoramidite chemistry. Fully modified CeNA oligomeric compounds and oligonucleotides with specific CeNA modifications have been prepared and studied (see Wang et al., J. Am. Chem. Soc., 122:8595-8602 (2000)). Incorporation of CeNA monomers into DNA strands improves the stability of DNA / RNA hybrids. CeNA oligoadenylates formed complexes with RNA and DNA complements with stabilities similar to those of the native complexes. Incorporation of the CeNA structure into native nucleic acid structures has been shown to promote conformational adaptation by NMR and circular dichroism.

[0132] Also suitable as modified nucleic acids are locked nucleic acids (LNA) and / or LNA analogs. In LNA, the 2'-hydroxyl group is linked to the 4' carbon atom of the sugar ring to form a 2'-C,4'-C-oxymethylene bond, thereby forming a bicyclic sugar moiety. The bond may be a methylene (--CH2--), a group bridging the 2' oxygen atom and the 4' carbon atom, where n is 1 or 2 (Singh et al., Chem. Commun., 4:455-456 (1998)). LNA and LNA analogs exhibit extremely high thermal stability (Tm = +3 to +10°C) when duplexed with complementary DNA and RNA, and exhibit stability against 3'-exonuclease degradation and good solubility properties. Potent, non-toxic oligonucleotides containing LNA have been described (Wahlestedt et al., Proc. Natl. Acad. Sci. USA, 97:5633-38 (2000)).

[0133] The synthesis and preparation of LNA monomers adenine, cytosine, guanine, 5-methyl-cytosine, thymine and uracil, as well as their oligomerization and nucleic acid recognition properties have been described (Koshkin et al., Tetrahedron, 54:3607-30 (1998)). LNA and its preparation are also described in WO98 / 39352 and WO99 / 14226, both of which are incorporated herein by reference in their entirety. Exemplary LNA analogs are described in U.S. Patent Nos. 7,399,845 and 7,569,686, both of which are incorporated herein by reference in their entirety.

[0134] Nucleic acids may also contain one or more substituted sugar moieties. Suitable polynucleotides contain sugar substituents selected from OH, F, O-, S-, or N-alkyl, O-, S-, or N-alkenyl, O-, S-, or N-alkynyl, or O-alkyl-O-alkyl, where the alkyl, alkenyl, and alkynyl can be substituted or unsubstituted C1-C10 alkyl or C2-C10 alkenyl and alkynyl. Also suitable are O((CH2) n O) mCH3, O(CH2) n OCH3, O(CH2) n NH2, O(CH2) n CH3, O(CH2) n ONH2 and O(CH2) n ON((CH2) n CH3)2, where n and m are from 1 to about 10. Other suitable polynucleotides include sugar substituents selected from C1-C10 lower alkyl, substituted lower alkyl, alkenyl, alkynyl, alkaryl, aralkyl, O-alkaryl or O-aralkyl, SH, SCH3, OCN, Cl, Br, CN, CF3, OCF3, SOCH3, SO2CH3, ONO2, NO2, N3, NH2, heterocycloalkyl, heterocycloalkaryl, aminoalkylamino, polyalkylamino, substituted silyl, RNA cleaving groups, reporter groups, intercalators, and other substituents with similar properties. Suitable modifications include 2'-methoxyethoxy (2'-O--CH2C H2 OCH, also known as 2'-O-(2-methoxyethoxy) or 2'-MOE (Martin et al., Helv. Chim. Acta, 78:486-504 (1995)), i.e., an alkoxyalkoxy group. Suitable modifications include the 2'-dimethylaminooxyethoxy, i.e., O(CH)ON(CH), also known as 2'-DMAOE, and 2'-dimethylaminoethoxyethoxy (also known as 2'-O-dimethyl-amino-ethoxy-ethyl or 2'-DMAEOE), i.e., 2'-O-CH-O-CH-N(CH).

[0135] Other suitable sugar substituents include methoxy (--O-CH), aminopropoxy (--OCH2CH2CH2NH2), allyl (--CH2-CH=CH2), --O-allyl (--O--CH2-CH=CH2), and fluoro (F). The 2'-sugar substituent can be in the arabino (up) or ribo (down) position. A suitable 2'-arabino modification is 2'-F. Similar modifications can also be made at other positions in the oligomeric compound, particularly the 3' position of the sugar of the 3'-terminal nucleoside or 2'-5'-linked oligonucleotide and the 5' position of the 5'-terminal nucleotide. Oligomeric compounds can also have sugar mimetics, such as cyclobutyl moieties, in place of the pentofuranosyl sugar.

[0136] Nucleic acids may also contain modifications or substitutions of nucleobases (also referred to as "bases"). As used herein, "unmodified" or "natural" nucleobases include the purine bases adenine (A) and guanine (G), and the pyrimidine bases thymine (T), cytosine (C), and uracil (U). Modified nucleobases include other synthetic and natural nucleobases, such as 5-methylcytosine (5-me-C), 5-hydroxymethylcytosine, xanthine, hypoxanthine, 2-aminoadenine, 6-methyl and other alkyl derivatives of adenine and guanine, 2-propyl and other alkyl derivatives of adenine and guanine, 2-thiouracil, 2-thiothymine and 2-thiocytosine, 5-halouracil and cytosine, 5-propynyl (--C=C--CH3) uracil and cytosine and other alkynyl derivatives of pyrimidine bases, 6-azouracil, cytosine, and the like. These include tosine and thymine, 5-uracil (pseudouracil), 4-thiouracil, 8-halo, 8-amino, 8-thiol, 8-thioalkyl, 8-hydroxy and other 8-substituted adenines and guanines, 5-halo, especially 5-bromo, 5-trifluoromethyl and other 5-substituted uracils and cytosines, 7-methylguanine and 7-methyladenine, 2-F-adenine, 2-amino-adenine, 8-azaguanine and 8-azaadenine, 7-deazaguanine and 7-deazaadenine and 3-deazaguanine and 3-deazaadenine. Modified nucleobases also include tricyclic pyrimidines such as phenoxazine cytidine (1H-pyrimido(5,4-b)(1,4)benzoxazin-2(3H)-one), phenothiazine cytidine (1H-pyrimido(5,4-b)(1,4)benzothiazin-2(3H)-one), G clamps such as substituted phenoxazine cytidines (e.g., 9-(2-aminoethoxy)-H-pyrimido(5,4-(b)(1,4)benzoxazin-2(3H)-one), carbazole cytidine (2H-pyrimido(4,5-b)indol-2-one), and pyridoindole cytidine (H-pyrido(3',2':4,5)pyrrolo(2,3-d)pyrimidin-2-one).

[0137] Heterocyclic base moieties may also include those in which the purine or pyrimidine base is replaced with other heterocycles, such as 7-deaza-adenine, 7-deazaguanosine, 2-aminopyridine, and 2-pyridone. Additional nucleobases include those disclosed in U.S. Pat. No. 3,687,808, The Concise Encyclopedia of Polymer Science and Engineering, pages 858-859, Kroschwitz, JI, ed. John Wiley & Sons, 1990, Englisch et al., Angewandte Chemie, International Edition, 1991, 30, 613, and Sanghvi, YS, Chapter 15, Antisense Research and Applications, pages 289-302, Crooke, ST and Lebleu, B., ed., CRC Press, 1993. Some of these nucleobases are useful for increasing the binding affinity of oligomeric compounds. These include 5-substituted pyrimidines, 6-azapyrimidines, and N-2, N-6, and O-6 substituted purines, including 2-aminopropyladenine, 5-propynyluracil, and 5-propynylcytosine. 5-Methylcytosine substitutions have been shown to increase the stability of nucleic acid duplexes by 0.6 to 1.2°C (Sanghvi et al., eds., Antisense Research and Applications, CRC Press, Boca Raton, 1993, pp. 276-278), making them suitable base substitutions when combined with, for example, 2'-O-methoxyethyl sugar modifications.

[0138] Lipids (e.g., lipid(s), lipid precursor(s), or oleochemical(s)) can be produced by any chemical or biochemical method (see, e.g., US 9,896,691, US 9,598,710, US 9,499,829, US 9,428,779, US 9,127,288, and Kinney, 1997, Genetic Engineering, Ed.: J.K. Setlow, 19:149-166; Ohlrogge and Browse, 1995, Plant Cell 7:957-970; Shanklin and Cahoon, 1998, Annu. Rev. Plant Physiol. Plant Mol. Biol. 49:611-641; Voelker, 1996, Genetic Engineering, Ed.: J.K. Setlow,18:111-13, Gerhardt,1992,Prog.Lipid R.31:397-417,Guhnemann-Schafer & Kindl,1995,Biochim.Biophys Acta 1256:181-186,Kunau et al.,1995,Prog.Lipid Res.34:267-342,Stymne et al. al.,1993,in:Biochemistry and Molecular Biology of Membrane and Storage Lipids of Plants,Ed.:Murata and Somerville,Rockville,American Society of Plant Physiologists,150-158,Murphy & Ross 1998,Plant Journal. 13(1):1-16), and can be collected by any convenient method (e.g., centrifugation of extracellular secreted lipids, exposure to solvent, whole cell extraction (e.g., cell disruption and collection), hydrophobic solvent extraction (e.g., hexane), liquefaction, supercritical carbon dioxide extraction, lyophilization, mechanical disruption, secretion (e.g., by addition of an effective exporter protein), or a combination thereof). In some embodiments, the lipids can be extracted and purified from, for example, plants, bacteria, or oleaginous yeast or fungi.

[0139] In some embodiments, the lipid is covalently linked to the oligonucleotide disclosed herein. In some embodiments, the lipid is crosslinked to the oligonucleotide disclosed herein. The method of binding the lipid to the oligonucleotide is not particularly limited. The lipid and the oligonucleotide may be directly bound or may be bound via a linker (binding region). In some embodiments, the linker used to bind the lipid to the oligonucleotide comprises a nucleic acid. In some embodiments, the linker used to bind the lipid to the oligonucleotide does not comprise a nucleic acid.

[0140] As long as the lipid and the oligonucleotide are covalently linked to each other, the linker that can be used is not particularly limited.Examples of usable linkers include the following structures: --OP(=O)(OH)-O--, --O--CO--O--, --NH--CO--O--, --NH--CO--NH--, --NH--(CH2) n1 --, --S--(CH2) n1 --, --CO--(CH2) n1 --CO--, --CO--(CH2) n1 --NH--, --NH--(CH2) n1 --NH--, --CO--NH--(CH2) n1 --NH--CO--, --C(=S)--NH--(CH2) n1 --NH--CO--, --C(=S)--NH--(CH2) n1 --NH--C--(=S)--, --CO--O--(CH2) n1 --O--CO--, --C(.=S)--O--(CH2) n1 --O--CO--, --C(=S)--O--(CH2) n1 --O--C--(=S)--, --CO--NH--(CH2) n1 --O--CO--, --C(=S)--NH--(CH2) n1 --O--CO--, --C(=S)--NH--(CH2) n1--O--C--(=S)--, --CO--NH--(CH2) n1 --O--CO--, --C(=S)--NH--(CH2) n1 --CO--, --C(=S)--O--(CH2) n1 --NH--CO--, --C(=S)--NH--(CH2) n1 --O--C--(=S)--, --NH--(CH2CH2O) n2 --CH(CH2OH)--, --NH--(CH2CH2O) n2 --CH2--, --NH--(CH2CH2O) n2 --CH2--CO--, --O--(CH2) n3 --S--S--(CH2) n4 --O--P(=O)2--, --CO--(CH2) n3 --O--CO--NH--(CH2) n4 --, --CO--(CH2) n3 --CO--NH--(CH2) n4 --, where n1 is an integer from about 1 to about 40, n2 is an integer from about 1 to about 20, and n3 and n4, which may be the same or different, are integers from about 1 to about 20. In some embodiments, the linker is a phosphate group (--OP(=O)(OH)--).

[0141] How to use Lipid-modified or hydrophobic-anchored oligonucleotide compounds and compositions containing lipid-modified or hydrophobic-anchored oligonucleotides can be used in a variety of different pharmaceutical, cosmeceutical, diagnostic, and biomedical applications. For example, lipid-modified or hydrophobic-anchored oligonucleotide compounds and compositions containing lipid-modified or hydrophobic-anchored oligonucleotide compounds are used in research and therapeutic applications, including the study of cell-cell interactions, membrane structure, bottom-up assembly of tissues, quantitative imaging of non-adherent cells, or the study of biological processes occurring near the cell surface. Lipid-modified or hydrophobic-anchored oligonucleotide compounds and compositions containing lipid-modified or hydrophobic-anchored oligonucleotides can also be used to study the spatial expression profile of genes in samples or tissues from subjects.

[0142] In carrying out such a method, a composition comprising lipid-modified or hydrophobic anchor oligonucleotide may first be contacted with a membrane, for example, a cell membrane, the plasma membrane separating the cytoplasm from the outside of the cell, or a nuclear membrane, under conditions that allow the composition to be inserted into the membrane.In some embodiments, the method comprises contacting a membrane with a composition comprising lipid-modified or hydrophobic anchor oligonucleotide and incubating the composition with the membrane under conditions that allow the composition to be inserted into the membrane.In some embodiments, the method comprises contacting a membrane with a composition comprising lipid-modified or hydrophobic anchor oligonucleotide and incubating the composition with the membrane for a sufficient time that the anchor lipid-modified oligonucleotide embeds itself in the membrane.

[0143] In some embodiments, the anchor lipid-modified or hydrophobic anchor oligonucleotide and the co-anchor lipid-modified or hydrophobic anchor oligonucleotide are added to cells or membranes simultaneously. In some embodiments, the anchor lipid-modified or hydrophobic anchor oligonucleotide and the co-anchor lipid-modified or hydrophobic anchor oligonucleotide are added to cells or membranes sequentially. In some embodiments, the anchor lipid-modified or hydrophobic anchor oligonucleotide is added to cells or membranes first, and then the co-anchor lipid-modified or hydrophobic anchor oligonucleotide is added to the cells or membranes. In some embodiments, the co-anchor lipid-modified or hydrophobic anchor oligonucleotide is added to cells or membranes first, and then the anchor lipid-modified or hydrophobic anchor oligonucleotide is added to the cells or membranes. In some embodiments, the anchor lipid-modified or hydrophobic anchor oligonucleotide is hybridized to the barcode oligonucleotide in advance and then added to the cells or membranes. In such embodiments, the anchor lipid-modified or hydrophobic anchor oligonucleotide is hybridized to the barcode oligonucleotide via its respective primer region.

[0144] In some embodiments, the present disclosure relates to the method for labeling cell or cell sample, the method for isolating endogenous DNA from cell sample, or the method for sequencing nucleic acid sequence from cell sample, said method comprises exposing said cell sample to one or more anchor lipid-modified oligonucleotides disclosed herein, and then adding at least the first labeled oligonucleotide sequence that is complementary to said anchor lipid-modified oligonucleotide, and said first labeled oligonucleotide sequence comprises known nucleic acid sequence portion in its 3 ' region.In some embodiments, said sample is derived from human.

[0145] In some embodiments, the first labeled oligonucleotide is complementary to the 3' region of the anchor lipid-modified oligonucleotide, and therefore, at least a portion of the known nucleic acid sequence at the 3' end is single-stranded.In some embodiments, the method further comprises exposing the single-stranded 3' end of the first labeled oligonucleotide to a ligase buffer and a ligase, covalently linking the first labeled oligonucleotide to the anchor lipid-modified oligonucleotide.The anchor lipid-modified oligonucleotide may be successively exposed to at least a second, third, or fourth or more labeled oligonucleotides, and the first, second, third, fourth or more labeled oligonucleotides each contain a unique identifying nucleic acid sequence at the 3' region of the molecule. In some embodiments, the method further comprises exposing the anchor lipid-modified oligonucleotide and the first or multiple labeled oligonucleotides to a first linker, wherein the oligonucleotide sequence is complementary to a portion of the first labeled oligonucleotide and the second labeled oligonucleotide, and thus when exposed to a ligase and free dNTPs, the first linker serves as a template nucleic acid strand for ligation and formation of a complementary nucleic acid sequence along the nucleic acid strand that is the single-stranded region at the 3 '-end of each labeled oligonucleotide. In some embodiments, the present disclosure relates to a method for labeling a cell sample, a method for isolating endogenous DNA from a cell sample, or a method for sequencing a nucleic acid sequence from a cell sample, the method comprising: (a) exposing the cell sample to one or more of the anchor lipid-modified oligonucleotides disclosed herein or a composition comprising same for a period of time sufficient for the anchor lipid-modified oligonucleotides to embed themselves within the cell membrane of the cells; (b) exposing the cell sample to one or more labeled oligonucleotides complementary to the anchor lipid-modified oligonucleotides for a time sufficient to allow the anchor lipid-modified oligonucleotides to form complementary strands of nucleic acid with the one or more labeled oligonucleotides; (c) ligating one or more of the labeled oligonucleotides to one or more of the anchor lipid-modified oligonucleotides, and optionally (d) detecting the presence of one or more of the labeled oligonucleotides by detecting one or more unique nucleotide sequences corresponding to one or more of the labeled oligonucleotides; and / or (e) isolating the cell based on the presence of one or more of the labeled oligonucleotides, wherein the presence of one or more of the labeled oligonucleotides is determined by detection of one or more unique nucleotide sequences corresponding to one or more of the labeled oligonucleotides.

[0146] In some embodiments, the method of labeling the cell sample, isolating endogenous DNA from the cell sample, or sequencing a nucleic acid sequence from the cell sample comprises: (a) exposing a cell to one or more of the anchor lipid-modified oligonucleotides for a period of time sufficient for the anchor lipid-modified oligonucleotides to embed themselves within the cell membrane of the cell; (b) exposing the cells to a first labeled oligonucleotide complementary to the anchor lipid-modified oligonucleotide for a time sufficient to allow the anchor lipid-modified oligonucleotide to form a complementary strand of nucleic acid with one or more of the labeled oligonucleotides; (c) ligating the first labeled oligonucleotide to one or more of the anchor lipid-modified oligonucleotides, and optionally (d) detecting the presence of the first labeled oligonucleotide by detecting one or more unique nucleotide sequences corresponding to the first labeled oligonucleotide; and / or (e) isolating the cell based on the presence of the first labeled oligonucleotide, wherein the presence of the first labeled oligonucleotide is determined by detection of one or more unique nucleotide sequences corresponding to the first labeled oligonucleotide.

[0147] In some embodiments, the method of labeling the cell sample, isolating endogenous DNA from the cell sample, or sequencing a nucleic acid sequence from the cell sample comprises: (a) exposing a cell to one or more of the anchor lipid-modified oligonucleotides for a period of time sufficient for the anchor lipid-modified oligonucleotides to embed themselves within the cell membrane of the cell; (b) exposing the cells to first, second, third, fourth or more labeled oligonucleotides, wherein the first labeled oligonucleotide is complementary to one or more of the anchor lipid-modified oligonucleotides and is exposed to one or more of the anchor lipid-modified oligonucleotides for a time sufficient for the anchor lipid-modified oligonucleotides to form complementary strands of nucleic acid with the first labeled oligonucleotide, and wherein the second, third, and / or fourth or more labeled oligonucleotides are sequentially exposed to the 3' portions of the labeled oligonucleotides exposed to the cells immediately followed by exposure of the second, third, and / or fourth or more labeled oligonucleotides to the cells for a time sufficient for the second, third, and / or fourth or more labeled oligonucleotides to bind covalently or non-covalently to the previously exposed labeled oligonucleotides; (c) ligating one or more of the labeled oligonucleotides to the first anchor lipid-modified oligonucleotide, and in the case of the second, third, fourth or more labeled oligonucleotides, ligating the labeled oligonucleotides simultaneously or one after the other, and optionally (d) detecting the presence of the first, second, third, fourth, and / or more labeled oligonucleotides by detecting one or more unique nucleotide sequences corresponding to each of the first, second, third, and fourth labeled oligonucleotides; and / or (e) isolating the cells based on the presence of the first, second, third, fourth and / or more labeled oligonucleotides or sequential combinations of each unique nucleotide sequence of the first, second, third, fourth and / or more labeled oligonucleotides, wherein the presence of the first, second, third, fourth and / or more labeled oligonucleotides is determined by detection of one or more unique nucleotide sequences corresponding to one or each combination of the labeled oligonucleotides.

[0148] In some embodiments, the present disclosure relates to a method for spatially localizing expression of an endogenous nucleic acid in a cell sample, isolating endogenous DNA from a cell sample, and / or sequencing an endogenously expressed nucleic acid sequence from a cell sample, the method comprising: (a) exposing a cell to one or more of the anchor lipid-modified oligonucleotides for a period of time sufficient for the anchor lipid-modified oligonucleotides to embed themselves within the cell membrane of the cell; (b) exposing the cells to first, second, third, fourth or more labeled oligonucleotides, wherein the first labeled oligonucleotide is complementary to one or more of the anchor lipid-modified oligonucleotides and is exposed to one or more of the anchor lipid-modified oligonucleotides for a time sufficient for the anchor lipid-modified oligonucleotides to form complementary strands of nucleic acid with the first labeled oligonucleotide, and wherein the second, third, and / or fourth or more labeled oligonucleotides are sequentially exposed to the 3' portions of the labeled oligonucleotides exposed to the cells immediately followed by exposure of the second, third, and / or fourth or more labeled oligonucleotides to the cells for a time sufficient for the second, third, and / or fourth or more labeled oligonucleotides to bind covalently or non-covalently to the previously exposed labeled oligonucleotides; (c) ligating one or more of the labeled oligonucleotides to the first anchor lipid-modified oligonucleotide, and in the case of the second, third, fourth or more labeled oligonucleotides, ligating the labeled oligonucleotides simultaneously or one after the other, and optionally (d) detecting the presence of the first, second, third, fourth, and / or more labeled oligonucleotides by detecting one or more unique nucleotide sequences corresponding to each of the first, second, third, and fourth labeled oligonucleotides; and / or (e) analyzing regional expression of the cell based on the spatial location of the first, second, third, fourth and / or more labeled oligonucleotides or based on sequential combinations of the unique nucleotide sequences of each of the first, second, third, fourth and / or more labeled oligonucleotides, wherein the spatial location of the cell affected by endogenous expression of the nucleic acid is determined by detecting the spatial location of the first, second, third, fourth and / or more labeled oligonucleotides corresponding to one or each combination of the labeled oligonucleotides.

[0149] In some embodiments, the cells or cell samples used in the method of labeling a cell sample, the method of isolating endogenous DNA from the cell sample, or the method of sequencing nucleic acid sequences from the cell sample of the present disclosure are obtained by dividing the cells corresponding to the region of the sample into containers.In some embodiments, the division of the cells is carried out by placing a slice or a three-dimensional preparation of a tissue sample on top of the container and positioning the cells corresponding to the region of the sample in the container.In some embodiments, the container is a well.In some embodiments, the container is a microwell.

[0150] In some embodiments, the present disclosure provides a method for spatially localizing a pattern of nucleic acid expression within a sample or tissue of a subject, the method comprising: (a) dividing one or more cells from a sample or tissue corresponding to a region of the sample into one of a plurality of containers by placing a slice or three-dimensional preparation of the sample into one of a plurality of containers; (b) exposing one or more cells corresponding to a region of the sample to known oligonucleotides disclosed herein for a time sufficient to allow incorporation of the known oligonucleotides into the one or more cells, each oligonucleotide being unique to and corresponding to one of the plurality of containers to which the one or more cells are exposed; (c) isolating and / or sequencing nucleic acid from the one or more cells in response to the known oligonucleotides; and (d) correlating the expression profile of the one or more cells to the spatial location of the cells within the sample, relative to the tissue, or within the tissue.

[0151] In such embodiments, the known oligonucleotides in each of the plurality of containers are independently selected from one or a combination of the following: (x) a first lipid-conjugated DNA oligonucleotide comprising a first lipid moiety, a first hybridization region, and a first primer region; and / or (y) a second lipid-conjugated DNA oligonucleotide comprising a second hybridization region and a second lipid moiety, wherein the second hybridization region is the reverse complement of the first hybridization region; and / or (z) a third DNA oligonucleotide comprising a second primer region, a barcode region, and a capture sequence, wherein the second primer region is the reverse complement of the first primer region.

[0152] In some embodiments, the present disclosure provides a method of identifying spatial expression patterns of a nucleic acid in a tissue of a subject, the method comprising: (a) dividing one or more cells from a sample corresponding to the region of tissue into one of a plurality of containers; (b) exposing the one or more cells to a lipid-conjugated DNA oligonucleotide comprising a lipid moiety as defined elsewhere herein, a barcode region as defined elsewhere herein, and a capture sequence as defined elsewhere herein for a time sufficient to allow the lipid moiety to embed itself within the cell membrane of the one or more cells, wherein the barcode region of the lipid-conjugated DNA oligonucleotide is unique to each one of the plurality of containers to which the one or more cells are exposed; (c) sequencing the nucleic acid captured by the capture sequence of the lipid-conjugated DNA oligonucleotide in the one or more cells; and (d) correlating the sequenced nucleic acids from the one or more cells with the spatial location of the one or more cells within and / or relative to the tissue according to the barcode region contained in each of the sequenced nucleic acids.

[0153] In some embodiments, the lipid-conjugated DNA oligonucleotide used in such methods comprises: (i) a first lipid-conjugated DNA oligonucleotide comprising a lipid moiety and a first primer region as defined elsewhere herein; and (ii) a second DNA oligonucleotide comprising a second primer region, a barcode region, and a capture sequence as defined elsewhere herein, wherein the second primer region is the reverse complement of the first primer region. In some embodiments, the first lipid-conjugated DNA oligonucleotide further comprises a first hybridization region as defined elsewhere herein. In some embodiments, such methods further comprise exposing the one or more cells to a second lipid-conjugated DNA oligonucleotide prior to the sequencing step, wherein the second lipid-conjugated DNA oligonucleotide comprises a second hybridization region as defined elsewhere herein and a second lipid moiety as defined elsewhere herein, wherein the second hybridization region is the reverse complement of the first hybridization region.

[0154] In some embodiments, the cell division is performed by placing a tissue slice or a three-dimensional tissue preparation having an appropriate thickness, as described elsewhere herein, into one of a plurality of containers. In some embodiments, the cell division is performed by placing the tissue slice or the three-dimensional tissue preparation on top of one of the plurality of containers and pressing a cell(s) corresponding to a predetermined region of the tissue sample into one of the plurality of containers. In some embodiments, one of the plurality of containers is a microwell. Microwell arrays can be created directly in a substrate such as SU-8 using soft lithography. A lithographically fabricated SU-8 mold can also be used to create microwells in PDMS. PDMS stamps can also be used as molds for other polymers, such as TPE, NOA81, and PUMA. Alternatively, arrays can be fabricated using 3D printing methods, such as Nanoscribe. Any other options known in the art for creating microwell arrays, including, but not limited to, injection molding or hot embossing, can also be used.

[0155] Depending on the size of the microwells, the spatial resolution of the wells used in the methods of the present disclosure can be about 100, 90, 80, 70, 60, 50, 40, 30, 20, or 10 microns. In some embodiments, the spatial resolution of the wells is less than about 10 microns. In some embodiments, the spatial resolution of the wells is about 10 microns. In some embodiments, the spatial resolution of the wells is about 20 microns. In some embodiments, the spatial resolution of the wells is about 30 microns. In some embodiments, the spatial resolution of the wells is about 40 microns. In some embodiments, the spatial resolution of the wells is about 50 microns. In some embodiments, the spatial resolution of the wells is about 60 microns. In some embodiments, the spatial resolution of the wells is about 70 microns. In some embodiments, the spatial resolution of the wells is about 80 microns. In some embodiments, the spatial resolution of the wells is about 90 microns. In some embodiments, the spatial resolution of the wells is about 100 microns. In some embodiments, the spatial resolution of the wells is greater than about 100 microns.

[0156] In some embodiments, the detection and / or analysis of the barcode oligonucleotides is performed by detecting or quantifying one or more of the probes or labeled oligonucleotides. For example, instruments such as the Biorad ddSEQ reader can be used to detect and analyze single or multiple cell samples in a multiplexed format. For example, Nature Biotechnology volume 37, pages 916-924 (2019) discloses known methods for quantifying, detecting, and analyzing the presence or absence of barcoded samples. Such techniques are incorporated by reference in their entirety.

[0157] In some embodiments, the cell sample is a three-dimensional preparation of cells derived from a tissue sample of a subject. In some embodiments, the cell sample is a tissue slice. In some embodiments, the cell sample has a thickness of about 0.05 microns to about 150 microns. In some embodiments, the cell sample has a thickness of about 0.1 microns to about 99 microns. In some embodiments, the cell sample has a thickness of about 0.5 microns to about 80 microns. In some embodiments, the cell sample has a thickness of about 1 micron to about 50 microns. In some embodiments, the cell sample has a thickness of about 2 microns to about 40 microns. In some embodiments, the cell sample has a thickness of about 3 microns to about 30 microns. In some embodiments, the cell sample has a thickness of about 4 microns to about 20 microns. In some embodiments, the cell sample has a thickness of about 0.05 microns. In some embodiments, the cell sample has a thickness of about 0.1 microns. In some embodiments, the cell sample has a thickness of about 0.5 microns. In some embodiments, the cell sample has a thickness of about 1 micron. In some embodiments, the cell sample has a thickness of about 5 microns. In some embodiments, the cell sample has a thickness of about 10 microns. In some embodiments, the cell sample has a thickness of about 20 microns. In some embodiments, the cell sample has a thickness of about 30 microns. In some embodiments, the cell sample has a thickness of about 40 microns. In some embodiments, the cell sample has a thickness of about 50 microns. In some embodiments, the cell sample has a thickness of about 60 microns. In some embodiments, the cell sample has a thickness of about 70 microns. In some embodiments, the cell sample has a thickness of about 80 microns. In some embodiments, the cell sample has a thickness of about 90 microns. In some embodiments, the cell sample has a thickness of about 100 microns. In some embodiments, the cell sample has a thickness of about 150 microns.In some embodiments, the cell sample has a thickness of greater than about 150 microns.

[0158] In some embodiments, the present disclosure relates to a method of labeling a plurality of cells from a sample, wherein the method comprises exposing the cells to a plurality of labeled oligonucleotides, and further comprises pooling the cells into a single container, and then subjecting the cells to sequential steps of exposure to a second, third, fourth, or more labeled oligonucleotides. In some embodiments, the method further comprises pooling the cells into a single container after subjecting the cells to sequential steps of exposure to a second, third, fourth, or more labeled oligonucleotides.

[0159] In one embodiment, microfluidic droplets are used to contain cells, for example.For example, microfluidic droplets can be used to keep the cells of the plurality of cells distinct and distinguishable, so that the differences between different cells can be distinguished.The plurality of cells, some or all of which may contain different differences, can be studied at the resolution of single cell level, for example, by using lipid-modified or hydrophobic anchor oligonucleotides disclosed herein.

[0160] The cell sample can be derived from any subject, for example, a human or a non-human animal, for example, an invertebrate cell (e.g., a cell from a fruit fly), a fish cell (e.g., a zebrafish cell), an amphibian cell (e.g., a frog cell), a reptile cell, a bird cell, or a mammalian cell, for example, a monkey, ape, cow, sheep, goat, horse, donkey, camel, llama, alpaca, rabbit, pig, mouse, rat, guinea pig, hamster, dog, cat, etc. If the cell sample is derived from a multicellular organism, the cell sample can be derived from any part of the organism. In some embodiments, tissue can be examined. For example, the tissue of an organism can be processed (e.g., by slicing the tissue) to generate a cell sample so that the differences in the tissue discussed herein can be identified.

[0161] The cell sample or tissue can be derived from a healthy subject, or can be derived from a diseased subject or a subject suspected of disease.For example, the cell sample or tissue of the subject can be taken and examined, and the difference or change in the profile of the cell sample or tissue can be identified, for example, the subject can be determined to be healthy or have disease, for example, whether the animal has cancer (for example, by measuring the cancer-specific expression profile in the cell sample or tissue).In some cases, tumor can be examined (for example, by using biopsy), and the profile of the tumor can be identified.

[0162] The droplets can be contained in a microfluidic channel. For example, in certain embodiments, the droplets can have an average size or average diameter of less than about 1 mm, less than about 500 micrometers, less than about 300 micrometers, less than about 200 micrometers, less than about 100 micrometers, less than about 75 micrometers, less than about 50 micrometers, less than about 30 micrometers, less than about 25 micrometers, less than about 10 micrometers, less than about 5 micrometers, less than about 3 micrometers, or in some cases less than about 1 micrometer. The average diameter can also in some cases be at least about 1 micrometer, at least about 2 micrometers, at least about 3 micrometers, at least about 5 micrometers, at least about 10 micrometers, at least about 15 micrometers, or at least about 20 micrometers. The droplets can be spherical or non-spherical. When the droplets are non-spherical, the average diameter or average dimension of the droplets can be considered as the diameter of a perfect sphere having the same volume as the non-spherical droplet.

[0163] The droplets may be generated using any suitable technique. For example, a channel junction may be used to create droplets. The junction may be, for example, a T-junction, a Y-junction, a channel-internal junction (e.g., coaxially arranged, or including an inner channel and an outer channel surrounding at least a portion of the inner channel), a cross (or "X") junction, a flow-focus junction, or any other junction suitable for droplet creation. See, for example, WO 2004 / 091763 and WO 2004 / 002627, each of which is incorporated herein by reference in its entirety. In some embodiments, the junction may be configured and arranged to generate substantially monodisperse droplets.

[0164] In some cases, the cells can be encapsulated within the droplets at a relatively high rate. For example, the rate of encapsulation of cells into droplets can be at least about 10 cells / second, at least about 30 cells / second, at least about 100 cells / second, at least about 300 cells / second, at least about 1,000 cells / second, at least about 3,000 cells / second, at least about 10,000 cells / second, at least about 30,000 cells / second, at least about 100,000 cells / second, at least about 300,000 cells / second, or at least about 10 6 cells / sec.

[0165] PCR reactions (e.g., including reverse transcription PCR and primer extension PCR utilizing lipid-modified or hydrophobic anchor oligonucleotides disclosed herein) can be performed using, for example, any microfluidic device (e.g., including microfluidic devices in conjunction with multi-sample nanodispensers). Microfluidic devices are fluidic systems that handle small volumes of fluid, typically on the order of microliters to nanoliters. In some embodiments, the microfluidics can process tens to thousands of samples in small volumes. Microfluidics can be active or passive. The use of active elements, such as valves, in microfluidic devices can create microfluidic circuits. This allows for high task parallelism as well as the use of small amounts of reagents, as several steps can be performed on the same chip and physically mounted on the same chip.

[0166] In microfluidic channels, liquid flow can be completely laminar, meaning all fluids move in the same direction and at the same speed. Unlike turbulent flow, this makes the transport of molecules within the fluid highly predictable. Microfluidic devices can be fabricated from glass or plastic. In some embodiments, polydimethylsiloxane (PDMS), a type of silicone, can be used. Some advantages of PDMS include its low cost, optical transparency, and permeability to several substances, including gases. In some embodiments, soft lithography or micromolding can be used to create PDMS-based microfluidic devices. The devices can use pressure-driven flow, electrodynamic flow, or wetting-driven flow.

[0167] In some embodiments, the microfluidic device has multiple chambers, each chamber having a real-time microarray. In some embodiments, the array is incorporated into the microfluidic device. In some embodiments, the microfluidic device is formed by adding features to a plane having multiple real-time microarrays to create chambers corresponding to the real-time microarrays. In some embodiments, a substrate having three-dimensional features, such as a PDMS surface with wells, microwells, or channels, is placed in contact with the surface having multiple real-time microarrays to form a microfluidic device having multiple arrays within multiple chambers.

[0168] A device having multiple chambers, each equipped with a real-time microarray, can be used to simultaneously analyze multiple samples. In some embodiments, the multiple chambers contain sample fluids derived from the same sample. Having sample fluids derived from the same sample in multiple chambers can be useful, for example, for measuring each in a different array to analyze different aspects of the same sample, or for example, for parallel measurement on the same array to increase accuracy. In some cases, different amplicons in the same sample have different optimal temperature profile conditions. Thus, in some embodiments, the same sample is divided into different fluid volumes, and the different fluid volumes are in different chambers equipped with a real-time microarray. Also, at least some of the different fluid volumes are subjected to different temperature cycles.

[0169] In some embodiments, the multiple chambers contain sample fluids from different sources. Having sample fluids from different sources can be useful for increasing throughput by measuring more samples in a given instrument at a given time. In some embodiments, multi-chambered microfluidics containing real-time microarrays can be used for diagnostic applications. The device can have about 2, 3, 4, 5, 6, 7, 8, 9, 10, 10-15, 15-20, 20-30, 30-50, 50-75, 75-100, or 100 or more chambers, each containing a real-time microarray.

[0170] Some embodiments relate to the use of the lipid-modified or hydrophobic anchor oligonucleotide in conjunction with a solid support. "Solid support" refers to any substrate having a surface to which molecules can be directly or indirectly bound, either through covalent or non-covalent bonds. The solid support can include any substrate material capable of providing physical support for the probes bound to the surface. The material is generally capable of withstanding the conditions associated with the binding of the barcode oligonucleotide to the surface and any subsequent treatments, manipulations, or processes during the performance of the assay. The material can be naturally occurring, synthetic, or modified from a naturally occurring material. Suitable solid support materials may include silicon, graphite, mirrors, laminates, ceramics, plastics (including polymers such as poly(vinyl chloride), cycloolefin copolymers, polyacrylamide, polyacrylate, polyethylene, polypropylene, poly(4-methylbutene), polystyrene, polymethacrylate, poly(ethylene terephthalate), polytetrafluoroethylene (PTFE or Teflon®), nylon, poly(vinyl butyrate)), germanium, gallium arsenide, gold, silver, and the like, used alone or in combination with other materials. Additional rigid materials may be considered, such as glass containing silica, including glass available as bioglass. Other materials that may be used may include porous materials, such as controlled pore glass beads. Any other material known in the art that can have one or more functional groups incorporated into its surface, such as amino, carboxyl, thiol, or hydroxyl functional groups, is also contemplated.

[0171] The materials used for solid supports can take on a variety of configurations, from simple to complex. The solid supports can have any one of several shapes, including strips, plates, disks, rods, particles, including beads, tubes, wells, and the like. Typically, the materials are relatively flat, such as slides, but they can also be spherical, such as beads, or cylindrical (e.g., columns). In many embodiments, the materials are generally shaped as rectangular parallelepipeds. Multiple predetermined arrangements, such as arrays of probes, can be synthesized on a sheet, which is then diced, i.e., cut by breaking along parting lines, into single array substrates. Exemplary solid supports that can be used include microtiter wells, microscope slides, membranes, magnetic beads, charged paper, Langmuir-Blodgett membranes, silicon wafer chips, flow-through chips, and microbeads. In some embodiments, the beads are plastic or polystyrene. In some embodiments, the beads are magnetic beads. After exposing the beads to magnetic force, nucleic acid molecules or sequences can be isolated from the cells.

[0172] Individual DNA and RNA In some embodiments, a barcode oligonucleotide can be directly conjugated to the solid support via its capture sequence, e.g., a poly(A) tail. In such embodiments, cells or membranes containing the anchor and, optionally, a co-anchor, lipid-modified, or hydrophobic anchor oligonucleotide can be exposed to the solid support under hybridization conditions such that the cells or membranes containing the anchor and, optionally, a co-anchor, lipid-modified, or hydrophobic anchor oligonucleotide bind to the barcode oligonucleotide. The solid support is then washed free of unbound material, and the solid support containing the remaining bound cells and membranes can be further analyzed by sequencing or other identification methods.

[0173] In some embodiments, the solid support is initially unbound and binds to the barcode oligonucleotide via the capture sequence of the barcode oligonucleotide. In such embodiments, cells or membranes containing the anchor, and optionally co-anchor, lipid-modified, or hydrophobic anchor oligonucleotide, have already hybridized to the barcode oligonucleotide via the first and second primer regions. The solid support is then washed free of unbound material, and the solid support containing the remaining bound cells and membranes can be further analyzed by sequencing or other identification methods.

[0174] In some embodiments, the present disclosure relates to a method for labeling cells with the barcodes or lipid-modified oligonucleotides disclosed herein. In some embodiments, the method comprises contacting one or more homogenous or heterogeneous mixtures of the disclosed oligonucleotides with one or more cells derived from a sample or tissue. In some embodiments, the cells are contacted with at least first and second lipid-modified oligonucleotides, such that the lipid-modified oligonucleotides hybridize with RNA, DNA, or RNA / DNA hybrids present in the cells. In some embodiments, the capture region of the oligonucleotide is used to isolate the oligonucleotide or another immobilized oligonucleotide to a solid support so that unhybridized elements or components from the cells can be removed and the remainder of the captured RNA / DNA from the cells can be preserved. In some embodiments, the captured RNA / DNA from a single cell can be isolated in a culture vessel. In some embodiments, a plurality of captured DNA / RNA sequences can be maintained in a library corresponding to the cell from which it was isolated. In some embodiments, the plurality of captured DNA / RNA sequences is sequenced so that the sequenced DNA or RNA corresponds to the expression pattern of the cell.

[0175] The present disclosure relates to a method for creating a library of oligonucleotides that are expressed by a single cell or multiple cells separately, and the method for generating the library comprises exposing the single cell or multiple cells to one or more lipid-modified oligonucleotides disclosed herein, and then sequencing the RNA or DNA from the single cell or multiple cells.In some embodiments, the endogenous nucleotides from one or multiple cells can be isolated and / or identified by correlating the known signal or frequency of the probe bound to the lipid-modified oligonucleotide with the cell to which the oligonucleotide is bound.The signal or frequency of the probe can be combined with the origin of the endogenous DNA and / or RNA.In some embodiments, the endogenous nucleotides are mRNA expressed from one or multiple cells, or cDNA formed by constructing a library of complementary strands of mRNA, for example, by PCR or other known techniques for creating a cDNA library from isolated mRNA. In some embodiments, cDNA or mRNA is isolated from a single cell by contacting the cell with one or more of the lipid-modified oligonucleotides disclosed herein, so that each captured cell can be correlated with the identity, number, or detection of the probe bound to the lipid-modified oligonucleotide(s). The differential distribution of one or more lipid-modified oligonucleotides on one or more cells can be used to correlate the cell with its specific set of endogenous expression patterns. Cells bearing one or more lipid-modified oligonucleotides can be isolated using antibodies specific to antigens or other proteins on the surface of the cells. The expression pattern of a cell (created by sequencing one or more of the endogenous sequences expressed by the cell) can be correlated with the identity of the corresponding cell by a known antigen identified by adhesion of an antibody known to bind to the antibody's target sequence. [Example]

[0176] General spatial information LMOs can also be used to barcode cells based on their relative spatial orientation. Broadly, there are two approaches to achieving spatial barcoding: (1) physical separation of cell / tissue regions followed by barcoding as previously described, and (2) the addition of lipid-modified oligonucleotide anchors to cells followed by the addition of spatially defined barcode oligonucleotides. In the first case, the physical separation can be achieved over a wide range of length scales by scalpel dissection, isolation in microwells, or laser capture microdissection. After cell separation, barcodes can be introduced into each unique sample to indicate the relative location of cells within that sample. In the second case, all cells receive anchors and co-anchors, after which barcode oligonucleotides are added. The barcodes are location-specific and diffuse from their introduction location to the anchor strands, where they are captured by cells via hybridization. The relative location of cells is determined by the amount and relative proportion of spatial barcodes. Introduction of spatial barcodes can be achieved by several methods, including microarrayers, inkjet printers, acoustic liquid handlers, and cutting from solid supports (eg, arrays or beads).

[0177] Spatial barcoding of the developing gut (prophetic) To apply MULTI-seq to mesoscale spatial barcoding of scRNA-seq samples, we first surgically resected the small intestine from a freshly euthanized and dissected mouse or from a dissected embryo. The small intestine was then stretched along its surface, after which connective tissue and fat were removed with a scalpel. The small intestine was then opened and subsequently washed four times with ice-cold PBS while shaking. After washing, each intestine was cut into equally sized fragments (i.e., approximately 1 to 100 microns in length along the proximal-distal axis for adult intestines and approximately 1 to 100 microns for developing intestines). The fragments were then dissociated at room temperature for 20 minutes with shaking in 2 mL of dissociation medium (RPMI 1640 containing 3% FBS, 1% pen / strep, 1% sodium pyruvate, 1% MEM non-essential amino acids, 1% L-glutamine, 2.5% HEPES, 5 mM EDTA, and 10 mM DTT).

[0178] Following dissociation and manual agitation with a p1000 pipette, the dissociation solution is strained through a 100 mm filter into a 15 mL conical vial on ice. The tissue mass remaining on the strainer is then transferred to another 15 mL conical vial containing 4 mL of EDTA for further digestion (e.g., vigorously shaking for 30 seconds), followed by filtering again. This process is repeated twice to produce a crudely filtered cell suspension. This crudely filtered suspension is then filtered through a 70 mm filter into a new 15 mL conical vial, followed by centrifugation at 1500 rpm for 8 minutes.

[0179] Each cell suspension was then washed once with 10 mL of ice-cold PBS and then resuspended in 160 mL of ice-cold PBS until a single-cell suspension was obtained. The single-cell suspension was then transferred to individual wells of a 48-well plate, followed by the addition of 20 mL of 2.5 mM anchor LMO pre-hybridized to a single MULTI-seq sample barcode. The cell suspension was then manually stirred and incubated on ice for 5 minutes, after which 20 mL of 2.5 mM co-anchor LMO was added, followed by another 5 minutes of incubation on ice. Following MULTI-seq labeling, the labeling solution was diluted with 300 mL of 5% BSA to quench the surrounding LMO. Live cells were then pooled, antibody stained, and FACS-enriched. The live cells were then subjected to the standard 10xGenomics scRNA-seq workflow using droplet microfluidics.

[0180] Physical separation using microwells Microwell arrays can be created directly in substrates such as SU-8 using soft lithography. Lithographically fabricated SU-8 molds can also be used to create microwells in PDMS. Furthermore, PDMS stamps can be used as molds for other polymers, such as TPE, NOA81, and PUMA. Alternatively, arrays can be fabricated using 3D printing methods, such as Nanoscribe. Other options for creating arrays are injection molding or hot embossing. The spatial resolution of the wells can be 100, 90, 80, 70, 60, 50, 40, 30, 20, or 10 microns.

[0181] After fabrication, tissue dissociation reagents can be coated onto the wells. Additionally, 20–500 nM anchors bearing different predefined barcodes or barcode combinations are added to each well to obtain spatial information. A tissue slice is placed onto the array and secured with a seal, which presses the tissue into the well. After dissociation, the seal is removed, and the array is placed in a solution containing a quenching reagent and a non-quenching co-anchor. The cells are washed and filtered through a cell strainer to select intact cells.

[0182] Cells can be further processed using any number of commercially available single-cell sequencing platforms, e.g., 10X Genomics, Bio-Rad ddSEQ, CelSee Genesis, etc. The MULTI-seq barcode or combination of barcodes associated with each cell reveals the spatial location of that cell in the original array and can be linked back to the transcriptional information from that cell.

[0183] Diffusion-based methods using oligonucleotide barcode arrays A tissue slice is placed on a microarray with oligonucleotide barcodes at sub-10 μM resolution. Anchors (100 nM–2 μM) are washed over the tissue slice and allowed to diffuse through the tissue, after which an equimolar co-anchor is added. The spatial barcodes are then released (e.g., by cleavage of the disulfide linker) and diffuse through the tissue. The tissue is then dissociated and can be further processed using any number of commercially available single-cell sequencing platforms, such as 10X Genomics, Bio-Rad ddSEQ, or CelSee Genesis. Cell location information can be determined by the relative abundance of each spatial barcode.

[0184] Alternative ways to implement barcodes for diffusion-based methods Instead of using oligonucleotide arrays, barcoding reagents can be delivered directly into the tissue-containing solution with an array of printed pins using an acoustic liquid handler, such as the Labcyte Echo or EDC ATS. Barcoding reagents can also be delivered into the tissue by flowing different barcodes along different axes into different regions of the tissue, toward a microfluidic device.

[0185] Sequence information SEQUENCE LISTING <110> The Regents of the University of California <120> LIPID-MODIFIED OLIGONUCLEOTIDES AND METHODS OF USING THE SAME <150> US 62 / 871,702 <151> 2019-07-08 <160> 7 <170> PatentIn version 3.5 <210> 1 <211> 41 <212> DNA <213> Artificial Sequence <220> <223> DNA oligonucleotide <400> 1 gtaacgatcc agctgtcact tggaattctc gggtgccaag g 41 <210> 2 <211> 20 <212> DNA <213> Artificial Sequence <220> <223> Second hybridization region <400> 2 agtgacagct ggatcgttac 20 <210> 3 <211> 59 <212> DNA <213> Artificial Sequence <220> <223> DNA oligonucleotide <220> <221> misc_feature <222> (22)..(27) <223> n is a, c, g, or t <400> 3 ccttggcacc cgagaattcc annnnnnaaa aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa <210> 4 <211> 20 <212> DNA <213> Artificial Sequence <220> <223> First hybridization region <400> 4 gtaacgatcc agctgtcact <210> 5 <211> 21 <212> DNA <213> Artificial Sequence <220> <223> First first region <400> 5 tggaattctc gggtgccaag g <210> 6 <211> 21 <212> DNA <213> Artificial Sequence <220> <223> Second first region <400> 6 ccttggcacc cgagaattcc a <210> 7 <211> 32 <212> DNA <213> Artificial Sequence <220> <223> Capture sequence <400> 7 32. AAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAA

Claims

1. (a) a first lipid-conjugated DNA oligonucleotide comprising a first lipid moiety, a first hybridization region, and a first primer region; (b) a second lipid-conjugated DNA oligonucleotide comprising a second hybridization region and a second lipid moiety, wherein the second hybridization region is the reverse complement of the first hybridization region; and (c) a third DNA oligonucleotide comprising a second primer region, a barcode region, and a capture sequence, wherein the second primer region is a reverse complement of the first primer region. A composition comprising:

2. 2. The composition of claim 1, wherein the first lipid-conjugated DNA oligonucleotide comprises, in a 5' to 3' direction, the first lipid moiety, the first hybridization region, and the first primer region.

3. 2. The composition of claim 1, wherein the first lipid-conjugated DNA oligonucleotide comprises, in a 3' to 5' direction, the first lipid moiety, the first hybridization region, and the first primer region.

4. The composition of any one of claims 1 to 3, wherein the second lipid-conjugated DNA oligonucleotide comprises, in a 5' to 3' direction, the second hybridization region and the second lipid moiety.

5. The composition of any one of claims 1 to 3, wherein the second lipid-conjugated DNA oligonucleotide comprises, in a 3' to 5' direction, the second hybridization region and the second lipid moiety.

6. The composition of any one of claims 1 to 5, wherein the third DNA oligonucleotide comprises, in a 5' to 3' direction, the second primer region, the barcode region, and the capture sequence.

7. The composition of any one of claims 1 to 5, wherein the third DNA oligonucleotide comprises, in a 3' to 5' direction, the second primer region, the barcode region, and the capture sequence.

8. 8. The composition of claim 1, wherein the first lipid portion and the second lipid portion comprise fatty acids having from about 12 to about 28 carbons.

9. The first lipid moiety is a compound of Formula I: or a physiologically acceptable salt thereof, In the formula, n 1 is 5 to 25, and n 2 is 1 to 25, and X is NH, CH 2 9. The composition of claim 1, wherein the alkyl group is selected from the group consisting of C12-C28 monoglyceride, alkenyl, alkyl, aryl, or aralkyl.

10. The second lipid moiety is a compound of Formula II: or a physiologically acceptable salt thereof, In the formula, n 1 is 5 to 25, and n 2 is from about 0 to about 24, and X is NH, CH 2 10. The composition of claim 1, wherein the alkyl group is selected from the group consisting of C12-C28 monoglyceride, alkenyl, alkyl, aryl, or aralkyl.

11. The composition of any one of claims 1 to 10, wherein the first lipid moiety comprises a lipid selected from lignoceric acid and cholesterol.

12. The composition of any one of claims 1 to 11, wherein the second lipid portion comprises a lipid selected from palmitic acid and cholesterol.

13. 13. The composition of claim 11 or 12, wherein the cholesterol is cholesterol-triethylene glycol (TEG).

14. The composition of any one of claims 1 to 13, wherein the capture sequence is a polyadenylation region.

15. The composition of any one of claims 1 to 14, wherein the first lipid-conjugated DNA oligonucleotide comprises a nucleic acid sequence having at least about 70% sequence identity to the nucleic acid sequence of SEQ ID NO: 1 (GTAACGATCCAGCTGTCACTTGGAATTCTCGGGTGCCAAGG).

16. The composition of any one of claims 1 to 15, wherein the second lipid-conjugated DNA oligonucleotide comprises a nucleic acid sequence having at least about 70% sequence identity to the nucleic acid sequence of SEQ ID NO: 2 (AGTGACAGCTGGATCGTTAC).

17. 17. The composition of any one of claims 1 to 16, wherein the third DNA oligonucleotide comprises a nucleic acid sequence having at least 70% sequence identity to the nucleic acid sequence of SEQ ID NO: 3 (CCTTGGCACCCGAGAATTCCANNNNNNNAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAA).

18. 18. The composition of any one of claims 1 to 17, wherein one or more of the first lipid-conjugated DNA oligonucleotide, the second lipid-conjugated DNA oligonucleotide, and the third lipid-conjugated DNA oligonucleotide are bound to a solid support.

19. 20. The composition of claim 18, wherein the solid support is a bead.

20. 20. The composition of claim 19, wherein the capture sequence comprises a polyadenylation region and the bead comprises a poly(T) region that hybridizes to the polyadenylation region of the third DNA oligonucleotide.

21. A composition comprising a lipid-conjugated DNA oligonucleotide comprising a lipid moiety, a barcode region, and a capture sequence.

22. (a) a first lipid-conjugated DNA oligonucleotide comprising a lipid moiety and a first primer region; and (b) a second DNA oligonucleotide comprising a second primer region, a barcode region, and a capture sequence, wherein the second primer region is a reverse complement of the first primer region. A composition comprising:

23. 1. A method for labeling a cell sample, isolating endogenous DNA from a cell sample, or sequencing a nucleic acid sequence from a cell sample, comprising: (a) exposing the cell sample to one or more of the anchor lipid-modified oligonucleotides disclosed herein or the composition of any one of claims 1 to 22 for a time sufficient for the anchor lipid-modified oligonucleotides to embed themselves within the cell membrane of the cells; (b) exposing the cell sample to one or more labeled oligonucleotides complementary to the anchor lipid-modified oligonucleotides for a time sufficient to allow the anchor lipid-modified oligonucleotides to form complementary strands of nucleic acid with the one or more labeled oligonucleotides; (c) ligating one or more of the labeled oligonucleotides to one or more of the anchor lipid-modified oligonucleotides, and optionally (d) detecting the presence of one or more of said labeled oligonucleotides by detecting one or more unique nucleotide sequences corresponding to one or more of said labeled oligonucleotides; and / or (e) isolating the cells based on the presence of one or more of the labeled oligonucleotides, wherein the presence of one or more of the labeled oligonucleotides is determined by detection of one or more unique nucleotide sequences corresponding to one or more of the labeled oligonucleotides. The method comprising:

24. 24. The method of claim 23, wherein the cell sample is a three-dimensional preparation of cells derived from a tissue sample of a subject and has a thickness of about 0.1 microns to about 99 microns.

25. 25. The method of claim 23 or 24, wherein the cell membrane is the plasma membrane that separates the cytoplasm from the exterior of the cell, or the cell membrane is the nuclear membrane.

26. 1. A method of sequencing nucleic acid from one or more cells of a sample from a subject, comprising: (a) dividing the one or more cells in a three-dimensional preparation or slice having a thickness of about 0.1 micron to about 99 microns corresponding to a region of the sample into containers; (b) labeling the one or more cells corresponding to a region of the sample with an oligonucleotide disclosed herein; (c) isolating the nucleic acid from the one or more cells; and (d) sequencing the nucleic acid from the one or more cells. The method comprising:

27. (e) compiling sequence information from each of said one or more cells.

27. The method of claim 26, further comprising:

28. 28. The method of claim 27, wherein the compiling step comprises generating an expression profile for each of the one or more cells corresponding to a region of the sample.

29. 29. The method of claim 28, further comprising correlating the sequence information and / or the expression profile of the nucleic acid with the spatial location of the one or more cells within the sample or the subject.

30. The labeling step comprises: (w) exposing the one or more cells to one or more of the anchor lipid-modified oligonucleotides for a time sufficient to allow the anchor lipid-modified oligonucleotides to embed themselves within the cell membrane of the cells; and / or (x) exposing the one or more cells to one or more labeled oligonucleotides complementary to the anchor lipid-modified oligonucleotides for a time sufficient to allow the anchor lipid-modified oligonucleotides to form complementary strands of nucleic acid with the one or more labeled oligonucleotides; and / or (y) ligating one or more of the labeled oligonucleotides to one or more of the anchor lipid-modified oligonucleotides, and optionally (z) detecting the presence of one or more of said labeled oligonucleotides by detecting one or more unique nucleotide sequences corresponding to one or more of said labeled oligonucleotides.

30. The method of any one of claims 26 to 29, comprising:

31. 31. The method of claim 30, wherein the step of isolating the cells comprises isolating the cells based on the presence of one or more of the labeled oligonucleotides, wherein the presence of one or more of the labeled oligonucleotides is determined by detection of one or more unique nucleotide sequences corresponding to one or more of the labeled oligonucleotides.

32. one or more of said anchor lipid-modified oligonucleotides (a) a first lipid-conjugated DNA oligonucleotide comprising a first lipid moiety, a first hybridization region, and a first primer region; (b) a second lipid-conjugated DNA oligonucleotide comprising a second hybridization region and a second lipid moiety, wherein the second hybridization region is the reverse complement of the first hybridization region; and (c) a third DNA oligonucleotide comprising a second primer region, a barcode region, and a capture sequence, wherein the second primer region is a reverse complement of the first primer region.

31. The method of claim 30, comprising:

33. 33. The method of claim 32, wherein the first lipid-conjugated DNA oligonucleotide comprises, in a 5' to 3' direction, the first lipid moiety, the first hybridization region, and the first primer region.

34. 33. The method of claim 32, wherein the first lipid-conjugated DNA oligonucleotide comprises, in a 3' to 5' direction, the first lipid moiety, the first hybridization region, and the first primer region.

35. 35. The method of any one of claims 32 to 34, wherein the second lipid-conjugated DNA oligonucleotide comprises, in a 5' to 3' direction, the second hybridization region and the second lipid moiety.

36. 35. The method of any one of claims 32 to 34, wherein the second lipid-conjugated DNA oligonucleotide comprises, in a 3' to 5' direction, the second hybridization region and the second lipid moiety.

37. The method of any one of claims 32 to 36, wherein the third DNA oligonucleotide comprises, in a 5' to 3' direction, the second primer region, the barcode region, and the capture sequence.

38. The method of any one of claims 32 to 36, wherein the third DNA oligonucleotide comprises, in a 3' to 5' direction, the second primer region, the barcode region, and the capture sequence.

39. 39. The method of any one of claims 32-38, wherein the first lipid portion and the second lipid portion comprise fatty acids having from about 12 to about 28 carbons.

40. The first lipid moiety is a compound of Formula I: or a physiologically acceptable salt thereof, 40. The method of any one of claims 32 to 39, wherein n1 is 5 to 25, n2 is 1 to 25, X is selected from the group consisting of NH, CH2, O, and CH-R, and R is a C12 to C28 monoglyceride, alkenyl, alkyl, aryl, or aralkyl.

41. The second lipid moiety is a compound of Formula II: or a physiologically acceptable salt thereof, 41. The method of any one of claims 32 to 40, wherein n1 is 5 to 25, n2 is 0 to 24, X is selected from the group consisting of NH, CH2, O, and CH-R, and R is a C12 to C28 monoglyceride, alkenyl, alkyl, aryl, or aralkyl.

42. 42. The method of any one of claims 32 to 41, wherein the first lipid moiety comprises a lipid selected from lignoceric acid and cholesterol.

43. 43. The method of any one of claims 32 to 42, wherein the second lipid moiety comprises a lipid selected from palmitic acid and cholesterol.

44. 44. The method of claim 42 or 43, wherein the cholesterol is cholesterol-triethylene glycol (TEG).

45. The method of any one of claims 32 to 44, wherein the capture sequence is a polyadenylation region.

46. 46. ​​The method of any one of claims 32 to 45, wherein the first lipid-conjugated DNA oligonucleotide comprises a nucleic acid sequence having at least about 70% sequence identity to the nucleic acid sequence of SEQ ID NO: 1 (GTAACGATCCAGCTGTCACTTGGAATTCTCGGGTGCCAAGG).

47. 47. The method of any one of claims 32 to 46, wherein the second lipid-conjugated DNA oligonucleotide comprises a nucleic acid sequence having at least about 70% sequence identity to the nucleic acid sequence of SEQ ID NO: 2 (AGTGACAGCTGGATCGTTAC).

48. 48. The method of any one of claims 32 to 47, wherein the third DNA oligonucleotide comprises a nucleic acid sequence having at least about 70% sequence identity to the nucleic acid sequence of SEQ ID NO: 3 (CCTTGGCACCCGAGAATTCCANNNNNNNAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAA, where N is any nucleic acid).

49. 49. The method of any one of claims 32-48, wherein one or more of the first lipid-conjugated DNA oligonucleotide, the second lipid-conjugated DNA oligonucleotide, and the third lipid-conjugated DNA oligonucleotide are bound to a solid support.

50. 50. The method of claim 49, wherein the solid support is a bead.

51. 51. The method of claim 50, wherein the capture sequence comprises a polyadenylation region and the bead comprises a poly(T) region that hybridizes to the polyadenylation region of the third DNA oligonucleotide.

52. 1. A method for determining the spatial location of a pattern of nucleic acid expression within a sample or tissue of a subject, comprising: (a) dividing one or more cells from the sample or tissue corresponding to a region of the sample into one of a plurality of containers; (b) exposing the one or more cells corresponding to a region of the sample to known oligonucleotides disclosed herein for a time sufficient to allow incorporation of the known oligonucleotides into the one or more cells, wherein each oligonucleotide is unique to and corresponds to one of the plurality of containers to which the one or more cells are exposed; (c) isolating nucleic acid from the one or more cells in response to the known oligonucleotide; (d) quantifying expression of nucleic acids from said one or more cells and / or sequencing said nucleic acids; (e) normalizing the expression of the nucleic acid in the expression profile; and (f) correlating the expression profile of the one or more cells to the spatial location of cells within the sample, with respect to the tissue, or within the tissue. Including, The known oligonucleotide disclosed herein in each of the plurality of containers comprises: (x) a first lipid-conjugated DNA oligonucleotide comprising a first lipid moiety, a first hybridization region, and a first primer region; and / or (y) a second lipid-conjugated DNA oligonucleotide comprising a second hybridization region and a second lipid moiety, wherein the second hybridization region is the reverse complement of the first hybridization region; and / or (z) a third DNA oligonucleotide comprising a second primer region, a barcode region, and a capture sequence, wherein the second primer region is the reverse complement of the first primer region. and independently selected from one or a combination of: the dividing step comprises placing a thin slice or three-dimensional preparation of the sample having a thickness of about 0.1 to about 99 microns into one of a plurality of containers; The method.

53. (g) subjecting the cells to flow cytometry 53. The method of claim 52, further comprising:

54. 54. The method of claim 52 or 53, wherein the step of isolating the cells comprises subjecting the cells to flow cytometry.

55. 1. A method of sequencing nucleic acid from one or more cells of a sample from a subject, comprising: (a) dividing the one or more cells corresponding to a region of the sample into containers; (b) labeling the one or more cells corresponding to a region of the sample with an oligonucleotide disclosed herein; (c) isolating the nucleic acid from the one or more cells; and (d) sequencing the nucleic acid from the one or more cells. Including, the dividing step comprises placing a thin slice or three-dimensional preparation of the sample having a thickness of about 0.1 to about 99 microns into one of a plurality of containers; The method.

56. (e) compiling sequence information from each of said one or more cells.

56. The method of claim 55, further comprising:

57. 57. The method of claim 56, wherein said compiling step comprises generating an expression profile for each of said one or more cells corresponding to a region of said sample.

58. 58. The method of any one of claims 55 to 57, further comprising correlating the sequence information and / or the expression profile of the nucleic acid with the spatial location of the one or more cells within the sample or the subject.

59. The labeling step comprises: (w) exposing the one or more cells to one or more of the anchor lipid-modified oligonucleotides for a time sufficient to allow the anchor lipid-modified oligonucleotides to embed themselves within the cell membrane of the cells; and / or (x) exposing the one or more cells to one or more labeled oligonucleotides complementary to the anchor lipid-modified oligonucleotides for a time sufficient to allow the anchor lipid-modified oligonucleotides to form complementary strands of nucleic acid with the one or more labeled oligonucleotides; and / or (y) ligating one or more of the labeled oligonucleotides to one or more of the anchor lipid-modified oligonucleotides, and optionally (z) detecting the presence of one or more of said labeled oligonucleotides by detecting one or more unique nucleotide sequences corresponding to one or more of said labeled oligonucleotides.

59. The method of any one of claims 55 to 58, comprising:

60. 60. The method of Claim 59, wherein isolating the nucleic acid of the one or more cells comprises isolating the one or more cells based on the presence of one or more labeled oligonucleotides, wherein the presence of the one or more labeled oligonucleotides is determined by detection of one or more unique nucleotide sequences corresponding to one or more of the labeled oligonucleotides.

61. 61. The method of any one of claims 52 to 60, wherein the first lipid-conjugated DNA oligonucleotide comprises, in a 5' to 3' direction, the first lipid moiety, the first hybridization region, and the first primer region.

62. 62. The method of any one of claims 52 to 61, wherein the first lipid-conjugated DNA oligonucleotide comprises, in a 3' to 5' direction, the first lipid moiety, the first hybridization region, and the first primer region.

63. 63. The method of any one of claims 52 to 62, wherein the second lipid-conjugated DNA oligonucleotide comprises, in a 5' to 3' direction, the second hybridization region and the second lipid moiety.

64. 63. The method of any one of claims 52 to 62, wherein the second lipid-conjugated DNA oligonucleotide comprises, in a 3' to 5' direction, the second hybridization region and the second lipid moiety.

65. 65. The method of any one of claims 52 to 64, wherein the third DNA oligonucleotide comprises, in a 5' to 3' direction, the second primer region, the barcode region, and the capture sequence.

66. 65. The method of any one of claims 52 to 64, wherein the third DNA oligonucleotide comprises, in a 3' to 5' direction, the second primer region, the barcode region, and the capture sequence.

67. 67. The method of any one of claims 52-66, wherein the first lipid portion and the second lipid portion comprise fatty acids having from about 12 to about 28 carbons.

68. The first lipid moiety is a compound of Formula I: or a physiologically acceptable salt thereof, 68. The method of any one of claims 52 to 67, wherein n1 is 5 to 25, n2 is 1 to 25, X is selected from the group consisting of NH, CH2, O, and CH-R, and R is a C12 to C28 monoglyceride, alkenyl, alkyl, aryl, or aralkyl.

69. The second lipid moiety is a compound of Formula II: or a physiologically acceptable salt thereof, 69. The method of any one of claims 52 to 68, wherein n1 is 5 to 25, n2 is 0 to 24, X is selected from the group consisting of NH, CH2, O, and CH-R, and R is a C12 to C28 monoglyceride, alkenyl, alkyl, aryl, or aralkyl.

70. 70. The method of any one of claims 52 to 69, wherein the first lipid moiety comprises a lipid selected from lignoceric acid and cholesterol.

71. 71. The method of any one of claims 52 to 70, wherein the second lipid moiety comprises a lipid selected from palmitic acid and cholesterol.

72. 72. The method of claim 70 or 71, wherein the cholesterol is cholesterol-triethylene glycol (TEG).

73. The method of any one of claims 52 to 72, wherein the capture sequence is a polyadenylation region.

74. 74. The method of any one of claims 52 to 73, wherein the first lipid-conjugated DNA oligonucleotide comprises a nucleic acid sequence having at least about 70% sequence identity to the nucleic acid sequence of SEQ ID NO: 1 (GTAACGATCCAGCTGTCACTTGGAATTCTCGGGTGCCAAGG).

75. 75. The method of any one of claims 52 to 74, wherein the second lipid-conjugated DNA oligonucleotide comprises a nucleic acid sequence having at least about 70% sequence identity to the nucleic acid sequence of SEQ ID NO: 2 (AGTGACAGCTGGATCGTTAC).

76. 76. The method of any one of claims 52 to 75, wherein the third DNA oligonucleotide comprises a nucleic acid sequence having at least about 70% sequence identity to the nucleic acid sequence of SEQ ID NO: 3 (CCTTGGCACCCGAGAATTCCANNNNNNAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAA, where N is any nucleic acid).

77. 77. The method of any one of claims 52 to 76, wherein sequencing nucleic acid expression of the one or more cells or quantifying the expression of the nucleic acid of the one or more cells comprises performing droplet microfluidics on a sequencing device.

78. 1. A method for identifying spatial expression patterns of a nucleic acid in a tissue of a subject, comprising: (a) dividing one or more cells from a sample corresponding to the region of tissue into one of a plurality of containers; (b) exposing the one or more cells with a lipid-conjugated DNA oligonucleotide comprising a lipid moiety, a barcode region, and a capture sequence for a time sufficient for the lipid moiety to embed itself within the cell membrane of the one or more cells, wherein the barcode region of the lipid-conjugated DNA oligonucleotide is unique to each one of the plurality of containers to which the one or more cells are exposed; (c) sequencing the nucleic acids captured by the capture sequences of the lipid-conjugated DNA oligonucleotides in the one or more cells; and (d) correlating the sequenced nucleic acids from the one or more cells to the spatial location of the one or more cells within and / or relative to the tissue according to the barcode region contained in each of the sequenced nucleic acids. The method comprising:

79. the lipid-conjugated DNA oligonucleotide is (i) a first lipid-conjugated DNA oligonucleotide comprising the lipid moiety and a first primer region; and (ii) a second DNA oligonucleotide comprising a second primer region, the barcode region, and the capture sequence, wherein the second primer region is a reverse complement of the first primer region.

79. The method of claim 78, comprising:

80. 80. The method of claim 79, wherein the first lipid-conjugated DNA oligonucleotide further comprises a first hybridization region.

81. further comprising exposing the one or more cells to a second lipid-conjugated DNA oligonucleotide prior to said sequencing; the second lipid-conjugated DNA oligonucleotide comprises a second hybridization region and a second lipid moiety; the second hybridization region is the reverse complement of the first hybridization region; 81. The method of claim 80.

82. 82. The method of any one of claims 78-81, wherein the first lipid portion and the second lipid portion comprise fatty acids having from about 12 to about 28 carbons.

83. The first lipid moiety is a compound of Formula I: or a physiologically acceptable salt thereof, In the formula, n 1 is 5 to 25, and n 2 is 1 to 25, and X is NH, CH 2 83. The method of any one of claims 78-82, wherein the alkyl group is selected from the group consisting of C12-C28 monoglyceride, alkenyl, alkyl, aryl, or aralkyl.

84. The second lipid moiety is a compound of Formula II: or a physiologically acceptable salt thereof, In the formula, n 1 is 5 to 25, and n 2 is 0 to 24, and X is NH, CH 2 84. The method of any one of claims 78-83, wherein the alkyl group is selected from the group consisting of C12-C28 monoglyceride, alkenyl, alkyl, aryl, or aralkyl.

85. 85. The method of any one of claims 78 to 84, wherein the first lipid moiety comprises a lipid selected from lignoceric acid and cholesterol.

86. 86. The method of any one of claims 78 to 85, wherein the second lipid moiety comprises a lipid selected from palmitic acid and cholesterol.

87. 87. The method of claim 85 or 86, wherein the cholesterol is cholesterol-triethylene glycol (TEG).

88. 88. The method of any one of claims 78 to 87, wherein the capture sequence is a polyadenylation region.

89. 89. The method of any one of claims 78 to 88, wherein the first lipid-conjugated DNA oligonucleotide comprises a nucleic acid sequence having at least about 70% sequence identity to the nucleic acid sequence of SEQ ID NO: 1 (GTAACGATCCAGCTGTCACTTGGAATTCTCGGGTGCCAAGG).

90. 90. The method of any one of claims 78 to 89, wherein the second lipid-conjugated DNA oligonucleotide comprises a nucleic acid sequence having at least about 70% sequence identity to the nucleic acid sequence of SEQ ID NO: 2 (AGTGACAGCTGGATCGTTAC).

91. 91. The method of any one of claims 78 to 90, wherein the third DNA oligonucleotide comprises a nucleic acid sequence having at least 70% sequence identity to the nucleic acid sequence of SEQ ID NO: 3 (CCTTGGCACCCGAGAATTCCANNNNNNNAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAA).

92. (e) subjecting the cells to flow cytometry 92. The method of any one of claims 78 to 91, further comprising:

93. 93. The method of any one of claims 78-92, wherein the dividing step comprises placing a tissue slice or three-dimensional preparation of the sample having a thickness of from about 0.1 to about 99 microns into one of the plurality of containers.

94. 94. The method of any one of claims 78 to 93, wherein the one or more containers are multi-wells.

95. 95. The method of any one of claims 78 to 94, wherein the sample is a tissue slice.