RNA detection and amplification method for single cell analysis on fixed cells
By using more than one probe oligonucleotide to contact the nucleic acid target and barcode it, the problem of gene expression analysis of fixed samples in the existing technology is solved, and accurate quantification of nucleic acid target copy number and spatial location is achieved, improving the accuracy and efficiency of analysis.
Patent Information
- Application Number
- CN202480033772.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-05-23
- Filing Date
- 2024-05-22
- Publication Date
- 2025-12-16
AI Technical Summary
Existing gene expression profiling methods are incompatible with fixed samples, cannot effectively analyze mRNA in fixed samples, and are difficult to perform spatial analysis of gene expression in fixed samples.
Multiple probe oligonucleotides are contacted with the nucleic acid target, and multiple barcoded probe oligonucleotides are generated through extension and barcoding to obtain sequencing data to determine the copy number and spatial location of the nucleic acid target. Specific reagents and devices are then used for sample processing and amplification.
It enables accurate quantification of the copy number and spatial location of nucleic acid targets in fixed samples, supports gene expression analysis of fixed samples, and improves the accuracy and efficiency of the analysis.
Smart Images

Figure CN121152885A_ABST
Abstract
Description
Related Applications
[0001] This application claims the benefit under 35 U.S.C. § 119(e) of U.S. Provisional Patent Application Serial No. 63 / 503,856, filed May 23, 2023, the content of which is incorporated by reference herein in its entirety for all purposes. BACKGROUND
[0002] Field The present disclosure relates generally to the field of molecular biology, for example, determining gene expression using molecular barcoding.
[0003] Description of the Related Art Many cell samples contain fixed cells, including patient samples that need to be stored in a fixative. However, current gene expression profiling compositions and methods, including single cell RNA sequencing workflows and platforms, are not compatible with fixed samples because the mRNA is cross-linked and not available for capture by conventional mRNA capture methods (e.g., via poly A / dT capture). In situ hybridization and FISH methods have visualized mRNA in fixed samples; however, the number of probes that can be used to detect mRNA is limited due to the limited number of fluorescence detectors, making transcriptome level analysis difficult. There is a need for compositions, methods, systems, and kits that can analyze gene expression in fixed samples. Further, there is a need for compositions, methods, systems, and kits for spatial analysis of gene expression in fixed samples. SUMMARY
[0004] The disclosure herein includes methods for labeling nucleic acid targets in a sample. In some embodiments, the method includes: contacting a sample comprising copies of a nucleic acid target with a plurality of probe oligonucleotides, wherein each probe oligonucleotide comprises a coupling sequence and a probe sequence configured to hybridize to the nucleic acid target. The method can include: extending the plurality of probe oligonucleotides hybridized to copies of the nucleic acid target to generate a plurality of extended probe oligonucleotides each comprising a sequence complementary to at least a portion of the nucleic acid target. The method can include: barcoding the plurality of extended probe oligonucleotides or products thereof using a plurality of oligonucleotide barcodes to generate a plurality of barcoded probe oligonucleotides, wherein each oligonucleotide barcode of the plurality of oligonucleotide barcodes comprises a molecular label, and wherein each of the plurality of barcoded probe oligonucleotides comprises the molecular label, the probe sequence, and the sequence complementary to at least a portion of the nucleic acid target. The method can include: obtaining sequencing data comprising a plurality of sequencing reads of the barcoded probe oligonucleotides or products thereof, wherein each of the plurality of sequencing reads comprises a molecular label sequence and a subsequence of the nucleic acid target. The method can include: determining a copy number of the nucleic acid target in the sample based on a number of molecular labels associated with the plurality of barcoded probe oligonucleotides or products thereof.
[0005] The disclosure herein includes methods for determining a copy number of a nucleic acid target in a sample. In some embodiments, the method includes: contacting a sample comprising copies of a nucleic acid target with a plurality of probe oligonucleotides, wherein each probe oligonucleotide comprises a coupling sequence and a probe sequence configured to hybridize to the nucleic acid target. The method can include: extending the plurality of probe oligonucleotides hybridized to copies of the nucleic acid target to generate a plurality of extended probe oligonucleotides each comprising a sequence complementary to at least a portion of the nucleic acid target. The method can include: barcoding the plurality of extended probe oligonucleotides or products thereof using a plurality of oligonucleotide barcodes to generate a plurality of barcoded probe oligonucleotides, wherein each oligonucleotide barcode of the plurality of oligonucleotide barcodes comprises a molecular label, and wherein each of the plurality of barcoded probe oligonucleotides comprises the molecular label, the probe sequence, and the sequence complementary to at least a portion of the nucleic acid target. The method can include: obtaining sequencing data comprising a plurality of sequencing reads of the barcoded probe oligonucleotides or products thereof, wherein each of the plurality of sequencing reads comprises a molecular label sequence and a subsequence of the nucleic acid target. The method can include: determining a copy number of the nucleic acid target in the sample based on a number of molecular labels associated with the plurality of barcoded probe oligonucleotides or products thereof.
[0006] The present disclosure includes methods for determining spatial location and copy number of a nucleic acid target in a sample. In some embodiments, the method comprises: contacting each of two or more spatial locations of a sample comprising copies of a nucleic acid target with a plurality of probe oligonucleotides, wherein each probe oligonucleotide comprises a coupling sequence, a probe sequence configured to hybridize to the nucleic acid target, and a predetermined spatial marker. In some embodiments, the probe oligonucleotides contacted with the same spatial location comprise the same spatial marker sequence, and wherein the probe oligonucleotides contacted with different spatial locations of the sample comprise different spatial marker sequences. The method can comprise: extending the plurality of probe oligonucleotides hybridized to copies of the nucleic acid target to generate a plurality of extended probe oligonucleotides each comprising a sequence complementary to at least a portion of the nucleic acid target. The method can comprise: barcoding the plurality of extended probe oligonucleotides or products thereof using a plurality of oligonucleotide barcodes to generate a plurality of barcoded probe oligonucleotides, wherein each oligonucleotide barcode of the plurality of oligonucleotide barcodes comprises a molecular marker, and wherein each of the plurality of barcoded probe oligonucleotides comprises a molecular marker, a probe sequence, and a sequence complementary to at least a portion of the nucleic acid target. The method can comprise: obtaining sequencing data comprising a plurality of sequencing reads of the barcoded probe oligonucleotides or products thereof, wherein each of the plurality of sequencing reads comprises a spatial marker sequence, a molecular marker sequence, and a subsequence of the nucleic acid target. The method can comprise: for each unique spatial marker sequence, which is associated with a different spatial location of the sample, counting the number of molecular markers having different sequences associated with the nucleic acid target to determine the copy number of the nucleic acid target at each spatial location of the sample.
[0007] In some embodiments, barcoding more than one extended probe oligonucleotide or its product using more than one oligonucleotide barcode includes: providing a coupled oligonucleotide comprising a 5' complement of a coupling sequence and a 3' complement of a capture sequence; hybridizing the coupling sequence of the extended probe oligonucleotide with the 5' complement of the coupling sequence of the coupled oligonucleotide; hybridizing the 3' complement of the capture sequence of the coupled oligonucleotide with the capture sequence of the oligonucleotide barcode in more than one oligonucleotide barcode; and / or ligating the extended probe oligonucleotide to the hybridized oligonucleotide barcode. The method may include: filling the gap between the extended probe oligonucleotide and the hybridized oligonucleotide barcode with a DNA polymerase lacking at least one of 5' to 3' exonuclease activity and 3' to 5' exonuclease activity before ligating the extended probe oligonucleotide to the oligonucleotide barcode. In some embodiments, the ligation of the extended probe oligonucleotide to the hybridized oligonucleotide barcode is performed using a DNA ligase. In some embodiments, the coupled oligonucleotide is a single-stranded oligonucleotide. In some embodiments, the coupled oligonucleotide comprises at least 6 nucleotides. In some embodiments, the coupling sequence comprises at least 4 nucleotides. In some embodiments, the 5' end of each probe oligonucleotide is phosphorylated. In some embodiments, the probe oligonucleotide is capable of entering the cells and / or nuclei of the sample (e.g., permeable cells and / or permeable nuclei of the sample). The method may include: after contacting the probe oligonucleotide with the sample, removing one or more probe oligonucleotides from more than one probe oligonucleotide that have not been contacted with the sample; optionally, removing one or more probe oligonucleotides that have not been contacted with the sample includes: removing one or more probe oligonucleotides that have not entered the cells of the sample. In some embodiments, the contacting step includes contacting the sample with a device (e.g., an inkjet device) configured to deposit the probe oligonucleotide. In some embodiments, the device is a needle, needle array, tube, aspiration device, injection device, electroporation device, fluorescence-activated cell sorting device, inkjet device, microfluidic device, or any combination thereof. In some embodiments, the device contacts different spatial locations of the sample at a specified rate. In some embodiments, the spatial marker is 6-60 nucleotides in length. In some embodiments, the two or more spatial locations include at least about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 12, about 14, about 16, about 18, about 20, about 30, about 40, about 50, about 60, about 70, about 80, about 90, or about 100 different spatial locations of the sample.In some implementations, the spatial location of the sample corresponds to a region containing no more than about 50 cells, about 45 cells, about 40 cells, about 35 cells, about 30 cells, about 25 cells, about 20 cells, about 15 cells, about 10 cells, about 9 cells, about 8 cells, about 7 cells, about 6 cells, about 5 cells, about 4 cells, about 3 cells, about 2 cells, or about 1 cell.
[0008] The method may include: contacting a sample with an extension reagent. In some embodiments, at least a portion of the contacting step is performed in the presence of the extension reagent. In some embodiments, the entire contacting step is performed in the presence of the extension reagent. In some embodiments, the contacting step and the extension step are simultaneous. In some embodiments, the extension is performed in situ. In some embodiments, the extension includes in situ reverse transcription. In some embodiments, the cells of the sample remain intact during the extension step. In some embodiments, the extension reagent includes a reverse transcription reagent. In some embodiments, the reverse transcription reagent comprises reverse transcriptase and dNTPs. In some embodiments, the reverse transcriptase includes viral reverse transcriptase. In some embodiments, the viral reverse transcriptase is murine leukemia virus (MLV) reverse transcriptase or Moloney murine leukemia virus (MMLV) reverse transcriptase.
[0009] In some embodiments, the sample is physically fragmented or remains intact during the contact step. In some embodiments, the sample comprises a single cell. In some embodiments, the sample comprises more than one single cell. In some embodiments, the sample comprises more than one cell, and optionally the method includes: dissociating the sample to produce more than one single cell, optionally said dissociation including chemical dissociation, enzymatic dissociation, and / or mechanical dissociation, optionally said dissociation employing one or more of collagenase, chymotrypsin, dispersin, elastase, hyaluronidase, trypsin, papain, and trypsin. The method may include: prior to the barcoding step: partitioning more than one single cell into more than one partition, wherein the partitions of the more than one partition contain single cells from more than one single cell; and in the partition containing the single cell, contacting an extended probe oligonucleotide with more than one oligonucleotide barcode. In some embodiments, in the partition containing the single cell, contacting the single cell with a lysis buffer at 15°C–65°C to lyse the single cell. In some embodiments, the lysis buffer contains an agent capable of dissociating protein-nucleic acid complexes.
[0010] In some embodiments, each oligonucleotide barcode in more than one oligonucleotide barcode contains a first universal sequence. In some embodiments, obtaining sequencing data includes amplifying more than one barcoded probe oligonucleotide using a first primer capable of hybridizing with the first universal sequence or its complement and an amplification primer capable of hybridizing with a nucleic acid target or its complement, thereby generating more than one amplified barcoded probe oligonucleotide, wherein obtaining sequencing data includes obtaining sequencing data containing more than one sequencing read comprising the amplified barcoded probe oligonucleotide or its product. In some embodiments, obtaining sequencing data includes attaching binding sites of sequencing primers and / or sequencing adaptors to more than one barcoded probe oligonucleotide or its product. In some embodiments, the amplification primers contain a second universal sequence and / or wherein the first primer contains a third universal sequence. In some embodiments, the first universal sequence, the second universal sequence, and / or the third universal sequence are the same. In some embodiments, the first universal sequence, the second universal sequence, and / or the third universal sequence are different. In some embodiments, the first universal sequence, the second universal sequence, and / or the third universal sequence comprises the binding site of the sequencing primer and / or the sequencing adaptor, its complementary sequence, and / or a portion thereof. In some embodiments, the sequencing adaptor comprises the P5 sequence, the P7 sequence, its complementary sequence, and / or a portion thereof. In some embodiments, the sequencing primer comprises a read 1 sequencing primer, a read 2 sequencing primer, its complementary sequence, and / or a portion thereof.
[0011] In some embodiments, the sample contains more than one nucleic acid target, such as, for example, a target group of at least about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 12, about 14, about 16, about 18, about 20, about 30, about 40, about 50, about 60, about 70, about 80, about 90, about 100, about 125, about 150, about 175, about 200, about 225, about 250, about 275, about 300, about 325, about 350, about 375, about 400, about 425, about 450, about 475, or about 500 different nucleic acid targets. In some embodiments, two or more nucleic acid targets in the target group are biomarkers. In some embodiments, the biomarker is a biomarker of a disease or condition. In some embodiments, the disease or condition is cancer, infection, viral infection, inflammatory disease, neurodegenerative disease, fungal disease, bacterial infection, or any combination thereof. In some embodiments, the contact step includes contacting the sample with a group of probe oligonucleotides comprising two or more probe oligonucleotides, wherein each or more of them comprises a probe sequence configured to hybridize with nucleic acid targets in more than one nucleic acid target. In some embodiments, determining the copy number of a nucleic acid target in the sample comprises determining the copy number of each of more than one nucleic acid target in the sample based on the number of molecular markers having distinct sequences associated with more than one barcoded probe oligonucleotide or its product, said more than one barcoded probe oligonucleotide or its product comprising sequences of more than one nucleic acid target. The method may include: for each unique spatial marker sequence associated with a different spatial location in the sample, counting the number of molecular markers having distinct sequences associated with each of more than one nucleic acid target to determine the copy number of each of more than one nucleic acid target at each spatial location in the sample. In some embodiments, the amplification primers include a set of amplification primers configured to hybridize with more than one nucleic acid target or its complement, such as, for example, a set of at least about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 12, about 14, about 16, about 18, about 20, about 30, about 40, about 50, about 60, about 70, about 80, about 90, about 100, about 125, about 150, about 175, about 200, about 225, about 250, about 275, about 300, about 325, about 350, about 375, about 400, about 425, about 450, about 475, or about 500 different amplification primers. In some embodiments, the nucleic acid target includes a nucleic acid molecule.In some implementations, nucleic acid molecules include ribonucleic acid (RNA), messenger RNA (mRNA), microRNA, small interfering RNA (siRNA), RNA degradation products, RNA containing multiple (A) tails, sample index oligonucleotides, cell component binding reagent-specific oligonucleotides, or any combination thereof.
[0012] In some embodiments, more than one cell comprises one or more cell types. In some embodiments, the one or more cell types are selected from the group consisting of: brain cells, heart cells, cancer cells, circulating tumor cells, organ cells, epithelial cells, metastatic cells, benign cells, primary cells, and circulating cells, or any combination thereof. In some embodiments, the sample comprises a biological sample, a clinical sample, an environmental sample, a biological fluid, tissue, a tissue section derived from a subject, or any combination thereof. In some embodiments, the subject is a human, mouse, dog, rat, or vertebrate. The method may include: determining the subject's genotype, phenotype, or one or more gene mutations based on the spatial location of nucleic acid targets in the sample. The method may include: predicting the subject's susceptibility to one or more diseases, such as, for example, cancer or a hereditary disease. The method may include: determining the cell type of more than one cell in the sample. In some embodiments, selecting a drug based on the predicted reactivity of the cell type of more than one cell in the sample. The method may include: imaging the sample, optionally before and / or after the contact step, optionally generating imaging data. In some embodiments, imagery of the sample includes staining the sample with a staining agent, wherein the staining agent is a fluorescent staining agent, a negative staining agent, an antibody staining agent, or any combination thereof. In some embodiments, staining includes immunocytochemistry (ICC), immunohistochemistry (IHC), immunofluorescence (IF), or any combination thereof. In some embodiments, imaging includes microscopy, confocal microscopy, time-lapse imaging microscopy, fluorescence microscopy, multiphoton microscopy, quantitative phase microscopy, surface-enhanced Raman spectroscopy, photography, manual visual analysis, automated visual analysis, or any combination thereof. The method may include: correlating imaging data and sequencing data at one or more spatial locations of the sample. The method may include: correlation analysis of imaging data and sequencing data at spatial locations. In some embodiments, the correlation analysis identifies one or more of the following: candidate biomarkers, candidate therapeutic agents, candidate doses of therapeutic agents, and / or cellular targets of candidate therapeutic agents. In some embodiments, the imaging produces an image for constructing a physical representation of the sample. In some embodiments, the image is two-dimensional or three-dimensional. This method may include mapping nucleic acid targets and / or cellular component targets onto a sample atlas. This method may also include mapping one or more single cells from a plurality of cells onto a sample atlas.
[0013] In some embodiments, the sample has been contacted with one or more fixatives and / or permeabilizers. In some embodiments, the sample comprises tissue, cell monolayers, fixed cells, tissue sections, or any combination thereof. In some embodiments, the sample comprises fresh tissue sections, frozen tissue sections, fixed tissue sections, formalin-fixed tissue sections, formalin-fixed paraffin-embedded (FFPE) tissue sections, acetone-fixed tissue sections, paraformaldehyde (PFA)-fixed tissue sections, and / or methanol-fixed tissue sections. In some embodiments, the sample comprises a cell nuclear suspension, such as, for example, a fixed cell nuclear suspension and / or a permeabilized cell nuclear suspension. In some embodiments, the sample comprises cells, such as, for example, fresh cells, frozen cells, fixed cells, formalin-fixed cells, formalin-fixed paraffin-embedded (FFPE) cells, acetone-fixed cells, paraformaldehyde (PFA)-fixed cells, and / or methanol-fixed cells. The method may include: permeabilizing the sample and / or fixing the sample. In some embodiments, fixing the sample includes contacting the sample with a fixative. In some embodiments, the fixative comprises a non-crosslinking fixative (e.g., methanol). In some embodiments, the fixative includes a crosslinking agent. In some embodiments, the crosslinking agent includes a degradable crosslinking agent. In some embodiments, the degradable crosslinking agent includes or is derived from dithiobis(succinimide propionate) (DSP), disuccinimide tartrate (DST), bis[2-(succinimideoxycarbonyloxy)ethyl] sulfone (BSOCOES), ethylene glycol bis(succinimide succinate) (EGS), dimethyl 3,3'-dithiobispropionylimide (DTBP), and succinimide 3-(2-pyridyldithio)propionate (SPD). P), succinimidyl 6-(3(2-pyridyldithio)propionamido)hexanoate (LC-SPDP), 4-succinimidyloxycarbonyl-α-methyl-α-(2-pyridyldithio)toluene (SMPT), 3-(2-pyridyldithio)propionylhydrazine (PDPH), succinimidyl 2-((4,4'-azidopentamido)ethyl)-1,3'-dithiopropionate (SDAD, NHS-SS-diazadiazide), or any combination thereof. In some embodiments, the cleavable crosslinker includes cleavable links selected from the group consisting of: chemically cleavable links, photocleavable links, acid-labile links, heat-sensitive links, enzyme-cleavable links, and any combination thereof. In some embodiments, the cleavable crosslinker is a thiol-cleavable crosslinker or contains a disulfide linker. In some embodiments, the fixative includes paraformaldehyde (PFA), dithiobis(succinimide propionate) (DSP), succinimide-3-(2-pyridyldithio)propionate (SPDP), CellCover, or combinations thereof. In some embodiments, sample fixation and permeation are performed simultaneously.In some embodiments, sample fixation and permeation are performed in the presence of a dual-functional agent capable of both fixing and permeating the sample. In some embodiments, the dual-functional agent is methanol.
[0014] In some embodiments, permeabilizing the sample includes contacting the sample with a permeabilizing agent. This method may include removing the permeabilizing agent from the sample after contacting it with more than one probe oligonucleotide or more than one cell component binding reagent. In some embodiments, the permeabilizing agent is capable of (i) permeating the cell membrane of a cell, and (ii) making the cell membrane of the cell permeable to the probe oligonucleotide or cell component binding reagent, or both. In some embodiments, the permeabilizing agent includes (i) a solvent, detergent, or surfactant; (ii) BD Cytoperm; (iii) a saponin or a derivative thereof; (iv) Triton X-100; (v) methanol or a derivative thereof; and / or (vi) digitalis saponin or a derivative thereof. In some embodiments, the agent capable of dissociating the protein-nucleic acid complex includes a broad-spectrum serine protease. In some embodiments, the broad-spectrum serine protease is proteinase K. In some embodiments, the lysis buffer includes a defixing agent. In some embodiments, the defixing agent includes thiols, hydroxylamine, periodate, bases, or any combination thereof. In some embodiments, the lysis buffer contains DTT. This method may include reversing the fixation of a sample and / or a single cell. In some embodiments, reversing the fixation of a sample and / or a single cell includes UV photolysis, chemical treatment, heating, enzymatic treatment, or any combination thereof.
[0015] The sample may contain more than one cell component target, and the method further includes: contacting the sample with more than one cell component binding agent, wherein each of the more than one cell component binding agent contains a cell component binding agent-specific oligonucleotide, the cell component binding agent-specific oligonucleotide containing a unique identifier sequence of the cell component binding agent, and wherein the cell component binding agent is capable of specifically binding to at least one of the more than one cell component target; barcoding the cell component binding agent-specific oligonucleotide to generate more than one barcoded cell component binding agent-specific oligonucleotide, each of the more than one barcoded cell component binding agent-specific oligonucleotide containing a sequence complementary to at least a portion of the unique identifier sequence and a molecular marker sequence; and obtaining sequencing data comprising more than one sequencing read containing more than one barcoded cell component binding agent-specific oligonucleotide or its product, wherein each of the more than one sequencing read contains at least a portion of the molecular marker sequence and the unique identifier sequence. In some embodiments, obtaining sequencing data includes attaching the binding sites of sequencing primers and / or sequencing adaptors to the barcoded cell component binding agent-specific oligonucleotide or its product.
[0016] The method may include: after contacting a sample with more than one cell component binding agent, removing one or more cell component binding agents from the more than one cell component binding agent that have not been contacted with the sample; optionally, removing one or more cell component binding agents that have not been contacted with the sample includes removing one or more cell component binding agents that have not been contacted with at least one corresponding cell component target. In some embodiments, cell component targets include intracellular proteins, carbohydrates, lipids, proteins, extracellular proteins, cell surface proteins, cell markers, B cell receptors, T cell receptors, major histocompatibility complex, tumor antigens, receptors, intracellular proteins, or any combination thereof. In some embodiments, the cell component binding agent-specific oligonucleotides contain a second molecular marker, optionally at least 10 of the more than one cell component binding agent-specific oligonucleotides contain different second molecular marker sequences. In some embodiments, at least two cell component binding agent-specific oligonucleotides have different second molecular marker sequences, and the unique identifier sequences of said at least two cell component binding agent-specific oligonucleotides are identical. In some embodiments, the second molecular marker sequences of at least two cell component binding reagent-specific oligonucleotides are different, and the unique identifier sequences of said at least two cell component binding reagent-specific oligonucleotides are different. In some embodiments, the number of unique molecular marker sequences associated with the unique identifier sequences for the cell component binding reagent in the sequencing data indicates the copy number of at least one cell component target in the sample, said cell component binding reagent being capable of specifically binding to at least one cell component target. In some embodiments, the number of unique second molecular marker sequences associated with the unique identifier sequences for the cell component binding reagent in the sequencing data indicates the copy number of at least one cell component target in the sample, said cell component binding reagent being capable of specifically binding to at least one cell component target.
[0017] The method may include contacting the sample with a blocking agent, one or more decoy oligonucleotides, and / or one or more blocking oligonucleotides before contacting the sample with more than one cell component binding agent and / or contacting the sample with more than one probe oligonucleotide. In some embodiments, contacting the sample with more than one cell component binding agent is performed in the presence of a blocking agent. In some embodiments, the blocking agent comprises more than one oligonucleotide complementary to at least a portion of a cell component binding agent-specific oligonucleotide. In some embodiments, the blocking agent comprises an antibody or fragment thereof derived from a first species, and wherein the blocking agent comprises serum derived from the first species. In some embodiments, the sample comprises one or more non-target nucleic acids, wherein the blocking agent comprises more than one decoy oligonucleotide capable of hybridizing with at least one of the one or more non-target nucleic acids. In some embodiments, each of the more than one decoy oligonucleotide is capable of hybridizing with at least a portion of the non-target nucleic acid. In some embodiments, the decoy oligonucleotide: comprises a sequence complementary to at least a portion of a non-target nucleic acid; comprises a sequence identical or substantially similar to that of a cell component binding reagent-specific oligonucleotide, optionally having a length of 3 to 40 nucleotides; has at most 50% sequence identity with the cell component binding reagent-specific oligonucleotide; does not contain a UMI; comprises a random sequence, optionally having a length of about four, five, six, seven, eight, nine, ten, eleven, twelve, thirteen, fourteen, or fifteen nucleotides; does not contain any sequence having more than four, five, six, or seven consecutive T or A; contains at least one G or C in every four, five, six, or seven consecutive nucleotides; comprises one or more modified nucleotides; comprises a 5' modification, optionally comprising a 5' amino-modified C12 modification (5AmMC12); comprises a 3' modification, optionally comprising a 3' dideoxy-C modification (ddC); and / or has a length of 30 to 65 nucleotides. In some embodiments, the sample contains one or more undesirable nucleic acid substances, and the method includes: contacting a blocking oligonucleotide with the sample, wherein the blocking oligonucleotide specifically binds to at least one of the one or more undesirable nucleic acid substances; wherein reverse transcription of at least one of the one or more undesirable nucleic acid substances is reduced by the blocking oligonucleotide. In some embodiments, the blocking oligonucleotide is contacted with the sample: before contacting more than one probe oligonucleotide with the sample; after contacting more than one probe oligonucleotide with the sample; and / or when contacting more than one probe oligonucleotide with the sample. The method may include: providing a blocking oligonucleotide that specifically binds to two or more undesirable nucleic acid substances in the sample, optionally at least 10 or up to at least 100 undesirable nucleic acid substances.In some implementations, the blocking oligonucleotide is: locked nucleic acid (LNA), peptide nucleic acid (PNA), DNA, LNA / PNA chimera, LNA / DNA chimera, or PNA / DNA chimera; specifically binds to the 3' end of one or more undesirable nucleic acid substances within 100 nt, 50 nt, or 25 nt; specifically binds to the 5' end of one or more undesirable nucleic acid substances within 100 nt, or the blocking oligonucleotide specifically binds to the middle of one or more undesirable nucleic acid substances within 100 nt; contains or does not contain non-natural nucleotides; has a Tm of at least 50°C, at least 60°C, or at least 70°C; cannot be used as a primer for reverse transcriptase or polymerase; and / or is 8 nt to 100 nt long, 10 nt to 50 nt long, 12 nt to 21 nt long, 20 nt to 30 nt long, or about 25 nt long. In some embodiments, one or more undesirable nucleic acid substances constitute approximately 50%, approximately 60%, approximately 70%, or approximately 80% of the nucleic acid content of the sample. In some embodiments, the undesirable nucleic acid substances are selected from the group consisting of ribosomal RNA, mitochondrial RNA, genomic DNA, intron sequences, high-abundance sequences, and combinations thereof. In some embodiments, one or more undesirable nucleic acid substances are mRNA molecules, and the blocking oligonucleotides specifically bind to the 3' multi(A) tail of one or more undesirable nucleic acid substances within 10 nt.
[0018] In some embodiments, each molecular marker of more than one oligonucleotide barcode contains at least 6 nucleotides. In some embodiments, each capture sequence of more than one oligonucleotide barcode contains at least 4 nucleotides. In some embodiments, more than one oligonucleotide barcode is associated with a solid support, and wherein a partition in more than one partition contains a single solid support. In some embodiments, each of the more than one oligonucleotide barcodes contains a cell marker. In some embodiments, each cell marker of more than one oligonucleotide barcode contains at least 6 nucleotides. In some embodiments, oligonucleotide barcodes associated with the same solid support in more than one oligonucleotide barcode contain the same cell marker. In some embodiments, oligonucleotide barcodes associated with different solid supports in more than one oligonucleotide barcode contain different cell markers. In some embodiments, the solid support comprises synthetic particles, a flat surface, or a combination thereof. The method may include associating synthetic particles containing more than one oligonucleotide barcode with cells in a partition. The method may include lysing cells after associating the synthetic particles with cells. In some embodiments, cell lysis includes heating the cells, contacting the cells with a detergent, altering the pH of the cells, or any combination thereof. In some embodiments, the synthetic particle and the single cell are in the same compartment, and optionally this compartment is a pore or a droplet. In some embodiments, at least one oligonucleotide barcode of more than one oligonucleotide barcode is immobilized or partially immobilized on the synthetic particle, or at least one oligonucleotide barcode of more than one oligonucleotide barcode is encapsulated or partially encapsulated within the synthetic particle. In some embodiments, the synthetic particle is destructible (e.g., a destructible hydrogel particle). In some embodiments, the synthetic particle comprises beads. In some embodiments, the beads comprise agarose gel (Sepharose) beads, streptoacidin beads, agarose beads, magnetic beads, conjugated beads, protein A conjugated beads, protein G conjugated beads, protein A / G conjugated beads, protein L conjugated beads, oligo(dT) conjugated beads, silica beads, silica-like beads, avidin microbeads, anti-fluorescent dye microbeads, or any combination thereof. In some embodiments, the synthetic particles comprise materials selected from the group consisting of: polydimethylsiloxane (PDMS), polystyrene, glass, polypropylene, agarose, gelatin, hydrogel, paramagnetic materials, ceramics, plastics, glass, methylstyrene, acrylic polymers, titanium, latex, agarose gel, cellulose, nylon, silicone, and any combination thereof. In some embodiments, each oligonucleotide barcode in more than one oligonucleotide barcode comprises a linker functional group. In some embodiments, the synthetic particles comprise solid support functional groups.In some embodiments, the support functional group and the connector functional group are associated with each other, and optionally the connector functional group and the support functional group are individually selected from the group consisting of: C6, biotin, streptavidin, one or more primary amines, one or more aldehydes, one or more ketones, and any combination thereof.
[0019] The disclosure herein includes compositions (e.g., kits). In some embodiments, the kit comprises: more than one probe oligonucleotide, wherein each probe oligonucleotide comprises a coupling sequence and a probe sequence configured to hybridize with a nucleic acid target, optionally, the probe oligonucleotide comprises a predetermined spatial marker; a coupling oligonucleotide comprising a 5' complement of the coupling sequence and a 3' complement of a capture sequence; more than one oligonucleotide barcode, wherein the 3' end of each of the more than one oligonucleotide barcodes is associated with a solid support, wherein the 5' end of each of the more than one oligonucleotide barcodes comprises a capture sequence; a first primer capable of hybridizing with a first universal sequence, optionally also comprising a third universal sequence; and a primer capable of hybridizing with a nucleic acid target or its complement. The amplification primers optionally further comprise a second universal sequence; a DNA ligase; an extension reagent, optionally a reverse transcription reagent, and also optionally reverse transcriptase and dNTPs; one or more immobilizers; one or more permeabilizers; a cross-linking agent; a deimmobilizing agent; a lysis buffer; more than one cell component binding agent, wherein each of the more than one cell component binding agent comprises a cell component binding agent-specific oligonucleotide, said cell component binding agent-specific oligonucleotide comprising a unique identifier sequence of the cell component binding agent, and wherein the cell component binding agent is capable of specifically binding at least one of more than one cell component target; a blocking agent; one or more decoy oligonucleotides; and / or one or more blocking oligonucleotides.
[0020] In some embodiments, more than one probe oligonucleotide comprises a probe oligonucleotide set, said probe oligonucleotide set comprising two or more types of more than one probe oligonucleotide, wherein each of more than one contains a probe sequence configured to hybridize with nucleic acid targets in more than one nucleic acid target, optionally at least about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 12, about 14, about 16, or about 18. Approximately 20, 30, 40, 50, 60, 70, 80, 90, 100, 125, 150, 175, 200, 225, 250, 275, 300, 325, 350, 375, 400, 425, 450, 475, or 500 different nucleic acid targets. In some implementations, the amplification primers comprise a set of amplification primers configured to hybridize with more than one nucleic acid target or its complement, optionally at least about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 12, about 14, about 16, about 18, about 20, about 30, about 40, about 50, about 60, about 70, about 80, about 90, about 100, about 125, about 150, about 175, about 200, about 225, about 250, about 275, about 300, about 325, about 350, about 375, about 400, about 425, about 450, about 475, or about 500 different amplification primers. Brief description of the attached diagram
[0021] Figure 1 The illustration shows a non-limiting exemplary barcode.
[0022] Figure 2 This illustrates a non-limiting exemplary workflow for barcode encoding and digital counting.
[0023] Figure 3 This is a schematic diagram illustrating a non-limiting exemplary process for generating an index library of targets with 3' end barcodes from more than one type of target.
[0024] Figures 4A-4D A non-limiting exemplary schematic workflow for gene expression analysis of fixed cells is described. Detailed Explanation
[0025] The following detailed description references the accompanying drawings, which form a part of this document. In the drawings, like symbols generally identify like components unless the context otherwise indicates. The illustrative embodiments described in the detailed description, drawings, and claims are not intended to be limiting. Other embodiments may be utilized, and other changes may be made, without departing from the spirit or scope of the subject matter presented herein. It will be readily understood that aspects of this disclosure as generally described herein and illustrated in the drawings can be arranged, substituted, combined, separated, and designed in a variety of different configurations, all of which are expressly contemplated herein and form part of this disclosure.
[0026] All patents, published patent applications, other publications, and sequences from GenBank and other databases relating to the relevant technology mentioned herein are incorporated herein by reference in their entirety.
[0027] Quantifying small numbers of nucleic acid molecules (e.g., messenger ribonucleotide (mRNA) molecules) is clinically important for identifying genes expressed in cells, for example, at different developmental stages or under different environmental conditions. However, determining the absolute number of nucleic acid molecules (e.g., mRNA molecules) can also be very challenging, especially when the number of molecules is very small. One method for determining the absolute number of molecules in a sample is digital polymerase chain reaction (PCR). Ideally, PCR produces the same copy of the molecule in each cycle. However, PCR can have drawbacks, causing each molecule to replicate with a random probability that varies depending on the PCR cycle and the gene sequence, leading to amplification bias and inaccurate gene expression measurements. Random barcodes with unique molecular labels (also known as molecular indexes (MI)) can be used to count the number of molecules and correct for amplification bias. Random barcoding, such as Precise... TM Measurement (Cellular Research, Inc. (Palo Alto, CA)) and Rhapsody TM The assay (Becton, Dickinson and Company (Franklin Lakes, NJ)) can correct for biases induced by PCR and library preparation steps by using molecular markers (ML) to label mRNA during reverse transcription (RT).
[0028] Precise TMAssays can utilize a non-depleting pool of random barcodes containing a large number (e.g., 6561 to 65536) unique molecular marker sequences on multiple (T) oligonucleotides to hybridize with all multiple (A)-mRNAs in the sample during the RT step. The random barcodes may contain universal PCR initiation sites. During RT, target gene molecules react randomly with the random barcodes. Each target molecule can hybridize with a random barcode, resulting in the production of randomly barcoded complementary ribonucleotide (cDNA) molecules. After labeling, the randomly barcoded cDNA molecules from the wells of a microplate can be pooled into a single tube for PCR amplification and sequencing. The raw sequencing data can be analyzed to produce the number of reads, the number of random barcodes with unique molecular marker sequences, and the number of mRNA molecules.
[0029] This disclosure includes methods for labeling nucleic acid targets in a sample. In some embodiments, the method includes: contacting a sample containing a copy of the nucleic acid target with more than one probe oligonucleotide, wherein each probe oligonucleotide contains a coupling sequence and a probe sequence configured to hybridize with the nucleic acid target. The method may include: extending more than one probe oligonucleotide hybridized with a copy of the nucleic acid target to produce more than one extended probe oligonucleotide, each of the more than one extended probe oligonucleotide containing a sequence complementary to at least a portion of the nucleic acid target. The method may include: barcoding more than one extended probe oligonucleotide or its product using more than one oligonucleotide barcoding to generate more than one barcoded probe oligonucleotide, wherein each of the more than one oligonucleotide barcodes contains a molecular marker, and wherein each of the more than one barcoded probe oligonucleotides contains a molecular marker, a probe sequence, and a sequence complementary to at least a portion of the nucleic acid target. The method may include: obtaining sequencing data containing more than one sequencing read of the barcoded probe oligonucleotide or its product, wherein each of the more than one sequencing read contains a molecular marker sequence and a subsequence of the nucleic acid target. This method may include determining the copy number of a nucleic acid target in a sample based on the number of molecular markers associated with more than one barcoded probe oligonucleotide or its product.
[0030] This disclosure includes methods for determining the copy number of a nucleic acid target in a sample. In some embodiments, the method includes contacting a sample containing copies of the nucleic acid target with more than one probe oligonucleotide, wherein each probe oligonucleotide contains a coupling sequence and a probe sequence configured to hybridize with the nucleic acid target. The method may include extending the more than one probe oligonucleotide hybridized with copies of the nucleic acid target to produce more than one extended probe oligonucleotide, each of the more than one extended probe oligonucleotide containing a sequence complementary to at least a portion of the nucleic acid target. The method may include barcoding more than one extended probe oligonucleotide or its product using more than one oligonucleotide barcoding to generate more than one barcoded probe oligonucleotide, wherein each oligonucleotide barcode in the more than one oligonucleotide barcode contains a molecular marker, and wherein each of the more than one barcoded probe oligonucleotide contains a molecular marker, a probe sequence, and a sequence complementary to at least a portion of the nucleic acid target. The method may include obtaining sequencing data containing more than one sequencing read of the barcoded probe oligonucleotide or its product, wherein each of the more than one sequencing read contains a molecular marker sequence and a subsequence of the nucleic acid target. This method may include determining the copy number of a nucleic acid target in a sample based on the number of molecular markers associated with more than one barcoded probe oligonucleotide or its product.
[0031] This disclosure includes methods for determining the spatial location and copy number of a nucleic acid target in a sample. In some embodiments, the method includes contacting each of two or more spatial locations of a sample containing a copy of the nucleic acid target with more than one probe oligonucleotide, wherein each probe oligonucleotide comprises a coupling sequence, a probe sequence configured to hybridize with the nucleic acid target, and a predetermined spatial marker. In some embodiments, probe oligonucleotides contacting the same spatial location comprise the same spatial marker sequence, and probe oligonucleotides contacting different spatial locations of the sample comprise different spatial marker sequences. The method may include extending more than one probe oligonucleotide hybridizing with a copy of the nucleic acid target to produce more than one extended probe oligonucleotide, each of the more than one extended probe oligonucleotides comprising a sequence complementary to at least a portion of the nucleic acid target. This method may include: barcoding more than one extended probe oligonucleotide or its product using more than one oligonucleotide barcode to generate more than one barcoded probe oligonucleotide, wherein each oligonucleotide barcode in the more than one oligonucleotide barcode contains a molecular marker, and wherein each of the more than one barcoded probe oligonucleotide contains a molecular marker, a probe sequence, and a sequence complementary to at least a portion of a nucleic acid target. This method may include: obtaining sequencing data containing more than one sequencing read containing the barcoded probe oligonucleotide or its product, wherein each of the more than one sequencing read contains a spatial marker sequence, a molecular marker sequence, and a subsequence of the nucleic acid target. This method may include: for each unique spatial marker sequence associated with a different spatial location in the sample, counting the number of molecular markers having different sequences associated with the nucleic acid target to determine the copy number of the nucleic acid target at each spatial location in the sample.
[0032] Definitions Unless otherwise defined, the technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains. See, for example, Singleton et al., Dictionary of Microbiology and Molecular Biology, 2nd edition, J. Wiley & Sons (New York, NY 1994); Sambrook et al., Molecular Cloning, A Laboratory Manual, Cold Spring Harbor Press (Cold Spring Harbor, NY 1989). For the purposes of this disclosure, the following terms are defined below.
[0033] As used herein, the term "adaptor" may refer to a sequence that facilitates the amplification or sequencing of an associated nucleic acid. The associated nucleic acid may include a target nucleic acid. The associated nucleic acid may include one or more of spatial markers, target markers, sample markers, index markers, or barcode sequences (e.g., molecular markers). The adaptor may be linear. The adaptor may be a pre-adenylated adaptor. The adaptor may be double-stranded or single-stranded. One or more adaptors may be located at the 5' or 3' end of the nucleic acid. When the adaptor contains known sequences at both the 5' and 3' ends, the known sequences may be the same or different sequences. Adaptors located at the 5' and / or 3' ends of a polynucleotide may be able to hybridize with one or more oligonucleotides immobilized on a surface. In some embodiments, the adaptor may contain a universal sequence. The universal sequence may be a region of a nucleotide sequence common to two or more nucleic acid molecules. The two or more nucleic acid molecules may also have regions with different sequences. Thus, for example, the 5' adaptor may contain the same and / or universal nucleic acid sequence, and the 3' adaptor may contain the same and / or universal sequence. A universal sequence that can be present in different members of more than one nucleic acid molecule allows for the replication or amplification of more than one different sequence using a single universal primer complementary to the universal sequence. Similarly, at least one, two (e.g., a pair) or more universal sequences that can be present in different members of a set of nucleic acid molecules allow for the replication or amplification of more than one different sequence using at least one, two (e.g., a pair) or more single universal primers complementary to the universal sequence. Therefore, universal primers contain sequences that can hybridize with such universal sequences. Molecules having target nucleic acid sequences can be modified to attach universal adaptors (e.g., non-target nucleic acid sequences) to one or both ends of different target nucleic acid sequences. One or more universal primers attached to the target nucleic acid can provide sites for universal primer hybridization. One or more universal primers attached to the target nucleic acid can be the same as or different from each other.
[0034] As used herein, the term "association" or "associated with" can mean that two or more substances can be identified as co-located at a point in time. Association can mean that two or more substances are or were in similar containers. Association can be an informatics association. For example, digital information about two or more substances can be stored and used to determine that one or more substances are co-located at a point in time. Association can also be a physical association. In some embodiments, two or more associated substances are "tethered," "attached," or "fixed" to each other or to a common solid or semi-solid surface. Association can refer to a covalent or non-covalent manner used to attach a marker to a solid or semi-solid support, such as a bead. Association can be a covalent bond between a target and a marker. Association can include hybridization between two molecules, such as a target molecule and a marker.
[0035] As used herein, the term "complementary" can refer to the ability of two nucleotides to pair precisely. For example, if a nucleotide at a given position in a nucleic acid can hydrogen-bond with a nucleotide of another nucleic acid, the two nucleic acids are considered complementary to each other at that position. Complementarity between two single-stranded nucleic acid molecules can be "partial," where only some nucleotides bind, or it can be complete when there is full complementarity between the single-stranded molecules. If a first nucleotide sequence is complementary to a second nucleotide sequence, the first nucleotide sequence can be referred to as the "complement" of the second sequence. If a first nucleotide sequence is complementary to a sequence opposite to the second sequence (i.e., the nucleotide sequence is reversed), the first nucleotide sequence can be referred to as the "reverse complement" of the second sequence. As used herein, a "complementary" sequence can refer to either the "complement" or the "reverse complement" of a sequence. From this disclosure, it is understood that if a molecule can hybridize with another molecule, it can be complementary or partially complementary to the molecule it hybridizes with.
[0036] As used herein, the term "numerical counting" can refer to a method used to estimate the number of target molecules in a sample. Numerical counting may include steps to determine the number of unique markers already associated with a target in the sample. This method (which can be random in nature) transforms the problem of counting molecules from one of the localization and identification of the same molecules to a series of yes / no numerical questions about detecting a predefined set of markers.
[0037] As used herein, the terms "a label" or "more than one label" can refer to a nucleic acid code associated with a target in a sample. A label can be, for example, a nucleic acid label. A label can be a fully or partially amplifiable label. A label can be a fully or partially sequenceable label. A label can be a portion of a naturally occurring nucleic acid that can be identified as distinct. A label can be a known sequence. A label can include a linker of nucleic acid sequences, such as a linker between natural and non-natural sequences. As used herein, the term "label" can be used interchangeably with the terms "index," "tag," or "label-tag." A label can convey information. For example, in various embodiments, a label can be used to determine the identity of a sample, the origin of the sample, the identity of the cells, and / or the target.
[0038] As used herein, the term "non-depleting reservoir" can refer to a pool of barcodes (e.g., random barcodes) composed of many different labels. A non-depleting reservoir can include a large number of different barcodes, such that when the non-depleting reservoir is associated with a target pool, each target may be associated with a unique barcode. The uniqueness of the target molecules for each label can be determined by statistically selected randomness and depends on the copy number of identical target molecules in the set relative to the diversity of the labels. The size of the resulting set of labeled target molecules can be determined by the randomness of the barcode processing, and then analysis of the number of detected barcodes allows for the calculation of the number of target molecules present in the original set or sample. When the ratio of the copy number of present target molecules to the number of unique barcodes is low, the labeled target molecules are highly unique (i.e., the probability that more than one target molecule is labeled by a given label is very low).
[0039] As used herein, the term "nucleic acid" refers to a polynucleotide sequence or a fragment thereof. Nucleic acids may include nucleotides. Nucleic acids may be exogenous or endogenous to cells. Nucleic acids may exist in cell-free environments. Nucleic acids may be genes or fragments thereof. Nucleic acids may be DNA. Nucleic acids may be RNA. Nucleic acids may include one or more analogs (e.g., modified backbone, sugar, or nucleobases). Some non-limiting examples of analogs include: 5-bromouracil, peptide nucleic acids, xenonucleic acids, morpholinos, locked nucleic acids, diol nucleic acids, threonine, dideoxynucleotides, cordycepin, 7-denitro-GTP, fluorophores (e.g., rhodamine or sugar-linked fluorescein), thiol-containing nucleotides, biotin-linked nucleotides, fluorescent base analogs, CpG islands, methyl-7-guanosine, methylated nucleotides, inosine, thiouridine, pseudouridine, dihydrouridine, queuosine, and wyosine. "Nucleic acid", "polynucleotide", "target polynucleotide" and "target nucleic acid" can be used interchangeably.
[0040] Nucleic acids can include one or more modifications (e.g., base modifications, backbone modifications) to provide new or enhanced characteristics (e.g., improved stability). Nucleic acids can contain nucleic acid affinity tags. Nucleosides can be base-sugar combinations. The base moiety of a nucleoside can be a heterocyclic base. The two most common classes of such heterocyclic bases are purines and pyrimidines. Nucleotides can also include a phosphate group covalently linked to the sugar moiety of the nucleoside. For those nucleosides that include furanopentoses, the phosphate group can be linked to the 2', 3', or 5' hydroxyl moiety of the sugar. In the formation of nucleic acids, phosphate groups can covalently link adjacent nucleosides to each other to form a linear polymeric compound. Subsequently, the ends of this linear polymeric compound can be further linked to form a cyclic compound; however, linear compounds are generally preferred. Furthermore, linear compounds can have internal nucleotide base complementarity and can therefore fold in a manner that produces fully or partially double-stranded compounds. In nucleic acids, the phosphate group can generally be referred to as the inter-nucleoside backbone that forms the nucleic acid. The linkage or backbone can be a 3' to 5' phosphodiester linkage.
[0041] Nucleic acids may include modified backbones and / or modified nucleoside linkages. Modified backbones may include backbones that retain phosphorus atoms and backbones that do not contain phosphorus atoms. Suitable nucleic acid backbones containing phosphorus atoms may include, for example, thiophosphates, chiral thiophosphates, dithiophosphates, phosphate triesters, aminoalkyl phosphate triesters, methyl and other alkylphosphonates such as 3'-alkylene phosphonates, 5'-alkylene phosphonates, chiral phosphonates, phosphonates, phosphoramidates (including 3'-aminophosphatases and aminoalkylphosphatases, phosphorodiamidates, thionophosphatases), thioalkylphosphonates, thioalkyl phosphate triesters, selenophosphates and borophosphates, analogs having normal 3'-5' linkages, 2'-5' linkages, and analogs having reverse polarity (where one or more nucleotide linkages are 3' to 3', 5' to 5', or 2' to 2' linkages).
[0042] Nucleic acids can include polynucleotide backbones formed by short-chain alkyl or cycloalkyl nucleosides, mixed heteroatoms, and alkyl or cycloalkyl nucleosides, or one or more short-chain heteroatoms or heterocyclic nucleosides. These can include those with morpholino bonds (partially formed from the sugar moiety of the nucleoside); siloxane backbones; sulfide, sulfoxide, and sulfone backbones; formacetyl and thioformacetyl backbones; methyleneformacetyl and thioformacetyl backbones; riboacetyl backbones; olefin-containing backbones; aminosulfonate backbones; methyleneimino and methylenehydrazine backbones; sulfonate and sulfonamide backbones; amide backbones; and others with mixed N, O, S, and CH2 component moieties.
[0043] Nucleic acids can include nucleic acid mimics. The term "mimic" can be intended to include polynucleotides in which only the furanose ring or both the furanose ring and the internucleotide linkage are replaced by non-furanose groups; substitution of only the furanose ring can also be called a sugar substitute (surrogate). The heterocyclic base moiety or modified heterocyclic base moiety can be maintained to hybridize with a suitable target nucleic acid. One such nucleic acid can be a peptide nucleic acid (PNA). In a PNA, the sugar backbone of the polynucleotide can be replaced by an amide-containing backbone, particularly an aminoethylglycine backbone. The nucleotide can be retained and directly or indirectly bound to the aza-nitrogen atom of the amide moiety of the backbone. The backbone in a PNA compound can contain two or more linked aminoethylglycine units, giving the PNA an amide-containing backbone. The heterocyclic base moiety can directly or indirectly bind to the aza-nitrogen atom of the amide moiety of the backbone.
[0044] Nucleic acids may include a morpholine backbone structure. For example, a nucleic acid may contain a 6-membered morpholine ring instead of a ribose ring. In some of these embodiments, a phosphate diamide ester or other non-phosphodiester nucleoside linker may replace the phosphate diester linker.
[0045] Nucleic acids can include morpholino units (e.g., morpholinonucleotides) with heterocyclic bases attached to a morpholino ring. Linking groups can connect the morpholino monomer units within the morpholinonucleotide. Nonionic morpholino-based oligomers can exhibit fewer undesirable interactions with cellular proteins. Morpholino-based polynucleotides can be nonionic mimics of nucleic acids. Various compounds within the morpholino category can be linked using different linking groups. Another category of polynucleotide mimics can be called cyclohexenyl nucleic acids (CeNA). The furanose ring typically present in nucleic acid molecules can be replaced by a cyclohexenyl ring. CeNA DMT-protected phosphoramidite monomers can be prepared using phosphoramidite chemistry and used in oligomer synthesis. Incorporating CeNA monomers into nucleic acid chains can increase the stability of DNA / RNA hybrids. CeNA oligoadenylates can form complexes with nucleic acid complements exhibiting similar stability to natural complexes. Other modifications can include locked nucleic acids (LNAs), in which the 2'-hydroxyl group is linked to the 4' carbon atom of the sugar ring, forming a 2'-C, 4'-C-oxomethylene bond, thus forming a bicyclic sugar moiety. The bond can be a methylene (-CH2-). n A group bridging the 2' oxygen atom and the 4' carbon atom, wherein n It is 1 or 2. LNA and LNA analogs can exhibit very high double-stranded thermal stability (Tm = +3°C to +10°C) with complementary nucleic acids, stability against 3'-exonuclease degradation, and good solubility.
[0046] Nucleic acids may also include nucleobase (usually referred to simply as "bases") modifications or substitutions. As used herein, "unmodified" or "natural" nucleobases may include purine bases (e.g., adenine (A) and guanine (G)) and pyrimidine bases (e.g., thymine (T), cytosine (C) and uracil (U)). Modified nucleobases may include other synthetic and natural nucleobases, such as 5-methylcytosine (5-me-C), 5-hydroxymethylcytosine, xanthine, hypoxanthine, 2-aminoadenine, adenine, and guanine 6-methyl derivatives and other alkyl derivatives, adenine and guanine 2-propyl derivatives and other alkyl derivatives, 2-thiouracil, 2-thiothymine and 2-thiocytosine, 5-halouracil and cytosine, 5-propynyl (-C≡C-CH3)uracil and cytosine, and other pyrimidine bases. Alkyne derivatives, 6-azouracil, cytosine and thymine, 5-uracil (pseudouracil), 4-thiouracil, 8-halogen, 8-amino, 8-thio, 8-thioalkyl, 8-hydroxy and other 8-substituted adenine and guanine, 5-halogen, especially 5-bromo, 5-trifluoromethyl and other 5-substituted uracil and cytosine, 7-methylguanine and 7-methyladenine, 2-F-adenine, 2-aminoadenine, 8-azaguanine and 8-azaadenine, 7-deadenine and 7-deadenine and 3-deadenine and 3-deadenine. Modified nucleobases can include tricyclic pyrimidines, such as phenoxazincytidine (1H-pyrimido(5,4-b)(1,4)benzoxazin-2(3H)-one), phenthiazincytidine (1H-pyrimido(5,4-b)(1,4)benzoxazin-2(3H)-one), and G-clamps such as substituted phenoxazincytidines (e.g., 9-(2-aminoethoxy)-H-pyrimido(5,4-(b)(1,4)benzoxazin-2(3H)-one), phenoxazincytidine (1 ...b)(1,4)benzoxazin-2(3H)-one), phenoxazincytidine (1H-pyrimido(5,4-b)(1,4)benzoxazin-2(3H)-one), phenoxazincytidine (1H-pyrimido(5,4-b)(1,4)benzoxazin-2(3H)-one), phenoxazincytidine (1H-pyrimido(5,4-b)(1,4)benzoxazin-2(3H)-one), phenoxazincytidine (1H-pyrimido(5,4 Thiazicytidine (1H-pyrimido(5,4-b)(1,4)benzothiazine-2(3H)-one), G-clamp such as substituted phenoxazincytidine (e.g., 9-(2-aminoethoxy)-H-pyrimido(5,4-(b)(1,4)benzothiazine-2(3H)-one), carbazolecytidine (2H-pyrimido(4,5-b)indole-2-one), pyridoindolecytidine (H-pyrido(3',2':4,5)pyrrolo[2,3-d]pyrimido-2-one).
[0047] As used herein, the term "sample" can refer to a composition containing a target. Suitable samples for analysis using the disclosed methods, apparatus, and systems include cells, tissues, organs, or organisms.
[0048] As used herein, the term "sampling device" or "device" can refer to a device that can take a sample slice and / or place said slice on a substrate. A sampling device can refer to, for example, a fluorescence activated cell sorting (FACS) machine, a cell sorter, a biopsy needle, a biopsy device, a tissue sectioning device, a microfluidic device, a blade grid, and / or an ultramicrotome.
[0049] As used herein, the term "solid support" can refer to a discrete solid or semi-solid surface on which more than one barcode (e.g., a random barcode) can be attached. Solid supports can include any type of solid, porous, or hollow spheres, balls, bearings, cylinders, or other similar configurations comprising plastic, ceramic, metallic, or polymeric materials (e.g., hydrogels) on which nucleic acids (e.g., covalently or non-covalently) can be immobilized. Solid supports can include discrete particles that can be spherical (e.g., microspheres) or have non-spherical or irregular shapes, such as cubic, rectangular, conical, cylindrical, elliptical, or disk-shaped. Beads can be non-spherical. More than one solid support spaced apart in an array may not include a substrate. The term "solid support" is used interchangeably with the term "bead."
[0050] As used herein, the term "random barcode" can refer to a labeled polynucleotide sequence of this disclosure. A random barcode can be a polynucleotide sequence that can be randomly barcoded. Random barcodes can be used for target quantification in a sample. Random barcodes can be used to control for errors that may occur after a label is associated with a target. For example, random barcodes can be used to assess amplification or sequencing errors. A random barcode associated with a target can be referred to as random barcode-target or random barcode-tag-target.
[0051] As used herein, the term "gene-specific random barcode" can refer to a polynucleotide sequence containing a marker and a gene-specific target-binding region. A random barcode can be a polynucleotide sequence that can be randomly barcoded. Random barcodes can be used to quantify a target in a sample. Random barcodes can be used to control for errors that may occur after a marker is associated with a target. For example, random barcodes can be used to assess amplification or sequencing errors. A random barcode associated with a target can be referred to as a random barcode-target or a random barcode-tag-target.
[0052] As used herein, the term "random barcoding" can refer to the random labeling (e.g., barcoding) of nucleic acids. Random barcoding can utilize a recursive Poisson strategy to associate and quantify the label associated with the target. As used herein, the term "random barcoding" can be used interchangeably with "random labeling".
[0053] As used herein, the term "target" can refer to a composition that can be associated with a barcode (e.g., a random barcode). Exemplary suitable targets for analysis using the disclosed methods, apparatus, and systems include oligonucleotides, DNA, RNA, mRNA, microRNA, tRNA, etc. Targets can be single-stranded or double-stranded. In some embodiments, the target can be a protein, peptide, or polypeptide. In some embodiments, the target is a lipid. As used herein, "target" may be used interchangeably with "species".
[0054] As used herein, the term "reverse transcriptase" can refer to a group of enzymes that possess reverse transcriptase activity (i.e., catalyze the synthesis of DNA from an RNA template). Generally, such enzymes include, but are not limited to, retroviral reverse transcriptases, retrotransposon reverse transcriptases, retroplasmid reverse transcriptases, retrotranscripton reverse transcriptases, bacterial reverse transcriptases, group II intron-derived reverse transcriptases, and their mutants, variants, or derivatives. Non-retroviral reverse transcriptases include non-LTR retrotransposon reverse transcriptases, retroplasmid reverse transcriptases, retrotranscripton reverse transcriptases, and group II intron reverse transcriptases. Examples of group II intron reverse transcriptases include Lactococcus lactis (… Lactococcus lactis ) LI.LtrB intron reverse transcriptase, slender thermophilic Synechococcus ( Thermosynechococcus elongatus TeI4c intron reverse transcriptase or thermophilic Bacillus stearothermophilus ( Geobacillus stearothermophilus GsI-IIC intron reverse transcriptases. Other classes of reverse transcriptases can include many types of non-retroviral reverse transcriptases (i.e., especially retrotranscriptants, group II introns, and diversity-generating reverse transcription elements).
[0055] The terms “universal adaptor primer,” “universal primer adaptor,” or “universal adaptor sequence” are used interchangeably to refer to a nucleotide sequence that can be used to hybridize with a barcode (e.g., a random barcode) to produce a gene-specific barcode. A universal adaptor sequence can be, for example, a known sequence universally applicable to all barcodes and used in the methods of this disclosure. For example, when labeling more than one target using the methods disclosed herein, each target-specific sequence can be ligated to the same universal adaptor sequence. In some embodiments, more than one universal adaptor sequence can be used in the methods disclosed herein. For example, when labeling more than one target using the methods disclosed herein, at least two target-specific sequences are ligated to different universal adaptor sequences. The universal adaptor primer and its complement can be included in two oligonucleotides, one of which contains the target-specific sequence and the other contains the barcode. For example, the universal adaptor sequence can be part of an oligonucleotide containing the target-specific sequence to produce a nucleotide sequence complementary to the target nucleic acid. A second oligonucleotide containing the barcode and the complement of the universal adaptor sequence can hybridize with the nucleotide sequence to produce a target-specific barcode (e.g., a target-specific random barcode). In some implementations, the universal adaptor primers have sequences different from those of the universal PCR primers used in the methods of this disclosure.
[0056] Barcodes Barcoding, such as random barcoding, has been described in, for example, by Fu et al. Proc Natl Acad Sci U.S.A. , May 31, 2011, 108(22):9026-31; U.S. Patent Application Publication No. US2011 / 0160078; Fan et al., ScienceFebruary 6, 2015, 347(6222):1258367; U.S. Patent Application Publication No. US2015 / 0299784 and PCT Application Publication No. WO2015 / 031691; the contents of each of these, including any supporting or supplementary information or material, are incorporated herein by reference in their entirety. In some embodiments, the barcode disclosed herein may be a random barcode, which may be a polynucleotide sequence that can be used to randomly label (e.g., barcode, tag) a target. A barcode can be called a random barcode if the ratio of the number of different barcode sequences in a random barcode to the number of occurrences of any target to be labeled can be, or approximately, one of the following: 1:1, 2:1, 3:1, 4:1, 5:1, 6:1, 7:1, 8:1, 9:1, 10:1, 11:1, 12:1, 13:1, 14:1, 15:1, 16:1, 17:1, 18:1, 19:1, 20:1, 30:1, 40:1, 50:1, 60:1, 70:1, 80:1, 90:1, 100:1, or any number or range between any two of these values. A target can be mRNA material comprising mRNA molecules having the same or nearly identical sequences. A barcode can be called a random barcode if the ratio of the number of different barcode sequences in a random barcode to the number of occurrences of any target to be labeled is at least or at most the following: 1:1, 2:1, 3:1, 4:1, 5:1, 6:1, 7:1, 8:1, 9:1, 10:1, 11:1, 12:1, 13:1, 14:1, 15:1, 16:1, 17:1, 18:1, 19:1, 20:1, 30:1, 40:1, 50:1, 60:1, 70:1, 80:1, 90:1, or 100:1. The barcode sequences of a random barcode can be called molecular markers.
[0057] Barcodes (e.g., random barcodes) may include one or more types of markers. Exemplary markers may include universal markers, cell markers, barcode sequences (e.g., molecular markers), sample markers, plate markers, spatial markers, and / or pre-spatial labels. Figure 1 An exemplary barcode 104 with spatial markers is illustrated. Barcode 104 may contain a 5' amine that can link the barcode to a solid support 105. The barcode may contain universal markers, dimensional markers, spatial markers, cellular markers, and / or molecular markers. The order of the different markers (including but not limited to universal markers, dimensional markers, spatial markers, cellular markers, and molecular markers) in the barcode may vary. For example, as... Figure 1As shown, the universal label can be the 5'-most label, and the molecular label can be the 3'-most label. Spatial, dimensional, and cellular labels can be in any order. In some embodiments, the universal, spatial, dimensional, cellular, and molecular labels are in any order. The barcode can contain a target-binding region. The target-binding region can interact with a target in the sample (e.g., target nucleic acid, RNA, mRNA, DNA). For example, the target-binding region can contain an oligo(dT) sequence that can interact with the multiple (A) tail of mRNA. In some cases, the barcode labels (e.g., universal label, dimensional label, spatial label, cellular label, and barcode sequence) can be separated by 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 or more nucleotides.
[0058] A marker (e.g., a cell marker) can comprise a unique set of nucleic acid subsequences of defined length, for example, seven nucleotides each (equivalent to the number of bits used in some Hamming error-correcting codes), which can be designed to provide error-correcting capabilities. A set of error-correcting subsequences comprising seven nucleotide sequences can be designed such that any pairwise combination of sequences in the set exhibits a defined “genetic distance” (or number of mismatched bases); for example, a set of error-correcting subsequences can be designed to exhibit a genetic distance of three nucleotides. In this case, review of the error-correcting sequences in the sequence data set of the marked target nucleic acid molecule (described in more detail below) allows for the detection or correction of amplification or sequencing errors. In some embodiments, the length of the nucleic acid subsequences used to generate the error-correcting code can vary; for example, their length can be the following or can be approximately the following: 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 30, 31, 40, 50, or any number or range of nucleotides between any two of these values. In some implementations, nucleic acid subsequences of other lengths can be used to generate error correction codes.
[0059] The barcode may contain a target-binding region. This region can interact with a target in the sample. The target may be, or include, ribonucleic acid (RNA), messenger RNA (mRNA), microRNA, small interfering RNA (siRNA), RNA degradation products, RNAs each containing multiple (A) tails, or any combination thereof. In some embodiments, more than one target may include deoxyribonucleic acid (DNA).
[0060] In some embodiments, the target binding region may include an oligo(dT) sequence that can interact with the multiple (A) tails of mRNA. One or more markers of the barcode (e.g., universal markers, dimensional markers, spatial markers, cellular markers, and barcode sequences (e.g., molecular markers)) may be separated from another or two remaining markers of the barcode by spacers. Spacers may be, for example, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 or more nucleotides. In some embodiments, none of the markers in the barcode are separated by spacers.
[0061] Universal Markers Barcodes may contain one or more universal markers. In some embodiments, one or more universal markers may be the same for all barcodes in a group of barcodes attached to a given solid support. In some embodiments, one or more universal markers may be the same for all barcodes attached to more than one bead. In some embodiments, the universal marker may include a nucleic acid sequence capable of hybridizing with sequencing primers. Sequencing primers may be used to sequence barcodes including universal markers. Sequencing primers (e.g., universal sequencing primers) may include sequencing primers associated with a high-throughput sequencing platform. In some embodiments, the universal marker may include a nucleic acid sequence capable of hybridizing with PCR primers. In some embodiments, the universal marker may include a nucleic acid sequence capable of hybridizing with both sequencing primers and PCR primers. The nucleic acid sequence of the universal marker capable of hybridizing with sequencing primers or PCR primers may be referred to as a primer binding site. The universal marker may include a sequence that can be used to initiate barcode transcription. The universal marker may include a sequence that can be used to extend the barcode or a region within the barcode. The length of the universal marker can be, or can be about, one, two, three, four, five, ten, fifteen, twenty, twenty, twenty, five, thirty, thirty, thirty, forty, forty, forty, fifty, or any number or range of nucleotides between any two of these values. For example, the universal marker may include at least about ten nucleotides. The length of the universal marker can be at least, or can be at most, one, two, three, four, five, ten, fifteen, twenty, twenty, fifty, thirty, thirty, forty ...
[0062] Dimensional Markers Barcodes can contain one or more dimensional markers. In some implementations, dimensional markers may include nucleic acid sequences that provide information about the dimension in which the marking (e.g., random marking) occurred. For example, a dimensional marker can provide information about the timing of target barcoding. Dimensional markers can be associated with the timing of barcoding (e.g., random barcoding) in the sample. Dimensional markers can be activated at the time of marking. Different dimensional markers can be activated at different times. Dimensional markers provide information about the order in which targets, groups of targets, and / or samples are barcoded. For example, a population of cells can be barcoded during G0 phase of the cell cycle. During G1 phase of the cell cycle, cells can be pulsed again with barcodes (e.g., random barcodes). During S phase of the cell cycle, cells can be pulsed again with barcodes, and so on. The barcode at each pulse (e.g., each phase of the cell cycle) can contain different dimensional markers. In this way, dimensional markers provide information about which targets were marked at which phase of the cell cycle. Dimensional markers can probe many different biological times. Exemplary biological timeframes may include, but are not limited to, the cell cycle, transcription (e.g., transcription initiation), and transcript degradation. In another instance, samples (e.g., cells, cell populations) may be labeled before and / or after treatment with drugs and / or therapies. Changes in copy number of different targets may indicate a sample’s response to drugs and / or therapies.
[0063] Dimension markers can be activated. Activatable dimension markers can be activated at specific time points. Activatable markers can be, for example, constitutively activated (e.g., not turned off). Activatable dimension markers can be, for example, reversibly activated (e.g., activated dimension markers can be turned on and off). Dimension markers can be, for example, reversibly activated at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more times. Dimension markers can be reversibly activated, for example, at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more times. In some embodiments, dimension markers can be activated by fluorescence, light, chemical events (e.g., cleavage, linking to another molecule, addition of modifications (e.g., PEGylation, sumoylate, acetylation, methylation, deacetylation, demethylation), photochemical events (e.g., photocaging), and the introduction of non-natural nucleotides.
[0064] In some embodiments, the dimension markers may be the same for all barcodes (e.g., random barcodes) attached to a given solid support (e.g., beads), but different for different solid supports (e.g., beads). In some embodiments, at least 60%, 70%, 80%, 85%, 90%, 95%, 97%, 99%, or 100% of the barcodes on the same solid support may contain the same dimension markers. In some embodiments, at least 60% of the barcodes on the same solid support may contain the same dimension markers. In some embodiments, at least 95% of the barcodes on the same solid support may contain the same dimension markers.
[0065] Up to 10 can be presented in more than one solid support (e.g., beads). 6 One or more unique dimension marker sequences. The length of the dimension marker can be the following or can be about the following: 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, or any two of these values of nucleotides. The length of the dimension marker can be at least the following or can be at most the following: 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 100, 200, or 300 nucleotides. The dimension marker can contain between about 5 and about 200 nucleotides. The dimension marker can contain between about 10 and about 150 nucleotides. The dimension marker can contain nucleotides of length between about 20 and about 125.
[0066] Spatial Markers A barcode may contain one or more spatial markers. In some embodiments, the spatial marker may contain a nucleic acid sequence that provides information about the spatial orientation of a target molecule associated with the barcode. The spatial marker may be associated with coordinates in a sample. The coordinates may be fixed coordinates. For example, the coordinates may be fixed with reference to a substrate. The spatial marker may be referenced to a two-dimensional or three-dimensional grid. The coordinates may be fixed with reference to a landmark. The landmark may be identifiable in space. The landmark may be an imageable structure. The landmark may be a biological structure, such as an anatomical landmark. The landmark may be a cellular landmark, such as an organelle. The landmark may be a non-natural landmark, such as a structure with an identifiable identifier (such as a color code, barcode, magnetic property, fluorescence, radioactivity, or unique size or shape). The spatial marker may be associated with physical partitions (e.g., pores, containers, or droplets). In some embodiments, more than one spatial marker may be used together to encode one or more locations in space.
[0067] The spatial marker can be the same for all barcodes attached to a given solid support (e.g., beads), but different for different solid supports (e.g., beads). In some embodiments, the percentage of barcodes containing the same spatial marker on the same solid support can be or can be about the following: 60%, 70%, 80%, 85%, 90%, 95%, 97%, 99%, 100%, or a number or range between any two of these values. In some embodiments, the percentage of barcodes containing the same spatial marker on the same solid support can be at least or at most the following: 60%, 70%, 80%, 85%, 90%, 95%, 97%, 99%, or 100%. In some embodiments, at least 60% of the barcodes on the same solid support can contain the same spatial marker. In some embodiments, at least 95% of the barcodes on the same solid support can contain the same spatial marker.
[0068] Up to 10 can be presented in more than one solid support (e.g., beads). 6 One or more unique spatial marker sequences. The length of the spatial marker can be the following or can be about the following: 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, or any two of these values of nucleotides. The length of the spatial marker can be at least the following or at most the following: 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 100, 200, or 300 nucleotides. The spatial marker can contain between about 5 and about 200 nucleotides. The spatial marker can contain between about 10 and about 150 nucleotides. The spatial marker can contain nucleotides with a length between about 20 and about 125 nucleotides.
[0069] Cellular Markers Barcodes (e.g., random barcodes) may contain one or more cell markers. In some embodiments, the cell markers may contain nucleic acid sequences that provide information for determining which target nucleic acid originates from which cell. In some embodiments, the cell markers are the same for all barcodes attached to a given solid support (e.g., beads), but different for different solid supports (e.g., beads). In some embodiments, the percentage of barcodes containing the same cell marker on the same solid support may be or about the following: 60%, 70%, 80%, 85%, 90%, 95%, 97%, 99%, 100%, or a number or range between any two of these values. In some embodiments, the percentage of barcodes containing the same cell marker on the same solid support may be or about the following: 60%, 70%, 80%, 85%, 90%, 95%, 97%, 99%, or 100%. For example, at least 60% of the barcodes on the same solid support may contain the same cell marker. As another example, at least 95% of the barcodes on the same solid support may contain the same cell marker.
[0070] Up to 10 can be presented in more than one solid support (e.g., beads). 6 One or more unique cell marker sequences. The length of the cell marker can be the following or can be about the following: 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, or any two of these values of nucleotides. The length of the cell marker can be at least the following or can be at most the following: 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 100, 200, or 300 nucleotides. For example, a cell marker can contain between about 5 and about 200 nucleotides. As another example, a cell marker can contain between about 10 and about 150 nucleotides. As yet another example, a cell marker can contain nucleotides with a length between about 20 and about 125.
[0071] Barcode Sequences A barcode may contain one or more barcode sequences. In some embodiments, the barcode sequence may contain a nucleic acid sequence that provides identification information for a specific type of target nucleic acid substance that hybridizes with the barcode. The barcode sequence may contain a nucleic acid sequence that provides a counter (e.g., provides a rough estimate) for a specific occurrence of target nucleic acid substance that hybridizes with the barcode (e.g., a target binding region).
[0072] In some embodiments, a set of diverse barcode sequences is attached to a given solid support (e.g., beads). In some embodiments, there may be the following, or approximately the following, unique molecular marker sequences: 10 2 10 species 3 10 species 4 10 species 5 10 species 6 10 species 7 10 species 8 10 species 9 A number or range between any two of these values. For example, more than one barcode may include approximately 6,561 barcode sequences with different sequences. As another example, more than one barcode may include approximately 65,536 barcode sequences with different sequences. In some implementations, there may be at least the following, or at most the following, unique barcode sequences: 10 2 10 species 3 10 species 4 10 species 5 10 species 6 10 species 7 10 species 8 species or 10 9 A unique molecular marker sequence can be attached to a given solid support (e.g., beads). In some embodiments, the unique molecular marker sequence is partially or wholly contained in the particles (e.g., hydrogel beads).
[0073] In different implementations, the length of the barcode can be different. For example, the length of the barcode can be the following or can be about the following: 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, or any two of these values of nucleotides or a range thereof. As another example, the length of the barcode can be at least the following or can be at most the following: 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 100, 200, or 300 nucleotides.
[0074] Molecular Markers A barcode (e.g., a random barcode) may contain one or more molecular markers. The molecular marker may contain a barcode sequence. In some embodiments, the molecular marker may contain a nucleic acid sequence that provides identification information for a specific type of target nucleic acid substance hybridizing with the barcode. The molecular marker may contain a nucleic acid sequence that provides a counter for the specific occurrence of target nucleic acid substance hybridizing with the barcode (e.g., a target binding region).
[0075] In some embodiments, a set of dissimilar molecular markers are attached to a given solid support (e.g., beads). In some embodiments, there may be the following, or approximately the following, unique molecular marker sequences: 10 2 10 species 3 10 species 4 10 species 5 10 species 6 10 species 7 10 species 8 10 species 9 A number or range between any two of these values. For example, more than one barcode may include approximately 6,561 molecular markers with different sequences. As another example, more than one barcode may include approximately 65,536 molecular markers with different sequences. In some embodiments, there may be at least the following, or at most the following, unique molecular marker sequences: 10 2 10 species 3 10 species 4 10 species 5 10 species 6 10 species 7 10 species 8 species or 10 9 A type of barcode with a unique molecular marker sequence can be attached to a given solid support (e.g., beads).
[0076] For barcoding using more than one random barcode (e.g., random barcoding), the ratio of the number of different molecular marker sequences to the frequency of any target can be, or approximately, one of the following: 1:1, 2:1, 3:1, 4:1, 5:1, 6:1, 7:1, 8:1, 9:1, 10:1, 11:1, 12:1, 13:1, 14:1, 15:1, 16:1, 17:1, 18:1, 19:1, 20:1, 30:1, 40:1, 50:1, 60:1, 70:1, 80:1, 90:1, 100:1, or any number or range between two of these values. A target can be an mRNA substance comprising mRNA molecules having the same or nearly identical sequences. In some implementations, the ratio of the number of different molecular marker sequences to the number of occurrences of any target is at least or at most the following: 1:1, 2:1, 3:1, 4:1, 5:1, 6:1, 7:1, 8:1, 9:1, 10:1, 11:1, 12:1, 13:1, 14:1, 15:1, 16:1, 17:1, 18:1, 19:1, 20:1, 30:1, 40:1, 50:1, 60:1, 70:1, 80:1, 90:1, or 100:1.
[0077] The length of the molecular marker can be the following or approximately the following: 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, or any number or range of nucleotides between any two of these values. The length of the molecular marker can be at least the following or at most the following: 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 100, 200, or 300 nucleotides.
[0078] Target Binding Regions The barcode may contain one or more target-binding regions, such as capture probes. In some embodiments, the target-binding region may hybridize with a target of interest. In some embodiments, the target-binding region may contain a nucleic acid sequence that specifically hybridizes with a target (e.g., a target nucleic acid, a target molecule, such as the cellular nucleic acid to be analyzed) (e.g., specifically hybridizes with a specific gene sequence). In some embodiments, the target-binding region may contain a nucleic acid sequence that can attach (e.g., hybridize) to a specific location on a specific target nucleic acid. In some embodiments, the target-binding region may contain a nucleic acid sequence capable of specifically hybridizing with a restriction enzyme site overhang (e.g., an EcoRI sticky end overhang). The barcode can then be linked to any nucleic acid molecule containing a sequence complementary to the restriction site overhang.
[0079] In some implementations, the target binding region may contain a nonspecific target nucleic acid sequence. A nonspecific target nucleic acid sequence can refer to a sequence that can bind to more than one target nucleic acid independently of a specific sequence of the target nucleic acid. For example, the target binding region may contain random multimeric sequences, multi-(dA) sequences, multi-(dT) sequences, multi-(dG) sequences, multi-(dC) sequences, or combinations thereof. For example, the target binding region may be an oligo(dT) sequence that hybridizes to multiple (A) tails on an mRNA molecule. Random multimeric sequences may be, for example, random dimers, trimers, tetramers, pentamers, hexamers, heptamers, octamers, nonamers, decamers, or higher multimeric sequences of any length. In some implementations, the target binding region is the same for all barcodes attached to a given bead. In some implementations, for more than one type of barcode attached to a given bead, the target binding region may include two or more different target binding sequences. The length of the target binding region can be the following or approximately the following: 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, or any number or range of nucleotides between any two of these values. The length of the target binding region can be at most approximately 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, or more nucleotides. For example, a reverse transcriptase such as Moloney murine leukemia virus (MMLV) reverse transcriptase can be used to reverse transcribe mRNA molecules to produce cDNA molecules with multiple (dC) tails. Barcodes can include target binding regions with multiple (dG) tails. After base pairing between the multiple (dG) tail of the barcode and the multiple (dC) tail of the cDNA molecule, the reverse transcriptase converts the template strand from the cellular RNA molecule to the barcode and continues replication towards the 5' end of the barcode. By doing so, the resulting cDNA molecule contains a barcode sequence (such as a molecular marker) at its 3' end.
[0080] In some implementations, the target binding region may contain an oligo(dT) that can hybridize with mRNA containing a polyadenylated terminus. The target binding region may be gene-specific. For example, the target binding region may be configured to hybridize with a specific region of a target. The length of the target binding region may be one of the following or approximately one of the following: 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30 nucleotides or a number or range between any two of these values. The length of the target binding region can be at least or at most the following: 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides. The length of the target binding region can be approximately 5-30 nucleotides. When a barcode contains a gene-specific target binding region, the barcode may be referred to as a gene-specific barcode in this paper.
[0081] Orientation Property A random barcode (e.g., a randomized barcode) may contain one or more orientation properties that can be used for orienting (e.g., alignment) the barcode. The barcode may include portions for isoelectric focusing. Different barcodes may contain different isoelectric focusing points. When these barcodes are introduced into a sample, the sample can undergo isoelectric focusing to orient the barcodes in a known manner. In this way, the orientation properties can be used to develop known mappings of barcodes within a sample. Exemplary orientation properties may include electrophoretic mobility (e.g., based on barcode size), isoelectric point, spin, conductivity, and / or self-assembly. For example, a barcode with self-assembly orientation properties can self-assemble into a specific orientation (e.g., nucleic acid nanostructures) upon activation.
[0082] Affinity Property Barcodes (e.g., random barcodes) may contain one or more affinity properties. For example, spatial markers may contain affinity properties. Affinity properties may include chemical and / or biological portions that can facilitate the binding of the barcode to another entity (e.g., a cell receptor). For example, affinity properties may include antibodies, such as antibodies specific to a particular portion (e.g., a receptor) on a sample. In some embodiments, antibodies may direct the barcode to a specific cell type or molecule. Targets at and / or near a specific cell type or molecule may be labeled (e.g., randomly labeled). In some embodiments, affinity properties may provide spatial information beyond the nucleotide sequence of the spatial marker because the antibody may direct the barcode to a specific location. Antibodies may be therapeutic antibodies, such as monoclonal or polyclonal antibodies. Antibodies may be humanized or chimeric. Antibodies may be naked antibodies or fusion antibodies.
[0083] Antibodies can be full-length (i.e., naturally occurring or formed through normal immunoglobulin gene fragment recombination processes) immunoglobulin molecules (e.g., IgG antibodies) or immunologically active (i.e., specific binding) portions of immunoglobulin molecules (such as antibody fragments).
[0084] Antibody fragments can be, for example, portions of an antibody, such as F(ab')2, Fab', Fab, Fv, sFv, etc. In some embodiments, antibody fragments can bind to the same antigen recognized by a full-length antibody. Antibody fragments can include separate fragments composed of variable regions of an antibody, such as an "Fv" fragment composed of variable regions of a heavy chain and a light chain, and a recombinant single-chain polypeptide molecule ("scFv protein") in which the light chain and heavy chain variable regions are linked by peptide linkers. Exemplary antibodies can include, but are not limited to, cancer cell antibodies, viral antibodies, antibodies that bind to cell surface receptors (CD8, CD34, CD45), and therapeutic antibodies.
[0085] Universal Adapter Primers Barcodes can contain one or more universal adaptor primers. For example, gene-specific barcodes (such as gene-specific random barcodes) can contain universal adaptor primers. A universal adaptor primer can refer to a common nucleotide sequence that is present in all barcodes. Universal adaptor primers can be used to construct gene-specific barcodes. The length of a universal adaptor primer can be the following or approximately the following: 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30 nucleotides or a number or range between any two of these values. The length of a universal adaptor primer can be at least the following or at most the following: 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides. The length of a universal adaptor primer can be 5-30 nucleotides.
[0086] Linkers When a barcode contains more than one type of marker (e.g., more than one cellular marker or more than one barcode sequence, such as a molecular marker), adapter marker sequences may be interspersed among the markers. The length of the adapter marker sequence can be at least about 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, or more nucleotides. The length of the adapter marker sequence can be at most about 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, or more nucleotides. In some cases, the length of the adapter marker sequence is 12 nucleotides. The adapter marker sequence can be used to facilitate barcode synthesis. The adapter marker may include error correction (e.g., Hamming) codes.
[0087] Solid Supports In some embodiments, the barcodes disclosed herein (such as random barcodes) may be associated with a solid support. The solid support may be, for example, synthetic particles. In some embodiments, some or all of the barcode sequences (such as molecular markers of random barcodes (e.g., the first barcode sequence) on the solid support differ by at least one nucleotide. The cell markers of barcodes on the same solid support may be identical. The cell markers of barcodes on different solid supports may differ by at least one nucleotide. For example, the first cell marker of a first barcode on a first solid support may have the same sequence, and the second cell marker of a second barcode on a second solid support may have the same sequence. The first cell marker of a first barcode on a first solid support and the second cell marker of a second barcode on a second solid support may differ by at least one nucleotide. The cell marker may be, for example, about 5-20 nucleotides long. The barcode sequence may be, for example, about 5-20 nucleotides long. The synthetic particles may be, for example, beads.
[0088] The beads can be, for example, silica gel beads, controlled-aperture glass beads, magnetic beads, Dynabead, Sephadex / agarose gel beads, cellulose beads, polystyrene beads, or any combination thereof. Beads can include materials such as polydimethylsiloxane (PDMS), polystyrene, glass, polypropylene, agarose, gelatin, hydrogel, paramagnetic materials, ceramics, plastics, glass, methylstyrene, acrylic polymers, titanium, latex, agarose gel, cellulose, nylon, silicone, or any combination thereof.
[0089] In some embodiments, the beads may be polymer beads (e.g., deformable beads or gel beads) functionalized with barcodes or random barcodes (such as gel beads from 10X Genomics (San Francisco, CA)). In some embodiments, the gel beads may comprise a polymer-based gel. Gel beads may be produced, for example, by encapsulating one or more polymer precursors into droplets. Gel beads can be produced after exposing the polymer precursor to a promoter (e.g., tetramethylethylenediamine (TEMED)).
[0090] In some embodiments, the particles may be destructible (e.g., soluble, degradable). For example, polymer beads may dissolve, melt, or degrade under desired conditions. Desired conditions may include environmental conditions. Desired conditions may cause the polymer beads to dissolve, melt, or degrade in a controlled manner. Gel beads may dissolve, melt, or degrade due to chemical, physical, biological, thermal, magnetic, electrical, or light stimulation, or any combination thereof.
[0091] For example, analytes and / or reagents (such as oligonucleotide barcodes) can be coupled / immobilized to the inner surface of a gel bead (e.g., the interior accessible via diffusion of the oligonucleotide barcode and / or the material used to generate the oligonucleotide barcode) and / or the outer surface of the gel bead or any other microcapsule described herein. Coupling / immobilization can be via any form of chemical bonding (e.g., covalent bonds, ionic bonds) or physical phenomena (e.g., van der Waals forces, dipole-dipole interactions, etc.). In some embodiments, the coupling / immobilization of the reagents described herein with the gel bead or any other microcapsule can be reversible, such as, for example, via an unstable portion (e.g., via a chemical crosslinker, including those described herein). Upon application of a stimulus, the unstable portion can be cleaved, releasing the immobilized reagent. In some embodiments, the unstable portion is a disulfide bond. For example, in the case of immobilizing an oligonucleotide barcode to a gel bead via a disulfide bond, exposing the disulfide bond to a reducing agent can cleave the disulfide bond and release the oligonucleotide barcode from the bead. The unstable portion may be included as part of a gel bead or microcapsule, as part of a chemical connector linking a reagent or analyte to the gel bead or microcapsule, and / or as part of the reagent or analyte. In some embodiments, at least one barcode of more than one type may be affixed to the particle, partially affixed to the particle, encapsulated in the particle, partially encapsulated in the particle, or any combination thereof.
[0092] In some embodiments, the gel beads may comprise a wide range of different polymers, including but not limited to: polymers, thermosensitive polymers, photosensitive polymers, magnetic polymers, pH-sensitive polymers, salt-sensitive polymers, chemically sensitive polymers, polyelectrolytes, polysaccharides, peptides, proteins and / or plastics. The polymer may include, but is not limited to, the following materials: such as poly(N-isopropylacrylamide) (PNIPAAm), poly(styrene sulfonate) (PSS), poly(allylamine) (PAAm), poly(acrylic acid) (PAA), poly(ethyleneimine) (PEI), poly(diallyldimethylammonium chloride) (PDADMAC), poly(pyrrole) (PPy), poly(vinylpyrrolidone) (PVPON), poly(vinylpyridine) (PVP), poly(methacrylic acid) (PMAA), poly(methyl methacrylate) (PMMA), polystyrene (PS), poly(tetrahydrofuran) (PTHF), poly(phthalaldehyde) (PTHF), poly(hexyl viologen) (PHV), poly(L-lysine) (PLL), poly(L-arginine) (PARG), and poly(lactic-co-hydroxyacetic acid) (PLGA).
[0093] Many chemical stimuli can be used to trigger the destruction, dissolution, or degradation of beads. Examples of these chemical changes include, but are not limited to, pH-mediated bead wall alterations, bead wall disintegration via chemical cleavage of cross-links, triggered depolymerization of the bead wall, and bead wall conversion reactions. Bulk changes can also be used to trigger bead destruction.
[0094] The ability to induce bulk or physical alterations in microcapsules through various stimuli also offers numerous advantages in designing capsules for reagent release. These bulk or physical alterations occur on a macroscopic scale, where bead rupture is the result of mechanical-physical forces induced by stimuli. These processes can include, but are not limited to, pressure-induced rupture, bead wall melting, or changes in bead wall porosity.
[0095] Biostimuli can also be used to trigger the destruction, dissolution, or degradation of beads. Typically, biotriggers are similar to chemical triggers, but many examples use molecules common in biomolecules or living systems, such as enzymes, peptides, sugars, fatty acids, and nucleic acids. For example, beads can contain polymers with peptide cross-links sensitive to cleavage by specific proteases. More specifically, one example may include microcapsules containing GFLGK peptide cross-links. Upon addition of a biotrigger (such as the protease cathepsin B), the peptide cross-links of the shell wall are cleaved and the contents of the bead are released. In other cases, the protease may be thermally activated. In another example, the beads comprise a shell wall containing cellulose. The addition of chitosan hydrolases acts as a biotrigger for cellulose bond cleavage, shell wall depolymerization, and the release of the internal contents.
[0096] Applying heat can also induce the beads to release their contents. Changes in temperature can cause various alterations in the beads. Changes in heat can cause the beads to melt, leading to the disintegration of the bead wall. In other cases, heat can increase the internal pressure of the bead's internal components, causing the bead to rupture or explode. In still other cases, heat can cause the beads to transform into a shrinking, dehydrated state. Heat can also act on the heat-sensitive polymers within the bead wall, thereby causing the bead to break.
[0097] Incorporating magnetic nanoparticles within the bead walls of microcapsules allows for triggered breakage of the beads and the guidance of the beads into an array. Devices of this disclosure may include magnetic beads for either purpose. In one example, Fe3O4 nanoparticles are incorporated into beads containing a polyelectrolyte, triggering breakage in the presence of an oscillating magnetic field stimulus.
[0098] Beads can also be broken, dissolved, or degraded as a result of electrical stimulation. Similar to the magnetic particles described in the previous section, electrosensitive beads can allow for triggered breakage and other functions, such as alignment in an electric field, conductivity, or redox reactions. In one example, beads containing electrosensitive materials align in an electric field, thereby allowing control over the release of internal reagents. In other examples, the electric field can induce redox reactions within the bead wall itself, which can increase porosity.
[0099] Photostimulation can also be used to disrupt beads. Many phototriggers are possible, and systems using a variety of molecules, such as nanoparticles and chromophores capable of absorbing photons in specific wavelength ranges, can be included. For example, metal oxide coatings can be used as capsule triggers. UV irradiation of polyelectrolyte capsules coated with SiO2 can cause the bead walls to disintegrate. In yet another example, photoswitching materials, such as azophenyl groups, can be incorporated into the bead walls. Upon application of UV or visible light, these chemicals undergo reversible cis-to-trans isomerization after absorbing photons. In this respect, the incorporation of a photon switch produces bead walls that can disintegrate or become more porous upon application of a phototrigger.
[0100] For example, in Figure 2 In a non-limiting example of barcoding (e.g., random barcoding) illustrated in the diagram, after introducing a cell (such as a single cell) onto more than one well of the microwell array at box 208, beads can be introduced onto more than one well of the microwell array at box 212. Each well may contain one bead. The bead may contain more than one type of barcode. The barcode may contain a 5' amine region attached to the bead. The barcode may contain a universal marker, a barcode sequence (e.g., a molecular marker), a target-binding region, or any combination thereof.
[0101] The barcodes disclosed herein can be associated (e.g., attached) to a solid support (e.g., beads). Each barcode associated with a solid support may contain a barcode sequence selected from the group consisting of at least 100 or 1000 barcode sequences having a unique sequence. In some embodiments, different barcodes associated with a solid support may contain barcodes with different sequences. In some embodiments, a certain percentage of the barcodes associated with a solid support contains the same cell marker. For example, the percentage may be or may be about the following: 60%, 70%, 80%, 85%, 90%, 95%, 97%, 99%, 100%, or a number or range between any two of these values. As another example, the percentage may be at least or at most the following: 60%, 70%, 80%, 85%, 90%, 95%, 97%, 99%, or 100%. In some embodiments, the barcodes associated with a solid support may have the same cell marker. Barcodes associated with different solid supports may have different cell markers selected from the group consisting of at least 100 or 1000 cell markers with unique sequences.
[0102] The barcodes disclosed herein can be associated (e.g., attached) to a solid support (e.g., beads). In some embodiments, more than one target in a sample can be barcoded using a solid support comprising more than one synthetic particle associated with more than one barcode. In some embodiments, the solid support may comprise more than one synthetic particle associated with more than one barcode. The spatial markers of more than one barcode on different solid supports may differ by at least one nucleotide. The solid support may, for example, comprise more than one barcode in two or three dimensions. The synthetic particle may be a bead. Beads may be silica beads, controlled-aperture glass beads, magnetic beads, Dynabead, Sephadex / agarose gel beads, cellulose beads, polystyrene beads, or any combination thereof. The solid support may comprise a polymer, matrix, hydrogel, needle array device, antibody, or any combination thereof. In some embodiments, the solid support may be free-floating. In some embodiments, the solid support may be embedded in a semi-solid or solid array. The barcode may not be associated with the solid support. The barcode may be a single nucleotide. The barcode may be associated with a substrate.
[0103] As used herein, the terms “tethered,” “attached,” and “fixed” are used interchangeably and can refer to covalent or non-covalent methods for attaching barcodes to solid supports. Any of a variety of solid supports can be used as a solid support for attaching pre-synthesized barcodes or for in-situ solid-phase synthesis of barcodes.
[0104] In some embodiments, the solid support is a bead. Beads may include one or more types of solid, porous, or hollow spheres, balls, supports, cylinders, or other similar configurations that can immobilize nucleic acids (e.g., covalently or non-covalently). Beads may comprise, for example, plastic, ceramic, metal, polymeric materials, or any combination thereof. Beads may be or comprise spherical (e.g., microspheres) or discrete particles with non-spherical or irregular shapes, such as cubic, rectangular, conical, cylindrical, elliptical, or disk-shaped. In some embodiments, the bead shape may be non-spherical.
[0105] Beads can include a variety of materials, including but not limited to paramagnetic materials (e.g., magnesium, molybdenum, lithium, and tantalum), superparamagnetic materials (e.g., ferrite (Fe3O4; magnetite) nanoparticles), ferromagnetic materials (e.g., iron, nickel, cobalt, some alloys thereof, and some rare earth metal compounds), ceramics, plastics, glass, polystyrene, silica, methyl styrene, acrylic polymers, titanium, latex, agarose gel, agarose, hydrogel, polymers, cellulose, nylon, or any combination thereof.
[0106] In some embodiments, the beads (e.g., the beads to which the marker is attached) are hydrogel beads. In some embodiments, the beads comprise hydrogel.
[0107] Some embodiments disclosed herein include one or more particles (e.g., beads). Each particle may contain more than one type of oligonucleotide (e.g., barcode). Each of the more than one oligonucleotide may contain a barcode sequence (e.g., molecular marker sequence), a cell marker, and a target-binding region (e.g., oligo(dT) sequence, gene-specific sequence, random multimer, or a combination thereof). The cell marker sequence for each of the more than one oligonucleotide may be identical. The cell marker sequences for oligonucleotides on different particles may be different, allowing identification of oligonucleotides on different particles. The number of different cell marker sequences may vary in different embodiments. In some implementations, the number of cell marker sequences may be, or may be approximately, the following: 10, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 20000, 30000, 40000, 50000, 60000, 70000, 80000, 90000, 100000, 10 6 10 7 10 8 10 9These values are numbers or ranges between any two of these values, or more. In some implementations, the number of cell marker sequences may be at least the following or at most the following: 10, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 20000, 30000, 40000, 50000, 60000, 70000, 80000, 90000, 100000, 10 6 10 7 10 8 Or 10 9 In some embodiments, no more than 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000 or more particles in more than one particle may comprise oligonucleotides having the same cell sequence. In some embodiments, more than one particle comprising oligonucleotides having the same cell sequence may be up to 0.1%, 0.2%, 0.3%, 0.4%, 0.5%, 0.6%, 0.7%, 0.8%, 0.9%, 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10% or more. In some embodiments, more than one particle may not all have the same cell marker sequence.
[0108] Each particle may contain more than one oligonucleotide that includes different barcode sequences (e.g., molecular markers). In some embodiments, the number of barcode sequences may be, or may be approximately, the following: 10, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 20000, 30000, 40000, 50000, 60000, 70000, 80000, 90000, 100000, 10 6 10 7 10 8 10 9Or, a number or range between any two of these values. In some implementations, the number of barcode sequences may be at least the following or at most the following: 10, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 20000, 30000, 40000, 50000, 60000, 70000, 80000, 90000, 100000, 10 6 10 7 10 8 Or 10 9 For example, at least 100 of more than one oligonucleotide contain different barcode sequences. As another example, in a single particle, at least 100, 500, 1000, 5000, 10000, 15000, 20000, 50000, or more oligonucleotides, any number or range between any two of these values, contain different barcode sequences. Some embodiments provide more than one particle containing barcodes. In some embodiments, the ratio of the target to be labeled to the occurrence (or copies or number) of different barcode sequences can be at least 1:1, 1:2, 1:3, 1:4, 1:5, 1:6, 1:7, 1:8, 1:9, 1:10, 1:11, 1:12, 1:13, 1:14, 1:15, 1:16, 1:17, 1:18, 1:19, 1:20, 1:30, 1:40, 1:50, 1:60, 1:70, 1:80, 1:90, or higher. In some embodiments, each of more than one oligonucleotide also includes a sample marker, a universal marker, or both. The particles can be, for example, nanoparticles or microparticles.
[0109] The size of the beads can vary. For example, the diameter of the beads can range from 0.1 micrometers to 50 micrometers. In some embodiments, the diameter of the beads can be the following or about the following: 0.1 micrometer, 0.5 micrometer, 1 micrometer, 2 micrometer, 3 micrometer, 4 micrometer, 5 micrometer, 6 micrometer, 7 micrometer, 8 micrometer, 9 micrometer, 10 micrometer, 20 micrometer, 30 micrometer, 40 micrometer, 50 micrometer, or any number or range between two of these values.
[0110] The diameter of the bead can be related to the diameter of the pores in the substrate. In some embodiments, the bead diameter can be longer or shorter than the pore diameter by at least or at most the following: 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, or any number or range between these values. The diameter of the bead can also be related to the diameter of a cell (e.g., a single cell trapped by the pores in the substrate). In some embodiments, the bead diameter can be longer or shorter than the pore diameter by at least or at most the following: 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 100%. The diameter of the bead can also be related to the diameter of a cell (e.g., a single cell trapped by the pores in the substrate). In some embodiments, the diameter of the bead may be longer or shorter than the diameter of the cell by or approximately the following: 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, 150%, 200%, 250%, 300%, or any number or range between these values. In some embodiments, the diameter of the bead may be longer or shorter than the diameter of the cell by at least or at most the following: 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, 150%, 200%, 250%, or 300%.
[0111] Beads may be attached to and / or embedded in a substrate. Beads may be attached to and / or embedded in a gel, hydrogel, polymer, and / or matrix. The spatial location of the bead within the substrate (e.g., gel, matrix, scaffold, or polymer) can be identified using a spatial marker present on a barcode on the bead, which can be used as a location address.
[0112] Examples of beads may include, but are not limited to, streptoacidin beads, agarose beads, magnetic beads, Dynabeads®, MACS® microbeads, antibody-conjugated beads (e.g., anti-immunoglobulin microbeads), protein A-conjugated beads, protein G-conjugated beads, protein A / G-conjugated beads, protein L-conjugated beads, oligo(dT)-conjugated beads, silica beads, silica-like beads, avidin microbeads, anti-fluorescent dye microbeads, and BcMag. TM Carboxyl-terminated magnetic beads.
[0113] Beads can be associated with quantum dots or fluorescent dyes (e.g., impregnated with quantum dots or fluorescent dyes) to make them fluoresce in one or more fluorescent optical channels. Beads can be associated with iron oxide or chromium oxide to make them paramagnetic or ferromagnetic. Beads can be identifiable. For example, beads can be imaged using a camera. Beads can have a detectable code associated with them. For example, beads can contain barcodes. Beads can change size, for example, due to swelling in organic or inorganic solutions. Beads can be hydrophobic. Beads can be hydrophilic. Beads can be biocompatible.
[0114] Solid supports (e.g., beads) can be visualized. Solid supports may contain visual labels (e.g., fluorescent dyes). Solid supports (e.g., beads) may be etched with identifiers (e.g., numbers). Identifiers can be visualized by imaging the beads.
[0115] Solid supports can include soluble, semi-soluble, or insoluble materials. A solid support may be referred to as "functionalized" when it includes attached connectors, supports, building blocks, or other reactive portions, and as "unfunctionalized" when it lacks such attached reactive portions. Solid supports can be free in solution, such as in microburette wells; in flow-through form, such as in columns; or used as dipsticks.
[0116] Solid supports can include membranes, paper, plastics, coated surfaces, flat surfaces, glass, glass slides, chips, or any combination thereof. Solid supports can take the form of resins, gels, microspheres, or other geometric configurations. Solid supports can include silica chips, micron-sized particles, nanoparticles, plates, arrays, capillaries, flat supports such as glass fiber filters, glass surfaces, metal surfaces (steel, gold, silver, aluminum, silicon, and copper), glass supports, plastic supports, silicon supports, chips, filters, membranes, microplates, glass slides, plastic materials including porous plates or membranes (e.g., formed from polyethylene, polypropylene, polyamide, polyvinylidene fluoride), and / or wafers, combs, needles, or needle tips (e.g., needle arrays suitable for combined synthesis or analysis) or beads, flat surfaces such as recessed or nanoporous arrays of wafers (e.g., silicon wafers), and wafers with recesses (with or without filter bottoms).
[0117] Solid supports may include polymer matrices (e.g., gels, hydrogels). Polymer matrices may be able to permeate intracellular spaces (e.g., around organelles). Polymer matrices may be able to be pumped throughout the circulatory system.
[0118] Substrates and Microwell Arrays As used herein, a substrate can refer to a type of solid support. A substrate can refer to a solid support that may contain a barcode or random barcode of this disclosure. A substrate may, for example, include more than one microwell. A substrate may, for example, be a pore array comprising two or more microwells. In some embodiments, the microwells may include small reaction chambers of defined volume. In some embodiments, the microwells may capture one or more cells. In some embodiments, the microwells may capture only one cell. In some embodiments, the microwells may capture one or more solid supports. In some embodiments, the microwells may capture only one solid support. In some embodiments, the microwells capture a single cell and a single solid support (e.g., a bead). The microwells may contain barcode reagents of this disclosure.
[0119] Methods of Barcoding This disclosure provides methods for estimating the number of distinct targets at different locations in a body sample (e.g., tissue, organ, tumor, cell). The methods may include placing a barcode (e.g., a random barcode) close to the sample, lysing the sample, associating different targets with the barcode, amplifying the targets, and / or digitally counting the targets. The methods may also include analyzing and / or visualizing information obtained from spatial markings on the barcode. In some embodiments, the methods include visualizing more than one target in the sample. Mapping more than one target onto a map of the sample may include generating a two-dimensional or three-dimensional map of the sample. The two-dimensional and three-dimensional maps may be generated before or after barcoding (e.g., random barcoding) the more than one target in the sample. Visualizing more than one target in the sample may include mapping the more than one target onto a map of the sample. Mapping more than one target onto a map of the sample may include generating a two-dimensional or three-dimensional map of the sample. The two-dimensional and three-dimensional maps may be generated before or after barcoding the more than one target in the sample. In some implementations, two-dimensional and three-dimensional mapping maps can be generated before or after pyrolyzing the sample. Pyrolyzing the sample before or after generating the two-dimensional or three-dimensional mapping map may include heating the sample, contacting the sample with a detergent, altering the pH of the sample, or any combination thereof.
[0120] In some implementations, barcoding more than one target includes hybridizing more than one barcode with more than one target to produce a barcoded target (e.g., a random barcoded target). Barcoding more than one target may include an index library that produces the barcoded target. The index library that produces the barcoded target can be produced using a solid support containing more than one barcode (e.g., a random barcode).
[0121] Contacting a Sample and Barcodes This disclosure provides methods for contacting a sample (e.g., cells) with a substrate of this disclosure. Samples, including, for example, thin sections of cells, organs, or tissues, can be contacted with barcodes (e.g., random barcodes). Cells can be contacted, for example, by gravity flow, where cells can settle and form a monolayer. The sample can be a thin section of tissue. The thin section can be placed on the substrate. The sample can be one-dimensional (e.g., forming a flat surface). The sample (e.g., cells) can be dispersed throughout the substrate, for example, by growing / culturing cells on the substrate.
[0122] When a barcode is brought close to a target, the target can hybridize with the barcode. The barcode can contact the target in an inexhaustible ratio, allowing each different target to be associated with a different barcode of this disclosure. To ensure effective association between the target and the barcode, the target and the barcode can be crosslinked.
[0123] Cell Lysis Following cell and barcode assignment, cells can be lysed to release target molecules. Cell lysis can be accomplished by any of a variety of means, such as chemical or biochemical methods, osmotic shock, or thermal, mechanical, or optical lysis. Cells can be lysed by adding a cell lysis buffer containing detergents (e.g., SDS, lithium dodecyl sulfate, Triton X-100, Tween 20, or NP-40), organic solvents (e.g., methanol or acetone), or digestive enzymes (e.g., proteinase K, pepsin, or trypsin), or any combination thereof. To increase the association between the target and the barcode, the diffusion rate of the target molecules can be altered, for example, by decreasing the temperature of the lysate and / or increasing the viscosity of the lysate.
[0124] In some implementations, filter paper can be used to lyse the sample. The filter paper can be soaked in lysis buffer. Pressure can be applied to the filter paper onto the sample, which can promote sample lysis and hybridization between the sample's target and substrate.
[0125] In some embodiments, lysis can be performed by mechanical lysis, thermal lysis, optical lysis, and / or chemical lysis. Chemical lysis may include the use of digestive enzymes such as proteinase K, pepsin, and trypsin. Lysis can be performed by adding a lysis buffer to the substrate. The lysis buffer may contain Tris HCl. The lysis buffer may contain at least about 0.01 M, 0.05 M, 0.1 M, 0.5 M, or 1 M or more of Tris HCl. The lysis buffer may contain up to about 0.01 M, 0.05 M, 0.1 M, 0.5 M, or 1 M or more of Tris HCl. The lysis buffer may contain about 0.1 M Tris HCl. The pH of the lysis buffer may be at least about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or higher. The pH of the lysis buffer may be up to about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or higher. In some embodiments, the pH of the lysis buffer is about 7.5. The lysis buffer may contain a salt (e.g., LiCl). The salt concentration in the lysis buffer can be at least about 0.1 M, 0.5 M, or 1 M or higher. The salt concentration in the lysis buffer can be at most about 0.1 M, 0.5 M, or 1 M or higher. In some embodiments, the salt concentration in the lysis buffer is about 0.5 M. The lysis buffer may contain a detergent (e.g., SDS, lithium dodecyl sulfate, Triton X, Tween, NP-40). The detergent concentration in the lysis buffer can be at least about 0.0001%, 0.0005%, 0.001%, 0.005%, 0.01%, 0.05%, 0.1%, 0.5%, 1%, 2%, 3%, 4%, 5%, 6%, or 7% or higher. The detergent concentration in the lysis buffer can be at most about 0.0001%, 0.0005%, 0.001%, 0.005%, 0.01%, 0.05%, 0.1%, 0.5%, 1%, 2%, 3%, 4%, 5%, 6%, or 7% or higher. In some embodiments, the detergent concentration in the lysis buffer is about 1% lithium dodecyl sulfate. The time used in the lysis method can depend on the amount of detergent used. In some embodiments, the more detergent used, the less time is required for lysis. The lysis buffer may contain a chelating agent (e.g., EDTA, EGTA). The concentration of the chelating agent in the lysis buffer can be at least about 1 mM, 5 mM, 10 mM, 15 mM, 20 mM, 25 mM, or 30 mM or higher. The concentration of the chelating agent in the lysis buffer can be at most about 1 mM, 5 mM, 10 mM, 15 mM, 20 mM, 25 mM, or 30 mM or higher. In some embodiments, the chelating agent concentration in the lysis buffer is about 10 mM. The lysis buffer may contain a reducing agent (e.g., β-mercaptoethanol, DTT).The concentration of the reducing agent in the lysis buffer can be at least about 1 mM, 5 mM, 10 mM, 15 mM, or 20 mM or higher. The concentration of the reducing agent in the lysis buffer can be at most about 1 mM, 5 mM, 10 mM, 15 mM, or 20 mM or higher. In some embodiments, the concentration of the reducing agent in the lysis buffer is about 5 mM. In some embodiments, the lysis buffer may contain about 0.1 M Tris HCl, about pH 7.5, about 0.5 M LiCl, about 1% lithium dodecyl sulfate, about 10 mM EDTA, and about 5 mM DTT.
[0126] Lysis can be performed at temperatures of approximately 4°C, 10°C, 15°C, 20°C, 25°C, or 30°C. Lysis can be performed for approximately 1 minute, 5 minutes, 10 minutes, 15 minutes, or 20 minutes or more. The lysed cells may include at least approximately 100,000, 200,000, 300,000, 400,000, 500,000, 600,000, or 700,000 or more target nucleic acid molecules. The lysed cells may include up to approximately 100,000, 200,000, 300,000, 400,000, 500,000, 600,000, or 700,000 or more target nucleic acid molecules.
[0127] Attaching Barcodes to Target Nucleic Acid Molecules Following cell lysis and the release of nucleic acid molecules from the cells, the nucleic acid molecules can be randomly associated with barcodes on a co-localized solid support. Association may include hybridizing the target recognition region of the barcode with a complementary portion of the target nucleic acid molecule (e.g., the oligo(dT) of the barcode may interact with the multi(A) tail of the target). Assay conditions for hybridization (e.g., buffer pH, ionic strength, temperature, etc.) can be selected to promote the formation of specific, stable hybrids. In some embodiments, nucleic acid molecules released from lysed cells can be associated with more than one probe on the substrate (e.g., hybridization with probes on the substrate). When the probe contains an oligo(dT), mRNA molecules can be hybridized to the probe and reverse transcribed. The oligo(dT) portion of the oligonucleotide can act as a primer for the first-strand synthesis of cDNA molecules. For example, in… Figure 2 In the non-restrictive example of barcoding illustrated in the diagram, at box 216, the mRNA molecule can hybridize with the barcode on the bead. For example, a single-stranded nucleotide fragment can hybridize with the target binding region of the barcode.
[0128] Attachment may also involve linking the target recognition region of the barcode to a portion of the target nucleic acid molecule. For example, the target binding region may contain a nucleic acid sequence capable of specifically hybridizing to restriction site overhangs (e.g., EcoRI sticky end overhangs). The assay procedure may also include treating the target nucleic acid with a restriction enzyme (e.g., EcoRI) to generate restriction site overhangs. The barcode can then be ligated to any nucleic acid molecule containing a sequence complementary to the restriction site overhang. A ligase (e.g., T4 DNA ligase) may be used to ligate the two fragments.
[0129] For example, in Figure 2 In a non-limiting example of barcoding illustrated in the diagram, at box 220, labeled targets (e.g., target-barcode molecules) from more than one cell (or more than one sample) can then be collected into, for example, a tube. The labeled targets can be collected, for example, by retrieving the barcodes and / or attaching beads to the target-barcode molecules.
[0130] The recovery of attached target-barcode molecules from solid-support-based assemblies can be achieved using magnetic beads and an externally applied magnetic field. After assembling the target-barcode molecules, all further processing can be performed in a single reaction vessel. Further processing may include, for example, reverse transcription, amplification, lysis, dissociation, and / or nucleic acid extension reactions. These further processing reactions can be carried out within microwells, i.e., without first assembling labeled target nucleic acid molecules from more than one cell.
[0131] Reverse Transcription or Nucleic Acid Extension This disclosure provides the use of reverse transcription (e.g., in...) Figure 2 Methods for generating target-barcode conjugates include using the sequence of the target RNA (frame 224) or nucleic acid extension. Target-barcode conjugates may contain a barcode and all or part of the complementary sequence of the target nucleic acid (i.e., a barcoded cDNA molecule, such as a randomly barcoded cDNA molecule). Reverse transcription of the associated RNA molecule can occur by adding a reverse transcription primer along with reverse transcriptase. Reverse transcription primers can be oligo(dT) primers, random hexanucleotide primers, or target-specific oligonucleotide primers. Oligo(dT) primers can be 12-18 nucleotides in length or about 12-18 nucleotides in length and bind to an endogenous poly(A) tail at the 3' end of mammalian mRNA. Random hexanucleotide primers can bind to mRNA at various complementary sites. Target-specific oligonucleotide primers typically selectively priming the mRNA of interest.
[0132] In some implementations, reverse transcription of mRNA molecules to labeled RNA molecules can occur by adding reverse transcription primers. In some implementations, the reverse transcription primers are oligo(dT) primers, random hexanucleotide primers, or target-specific oligonucleotide primers. Typically, oligo(dT) primers are 12-18 nucleotides in length and bind to an endogenous poly(A) tail at the 3' end of mammalian mRNA. Random hexanucleotide primers can bind to mRNA at various complementary sites. Target-specific oligonucleotide primers typically selectively priming the mRNA of interest.
[0133] In some implementations, the target is a cDNA molecule. For example, a reverse transcriptase such as Moloney murine leukemia virus (MMLV) reverse transcriptase can be used to reverse transcribe an mRNA molecule to produce a cDNA molecule with multiple (dC) tails. The barcode may include a target-binding region with multiple (dG) tails. After base pairing between the multiple (dG) tail of the barcode and the multiple (dC) tail of the cDNA molecule, the reverse transcriptase transfers the template strand from the cellular RNA molecule to the barcode and continues replication towards the 5' end of the barcode. By doing so, the resulting cDNA molecule contains a barcode sequence (such as a molecular marker) at the 3' end of the cDNA molecule.
[0134] Reverse transcription can occur repeatedly to produce more than one labeled cDNA molecule. The methods disclosed herein may include performing at least about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 reverse transcription reactions. Methods may also include performing at least about 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, or 100 reverse transcription reactions.
[0135] Amplification One or more nucleic acid amplification reactions can be performed (e.g., in...) Figure 2(See box 228) to generate more than one copy of the labeled target nucleic acid molecule. Amplification can be performed in a multiplexed manner, wherein more than one target nucleic acid sequence is amplified simultaneously. The amplification reaction can be used to add a sequencing adaptor to the nucleic acid molecule. The amplification reaction may include at least a portion of an amplified sample label (if present). The amplification reaction may include at least a portion of an amplified cellular label and / or barcode sequence (e.g., a molecular label). The amplification reaction may include at least a portion of an amplified sample tag, cellular label, spatial label, barcode sequence (e.g., a molecular label), target nucleic acid, or a combination thereof. The amplification reaction may include amplifying more than one nucleic acid at a rate of 0.5%, 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 97%, 100%, or any two of these values or a number. The method may also include performing one or more cDNA synthesis reactions to produce one or more cDNA copies of a target-barcode molecule containing sample markers, cell markers, spatial markers and / or barcode sequences (e.g., molecular markers).
[0136] In some implementations, amplification can be performed using polymerase chain reaction (PCR). As used herein, PCR can refer to a reaction used to amplify a specific DNA sequence in vitro by simultaneously extending primers with complementary strands of the DNA. As used herein, PCR can encompass derivative forms of the reaction, including but not limited to RT-PCR, real-time PCR, nested PCR, quantitative PCR, multiplex PCR, digital PCR, and assembly PCR.
[0137] Amplification of labeled nucleic acids can include non-PCR-based methods. Examples of non-PCR-based methods include, but are not limited to, multiple substitution amplification (MDA), transcription-mediated amplification (TMA), nucleic acid sequence-based amplification (NASBA), strand substitution amplification (SDA), real-time SDA, rolling circle amplification, or circle-to-circle amplification. Other non-PCR-based amplification methods include DNA-dependent RNA polymerase-driven RNA transcription amplification or RNA-directed DNA synthesis and transcription in more than one cycle to amplify DNA or RNA targets, ligase chain reaction (LCR) and Qβ replicase (Qβ) methods, the use of palindromic probes, strand substitution amplification, oligonucleotide-driven amplification using restriction endonucleases, amplification methods that hybridize primers to nucleic acid sequences and cleave the resulting duplexes before extension and amplification, strand substitution amplification using nucleic acid polymerases lacking 5' exonuclease activity, rolling circle amplification, and branch extension amplification (RAM). In some embodiments, amplification does not produce circularized transcripts.
[0138] In some embodiments, the methods disclosed herein further include performing a polymerase chain reaction on a labeled nucleic acid (e.g., labeled RNA, labeled DNA, labeled cDNA) to generate labeled amplicons (e.g., randomly labeled amplicons). The labeled amplicons may be double-stranded molecules. Double-stranded molecules may include double-stranded RNA molecules, double-stranded DNA molecules, or RNA molecules that hybridize with DNA molecules. One or both strands of the double-stranded molecule may contain sample markers, spatial markers, cellular markers, and / or barcode sequences (e.g., molecular markers). The labeled amplicons may be single-stranded molecules. Single-stranded molecules may include DNA, RNA, or combinations thereof. The nucleic acids of this disclosure may include synthetic or modified nucleic acids.
[0139] Amplification may involve the use of one or more non-natural nucleotides. Non-natural nucleotides may include light-labile or triggerable nucleotides. Examples of non-natural nucleotides may include, but are not limited to, peptide nucleic acids (PNAs), morpholino and locked nucleic acids (LNAs), and glycol nucleic acids (GNAs) and threononucleotides (TNAs). Non-natural nucleotides may be added to one or more cycles of the amplification reaction. The addition of non-natural nucleotides can be used to identify the products at specific cycles or time points in the amplification reaction.
[0140] Performing one or more amplification reactions may involve using one or more primers. One or more primers may include, for example, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15 or more nucleotides. One or more primers may include at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15 or more nucleotides. One or more primers may contain fewer than 12-15 nucleotides. One or more primers may anneal at least a portion of more than one labeled target (e.g., a randomly labeled target). One or more primers may anneal to the 3' or 5' end of more than one labeled target. One or more primers may anneal to the internal region of more than one labeled target. The internal region may be at least approximately 50, 100, 150, 200, 220, 230, 240, 250, 260, 270, 280, 290, 300, 310, 320, 330, 340, 350, 360, 370, 380, 390, or 400 units away from the 3' end of more than one marked target. 410, 420, 430, 440, 450, 460, 470, 480, 490, 500, 510, 520, 530, 540, 550, 560, 570, 580, 590, 600, 650, 700, 750, 800, 850, 900, or 1000 nucleotides. One or more primers may include a fixed set of primers. One or more primers may include at least one or more custom primers. One or more primers may include at least one or more control primers. One or more primers may include at least one or more gene-specific primers.
[0141] One or more primers may include universal primers. Universal primers may be annealed to universal primer binding sites. One or more custom primers may be annealed to a first sample marker, a second sample marker, a spatial marker, a cellular marker, a barcode sequence (e.g., a molecular marker), a target, or any combination thereof. One or more primers may include both universal and custom primers. Custom primers may be designed to amplify one or more targets. Targets may include a subset of total nucleic acids in one or more samples. Targets may include a subset of total labeled targets in one or more samples. One or more primers may include at least 96 or more custom primers. One or more primers may include at least 960 or more custom primers. One or more primers may include at least 9600 or more custom primers. One or more custom primers may be annealed to two or more different labeled nucleic acids. Two or more different labeled nucleic acids may correspond to one or more genes.
[0142] Any amplification protocol can be used in the methods described in this disclosure. For example, in one protocol, the first round of PCR can amplify the bead-attached molecule using gene-specific primers and primers targeting the universal Illumina sequencing primer 1 sequence. The second round of PCR can amplify the first PCR product using nested gene-specific primers flanked by the Illumina sequencing primer 2 sequence and primers targeting the universal Illumina sequencing primer 1 sequence. The third round of PCR adds P5 and P7, as well as a sample index, to transform the PCR product into an Illumina sequencing library. Sequencing using 150 bp × 2 sequencing can reveal cellular markers and barcode sequences (e.g., molecular markers) on read 1, genes on read 2, and sample indexes on index 1 reads.
[0143] In some embodiments, nucleic acids can be removed from the substrate using chemical cleavage. For example, chemical groups or modified bases present in the nucleic acids can be used to facilitate the removal of nucleic acids from the solid support. Enzymes can be used, for example, to remove nucleic acids from the substrate. For example, nucleic acids can be removed from the substrate by digestion with restriction endonucleases. For example, treatment of nucleic acids containing dUTP or ddUTP with uracil-d-glycosidase (UDG) can be used to remove nucleic acids from the substrate. For example, nucleic acids can be removed from the substrate using enzymes that perform nucleotide excision, such as base excision repair enzymes, such as apurinic / apyrimidinic (AP) endonucleases. In some embodiments, photolyzable groups and light can be used to remove nucleic acids from the substrate. In some embodiments, cleavable adapters can be used to remove nucleic acids from the substrate. For example, cleavable adapters can include at least one of the following: biotin / avidin, biotin / streptavitin, biotin / neutral avidin, Igprotein A, photostable adapters, acid- or base-labile adapter groups, or aptamers.
[0144] When the probe is gene-specific, the molecule can be hybridized to the probe and then reverse transcribed and / or amplified. In some implementations, the nucleic acid can be amplified after it has been synthesized (e.g., reverse transcribed). Amplification can be performed in multiplexes, where multiple target nucleic acid sequences are amplified simultaneously. Amplification can involve adding sequencing adaptors to the nucleic acid.
[0145] In some implementations, amplification can be performed on a substrate, for example, using bridging amplification. The cDNA can be given a homopolymer tail to produce compatible ends for bridging amplification using oligo(dT) probes on the substrate. In bridging amplification, the primer complementary to the 3' end of the template nucleic acid can be the first primer in each pair of primers covalently attached to the solid particle. When the sample containing the template nucleic acid is contacted with the particle and subjected to a single thermal cycle, the template molecule can be annealed to the first primer, and the first primer is extended forward by adding nucleotides to form a double-stranded molecule consisting of the template molecule and a newly formed DNA strand complementary to the template. In the heating step of the next cycle, the double-stranded molecule can denature, releasing the template molecule from the particle and leaving the complementary DNA strand attached to the particle via the first primer. In the annealing phase of the subsequent annealing and extension steps, the complementary strand can hybridize with a second primer, which is complementary to a segment of the complementary strand at the site where the first primer was removed. This hybridization results in the formation of a bridge between the first and second primers, with the first primer covalently linked and the second primer linked via hybridization. During the extension phase, the second primer can be extended in the reverse direction by adding nucleotides to the same reaction mixture, thus converting the bridge into a double-stranded bridge. The next cycle then begins, and the double-stranded bridge can be denatured to produce two single-stranded nucleic acid molecules, each with one end attached to the particle surface via the first and second primers, respectively, while the other end of each single-stranded nucleic acid molecule remains unattached. In the annealing and extension steps of this second cycle, each strand can hybridize with another previously unused complementary primer on the same particle to form a new single-stranded bridge. The two previously unused primers that are now hybridized are extended, thus converting the two new bridges into double-stranded bridges.
[0146] The amplification reaction may include amplifying at least 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 97%, or 100% of more than one nucleic acid.
[0147] Amplification of labeled nucleic acids can include PCR-based or non-PCR-based methods. Amplification of labeled nucleic acids can include exponential amplification of the labeled nucleic acid. Amplification of labeled nucleic acids can include linear amplification of the labeled nucleic acid. Amplification can be performed using polymerase chain reaction (PCR). PCR can refer to a reaction used to amplify a specific DNA sequence in vitro by simultaneously extending primers with complementary strands of the DNA. PCR can encompass derivative forms of the reaction, including but not limited to RT-PCR, real-time PCR, nested PCR, quantitative PCR, multiplex PCR, digital PCR, suppression PCR, semi-suppression PCR, and assembly PCR.
[0148] In some implementations, the amplification of labeled nucleic acids includes non-PCR-based methods. Examples of non-PCR-based methods include, but are not limited to, multiple displacement amplification (MDA), transcription-mediated amplification (TMA), nucleic acid sequence-based amplification (NASBA), strand displacement amplification (SDA), real-time SDA, rolling circle amplification, or loop-to-loop amplification. Other non-PCR-based amplification methods include DNA-dependent RNA polymerase-driven RNA transcriptional amplification or RNA-directed DNA synthesis and transcription in more than one cycle to amplify DNA or RNA targets, ligase chain reaction (LCR), Qβ replicase (Qβ) methods, the use of palindromic probes, strand displacement amplification, oligonucleotide-driven amplification using restriction endonucleases, amplification methods that hybridize primers to nucleic acid sequences and cleave the resulting duplexes before extension and amplification, strand displacement amplification using nucleic acid polymerases lacking 5' exonuclease activity, rolling circle amplification, and / or branched extension amplification (RAM).
[0149] In some embodiments, the methods disclosed herein further include a nested polymerase chain reaction (PCR) of the amplified amplicons (e.g., targets). The amplicons may be double-stranded molecules. Double-stranded molecules may include double-stranded RNA molecules, double-stranded DNA molecules, or RNA molecules hybridized to DNA molecules. One or both strands of the double-stranded molecule may contain a sample tag or molecular identifier. Optionally, the amplicons may be single-stranded molecules. Single-stranded molecules may include DNA, RNA, or combinations thereof. The nucleic acids of the present invention may include synthetic or modified nucleic acids.
[0150] In some embodiments, the method includes repeatedly amplifying the labeled nucleic acid to generate more than one amplicon. The methods disclosed herein may include performing at least about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 amplification reactions. Optionally, the method includes performing at least about 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, or 100 amplification reactions.
[0151] Amplification may also involve adding one or more control nucleic acids to one or more samples containing more than one nucleic acid. The control nucleic acid may contain a control marker.
[0152] Amplification may involve the use of one or more non-natural nucleotides. Non-natural nucleotides may include photostable and / or triggerable nucleotides. Examples of non-natural nucleotides include, but are not limited to, peptide nucleic acids (PNAs), morpholino and locked nucleic acids (LNAs), and glycol nucleic acids (GNAs) and threononucleotides (TNAs). Non-natural nucleotides may be added to one or more cycles of the amplification reaction. The addition of non-natural nucleotides can be used to identify the products at specific cycles or time points in the amplification reaction.
[0153] Performing one or more amplification reactions may involve using one or more primers. One or more primers may include one or more oligonucleotides. One or more oligonucleotides may contain at least about 7-9 nucleotides. One or more oligonucleotides may contain fewer than 12-15 nucleotides. One or more primers may anneal at least a portion of a labeled nucleic acid. One or more primers may anneal the 3' and / or 5' ends of a labeled nucleic acid. One or more primers may anneal the internal regions of a labeled nucleic acid. The internal region can be at least approximately 50, 100, 150, 200, 220, 230, 240, 250, 260, 270, 280, 290, 300, 310, 320, 330, 340, 350, 360, 370, 380, 390, or 400 from the 3' end of more than one labeled nucleic acid. 410, 420, 430, 440, 450, 460, 470, 480, 490, 500, 510, 520, 530, 540, 550, 560, 570, 580, 590, 600, 650, 700, 750, 800, 850, 900, or 1000 nucleotides. One or more primers may include a fixed set of primers. One or more primers may include at least one or more custom primers. One or more primers may include at least one or more control primers. One or more primers may include at least one or more housekeeping gene primers. One or more primers may include universal primers. Universal primers may be annealed to a universal primer binding site. One or more custom primers may be annealed to a first sample label, a second sample label, a molecular identifier marker, a nucleic acid, or a product thereof. One or more primers may include universal primers and custom primers. Custom primers may be designed to amplify one or more target nucleic acids. Target nucleic acids may include a subset of total nucleic acids in one or more samples. In some embodiments, the primers are probes attached to an array of the present disclosure.
[0154] In some implementations, barcoding more than one target in a sample (e.g., random barcoding) also includes generating an index library of barcoded targets (e.g., random barcoded targets) or barcoded fragments of targets. The barcode sequences of different barcodes (e.g., molecular markers of different random barcodes) can be different from each other. Generating an index library of barcoded targets involves generating more than one index polynucleotide from more than one target in the sample. For example, for an index library of barcoded targets including a first index target and a second index target, the labeled region of the first index polynucleotide and the labeled region of the second index polynucleotide can differ by less than, approximately less than, at least less than, or at most less than: 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, or any number or range of nucleotides between any two of these values. In some embodiments, generating an index library for barcoded targets includes contacting more than one target (e.g., an mRNA molecule) with more than one oligonucleotide comprising a multiple (T) region and a labeled region; and performing first-strand synthesis using reverse transcriptase to generate single-stranded labeled cDNA molecules (each comprising a cDNA region and a labeled region), wherein the more than one target comprises at least two different sequences of mRNA molecules, and the more than one oligonucleotide comprises at least two different sequences of oligonucleotides. Generating an index library for barcoded targets may also include amplifying single-stranded labeled cDNA molecules to generate double-stranded labeled cDNA molecules; and performing nested PCR on the double-stranded labeled cDNA molecules to generate labeled amplicons. In some embodiments, the method may include generating amplicons with adaptor labels.
[0155] Barcoding (e.g., random barcoding) can include using nucleic acid barcodes or tags to label individual nucleic acid (e.g., DNA or RNA) molecules. In some embodiments, it includes adding a DNA barcode or tag to the cDNA molecule when generating the cDNA molecule from mRNA. Nested PCR can be performed to minimize PCR amplification bias. Adaptors can be added for use in sequencing (e.g., next-generation sequencing (NGS)). Figure 2 At box 232, sequencing results can be used to determine the sequence of one or more copies of cellular markers, molecular markers, and nucleotide fragments of the target.
[0156] Figure 3This is a schematic diagram illustrating a non-limiting exemplary process for generating an index library of barcoded targets (e.g., random barcoded targets), such as an index library of barcoded mRNA or fragments thereof. As shown in step 1, the reverse transcription process can encode each mRNA molecule with unique molecular marker sequences, cellular marker sequences, and universal PCR sites. Specifically, RNA molecule 302 can be reverse transcribed to produce labeled cDNA molecule 304 (including cDNA region 306) by hybridizing a set of barcodes (e.g., random barcodes) 310 with a multi(A)-tailed region 308 of RNA molecule 302 (e.g., random hybridization). Each of the barcodes 310 may include a target-binding region, such as a multi(dT) region 312, a labeled region 314 (e.g., a barcode sequence or molecule), and a universal PCR region 316.
[0157] In some embodiments, the cell marker sequence may contain 3 to 20 nucleotides. In some embodiments, the molecular marker sequence may contain 3 to 20 nucleotides. In some embodiments, each of more than one random barcode further includes one or more of a universal marker and a cell marker, wherein the universal marker is identical for more than one random barcode on the solid support, and the cell marker is identical for more than one random barcode on the solid support. In some embodiments, the universal marker may contain 3 to 20 nucleotides. In some embodiments, the cell marker contains 3 to 20 nucleotides.
[0158] In some embodiments, the marker region 314 may include a barcode sequence or molecular marker 318 and a cell marker 320. In some embodiments, the marker region 314 may include one or more of a universal marker, a dimensional marker, and a cell marker. The length of the barcode sequence or molecular marker 318 may be the following, may be about the following, may be at least the following, or may be at most the following: 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, or any number or range of nucleotides between any two of these values. The length of cell marker 320 can be, approximately, at least, or at most, the following: 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, or any number or range of nucleotides between any two of these values. The length of universal marker can be, approximately, at least, or at most, the following: 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, or any number or range of nucleotides between any two of these values. Universal markers can be identical for more than one random barcode on a solid support, and cell markers can be identical for more than one random barcode on a solid support. The length of the dimension marker can be the following, can be about the following, can be at least the following, or can be at most the following: 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, or any number or range of nucleotides between any two of these values.
[0159] In some implementations, the marker region 314 may include, may include about, may include at least, or may include at most: 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, or any number or range between any two of these values, such as barcode sequences or molecular markers 318 and cell markers 320. The length of each marker can be, approximately, at least, or at most, the following: 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100 nucleotides or a number or range of nucleotides between any two of these values. A set of barcodes or random barcodes 310 can contain, approximately, at least, or at most, the following: 10, 20, 40, 50, 70, 80, 90, 10 2 10 species 3 10 species 4 10 species 5 10 species 6 10 species 7 10 species 8 10 species 9 10 species 10 10 species 11 10 species 12 10 species 13 10 species 14 10 species 15 10 species 20 A barcode or random barcode 310 representing a number or range between any two of these values. And groups of barcodes or random barcodes 310 may, for example, each contain a unique labeled region 314. The labeled cDNA molecule 304 can be purified to remove excess barcodes or random barcodes 310. Purification may include Ampure bead purification.
[0160] As shown in step 2, the product from the reverse transcription process in step 1 can be pooled into a single tube and PCR amplified using a first PCR primer pool and a first universal PCR primer. Pooling is possible because of the uniquely labeled region 314. Specifically, labeled cDNA molecules 304 can be amplified to generate nested PCR-labeled amplicons 322. Amplification can include multiplex PCR amplification. Amplification can include multiplex PCR amplification with 96 multiplex primers in a single reaction volume. In some embodiments, multiplex PCR amplification in a single reaction volume can utilize the following, about the following, at least the following, or at most the following: 10, 20, 40, 50, 70, 80, 90, 10 2 10 species 3 10 species 4 10 species 5 10 species 6 10 species 7 10 species 8 10 species 9 10 species 10 10 species 11 10 species 12 10 species 13 10 species 14 10 species 15 10 species 20 Multiplex primers can be used, or multiplex primers can be a number or range between any two of these values. Amplification may include a first PCR primer pool 324 comprising custom primers 326A-C targeting a specific gene and a universal primer 328. Custom primer 326 can hybridize with a region within the cDNA portion 306' of the labeled cDNA molecule 304. Universal primer 328 can hybridize with the universal PCR region 316 of the labeled cDNA molecule 304.
[0161] like Figure 3As shown in step 3, the product from the PCR amplification in step 2 can be amplified using a nested PCR primer pool and a second universal PCR primer. Nested PCR minimizes PCR amplification bias. In particular, the amplicon 322 of the nested PCR marker can be further amplified by nested PCR. Nested PCR can include multiplex PCR performed in a single reaction volume using a nested PCR primer pool 330 of nested PCR primers 332a-c and a second universal PCR primer 328'. Nested PCR primer pool 328 may contain, may contain about, may contain at least, or may contain at most: 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, or any number or range between these values for different nested PCR primers 330. Nested PCR primer 332 may contain adaptor 334 and hybridize with the region within the cDNA portion 306'' of the labeled amplicon 322. Universal primer 328' may contain adaptor 336 and hybridize with the universal PCR region 316 of the labeled amplicon 322. Thus, step 3 generates an adaptor-tagged amplicon 338. In some embodiments, nested PCR primer 332 and the second universal PCR primer 328' may not contain adaptors 334 and 336. Instead, adaptors 334 and 336 may be ligated to the nested PCR product to generate an adaptor-tagged amplicon 338.
[0162] As shown in step 4, the PCR product from step 3 can be amplified by PCR using library amplification primers for sequencing. Specifically, adaptor 334 and adaptor 336 can be used to perform one or more additional assays on the adaptor-tagged amplicon 338. Adaptor 334 and adaptor 336 can hybridize with primers 340 and 342. One or more primers 340 and 342 can be PCR amplification primers. One or more primers 340 and 342 can be sequencing primers. One or more adaptors 334 and adaptor 336 can be used for further amplification of the adaptor-tagged amplicon 338. One or more adaptors 334 and adaptor 336 can be used for sequencing the adaptor-tagged amplicon 338. Primer 342 may contain a plate index 344, allowing amplicon generated using the same set of barcodes or random barcodes 310 to be sequenced in a single sequencing reaction using next-generation sequencing (NGS).
[0163] In Situ RNA Probing mRNA from formalin-fixed cells is cross-linked, and these molecules are not easily released and captured by barcoded particles (e.g., Rhapsody beads). To address this prominent issue, some embodiments provided herein provide RNA probes that can bind to the mRNA, and in situ reverse transcription can produce short cDNA linked to the probe. These probes can be readily captured by barcoded particles (e.g., Rhapsody beads), and the use of specific targeting primers can enhance probe specificity and eliminate potential probe nonspecificity. Targeting probe sets and targeting primer sets, as well as in situ probe hybridization and RT kits, are provided herein. In some embodiments, methods and compositions are provided for RNA detection and amplification on a single-cell analysis system (e.g., Rhapsody) to expand sample types to fixed (e.g., formalin-fixed) cells. The methods and compositions disclosed herein enable single-cell transcriptome profiling of single cells from fixed samples via RNA detection assays and a single-cell analysis system (e.g., Rhapsody). In some embodiments, current single-cell analysis workflows (e.g., Rhapsody) cannot provide solutions for gene expression profiling from cells immobilized with formalin because the mRNA is cross-linked and cannot be captured by conventional mRNA capture methods (e.g., via multi-A / dT capture). In some embodiments, and not bound by any particular theory, small DNA probes can enter immobilized cells and bind to target mRNA in methods such as, for example, in situ hybridization or FISH to visualize the mRNA, despite the difficulty in releasing mRNA from immobilized cells. However, the number of probes currently available for mRNA detection is limited by the number of fluorescence detectors, and therefore it is difficult to provide transcriptome solutions using these methods. In some embodiments of the compositions and methods provided herein, mRNA sequence-specific probes can be used to bind target mRNA and can be extended from the mRNA by in situ reverse transcription to obtain the sequence for another layer of specificity. Furthermore, in some embodiments disclosed herein, a high number of target sets can be designed without limiting the detection methods. After probe / in situ RT, cells can be washed to remove unbound probes and loaded onto a single-cell analysis platform (e.g., Rhapsody). The cDNA-attached probe can be released from a single cell and can be captured by oligonucleotide barcoding associated with barcoded particles (e.g., Rhapsody beads) via a capture sequence added to the probe using TSO capture oligonucleotides. The user can then amplify the target gene, along with cell markers and UMIs, from the beads using a targeted primer set for single-cell targeted mRNA profiling analysis via sequencing. The disclosed compositions and methods open the use of archived formalin-fixed cells for single-cell RNA-seq analysis.Additionally, in some embodiments, methods and compositions are provided for spatial gene expression studies on FFPE-fixed tissue sections. Samples that can be used in current single-cell analysis systems (e.g., Rhapsody) are limited to live cells, freshly isolated cell nuclei, or short-term stored cells. The disclosed compositions and methods enable the use of long-term stored formalin-fixed samples in single-cell analysis systems (e.g., Rhapsody).
[0164] Figures 4A-4DA non-limiting exemplary schematic workflow for gene expression analysis of fixed cells is described. The workflow may include contacting a sample (e.g., a fixed sample containing fixed cells) with more than one probe oligonucleotide (step 400a). Each probe oligonucleotide may contain a coupling sequence and a probe sequence configured to hybridize with a nucleic acid target within the sample. The 5' end of each probe oligonucleotide may be phosphorylated. The probe oligonucleotide may be able to enter the cells and / or nuclei of the sample (e.g., permeabilized cells and / or permeabilized cell nuclei of the sample). The workflow may include removing one or more probe oligonucleotides that were not contacted with the sample after contacting the probe oligonucleotide with the sample. Removing one or more probe oligonucleotides that were not contacted with the sample may include removing one or more probe oligonucleotides that did not enter the cells of the sample. The workflow may include extending more than one probe oligonucleotide that hybridizes with a copy of the nucleic acid target to produce more than one extended probe oligonucleotide, each of which contains a sequence complementary to at least a portion of the nucleic acid target (step 400b). The workflow may include contacting the sample with an extension reagent. The extension reagent may include a reverse transcription reagent (e.g., reverse transcriptase and dNTPs). Extension may be performed in situ and may include in situ reverse transcription. In some embodiments, the cells of the sample remain intact during the extension step. The sample may include more than one cell, and the workflow may include dissociating the sample to produce more than one single cell. The workflow may include partitioning more than one single cell into more than one partition. The workflow may include contacting more than one extended probe oligonucleotide with barcoded particles (e.g., Rhapsody beads) (step 400c). The barcoded particles (e.g., beads) may be associated with more than one oligonucleotide barcode, said barcode comprising one or more of a cell marker (CL), a molecular marker (UMI), a 5' first universal sequence, and a 3' TSO (e.g., a capture sequence). The workflow may include barcoding more than one extended probe oligonucleotide or its product using more than one oligonucleotide barcode to generate more than one barcoded probe oligonucleotide. Barcoding more than one extended probe oligonucleotide may include: providing a splice oligonucleotide (e.g., a coupled oligonucleotide) comprising a 5' complement of a coupled sequence and a 3' complement of a capture sequence; hybridizing the coupled sequence of the extended probe oligonucleotide with the 5' complement of the coupled sequence of the coupled oligonucleotide; hybridizing the 3' complement of the capture sequence of the coupled oligonucleotide with the capture sequence of an oligonucleotide barcode in more than one oligonucleotide barcode; and / or linking the extended probe oligonucleotide to the hybridized oligonucleotide barcode.The workflow may include amplifying more than one barcoded probe oligonucleotide using a first primer capable of hybridizing with a first universal sequence or its complement and an amplification primer capable of hybridizing with a nucleic acid target or its complement, thereby generating more than one amplified barcoded probe oligonucleotide (step 400d). The amplification primer may contain a second universal sequence (e.g., an R2 primer sequence) and / or the first primer may contain a third universal sequence. The workflow may include obtaining sequencing data containing more than one sequencing read of the amplified barcoded probe oligonucleotide or its product (step 400e). Obtaining the sequencing data may include attaching binding sites of sequencing primers and / or sequencing adaptors to more than one barcoded probe oligonucleotide or its product. The workflow may include determining the copy number of the nucleic acid target in the sample based on the number of molecular markers associated with more than one amplified barcoded probe oligonucleotide or its product. In some embodiments, the probe oligonucleotide contains a predetermined spatial marker, and the workflow includes determining the spatial location and copy number of the nucleic acid target in the sample.
[0165] Some implementations provide methods for labeling nucleic acid targets in a sample. In some implementations, the method includes: contacting a sample containing a copy of the nucleic acid target with more than one probe oligonucleotide, wherein each probe oligonucleotide contains a coupling sequence and a probe sequence configured to hybridize with the nucleic acid target. The method may include: extending more than one probe oligonucleotide hybridized with a copy of the nucleic acid target to produce more than one extended probe oligonucleotide, each of the more than one extended probe oligonucleotide containing a sequence complementary to at least a portion of the nucleic acid target. The method may include: barcoding more than one extended probe oligonucleotide or its product using more than one oligonucleotide barcoding to generate more than one barcoded probe oligonucleotide, wherein each oligonucleotide barcode in the more than one oligonucleotide barcode contains a molecular marker, and wherein each of the more than one barcoded probe oligonucleotide contains a molecular marker, a probe sequence, and a sequence complementary to at least a portion of the nucleic acid target. The method may include: obtaining sequencing data containing more than one sequencing read of the barcoded probe oligonucleotide or its product, wherein each of the more than one sequencing read contains a molecular marker sequence and a subsequence of the nucleic acid target. This method may include determining the copy number of a nucleic acid target in a sample based on the number of molecular markers associated with more than one barcoded probe oligonucleotide or its product.
[0166] Some implementations provide methods for determining the copy number of a nucleic acid target in a sample. In some implementations, the method includes: contacting a sample containing copies of the nucleic acid target with more than one probe oligonucleotide, wherein each probe oligonucleotide contains a coupling sequence and a probe sequence configured to hybridize with the nucleic acid target. The method may include: extending more than one probe oligonucleotide hybridized with copies of the nucleic acid target to produce more than one extended probe oligonucleotide, each of the more than one extended probe oligonucleotide containing a sequence complementary to at least a portion of the nucleic acid target. The method may include: barcoding more than one extended probe oligonucleotide or its product using more than one oligonucleotide barcoding to generate more than one barcoded probe oligonucleotide, wherein each oligonucleotide barcode in the more than one oligonucleotide barcode contains a molecular marker, and wherein each of the more than one barcoded probe oligonucleotide contains a molecular marker, a probe sequence, and a sequence complementary to at least a portion of the nucleic acid target. The method may include: obtaining sequencing data containing more than one sequencing read of the barcoded probe oligonucleotide or its product, wherein each of the more than one sequencing read contains a molecular marker sequence and a subsequence of the nucleic acid target. This method may include determining the copy number of a nucleic acid target in a sample based on the number of molecular markers associated with more than one barcoded probe oligonucleotide or its product.
[0167] Some implementations provide methods for determining the spatial location and copy number of a nucleic acid target in a sample. In some implementations, the method includes contacting each of two or more spatial locations of a sample containing a copy of the nucleic acid target with more than one probe oligonucleotide, wherein each probe oligonucleotide comprises a coupling sequence, a probe sequence configured to hybridize with the nucleic acid target, and a predetermined spatial marker. In some implementations, probe oligonucleotides contacting the same spatial location contain the same spatial marker sequence, and probe oligonucleotides contacting different spatial locations of the sample contain different spatial marker sequences. The method may include extending more than one probe oligonucleotide hybridizing with a copy of the nucleic acid target to produce more than one extended probe oligonucleotide, each of the more than one extended probe oligonucleotides comprising a sequence complementary to at least a portion of the nucleic acid target. This method may include: barcoding more than one extended probe oligonucleotide or its product using more than one oligonucleotide barcode to generate more than one barcoded probe oligonucleotide, wherein each oligonucleotide barcode in the more than one oligonucleotide barcode contains a molecular marker, and wherein each of the more than one barcoded probe oligonucleotide contains a molecular marker, a probe sequence, and a sequence complementary to at least a portion of a nucleic acid target. This method may include: obtaining sequencing data containing more than one sequencing read containing the barcoded probe oligonucleotide or its product, wherein each of the more than one sequencing read contains a spatial marker sequence, a molecular marker sequence, and a subsequence of the nucleic acid target. This method may include: for each unique spatial marker sequence associated with a different spatial location in the sample, counting the number of molecular markers having different sequences associated with the nucleic acid target to determine the copy number of the nucleic acid target at each spatial location in the sample.
[0168] The method may include: contacting a sample with an extension reagent. At least a portion of the contact step may be performed in the presence of the extension reagent. The entire contact step may be performed in the presence of the extension reagent. The contact step and the extension step may be performed simultaneously. Extension may be performed in situ. The extension may include in situ reverse transcription. In some embodiments, the cells of the sample remain intact during the extension step. The extension reagent may include a reverse transcription reagent. The reverse transcription reagent may include reverse transcriptase and dNTPs. The reverse transcriptase may include a viral reverse transcriptase. The viral reverse transcriptase may be murine leukemia virus (MLV) reverse transcriptase or Moloney murine leukemia virus (MMLV) reverse transcriptase.
[0169] Barcoding more than one extended probe oligonucleotide or its product may include: providing a coupled oligonucleotide comprising a 5' complement of a coupling sequence and a 3' complement of a capture sequence; hybridizing the coupling sequence of the extended probe oligonucleotide with the 5' complement of the coupling sequence of the coupled oligonucleotide; hybridizing the 3' complement of the capture sequence of the coupled oligonucleotide with the capture sequence of an oligonucleotide barcode in more than one oligonucleotide barcode; and / or ligating the extended probe oligonucleotide to the hybridized oligonucleotide barcode. The method may include: filling the gap between the extended probe oligonucleotide and the hybridized oligonucleotide barcode with a DNA polymerase lacking at least one of 5' to 3' exonuclease activity and 3' to 5' exonuclease activity before ligating the extended probe oligonucleotide to the oligonucleotide barcode. Ligating the extended probe oligonucleotide to the hybridized oligonucleotide barcode may be performed using a DNA ligase. The coupled oligonucleotide may be a single-stranded oligonucleotide, a double-stranded oligonucleotide, or a mixture thereof. The coupled oligonucleotide may contain non-natural nucleotides. Coupled oligonucleotides can contain at least about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 5 Nucleotides of 3, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, or any number or range between any two of these values.The coupled sequences can contain at least approximately 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, and 53. The number of nucleotides can be 1, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, or any number or range between these values. The 5' end of each probe oligonucleotide may be phosphorylated. The probe oligonucleotide may be able to enter the cells and / or nucleus of the sample (e.g., permeabilized cells and / or permeabilized cell nuclei of the sample). The method may include: after contacting the probe oligonucleotide with the sample, removing one or more probe oligonucleotides from more than one probe oligonucleotide that have not been contacted with the sample, optionally, removing one or more probe oligonucleotides that have not been contacted with the sample includes: removing one or more probe oligonucleotides from cells that have not entered the sample.
[0170] The contact step may include contacting the sample with a device (e.g., an inkjet device) configured to deposit probe oligonucleotides. The device may be a needle, needle array, tube, aspiration device, injection device, electroporation device, fluorescence-activated cell sorting device, inkjet device, microfluidic device, or any combination thereof. In some embodiments, the device contacts different spatial locations of the sample at a specified rate. The length of the spatial markers can be at least approximately 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 5 Nucleotides of 3, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, or any number or range between any two of these values. The two or more spatial locations may include at least about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 12, about 14, about 16, about 18, about 20, about 30, about 40, about 50, about 60, about 70, about 80, about 90, or about 100 different spatial locations of the sample. In some embodiments, the spatial locations of the sample correspond to regions containing no more than about 50 cells, about 45 cells, about 40 cells, about 35 cells, about 30 cells, about 25 cells, about 20 cells, about 15 cells, about 10 cells, about 9 cells, about 8 cells, about 7 cells, about 6 cells, about 5 cells, about 4 cells, about 3 cells, about 2 cells, about 1 cell, or a number or range between any two of these values.
[0171] Each oligonucleotide barcode in more than one oligonucleotide barcode may contain a first universal sequence. In some embodiments, obtaining sequencing data includes: amplifying more than one barcoded probe oligonucleotide using a first primer capable of hybridizing with the first universal sequence or its complement and an amplification primer capable of hybridizing with a nucleic acid target or its complement, thereby producing more than one amplified barcoded probe oligonucleotide. Obtaining sequencing data may include obtaining sequencing data containing more than one sequencing read comprising the amplified barcoded probe oligonucleotide or its product. Obtaining sequencing data may include attaching the binding sites of sequencing primers and / or sequencing adaptors to more than one barcoded probe oligonucleotide or its product. The amplification primers may contain a second universal sequence and / or the first primer may contain a third universal sequence. The first universal sequence, the second universal sequence, and / or the third universal sequence may be the same. The first universal sequence, the second universal sequence, and / or the third universal sequence may be different. The first universal sequence, the second universal sequence, and / or the third universal sequence may contain the binding site of the sequencing primer and / or the sequencing adaptor, its complement, and / or a portion thereof. Sequencing adaptors may include the P5 sequence, the P7 sequence, their complementary sequences, and / or portions thereof. Sequencing primers may include read 1 sequencing primers, read 2 sequencing primers, their complementary sequences, and / or portions thereof.
[0172] A sample may contain more than one nucleic acid target, such as, for example, a group of different nucleic acid targets, including at least about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 12, about 14, about 16, about 18, about 20, about 30, about 40, about 50, about 60, about 70, about 80, about 90, about 100, about 125, about 150, about 175, about 200, about 225, about 250, about 275, about 300, about 325, about 350, about 375, about 400, about 425, about 450, about 475, about 500, or any two of these values. Two or more nucleic acid targets in a target group can be biomarkers. Biomarkers can be biomarkers of diseases or conditions. Diseases or conditions can be cancer, infection, viral infection, inflammatory disease, neurodegenerative disease, fungal disease, bacterial infection, or any combination thereof. The contact step can include contacting the sample with a set of probe oligonucleotides comprising two or more probe oligonucleotides, wherein each of the multiple probe oligonucleotides comprises a probe sequence configured to hybridize with nucleic acid targets in more than one nucleic acid target. Determining the copy number of the nucleic acid targets in the sample can include determining the copy number of each of the multiple nucleic acid targets in the sample based on the number of molecular markers having distinct sequences associated with more than one barcoded probe oligonucleotide or its product, said multiple barcoded probe oligonucleotide or its product comprising sequences of more than one nucleic acid target. The method can include: for each unique spatial marker sequence associated with a different spatial location in the sample, counting the number of molecular markers having distinct sequences associated with each of the multiple nucleic acid targets to determine the copy number of each of the multiple nucleic acid targets at each spatial location in the sample. Amplification primers may include a set of amplification primers configured to hybridize with more than one nucleic acid target or its complement, such as, for example, a set of different amplification primers of at least about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 12, about 14, about 16, about 18, about 20, about 30, about 40, about 50, about 60, about 70, about 80, about 90, about 100, about 125, about 150, about 175, about 200, about 225, about 250, about 275, about 300, about 325, about 350, about 375, about 400, about 425, about 450, about 475, about 500, or any number or range between any two of these values. As used herein, the term “panel” should be given its general meaning and should also refer to a group of nucleic acids designed to hybridize with a set of target nucleic acid sequences of interest or their products.For example, in some embodiments, expression analysis of the genes can be performed using a set of 200 different probe oligonucleotides designed to bind to 200 transcripts of 200 genes (e.g., a set of probe oligonucleotides containing more than one probe oligonucleotide). In some such embodiments, after in situ extension of these probe oligonucleotides, the extension products (extended probe oligonucleotides) can be barcoded as described herein, and then amplified using a set of 200 different amplification primers designed to bind sequences of 200 nucleic acid targets (or their complements). Nucleic acid targets may include nucleic acid molecules (e.g., ribonucleic acid (RNA), messenger RNA (mRNA), microRNA, small interfering RNA (siRNA), RNA degradation products, RNA containing multiple (A) tails, sample index oligonucleotides, cell component binding reagent-specific oligonucleotides, or any combination thereof).
[0173] Each molecular marker of more than one oligonucleotide barcode may contain at least 6 nucleotides. Each capture sequence of more than one oligonucleotide barcode may contain at least 4 nucleotides. More than one oligonucleotide barcode may be associated with a solid support, and partitions in more than one partition may contain a single solid support. Each of the more than one oligonucleotide barcodes may contain a cell marker. Each cell marker of more than one oligonucleotide barcode may contain at least 6 nucleotides. Oligonucleotide barcodes associated with the same solid support in more than one oligonucleotide barcode may contain the same cell marker. Oligonucleotide barcodes associated with different solid supports in more than one oligonucleotide barcode may contain different cell markers. The solid support may include synthetic particles, a flat surface, or a combination thereof. The method may include associating synthetic particles containing more than one oligonucleotide barcode with cells in a partition. The method may include lysing cells after associating synthetic particles with cells. Lysing cells may include heating cells, contacting cells with a detergent, altering the pH of cells, or any combination thereof. Synthetic particles and single cells can be located in the same partition, which can be a pore or a droplet. At least one oligonucleotide barcode of more than one type can be immobilized or partially immobilized on the synthetic particle, and / or at least one oligonucleotide barcode of more than one type can be encapsulated or partially encapsulated within the synthetic particle. The synthetic particle can be destructible (e.g., a destructible hydrogel particle). The synthetic particle can include beads. Beads can include agarose gel beads, streptavidin beads, agarose beads, magnetic beads, conjugated beads, protein A conjugated beads, protein G conjugated beads, protein A / G conjugated beads, protein L conjugated beads, oligo(dT) conjugated beads, silica beads, silica-like beads, avidin microbeads, anti-fluorescent dye microbeads, or any combination thereof. The synthetic particles may contain materials selected from the group consisting of: polydimethylsiloxane (PDMS), polystyrene, glass, polypropylene, agarose, gelatin, hydrogel, paramagnetic material, ceramic, plastic, methylstyrene, acrylic polymer, titanium, latex, agarose gel, cellulose, nylon, silicone, and any combination thereof. Each oligonucleotide barcode in more than one oligonucleotide barcode may contain a linker functional group. The synthetic particles may contain a solid support functional group. The support functional group and the linker functional group may be associated with each other, and the linker functional group and the support functional group may individually be selected from the group consisting of: C6, biotin, streptavidin, one or more primary amines, one or more aldehydes, one or more ketones, and any combination thereof.
[0174] Sample Analysis During the contact step, the sample may be physically fragmented or may be intact. The sample may include single cells. The sample may include more than one single cell. The sample may include more than one cell, and the method may include: dissociating the sample to generate more than one single cell. The dissociation may include chemical dissociation, enzymatic dissociation, and / or mechanical dissociation. The dissociation may employ one or more of collagenase, chymotrypsin, dispersase, elastase, hyaluronidase, trypsin, papain, and trypsin. The method may include: prior to the barcoding step: partitioning more than one single cell into at least one partition, wherein the partitions in the more than one partition contain single cells from more than one single cell; and in the partition containing the single cell, contacting an extended probe oligonucleotide with more than one oligonucleotide barcode. In some embodiments, the method includes: in the partition containing the single cell, contacting the single cell with a lysis buffer at 15°C–65°C to lyse the single cell. The lysis buffer may contain an agent capable of dissociating protein-nucleic acid complexes.
[0175] More than one cell may include one or more cell types. The one or more cell types may be selected from the group consisting of: brain cells, heart cells, cancer cells, circulating tumor cells, organ cells, epithelial cells, metastatic cells, benign cells, primary cells, and circulating cells, or any combination thereof. The sample may include a biological sample, a clinical sample, an environmental sample, a biological fluid, tissue, a tissue section derived from the subject, or any combination thereof. The subject may be a human, mouse, dog, rat, or vertebrate. The method may include: determining the subject's genotype, phenotype, or one or more gene mutations based on the spatial location of nucleic acid targets in the sample. The method may include: predicting the subject's susceptibility to one or more diseases, such as, for example, cancer or a hereditary disease. The method may include: identifying more than one cell type in the sample. In some embodiments, a drug may be selected based on the predicted reactivity of more than one cell type in the sample.
[0176] This method may include imaging the sample, optionally before and / or after the contact step, optionally generating imaging data. Imaging the sample may include staining the sample with a staining agent, which may be a fluorescent staining agent, a negative staining agent, an antibody staining agent, or any combination thereof. Staining may include immunocytochemistry (ICC), immunohistochemistry (IHC), immunofluorescence (IF), or any combination thereof. In some embodiments, imaging may include microscopy, confocal microscopy, time-of-flight imaging microscopy, fluorescence microscopy, multiphoton microscopy, quantitative phase microscopy, surface-enhanced Raman spectroscopy, photography, manual visual analysis, automated visual analysis, or any combination thereof. This method may include correlating imaging data and sequencing data at one or more spatial locations of the sample. This method may include correlation analysis of the spatial location imaging data and sequencing data. The correlation analysis may identify one or more of the following: candidate biomarkers, candidate therapeutic agents, candidate doses of therapeutic agents, and / or cellular targets of candidate therapeutic agents. In some embodiments, the imaging produces an image for constructing a physical representation of the sample. In some implementations, the atlas may be two-dimensional or three-dimensional. The method may include mapping nucleic acid targets and / or cellular component targets onto the sample atlas. The method may also include mapping one or more single cells from more than one cell onto the sample atlas. Sequencing reads derived from the same single cell from more than one cell may contain the same cell markers. These sequence reads may also contain sequences of spatial markers. Users can associate single cells of a sample with the spatial location of the sample based on the association between cell markers and spatial markers.
[0177] Cellular Component Target Profiling The sample may contain more than one cell component target, and the method may include: contacting the sample with more than one cell component binding agent, wherein each of the more than one cell component binding agent contains a cell component binding agent-specific oligonucleotide, the cell component binding agent-specific oligonucleotide containing a unique identifier sequence of the cell component binding agent, and wherein the cell component binding agent is capable of specifically binding to at least one of the more than one cell component target; barcoding the cell component binding agent-specific oligonucleotide to generate more than one barcoded cell component binding agent-specific oligonucleotide, each of the more than one barcoded cell component binding agent-specific oligonucleotide containing a sequence complementary to at least a portion of the unique identifier sequence and a molecular marker sequence; and obtaining sequencing data comprising more than one sequencing read containing more than one barcoded cell component binding agent-specific oligonucleotide or its product, wherein each of the more than one sequencing read contains at least a portion of the molecular marker sequence and the unique identifier sequence. Obtaining the sequencing data may include attaching the binding sites of sequencing primers and / or sequencing adaptors to the barcoded cell component binding agent-specific oligonucleotide or its product.
[0178] The method may include: after contacting a sample with more than one cell component binding agent, removing one or more cell component binding agents from the more than one cell component binding agent that have not been contacted with the sample; optionally, removing one or more cell component binding agents that have not been contacted with the sample includes: removing one or more cell component binding agents that have not been contacted with at least one corresponding cell component target. Cell component targets may include intracellular proteins, carbohydrates, lipids, proteins, extracellular proteins, cell surface proteins, cell markers, B cell receptors, T cell receptors, major histocompatibility complex, tumor antigens, receptors, intracellular proteins, or any combination thereof. Cell component binding agent-specific oligonucleotides may contain a second molecular marker, optionally at least 10 of the more than one cell component binding agent-specific oligonucleotides containing different second molecular marker sequences. The second molecular marker sequences of at least two cell component binding agent-specific oligonucleotides may be different, and the unique identifier sequences of at least two cell component binding agent-specific oligonucleotides may be the same. In some embodiments, the number of unique molecular marker sequences associated with a unique identifier sequence for a cell component binding agent in the sequencing data indicates the copy number of at least one cell component target in the sample, the cell component binding agent being capable of specifically binding to at least one cell component target. In some embodiments, the number of unique second molecular marker sequences associated with a unique identifier sequence for a cell component binding agent in the sequencing data indicates the copy number of at least one cell component target in the sample, the cell component binding agent being capable of specifically binding to at least one cell component target.
[0179] In some embodiments, the cell component binding reagent-specific oligonucleotides are barcoded using more than one oligonucleotide barcode, the same one used for barcoding the extended probe oligonucleotide. In some embodiments, the solid support contains two or more oligonucleotide barcodes, each containing a different 3' target binding region or capture sequence. For example, in some embodiments described herein, the extended probe oligonucleotide is barcoded using a first more than one oligonucleotide barcode having a 3' capture sequence configured to hybridize with a coupled oligonucleotide (e.g., a TSO decoy), and the cell component binding reagent-specific oligonucleotide is barcoded using a second more than one oligonucleotide barcode having a 3' multiple (dT) sequence. In some such embodiments, the cell component binding reagent-specific oligonucleotide contains a multiple (dA) sequence.
[0180] In some implementations, methods are provided for determining the spatial location and copy number of cellular component targets in a sample. Sequencing reads of more than one barcoded cellular component-binding reagent-specific oligonucleotide or its product may each contain a cellular marker sequence. As described above, a user can associate single cells of a sample with the spatial location of the sample based on the association between cellular markers and spatial markers. Therefore, a user can thus determine the spatial location and copy number of cellular component targets in a sample based on the spatial markers (and thus the spatial locations) associated with the cellular markers in the sequencing data.
[0181] Embodiments using cell component binding agents (e.g., protein binding agents) associated with oligonucleotides (e.g., oligonucleotide-conjugated antibodies (AbO) and oligonucleotide-conjugated aptamers) for barcoding and / or for determining protein expression profiles in single cells and for sample tracking (e.g., tracing sample origin) have been described in the following: US2018 / 0088112 and US2018 / 0346970; and WO / 2020 / 037065; the contents of each of these applications are incorporated herein by reference in their entirety. In some embodiments, the systems, methods, compositions, and kits provided herein may be used in conjunction with the systems, methods, compositions, and kits described in PCT application publication WO / 2021 / 163374, the contents of which are incorporated herein by reference in their entirety. In some embodiments, the systems, methods, compositions, and kits provided herein may be used in conjunction with the systems, methods, compositions, and kits described in PCT Application Publications WO / 2024 / 097719 and WO / 2024 / 097718, the contents of each of which are incorporated herein by reference in their entirety.
[0182] Fixing Agents, Unfixing Agents, and Permeabilizing Agents The sample may include tissue, cell monolayer, fixed cells, tissue sections, or any combination thereof. The sample may include fresh tissue sections, frozen tissue sections, fixed tissue sections, formalin-fixed tissue sections, formalin-fixed paraffin-embedded (FFPE) tissue sections, acetone-fixed tissue sections, paraformaldehyde (PFA)-fixed tissue sections, and / or methanol-fixed tissue sections. The sample may include cell nuclear suspensions, such as, for example, fixed cell nuclear suspensions and / or permeabilized cell nuclear suspensions. In some embodiments, the sample has been contacted with one or more fixatives and / or permeabilizers. The sample may include cells, such as, for example, fresh cells, frozen cells, fixed cells, formalin-fixed cells, formalin-fixed paraffin-embedded (FFPE) cells, acetone-fixed cells, paraformaldehyde (PFA)-fixed cells, and / or methanol-fixed cells. The method may include: permeabilizing the sample and / or fixing the sample. Fixing the sample may include contacting the sample with a fixative. The fixative may include a non-crosslinking fixative (e.g., methanol). The fixative may include a crosslinking agent. The crosslinking agent may include a cleavable crosslinking agent. Degradable crosslinking agents may include or be derived from dithiobis(succinimide propionate) (DSP), disuccinimide tartrate (DST), bis[2-(succinimideoxycarbonyloxy)ethyl] sulfone (BSOCOES), ethylene glycol bis(succinimide succinate) (EGS), dimethyl 3,3'-dithiobispropionylimine ester (DTBP), and succinimide 3-(2-pyridyldithio)propionate (SPDP). Succinimidyl 6-(3(2-pyridyldithio)propionamido)hexanoate (LC-SPDP), 4-succinimidyloxycarbonyl-α-methyl-α-(2-pyridyldithio)toluene (SMPT), 3-(2-pyridyldithio)propionylhydrazine (PDPH), succinimidyl 2-((4,4'-azidopentamido)ethyl)-1,3'-dithiopropionate (SDAD, NHS-SS-diazadiazide), or any combination thereof. The cleavable crosslinker may include cleavable links selected from the group consisting of: chemically cleavable links, photocleavable links, acid-labile linkers, heat-sensitive linkers, enzyme-cleavable links, and combinations thereof. The cleavable crosslinker may be a thiol-cleavable crosslinker or may contain a disulfide linker. Fixatives may include paraformaldehyde (PFA), dithiobis(succinimide propionate) (DSP), succinimide 3-(2-pyridyl dithio)propionate (SPDP), CellCover, or combinations thereof. Fixation and permeation of samples can be performed simultaneously.
[0183] Sample fixation and permeabilization can be performed in the presence of a dual-function agent capable of both fixing and permeabilizing the sample. The dual-function agent may be methanol. Sample permeabilization may include contacting the sample with a permeabilizing agent. This method may include removing the permeabilizing agent from the sample after contacting it with more than one probe oligonucleotide or more than one cell component binding reagent. The permeabilizing agent may be capable of (i) permeating the cell membrane of a cell, and (ii) making the cell membrane of the cell permeable to the probe oligonucleotide or the cell component binding reagent, or both. The permeabilizing agent may include (i) a solvent, detergent, or surfactant; (ii) BDCytoperm; (iii) a saponin or a derivative thereof; (iv) Triton X-100; (v) methanol or a derivative thereof; and / or (vi) digitalis saponins or a derivative thereof. Agents capable of dissociating protein-nucleic acid complexes may include broad-spectrum serine proteases. Broad-spectrum serine proteases may be proteinase K. The lysis buffer may contain a defixing agent. The defixing agent may include thiols, hydroxylamine, periodate, bases, or any combination thereof. The lysis buffer may contain DTT. The method may include: reversing the fixation of a sample and / or a single cell. Reversing the fixation of a sample and / or a single cell may include UV photolysis, chemical treatment, heating, enzymatic treatment, or any combination thereof.
[0184] Blocking Reagents, Decoy Oligonucleotides, and Blocking Oligonucleotides This method may include contacting the sample with a blocking agent, one or more decoy oligonucleotides, and / or one or more blocking oligonucleotides before contacting the sample with more than one cell component binding agent and / or contacting the sample with more than one probe oligonucleotide. The methods and compositions provided herein can be used in conjunction with the methods and compositions described in PCT patent application No. PCT / US22 / 75661, filed August 30, 2022, entitled “RNA PRESERVATION AND RECOVERY FROM FIXED CELLS,” the entire contents of which are incorporated herein by reference. The methods and compositions provided herein can be used in conjunction with blocking agents, such as those described in PCT patent application No. PCT / US22 / 75656, filed August 30, 2022, entitled “USE OF DECOYPOLYNUCLEOTIDES IN SINGLE CELL MULTIOMICS,” the entire contents of which are incorporated herein by reference. Contacting the sample with more than one cell component binding agent can be performed in the presence of the blocking agent. The blocking agent may include more than one oligonucleotide that is complementary to at least a portion of a cell component-specific oligonucleotide. The blocking agent may include an antibody or fragment thereof derived from a first species, and may also include serum derived from the first species. The sample may contain one or more non-target nucleic acids, and the blocking agent may contain more than one bait oligonucleotide capable of hybridizing with at least one of the one or more non-target nucleic acids. Each of the more than one bait oligonucleotide may be capable of hybridizing with at least a portion of the non-target nucleic acid. In some embodiments, the decoy oligonucleotide: comprises a sequence complementary to at least a portion of a non-target nucleic acid; comprises a sequence identical or substantially similar to that of a cell component binding reagent-specific oligonucleotide, optionally having a length of 3 to 40 nucleotides; has at most 50% sequence identity with the cell component binding reagent-specific oligonucleotide; does not contain a UMI; comprises a random sequence, optionally having a length of about four, five, six, seven, eight, nine, ten, eleven, twelve, thirteen, fourteen, or fifteen nucleotides; does not contain any sequence having more than four, five, six, or seven consecutive T or A; contains at least one G or C in every four, five, six, or seven consecutive nucleotides; comprises one or more modified nucleotides; comprises a 5' modification, optionally comprising a 5' amino-modified C12 modification (5AmMC12); comprises a 3' modification, optionally comprising a 3' dideoxy-C modification (ddC); and / or has a length of 30 to 65 nucleotides.
[0185] The sample may contain one or more undesirable nucleic acid substances, and the method may include: contacting a blocking oligonucleotide with the sample, wherein the blocking oligonucleotide specifically binds to at least one of the one or more undesirable nucleic acid substances, and wherein reverse transcription of at least one of the one or more undesirable nucleic acid substances is reduced by the blocking oligonucleotide. In some embodiments, the blocking oligonucleotide is contacted with the sample: before contacting more than one probe oligonucleotide with the sample; after contacting more than one probe oligonucleotide with the sample; and / or when contacting more than one probe oligonucleotide with the sample. The method may include: providing a blocking oligonucleotide that specifically binds to two or more undesirable nucleic acid substances in the sample, optionally at least 10 or up to at least 100 undesirable nucleic acid substances. In some implementations, the blocking oligonucleotide is: locked nucleic acid (LNA), peptide nucleic acid (PNA), DNA, LNA / PNA chimera, LNA / DNA chimera, or PNA / DNA chimera; specifically binds to the 3' end of one or more undesirable nucleic acid substances within 100 nt, 50 nt, or 25 nt; specifically binds to the 5' end of one or more undesirable nucleic acid substances within 100 nt, or the blocking oligonucleotide specifically binds to the middle of one or more undesirable nucleic acid substances within 100 nt; contains or does not contain non-natural nucleotides; has a Tm of at least 50°C, at least 60°C, or at least 70°C; cannot be used as a primer for reverse transcriptase or polymerase; and / or is 8 nt to 100 nt long, 10 nt to 50 nt long, 12 nt to 21 nt long, 20 nt to 30 nt long, or about 25 nt long. One or more undesirable nucleic acid substances may constitute approximately 50%, 60%, 70%, or 80% of the nucleic acid content of the sample. Undesirable nucleic acid substances may be selected from the group consisting of ribosomal RNA, mitochondrial RNA, genomic DNA, intron sequences, high-abundance sequences, and combinations thereof. One or more undesirable nucleic acid substances may be mRNA molecules, and blocking oligonucleotides may specifically bind to the 3' multi(A) tail of one or more undesirable nucleic acid substances within 10 nt.
[0186] Kits In some embodiments, a composition (e.g., a kit) is provided. In some embodiments, the kit comprises: more than one probe oligonucleotide, wherein each probe oligonucleotide comprises a coupling sequence and a probe sequence configured to hybridize with a nucleic acid target, optionally, the probe oligonucleotide comprises a predetermined spatial marker; a coupling oligonucleotide comprising a 5' complement of the coupling sequence and a 3' complement of the capture sequence; more than one oligonucleotide barcode, wherein the 3' end of each of the more than one oligonucleotide barcodes is associated with a solid support, wherein the 5' end of each of the more than one oligonucleotide barcodes comprises a capture sequence; a first primer capable of hybridizing with a first universal sequence, optionally further comprising a third universal sequence; and a primer capable of hybridizing with a nucleic acid target or its complement. The amplification primers optionally further comprise a second universal sequence; a DNA ligase; an extension reagent, optionally a reverse transcription reagent, and also optionally reverse transcriptase and dNTPs; one or more immobilizers; one or more permeabilizers; a cross-linking agent; a deimmobilizing agent; a lysis buffer; more than one cell component binding agent, wherein each of the more than one cell component binding agent comprises a cell component binding agent-specific oligonucleotide, said cell component binding agent-specific oligonucleotide comprising a unique identifier sequence of the cell component binding agent, and wherein the cell component binding agent is capable of specifically binding at least one of more than one cell component target; a blocking agent; one or more decoy oligonucleotides; and / or one or more blocking oligonucleotides.
[0187] More than one probe oligonucleotide may comprise a probe oligonucleotide set comprising two or more types of more than one probe oligonucleotide, wherein more than one of each type comprises a probe sequence configured to hybridize with nucleic acid targets in more than one nucleic acid target, optionally at least about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 12, about 14, about 16, about 18, about 20, about 30, or about A group of different nucleic acid targets, including 40, approximately 50, approximately 60, approximately 70, approximately 80, approximately 90, approximately 100, approximately 125, approximately 150, approximately 175, approximately 200, approximately 225, approximately 250, approximately 275, approximately 300, approximately 325, approximately 350, approximately 375, approximately 400, approximately 425, approximately 450, approximately 475, approximately 500, or any two of these values.
[0188] Amplification primers may include a set of amplification primers configured to hybridize with more than one nucleic acid target or its complement, optionally at least about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 12, about 14, about 16, about 18, about 20, about 30, about 40, about 50, about 60, about 70, about 80, about 90, about 100, about 125, about 150, about 175, about 200, about 225, about 250, about 275, about 300, about 325, about 350, about 375, about 400, about 425, about 450, about 475, about 500, or any two of these values, or a range of different amplification primers.
[0189] Terminology In at least some of the previously described embodiments, one or more elements used in one embodiment may be used interchangeably in another embodiment, unless such substitution is technically impractical. Those skilled in the art will understand that various other omissions, additions, and modifications can be made to the methods and structures described above without departing from the scope of the claimed subject matter. All such modifications and changes are intended to fall within the scope of the subject matter defined by the appended claims.
[0190] Those skilled in the art will understand that the functions performed in this and other processes and methods disclosed herein can be implemented in different orders. Furthermore, the steps and operations outlined are provided by way of example only, and some of these steps and operations may be optional, combined into fewer steps and operations, or expanded into other steps and operations without departing from the essence of the disclosed embodiments.
[0191] Regarding the use of substantially any plural and / or singular terms herein, those skilled in the art can convert from plural to singular and / or from singular to plural where appropriate for the context and / or application. For clarity, various singular / plural arrangements may be explicitly set forth herein. As used in this specification and the appended claims, the singular forms “a,” “an,” and “the” include plural referents unless the context explicitly indicates otherwise. Unless otherwise stated, any reference to “or” herein is intended to cover “and / or.”
[0192] Those skilled in the art will understand that, in general, the terminology used herein, and especially in the appended claims (e.g., the body of the appended claims), is typically intended as “open-ended” terms (e.g., the term “including” should be interpreted as “including but not limited to”, the term “having” should be interpreted as “having at least”, the term “includes” should be interpreted as “including but not limited to”, etc.). Those skilled in the art will further understand that if a particular number of claims is intended to be presented, such intention will be explicitly stated in the claims, and if such a statement is absent, such intention does not exist. For example, to aid understanding, the appended claims may include the introductory phrases “at least one” and “one or more” to introduce the claims. However, the use of such wording should not be construed as meaning that introducing a claim statement with the indefinite article “a(a)” or “an” would limit any specific claim in a claim statement containing such an introduction to an embodiment containing only one such statement, even when the same claim includes the introductory wording “one or more” or “at least one” and indefinite articles such as “a(a)” or “an” (e.g., “a(a)” and / or “an” should be interpreted as meaning “at least one” or “one or more”); the same applies to the use of definite articles to introduce a claim statement. Furthermore, even if a specific number in an introductory claim statement is explicitly stated, those skilled in the art will recognize that such a statement should be interpreted as meaning at least the number stated (e.g., simply stating “two statements” without other modifiers means at least two statements or two or more statements). Furthermore, in cases where conventions such as "at least one of A, B, and C" are used, this syntactic structure is generally intended to be understood by a person skilled in the art (e.g., "a system having at least one of A, B, and C" will include, but is not limited to, systems having a single A, having a single B, having a single C, A and B together, A and C together, B and C together, and / or A, B, and C together, etc.). In cases where conventions such as "at least one of A, B, or C" are used, this syntactic structure is generally intended to be understood by a person skilled in the art (e.g., "a system having at least one of A, B, or C" will include, but is not limited to, systems having a single A, having a single B, having a single C, A and B together, A and C together, B and C together, and / or A, B, and C together, etc.).Those skilled in the art will further understand that, in practice, any separate words and / or wording presenting two or more alternative terms, whether in the specification, claims, or drawings, should be understood to account for the possibility of including one, any, or both terms. For example, the phrase "A or B" should be understood to include the possibility of including "A" or "B" or "A and B".
[0193] Furthermore, when features or aspects of this disclosure are described in terms of the Markush group, those skilled in the art will recognize that this disclosure is also described in terms of any individual member or subgroup of the Markush group.
[0194] As those skilled in the art will understand, for any and all purposes, such as providing a written description, all scopes disclosed herein also include any and all possible subscopes and combinations of subscopes. Any listed scope can be readily identified as sufficiently descriptive and such scopes can be decomposed into at least equal halves, thirds, quarters, fifths, tenths, etc. As a non-limiting example, each scope discussed herein can be readily decomposed into a lower third, middle third, and upper third, etc. As those skilled in the art will also understand, all language such as “up to,” “at least,” “greater than,” “less than,” etc., includes the stated numbers and refers to a scope that can subsequently be decomposed into subscopes as discussed above. Finally, as those skilled in the art will understand, a scope includes members of each individual. Thus, for example, a group having 1-3 items means a group having 1, 2, or 3 items. Similarly, a group having 1-5 items means a group having 1, 2, 3, 4, or 5 items, and so on.
[0195] From the foregoing, it should be understood that various embodiments of this disclosure have been described herein for illustrative purposes, and various modifications can be made without departing from the scope and spirit of this disclosure. Therefore, the various embodiments disclosed herein are not intended to be limiting, and the true scope and spirit are indicated by the following claims.
Claims
1. A method for labeling nucleic acid targets in a sample, the method comprising: A sample containing a copy of a nucleic acid target is contacted with more than one probe oligonucleotide, wherein each probe oligonucleotide contains a coupling sequence and a probe sequence configured to hybridize with the nucleic acid target. Extending the more than one probe oligonucleotide hybridized with the nucleic acid target copy to generate more than one extended probe oligonucleotide, each of the more than one extended probe oligonucleotide containing a sequence complementary to at least a portion of the nucleic acid target; and The more than one extended probe oligonucleotide or its product is barcoded using more than one oligonucleotide barcode to generate more than one barcoded probe oligonucleotide, wherein each of the more than one oligonucleotide barcodes contains a molecular marker, and wherein each of the more than one barcoded probe oligonucleotides contains a molecular marker, a probe sequence, and a sequence complementary to at least a portion of the nucleic acid target.
2. The method according to claim 1, further comprising: Obtain sequencing data containing more than one sequencing read of the barcoded probe oligonucleotide or its product, wherein each of the more than one sequencing read contains a molecular marker sequence and a subsequence of the nucleic acid target; and The copy number of the nucleic acid target in the sample is determined based on the number of molecular markers associated with the more than one barcoded detection oligonucleotide or its product.
3. A method for determining the copy number of a nucleic acid target in a sample, the method comprising: A sample containing a copy of a nucleic acid target is contacted with more than one probe oligonucleotide, wherein each probe oligonucleotide contains a coupling sequence and a probe sequence configured to hybridize with the nucleic acid target. The more than one probe oligonucleotide hybridized with the nucleic acid target copy is extended to produce more than one extended probe oligonucleotide, each of the more than one extended probe oligonucleotide containing a sequence complementary to at least a portion of the nucleic acid target. The more than one extended probe oligonucleotide or its product is barcoded using more than one oligonucleotide barcode to generate more than one barcoded probe oligonucleotide, wherein each of the more than one oligonucleotide barcodes contains a molecular marker, and each of the more than one barcoded probe oligonucleotides contains a molecular marker, a probe sequence, and a sequence complementary to at least a portion of the nucleic acid target. Obtain sequencing data containing more than one sequencing read of the barcoded probe oligonucleotide or its product, wherein each of the more than one sequencing read contains a molecular marker sequence and a subsequence of the nucleic acid target; and The copy number of the nucleic acid target in the sample is determined based on the number of molecular markers associated with the more than one barcoded detection oligonucleotide or its product.
4. A method for determining the spatial location and copy number of a nucleic acid target in a sample, comprising: Each of two or more spatial locations of a sample containing a copy of the nucleic acid target is contacted with more than one probe oligonucleotide, wherein each probe oligonucleotide comprises a coupling sequence, a probe sequence configured to hybridize with the nucleic acid target, and a predetermined spatial marker. The probe oligonucleotides that are in contact with the same spatial location contain the same spatial marker sequence, and the probe oligonucleotides that are in contact with different spatial locations of the sample contain different spatial marker sequences; The more than one probe oligonucleotide hybridized with the nucleic acid target copy is extended to produce more than one extended probe oligonucleotide, each of the more than one extended probe oligonucleotide containing a sequence complementary to at least a portion of the nucleic acid target. The more than one extended probe oligonucleotide or its product is barcoded using more than one oligonucleotide barcode to generate more than one barcoded probe oligonucleotide, wherein each of the more than one oligonucleotide barcodes contains a molecular marker, and each of the more than one barcoded probe oligonucleotides contains a molecular marker, a probe sequence, and a sequence complementary to at least a portion of the nucleic acid target. Obtain sequencing data containing more than one sequencing read of the barcoded detection oligonucleotide or its product, wherein each of the more than one sequencing read contains a spatial marker sequence, a molecular marker sequence, and a subsequence of the nucleic acid target; and For each unique spatial marker sequence, which is associated with a different spatial location in the sample, the number of molecular markers with different sequences associated with the nucleic acid target is counted to determine the copy number of the nucleic acid target at each spatial location in the sample.
5. The method according to any one of claims 1-4, wherein barcoding the more than one extended probe oligonucleotide or its product using more than one oligonucleotide barcode comprises: Provide a coupled oligonucleotide comprising the 5' complement of the coupling sequence and the 3' complement of the capture sequence; Hybridize the extended probe oligonucleotide's coupling sequence with the 5' complement of the coupling sequence of the oligonucleotide. The 3' complement of the capture sequence of the coupled oligonucleotide is hybridized with the capture sequence of the oligonucleotide barcode in the more than one oligonucleotide barcode; and The extended probe oligonucleotide is linked to the hybridized oligonucleotide barcode.
6. The method according to any one of claims 1-5, comprising, prior to ligating the extended probe oligonucleotide to the oligonucleotide barcode, filling the gap between the extended probe oligonucleotide and the hybridized oligonucleotide barcode with a DNA polymerase lacking at least one of 5' to 3' exonuclease activity and 3' to 5' exonuclease activity.
7. The method according to any one of claims 1-6, wherein the extended probe oligonucleotide is ligated to the hybridized oligonucleotide barcode using a DNA ligase.
8. The method according to any one of claims 1-7, wherein the coupled oligonucleotide is a single-stranded oligonucleotide, and optionally the coupled oligonucleotide comprises at least 6 nucleotides.
9. The method according to any one of claims 1-8, wherein the coupling sequence comprises at least 4 nucleotides.
10. The method according to any one of claims 1-9, wherein the 5' end of each probe oligonucleotide is phosphorylated.
11. The method according to any one of claims 1-10, wherein the probe oligonucleotide is capable of entering the cells and / or nuclei of the sample, optionally entering the permeable cells and / or nuclei of the sample.
12. The method according to any one of claims 1-11, further comprising, after contacting the probe oligonucleotide with the sample, removing one or more probe oligonucleotides from the more than one probe oligonucleotide that were not in contact with the sample, optionally, removing the one or more probe oligonucleotides that were not in contact with the sample comprises: Remove one or more probe oligonucleotides from cells that have not entered the sample.
13. The method according to any one of claims 1-12, wherein the contact step comprises contacting the sample with a device configured to deposit probe oligonucleotides, optionally an inkjet device.
14. The method according to any one of claims 1-13, wherein the device is a needle, needle array, tube, aspiration device, injection device, electroporation device, fluorescence-activated cell sorting device, inkjet device, microfluidic device, or any combination thereof.
15. The method according to any one of claims 1-14, wherein the device contacts different spatial locations of the sample at a specified rate.
16. The method according to any one of claims 1-15, wherein the length of the spatial marker is 6-60 nucleotides.
17. The method according to any one of claims 1-16, wherein the two or more spatial locations comprise at least about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 12, about 14, about 16, about 18, about 20, about 30, about 40, about 50, about 60, about 70, about 80, about 90, or about 100 different spatial locations of the sample.
18. The method according to any one of claims 1-17, wherein the spatial location of the sample corresponds to a region containing no more than about 50 cells.
19. The method according to any one of claims 1-18, comprising contacting the sample with an extension reagent.
20. The method according to any one of claims 1-19, wherein at least a portion of the contacting step is carried out in the presence of the extending agent, optionally the entire contacting step is carried out in the presence of the extending agent, and optionally the contacting step and the extending step are simultaneous.
21. The method according to any one of claims 1-20, wherein the extension is performed in situ, optionally the extension includes in situ reverse transcription, and also optionally the cells of the sample remain intact during the extension step.
22. The method according to any one of claims 1-21, wherein the extension reagent comprises a reverse transcription reagent, optionally comprising reverse transcriptase and dNTPs.
23. The method according to any one of claims 1-22, wherein the reverse transcriptase comprises a viral reverse transcriptase, optionally wherein the viral reverse transcriptase is murine leukemia virus (MLV) reverse transcriptase or Moloney murine leukemia virus (MMLV) reverse transcriptase.
24. The method according to any one of claims 1-23, wherein the sample is physically divided or kept intact during the contact step.
25. The method according to any one of claims 1-24, wherein the sample comprises a single cell.
26. The method according to any one of claims 1-25, wherein the sample comprises more than one single cell.
27. The method according to any one of claims 1-26, wherein the sample comprises more than one cell, and optionally the method comprises: Dissociate the sample to produce more than one single cell. Optionally, the dissociation includes chemical dissociation, enzymatic dissociation, and / or mechanical dissociation. Optionally, the dissociation employs one or more of collagenase, chymotrypsin, dispersin, elastase, hyaluronidase, pancreatin, papain, and trypsin.
28. The method according to any one of claims 1-27, comprising, prior to the barcode generation step: The more than one single-cell partition, wherein the partition in the more than one partition contains single cells from the more than one single cell; and In the partition containing the single cell, the extended probe oligonucleotide is brought into contact with more than one oligonucleotide barcode.
29. The method according to any one of claims 1-28, wherein in a partition containing the single cell, the single cell is contacted with a lysis buffer at 15°C-65°C to lyse the single cell, optionally the lysis buffer containing an agent capable of dissociating protein-nucleic acid complexes.
30. The method according to any one of claims 1-29, wherein each oligonucleotide barcode in the more than one oligonucleotide barcode comprises a first universal sequence.
31. The method according to any one of claims 1-30, wherein obtaining sequencing data comprises: The more than one barcoded probe oligonucleotide is amplified using a first primer capable of hybridizing with the first universal sequence or its complement and an amplification primer capable of hybridizing with the nucleic acid target or its complement, thereby generating more than one amplified barcoded probe oligonucleotide. Obtaining sequencing data includes obtaining sequencing data containing more than one sequencing read of the barcoded probe oligonucleotide or its product that has been amplified.
32. The method according to any one of claims 1-31, wherein obtaining sequencing data comprises attaching the binding sites of sequencing primers and / or sequencing adaptors to the more than one barcoded probe oligonucleotide or its product.
33. The method according to any one of claims 1-32, wherein the amplification primer comprises a second universal sequence and / or wherein the first primer comprises a third universal sequence.
34. The method according to any one of claims 1-33, wherein: (i) The first universal sequence, the second universal sequence, and / or the third universal sequence are identical; and / or (ii) The first universal sequence, the second universal sequence and / or the third universal sequence are different.
35. The method according to any one of claims 1-34, wherein the first universal sequence, the second universal sequence and / or the third universal sequence comprises a binding site of a sequencing primer and / or a sequencing adaptor, its complementary sequence and / or a portion thereof.
36. The method according to any one of claims 1-35, wherein the sequencing adaptor comprises a P5 sequence, a P7 sequence, a complementary sequence thereof, and / or a portion thereof.
37. The method according to any one of claims 1-36, wherein the sequencing primers comprise a read 1 sequencing primer, a read 2 sequencing primer, their complementary sequences and / or portions thereof.
38. The method according to any one of claims 1-37, wherein the sample comprises more than one nucleic acid target, optionally at least about two different nucleic acid target groups.
39. The method according to any one of claims 1-38, wherein two or more nucleic acid targets of the target group are biomarkers, optionally the biomarkers are biomarkers of diseases or conditions, and even optionally the diseases or conditions are cancer, infection, viral infection, inflammatory disease, neurodegenerative disease, fungal disease, bacterial infection or any combination thereof.
40. The method of any one of claims 1-39, wherein the contacting step comprises contacting the sample with a group of probe oligonucleotides comprising two or more probe oligonucleotides, wherein each or more of them comprises a probe sequence configured to hybridize with a nucleic acid target in the more than one nucleic acid target.
41. The method according to any one of claims 1-40, wherein determining the copy number of the nucleic acid target in the sample comprises determining the copy number of each of the more than one nucleic acid target in the sample based on the number of molecular markers having different sequences associated with the more than one barcoded probe oligonucleotide or its product, the more than one barcoded probe oligonucleotide or its product comprising the sequence of each of the more than one nucleic acid target.
42. The method according to any one of claims 1-41, comprising, for each unique spatial marker sequence associated with a different spatial location of the sample, counting the number of molecular markers having a different sequence associated with each of the more than one nucleic acid target, to determine the copy number of each of the more than one nucleic acid target at each spatial location of the sample.
43. The method according to any one of claims 1-42, wherein the amplification primers comprise a set of amplification primers configured to hybridize with the more than one nucleic acid target or its complement, optionally comprising a set of at least about two different amplification primers.
44. The method according to any one of claims 1-43, wherein the nucleic acid target comprises a nucleic acid molecule, optionally wherein the nucleic acid molecule comprises ribonucleic acid (RNA), messenger RNA (mRNA), microRNA, small interfering RNA (siRNA), RNA degradation products, RNA containing multiple (A) tails, sample index oligonucleotides, cell component binding reagent-specific oligonucleotides, or any combination thereof.
45. The method according to any one of claims 1-44, wherein the more than one cell comprises one or more cell types, optionally, the one or more cell types being selected from the group consisting of: brain cells, heart cells, cancer cells, circulating tumor cells, organ cells, epithelial cells, metastatic cells, benign cells, primary cells and circulating cells, or any combination thereof.
46. The method according to any one of claims 1-45, wherein the sample comprises a biological sample, a clinical sample, an environmental sample, a biological fluid, a tissue, a tissue section derived from a subject, or any combination thereof, optionally the subject being a human, mouse, dog, rat, or vertebrate.
47. The method according to any one of claims 1-46, further comprising determining the subject's genotype, phenotype, or one or more gene mutations based on the spatial location of the nucleic acid target in the sample.
48. The method according to any one of claims 1-47, further comprising predicting the subject's susceptibility to one or more diseases, optionally cancer or a hereditary disease.
49. The method according to any one of claims 1-48, further comprising determining the cell type of the more than one cell in the sample.
50. The method according to any one of claims 1-49, wherein the drug is selected based on the predicted reactivity of the cell type of the more than one cell in the sample.
51. The method according to any one of claims 1-50, comprising imaging the sample, optionally before and / or after the contact step, wherein the imaging optionally generates imaging data.
52. The method according to any one of claims 1-51, wherein imaging the sample comprises staining the sample with a staining agent, wherein the staining agent is a fluorescent staining agent, a negative staining agent, an antibody staining agent or any combination thereof, and optionally staining comprises immunocytochemistry (ICC), immunohistochemistry (IHC), immunofluorescence (IF) or any combination thereof.
53. The method according to any one of claims 1-52, wherein imaging comprises microscopy, confocal microscopy, time-of-flight imaging microscopy, fluorescence microscopy, multiphoton microscopy, quantitative phase microscopy, surface-enhanced Raman spectroscopy, photography, manual visual analysis, automated visual analysis, or any combination thereof.
54. The method according to any one of claims 1-53, further comprising associating imaging data and sequencing data of one or more spatial locations of the sample.
55. The method according to any one of claims 1-54, comprising correlation analysis of imaging data and sequencing data of the said spatial location, optionally, the correlation analysis identifying one or more of the following: candidate biomarkers, candidate therapeutic agents, candidate doses of therapeutic agents, and / or cellular targets of candidate therapeutic agents.
56. The method according to any one of claims 1-55, wherein the imaging produces an image of a spectrum for constructing a physical representation of the sample, optionally the spectrum being two-dimensional or three-dimensional.
57. The method according to any one of claims 1-56, comprising mapping the nucleic acid target and / or cellular component target onto the map of the sample.
58. The method according to any one of claims 1-57, comprising mapping one or more single cells of the more than one cell onto the atlas of the sample.
59. The method according to any one of claims 1-58, wherein the sample has been contacted with one or more fixatives and / or permeating agents.
60. The method according to any one of claims 1-59, wherein the sample comprises tissue, cell monolayer, fixed cells, tissue section or any combination thereof, optionally fresh tissue section, frozen tissue section, fixed tissue section, formalin-fixed tissue section, formalin-fixed paraffin-embedded (FFPE) tissue section, acetone-fixed tissue section, paraformaldehyde (PFA)-fixed tissue section and / or methanol-fixed tissue section.
61. The method according to any one of claims 1-60, wherein the sample comprises a cell nuclear suspension, optionally a fixed cell nuclear suspension and / or a permeabilized cell nuclear suspension.
62. The method according to any one of claims 1-61, wherein the sample comprises cells, optionally fresh cells, frozen cells, fixed cells, formalin-fixed cells, formalin-fixed paraffin-embedded (FFPE) cells, acetone-fixed cells, paraformaldehyde (PFA)-fixed cells, and / or methanol-fixed cells.
63. The method according to any one of claims 1-62, comprising: Permeate the sample and / or fix the sample.
64. The method according to any one of claims 1-63, wherein fixing the sample comprises contacting the sample with a fixative; optionally, wherein the fixative comprises a non-crosslinking fixative, and further optionally, the non-crosslinking fixative comprises methanol; or optionally, wherein the fixative comprises a crosslinking agent.
65. The method of claim 64, wherein the crosslinking agent comprises a degradable crosslinking agent, and optionally wherein (a) the degradable crosslinking agent comprises or is derived from the following: dithiobis(succinimide propionate) (DSP), disuccinimide tartrate (DST), bis[2-(succinimideoxycarbonyloxy)ethyl] sulfone (BSOCOES), ethylene glycol bis(succinimide succinate) (EGS), dimethyl 3,3'-dithiobispropionylimide (DTBP), succinimide 3-(2-pyridyldithio)propionate (SPDP), succinimide 6-(3(2-pyridyldithio)propamido)hexanoate (LC-SPDP), 4-succinimide-oxycarbonyl-α-methyl-α(2-pyridyldithio)toluene (SMPT), 3-(2-pyridyldithio)propionylhydrazine (PDPH), succinimide-2-((4,4'-azidopentamido)ethyl)-1,3'-dithiopropionate (SDAD, NHS-SS-diazolidine) or any combination thereof; (b) the cleavable crosslinker comprises a cleavable linker selected from the group consisting of: chemically cleavable links, photocleavable links, acid-labile linkers, heat-sensitive links, enzyme-cleavable links, and combinations thereof; and / or (c) the cleavable crosslinker is a thiol cleavable crosslinker or contains a disulfide linker.
66. The method according to any one of claims 1-65, wherein the fixative comprises paraformaldehyde (PFA), dithiobis(succinimide propionate) (DSP), succinimide 3-(2-pyridyl dithio)propionate (SPDP), CellCover, or a combination thereof.
67. The method according to any one of claims 1-66, wherein fixing the sample and permeating the sample are performed simultaneously.
68. The method according to any one of claims 1-67, wherein the fixation and permeation of the sample are carried out in the presence of a dual-functional agent capable of fixing and permeating the sample, optionally wherein the dual-functional agent is methanol.
69. The method according to any one of claims 1-68, wherein permeating the sample comprises contacting the sample with a permeating agent.
70. The method according to any one of claims 1-69, comprising removing the permeabilizer from the sample after contacting the more than one probe oligonucleotide or the more than one cell component binding reagent with the sample.
71. The method according to any one of claims 1-70, wherein: The permeabilizer is capable of (i) permeating the cell membrane of the cell, and (ii) making the cell membrane of the cell permeable to the probe oligonucleotide or the cell component binding reagent or both; The permeating agent includes (i) a solvent, detergent or surfactant; (ii) BD Cytoperm; (iii) a saponin or a derivative thereof; (iv) Triton X-100; (v) methanol or a derivative thereof; and / or (vi) digitalis saponin or a derivative thereof. Agents capable of dissociating protein-nucleic acid complexes include broad-spectrum serine proteases, optionally said broad-spectrum serine proteases being proteinase K; and / or The lysis buffer contains (a) a defixing agent, optionally including thiols, hydroxylamine, periodate, bases or any combination thereof; and / or (b) DTT.
72. The method according to any one of claims 1-71, comprising reversing the fixation of the sample and / or single cells, optionally wherein reversing the fixation of the sample and / or single cells comprises UV photolysis, chemical treatment, heating, enzyme treatment, or any combination thereof.
73. The method according to any one of claims 1-72, wherein the sample comprises more than one cellular component target, and wherein the method further comprises: The sample is contacted with more than one cell component binding agent, each of the more than one cell component binding agent comprising a cell component binding agent-specific oligonucleotide, the cell component binding agent-specific oligonucleotide comprising a unique identifier sequence of the cell component binding agent, and wherein the cell component binding agent is capable of specifically binding to at least one of the more than one cell component targets; The cell component binding reagent-specific oligonucleotides are barcoded to generate more than one type of barcoded cell component binding reagent-specific oligonucleotide, each of the more than one barcoded cell component binding reagent-specific oligonucleotides comprising a sequence complementary to at least a portion of the unique identifier sequence and a molecular marker sequence. and Obtain sequencing data, the sequencing data comprising more than one sequencing read of more than one barcoded cell component binding reagent-specific oligonucleotide or its product, wherein each of the more than one sequencing read contains at least a portion of a molecular marker sequence and the unique identifier sequence.
74. The method according to any one of claims 1-73, wherein obtaining sequencing data comprises attaching the binding sites of sequencing primers and / or sequencing adaptors to the barcoded cell component binding reagent-specific oligonucleotides or their products.
75. The method according to any one of claims 1-74, the method comprising, after contacting the more than one cell component binding agent with the sample, removing one or more cell component binding agents from the more than one cell component binding agent that have not been in contact with the sample, optionally, removing one or more cell component binding agents that have not been in contact with the sample comprises: Remove one or more cell component binding agents that have not been in contact with at least one of the corresponding cell component targets.
76. The method according to any one of claims 1-75, wherein the cellular component target comprises intracellular proteins, carbohydrates, lipids, proteins, extracellular proteins, cell surface proteins, cell markers, B cell receptors, T cell receptors, major histocompatibility complex, tumor antigens, receptors, intracellular proteins, or any combination thereof.
77. The method according to any one of claims 1-76, wherein the cell component binding reagent-specific oligonucleotide comprises a second molecular marker, and optionally at least 10 of the more than one cell component binding reagent-specific oligonucleotide comprises different second molecular marker sequences.
78. The method according to any one of claims 1-77, wherein: At least two cell component-binding reagent-specific oligonucleotides have different second molecular marker sequences, and the unique identifier sequences of said at least two cell component-binding reagent-specific oligonucleotides are identical; and / or The second molecular marker sequences of at least two cell component binding reagent-specific oligonucleotides are different, and the unique identifier sequences of said at least two cell component binding reagent-specific oligonucleotides are different.
79. The method according to any one of claims 1-78, wherein: The number of unique molecular marker sequences in the sequencing data associated with the unique identifier sequence used for the cell component binding agent indicates the copy number of the at least one cell component target in the sample, the cell component binding agent being capable of specifically binding to the at least one cell component target; and / or The number of unique second molecular marker sequences associated with the unique identifier sequence used for the cell component binding reagent in the sequencing data indicates the copy number of the at least one cell component target in the sample, the cell component binding reagent being able to specifically bind to the at least one cell component target.
80. The method according to any one of claims 1-79, comprising contacting the sample with a blocking agent, one or more decoy oligonucleotides and / or one or more blocking oligonucleotides before contacting the sample with more than one cell component binding agent and / or contacting the sample with more than one probe oligonucleotide.
81. The method according to any one of claims 1-80, wherein contacting the sample with more than one cell component binding agent is performed in the presence of a blocking agent.
82. The method according to any one of claims 1-81, wherein the blocking agent comprises more than one oligonucleotide complementary to at least a portion of the cellular component binding agent-specific oligonucleotide.
83. The method according to any one of claims 1-82, wherein the blocking agent comprises an antibody or a fragment thereof derived from a first species, and wherein the blocking agent comprises serum derived from the first species.
84. The method according to any one of claims 1-83, wherein the sample comprises one or more non-target nucleic acids, and wherein the blocking agent comprises more than one decoy oligonucleotide capable of hybridizing with at least one of the one or more non-target nucleic acids.
85. The method according to any one of claims 1-84, wherein each of the more than one bait oligonucleotide is capable of hybridizing with at least a portion of a non-target nucleic acid.
86. The method according to any one of claims 1-85, wherein the bait oligonucleotide: It contains a sequence that is at least partially complementary to a non-target nucleic acid; The sequence comprises an oligonucleotide that is identical or substantially similar to the sequence of the cell component binding reagent, optionally having a length of 3 to 40 nucleotides; The specific oligonucleotides that bind to the cell components have up to 50% sequence identity; UMI is not included; It contains a random sequence, and optionally, the random sequence has a length of about 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14 or 15 nucleotides; It does not contain any sequence having more than 4, 5, 6, or 7 consecutive T or A sequences; It contains at least one G or C in every 4, 5, 6 or 7 consecutive nucleotides; Nucleotides containing one or more modifications; It includes a 5' modification, and optionally the 5' modification includes a 5' amino C12 modification (5AmMC12); Contains a 3' modification, and optionally the 3' modification contains a 3' dideoxy-C modification (ddC); and / or It is 30 to 65 nucleotides in length.
87. The method according to any one of claims 1-86, wherein the sample comprises one or more undesirable nucleic acid substances, the method comprising: The blocking oligonucleotide is contacted with the sample, wherein the blocking oligonucleotide specifically binds to at least one of the one or more undesirable nucleic acid substances; The reverse transcription of at least one of the one or more undesirable nucleic acid substances is reduced by the blocking oligonucleotide.
88. The method according to any one of claims 1-87, wherein the blocking oligonucleotide: The sample is contacted before the more than one probe oligonucleotide comes into contact with the sample; Contact the sample after contacting the sample with more than one probe oligonucleotide; and / or The sample is contacted when more than one probe oligonucleotide comes into contact with the sample.
89. The method according to any one of claims 1-88, the method comprising providing a blocking oligonucleotide that specifically binds to two or more undesirable nucleic acid substances in the sample, optionally at least 10 or up to at least 100 undesirable nucleic acid substances.
90. The method according to any one of claims 1-89, wherein the blocking oligonucleotide: It can be locked nucleic acid (LNA), peptide nucleic acid (PNA), DNA, LNA / PNA chimera, LNA / DNA chimera, or PNA / DNA chimera; It specifically binds to the 3' end of one or more undesirable nucleic acid substances within 100 nt; Specifically binds to the 5' end of one or more undesirable nucleic acid substances within 100 nt, or the blocking oligonucleotide specifically binds to the middle 100 nt of one or more undesirable nucleic acid substances. May or may not contain non-natural nucleotides; T with at least 50°C m ; It cannot be used as a primer for reverse transcriptase or polymerase; and / or The length ranges from 8 nt to 100 nt.
91. The method according to any one of claims 1-90, wherein: The one or more undesirable nucleic acid substances account for about 50% to about 80% of the nucleic acid content of the sample; The undesirable nucleic acid material is selected from the group consisting of: ribosomal RNA, mitochondrial RNA, genomic DNA, intron sequences, high-abundance sequences, and combinations thereof; and / or The one or more undesirable nucleic acid substances are mRNA molecules, and the blocking oligonucleotide specifically binds to 10 nt of the 3' multi(A) tail of the one or more undesirable nucleic acid substances.
92. The method according to any one of claims 1-91, wherein: Each molecular marker of the more than one oligonucleotide barcode contains at least 6 nucleotides; Each capture sequence of the more than one oligonucleotide barcode contains at least four nucleotides; The more than one oligonucleotide barcode is associated with a solid support, and the partition in the more than one partition contains a single solid support; Each of the more than one oligonucleotide barcodes comprises a cell marker, and optionally each cell marker of the more than one oligonucleotide barcode comprises at least 6 nucleotides; and / or The oligonucleotide barcodes associated with the same solid support in the more than one oligonucleotide barcode contain the same cell marker, and optionally the oligonucleotide barcodes associated with different solid supports in the more than one oligonucleotide barcode contain different cell markers.
93. The method according to any one of claims 1-92, wherein the solid support comprises synthetic particles, a flat surface, or a combination thereof.
94. The method according to any one of claims 1-93, comprising associating a synthetic particle containing the more than one oligonucleotide barcode with cells in the partition.
95. The method according to any one of claims 1-94, This includes lysing the cells after associating the synthetic particles with the cells, optionally including heating the cells, contacting the cells with a detergent, altering the pH of the cells, or any combination thereof; The synthetic particles and the single cells are in the same partition, and optionally the partition is a pore or a droplet; In the case of the above-mentioned oligonucleotide barcodes, at least one oligonucleotide barcode is fixed or partially fixed on the synthetic particle, or at least one oligonucleotide barcode is encapsulated or partially encapsulated in the synthetic particle; The synthetic particles are destructible, optionally destructible hydrogel particles; The synthetic particles described herein comprise beads, optionally including agarose gel beads, streptavidin beads, agarose beads, magnetic beads, conjugated beads, protein A conjugated beads, protein G conjugated beads, protein A / G conjugated beads, protein L conjugated beads, oligo(dT) conjugated beads, silica beads, silica-like beads, avidin microbeads, antifluorescent dye microbeads, or any combination thereof; and / or The synthetic particles described herein comprise materials selected from the group consisting of: polydimethylsiloxane (PDMS), polystyrene, glass, polypropylene, agarose, gelatin, hydrogel, paramagnetic material, ceramic, plastic, glass, methylstyrene, acrylic polymer, titanium, latex, agarose gel, cellulose, nylon, silicone, and any combination thereof.
96. The method according to any one of claims 1-95, Each of the more than one oligonucleotide barcodes contains a linker functional group. The synthesized particles contain solid support functional groups, and The support functional group and the connector functional group are associated with each other, and optionally the connector functional group and the support functional group are individually selected from the group consisting of: C6, biotin, streptavidin, one or more primary amines, one or more aldehydes, one or more ketones and any combination thereof.
97. A reagent kit comprising: More than one probe oligonucleotide, wherein each of the probe oligonucleotides comprises a coupling sequence and a probe sequence configured to hybridize with a nucleic acid target, and optionally, the probe oligonucleotide comprises a predetermined spatial marker; Coupled oligonucleotides comprising a 5' complement of the coupling sequence and a 3' complement of the capture sequence; More than one oligonucleotide barcode, wherein the 3' end of each oligonucleotide barcode is associated with a solid support, and the 5' end of each oligonucleotide barcode contains a capture sequence; A first primer capable of hybridizing with a first universal sequence, optionally further comprising a third universal sequence; Amplification primers capable of hybridizing with the nucleic acid target or its complement may optionally further comprise a second universal sequence; DNA ligase; Extension reagents, optionally reverse transcription reagents, and optionally reverse transcriptase and dNTPs; One or more fixatives; One or more permeable agents; Crosslinking agent; De-fixing agent; Lysis buffer; More than one cell component binding agent, wherein each of the more than one cell component binding agent comprises a cell component binding agent-specific oligonucleotide, the cell component binding agent-specific oligonucleotide comprising a unique identifier sequence of the cell component binding agent, and wherein the cell component binding agent is capable of specifically binding to at least one of more than one cell component targets; Blocking reagents; One or more bait oligonucleotides; and / or One or more blocking oligonucleotides.
98. The kit of claim 97, wherein the more than one probe oligonucleotide comprises a set of probe oligonucleotides, the set of probe oligonucleotides comprising two or more types of more than one probe oligonucleotide, wherein each of more than one contains a probe sequence configured to hybridize with nucleic acid targets in more than one nucleic acid target, optionally a target group of at least about two different nucleic acid targets.
99. The kit according to any one of claims 97-98, wherein the amplification primers comprise a set of amplification primers configured to hybridize with more than one nucleic acid target or its complement, optionally comprising a set of at least about two different amplification primers.
Citation Information
Patent Citations
Beet puller and topping machine.
US1258367A
Digital Counting of Individual Molecules by Stochastic Attachment of Diverse Labels
US20110160078A1
Massively parallel single cell analysis
US20150299784A1
Measurement of protein expression using reagents with barcoded oligonucleotide sequences
US20180088112A1
Sample indexing for single cells
US20180346970A1