Methods and systems for preparing a nucleic acid library
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-08-19
- Publication Date
- 2026-03-26
AI Technical Summary
Existing methods for nucleic acid amplification and sequencing in biological samples are limited by probe hybridization-based assays, which are not ideal for unbiased whole transcriptome analysis, and reverse transcription reactions often continue beyond template exhaustion, necessitating a workflow for controlled fragment generation and unbiased library preparation for de novo sequencing.
A method involving hybridization of an oligonucleotide to an RNA template, followed by extension with a reverse transcriptase and free nucleotides, including a terminating nucleotide to stop extension, digestion of the RNA template, and generation of a circularized template for rolling circle amplification, allowing for unbiased whole transcriptome sequencing in situ.
This approach enables the generation of uniform amplification products for sensitive and accurate in situ sequencing by controlling fragment length, reducing bias and improving assay sensitivity.
Smart Images

Figure US2025042580_26032026_PF_FP_ABST
Abstract
Description
202412023740METHODS AND SYSTEMS FOR PREPARING A NUCLEIC ACID LIBRARYCROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims priority to U.S. Provisional Patent Application No. 63 / 685,176, filed August 20, 2024, entitled “METHODS AND SYSTEMS FOR PREPARING A NUCLEIC ACID LIBRARY,” and U.S. Provisional Patent Application No. 63 / 819,236, filed June 6, 2025, entitled “METHODS AND SYSTEMS FOR PREPARING A NUCLEIC ACID LIBRARY,” each of which is herein incorporated by reference in its entirety for all purposes.FIELD
[0002] The present disclosure relates in some aspects to methods for in situ analysis of nucleic acids (e.g., RNAs) in a biological sample including generating a library of nucleic acids for analysis. In some aspects, a plurality of nucleic acids generated using RNA templates are processed and analyzed in situ in a sample or used to generate a plurality of barcoded nucleic acid molecules for analysis.BACKGROUND
[0003] Profiling of analytes in cells and tissue samples involve accurate and reliable methods for nucleic acid amplification. Certain existing analyte targeting methods can yield products that may not be ideal for downstream assay processing and detection. Improved methods for generating a library of nucleic acids for unbiased in situ analysis and / or sequencing (e.g., de novo sequencing of transcripts) are needed. The present disclosure addresses these and other needs.SUMMARY
[0004] Rolling circle amplification (RCA)-based detection methods provide a powerful tool for detection of analytes at their relative spatial locations (e.g., in situ) in biological samples. However, an RCA-based assay that relies on probe hybridization for in situ analysis may have certain limitations, for example compared to de novo sequencing of transcripts or for performing whole transcriptome analysis in an unbiased manner. In some embodiments, de novo sequencing of transcripts requires a suitable process for generating a library of transcripts. In some aspects, a non-targeted approach is desirable for unbiased whole transcriptome sequencing in situ. If a reverse transcription reaction is performed using an RNA as template, the reaction continues until the template ends and / or other reagents are exhausted.1MOFO-358213026202412023740There is a need for a workflow that allows fragments of controlled lengths or short fragments to be generated and processed. In some aspects, there is a need for a library preparation workflow suitable for de novo sequencing of transcripts in situ. In some aspects, there is a need for an unbiased whole transcriptome library preparation workflow suitable for in situ sequencing. In some instances, using reverse transcription to generate a molecule for downstream detection allows for unbiased detection. In some embodiments, the detection is performed using in situ sequencing (e.g., base-by-base sequencing). In some embodiments, the detection is performed by sequencing a barcoded nucleic acid molecule.
[0005] In some aspects, provided herein are methods and compositions for generating a library of nucleic acids for analysis. In some aspects, a plurality of different RNAs are processed and analyzed in situ in a sample. Also provided are oligonucleotides, a plurality of free nucleotides, detection reagents, compositions, and systems for use in accordance with the methods. In some aspects, provided herein is a library preparation workflow suitable for de novo sequencing of transcripts in situ. In some aspects, provided herein is an unbiased (e.g., whole transcriptome) library preparation workflow suitable for de novo sequencing of transcripts in situ. In some aspects, provided herein is a library preparation workflow suitable for unbiased analysis of sequences associated with variable or unknown analyte sequences. In some aspects, provided herein is a library preparation workflow suitable for unbiased analysis of sequences associated with a variable sequence (e.g., SNV detection), exogenous nucleic acids or perturbation agents introduced to a cell (e.g., CRISPR barcode scanning), or a sequence of or associated with an immune molecule (e.g., for VDJ detection).
[0006] Provided herein is a method of nucleic acid processing comprising: (a) hybridizing an oligonucleotide to a ribonucleic acid (RNA) template; (b) extending the oligonucleotide using a reverse transcriptase and a plurality of free nucleotides to generate an extended oligonucleotide, wherein the plurality of free nucleotides comprises a terminating nucleotide that prevents further extension of the extended oligonucleotide; (c) digesting the RNA template; and (d) using the extended oligonucleotide to generate a circularized template.
[0007] In some embodiments, the terminating nucleotide is a nucleotide lacking a 3' hydroxyl group in the deoxyribose sugar. In some instances, the terminating nucleotide is a dideoxyribonucleotide (ddNTP). In some instances, the plurality of free nucleotides comprises two, three or four different ddNTPs selected from the group consisting of ddGTP, ddATP,2MOFO-358213026202412023740 ddTTP and ddCTP. In some instances, the terminating nucleotide comprises a terminating group or a terminating modification. In some instances, the terminating modification is a reverse linkage. In some instances, the terminating nucleotide comprises a 3' inverted dT.
[0008] In some embodiments, the terminating nucleotide is present in an amount lower than an amount of a corresponding free nucleotide in the plurality of free nucleotides. In some embodiments, the terminating nucleotide is present in an amount greater than an amount of a corresponding free nucleotide in the plurality of free nucleotides. In some instances, the amount of the terminating nucleotide is at least 2-fold lower than the amount of the corresponding free nucleotide in the plurality of free nucleotides. In some instances, the amount of the terminating nucleotide is at least 2-fold greater than the amount of the corresponding free nucleotide in the plurality of free nucleotides.
[0009] In some embodiments, the plurality of free nucleotides comprises a cleavable nucleotide. In some examples, the cleavable nucleotide is a uracil. In some instances, the cleavable nucleotide is an inosine. In some instances, the cleavable nucleotide is a ribonucleotide and the extended oligonucleotide is DNA. In some instances, the method further comprises using Uracil- Specific Excision Reagent (USER) to cleave the extended oligonucleotide before using the extended oligonucleotide to generate a circularized template. In some instances, the method further comprises using an endonuclease (e.g., Endonuclease V) to cleave the extended oligonucleotide before using the extended oligonucleotide to generate a circularized template. In some instances, the endonuclease comprises Endonuclease V. In some instances, the terminating nucleotide is a reversible terminator nucleotide. In some examples, the reversible terminator nucleotide comprises an azidomethyl group, an amino group, a nitrobenzyl group, an allyl group, a carbonate, a functional photocleavable ether, a methyl group, or a cyanoethyl group. In some instances, the reversible terminator nucleotide comprises an azidomethyl group, and amino group, a nitrobenzyl group, or an allyl group. In some instances, the reversible terminator nucleotide is a 3'-O-blocked reversible terminator nucleotide. In some instances, the plurality of free nucleotides comprises two, three or four different ddNTPs comprising the terminating group or the terminating modification. In some instances, RNase H is used to digest the RNA template.
[0010] In some instances, the method further comprises removing the terminating nucleotide from the extended oligonucleotide before using the extended oligonucleotide to3MOFO-358213026202412023740 generate a circularized template. In some examples, the terminating nucleotide is cleaved from the extended oligonucleotide. In some aspects, the terminating group or the terminating modification is removed from the terminating nucleotide incorporated into the extended oligonucleotide. In some instances, removing the terminating group or the terminating modification from the terminating nucleotide is performed by cleaving a linker in the terminating nucleotide.
[0011] In some embodiments, the plurality of free nucleotides comprise all four canonical bases: adenine, thymine, guanine and cytosine; and at least two different bases of terminating nucleotides. In some cases, a kinase is used to phosphorylate a 5’ end of the extended oligonucleotide before using the extended oligonucleotide to generate a circularized template. In some embodiments, a 5’ end and a 3’ end of the extended oligonucleotide is ligated to generate the circularized template. In some embodiments, the extended oligonucleotide is used as a template to generate a copy of the extended oligonucleotide and a 5’ end and a 3’ end of the copy of the extended oligonucleotide is ligated to generate the circularized template.
[0012] Provided herein is a method comprising: (a) hybridizing an oligonucleotide to an RNA template; (b) extending the oligonucleotide using a reverse transcriptase and a plurality of free nucleotides to generate an extended oligonucleotide, wherein the plurality of free nucleotides comprises a ribonucleotide, an inosine, and / or a uracil; (c) digesting the RNA template and cleaving the extended oligonucleotide at the incorporated ribonucleotide, inosine, and / or uracil, thereby generating a cleaved extended sequence; and (d) using the cleaved extended sequence to generate a circularized template. In some embodiments, a 5’ end and a 3’ end of the cleaved extended sequence is ligated to generate the circularized template. In some embodiments, the inosine is incorporated into the extended oligonucleotide and an endonuclease to cleave the extended oligonucleotide. In some instances, the uracil is incorporated into the extended oligonucleotide and Uracil-Specific Excision Reagent (USER) is used to cleave the extended oligonucleotide.
[0013] In some embodiments, the plurality of free nucleotides comprise all four canonical bases: adenine, thymine, guanine and cytosine; and at least two different ribonucleotides. In some embodiments, the plurality of free nucleotides comprise inosine and all four canonical bases: adenine, thymine, guanine and cytosine. In some embodiments, the4MOFO-358213026202412023740 plurality of free nucleotides comprise uracil and all four canonical bases: adenine, thymine, guanine and cytosine.
[0014] In some embodiments, a kinase is used to phosphorylate a 5’ end of the cleaved extended sequence. In some instances, a 5’ end and a 3’ end of the cleaved extended sequence is ligated to generate the circularized template. In some instances, RNase H is used to perform the digesting and / or cleaving. In some instances, an enzymatic ligation is performed to generate the circularized template. In some instances, a chemical ligation is performed to generate the circularized template. In some examples, the enzymatic ligation is performed using CircLigase. In some cases, the ligation is performed using a splint.
[0015] In some instances, the circularized template does not comprise a barcode sequence.
[0016] In some instances, the oligonucleotide comprises a primer binding sequence.
[0017] In some embodiments, the method further comprises performing rolling circle amplification (RCA) of the circularized template to generate a rolling circle amplification product (RCP); and detecting the RCP in the biological sample. In some cases, performing RCA comprises binding a primer to the primer binding sequence of the circularized template, and extending the primer to generate an amplification product comprising multiple copies of the RNA template or a complement thereof. In some instances, RCA is performed using a polymerase having strand-displacement activity to generate the RCP. In some examples, the polymerase is a Phi29 polymerase.
[0018] In some embodiments, the extended oligonucleotide has a length of 50-200 nucleotides. In some embodiments, the extended oligonucleotide has a length of 70-100 nucleotides. In some embodiments, the method is performed in a biological sample. In some cases, the biological sample is a cell or tissue sample. In some embodiments, a sequence of the RCP is detected using sequencing by ligation, sequencing by synthesis, sequencing by binding, sequencing by avidity, or a combination thereof. In some cases, detecting the RCP comprises binding a sequencing primer to the RCP. In some cases, the sequencing primer binds to the primer binding sequence or a complement thereof.
[0019] In some embodiments, a biological sample is imaged to detect the RCP in situ in the biological sample or a matrix embedding the biological sample. In some embodiments, the sequence of the RCP or a complement thereof is sequenced using in situ sequencing by synthesis5MOFO-358213026202412023740(SBS) in the biological sample. In some embodiments, the sequence of the RCP or a complement thereof is detected using single nucleotide sequencing by synthesis. In some instances, the sequence of the RCP or a complement thereof is sequenced using in situ sequencing by ligation (SBL) in the biological sample. In some instances, the sequence of the RCP or a complement thereof is sequenced using in situ sequencing by binding (SBB) or sequencing-by-avidity (SBA) in the biological sample. In some embodiments, a sequencing primer is bound to the RCP and the sample is contacted with a polymerase and a first plurality of nucleotide molecules to form a complex comprising a 3’ terminus of the sequencing primer, the RCP, the polymerase, and a nucleotide molecule of the first plurality of nucleotide molecules and detecting presence of the nucleotide molecules in the complex to identify a complementary nucleotide in the RCP. In some cases, the sequencing primer comprises a 3’ terminal nucleotide that is reversibly blocked. In some embodiments, the method further comprises removing the complex, unblocking the reversibly blocked 3’ terminal nucleotide molecule and contacting the sequencing primer bound to the RCP with a polymerase and a second plurality of nucleotide molecules. In some embodiments, contacting the sequencing primer with an additional plurality of nucleotide molecules to identify additional complementary nucleotides of the sequence in the RCP is repeated for at least 2, 5, 10, 20, or 30 additional cycles.
[0020] In some embodiments, the RNA template is attached directly or indirectly to the biological sample or to a matrix embedding the biological sample. In some embodiments, the RNA template is crosslinked in the biological sample or in a matrix embedding the biological sample.
[0021] Provided herein is a method for nucleic acid processing comprising hybridizing an oligonucleotide to an RNA template in a biological sample; extending the oligonucleotide using a reverse transcriptase and a plurality of free nucleotides to generate an extended oligonucleotide, wherein the plurality of free nucleotides comprises a reversible terminator nucleotide and incorporation of the reversible terminator nucleotide stops further extension of the extended oligonucleotide; removing the terminating group from the reversible terminator nucleotide incorporated into the extended oligonucleotide; digesting the RNA template; ligating the 5’ and 3’ ends of the extended oligonucleotide to generate a circularized template; performing rolling circle amplification of the circularized template to generate a rolling circle amplification product (RCP); and detecting a sequence of the RCP in the biological sample6MOFO-358213026202412023740 using in situ sequencing by synthesis (SBS) in the biological sample, thereby detecting a sequence of the RCP or a complement thereof at a location in the biological sample. In some instances, the RNA template is digested before ligating the 5’ and 3’ ends of the extended oligonucleotide to generate a circularized template. In some instances, the RNA template is digested after ligating the 5’ and 3’ ends of the extended oligonucleotide to generate a circularized template.
[0022] In some embodiments, the biological sample is a cell or tissue sample comprising cells or cellular components. In some embodiments, the biological sample is a tissue section. In some embodiments, the biological sample is a formalin-fixed, paraffin-embedded (FFPE) sample, a frozen tissue sample, or a fresh tissue sample. In some embodiments, the biological sample is fixed and / or permeabilized. In some embodiments, the biological sample is crosslinked and / or embedded in a matrix. In some embodiments, the matrix comprises a hydrogel. In some embodiments, the biological sample is cleared.
[0023] Provided herein is a system comprising (i) an oligonucleotide; (ii) a plurality of free nucleotides that comprise different nucleic acid bases, wherein the plurality of free nucleotides comprises a terminating nucleotide that prevents further extension; and (iii) a reverse transcriptase for extending the oligonucleotide to generate an extended oligonucleotide. In some instances, the terminating nucleotide is a reversible terminator nucleotide. In some examples, the reversible terminator nucleotide comprises an azidomethyl group, an amino group, a nitrobenzyl group, an allyl group, a carbonate, a functional photocleavable ether, a methyl group, or a cyanoethyl group. In some instances, the reversible terminator nucleotide comprises an azidomethyl group, and amino group, a nitrobenzyl group, or an allyl group. In some instances, the reversible terminator nucleotide is a 3'-O-blocked reversible terminator nucleotide. In some instances, the plurality of free nucleotides comprises two, three or four different ddNTPs comprising a terminating group or a terminating modification.
[0024] Provided herein is a system comprising (i) an oligonucleotide; (ii) a plurality of free nucleotides that comprise different nucleic acid bases, wherein the plurality of free nucleotides comprises a cleavable nucleotide that prevents further extension; and (iii) a reverse transcriptase for extending the oligonucleotide to generate an extended oligonucleotide. In some examples, the cleavable nucleotide is a uracil. In some instances, the cleavable nucleotide is an inosine. In some instances, the cleavable nucleotide is a ribonucleotide and the extended7MOFO-358213026202412023740 oligonucleotide is DNA. In some instances, the system comprises Uracil-Specific Excision Reagent (USER). In some instances, the system comprises an Endonuclease V.
[0025] In some embodiments, the system comprises one or more reagents for circularizing the extended oligonucleotide. In some embodiments, the system comprises a polymerase having strand-displacement activity. In some instances, the polymerase is a Phi29 polymerase. In some instances, the oligonucleotide comprises a primer binding sequence. In some embodiments, the system comprises a plurality of a sequencing primer, a plurality of detectably labeled nucleotides, and a polymerase for performing sequencing in situ. In some instances, the oligonucleotide comprises a primer binding sequence configured to hybridize to the sequencing primer for performing sequencing in situ.
[0026] Provided herein is a method of nucleic acid processing, comprising in a biological sample, hybridizing a nucleic acid strand to a cleavage region in a ribonucleic acid (RNA) template; cleaving the RNA template in the cleavage region; hybridizing an oligonucleotide to the cleaved RNA template; and extending the oligonucleotide using a reverse transcriptase and a plurality of free nucleotides to generate an extended oligonucleotide comprising a sequence complementary to the cleaved RNA template. In some embodiments, the method further comprises detecting a sequence of the extended oligonucleotide or derivative thereof. In some embodiments, cleaving the RNA template comprises contacting the biological sample with an RNase H. In some instances, the cleavage region is 5’ to a region of interest in the RNA template. In some instances, the biological sample is contacted with the nucleic acid strand and with the RNase H simultaneously.
[0027] In some embodiments, the method further comprises using the extended oligonucleotide to generate a circularized template. In some embodiments, the method further comprises using a kinase to phosphorylate a 5’ end of the extended oligonucleotide prior to generating the circularized template. In some embodiments, generating the circularized template comprises ligating a 5’ end and a 3’ end of the extended oligonucleotide to generate the circularized template. In some embodiments, the extended oligonucleotide is used as a template to generate a copy of the extended oligonucleotide and a 5’ end and a 3’ end of the copy of the extended oligonucleotide is ligated to generate the circularized template.
[0028] In some embodiments, the RNase H comprises an RNase Hl and / or an RNase H2. In some embodiments, contacting the biological sample with the RNase H comprises8MOFO-358213026202412023740 contacting the biological sample with between about 0.5 enzyme units (U) and about 50 U of the RNase H. In some embodiments, the cleavage region is about 10 to about 22 nucleotides in length, or about 15 to about 20 nucleotides in length.
[0029] In some embodiments, the circularized template is generated by an enzymatic ligation. In some embodiments, the circularized template is generated by a chemical ligation. In some embodiments, the enzymatic ligation is performed using CircLigase. In some embodiments, the circularized template is generated by performing a ligation using a splint. In some embodiments, the circularized template does not comprise a barcode sequence.
[0030] In some embodiments, the oligonucleotide comprises a primer binding sequence. In some instances, the oligonucleotide is complementary to a sequence 3’ to the region of interest in the RNA template. In some instances, the oligonucleotide comprises a poly-T sequence or a randomer.
[0031] In some embodiments, the method further comprises performing rolling circle amplification (RCA) of the circularized template to generate a rolling circle amplification product (RCP). In some embodiments, performing RCA comprises binding a primer to the primer binding sequence of the circularized template, and extending the primer to generate an amplification product comprising multiple copies of the RNA template or a complement thereof. In some embodiments, the method comprises using a polymerase having strand-displacement activity to generate the RCP. In some examples, the polymerase is a Phi29 polymerase.
[0032] In some embodiments, the extended oligonucleotide has a length of 50-200 nucleotides. In some embodiments, the biological sample is a cell or tissue sample.
[0033] In some embodiments, the sequence of the extended oligonucleotide or derivative thereof comprises detecting a sequence of the RCP using sequencing by ligation, sequencing by synthesis, sequencing by binding, sequencing by avidity, or a combination thereof. In some embodiments, detecting the sequence of the RCP comprises binding a sequencing primer to the RCP. In some embodiments, the sequencing primer binds to the primer binding sequence or a complement thereof. In some embodiments, the method comprises imaging the biological sample to detect a sequence of the RCP in situ in the biological sample or a matrix embedding the biological sample. In some embodiments, the sequence of the RCP comprises the region of interest or a complement of a sequence in the region of interest. In some instances, the sequence of the RCP or a complement thereof is sequenced using in situ9MOFO-358213026202412023740 sequencing by synthesis (SBS) in the biological sample. In some instances, the sequence of the RCP or a complement thereof is sequenced using in situ sequencing by avidity (SBA) in the biological sample. In some instances, the sequence of the RCP or a complement thereof is sequenced using in situ sequencing by ligation (SBL) in the biological sample. In some instances, the sequence of the RCP or a complement thereof is sequenced using in situ sequencing by binding (SBB) in the biological sample.
[0034] In some embodiments, contacting the sequencing primer bound to the RCP with (i) a polymerase and (ii) a first plurality of nucleotide molecules to form a complex comprising a 3’ terminus of the sequencing primer, the RCP, the polymerase, and a nucleotide molecule of the first plurality of nucleotide molecules and detecting presence of the nucleotide molecules in the complex to identify a complementary nucleotide in the RCP. In some embodiments, the sequencing primer comprises a 3’ terminal nucleotide that is reversibly blocked. In some embodiments, the method further comprises removing the complex, unblocking the reversibly blocked 3’ terminal nucleotide molecule and contacting the sequencing primer bound to the RCP with a polymerase and a second plurality of nucleotide molecules. In some embodiments, the method further comprises repeating contacting the sequencing primer with an additional plurality of nucleotide molecules to identify additional complementary nucleotides of the sequence in the RCP for at least 2, at least 5, at least 10, at least 20, or at least 30 additional cycles.
[0035] In some embodiments, the RNA template is attached directly or indirectly to the biological sample or to a matrix embedding the biological sample. In some embodiments, the RNA template is crosslinked in the biological sample or in a matrix embedding the biological sample.
[0036] In some embodiments, the method further comprises digesting the cleaved RNA template after the extended oligonucleotide is generated. In some embodiments, additional RNase H is used to digest the cleaved RNA template.
[0037] In some embodiments, the extended oligonucleotide is a barcoded nucleic acid molecule. In some embodiments, the extended oligonucleotide is used to generate a barcoded nucleic acid molecule. In some embodiments, the oligonucleotide or a derivative of the extended oligonucleotide comprises a capture sequence. In some embodiments, the capture sequence is complementary to a capture domain of a capture probe. In some embodiments, the capture probe10MOFO-358213026202412023740 is immobilized on a substrate. In some embodiments, the oligonucleotide is immobilized on a substrate. In some embodiments, the capture probe comprises a spatial barcode. In some embodiments, the substrate is a bead. In some embodiments, the bead and the RNA template is in a partition. In some embodiments, the extended oligonucleotide is partitioned in a droplet or a well. In some embodiments, the method further comprises processing the barcoded nucleic acid molecule to append a functional sequence. In some embodiments, detecting the sequence of the extended oligonucleotide or a derivative thereof comprises sequencing the barcoded nucleic acid molecule or a derivative thereof.
[0038] Provided herein is a method of nucleic acid processing, comprising hybridizing an oligonucleotide to a ribonucleic acid (RNA) template; extending the oligonucleotide using a reverse transcriptase and a plurality of free nucleotides to generate an extended oligonucleotide, wherein the plurality of free nucleotides lacks nucleotides of at least one of four canonical bases: adenine (A), thymine (T), guanine (G), uracil (U), and cytosine (C); digesting the RNA template; and using the extended oligonucleotide to generate a circularized template. In some embodiments, the plurality of free nucleotides lacks nucleotides of at least two of four canonical bases: adenine (A), thymine (T), guanine (G), uracil (U), and cytosine (C). In some embodiments, the plurality of free nucleotides lacks nucleotides of three of four canonical bases: adenine (A), thymine (T), guanine (G), uracil (U), and cytosine (C). In some embodiments, the method comprises using an additional plurality of free nucleotides, wherein the additional plurality of free nucleotides comprises nucleotides of the canonical base lacking in the plurality of free nucleotides (e.g., to resume extension). In some embodiments, the system comprises a polymerase having strand-displacement activity. In some instances, the polymerase is a Phi29 polymerase. In some instances, the oligonucleotide comprises a primer binding sequence. In some embodiments, the system comprises a plurality of sequencing primers, a plurality of detectably labeled nucleotides, and a polymerase for performing sequencing in situ. In some instances, the oligonucleotide comprises a primer binding sequence configured to hybridize to the sequencing primer for performing sequencing in situ. In some instances, the system comprises a biological sample comprising the RNA template (e.g., an endogenous RNA template).
[0039] Provided herein is a method of nucleic acid processing comprising contacting a biological sample with a nucleic acid comprising one or more catalytic domains, wherein the11MOFO-358213026202412023740 nucleic acid hybridizes to a ribonucleic acid (RNA) template, and cleaves the RNA template; hybridizing an oligonucleotide to the RNA template cleaved by the nucleic acid; extending the oligonucleotide using a reverse transcriptase and a plurality of free nucleotides to generate an extended oligonucleotide; using the extended oligonucleotide to generate a circularized template. In some embodiments, the circularized template is generated by ligating the 5’ and 3’ ends of the extended oligonucleotide. In some instances, the method comprises performing rolling circle amplification of the circularized template to generate a rolling circle amplification product (RCP). In some instances, the method comprises detecting a sequence of the RCP in the biological sample. In some instances, detecting the sequence of the RCP comprises performing sequencing by synthesis (SBS), sequencing by avidity (SB A), sequencing by ligation (SBL) or sequencing by binding (SBB) in the biological sample. In some instances, detecting the sequence of the RCP comprises imaging the biological sample to detect the RCP in situ in the biological sample.
[0040] Provided herein is a kit or system comprising an oligonucleotide; a plurality of free nucleotides that comprise different nucleic acid bases, wherein the plurality of free nucleotides is lacking at least one of the canonical bases: adenine, thymine, guanine and cytosine; and a reverse transcriptase for extending the oligonucleotide to generate an extended oligonucleotide. In some instances, the system or kit comprises nucleotide pools of one nucleotide, e.g., only one of A, C, G, and T nucleotides in a single pool. In some instances, the system or kit comprises two or more pools of free nucleotides and the two or more pools are different. In some instances, a first pool is configured to halt extension of the oligonucleotide and a second pool is configured to allow extension of the oligonucleotide to resume.
[0041] Provided herein is a kit or system comprising a nucleic acid comprising one or more catalytic domains (e.g., DNAzyme or RNAzyme) configured to cleave a ribonucleic acid (RNA) template; an oligonucleotide; and a reverse transcriptase for extending the oligonucleotide using the RNA template to generate an extended oligonucleotide. Provided herein is a kit or system comprising an RNA-cutting enzyme (e.g., CRISPR effector proteins or Argonaute proteins) configured to cleave a ribonucleic acid (RNA) template; an oligonucleotide; and a reverse transcriptase for extending the oligonucleotide using the RNA template to generate an extended oligonucleotide. In some embodiments, the kit or system further comprises a biological sample comprising the RNA template.12MOFO-358213026202412023740
[0042] In some embodiments, the kit or system further comprises reagents (e.g., ligase) for generating a circularized template using the extended oligonucleotide. In some embodiments, the circularized template comprises at least a portion of the extended oligonucleotide. In some embodiments, the circularized template comprises a sequence of the extended oligonucleotide or a sequence complementary to the extended oligonucleotide. In some embodiments, the biological sample is a cell or tissue sample comprising cells or cellular components. In some embodiments, the biological sample is crosslinked and / or embedded in a matrix.
[0043] Provided herein is a kit, comprising: a nucleic acid strand, wherein the nucleic acid strand hybridizes to a cleavage region in a ribonucleic acid (RNA) template; an RNase H for cleaving the RNA template; a reverse transcriptase for generating an extended oligonucleotide; and an oligonucleotide configured to bind to the RNA template. In some embodiments, the kit further comprises a ligase for generating a circularized template and a polymerase for performing rolling circle amplification (RCA). In some embodiments, the kit further comprises a kinase. In some embodiments, the kit further comprises a plurality of free nucleotides that comprise different nucleic acid bases.
[0044] In some embodiments, the ligase for generating the circularized template is a CircLigase. In some embodiments, the RNase H comprises an RNase Hl and / or an RNAse H2. In some embodiments, the nucleic acid strand is about 10 to about 22 nucleotides in length, or about 15 to about 20 nucleotides in length. In some embodiments, the oligonucleotide comprises a primer binding sequence. In some embodiments, the oligonucleotide is complementary to a sequence 3’ to a region of interest in the RNA template. In some embodiments, the oligonucleotide comprises a poly-T sequence or a randomer. In some embodiments, the polymerase for performing rolling circle amplification is a Phi29 polymerase.
[0045] In some embodiments, the kit comprises a plurality of a sequencing primer, a plurality of detectably labeled nucleotides, and a polymerase.
[0046] Provided herein is a system, comprising: a biological sample comprising a ribonucleic acid (RNA) template; a nucleic acid strand, wherein the nucleic acid strand hybridizes to a cleavage region in the RNA template; an RNase H for cleaving the RNA template; a reverse transcriptase for generating an extended oligonucleotide; and an oligonucleotide configured to bind to the RNA template. In some embodiments, the system13MOFO-358213026202412023740 further comprises a ligase for generating a circularized template; and a polymerase for performing rolling circle amplification (RCA). In some instances, the biological sample is a cell or tissue sample provided on a solid support. In some embodiments, the system further comprises a kinase. In some embodiments, the system further comprises a plurality of free nucleotides that comprise different nucleic acid bases.
[0047] In some embodiments, the ligase for generating the circularized template is a CircLigase. In some embodiments, the RNase H comprises an RNase Hl and / or an RNase H2. In some embodiments, the nucleic acid strand is about 10 to about 22 nucleotides in length, or about 15 to about 20 nucleotides in length. In some embodiments, the oligonucleotide comprises a primer binding sequence. In some instances, the oligonucleotide is complementary to a sequence 3’ to the region of interest in the RNA template. In some instances, the oligonucleotide comprises a poly-T sequence or a randomer. In some embodiments, the polymerase for performing rolling circle amplification is a Phi29 polymerase.
[0048] In some embodiments, the system comprises reagents for performing sequencing by ligation, sequencing by synthesis, sequencing by binding, sequencing by avidity, or a combination thereof. In some embodiments, the system comprises a plurality of a sequencing primer, a plurality of detectably labeled nucleotides, and a polymerase. In some embodiments, the system comprises an optical detection system configured to detect a sequence of or associated with the RNA template or a complement thereof.BRIEF DESCRIPTION OF THE DRAWINGS
[0049] The drawings illustrate certain features and advantages of this disclosure. These embodiments are not intended to limit the scope of the appended claims in any manner.
[0050] FIG. 1 depicts an example workflow for the control of a reverse transcription reaction up to a terminating nucleotide to generate an extended oligonucleotide that is circularized for rolling circle amplification (RCA).
[0051] FIG. 2 depicts an example workflow for a reverse transcription reaction to generate an extended oligonucleotide with an incorporated ribonucleotide, inosine or uracil, cleaving the extended oligonucleotide for downstream processing and rolling circle amplification (RCA).14MOFO-358213026202412023740
[0052] FIG. 3 depicts an example workflow for cleaving an RNA template followed by a reverse transcription reaction to generate an extended oligonucleotide for downstream rolling circle amplification (RCA).
[0053] FIG. 4 depicts an example workflow for cleaving an RNA template followed by a reverse transcription reaction to generate an extended oligonucleotide associated with a spatial array or a barcode carrying bead.
[0054] FIG. 5 depicts an example workflow for using a plurality (e.g., a pool) of free nucleotides lacking nucleotides of a particular base in a reverse transcription reaction to generate an extended oligonucleotide.DETAILED DESCRIPTION
[0055] All publications, comprising patent documents, scientific articles and databases, referred to in this application are incorporated by reference in their entirety for all purposes to the same extent as if each individual publication were individually incorporated by reference. If a definition set forth herein is contrary to or otherwise inconsistent with a definition set forth in the patents, applications, published applications and other publications that are herein incorporated by reference, the definition set forth herein prevails over the definition that is incorporated herein by reference.
[0056] The section headings used herein are for organizational purposes only and are not to be construed as limiting the subject matter described.I. Overview
[0057] Profiling of analytes in cells and tissue samples involve accurate and reliable methods for nucleic acid amplification. Certain existing analyte targeting methods can yield products that may not be ideal for downstream assay processing and detection. RCA-based detection methods provide a powerful tool for detection of analytes at their relative spatial locations (e.g., in situ) in biological samples. However, an RCA-based assay that relies on probe hybridization for in situ analysis may be limited for whole transcriptome analysis in an unbiased manner. In some aspects, a non-targeted approach is desirable for sequencing in situ. If a reverse transcription reaction is performed using an RNA as template, the reaction continues until the template ends and / or other reagents are exhausted. There is a need for a workflow that allows short fragments to be generated. In some aspects, there is a need for a library preparation workflow suitable for de novo sequencing of transcripts in situ. In some aspects, there is a need15MOFO-358213026202412023740 for an unbiased (e.g., whole transcriptome) library preparation workflow suitable for in situ sequencing. In some embodiments, the detection is performed using in situ sequencing (e.g., base-by-base) sequencing. In some embodiments, the detection is performed by imaging.
[0058] In some embodiments, to detect a region of interest positioned within an RNA template, an amplification product associated with the RNA template is generated and detected. In some cases, a plurality of amplification products are generated using a plurality of RNA template molecules (e.g., different mRNAs). In some cases, extension products generated using RNA templates of various lengths may result in variability in the size of the generated amplification products. In some cases, various sizes of generated amplification products produce signals of different brightness levels when detected. In some cases, some amplification products result in dimmer signals for detection and limitations in the detection (e.g., by the camera) result in loss of signal from dimmer amplification products. In some cases, dimmer signals in close proximity to brighter signals are challenging to detect and may be lost. In some embodiments, controlling the RNA template length provides an advantage and allows for sensitivity gain by generating more uniform amplification products (e.g., rolling circle amplification products) in size and / or brightness for detection. In comparison to extension workflows where the length of the RNA template is not controlled, the provided methods mitigate problems caused by over extension or under extension which may affect assay sensitivity. In some cases, when generating circularized DNA in situ (e.g., a tissue section), there is a length bias that favors circularization of shorter (e.g., 15-100 nucleotide) substrates. In some cases, improved methods for limiting the length of reverse transcribed products to reduce the effect of the circularization bias are needed. Improved methods for generating a library of nucleic acids for in situ analysis (e.g., sequencing) are needed.
[0059] Provided herein are methods for nucleic acid processing comprising extending an oligonucleotide hybridized to a RNA template using a reverse transcriptase and a plurality of free nucleotides to generate an extended oligonucleotide. In some embodiments, a plurality of different extended oligonucleotides are generated corresponding to a plurality of different RNA templates (e.g., mRNAs in a cell or tissue sample). In some aspects, provided are methods for generation of a library of nucleic acids for sequencing in situ. In some aspects, provided are methods for generation of a library of nucleic acids for processing (e.g., capture and barcoding) and analysis.16MOFO-358213026202412023740
[0060] Provided herein are methods for nucleic acid processing comprising: hybridizing an oligonucleotide to a RNA template; extending the oligonucleotide using a reverse transcriptase and a plurality of free nucleotides to generate an extended oligonucleotide, wherein the plurality of free nucleotides comprises a terminating nucleotide that prevents further extension of the extended oligonucleotide; digesting the RNA template; and using the extended oligonucleotide to generate a circularized template.
[0061] Provided herein are methods for nucleic acid processing comprising: hybridizing an oligonucleotide to a RNA template; extending the oligonucleotide using a reverse transcriptase and a plurality of free nucleotides to generate an extended oligonucleotide, wherein the plurality of free nucleotides comprises a ribonucleotide, an inosine, or a uracil; digesting the RNA template and the extended oligonucleotide; and using the cleaved extended sequence to generate a circularized template. In some aspects, before performing ligation to generate the circularized template which requires a 3 ’OH and 5’ monophosphate for the ligase, the extended sequence is treated with a kinase. In some embodiments, the extended oligonucleotide is processed (e.g., treated to remove any terminating groups or modifications) to prepare the molecule for ligation and / or downstream amplification. In some aspects, the circularized template comprising the extended sequence is used as a template to generate a rolling circle amplification (RCA) product.
[0062] Provided herein are methods for nucleic acid processing comprising: contacting a biological sample with a nucleic acid strand, wherein the nucleic acid strand hybridizes to a cleavage region in a ribonucleic acid (RNA) template, and cleaving the RNA template in the cleavage region; hybridizing an oligonucleotide to the cleaved RNA template and extending the oligonucleotide using a reverse transcriptase and a plurality of free nucleotides to generate an extended oligonucleotide; digesting the RNA template; and using the extended oligonucleotide to generate a circularized template. In some aspects, the circularized template comprising the extended sequence is used as a template to generate a rolling circle amplification (RCA) product.
[0063] Provided herein are methods for nucleic acid processing comprising: contacting a biological sample with a nucleic acid comprising one or more catalytic domains (e.g., a DNAzyme or an RNAzyme), wherein the nucleic acid hybridizes to a cleavage region in a ribonucleic acid (RNA) template, and cleaving the RNA template in the cleavage region;17MOFO-358213026202412023740 hybridizing an oligonucleotide to the cleaved RNA template and extending the oligonucleotide using a reverse transcriptase and a plurality of free nucleotides to generate an extended oligonucleotide; digesting the RNA template; and using the extended oligonucleotide to generate a circularized template. In some aspects, the circularized template comprising the extended sequence is used as a template to generate a rolling circle amplification (RCA) product. In some embodiments, the nucleic acid comprising one or more catalytic domains comprises a binding domain. In some instances, the one or more catalytic domains comprise nucleic acid hydrolysis activity.
[0064] Provided herein are methods for nucleic acid processing comprising: contacting a biological sample with a guide nucleic acid and an RNA-cutting enzyme, wherein the guide nucleic acid hybridizes to a cleavage region in a ribonucleic acid (RNA) template, and cleaving the RNA template in the cleavage region; hybridizing an oligonucleotide to the cleaved RNA template and extending the oligonucleotide using a reverse transcriptase and a plurality of free nucleotides to generate an extended oligonucleotide; digesting the RNA template; and using the extended oligonucleotide to generate a circularized template. In some aspects, the circularized template comprising the extended sequence is used as a template to generate a rolling circle amplification (RCA) product. In some embodiments, the RNA-cutting enzyme is a CRISPR effector protein. In some embodiments, the RNA-cutting enzyme is an Argonaute protein.
[0065] Provided herein are methods for nucleic acid processing comprising: contacting the biological sample with a nucleic acid strand, wherein the nucleic acid strand hybridizes to a cleavage region in a ribonucleic acid (RNA) template, and cleaving the RNA template in the cleavage region; hybridizing an oligonucleotide to the cleaved RNA template and extending the oligonucleotide using a reverse transcriptase and a plurality of free nucleotides to generate an extended oligonucleotide; digesting the RNA template; and using the extended oligonucleotide to generate a barcoded nucleic acid molecule. In some aspects, the barcoded nucleic acid molecule is generated in a partition (e.g., as described in Section III.B). In some aspects, the barcoded nucleic acid molecule is generated or captured on a spatial array (e.g., as described in Section III.C).
[0066] Provided herein is a method comprising contacting a biological sample with a nucleic acid strand, wherein the nucleic acid strand hybridizes to a cleavage region in a18MOFO-358213026202412023740 ribonucleic acid (RNA) template, and cleaving the RNA template in the cleavage region; hybridizing an oligonucleotide to the cleaved RNA template and extending the oligonucleotide using a reverse transcriptase and a plurality of free nucleotides to generate an extended oligonucleotide, wherein the oligonucleotide is attached to a substrate and comprises a barcode sequence. In some instances, the generated extended oligonucleotide comprises a sequence of the region of interest in the RNA template or a complement thereof, and the barcode sequence or complement thereof. In some embodiments, the barcode sequence or complement thereof is a spatial barcode. In some embodiments, the barcode sequence is used to identify the cell of origin of the RNA template.
[0067] Provided herein is a method comprising contacting a biological sample with a nucleic acid strand, wherein the nucleic acid strand hybridizes to a cleavage region in a ribonucleic acid (RNA) template, and cleaving the RNA template in the cleavage region; hybridizing an oligonucleotide to the cleaved RNA template and extending the oligonucleotide using a reverse transcriptase and a plurality of free nucleotides to generate an extended oligonucleotide; contacting the extended oligonucleotide or a derivative thereof with a nucleic acid barcode molecule comprising a capture sequence and a barcode sequence; and generating a barcoded nucleic acid molecule. In some instances, the generated barcoded nucleic acid molecule comprises a sequence of the region of interest in the RNA template or a complement thereof, and the barcode sequence or complement thereof. In some embodiments, the barcode sequence or complement thereof is a spatial barcode. In some embodiments, the barcode sequence is used to identify the cell of origin of the RNA template.II. Methods for Library Preparation
[0068] In some aspects, a non-targeted approach is desirable for generating and processing a library of nucleic acids for unbiased whole transcriptome sequencing in situ. If a reverse transcription reaction is performed using an RNA as template, the reaction continues until the template ends and / or other reagents are exhausted. Provided herein is a workflow for generating a fragment with a controlled length using a RNA template. Provided herein is a workflow for generating short fragments using a template RNA by stopping extension using terminating nucleotides or by cleavage. Provided herein is a workflow for generating a fragment with a length of 50-200 nucleotides using a RNA template. Provided herein is a workflow for generating a fragment with a length of less than 50 nucleotides using a RNA template. Provided19MOFO-358213026202412023740 herein is a workflow for generating a library of short fragments (e.g., with a length of no more than 50-200 nucleotides in length) using a plurality of cleaved RNA templates. In some aspects, provided herein is a method for processing nucleic acids for an unbiased whole transcriptome library preparation workflow suitable for in situ sequencing. In some aspects, provided herein is a library preparation workflow comprising processing nucleic acids for downstream processing (e.g., barcoding) and analysis (e.g., sequencing).
[0069] In some embodiments, provided are nucleic acids (e.g., RNA templates, oligonucleotides, nucleic acid strands) for binding to other nucleic acids. In some instances, the nucleic acids bind via hybridization, typically by Watson-Crick base pairing, such as DNA, RNA, LNA, PNA, etc., depending on the application. In some embodiments, nucleic acids (e.g., RNA templates, oligonucleotides, nucleic acid strands) are able to bind to at least a portion of another nucleic acid. In some embodiments, the nucleic acids (e.g., RNA templates, oligonucleotides, nucleic acid strands) bind to a specific target sequence.
[0070] Provided herein are methods for nucleic acid processing comprising: hybridizing an oligonucleotide to a RNA template; extending the oligonucleotide using a reverse transcriptase and a plurality of free nucleotides to generate an extended oligonucleotide, wherein the plurality of free nucleotides comprises a terminating nucleotide that prevents further extension of the extended oligonucleotide; digesting the RNA template; and using the extended oligonucleotide to generate a circularized template. In some aspects, the oligonucleotide comprises a sequencing primer binding sequence. In some embodiments, the extension is initiated by a primer that binds to the sequencing primer binding sequence on the oligonucleotide. In some aspects, as illustrated in FIG. 1, the plurality of free nucleotides comprises a low amount of a terminating nucleotide (e.g., ddNTPs or inverted dT) that causes extension to stop. In some aspects, the amount of terminating nucleotide in the plurality of free nucleotides can be tested and tuned. In some aspects, the amount of terminating nucleotide is tuned to allow for the extension to continue for generation of an extended oligonucleotide that is about 50 to 200 nucleotides in total length. In some embodiments, the terminating nucleotide is cleavable from the extended oligonucleotide. In some embodiments, the terminating nucleotide comprises a terminating modification that is reversible or removable.
[0071] In some aspects, provided herein is a library preparation workflow suitable for unbiased analysis of sequences associated with variable or unknown analyte sequences. In some20MOFO-358213026202412023740 aspects, provided herein is a library preparation workflow suitable for unbiased analysis of sequences associated with a variable sequence (e.g., SNV detection), perturbation agents introduced to a cell (e.g., CRISPR barcode scanning), or a sequence of or associated with an immune molecule(e.g., for VDJ detection). In some embodiments, the method comprises hybridizing an oligonucleotide to a template, wherein the template is an endogenous sequence (e.g., RNA transcript). In some embodiments, an oligonucleotide is hybridized to a template, wherein the template is an exogenous sequence introduced to the biological sample (e.g., introduced into cells). In some embodiments, an exogenous molecule is introduced into the biological sample (e.g., a viral transcript or an exogenous library of nucleic acid constructs is introduced to the biological sample and the exogenous molecule or a derivative thereof is hybridized by the oligonucleotide. In some embodiments, an exogenous analyte or a nucleic acid molecule associated with an exogenous analyte is used as a template for extending the oligonucleotide to generate the extended oligonucleotide.
[0072] In some embodiments, the RNA template used for extending the oligonucleotide comprises a region of interest. In some embodiments, the RNA template used for extending the oligonucleotide comprises a nucleotide variation, a nucleotide polymorphism, a mutation, a substitution, an insertion, a deletion, a translocation, a duplication, an inversion, a rearrangement and / or a repetitive sequence, for identifying a variant sequence among a plurality of different sequences in situ in a biological sample. In some embodiments, a nucleotide variation, a nucleotide polymorphism, a mutation, a substitution, an insertion, a deletion, a translocation, a duplication, an inversion, a rearrangement and / or a repetitive sequence is in the region of interest. In some embodiments, the template used for extending the oligonucleotide comprises a variant sequence of a single nucleotide, for instance, a single nucleotide variation (SNV), a single nucleotide polymorphism (SNP), a point mutation, a single nucleotide substitution, a single nucleotide insertion, or a single nucleotide deletion. In some embodiments, the region of interest comprises a variant sequence of a single nucleotide, for instance, a single nucleotide variation (SNV), a single nucleotide polymorphism (SNP), a point mutation, a single nucleotide substitution, a single nucleotide insertion, or a single nucleotide deletion. In some embodiments, the template used for extending the oligonucleotide comprises one or more exonexon boundaries. In some embodiments, the template used for extending the oligonucleotide21MOFO-358213026202412023740 comprises a sequence of an exon-exon boundary. In some embodiments, the region of interest comprises a sequence of an exon-exon boundary.
[0073] In some embodiments, the template (e.g., RNA template) used for extending the oligonucleotide comprises one or more hotspots for mutation. In some embodiments, the template used for extending the oligonucleotide comprises a variant sequence among a plurality of different variant sequences. In some embodiments, the template used for extending the oligonucleotide comprises a variant sequence among a plurality of different variant sequences.
[0074] In some aspects, the template used for extending the oligonucleotide comprises a sequence of or associated with an exogenous construct or a complement thereof.
[0075] In some embodiments, the template used for extending the oligonucleotide comprises a sequence of or associated with an immune molecule. In TCR and BCR RNA transcripts, the V(D)J sequences are 5’ to the constant region exon(s) and the 3’ poly(A) tail of the transcript. The number of V(D)J transcripts of a particular V(D)J join sequence in a sample comprising T cells or B cells of various antigen specificities can be low, and sequence information in particular V(D)J joins can be lost and become unavailable for subsequent in situ detection. In some embodiments, the present disclosure provides methods for high-throughput profiling of V(D)J transcripts in a large number of clonal T cell populations comprising TCRs with varying antigenic specificities. The methods and compositions disclosed herein may be used in research, diagnostics, and drug target discovery. Analyzing the spatial distribution of V(D)J transcripts in situ in various tissues could be used for development of therapeutic and / or prophylactic agents, e.g., TCR therapeutic treatment modalities and / or anti-disease vaccination.
[0076] Provided herein is a method for analysis of immune molecule sequences by contacting the biological sample with an oligonucleotide and extending the oligonucleotide using a sequence of an immune molecule as template to generate an extended oligonucleotide. In some aspects, the provided methods for immune molecule analysis allows for sensitive detection even if the number of particular transcripts are low in the biological sample. In some embodiments, the template used for extending the oligonucleotide comprises a sequence of an immune molecule. For example, the template used for extending the oligonucleotide is an antigen receptor transcript. In some cases, the antigen receptor transcript is a T cell receptor (TCR) transcript, optionally wherein the TCR transcript comprises a TCRa VJ join, a TCRP VDJ join, a TCRy VJ join, or a TCRS VDJ join. In some cases, the antigen receptor transcript is an22MOFO-358213026202412023740 immunoglobulin (Ig) transcript, optionally wherein the Ig transcript comprises an IgK VJ join, an IgZ. VJ join, or an IgH VDJ join. In some embodiments, the methods are used for identifying multiple different antigen receptor transcripts present at a plurality of locations in the biological sample. In some embodiments, amplicons (e.g., RCA products) comprising V(D)J sequences or complements thereof are detected in situ using sequencing.
[0077] In some embodiments, the TCR transcript disclosed herein comprises a TCRa VJ join. In some embodiments, the TCR transcript disclosed herein comprises a TCRP VDJ join. In some embodiments, the TCR transcript disclosed herein comprises a TCRy VJ join. In some embodiments, the TCR transcript disclosed herein comprises a TCR5 VDJ join.
[0078] In some embodiments, disclosed herein is a method involving detecting a sequence of a V(D)J transcript in a biological sample. In some embodiments, V(D)J sequences include those in V(D)J transcripts comprising V(D)J joins. In some embodiments, V(D)J sequences are used as the template used for extending the oligonucleotide and V(D)J sequences or a complement thereof are present in amplification products of V(D)J transcripts.
[0079] In some aspects, methods provided herein further comprise generating rolling circle amplification (RCA) products of the circularized template generated from the oligonucleotide extended using a sequence of an immune molecule as template and corresponding products thereof (e.g., RCA products) are detected for analyzing the spatial organization of V(D)J sequences in samples (e.g., tissues such as tumors comprising infiltrating immune cells). Such insights can be crucial to understanding disease development and establishing new treatment strategies.
[0080] In some embodiments, the template used for extending the oligonucleotide comprises a sequence associated with a constant region of an antibody or a fragment thereof. In some embodiments, the template used for extending the oligonucleotide comprises encoding a constant region of an immune cell receptor. In some embodiments, the template used for extending the oligonucleotide comprises a sequence encoding a constant region of a B cell receptor. In some embodiments, the template used for extending the oligonucleotide comprises a sequence encoding a constant region of a T cell receptor.
[0081] Provided herein is a method for analysis of one or more perturbation agents introduced to a cell by contacting the biological sample with an oligonucleotide and extending the oligonucleotide using a sequence of the perturbation agent or a corresponding molecule as23MOFO-358213026202412023740 template to generate an extended oligonucleotide. In some aspects, the biological sample is contacted with a library of perturbation agents. In some aspects, the assays described herein are used for detecting CRISPR guides, e.g., guide RNAs (gRNAs). In some embodiments, the assays described herein are used for detecting perturbations introduced by CRISPR libraries and / or cellular RNA transcripts. In some cases, the template used for extending the oligonucleotide comprise a sequence of a CRISPR guide RNA.
[0082] In some embodiments, the template used for extending the oligonucleotide comprise a sequence of or associated with a perturbation agent introduced to the biological sample before the oligonucleotide is introduced. In some embodiments, a CRISPR molecule (e.g., a CRISPR RNA), a nucleic acid molecule edited using the CRISPR molecule, and / or a precursor or derivative thereof is detected. In some aspects, the template used for extending the oligonucleotide comprises a sequence of a CRISPR molecule (e.g., a CRISPR RNA), a nucleic acid molecule edited using the CRISPR molecule, and / or a precursor or derivative thereof. In some instances, the template used for extending the oligonucleotide is an RNA molecule derived from an exogenously introduced nucleic acid molecule. In some embodiments, the exogenously introduced nucleic acid molecule is an RNA derived from a plasmid, an integrated DNA sequence (e.g. using viral transduction in a cell), a gRNA from a CRISPR genetic element, etc. In some embodiments, the perturbation agent comprises a spacer sequence that is an element (e.g., about 20 nucleotides) that can be found as a component of gRNA.
[0083] In some embodiments, the oligonucleotide binds to or hybridizes to a protospacer sequence or hybridization regions upstream of a protospacer sequence, or a complement thereof. In some embodiments, the oligonucleotide binds to a conserved region of the guide RNA (e.g., a common sequence shared by a plurality of different guide RNAs). In some aspects, CRISPR libraries are generated in cells of a biological sample. In some aspects, a CRISPR library may comprise hundreds, thousands, or tens of thousands of different spacer sequences. In some embodiments, an oligonucleotide hybridizes to the complement or reverse complement of a guide RNA spacer sequence. In some examples, a nucleic acid molecule to be analyzed is introduced and / or delivered into a cell or a cell constituent (e.g., a nucleus of a cell) using any of a variety of techniques.
[0084] In some embodiments, the template used for extending the oligonucleotide is a transcript comprising a unique barcode specific to the perturbation agent. In some24MOFO-358213026202412023740 embodiments, the template used for extending the oligonucleotide is a transcript comprising a unique barcode specific to a guide RNA. In some embodiments, the template used for extending the oligonucleotide is a transcript comprising a guide RNA sequence. In some instances, a guide RNA and guide RNA barcode is expressed from the same vector and the barcode or a complement thereof is used as template to generate an extended probe comprising a gap filled sequence. For example, perturbation agents are described in U.S. Patent Application Publication No. 2021 / 0171938, which is incorporated herein by reference in its entirety.
[0085] In some embodiments, the template used for extending the oligonucleotide is a transcript from a exogenous source (e.g., a virus).
[0086] Provided herein are methods for nucleic acid processing comprising: hybridizing an oligonucleotide to a RNA template; extending the oligonucleotide using a reverse transcriptase and a plurality of free nucleotides to generate an extended oligonucleotide, wherein the plurality of free nucleotides comprises a ribonucleotide, an inosine, or a uracil; digesting the RNA template and cleaving the extended oligonucleotide; and using the cleaved extended sequence to generate a circularized template. In some aspects, the plurality of free nucleotides comprises a low amount of a cleavable nucleotide (e.g., a ribonucleotide, an inosine, or a uracil) to facilitate cleavage of the extended oligonucleotide as illustrated in FIG. 2. In some aspects, the amount of cleavable nucleotide in the plurality of free nucleotides can be tested and tuned. In some aspects, the amount of cleavable nucleotide is tuned to allow for generation of an extended oligonucleotide that is about 50 to 200 nucleotides in total length after cleavage of the extended oligonucleotide at the incorporated cleavable nucleotide.
[0087] In some embodiments, provided herein is a method of analyzing a biological sample, comprising: (a) hybridizing an oligonucleotide to a RNA template in a biological sample; (b) extending the oligonucleotide using a reverse transcriptase and a plurality of free nucleotides to generate an extended oligonucleotide, wherein the plurality of free nucleotides comprises a reversible terminator nucleotide and incorporation of the reversible terminator nucleotide stops further extension of the extended oligonucleotide; (c) removing the terminating group from the reversible terminator nucleotide incorporated into the extended oligonucleotide; (d) digesting the RNA template; (e) ligating the 5’ and 3’ ends of the extended oligonucleotide to generate a circularized template; (f) performing rolling circle amplification of the circularized template to generate a rolling circle amplification product (RCP); and; (g) detecting a sequence25MOFO-358213026202412023740 of the RCP in the biological sample using in situ sequencing by synthesis (SBS) in the biological sample, thereby detecting a sequence of the RCP or a complement thereof at a location in the biological sample.
[0088] Provided herein are methods for nucleic acid processing comprising: contacting a biological sample with a nucleic acid strand, wherein the nucleic acid strand hybridizes to a cleavage region in a ribonucleic acid (RNA) template, and cleaving the RNA template in the cleavage region; hybridizing an oligonucleotide to the cleaved RNA template and extending the oligonucleotide using a reverse transcriptase and a plurality of free nucleotides to generate an extended oligonucleotide. In some instances, the method comprises digesting the RNA template after generating the extended oligonucleotide. In some embodiments, the extended oligonucleotide is used to generate a circularized template. In some aspects, the circularized template comprising the extended sequence is used as a template to generate a rolling circle amplification (RCA) product. In some aspects, a cleaved RNA template is used for performing extension of an oligonucleotide as illustrated in FIG. 3. In some embodiments, the extended oligonucleotide is further processed. In some embodiments, the further processing comprises additional hybridization, extension, and / or polymerization reactions. In some embodiments, the further processing comprises contacting the extended oligonucleotide or a derivative thereof (e.g., a oligonucleotide or polynucleotide generated from the extended oligonucleotide) with a nucleic acid barcode molecule (e.g., coupled to a support). In some embodiments, the further processing comprises using the extended oligonucleotide to generate a barcoded nucleic acid molecule. In some instances, cleavage occurs in the biological sample at the site of a hybridized nucleic acid strand and RNA template. In some aspects, the nucleic acid strand determines the site of cleavage.
[0089] In some embodiments, a nucleic acid strand is used to guide cleavage of the RNA template. In some embodiments, the nucleic acid strand hybridizes to a cleavage region in a ribonucleic acid (RNA) template at a position 5’ to a region of interest in the RNA template. In some instances, an RNA template comprises from 5’ to 3’: a cleavage region - a region of interest - a sequence complementary to an oligonucleotide. In some instances, an RNA template comprises from 5’ to 3’: a cleavage region - a region of interest - a polyA tail.
[0090] In some aspects, a cleaved RNA template is used for performing extension of an oligonucleotide. In some instances, the cleaved RNA template is generated by contacting the26MOFO-358213026202412023740 biological sample with an RNase H. In some embodiments, the cleaved RNA template is generated by using RNase H and a nucleic acid strand to provide a DNA-RNA duplex upon hybridization to the RNA template, for cleavage of the RNA template by the RNase H. As illustrated in FIG. 3, in some embodiments, the method comprises hybridizing a nucleic acid strand to a cleavage region in a RNA template to form a DNA-RNA duplex. RNase H can then cleave the RNA template within the cleavage region. In some embodiments, one or more washes are then performed, and the biological sample is contacted with an oligonucleotide. In some instances, the oligonucleotide is extended using a reverse transcriptase and a plurality of free nucleotides using the cleaved RNA template to generate an extended oligonucleotide.
[0091] In some instances, the cleaved RNA template is generated by contacting the biological sample with a nucleic acid comprising one or more catalytic domains (e.g., a nucleic acid enzyme or NAzyme, for example a DNAzyme or an RNAzyme). In some embodiments, the biological sample is contacted with the nucleic acid comprising one or more catalytic domains before contacting the sample with the oligonucleotide for extension. In some embodiments, the nucleic acid comprising one or more catalytic domains comprises a binding domain. In some instances, the binding domain of the nucleic acid enzyme in configured to bind to a cleavage region that is 5’ to a region of interest in the RNA template. In some embodiments, the nucleic acid comprising one or more catalytic domains is a DNA molecule. In some embodiments, the nucleic acid comprising one or more catalytic domains is a RNA molecule. In some instances, the one or more catalytic domains comprise nucleic acid hydrolysis activity, wherein the nucleic acid hybridizes to a cleavage region in a ribonucleic acid (RNA) template, and cleaves the RNA template in the cleavage region. In some embodiments, the nucleic acid comprising one or more catalytic domains is a DNA molecule capable of catalyzing a chemical reaction. In some instances, the nucleic acid comprising one or more catalytic domains (e.g., DNAzyme or RNAzyme) catalyzes nucleic acid cleavage (such as oxidative cleavage or hydrolytic cleavage, for example, phosphodiester hydrolytic cleavage), nucleoside excision, phosphorylation (or dephosphorylation), ligation, or other reactions. In some instances, the nucleic acid comprising one or more catalytic domains (e.g., DNAzyme or RNAzyme) comprises one or more unnatural chemical modifications. In some instances, the cleaved RNA template is generated by contacting the biological sample with DNAzymes (catalytic DNA or deoxyribozymes). In some instances, the cleaved RNA template is generated by contacting the biological sample with a RNAzyme27MOFO-358213026202412023740(e.g., a ribozyme). In some instances, the nucleic acid comprising one or more catalytic domains (e.g., DNAzyme or RNAzyme) is a natural or unnatural nucleic acid molecule. In some instances, the nucleic acid comprising one or more catalytic domains (e.g., DNAzyme or RNAzyme) is a modified nucleic acid molecule.
[0092] In some embodiments, a sample is contacted with a plurality of different nucleic acids comprising different catalytic domains for binding to different RNA templates. In some embodiments, the nucleic acid comprises a catalytic domain flanked by two binding domains (e.g., substrate binding arms). In some instances, the binding arms recognize and hybridize to a specific target sequence in a RNA template as described here. In some embodiments, the position of the binding arms are configured to position the catalytic core to cleave the RNA template at the cleavage region. In some instances, the catalytic domain is configured to cleave a phosphodiester bond within the RNA template. In some embodiments, the binding domain is each independently about 10 to about 30, about 15 to about 30, about 5 to about 25, about 10 to about 20, about 15 to about 25, about 5 to about 15, about 8 to about 18, about 10 to about 18, or about 15 to about 20 nucleotides in length. In some embodiments, the binding domain is each independently about 5 to about 25, about 5 to about 20, about 5 to about 15, about 5 to about 10, about 8 to about 18, about 10 to about 18, or about 15 to about 20 nucleotides in length. In some instances, the binding domain is each independently about 5 to about 15 nucleotides in length, or about 5 to about 20 nucleotides in length. In some aspects, the binding domain is designed such that upon hybridization to the cleavage region in the RNA template, the catalytic domain of the nucleic acid cleaves the RNA template at a position at least 5-8 nucleotides from the region of interest in the RNA template.
[0093] In some embodiments, the cleavage of the RNA template is performed in the presence of one or more metal ions. In some instances, the catalytic function of the nucleic acid comprising one or more catalytic domains (e.g., DNAzyme or RNAzyme) is dependent on one or more divalent metal ions. In some embodiments, the cleavage of the RNA template is performed in a buffer comprising at least one divalent cation (e.g., Fe2+, Co2+, Ni2+, Cu2+, Zn2+, Mg2+, Mn2+, or Ca2+). In some embodiments, the one or more divalent metal ions comprise magnesium (Mg2+), manganese (Mn2+), and / or calcium (Ca2+). In some embodiments, the cleavage of the RNA template is performed in a buffer comprising Mg2+and / or Mn2+. Provided herein are methods for nucleic acid processing comprising contacting a biological sample with a nucleic28MOFO-358213026202412023740 acid comprising one or more catalytic domains (e.g., a DNAzyme or an RNAzyme) and a divalent cation, wherein the nucleic acid hybridizes to a cleavage region in a ribonucleic acid (RNA) template, and cleaving the RNA template in the cleavage region.
[0094] In some embodiments, the biological sample is contacted with the nucleic acid strand and with the RNase H simultaneously or sequentially (in either order) before contacting the sample with the oligonucleotide for extension. In some embodiments, the biological sample is contacted with the nucleic acid strand and with the RNase H before contacting the sample with the oligonucleotide for extension. In some embodiments, the method comprises washing the biological sample after contacting the biological sample with the RNase H and before contacting the biological sample with the oligonucleotide for extension. In some embodiments, RNase inactivating agents or inhibitors can be added to the sample after cleaving the RNA template. In some aspects, the present application provides designs for nucleic acid strands capable of forming DNA-RNA duplexes for RNase H cutting of an RNA template. In some instances, provided herein are nucleic acid strands configured for controlling the length of an extended oligonucleotide generated using a cleaved RNA template.
[0095] In some embodiments, hybridizing a nucleic acid strand to a cleavage region in a ribonucleic acid (RNA) template and cleaving the RNA template in the cleavage region occurs in the same mixture. In some embodiments, hybridizing a nucleic acid strand to a cleavage region in a ribonucleic acid (RNA) template and cleaving the RNA template in the cleavage region occurs in the same reaction.
[0096] In some embodiments, the nucleic acid strand is single-stranded. In some embodiments, the nucleic acid strand comprises at least 4, 5, 6, 7, or 8 consecutive deoxyribonucleotides. In some embodiments, the nucleic acid strand is a deoxyribonucleic acid (DNA) oligonucleotide. In some cases, the nucleic acid strand is a single- stranded DNA (ssDNA) oligonucleotide.
[0097] In some embodiments, the nucleic acid strand is designed to hybridize to a cleavage region in an RNA template. In some instances, the cleavage region is 5’ to a region of interest in the RNA template (e.g., as illustrated in FIG. 3). In some instances, the cleavage region is at least 1, at least 2, at least 3, at least 4, at least 5, at least 10, at least 15, or at least 20 nucleotides 5’ to a region of interest in the RNA template. In some instances, the cleavage region is 5’ to a region of interest and a sequence for hybridizing an oligonucleotide in the RNA29MOFO-358213026202412023740 template. In some embodiments, the cleavage region is about 10 to about 30, about 15 to about 30, about 5 to about 25, about 10 to about 20, about 15 to about 25, about 5 to about 15, about 8 to about 18, about 10 to about 18, or about 15 to about 20 nucleotides in length. In some embodiments, the nucleic acid strand is about 10 to about 30, about 15 to about 30, about 5 to about 25, about 10 to about 20, about 15 to about 25, about 5 to about 15, about 8 to about 18, about 10 to about 18, or about 15 to about 20 nucleotides in length. In some instances, the cleavage region is about 10 to about 22 nucleotides in length, or about 15 to about 20 nucleotides in length. In some aspects, the nucleic acid strand is designed such that upon hybridization to the cleavage region in the RNA template, the RNase H cleaves the RNA template at a position at least 5-8 nucleotides from the region of interest in the RNA template.
[0098] In some embodiments, cleaving the RNA template in the cleavage region comprises contacting the biological sample with an RNase H. Any suitable RNase H for cleaving RNA in an nucleic acid duplex (e.g., within the oligonucleotide hybridization region hybridized to the nucleic acid oligonucleotide) can be used. The RNase H enzyme and its family of enzymes include two classes, type 1 and type 2 RNase H based on the difference in their amino acid sequence. Type 1 RNases H include prokaryotic and eukaryotic RNases Hl and retroviral RNase H. Type 2 RNases H include prokaryotic and eukaryotic RNases H2 and bacterial RNase H3. These RNases H exist in a monomeric form, except for eukaryotic RNases H2, which exist in a heterotrimeric form. All of these enzymes share the characteristic that they are able to cleave the RNA component of an RNA:DNA heteroduplex or within a DNA:DNA duplex containing RNA base(s) within one or both of the strands. The cleaved product yields a free 3'-OH for both classes of RNase H. In some embodiments, RNase Hl requires more than a single RNA base within an RNA:DNA duplex for optimal activity. In some embodiments, RNase HII requires only a single RNA base in an RNA:DNA duplex.
[0099] In some embodiments, the RNase H enzyme comprises RNase Hl or RNase H3. An example of an RNase H enzyme includes E. coli RNase HII. In some embodiments, the RNase is an RNase HII. In some embodiments, the RNase is E. coli RNase H, which also cleaves a ribo-base when hybridized to DNA and leaves a 3 '-hydroxyl end. In some embodiments, the RNase H is a thermostable RNase H. In some embodiments, the RNase H is a mammalian RNase H, such as any of those described in U.S. Patent Application Publication No. 2005 / 0164234, the content of which is herein incorporated by reference in its entirety. For30MOFO-358213026202412023740 example, an enzyme with RNase HII characteristics has been purified to near homogeneity from human placenta (Frank et al., Nucleic Acids Res., 1994, 22, 5247-5254). This protein has a molecular weight of approximately 33 kDa and is active in a pH range of 6.5-10, with a pH optimum of 8.5-9. The enzyme requires Mg2+and is inhibited by Mn2+and n-ethyl maleimide. The products of cleavage reactions have 3' hydroxyl and 5' phosphate termini. In some embodiments, the cleaving is performed at a temperature below 60° C. (e.g. at room temperature or about 20-50° C„ 20-30°C„ 20-40°C„ 25-40°C„ 30-40°C„ 35-40°C„ or 40-50°C.). In some embodiments, the cleaving is performed by incubating the RNA template with RNase H at about 37°C for a duration of between about 10 minutes and about 2 hours, between about 10-120, between about 10-90, between about 10-60, between about 10-30, between about 20-90, between about 20-60, between about 20-30, or between about 30-60 minutes.
[0100] In some embodiments, the provided methods comprise contacting the RNA template (e.g., in a biological sample) with RNase H at a concentration of about at least 1 x IO"5, 1 x IO"4, 1 x 10"3, 1 x IO"2, 1 x 10"1, 1 U / pL, or higher. In some embodiments, the RNase H concentration is less than 1, 1 x 10"1, 1 x IO"2, 1 x 10"3, 1 x IO"4, 1 x IO"5U / pL, or less. In some embodiments, the RNase H concentration is between about 1 x IO"5and about 1 x IO"4, between about 1 x IO"4and about 1 x 10"3, between about 1 x 10"3and about 1 x IO"2, between about 1 x IO"2and about 1 x 10"1, between about 1 x 10"1and about 1, between about 1 x IO"5and about 1 x 10"3, between about 1 x IO"4and about 1 x IO"2, between about 1 x 10"3and about 1 x 10"1U / pL, or between about 1 x IO"2and about 1 U / pL. In some embodiments, the RNase H concentration is about 1 x ICT4, 3 x ICT4, 1 x 10’3, 3 x 10’3, 1 x IO’2, 3 x ICT2, or 1 x ICT1U / pL.
[0101] In some embodiments, the provided methods comprise contacting the RNA template (e.g., in a biological sample) with between about 0.5 enzyme units (U) and about 100 U of the RNase H. In some embodiments, the biological sample is contacted with between about 0.5 and about 100, between about 0.5 and about 80, between about 0.5 and about 60, between about 0.5 and about 50 , between about 0.5 and about 40, between about 0.5 and about 30, between about 0.5 and about 20, between about 0.5 and about 10, between about 0.5 and about 8, between about 0.5 and about 5 , between about 0.5 and about 5, between about 0.5 and about 3, between about 2 and about 10, between about 2 and about 8, between about 2 and about 6, between about 2 and about 5, between about 2 and about 4, between about 3 and about 8, between about 3 and about 6, between about 4 and about 8, between about 4 and about 6,31MOFO-358213026202412023740 between about 4 and about 5.5, between about 4.5 and about 5.5, between about 4.5 and about 6, between about 10 and about 100, between about 10 and about 50, between about 10 and about 30, between about 10 and about 20, between about 20 and about 100, between about 20 and about 50, between about 20 and about 40, between about 20 and about 30, between about 30 and about 100, between about 30 and about 80, or between about 30 and about 50 U of RNase H. In some embodiments, the amount of RNase H contacted with the RNA template is dependent on the amount of nucleic acid to be cleaved.
[0102] In some embodiments, the RNase H is incubated with the RNA template (e.g., in a biological sample) in a buffer comprising magnesium chloride. In some embodiments, the RNase H is incubated with the biological sample in a buffer comprising magnesium chloride, potassium chloride, dithiothreitol (DTT), and a buffering agent (e.g., Tris-HCL).
[0103] In some embodiments, the RNase H comprises an RNase Hl and / or an RNase H2. In some embodiments, the method comprises contacting the biological sample with an RNase Hl and an RNase H2. In some embodiments, the RNase H is RNase Hl. In some embodiments, the RNase H is an endoribonuclease that specifically hydrolyzes the phosphodiester bonds of RNA which is hybridized to DNA. In some embodiments, the RNase H does not digest single or double- stranded DNA. In some embodiments, the RNase H requires at least four contiguous bases of RNA for digestion of the RNA hybridized to DNA. In some embodiments, the RNase H does not digest a di-ribonucleotide-containing DNA sequence (e.g., one that is DNA-annealed). In some embodiments, the RNase H does not digest a monoribonucleotide-containing DNA sequence (e.g., one that is DNA-annealed).
[0104] In some embodiments, the RNase H is RNase H2. In some embodiments, the RNase H is an endoribonuclease that preferentially nicks 5' to one or more ribonucleotides (e.g., a single ribonucleotide, a diribonucleotide sequence, etc.) within the context of a DNA duplex, leaving 5' phosphate and 3' hydroxyl ends. In some embodiments, the RNase H nicks at multiple sites along an RNA portion hybridized to a DNA. In some embodiments, the RNase H digests a DNA-annealed di-ribonucleotide-containing DNA sequence, whereas the RNA- annealed di-ribonucleotide-containing DNA sequence is not digested. In some embodiments, the RNase H digests a DNA-annealed mono-ribonucleotide-containing DNA sequence, whereas RNA-annealed mono-ribonucleotide-containing DNA sequence is not digested.32MOFO-358213026202412023740
[0105] In some embodiments, the RNase H cleaves RNA in RNA-DNA duplexes. In some embodiments, the RNase H is a bacterial RNase H or analog or derivative thereof (e.g., Escherichia coli RNase H or an analog or derivative thereof). In some embodiments, the RNase H is a eukaryotic RNase H or analog or derivative thereof (e.g., human RNase H or an analog or derivative thereof). In some embodiments, the RNase H is a viral RNase H or analog or derivative thereof, such as an HIV-derived RNase H. RNase H sequence preferences are described, for example, in Kielpinski et al., “RNase H sequence preferences influence antisense oligonucleotide efficiency.” Nucleic Acids Res. 2017 Dec 15;45(22):12932-12944, the content of which is herein incorporated by reference in its entirety.
[0106] In some instances, the cleaved RNA template is generated by contacting the biological sample with an RNA-cutting enzyme (e.g., CRISPR effector proteins or Argonaute proteins). In some instances, the biological sample is contacted with the RNA-cutting enzyme before contacting the sample with the oligonucleotide for extension. In some embodiments, an oligonucleotide hybridizes to the cut RNA template to perform a reverse transcription reaction. In some embodiments, the method comprises contacting the biological sample with a guide nucleic acid and an RNA-cutting enzyme. In some embodiments, the RNA-cutting enzyme is an Argonaute protein. Any suitable Argonaute protein for cutting RNA in a nucleic acid duplex (e.g., within the guide target sequence bound to the guide nucleic acid) can be used. Generally, Argonaute proteins contain 6 main domains (N-terminal, LI (Linker 1), PAZ (Piwi-Argonaute- Zwille), L2 (Linker 2), MID (Middle) and PIWI (P-element induced wimpy testis) responsible for binding of a guide nucleic acid and recognition of a guide target sequence. More specifically, the PIWI domain can possess a nuclease active site with a catalytic tetrad (e.g., amino acid sequence DEDX, wherein X is the amino acid D, H, or K), wherein the catalytic tetrad coordinates two divalent metal cations (e.g., Mn2+, Mg2+, etc.) essential for target cleavage. In some embodiments, the Argonaute protein is an RNA-guided Argonaute, and the guide nucleic acid is an RNA molecule. In some embodiments, the Argonaute protein is a DNA-guided Argonaute, and the guide nucleic acid is a DNA molecule.
[0107] In some embodiments, the Argonaute protein is a naturally-occurring protein (e.g., naturally occurs in prokaryotic or eukaryotic cells). In some embodiments, the Argonaute protein is not a naturally-occurring protein (e.g., a variant or mutant protein). In some embodiments, the Argonaute protein is a recombinant protein. In some embodiments, the33MOFO-358213026202412023740Argonaute protein is genetically engineered (such as described in WO 2019 / 222036 and US20160289734, each of which herein incorporated by reference in their entireties).
[0108] In some embodiments, the Argonaute protein is a eukaryotic Argonaute protein. Generally, eukaryotic Argonaute proteins can mediate cutting of a target RNA template with a guide nucleic acid of RNA. In some embodiments, an Argonaute protein is of plant, algal, fungal (e.g., yeast), or animal (e.g., human, rodent, fruit fly, cnidarian, echinoderm, nematode, fish, amphibian, reptile, bird, etc.) origin. In some embodiments, the Argonaute protein is Agol, Ago2, Ago3, Ago4, PIWI 1, PIWIL 2, PIWI 3, or PIWI 4 (such as the Argonaute proteins described in WO 2007 / 048629, the content of which is herein incorporated by reference in its entirety). In some embodiments, the Argonaute protein is Ago2. In some embodiments, the Ago2 is Drosophila Ago2. In some embodiments, the Argonaute protein is a recombinant Drosophila Argonaute protein. In some embodiments, the Argonaute protein is expressed in a mammalian cell line. In some embodiments, the Argonaute protein is a Drosophila Argonaute protein expressed in a mammalian cell line. In some embodiments, a Drosophila Argonaute protein is expressed using a method such that a loading complex specific to Drosophila species is not provided to obtain guide-free proteins. In some embodiments, the Argonaute protein is a purified recombinant Drosophila Argonaute protein. In some embodiments, the Argonaute protein is expressed in an insect cell line, such as a Schneider 2 (S2) cell line. In some embodiments, the Argonaute protein is a Drosophila Argonaute protein expressed in an insect cell line, such as a S2 cell line. In some embodiments, the Drosophila Argonaute protein is loaded with the guide nucleic acid prior to contacting the biological sample. In some embodiments, the Argonaute protein is from Thermomyces thermophilus (such as an Argonaute protein described in U.S. Patent Application Publication No. 2023 / 0235306, the content of which is herein incorporated by reference in its entirety). In some embodiments, an Argonaute protein is from Vanderwaltozyma polyspora (also known as Kluyveromyces polysporus) (such as an Argonaute protein described in WO 2018 / 112336, the content of which is herein incorporated by reference in its entirety).
[0109] In some embodiments, the Argonaute protein is a prokaryotic Argonaute protein or a variant thereof. Generally, prokaryotic Argonaute proteins can mediate cutting of a RNA template with a guide oligonucleotide. In some cases, the prokaryotic Argonaute protein uses RNA as a guide oligonucleotide. In some cases, the prokaryotic Argonaute protein uses DNA as a guide oligonucleotide. In some embodiments, the Argonaute protein is a34MOFO-358213026202412023740Nitratireductor (optionally Nitratirereductor sp. XY-223), Enhydrobacter (optionally Enhydrobacter ae rosace us). Mesorhizobium (optionally Mesorhizobium sp. CNPSo 3140), Hyphomonas (optionally Hyphomonas sp. T16B2), Pseudooceanicola (optionally Pseudooceanicola lipolyticus), Tateyamaria (optionally Tateyamaria omphalii), Bradyrhizobium (optionally Bradyrhizobium sp. ORS 3257), Dehalococcoides (optionally Dehalococcoides mccartyi), Chroococcidiopsis (optionally Chroococcidiopsis cubana), Runella (optionally Runella slithyformis), Roseivirga (optionally Rosevirga seohaensis), Spirosoma (optionally Spirosoma endophy ticum), Pedobacter (optionally Pedobacter yonginense, Pedobacter insulae, or Pedobacter nyackensis), Planctomycetes bacterium (optionally Planctomycetes bacterium TBKlr or Planctomycetes bacterium V6), Dyadobacter (optionally Dyadobacter sp. QTA69), Mucilaginibacter (optionally Mucilaginibacter gotjawali, Mucilaginibacter polytichastri or Mucilaginibacter paludis), Hydrobacter (optionally Hydrobacter penzbergensis), Chitinophaga (optionally Chitinophaga costaii), Cytophagaceae bacterium (optionally Cytophagaceae bacterium SJW1-29), Emticicia (optionally Emticicia oligotrophica), Runella (optionally Runella sp. YX9), or Spirosoma (optionally Spirosoma pollinicola) Argonaute protein (See Li et al., “A programmable pAgo nuclease with RNA target preference from the psychrotolerant bacterium Mucilaginibacter paludis” Nucleic Acids Res. 2022 May 20;50(9):5226-5238; Lisitskaya et al., “Programmable RNA targeting by bacterial Argonaute nucleases with unconventional guide binding and cleavage specificity.” Nat Commun. 2022 Aug 8;13(1):4624; Sun et al., “An Argonaute from Thermus parvatiensis exhibits endonuclease activity mediated by 5' chemically modified DNA guides.” Acta Biochim Biophys Sin (Shanghai). 2022 May 25;54(5):686-695;U.S. Patent No. 10,253,311; US 15 / 089,243; US 17 / 575,957; US 17 / 854,897; and WO 2022 / 222920 each of which herein incorporated by reference in their entireties). In some embodiments, the Argonaute protein is from Thermus thermophilus.
[0110] In some embodiments, the Argonaute protein is a variant of a DNA-cutting Argonaute protein. In some cases, a DNA-cutting Argonaute protein is mutated to cut an RNA substrate via selection and / or directed evolution.
[0111] In some embodiments, the Argonaute protein cuts between any of the two positions between 9 and 12 of the guide target sequence. In some embodiments, the Argonaute protein cuts between positions 9 and 10 of the guide target sequence. In some embodiments, the Argonaute protein cuts between positions 11 and 12 of the guide target sequence.35MOFO-358213026202412023740
[0112] In some embodiments, the biological sample is incubated at a temperature below 60 °C (e.g., at room temperature or about 20-50 °C, about 20-30 °C, about 20-40 °C, about 25-40 °C, about 30-40 °C, about 35-40 °C, or about 40-50 °C) to allow the cutting of the guide target sequence by the Argonaute protein. In some embodiments, the biological sample is incubated at a temperature between 20 °C and 50 °C to allow the cutting of the guide target sequence by the Argonaute protein. In some embodiments, the biological sample is incubated at a temperature between 30 °C and 44 °C. In some embodiments, the biological sample is incubated at a temperature at about 37°C.
[0113] In some embodiments, the Argonaute protein possesses nuclease activity in a buffer comprising divalent cations. In some embodiments, the cutting of the RNA template by the Argonaute protein is performed in a buffer comprising at least one divalent cation (e.g., Fe2+, Co2+, Ni2+, Cu2+, Zn2+, Mg2+, Mn2+, or Ca2+). In some embodiments, the cutting of the RNA template by the Argonaute protein is performed in a buffer comprising Mg2+and / or Mn2+.
[0114] In some embodiments, the method comprises contacting the biological sample with a guide nucleic acid and an RNA-cutting enzyme. In some embodiments, the RNA-cutting enzyme is a CRISPR effector protein. Generally, a CRISPR effector protein can form a complex with a guide nucleic acid, and the complex functions as a CRISPR-Cas system. In some embodiments, the guide nucleic acid is a CRISPR guide RNA comprising a spacer sequence, wherein the spacer sequence hybridizes to the guide target sequence. Any suitable CRISPR-Cas systems can be used for cutting RNA in a nucleic acid duplex, and exemplary Cas effector proteins are described in herein.
[0115] In general, a CRISPR-Cas system is characterized by elements that promote the formation of a CRISPR complex at the site of a target RNA sequence (also referred to as a protospacer in the context of an endogenous CRISPR system). CRISPR-Cas systems form two major classes that differ in the organization of their effector modules. In Class 1 systems, multiple protein units form an effector complex together with the CRISPR RNA (crRNA) to recognize and cut a target RNA sequence, whereas a single protein complexing with crRNA does the job in a Class 2 system. To date, there are six types of CRISPR-Cas systems discovered: type I, type III, and type IV are identified as Class 1 systems, while type II, type V, and type VI are classified as Class 2. The specificity of cutting in CRISPR-Cas systems is conferred by RNA-36MOFO-358213026202412023740 based guidance through base-pairing, and the guide sequences can be adjusted to cut a new sequence.
[0116] In some embodiments, the CRISPR effector protein is a Class 2, Type VI Cas protein. In some embodiments, the CRISPR effector protein is a Class 2, Type II Cas protein. In some embodiments the CRISPR effector protein is a Cas 13 protein. In some embodiments the CRISPR effector protein is a Cas9 protein.
[0117] Type VI effectors are large proteins that contain two RNase domains of the higher eukaryotes and prokaryotes nucleotide-binding domain (HEPN) superfamily and that have been shown to, or are predicted to, specifically target RNA. In type VI CRISPR-Cas systems that target RNA, the Cas proteins usually comprise two conserved HEPN domains which are involved in RNA cleavage. In certain embodiments, the Cas protein processes crRNA to generate mature crRNA. The guide sequence of the crRNA recognizes target RNA with a complementary sequence and the Cas protein cuts the target RNA strand. More particularly, in certain embodiments, upon target binding, the Cas protein undergoes a structural rearrangement that brings two HEPN domains together to form an active HEPN catalytic site and the target RNA is then cut. The location of the catalytic site near the surface of the Cas protein allows nonspecific collateral ssRNA cutting.
[0118] Members of the CRISPR-Cas 13 system work as dual-component systems, in which a crRNA forms a complex with the Cas 13 protein without involving any tracrRNA. The flanking sequence(s) of protospacers, termed as “protospacer-flanking site” (PFS) and comparable to the “PAM” for Cas9, is essential for the RNA-targeting process.
[0119] In some embodiments, the Cas 13 protein is from a species of the genus Alistipes, Anaerosalibacter, Bacteroides, Bacteroidetes, Bergeyella, Blautia, Butyrivibrio, Capnocytophaga, Camobacterium, Chloroflexus, Chryseobacterium, Clostridium, Demequina, Eubacteriaceae , Eubacterium, Flavobacterium, Fusobacterium, Herbinix, Insolitispirillum, Lachnospiraceae, Leptotrichia, Listeria, Myroides, Paludibacter, Phaeodactylibacter, Porphyromonadaceae, Porphyromonas, Prevotella, Pseudobutyrivibrio, Psychroflexus, Reichenbachiella, Rhodobacter, Riemerella, Sinomicrobium, Thalassospira, Ruminococcus. In some embodiments, the Cas 13 protein is from Leptotrichia shahii, Listeria seeligeri, Lachnospiraceae bacterium (such as Lb MA2020, Lb NK4A1 79, Lb NK4A144), Clostridium aminophilum (such as Ca DSM 10710), Camobacterium gallinarum (such as Cg DSM 4847),37MOFO-358213026202412023740Paludibacter propionici genes (such as Pp WB4), Listeria weihenstephanensis (such as Lw FSL R9-03 17), Listeriaceae bacterium (such as Lb FSL M6-0635), Leptotrichia wadei (such as Lw F0279), Rhodobacter capsulatus (such as Re SB 1003, Re R121, Re DE442), Leptotrichia buccalis (such as Lb C-1013-b), Herbinix hemicellulosilytica, Eubacteriaceae bacterium (such as Eb CHKCI004), Blautia. sp Marseille-P2398, Leptotrichia sp. oral taxon 879 str. F0557, Chloroflexus aggregans, Demequina aurantiaca, Thalassospira sp. TSLS-1, Pseudobutyrivibrio sp. OR37, Butyrivibrio sp. Y AB3001, Leptotrichia sp. Marseille-P3007, Bacteroides ihuae, Porphyromonadaceae bacterium (such as Pb KH3CP3RA), Listeria riparia, Insolitispirillum peregrinum, Alistipes sp. ZOR0009, Bacteroides pyogenes (such as Bp F0041), Bacteroidetes bacterium (such as Bb GWA2_31_9), Bergeyella zoohelcum (such as Bz ATCC 43767), Capnocytophaga cammorsus, Capnocytophaga cynodegmi, Chryseobacterium carni pull orum, Chryseobacterium jejuense, Chryseobacterium ureilyticum, Flavobacterium branchiophilum, Flavobacterium columnare, Flavobacterium sp. 316, Myroides odoratimimus (such as Mo CCUG 10230, Mo CCUG 12901, Mo CCUG 3837), Paludibacter propionicigenes, Phaeodactylibacter xiamenensis, Porphyromonas gingivalis (such as Pg F0185, Pg F0568, Pg JCVI SC001, Pg W4087, Porphyromonas gulae, Porphyromonas sp. COT-052 OH4946, Prevotella aurantiaca, Prevotella buccae (such as Pb ATCC 33574), Prevotella falsenii, Prevotella intermedia (such as Pi 17, Pi ZT), Prevotella pallens (such as Pp ATCC 700821), Prevotella pleuritidis, Prevotella saccharolytica (such as Ps F0055), Prevotella sp. MA2016, Prevotella sp. MSX73, Prevotella sp. P4-76, Prevotella sp. PS-119, Prevotella sp. PS-125, Prevotella sp. PS-60, Psychroflexus torquis, Reichenbachiella agariperforans, Riemerella anatipestifer, Sinomicrobium oceam, Fusobacterium necrophorum (such as Fn subsp. funduliforme ATCC 51357, Fn DJ-2, Fn BFTR-1, Fn subsp. Funduliforme), Fusobacterium perfoetens (such as Fp ATCC 29250), Fusobacterium ulcerans (such as Fu ATCC 49185), Anaerosalibacter sp. ND1, Eubacterium siraeum, Ruminococcus flavefaciens (such as Rfx XPD3002), or Ruminococcus albus.
[0120] In some embodiments, the Casl3 protein is a Casl3a (C2c2) protein. In some embodiments, the Casl3a protein is from a species of the genus Bacteroides, Blautia, Butyrivibrio, Camobacterium, Chloroflexus, Clostridium, Demequina, Eubacterium, Herbinix, Insolitispirillum, Lachnospiraceae, Leptotrichia, Listeria, Paludibacter, Porphyromonadaceae, Pseudobutyrivibrio, Rhodobacter, or Thalassospira. In some embodiments, the Cas 13a protein is38MOFO-358213026202412023740 from Leptotrichia shahii, Listeria seeligeri, Lachnospiraceae bacterium (such as Lb MA2020, Lb NK4A179, Lb NK4A144), Clostridium aminophilum (such as Ca DSM 10710), Camobacterium gallinarum (such as Cg DSM 4847), Paludibacter propionicigenes (such as Pp WB4), Listeria weihenstephanensis (such as Lw FSL R9-03 17), Listeriaceae bacterium (such as Lb FSL M6-0635), Leptotrichia wadei (such as Lw F0279), Rhodobacter capsulatus (such as Re SB 1003, Re R121, Re DE442), Leptotrichia buccalis (such as Lb C-1013-b), Herbinix hemicellulosilytica, Eubacteriaceae bacterium (such as Eb CHKCI004), Blautia. sp Marseille-P2398, Leptotrichia sp. oral taxon 879 str. F0557, Chloroflexus aggregans.Demequina aurantiaca, Thalassospira sp. TSL5-1, Pseudobutyrivibrio sp. OR37, Butyrivibrio sp. Y AB3001, Leptotrichia sp. Marseille-P3007, Bacteroides ihuae, Porphyromonadaceae bacterium (such as Pb KH3CP3RA), Listeria riparia, or Insolitispirillum peregrinum.
[0121] In some embodiments, the Casl3 protein is a Casl3b protein. In some embodiments, the Casl3b protein is from a species of the genus Alistipes, Bacteroides, Bacteroidetes, Bergeyella, Capnocytophaga, Chryseobacterium, Flavobacterium, Myroides, Paludibacter, Phaeodactylibacter, Porphyromonas, Prevotella, Psychroflexus, Reichenbachiella, Riemerella, or Sinomicrobium; In some embodiments, the Casl3abprotein is from Alistipes sp. ZOR0009, Bacteroides pyogenes (such as Bp F0041 ), Bacteroidetes bacterium (such as Bb GW A2 _31 _9), Bergeyella zoohelcum (such as Bz ATCC 43767), Capnocytophaga canimorsus, Capnocytophaga cynodegmi, Chryseobacterium camipullorum, Chryseobacterium jejuense, Chryseobacterium ureilyticum, Flavobacterium branchiophilum, Flavobacterium columnare, Flavobacterium sp. 316, Myroides odoratimimus (such as Mo CCUG 10230, Mo CCUG 12901, Mo CCUG 3837), Paludibacter propionicigenes, Phaeodactylibacter xiamenensis, Porphyromonas gingivalis (such as Pg F0185, Pg F0568, Pg JCVI SC001, Pg W4087, Porphyromonas gulae, Porphyromonas sp. COT-052 OH4946, Prevotella aurantiaca, Prevotella buccae (such as Pb ATCC 33574), Prevotella falsenii, Prevotella intermedia (such as Pi 17, Pi ZT), Prevotella pallens (such as Pp ATCC 700821), Prevotella pleuritidis, Prevotella saccharolytica (such as Ps F0055), Prevotella sp. MA2016, Prevotella sp. MSX73, Prevotella sp. P4-76, Prevotella sp. P5-1 19, Prevotella sp. P5-125, Prevotella sp. P5-60, Psychroflexus torquis, Reichenbachiella agariperforans, Riemerella anatipestifer, or Sinomicrobium oceani.
[0122] In some embodiments, the Casl3 protein is a Casl3c protein. In some embodiments, the Casl3c protein is from a species of the genus Fusobacterium or39MOFO-358213026202412023740Anaerosalibacter. In some embodiments, the Cast 3c protein is from Fusobacterium necrophorum (such as Fn subsp. funduliforme ATCC 51357, Fn DJ-2, Fn BFTR-1, Fn subsp. Funduliforme), Fusobacterium perfoetens (such as Fp ATCC 29250), Fusobacterium ulcerans (such as Fu ATCC 49185), or Anaerosalibacter sp. NDI.
[0123] In some embodiments, the Casl3 protein is a Casl3d protein. In some embodiments, the Casl3d protein is from a species of the genus Eubacterium or Ruminococcus . In some embodiments, the Casl3d protein is from Eubacterium siraeum, Ruminococcus flavefaciens (such as Rfx XPD3002), or Ruminococcus albus.
[0124] The guide nucleic acid or guide RNA of a Class 2 type V CRISPR-Cas protein comprises a tracr-mate sequence (encompassing a “direct repeat” in the context of an endogenous CRISPR system) and a guide sequence (also referred to as a “spacer” in the context of an endogenous CRISPR system). Indeed, in contrast to the type II CRISPR-Cas proteins, the Casl3 protein does not rely on the presence of a tracr sequence. In some embodiments, the CRISPR-Cas system or complex as described herein does not comprise and / or does not rely on the presence of a tracr sequence (e.g., if the Cas protein is Casl3). In certain embodiments, the guide nucleic acid comprises, consists essentially of, or consists of a direct repeat sequence fused or linked to spacer sequence.
[0125] In some embodiments, the RNA targeting Cas protein is a Cas9 protein, which in some instances is referred to as RNA-targeting Cas9 (RCas9). In some embodiments, the Cas9 protein comprises a mutation in the naturally occurring Cas9. In some embodiments, a Cas9 protein is engineered to target RNA instead of DNA. In some embodiments, an engineered nucleoprotein complex comprises a Cas9 protein and a single guide RNA (sgRNA) to recognize a target RNA sequence. Optionally, in such systems, an (chemically-modified or synthetic) antisense PAMmer oligonucleotide is included to simulate a DNA substrate for recognition by Cas9 via hybridization to the target RNA. In some embodiments, the Cas9 protein is a 5. pyogenes Cas9 (SpyCas9) and the method comprises contacting the biological sample with a DNA oligonucleotide comprising the cognate PAM sequence (a PAMmer).
[0126] Programmable targeting of RNAs with Cas9 is possible by providing the PAM as part of an oligonucleotide (PAMmer) that hybridizes to the target RNA (O’Connell et al., “Programmable RNA recognition and cleavage by CRISPR / Cas9,” Nature 2014; 516(7530): 263-266, incorporated herein by references in its entirety for all purposes). By taking advantage40MOFO-358213026202412023740 of the Cas9 target search mechanism that relies on PAM sequences (Sternberg et al., “DNA interrogation by the CRISPR RNA-guided endonuclease Cas9. Nature. 2014 Mar 6;507(7490):62-72014, incorporated herein by references in its entirety for all purposes), a mismatched PAM sequence in the PAMmer / RNA hybrid allows exclusive targeting of RNA and not the encoding DNA. By separating the PAM and sgRNA target among two molecules (the PAMmer oligonucleotide and the target mRNA) that only associate in the presence of a target mRNA, the Cas9 allows recognition of RNA while avoiding the encoding DNA. In some instances, the PAMmer is dispensable for RNA targeting by Cas9, and RNA recognition does not require, but is enhanced by, the PAMmer (Batra et al., “Elimination of Toxic Micro satellite Repeat Expansion RNA by RNA-Targeting Cas9,” Cell 2017; 170(5): 899-912, incorporated herein by references in its entirety for all purposes).
[0127] In some embodiments, a CRISPR-Cas9 system for RNA-dependent RNA targeting is used in methods described herein. In some embodiments, Cas9 enzymes from subtype II-A or subtype II-C are used to recognize single-stranded RNA (ssRNA) by an RNA- guided mechanism that is independent of a protospacer-adjacent motif (PAM) sequence in the target RNA. In some embodiments, the Cas9 protein is a .S'. aureus Cas9 (SauCas9) or a C. jejuni Cas9 (CjeCas9). In some embodiments, a Cas protein disclosed herein is a nuclease Cas9 protein from subtype II-A or subtype II-C, and includes those described in Strutt et al., “RNA- dependent RNA targeting by CRISPR-Cas9,” eLife 2018; 7:e32724, incorporated herein by references in its entirety for all purposes.
[0128] In some embodiments, the guide nucleic acid is designed such that upon hybridization of the spacer sequence to the RNA template, the CRISPR effector protein cuts the RNA template at a position about 3 to about 15, about 5 to about 8, about 8 to about 15, about 12 to about 20, about 15 to about 25, or about 25 to about 35 nucleotides from the 3' end of the hybridized spacer sequence.
[0129] In some embodiments, the biological sample is incubated at a temperature to allow the cutting of the template RNA by the CRISPR effector protein. In some embodiments, the biological sample is incubated at a temperature below 60 °C. In some embodiments, the biological sample is incubated at room temperature. In some embodiments, the biological sample is incubated at a temperature of about 4-10 °C, about 10-20 °C, about 20-50 °C, about 20-30 °C,41MOFO-358213026202412023740 about 20-40 °C, about 25-40 °C, about 30-40 °C, about 35-40 °C, or about 40-50 °C. In some embodiments, the biological sample is incubated at a temperature at about 37°C.A. Oligonucleotide Extension and Processing
[0130] Provided herein is a workflow for generating short fragments using a template RNA by performing an extension reaction of an oligonucleotide hybridized to a RNA template and stopping extension using terminating nucleotides or by cleavage. In some embodiments, the biological sample is contacted with the oligonucleotide and with a reverse transcriptase simultaneously. In some embodiments, the biological sample is contacted with the oligonucleotide before the reverse transcriptase is provided. In some embodiments, the biological sample is contacted with a plurality of free nucleotides to generate an extended oligonucleotide, wherein the plurality of free nucleotides comprises a terminating nucleotide that prevents further extension of the extended oligonucleotide. In some embodiments, the method comprises washing the biological sample after extension. In some embodiments, the biological sample is contacted with a plurality of free nucleotides comprising a ribonucleotide, an inosine, or a uracil.
[0131] In some embodiments, an RNA template is cleaved prior to contacting the biological sample with the oligonucleotide for extension. In some embodiments, the extended oligonucleotide generated using the cleaved RNA template is about 1 to about 500 nucleotides in length. In some embodiments, the extended oligonucleotide generated using the cleaved RNA template is about 5 to about 500 nucleotides in length. In some embodiments, the extended oligonucleotide generated using the cleaved RNA template is about 10 to about 500 nucleotides in length. In some embodiments, the extended oligonucleotide generated using the cleaved RNA template is about 20 to about 500 nucleotides in length. In some embodiments, the extended oligonucleotide generated using the cleaved RNA template is about 40 to about 500 nucleotides in length. In some embodiments, the extended oligonucleotide generated using a cleaved RNA template with a length of between about 2 and about 500, between about 2 and 200, between about 2 and 100, between about 2 and 100, between about 2 and 50, between about 2 and 20, between about 10 and about 500, between about 10 and 200, between about 10 and 100, between about 10 and 150, between about 10 and 50, between about 10 and 20, between about 50 and about 500, between about 50 and 200, between about 50 and 150, between about 50 and 100, between about 50 and 80, or between about 100 and about 500 nucleotides in length. In some embodiments, the extended oligonucleotide generated using the cleaved RNA template is42MOFO-358213026202412023740 between about 10 and about 600, between about 10 and 500, or between about 10 and about 450 nucleotides in length. In some embodiments, the extended oligonucleotide generated using the cleaved RNA template is about 10, about 20, about 50, about 100, about 150, about 200, about 250, about 300, about 350, about 400, or about 450 nucleotides in length, or any length in a range having endpoints selected from the group consisting of 10, 20, 50, 100, 150, 200, 250, 300, 350, 400, or 450 nucleotides in length. In some embodiments, the extended oligonucleotide generated using the cleaved RNA template is between about 10 and about 300, between about 10 and about 200, between about 10 and about 100, between about 50 and about 300, between about 50 and about 200, between about 50 and about 100, between about 70 and about 300, between about 70 and about 200, between about 70 and about 100, between about 100 and about 300, between about 100 and about 250, between about 150 and about 300, between about 150 and about 250, between about 100 and about 200, or between about 150 and about 250 nucleotides in length.
[0132] In some embodiments, an RNA template is not cleaved prior to extending the oligonucleotide using a reverse transcriptase. In some embodiments, an RNA template cleaved prior to extending the oligonucleotide using a reverse transcriptase.
[0133] In some embodiments, the oligonucleotide is configured to bind to an RNA template. In some embodiments, the oligonucleotide is configured to be extended using an RNA template. In some embodiments, the oligonucleotide serves as a priming sequence for extension to generate an extended oligonucleotide. In some embodiments, the oligonucleotide is about 1 to about 50 nucleotides in length. In some embodiments, the oligonucleotide is between about 20 and about 60, between about 20 and 50, or between about 30 and about 45 nucleotides in length. In some embodiments, the oligonucleotide is about 25, about 30, about 35, about 40, or about 45 nucleotides in length, or any length in a range having endpoints selected from the group consisting of 25, 30, 35, 40, or 45 nucleotides in length. In some embodiments, the oligonucleotide is between about 10 and about 40, between about 10 and about 30, between about 15 and about 30, between about 15 and about 25, between about 10 and about 20, or between about 15 and about 25 nucleotides in length. In some embodiments, the oligonucleotide is between about 25 and about 35 nucleotides in length.
[0134] In some embodiments, the oligonucleotide is a single- stranded sequence comprising at least 4, 5, 6, 7, or 8 consecutive deoxyribonucleotides. In some embodiments, the 43MOFO-358213026202412023740 oligonucleotide is a deoxyribonucleic acid (DNA) oligonucleotide. In some cases, the oligonucleotide is a single- stranded DNA (ssDNA) oligonucleotide.
[0135] In some embodiments, the oligonucleotide hybridizes to about 10 to about 35, about 15 to about 30, about 25 to about 35, about 5 to about 25, about 10 to about 20, about 15 to about 25, about 5 to about 15, about 8 to about 18, about 10 to about 18, or about 15 to about 20 nucleotides of the RNA template. In some embodiments, the oligonucleotide hybridizes to about 10 to about 30 nucleotides of the RNA template. In some embodiments, the oligonucleotide hybridizes to about 5 to about 20 nucleotides of the RNA template.
[0136] In some embodiments, the extended oligonucleotide generated using a plurality of free nucleotides comprising a terminating nucleotide is about 40 to about 500 nucleotides in length. In some embodiments, the extended oligonucleotide generated using a plurality of free nucleotides comprising a terminating nucleotide is between about 40 and about 600, between about 40 and 500, or between about 40 and about 450 nucleotides in length. In some embodiments, the extended oligonucleotide generated using a plurality of free nucleotides comprising a terminating nucleotide is about 100, about 150, about 200, about 250, about 300, about 350, about 400, or about 450 nucleotides in length, or any length in a range having endpoints selected from the group consisting of 100, 150, 200, 250, 300, 350, 400, or 450 nucleotides in length. In some embodiments, the extended oligonucleotide generated using a plurality of free nucleotides comprising a terminating nucleotide is between about 50 and about 300, between about 50 and about 200, between about 50 and about 100, between about 70 and about 300, between about 70 and about 200, between about 70 and about 100, between about 100 and about 300, between about 100 and about 250, between about 150 and about 300, between about 150 and about 250, between about 100 and about 200, or between about 150 and about 250 nucleotides in length.
[0137] In some aspects, provided herein is a method of analyzing a biological sample, comprising contacting the biological sample with plurality of oligonucleotides. In some instances, the oligonucleotide comprises a common sequence for binding to a plurality of different RNAs. In some instances, the oligonucleotide comprises a sequence configured for binding to a target RNA template. In some instances, a plurality of different oligonucleotides comprising different sequences are used for binding to a plurality of different RNA templates. In some instances, the oligonucleotide comprises a sequence configured for binding to a sequence 44MOFO-3582130262024120237403’ of a region of interest in a RNA template. In some embodiments, each oligonucleotide of the plurality of oligonucleotides hybridizes to a ribonucleic acid (RNA) of a plurality of RNAs (e.g., different analytes) in the biological sample. In some embodiments, the plurality of RNAs comprises at least 2, at least 5, at least 10, at least 25, at least 50, at least 75, at least 100, at least 300, at least 1,000, at least 3,000, at least 10,000, at least 30,000, at least 50,000, at least 100,000, at least 250,000, at least 500,000, or at least 1,000,000 distinguishable RNAs. In some embodiments, a plurality of oligonucleotides are configured to bind to a plurality of RNAs comprising at least 2, at least 5, at least 10, at least 25, at least 50, at least 75, at least 100, at least 300, at least 1,000, at least 3,000, at least 10,000, at least 30,000, at least 50,000, at least 100,000, at least 250,000, at least 500,000, or at least 1,000,000 distinguishable RNAs. In some embodiments, subsets of the plurality of oligonucleotides may share a common sequence for hybridizing to different RNA molecules. In some examples, the oligonucleotides hybridize to the 3’ poly(A) tail of RNA transcripts. In some instances, the oligonucleotides hybridize to a sequence in the RNA transcript that is 3’ to a region of interest (e.g., a variant nucleotide or short variant sequence).
[0138] In some instances, the oligonucleotide comprises a random sequence for binding to an RNA template. In some embodiments, a plurality of oligonucleotides comprising a plurality of random sequences (e.g., randomers) are used to bind to a plurality of RNA templates. In some instances, an oligonucleotide comprises a sequence to target a specific analyte of interest. In some instances, an oligonucleotide comprises a targeted sequence to bind a specific analyte of interest (e.g., a specific mRNA).
[0139] In some instances, an oligonucleotide comprises a sequence complementary to an analyte of interest. In some embodiments, the oligonucleotide serves as a priming sequence. In some embodiments, the oligonucleotide serves as a primer for extension. For example, an oligonucleotide comprises a poly-T sequence complementary to a poly-A tail of an mRNA analyte. In some instances, an oligonucleotide comprises an anchoring sequence to ensure hybridization at the sequence end (e.g., of the mRNA). For example, the anchoring sequence can include a random short sequence of nucleotides, such as a 1-mer, 2-mer, 3-mer or longer sequence, which can ensure that a poly-T segment is more likely to hybridize at the sequence end of the poly-A tail of the mRNA. In some cases, reverse transcription results in a cDNA transcript of the mRNA and comprises the sequence(s) of the oligonucleotide. In some45MOFO-358213026202412023740 embodiments, the oligonucleotide comprises a sequencing primer binding sequence. In some embodiments, the oligonucleotide comprises a RCA primer binding sequence. In some instances, the primer binding sequence is the same as the RCA primer binding sequence. In some instances, the oligonucleotide comprises the primer binding sequence at the 5’ end of the oligonucleotide a poly-T sequence at the 3’ end of the oligonucleotide. In some instances, the oligonucleotide comprises from 5’ to 3’: primer binding sequence - poly-T sequence. In some embodiments, the oligonucleotide comprises a primer binding sequence that is used for both binding the RCA primer and binding a sequencing primer (e.g., for downstream detection, as described in Section III). In some aspects, the primer binding sequence or a complement thereof is used for binding the RCA primer and binding a sequencing primer. For example, the generated RCA product comprises a complement of the primer binding sequence of the oligonucleotide and a sequencing primer is capable of hybridizing to the complement of the primer binding sequence in the RCA product.
[0140] In some embodiments, extension of the oligonucleotide is performed by contacting the oligonucleotide with an enzyme for extension (e.g., reverse transcriptase) and a plurality of free nucleotides. In some cases, plurality of free nucleotides comprises a terminating nucleotide that prevents further extension of the extended oligonucleotide. In some instances, the terminating nucleotide is a nucleotide lacking a 3' hydroxyl group in the deoxyribose sugar. In some instances, the terminating nucleotide is a dideoxyribonucleotide (ddNTP). In some embodiments, the plurality of free nucleotides comprises two, three or four different ddNTPs selected from the group consisting of ddGTP, ddATP, ddTTP and ddCTP.
[0141] In some examples, the extension of the oligonucleotide is performed using a plurality of free nucleotides, wherein the plurality of free nucleotides lacks nucleotides of at least one of the four canonical bases: adenine (A), thymine (T), guanine (G), uracil (U), and cytosine (C). In some embodiments, the length of extension product (e.g., by reverse transcription or any polymerase extension) can be controlled by performing a stepwise extension. In some instances, in a particular cycle of extension, nucleotides are provided one by one to control the extension. For example, a particular cycle provides a plurality of free nucleotides with only A nucleotides, followed by a cycle that provides a plurality of free nucleotides with only G nucleotides, followed by a cycle that provides a plurality of free nucleotides with only C nucleotides, followed by a cycle that provides a plurality of free nucleotides with only T nucleotides, and repeat for a number of 46MOFO-358213026202412023740 cycles until the desired length is achieved. In some instances, the number of cycles necessary is reduced by using a mixture of free nucleotides, e.g., using a pool with A nucleotides, C nucleotides, and T nucleotides, followed by a cycle that provides a plurality of free nucleotides with only G nucleotides. In some instances, the controlled extension is repeated by performing cycles of providing free nucleotides lacking particular nucleotides. In some instances, the lacking nucleotide is switched between different cycles.
[0142] In some cases, a plurality of free nucleotides lack a free nucleotide that basepairs at the terminating nucleotide. In some embodiments, the free nucleotides are A, G, and C and is lacking the terminating nucleotide T. In some cases, the free nucleotides are G, A, and T and is lacking the terminating nucleotide C. In some instances, the free nucleotides are T, C, and G and is lacking the terminating nucleotide A. In some cases, the free nucleotides are G, A, and U and is lacking the terminating nucleotide C (as illustrated in FIG. 5, middle panel). In some embodiments, the free nucleotides are C, T, and A and is lacking the terminating nucleotide G (as illustrated in FIG. 5, bottom panel).
[0143] In some cases, resuming extension of the oligonucleotide to obtain a further extended nucleic acid molecule also comprises contacting the extended nucleic acid molecule with additional free nucleotides, where the additional free nucleotides comprise the lacking terminating nucleotide. In some examples, the additional free nucleotides comprise all canonical bases: adenine (A), thymine (T), guanine (G), uracil (U), and cytosine (C). In some examples, the additional free nucleotides of only one of adenine (A), thymine (T), guanine (G), uracil (U), or cytosine (C).
[0144] In some cases, the free nucleotides comprise A, T, and G, and the additional free nucleotides comprise C (e.g., as shown in FIG. 5). In some instances, the free nucleotides comprise T, G, and C, and the additional free nucleotides comprise A. In some examples, the free nucleotides comprise A, C, and G, and the additional free nucleotides comprise T. In some embodiments, the free nucleotides comprise A, U, and C, and the additional free nucleotides comprise G. In some cases, the free nucleotides comprise A, U, and G, and the additional free nucleotides comprise C. In some instances, the free nucleotides comprise U, G, and C, and the additional free nucleotides comprise A. In some examples, the free nucleotides comprise A, C, and G, and the additional nucleotides comprise U.47MOFO-358213026202412023740
[0145] In some embodiments, the extending of the oligonucleotide in template- directed fashion does not resume unless additional free nucleotides are added. In some embodiments, the additional free nucleotides comprises a complete mixture of free nucleotides comprising four different bases, and none of the four different bases are reversible terminator nucleotides. In some embodiments, subsequent to resuming extension of the extended nucleic acid molecule, unbound free nucleotides are removed from the extended oligonucleotide.
[0146] In some instances, extension of the oligonucleotide is performed by stepwise, providing nucleotide pools of one nucleotide at a time, e.g., only one of A, C, G, and T nucleotides. In some aspects, with each round of extension, a variable amount of nucleotides in the sample is extended. In some embodiments, the extension of oligonucleotides in the sample is on average in each cycle, no more than 3 nucleotides, no more than 4 nucleotides, no more than 5 nucleotides, or no more than 6 nucleotides. In some embodiments, the extension of oligonucleotides in the sample is on average in each cycle about 4 nucleotides. In some instances, the extension performed in a stepwise manner is performed until a desired length is reached.
[0147] In some instances, one single round of extension is performed to generate the extended oligonucleotide and the extended oligonucleotide is circularized (e.g., by ligation). In some instances, a polymerase chain reaction (PCR) reaction is not further performed using the extended oligonucleotide.
[0148] In some instances, the terminating nucleotide comprises a terminating group or a terminating modification. In some instances, the terminating modification is a reverse linkage. In some instances, the terminating nucleotide comprises a 3' inverted dT. An inverted dT is incorporated at the 3 ’-end of an oligo, leading to a 3 ’-3’ linkage which inhibits extension. In some embodiments, the terminator nucleotide comprises a modified nucleobase. In some embodiments, the modified nucleobase comprises a bulky moiety that inhibits extension of the oligonucleotide. In some embodiments, the modified nucleobase is alkylated. In some embodiments, the modified nucleobase is ’-alkylatcd dATP. In some embodiments, the modified nucleobase is 5-substituted aryloxy-methyl-dUTP or an analog thereof. In some embodiments, the modified nucleobase is a 5-propargyl-amino-dUTP scaffold. In some instances, the terminating nucleotide is a reversible terminator nucleotide. For example, the reversible terminator nucleotide comprises an azidomethyl group, an amino group, a nitrobenzyl48MOFO-358213026202412023740 group, an allyl group, a carbonate, a functionalized photocleavable ether, a methyl group, or a cyanoethyl group. In some cases, the reversible terminator nucleotide is a 3'-O-blocked reversible terminator nucleotide. In some embodiments, the plurality of free nucleotides comprises two, three or four different ddNTPs comprising the terminating group or the terminating modification. In any of the embodiments herein, the second nucleic acid strand is incapable of being extended by a polymerase. In some embodiments, the extended oligonucleotide comprises a 3’ block moiety. In some embodiments, the 3’ block moiety is an irreversible terminating group. In some embodiments, the irreversible terminating group is a 3’ dideoxy nucleotide. In some embodiments, the 3’ block moiety is a reversible terminating group. In some embodiments, the reversible terminating group is an azidomethyl group, an amino group, a nitrobenzyl group, or an allyl group.
[0149] In some embodiments, the reversible terminator nucleotide is 3’-O- azidomethyl-dNTP, such as 3’-O-azidomethy-dCTP. In some embodiments, the reversible terminator nucleotide is 3’-O-allyl-dNTP, such as 3’-O-allyl-dCTP. In some embodiments, the reversible terminator nucleotide is 3’-O-nitrobenzyl-dNTP, such as 3’-O-nitrobenzyl-dCTP. In some embodiments, the reversible terminator nucleotide is 3’-O-methyl-dNTP, such as 3’-O- methyl-dCTP. In some embodiments, the reversible terminator nucleotide is 3’-O-2-cyanoethyl- dNTP, such as 3’-O-2-cyanoethyl-dCTP. In some aspects, the nucleotide base comprising the reversible terminator is subject to the removal of the terminator, for example, via a reducing agent. In some aspects, the reducing agent for removing the terminator can include, but is not limited to, tris(2-carboxyethyl)phosphine (TCEP), dithiothreitol (DTT), or beta-mercaptoethanol (BME). In some instances, the sample is treated to chemically deblock the reversible terminator. In some cases, a linker in the reversible terminator is cleaved.
[0150] In some aspects, as illustrated in FIG. 1, the plurality of free nucleotides comprises an amount of a terminating nucleotide (e.g., ddNTPs or inverted dT) that causes extension to stop. In some embodiments, the amount of the terminating nucleotide is a low amount. Examples of an amount can refer to a weight, a mass, a concentration (e.g. a molar concentration), a ratio, or a density. In some embodiments, the amount of the terminating nucleotide is a molar concentration of the terminating nucleotide. In some embodiments, the amount of the terminating nucleotide within the plurality of free nucleotides is determined relative to an amount of a corresponding free nucleotide within the plurality of free nucleotides.49MOFO-358213026202412023740For example, the amount of the terminating nucleotide ddCTP within the plurality of free nucleotides is determined relative to the amount of the corresponding free nucleotide dCTP within the plurality of free nucleotides.
[0151] In some embodiments, the terminating nucleotide is present in an amount lower than an amount of the corresponding free nucleotide in the plurality of free nucleotides. In some embodiments, the amount of the terminating nucleotide is about 2-fold, 5-fold, 10-fold, or 30-fold lower than the amount of the corresponding free nucleotide within the plurality of free nucleotides. In some embodiments, the amount of the terminating nucleotide is about 1.5-fold, about 2-fold, about 2.5-fold, about 3-fold, about 3.5-fold, about 4-fold, about 4.5-fold, or about 5-fold lower than the amount of the corresponding free nucleotide within the plurality of free nucleotides. In some embodiments, the amount of the terminating nucleotide is about 2-fold lower than the amount of the corresponding free nucleotide within the plurality of free nucleotides. In some embodiments, the plurality of free nucleotides comprises two or more terminating nucleotides, wherein each terminating nucleotide of the two or more terminating nucleotides is present in a known amount, wherein the known amount of each terminating nucleotide is determined relative to each corresponding free nucleotide of the plurality of free nucleotides.
[0152] In some embodiments, the terminating nucleotide is present in an amount greater than an amount of the corresponding free nucleotide in the plurality of free nucleotides. In some embodiments, the amount of the terminating nucleotide is about 2-fold, 5-fold, 10-fold, or 30-fold greater than the amount of the corresponding free nucleotide within the plurality of free nucleotides. In some embodiments, the amount of the terminating nucleotide is about 1.5- fold, about 2-fold, about 2.5-fold, about 3-fold, about 3.5-fold, about 4-fold, about 4.5-fold, or about 5-fold greater than the amount of the corresponding free nucleotide within the plurality of free nucleotides. In some embodiments, the amount of the terminating nucleotide is about 2-fold greater than the amount of the corresponding free nucleotide within the plurality of free nucleotides. In some embodiments, the plurality of free nucleotides comprises two or more terminating nucleotides, wherein each terminating nucleotide of the two or more terminating nucleotides is present in a known amount, wherein the known amount of each terminating nucleotide is determined relative to each corresponding free nucleotide of the plurality of free nucleotides.50MOFO-358213026202412023740
[0153] In some embodiments, the ratio of terminating nucleotide to corresponding free nucleotide within the plurality of free nucleotides is at least 10:1, 100:1, or 300:1. In some embodiments, the ratio of terminating nucleotide to corresponding free nucleotide within the plurality of free nucleotides is no more than 10:1, 100:1, or 300:1. In some embodiments, the ratio of terminating nucleotide to corresponding free nucleotide within the plurality of free nucleotides is at least 1:10, 1:100, or 1:300. In some embodiments, the ratio of terminating nucleotide to corresponding free nucleotide within the plurality of free nucleotides is no more than 1:10, 1:100, or 1:300. In some embodiments, the ratio of terminating nucleotide to corresponding free nucleotide within the plurality of free nucleotides is at least 30:1, 5:1, or 3:1. In some embodiments, the ratio of terminating nucleotide to corresponding free nucleotide within the plurality of free nucleotides is at least 1:3, 1:5 or 1:30.
[0154] In some embodiments, the amount of terminating nucleotide in the plurality of free nucleotides is tested and / or tuned. For example, the amount of terminating nucleotide can be tuned and / or altered after testing (e.g., when initial tests form extended oligonucleotides either longer or shorter than about 50-200 nucleotides). In some embodiments, the amount of terminating nucleotide is decreased relative to an initial amount of the terminating nucleotide in the plurality of free nucleotides. In some embodiments, the amount of terminating nucleotide is decreased after testing identifies extended oligonucleotides of lengths shorter than about 50-200 nucleotides. In some embodiments, the amount of terminating nucleotide is decreased by about 5%, about 10%, about 15%, about 20%, about 25%, about 30%, about 35%, about 40%, about 45%, about 50%, about 55%, about 60%, about 65%, about 70%, about 75%, about 80%, about 85%, about 90% or about 95% from the initial amount. In some embodiments, the amount of terminating nucleotide is increased relative to an initial amount of the terminating nucleotide in the plurality of free nucleotides. In some embodiments, the amount of terminating nucleotide is increased after testing identifies extended oligonucleotides of lengths longer than about 50-200 nucleotides. In some embodiments, the amount of terminating nucleotide is increased by about 5%, about 10%, about 15%, about 20%, about 25%, about 30%, about 35%, about 40%, about 45%, about 50%, about 55%, about 60%, about 65%, about 70%, about 75%, about 80%, about 85%, about 90% or about 95% from the initial amount. In some embodiments, the testing and / or tuning comprises two or more rounds of increasing and / or decreasing the amount of terminating nucleotide within the plurality of free nucleotides.51MOFO-358213026202412023740
[0155] In some embodiments, the method comprises removing the terminating nucleotide from the extended oligonucleotide. In some embodiments, the method comprises cleaving the terminating nucleotide from the extended oligonucleotide. In some embodiments, the method comprises removing the terminating group or the terminating modification from the terminating nucleotide incorporated into the extended oligonucleotide.
[0156] In some aspects, as illustrated in FIG. 2, the plurality of free nucleotides comprises an amount of a cleavable nucleotide that may be cleaved to block extension of the oligonucleotide. In some embodiments, the cleavable nucleotide is a uracil. In some embodiments, the cleavable nucleotide is an inosine. In some embodiments, the cleavable nucleotide is a ribonucleotide and the extended oligonucleotide is DNA. In some examples, the extended oligonucleotide comprises one or more uracil incorporated during extension and the extended oligonucleotide is cleaved using Uracil- Specific Excision Reagent (USER). In some examples, the extended oligonucleotide comprises one or more inosine incorporated during extension and the extended oligonucleotide is cleaved using Endonuclease V. In some examples, the extended oligonucleotide is DNA and comprises one or more ribonucleotides incorporated during extension and the extended oligonucleotide is cleaved using RNAse H. In some aspects, an extended nucleotide is cleaved prior to circularization of the extended oligonucleotide to generate the circularized template.
[0157] In some embodiments, the oligonucleotide is extended using a plurality of free ribonucleotides that comprises all four canonical bases: adenine, uracil, guanine and cytosine. In some embodiments, the oligonucleotide is extended using a plurality of free ribonucleotides that comprises at least one of four canonical bases: adenine, uracil, guanine and cytosine. In some embodiments, the oligonucleotide is extended using a plurality of free ribonucleotides comprising only one of four canonical bases: adenine, uracil, guanine and cytosine. In some embodiments, the oligonucleotide is extended using a plurality of free ribonucleotides comprising no more than two of four canonical bases: adenine, uracil, guanine and cytosine. In some embodiments, the oligonucleotide is extended using a plurality of free ribonucleotides comprising no more than three of four canonical bases: adenine, uracil, guanine and cytosine.
[0158] In some embodiments, the extended oligonucleotide is generated using a plurality of free nucleotides that does not comprise a terminating nucleotide. In some embodiments, the extended oligonucleotide is generated using a plurality of free nucleotides that52MOFO-358213026202412023740 does not comprise a cleavable nucleotide. In some embodiments, the extended oligonucleotide generated using a plurality of free nucleotides that does not comprise a ribonucleotide, an inosine, or a uracil.
[0159] In some embodiments, the cleavable nucleotide is present in an amount lower than an amount of non-cleavable nucleotides in the plurality of free nucleotides. Without being bound by theory, a cleavable nucleotide can be any nucleotide within the plurality of free nucleotides that is specifically targeted for cleavage by a particular treatment, while a non- cleavable nucleotide is any nucleotide within the plurality of free nucleotides that cannot be targeted for cleavage by the same particular treatment. For example, the cleavable nucleotide could be a uracil and a non-cleavable nucleotide could be dCTP, wherein the cleavable nucleotide is targeted for cleavage by treatment with USER. In some aspects, a non-cleavable nucleotide is not configured to be cleaved in the extended oligonucleotide. In some embodiments, the amount of the cleavable nucleotide is about 2-fold, 5-fold, 10-fold, or 30-fold lower than the amount of the non-cleavable nucleotides within the plurality of free nucleotides. In some embodiments, the amount of the cleavable nucleotide is about 1.5-fold, about 2-fold, about 2.5-fold, about 3-fold, about 3.5-fold, about 4-fold, about 4.5-fold, or about 5-fold lower than the amount of the non-cleavable nucleotides within the plurality of free nucleotides. In some embodiments, the amount of the cleavable nucleotide is about 2-fold lower than the amount of the non-cleavable nucleotides within the plurality of free nucleotides. In some embodiments, the plurality of free nucleotides comprises two or more cleavable nucleotides, wherein each cleavable nucleotide of the two or more cleavable nucleotide is present in a known amount, wherein the known amount of each cleavable nucleotide is determined relative to the amount of the non-cleavable nucleotides within the plurality of free nucleotides.
[0160] In some embodiments, the cleavable nucleotide is present in an amount greater than an amount of non-cleavable nucleotides in the plurality of free nucleotides. In some embodiments, the amount of the cleavable nucleotide is about 2-fold, 5-fold, 10-fold, or 30-fold greater than the amount of non-cleavable nucleotides within the plurality of free nucleotides. In some embodiments, the amount of the cleavable nucleotide is about 1.5-fold, about 2-fold, about 2.5-fold, about 3-fold, about 3.5-fold, about 4-fold, about 4.5-fold, or about 5-fold greater than the amount of non-cleavable nucleotides within the plurality of free nucleotides. In some53MOFO-358213026202412023740 embodiments, the amount of the cleavable nucleotide is about 2-fold greater than the amount of non-cleavable nucleotides within the plurality of free nucleotides.
[0161] In some embodiments, the ratio of cleavable nucleotide to non-cleavable nucleotides within the plurality of free nucleotides is at least 10:1, 100:1, or 300:1. In some embodiments, the ratio of cleavable nucleotide to non-cleavable nucleotides within the plurality of free nucleotides is no more than 10:1, 100:1, or 300:1. In some embodiments, the ratio of cleavable nucleotide to non-cleavable nucleotides within the plurality of free nucleotides is at least 1:10, 1:100, or 1:300. In some embodiments, the ratio of cleavable nucleotide non- cleavable nucleotides within the plurality of free nucleotides is no more than 1:10, 1:100, or 1:300. In some embodiments, the ratio of cleavable nucleotide to non-cleavable nucleotides within the plurality of free nucleotides is at least 30:1, 5:1, or 3:1. In some embodiments, the ratio of cleavable nucleotide to non-cleavable nucleotides within the plurality of free nucleotides is at least 1:3, 1:5 or 1:30.
[0162] In some aspects, the amount of cleavable nucleotide in the plurality of free nucleotides can be tested and / or tuned. In some aspects, the amount of cleavable nucleotide is tuned to allow for generation of an extended oligonucleotide that is about 50 to 200 nucleotides in total length after cleavage of the extended oligonucleotide at the incorporated cleavable nucleotide. For example, the amount of cleavable nucleotide can be tuned and / or altered after testing (e.g., when initial tests form extended oligonucleotides either longer or shorter than about 50-200 nucleotides).
[0163] In some embodiments, the amount of cleavable nucleotide is decreased relative to an initial amount of the cleavable nucleotide in the plurality of free nucleotides. In some embodiments, the amount of cleavable nucleotide is decreased after testing identifies extended oligonucleotides of lengths shorter than about 50-200 nucleotides. In some embodiments, the amount of cleavable nucleotide is decreased by about 5%, about 10%, about 15%, about 20%, about 25%, about 30%, about 35%, about 40%, about 45%, about 50%, about 55%, about 60%, about 65%, about 70%, about 75%, about 80%, about 85%, about 90% or about 95% from the initial amount. In some embodiments, the amount of cleavable nucleotide is increased relative to an initial amount of the cleavable nucleotide in the plurality of free nucleotides. In some embodiments, the amount of cleavable nucleotide is increased after testing identifies extended oligonucleotides of lengths longer than about 50-200 nucleotides. In some 54MOFO-358213026202412023740 embodiments, the amount of cleavable nucleotide is increased by about 5%, about 10%, about 15%, about 20%, about 25%, about 30%, about 35%, about 40%, about 45%, about 50%, about 55%, about 60%, about 65%, about 70%, about 75%, about 80%, about 85%, about 90% or about 95% from the initial amount. In some embodiments, the testing and / or tuning comprises two or more rounds of increasing or decreasing the amount of cleavable nucleotide within the plurality of free nucleotides.
[0164] In some embodiments, extension of the oligonucleotide is performed by contacting the oligonucleotide with an enzyme for extension (e.g., reverse transcriptase) and a plurality of free nucleotides, wherein the plurality of free nucleotides comprises a ribonucleotide, an inosine, or a uracil. In some cases, the plurality of free nucleotide comprises a cleavable nucleotide that can be cleaved enzymatically. In some instances, the extended oligonucleotide is cleaved at the incorporated ribonucleotide, the inosine, or the uracil, thereby generating a cleaved extended sequence.
[0165] In some aspects, the extended oligonucleotide is cleaved at the incorporated ribonucleotide, inosine, or uracil, thereby generating a cleaved extended sequence. In some aspects, inosine is incorporated into the extended oligonucleotide and an endonuclease is used to cleave the extended oligonucleotide. In some aspects, the uracil is incorporated into the extended oligonucleotide and USER is used to cleave the extended oligonucleotide. In some cases, RNase H is used to perform the cleaving. In some aspects, the extended oligonucleotide comprises at least one ribonucleotide for RNase H cutting to generate a shorter molecule (e.g., a cleaved extended sequence) that is used to generate a circularized template.
[0166] In some embodiments, the cleaved extended sequence is about 40 to about 500 nucleotides in length. In some embodiments, the cleaved extended sequence is between about 40 and about 600, between about 40 and 500, or between about 40 and about 450 nucleotides in length. In some embodiments, the cleaved extended sequence is about 100, 150, 200, 250, 300, 350, 400, or 450 nucleotides in length, or any length in a range having endpoints selected from the group consisting of 100, 150, 200, 250, 300, 350, 400, or 450 nucleotides in length. In some embodiments, the cleaved extended sequence is between about 50 and about 300, between about 50 and about 200, between about 50 and about 100, between about 70 and about 300, between about 70 and about 200, between about 70 and about 100, between about 100 and about 300, between about 100 and about 250, between about 150 and about 300, between55MOFO-358213026202412023740 about 150 and about 250, between about 100 and about 200, or between about 150 and about 250 nucleotides in length. In some aspects, prior to cleavage, the extended oligonucleotide is at least 500, at least 600, at least 700, at least 800, at least 1,000, at least 2,000, or at least 5,000 nucleotides in length.
[0167] In some embodiments, extension of the oligonucleotide using a cleaved RNA template is performed by contacting the oligonucleotide with an enzyme for extension (e.g., reverse transcriptase) and a plurality of free nucleotides. In some embodiments, the cleavage region is 5’ to a region of interest in the RNA template and the oligonucleotide for extension comprises a poly-T sequence complementary to a poly-A tail of an mRNA analyte. In some instances, the region of interest in the RNA template is flanked by a cleavage region and a sequence complementary to the oligonucleotide for extension. In some instances, the generated extended oligonucleotide comprises the region of interest or a complement of a sequence in the region of interest and a poly-T sequence. In some instances, the generated extended oligonucleotide comprises the region of interest or a complement of a sequence in the region of interest and a sequence complementary to the oligonucleotide.
[0168] In some embodiments, after extending the oligonucleotide using the RNA template the RNA template is digested. In some embodiments, cleavage of the RNA template prior to extension of the oligonucleotide and digestion of the RNA template after extending the oligonucleotide is performed using the same enzyme. In some embodiments, cleavage of the RNA template prior to extension of the oligonucleotide and digestion of the RNA template after extending the oligonucleotide is performed using different enzymes. In some embodiments, the method comprises contacting the biological sample with an endonuclease. In some embodiments, cleavage of the RNA template prior to extension of the oligonucleotide and digestion of the RNA template after extending the oligonucleotide is performed using RNase H (e.g., provided at two different times in the workflow). In some embodiments, the method comprises contacting the biological sample with an RNase H. Any suitable RNase H for digesting RNA in a nucleic acid duplex (e.g., with the extended oligonucleotide hybridized to the RNA template) can be used. The RNase H enzyme and its family of enzymes include two classes, type 1 and type 2 RNase H based on the difference in their amino acid sequence. Type 1 RNases H include prokaryotic and eukaryotic RNases Hl and retroviral RNase H. Type 2 RNases H include prokaryotic and eukaryotic RNases H2 and bacterial RNase H3. These RNases H exist in a monomeric form,56MOFO-358213026202412023740 except for eukaryotic RNases H2, which exist in a heterotrimeric form. All of these enzymes share the characteristic that they are able to cleave the RNA component of an RNA:DNA heteroduplex or within a DNA:DNA duplex containing RNA base(s) within one or both of the strands. The cleaved product yields a free 3'-OH for both classes of RNase H. In some embodiments, RNase Hl requires more than a single RNA base within an RNA:DNA duplex for optimal activity. In some embodiments, RNase HII requires only a single RNA base in an RNA:DNA duplex. In some embodiments, RCA is performed using the cleaved RNA as a primer.
[0169] In some embodiments, the RNase H enzyme comprises RNase Hl (commercially available from NEB, Inc.) or RNase H3. An example of an RNase H enzyme includes E. coli RNase HII (available for example from NEB, Inc. (product M0288)). In some embodiments, the RNase is an RNase HII. In some embodiments, the RNase can be E. coli RNase H (available, for example, from NEB, Inc., for example product M0297), which also cleaves a ribo-base when hybridized to DNA and leaves a 3 '-hydroxyl end. In some embodiments, the RNase H is a thermostable RNase H (available, for example, from NEB, e.g., product M0523). In some embodiments, the RNase H is a mammalian RNase H, such as any of those described in US20050164234, the content of which is herein incorporated by reference in its entirety.
[0170] An enzyme with RNase HII characteristics has been purified to near homogeneity from human placenta (Frank et al., Nucleic Acids Res., 1994, 22, 5247-5254). This protein has a molecular weight of approximately 33 kDa and is active in a pH range of 6.5-10, with a pH optimum of 8.5-9. The enzyme requires Mg2+and is inhibited by Mn2+and n-ethyl maleimide. The products of cleavage reactions have 3' hydroxyl and 5' phosphate termini. In some embodiments, the digesting and / or cleaving is performed at a temperature below 60° C (e.g. at room temperature or about 20-50° C, 20-30°C, 20-40°C, 25-40°C, 30-40°C, 35-40°C, or 40-50°C). In some embodiments, the digesting and / or cleaving is performed by incubating the biological sample with RNase H at about 37°C. In some embodiments, the digesting and / or cleaving is performed by incubating the biological sample with RNase H at about 37°C for a duration of between about 10 minutes and about 2 hours, between about 10-120, 10-90, 10-60, 10-30, 20-90, 20-60, 20-30, or 30-60 minutes. In some embodiments, the digesting and / or cleaving is performed by incubating the biological sample with RNase H at about 37°C for a57MOFO-358213026202412023740 duration of less than 30 minutes. In some embodiments, the digesting and / or cleaving is performed by incubating the biological sample with RNase H at about 37 °C for a duration of about 20 minutes.
[0171] In some embodiments, the provided methods comprise digesting the RNA template by contacting the biological sample with RNase H at a concentration of about at least 1 x 10'5, 1 x 10'4, 1 x 10'3, 1 x 10'2, 1 x 10’1, 1 U / pL, or higher. In some embodiments, the RNase H concentration is less than 1, 1 x 10’1, 1 x 10'2, 1 x 10'3, 1 x 10'4, 1 x 10'5U / pL, or less. In some embodiments, the RNase H concentration is between about 1 x 10'5and about 1 x 10'4, between about 1 x 10'4and about 1 x 10'3, between about 1 x 10'3and about 1 x 10'2, between about 1 x 10'2and about 1 x 10’1, between about 1 x 10'1and about 1, between about 1 x 10'5and about 1 x 10'3, between about 1 x 10'4and about 1 x 10'2, between about 1 x 10'3and about 1 x 10'1U / pL, or between about 1 x 10'2and about 1 U / pL. In some embodiments, the RNase H concentration is about 1 x 10'4, 3 x 10'4, 1 x 10'3, 3 x 10'3, 1 x 10'2, 3 x 10'2, or 1 x 10'1U / pL. In some embodiments, the RNase H concentration is about 1 x 10'2, 3 x 10'2, or 1 x 10'1U / pL.
[0172] In some embodiments, the provided methods comprise contacting the biological sample with between about 0.5 enzyme units (U) and about 100 U of the RNase H. In some embodiments, the biological sample is contacted with between about 0.5 and about 100, between about 0.5 and about 80, between about 0.5 and about 60, between about 0.5 and about 50 , between about 0.5 and about 40, between about 0.5 and about 30, between about 0.5 and about 20, between about 0.5 and about 10, between about 0.5 and about 8, between about 0.5 and about 5 , between about 0.5 and about 5, between about 0.5 and about 3, between about 2 and about 10, between about 2 and about 8, between about 2 and about 6, between about 2 and about 5, between about 2 and about 4, between about 3 and about 8, between about 3 and about 6, between about 4 and about 8, between about 4 and about 6, between about 4 and about 5.5, between about 4.5 and about 5.5, between about 4.5 and about 6, between about 10 and about 100, between about 10 and about 50, between about 10 and about 30, between about 10 and about 20, between about 20 and about 100, between about 20 and about 50, between about 20 and about 40, between about 20 and about 30, between about 30 and about 100, between about 30 and about 80, or between about 30 and about 50 U of RNase H. In some embodiments, the amount of RNase H contacted with the biological sample is dependent on the amount of RNA template to be digested in the biological sample.58MOFO-358213026202412023740
[0173] In some embodiments, the RNase H is incubated with the biological sample in a buffer comprising magnesium chloride. In some embodiments, the RNase H is incubated with the biological sample in a buffer comprising magnesium chloride, potassium chloride, dithiothreitol (DTT), and a buffering agent (e.g., Tris-HCL).
[0174] In some embodiments, the RNase H comprises an RNase Hl and / or an RNase H2. In some embodiments, the method comprises contacting the biological sample with an RNase Hl and an RNase H2. In some embodiments, the RNase H is RNase Hl. In some embodiments, the RNase H is an endoribonuclease that specifically hydrolyzes the phosphodiester bonds of RNA which is hybridized to DNA.
[0175] In some embodiments, the endonuclease (e.g., RNase H) does not digest single or double-stranded DNA. In some embodiments, the RNase H requires at least four contiguous bases of RNA for digestion of the RNA hybridized to DNA. In some embodiments, the RNase H does not digest a di-ribonucleotide-containing DNA sequence (e.g., one that is DNA-annealed). In some embodiments, the RNase H does not digest a mono-ribonucleotidecontaining DNA sequence (e.g., one that is DNA-annealed). In some embodiments, the RNase H digests a RNA template but does not digest the extended oligonucleotide hybridized to the RNA. In some embodiments, the RNase H cleaves at a ribonucleotide in the extended oligonucleotide but does not digest the extended oligonucleotide hybridized to the RNA. In some embodiments, the RNase H cleaves at a ribonucleotide in the extended oligonucleotide and digests the RNA template but does not digest the extended oligonucleotide hybridized to the RNA.
[0176] In some embodiments, the RNase H is RNase H2. In some embodiments, the RNase H is an endoribonuclease that preferentially nicks 5' to one or more ribonucleotides (e.g., a single ribonucleotide, a diribonucleotide sequence, etc.) within the context of a DNA duplex, leaving 5' phosphate and 3' hydroxyl ends. In some embodiments, the RNase H nicks at multiple sites along an RNA portion hybridized to a DNA. In some embodiments, the RNase H cleaves an RNA portion of the extended oligonucleotide comprising one or more incorporated ribonucleotides. In some embodiments, the RNase H digests a DNA-annealed di-ribonucleotide- containing DNA sequence, whereas the RNA-annealed di-ribonucleotide-containing DNA sequence is not digested. In some embodiments, the RNase H digests a DNA-annealed mono-59MOFO-358213026202412023740 ribonucleotide-containing DNA sequence, whereas RNA-annealed mono-ribonucleotide- containing DNA sequence is not digested.
[0177] In some embodiments, the RNase H cleaves RNA in RNA-DNA duplexes. In some embodiments, the RNase H is a bacterial RNase H or analog or derivative thereof (e.g., Escherichia coli RNase H or an analog or derivative thereof). In some embodiments, the RNase H is a eukaryotic RNase H or analog or derivative thereof (e.g., human RNase H or an analog or derivative thereof). In some embodiments, the RNase H is a viral RNase H or analog or derivative thereof, such as an HIV-derived RNase H.
[0178] In some embodiments, digestion of the RNA template and cleavage of the extended oligonucleotide at the incorporated ribonucleotide, inosine, or uracil is performed using the same enzyme. In some embodiments, digestion of the RNA template and cleavage of the extended oligonucleotide at the incorporated ribonucleotide, inosine, or uracil is performed using different enzymes. In some embodiments, digestion of the RNA template and cleavage of the extended oligonucleotide at the incorporated ribonucleotide, inosine, or uracil is performed in the same step. In some embodiments, digestion of the RNA template and cleavage of the extended oligonucleotide at the incorporated ribonucleotide, inosine, or uracil is performed at two sequential steps.B. Kinase Treatment
[0179] In some cases, after extension of the oligonucleotide, the extended oligonucleotide is used to generate a circularized template. In some instances, the 5’ and 3’ ends of the extended oligonucleotide are ligated generate a circularized template. In some instances, a 5’ end and a 3’ end of the cleaved extended sequence is ligated to generate the circularized template.
[0180] In some aspects, before performing ligation to generate the circularized template which requires a 3 ’OH and 5’ monophosphate for the ligase, the extended sequence of the oligonucleotide is treated with a kinase to phosphorylate the end of the extended sequence. In some aspects, the extended sequence is processed to restore a 3 ’OH at the end of the extended sequence. In some cases, the extended oligonucleotide is processed (e.g., treated to remove any terminating groups or modifications) and contacted with a kinase. In some instances, a kinase is60MOFO-358213026202412023740 used to phosphorylate a 5’ end of the extended oligonucleotide prior to generating the circularized template.
[0181] In some embodiments, the method comprises contacting the biological sample with a polynucleotide kinase (PNK). In some embodiments, the sample is contacted with at least any of 1 U, 5 U, 10 U, 50 U, 100 U, 200 U, 300 U, 400 U, 500 U, 600 U, 700 U, 800 U, 900 U, or 1000 U of a kinase. In some embodiments, at least any of 1 U / pL, 2 U / pL, 3 U / pL, 4 U / pL, 5 U / pL, 6 U / pL, 7 U / pL, 8 U / pL, 9 U / pL, 10 U / pL, 20 U / pL, 30 U / pL, 40 U / pL, or 50 U / pL of a polynucleotide kinase as the final concentration is used. One unit of PNK activity is defined as the amount of enzyme (measured in units, U) that will catalyze the transfer of 1 nmol of phosphate from ATP to the 5 '-OH end of a polynucleotide in 30 minutes at an optimum temperature for the enzyme, usually 37 °C. In some embodiments, the PNK transfers the gamma phosphate group from adenosine triphosphate (ATP) to the 5’ hydroxyl termini of DNA or RNA (e.g., extended oligonucleotide described herein). In some embodiments, the kinase is an enzyme from the family of transferases that transfer phosphorus-containing groups to alcohol groups (phosphotransferases). Systematically, the kinase class may be from the enzyme class known as ATP:5 '-diphospho polynucleotide 5 '-phosphotransferase. In some embodiments, the kinase is a T4 PNK or a variant thereof.
[0182] In some embodiments, phosphorylation of the extended oligonucleotide comprises incubating the sample with the kinase (e.g., polynucleotide kinase (PNK)). In some embodiments, the method comprises incubating the sample with the kinase for at least 20 minutes, at least 30 minutes, at least 35 minutes, at least 40 minutes, at least 45 minutes, at least 50 minutes, at least 60 minutes, at least 80 minutes, at least 100 minutes, or at least 120 minutes. In some embodiments, the incubation with the kinase is performed for 20-60 minutes, 20-45 minutes, 20-120 minutes, 30-120 minutes, 30-60 minutes, or 30-90 minutes. In some embodiments, the incubation with kinase is performed at 30° C to 40° C, e.g., at 37° C.
[0183] In some embodiments, the extended oligonucleotide described herein, comprises a non-ligatable end before the ligation to generate a circularized template. In some instances, the kinase treatment converts one or more non-ligatable ends to a ligatable end with a 5’ phosphate. In some embodiments, a kinase treatment step is before and / or after one or more wash steps. In some embodiments, the wash step removes unused reagents (e.g., from the61MOFO-358213026202412023740 extension) or other products (e.g., from the processing to remove a terminating group or modification from the extended oligonucleotide) from the sample.
[0184] In some embodiments, the kinase is provided in excess to phosphorylate the extended oligonucleotide. In some embodiments, the kinase is not endogenous to the sample, e.g., the kinase is exogenously provided to the sample after generating the extended oligonucleotide.C. Circularization and / or Amplification
[0185] In some embodiments, provided herein are methods and compositions for analyzing one or more products of an endogenous analyte. In some embodiments, an endogenous RNA analyte is used as a RNA template to generate a rolling circle amplification (RCA) product after performing extension to generate the extended oligonucleotide. In some embodiments, a cleaved endogenous RNA analyte is used as a RNA template to generate the extended oligonucleotide. In some embodiments, provided herein are methods and compositions for analyzing one or more products of an exogenous molecule introduced into the biological sample (e.g., a viral transcript or an exogenous library of nucleic acid constructs introduced to the biological sample). In some embodiments, a generated extended oligonucleotide or a derivative thereof is processed (e.g., amplified and / or captured as described in Section III).
[0186] In some embodiments, an exogenous analyte or a nucleic acid molecule associated with an exogenous analyte is used as a template to generate a rolling circle amplification (RCA) product after performing extension to generate the extended oligonucleotide. In some embodiments, the method comprises ligating the ends of the extended oligonucleotide to generate the circularized template. In some embodiments, the circularized template comprises at least a portion of the extended oligonucleotide. In some embodiments, the circularized template comprises the sequence of at least a portion of the extended oligonucleotide. In some embodiments, the circularized template comprises a sequence complementary to at least a portion of the extended oligonucleotide. In some embodiments, the method comprises ligating the ends of a copy of the extended oligonucleotide to generate the circularized template. For example, the extended oligonucleotide is used as a template to generate a copy of the extended oligonucleotide and a 5’ end and a 3’ end of the copy of the extended oligonucleotide is ligated to generate the circularized template. In some instances, a 5’ end and a 3’ end of the cleaved extended sequence (e.g., as described in Section II. A) is ligated62MOFO-358213026202412023740 to generate the circularized template. In some instances, the ligating is performed enzymatically. In some instances, the ligating is performed by chemical ligation. For example, the ligating is performed using CircLigase. In some instances, the ligating is performed using a splint. In some aspects, the ligating is performed using reactive moieties to react to provide a linking moiety. For example, a click chemistry reaction involving an alkyne moiety and an azide moiety may be used to provide a triazole linking moiety.
[0187] In some embodiments, the method comprises introducing one or more functional sequences to the extended oligonucleotide or a copy thereof to generate the circularized template. For example, one or more functional sequences that can be used in subsequent processing, such as an adapter sequence, a unique molecular identifier (UMI) sequence, a primer or primer binding sequence, a sequencing primer or primer binding sequence is introduced to the extended oligonucleotide or a copy thereof. In some embodiments, the method comprises introducing a 5’ handle and / or a 3’ sequencing handle to the extended oligonucleotide or a copy thereof to generate the circularized template. In some embodiments, the one or more functional sequences are used to perform the ligation (e.g., ligating the ends of the extended oligonucleotide) to generate the circularized template. In some embodiments, the one or more functional sequences introduced to the extended oligonucleotide are used to perform the ligation using a splint with complementary sequences to the functional sequence(s) or a portion thereof to generate the circularized template.
[0188] In some embodiments, the ligation involves chemical ligation (e.g., click chemistry ligation). In some embodiments, the chemical ligation involves template dependent ligation. In some embodiments, the chemical ligation involves template independent ligation. In some embodiments, the click reaction is a template-independent reaction. In some embodiments, the click reaction is a template-dependent reaction or template-directed reaction. In some embodiments, the template-dependent reaction is sensitive to base pair mismatches such that reaction rate is significantly higher for matched versus unmatched templates. In some embodiments, the click reaction is a nucleophilic addition template-dependent reaction. In some embodiments, the click reaction is a cyclopropane-tetrazine template-dependent reaction.
[0189] In some embodiments, the ligation involves enzymatic ligation. In some embodiments, the enzymatic ligation involves use of a ligase. In some aspects, the ligase used herein comprises an enzyme that is commonly used to join polynucleotides together or to join the63MOFO-358213026202412023740 ends of a single polynucleotide. An RNA ligase, a DNA ligase, or another variety of ligase can be used to ligate two nucleotide sequences together. Ligases comprise ATP-dependent doublestrand polynucleotide ligases, NAD-i-dependent double-strand DNA or RNA ligases and singlestrand polynucleotide ligases, for example any of the ligases described in EC 6.5.1.1 (ATP- dependent ligases), EC 6.5.1.2 (NAD+-dependent ligases), EC 6.5.1.3 (RNA ligases). Specific examples of ligases comprise bacterial ligases such as E. coli DNA ligase, Tth DNA ligase, Thermococcus sp. (strain 9° N) DNA ligase (9°N™ DNA ligase, New England Biolabs), Taq DNA ligase, Ampligase™ (Epicentre Biotechnologies) and phage ligases such as T3 DNA ligase, T4 DNA ligase and T7 DNA ligase and mutants thereof. In some embodiments, the ligase is a T4 RNA ligase. In some embodiments, the ligase is a splintR ligase. In some embodiments, the ligase is a single stranded DNA ligase. In some embodiments, the ligase is a T4 DNA ligase. In some embodiments, the ligase is a ligase that has an DNA-splinted DNA ligase activity. In some embodiments, the ligase is a ligase that has an RNA-splinted DNA ligase activity.
[0190] In some embodiments, the ligation herein is a direct ligation. In some embodiments, the ligation herein is an indirect ligation. "Direct ligation" means that the ends of the polynucleotides hybridize immediately adjacently to one another to form a substrate for a ligase enzyme resulting in their ligation to each other (intramolecular ligation). Alternatively, "indirect" means that the ends of the polynucleotides hybridize non-adjacently to one another, e.g., separated by one or more intervening nucleotides or "gaps". In some embodiments, the ends are not ligated directly to each other, but instead occurs either via the intermediacy of one or more intervening (so-called "gap" or "gap-filling" (oligo)nucleotides) or by the extension to "fill" the "gap".
[0191] In some aspects, a high fidelity ligase, such as a thermostable DNA ligase (e.g., a Taq DNA ligase), is used. Thermostable DNA ligases are active at elevated temperatures, allowing further discrimination by incubating the ligation at a temperature near the melting temperature (Tm) of the DNA strands. This selectively reduces the concentration of annealed mismatched substrates (expected to have a slightly lower Tmaround the mismatch) over annealed fully base-paired substrates. Thus, high-fidelity ligation can be achieved through a combination of the intrinsic selectivity of the ligase active site and balanced conditions to reduce the incidence of annealed mismatched dsDNA.64MOFO-358213026202412023740
[0192] In some embodiments, after generating the circularized template, an amplification reaction is performed. In some embodiments, a rolling circle amplification (RCA) reaction is performed using the circularized template. In some instances, the circularized template comprises a primer binding sequence and the method comprises binding a primer to the primer binding sequence, and extending the primer to generate an amplification product comprising multiple copies of the RNA template or a complement thereof. In some embodiments, the amplification product comprises a sequence complementary to the extended oligonucleotide.
[0193] A primer is generally a single-stranded nucleic acid sequence having a 3’ end that, in some embodiments, is used as a substrate for a nucleic acid polymerase in a nucleic acid extension reaction. RNA primers are formed of RNA nucleotides, and are used in RNA synthesis, while DNA primers are formed of DNA nucleotides and used in DNA synthesis. Primers can also include both RNA nucleotides and DNA nucleotides (e.g., in a random or designed pattern). Primers can also include other natural or synthetic nucleotides described herein that can have additional functionality. In some examples, DNA primers can be used to prime RNA synthesis and vice versa (e.g., RNA primers can be used to prime DNA synthesis). Primers can vary in length. For example, primers can be about 6 bases to about 120 bases. For example, primers can include up to about 25 bases. A primer, in some cases, refers to a primer binding sequence. A primer extension reaction generally refers to any method where two nucleic acid sequences become linked (e.g., hybridized) by an overlap of their respective terminal complementary nucleic acid sequences (e.g., 3’ termini). Such linking can be followed by nucleic acid extension (e.g., an enzymatic extension) of one, or both termini using the other nucleic acid sequence as a template for extension. In some embodiments, enzymatic extension is performed by an enzyme including, but not limited to, a polymerase.
[0194] In some embodiments, the amplifying is achieved by performing rolling circle amplification (RCA). In some embodiments, a primer that hybridizes to the circularized template is added and used as such for amplification. In some embodiments, the RCA comprises a linear RCA, a branched RCA, a dendritic RCA, or any combination thereof. In some embodiments, rolling circle amplification of the circularized template is performed.
[0195] In some embodiments, the amplification is performed at a temperature between or between about 20°C and about 60°C. In some embodiments, the amplification is65MOFO-358213026202412023740 performed at a temperature between or between about 30°C and about 40°C. In some aspects, the amplification step, such as the rolling circle amplification (RCA) is performed at a temperature between at or about 25 °C and at or about 50°C, such as at or about 25 °C, at or about 27°C, at or about 29°C, at or about 31 °C, at or about 33°C, at or about 35°C, at or about 37°C, at or about 39°C, at or about 41°C, at or about 43°C, at or about 45°C, at or about 47°C, or at or about 49°C.
[0196] In some embodiments, upon addition of a DNA polymerase in the presence of appropriate dNTP precursors and other cofactors, a primer is elongated to produce multiple copies of the circular template. In some embodiments, this amplification step can utilize isothermal amplification or non-isothermal amplification. In some embodiments, a template is rolling-circle amplified to generate a cDNA nanoball (e.g., amplicon) containing multiple copies of the RNA template or a complement thereof. Techniques for rolling circle amplification (RCA) include linear RCA, a branched RCA, a dendritic RCA, or any combination thereof. Exemplary polymerases for use in RCA comprise DNA polymerase such phi29 (cp29) polymerase, Klenow fragment, Bacillus stearothermophilus DNA polymerase (BST), T4 DNA polymerase, T7 DNA polymerase, or DNA polymerase I. In some aspects, DNA polymerases that have been engineered or mutated to have desirable characteristics can be employed. In some embodiments, the polymerase is a phi29 DNA polymerase.
[0197] In some aspects, during the amplification step, modified nucleotides are added to the reaction to incorporate the modified nucleotides in the amplification product (e.g., nanoball). Exemplary of the modified nucleotides comprise amine-modified nucleotides. In some aspects of the methods, for example, for anchoring or cross-linking of the generated amplification product (e.g., nanoball) to a scaffold, to cellular structures and / or to other amplification products (e.g., other nanoballs). In some aspects, the amplification products comprises a modified nucleotide, such as an amine-modified nucleotide. In some embodiments, the amine-modified nucleotide comprises an acrylic acid N- hydroxy succinimide moiety modification. Examples of other amine-modified nucleotides comprise, but are not limited to, a5-Aminoallyl-dUTP moiety modification, a 5-Propargylamino-dCTP moiety modification, a N6-6-Aminohexyl-dATP moiety modification, or a 7-Deaza-7-Propargylamino-dATP moiety modification.66MOFO-358213026202412023740
[0198] In some aspects, the amplification products (e.g., amplicon) are anchored to a polymer matrix. For example, the polymer matrix can be a hydrogel. Examples of modification and polymer matrix that can be employed in accordance with the provided embodiments comprise those described in, for example, WO 2017 / 079406, US 2016 / 0024555, US 2018 / 0251833 and US 2017 / 0219465, which are herein incorporated by reference in their entireties. In some examples, the scaffold also contains modifications or functional groups that can react with or incorporate the modifications or functional groups of the amplification product. In some examples, the scaffold can comprise oligonucleotides, polymers or chemical groups, to provide a matrix and / or support structures.
[0199] The amplification products may be immobilized within the matrix generally at the location of the circularized template being amplified, thereby creating a localized colony of amplicons. The amplification products may be immobilized within the matrix by steric factors. The amplification products may also be immobilized within the matrix by covalent or noncovalent bonding. In this manner, the amplification products may be considered to be attached to the matrix. By being immobilized to the matrix, such as by covalent bonding or cross-linking, the size and spatial relationship of the original amplicons is maintained. By being immobilized to the matrix, such as by covalent bonding or cross-linking, the amplification products are resistant to movement or unraveling under mechanical stress.
[0200] In some aspects, the amplification products are copolymerized and / or covalently attached to the surrounding matrix thereby preserving their spatial relationship and any information inherent thereto. For example, if the amplification products are those generated from DNA or RNA within a cell embedded in the matrix, the amplification products can also be functionalized to form covalent attachment to the matrix preserving their spatial information within the cell thereby providing a subcellular localization distribution pattern. In some embodiments, the provided methods involve embedding the one or more of the amplification products in the presence of hydrogel subunits to form one or more hydrogel-embedded amplification products. In some embodiments, the hydrogel-tissue chemistry described comprises covalently attaching nucleic acids to in situ synthesized hydrogel for tissue clearing, enzyme diffusion, and multiple-cycle sequencing while an existing hydrogel-tissue chemistry method cannot. In some embodiments, to enable amplification product embedding in the tissuehydrogel setting, amine-modified nucleotides are comprised in the amplification step (e.g.,67MOFO-358213026202412023740RCA), functionalized with an acrylamide moiety using acrylic acid N-hydroxysuccinimide esters, and copolymerized with acrylamide monomers to form a hydrogel.
[0201] In some embodiments, the RCA template may comprise a sequence of the RNA template, or a part thereof, or it may be provided or generated as a proxy, or a marker, for the analyte. In some embodiments, different analytes are detected in situ in one or more cells using a RCA-based detection system, e.g., where the signal is provided by generating an RCA product from a circular RCA template which is provided or generated in the assay, and the RCA product is detected to detect the corresponding analyte. The RCA product may thus be regarded as a reporter which is detected to detect the RNA analyte. However, the RCA template may also be regarded as a reporter for the RNA analyte; the RCA product is generated based on the RCA template, and comprises complementary copies of the RCA template. The RCA template determines the signal which is detected, and is thus indicative of the RNA analyte. The RCA template used to generate the RCP may thus be a circular (e.g. circularized) reporter nucleic acid molecule, namely from any RCA-based detection assay which uses or generates a circular nucleic acid molecule as a reporter for the assay. Since the RCA template generates the RCP reporter, it may be viewed as part of the reporter system for the assay.
[0202] In some embodiments, provided herein is a method for detecting a molecule or a complex generated in a series of reactions, e.g., hybridization, ligation, extension, replication, transcription / reverse transcription, and / or amplification (e.g., rolling circle amplification), in any suitable combination. III. Detection and Analysis
[0203] In some embodiments, a method disclosed herein comprises processing and / or detecting amplification products of the extended oligonucleotides generated using the RNA templates as described in Section II. In some embodiments, a method disclosed herein comprises capturing the extended oligonucleotides generated using the RNA templates or derivatives thereof (e.g., cDNA generated using the RNA templates or an oligonucleotide or polynucleotide generated from the cDNA) on a support. In some instances, the extended oligonucleotides generated using the RNA templates or derivatives thereof are subject to spatial analysis methods and compositions that include, e.g., the use of a capture probe including a spatial barcode (e.g., a nucleic acid sequence that provides information as to the location or position of an analyte within a cell or a tissue sample) and a capture domain that is capable of68MOFO-358213026202412023740 binding to an analyte (e.g., directly or indirectly binding to the RNA template or a derivative thereof). In some instances, the extended oligonucleotides generated using the RNA templates or derivatives thereof are further processed with a capture probe having a capture domain that captures an intermediate agent for indirect detection of an analyte.
[0204] In some instances, the extended oligonucleotides generated using the RNA templates or derivatives thereof are processed with nucleic acid barcode molecules for single cell analysis. In some instances, the extended oligonucleotides are generated using the RNA templates by extending a barcoded oligonucleotide (e.g., a nucleic acid barcode molecule) for single cell analysis. In some instances, the process for generating extended oligonucleotides using RNA templates are used with systems and methods described herein that provide for the compartmentalization, depositing, or partitioning of one or more particles (e.g., biological particles, macromolecular constituents of biological particles, beads, reagents, etc.) into discrete compartments or partitions (referred to interchangeably herein as partitions), where each partition maintains separation of its own contents from the contents of other partitions.A. In situ Analysis
[0205] In some embodiments, a method disclosed herein comprises generating rolling circle amplification (RCA) products using the circularized template associated with one or more RNA templates in a sample. In some embodiments, the RCA products are detected in situ in a sample, thereby detecting the one or more endogenous nucleic acids. In some embodiments, each of the RCA products comprises multiple complementary copies of the RNA template. In some instances, a RCP is generated and detected at a location in the biological sample.
[0206] In some embodiments, a sequence of the RCP is analyzed at a location in the biological sample or a matrix embedding the biological sample. In some embodiments, the sequence of the RCP is analyzed by sequencing by ligation, sequencing by synthesis, sequencing by binding, sequencing by avidity or a combination thereof.
[0207] In some embodiments, the analysis comprises detecting a sequence e.g., an endogenous RNA sequence present in the sample. In some embodiments, the analysis includes quantification of puncta (e.g., if amplification products are detected). In some cases, the analysis includes determining whether particular cells and / or signals are present that correlate with one or more biomarkers from a particular panel. In some embodiments, the obtained information may be compared to a positive and negative control, or to a threshold of a feature to determine if the69MOFO-358213026202412023740 sample exhibits a certain feature or phenotype. In some cases, the information may comprise signals from a cell, a region, and / or comprise readouts from multiple detectable labels. In some case, the analysis further includes displaying the information from the analysis or detection step. In some embodiments, software may be used to automate the processing, analysis, and / or display of data.
[0208] In some embodiments, the detection is spatial, e.g., in two or three dimensions. In some embodiments, the detection is quantitative, e.g., the amount or concentration of an analyte is determined.
[0209] In some embodiments, the analysis and / or sequence determination comprises detecting all or a portion of the RCP(s). In some embodiments, the sequencing step involves sequencing by ligation and / or fluorescent in situ sequencing. In some embodiments, the detection or determination comprises imaging the RCP. In some embodiments, the analyte is an mRNA in a tissue sample, and the detection or determination is performed when the RCP is in situ at the location of the mRNA in the tissue sample.
[0210] In some aspects, the provided methods comprise imaging the RCP. In some embodiments, confocal microscopy is used for detection and imaging. Confocal microscopy uses point illumination and a pinhole in an optically conjugate plane in front of the detector to eliminate out-of-focus signal. As only light produced by fluorescence very close to the focal plane can be detected, the image’s optical resolution, particularly in the sample depth direction, is much better than that of wide-field microscopes. However, as much of the light from sample fluorescence is blocked at the pinhole, this increased resolution is at the cost of decreased signal intensity, so long exposures are often required. As only one point in the sample is illuminated at a time, 2D or 3D imaging requires scanning over a regular raster (e.g., a rectangular pattern of parallel scanning lines) in the specimen. The achievable thickness of the focal plane is defined mostly by the wavelength of the used light divided by the numerical aperture of the objective lens, but also by the optical properties of the specimen. The thin optical sectioning possible makes these types of microscopes particularly good at 3D imaging and surface profiling of samples. CLARITY™-optimized light sheet microscopy (COLM) provides an alternative microscopy for fast 3D imaging of large, clarified samples. COLM interrogates large immunostained tissues, permits increased speed of acquisition and results in a higher quality of generated data.70MOFO-358213026202412023740
[0211] Other types of microscopy that can be employed comprise bright field microscopy, oblique illumination microscopy, dark field microscopy, phase contrast, differential interference contrast (DIC) microscopy, interference reflection microscopy (also known as reflected interference contrast, or RIC), single plane illumination microscopy (SPIM), superresolution microscopy, laser microscopy, electron microscopy (EM), Transmission electron microscopy (TEM), Scanning electron microscopy (SEM), reflection electron microscopy (REM), Scanning transmission electron microscopy (STEM) and low- voltage electron microscopy (LVEM), scanning probe microscopy (SPM), atomic force microscopy (ATM), ballistic electron emission microscopy (BEEM), chemical force microscopy (CFM), conductive atomic force microscopy (C-AFM), electrochemical scanning tunneling microscope (ECSTM), electrostatic force microscopy (EFM), fluidic force microscope (FluidFM), force modulation microscopy (FMM), feature-oriented scanning probe microscopy (FOSPM), kelvin probe force microscopy (KPFM), magnetic force microscopy (MFM), magnetic resonance force microscopy (MRFM), near-field scanning optical microscopy (NSOM) (or SNOM, scanning near-field optical microscopy, SNOM, Piezoresponse Force Microscopy (PFM), PSTM, photon scanning tunneling microscopy (PSTM), PTMS, photothermal micro spectroscopy / microscopy (PTMS), SCM, scanning capacitance microscopy (SCM), SECM, scanning electrochemical microscopy (SECM), SGM, scanning gate microscopy (SGM), SHPM, scanning Hall probe microscopy (SHPM), SICM, scanning ion-conductance microscopy (SICM), SPSM spin polarized scanning tunneling microscopy (SPSM), SSRM, scanning spreading resistance microscopy (SSRM), SThM, scanning thermal microscopy (SThM), STM, scanning tunneling microscopy (STM), STP, scanning tunneling potentiometry (STP), SVM, scanning voltage microscopy (SVM), and synchrotron x-ray scanning tunneling microscopy (SXSTM), and intact tissue expansion microscopy (exM).
[0212] In some embodiments, sequencing is performed in situ. In situ sequencing typically involves incorporation of a labeled nucleotide (e.g., fluorescently labeled mononucleotides or dinucleotides) in a sequential, template-dependent manner or hybridization of a labeled primer (e.g., a labeled random hexamer) to a nucleic acid template such that the identities (e.g., nucleotide sequence) of the incorporated nucleotides or labeled primer extension products can be determined, and consequently, the nucleotide sequence of the corresponding template nucleic acid. Aspects of in situ sequencing are described, for example, in Mitra et al.,71MOFO-358213026202412023740(2003) Anal. Biochem. 320, 55-65, and Lee et al., (2014) Science, 343(6177), 1360-1363. In addition, examples of methods and systems for performing in situ sequencing are described in US 2016 / 0024555, US 2019 / 0194709, and in US 10,138,509, US 10,494,662 and US. 10,179,932, all of which are herein incorporated by reference in their entireties.
[0213] In some embodiments, analyzing, e.g., detecting or determining, one or more sequences present in the biological sample is performed using a base-by-base sequencing method, e.g., sequencing -by- synthesis (SBS), sequencing-by-avidity (SBA) or sequencing-by- binding (SBB). In some embodiments, the biological sample is contacted with a sequencing primer and base-by-base sequencing using a cyclic series of nucleotide incorporation or binding, respectively, thereby generating extension products of the sequencing primer is performed followed by removing, cleaving, or blocking the extension products of the sequencing primer.
[0214] Generally in sequencing-by- synthesis methods, a first population of detectably labeled nucleotides (e.g., dNTPs) are introduced to contact a template nucleotide (e.g., a barcode sequence in the RCP) hybridized to a sequencing primer, and a first detectably labeled nucleotide (e.g., A, T, C, or G nucleotide) is incorporated by a polymerase to extend the sequencing primer in the 5’ to 3’ direction using a complementary nucleotide (a first nucleotide residue) in the template nucleotide as template. A signal from the first detectably labeled nucleotide can then be detected. The first population of nucleotides may be continuously introduced, but in order for a second detectably labeled nucleotide to incorporate into the extended sequencing primer, nucleotides in the first population of nucleotides that have not incorporated into a sequencing primer are generally removed (e.g., by washing), and a second population of detectably labeled nucleotides are introduced into the reaction. Then, a second detectably labeled nucleotide (e.g., A, T, C, or G nucleotide) is incorporated by the same or a different polymerase to extend the already extended sequencing primer in the 5’ to 3’ direction using a complementary nucleotide (a second nucleotide residue) in the template nucleotide as template. Thus, in some embodiments, cycles of introducing and removing detectably labeled nucleotides are performed.
[0215] In some embodiments, the detectably labeled nucleotide comprises a fluorescent dye. Commercially available fluorescent dyes include, but are not limited to 4,7- dichlorofluorescein dyes, spectrally resolvable rhodamine dyes, 4,7- dichlororhodamine dyes, cyanine dyes, ether-substituted fluorescein dyes, energy transfer dyes, and xanthine dyes. Labeling can also be carried out with quantum dots. In some embodiments, a fluorescent label72MOFO-358213026202412023740 comprises a signaling moiety that conveys information through the fluorescent absorption and / or emission properties of one or more molecules. Exemplary fluorescent properties comprise fluorescence intensity, fluorescence lifetime, emission spectrum characteristics and energy transfer.
[0216] In some embodiments, the detectably labeled nucleotide is a fluorescent nucleotide analog. Without being bound by theory, a fluorescent nucleotide analog is readily incorporated into a nucleotide, polynucleotide, and / or oligonucleotide sequence during extension. Examples of commercially available fluorescent nucleotide analogues readily incorporated into nucleotide and / or polynucleotide sequences comprise, but are not limited to, Cy3™-dNTP (cyanine 3-dNTP), Cy3™-dNTP (cyanine 3-dNTP), Cy5™-dNTP (cyanine 5- dNTP), Cy5™-dNTP (cyanine 5 dNTP) (Amersham Biosciences, Piscataway, N .J.), fluorescein- 12-dNTP, tetramethylrhodamine-6-dNTP, TEXAS RED®-5-dNTP (red fluorescent dye-dNTP), CASCADE® BLUE-7-dNTP (blue fluorescent dye - dNTP), BODIPY™ FL-14-dNTP (green fluorescent dye-dNTP), BODIPY™ TMR-14-dNTP (orange fluorescent dye-dNTP), BODIPY™ TR- 14-dNTP (red fluorescent dye-dNTP), RHODAMINE GREEN™-5-dNTP (green fluorescent dye-dNTP), OREGON GREEN™ 488-5-dNTP (green fluorescent dye-dNTP), TEXAS RED™- 12-dNTP (red fluorescent dye-dNTP), BODIPY™ 630 / 650- 14-dNTP (far red fluorescent dye- dNTP), BODIPY™ 650 / 665- 14-dNTP (far red fluorescent dye-dNTP), ALEXA FLUOR™ 488- 5-dNTP (green fluorescent dye-dNTP), ALEXA FLUOR™ 532-5-dNTP (yellow fluorescent dye-dNTP), ALEXA FLUOR™ 568-5-dNTP (red / orange fluorescent dye-dNTP), ALEXA FLUOR™ 594-5-dNTP (red fluorescent dye-dNTP), ALEXA FLUOR™ 546- 14-dNTP (orange fluorescent dye-dNTP), fluorescein- 12-UTP, tetramethylrhodamine-6-UTP, TEXAS RED™-5- UTP (red fluorescent dye-UTP), mCherry, CASCADE® BLUE-7-UTP (blue fluorescent dye- UTP), BODIPY™ FL-14-UTP (green fluorescent protein-UTP), BODIPY™ TMR- 14-UTP (orange fluorescent dye-UTP), BODIPY™ TR-14-UTP (red fluorescent dye-UTP), RHODAMINE GREEN™-5-UTP (green fluorescent dye-UTP), ALEXA FLUOR™ 488-5-UTP (green fluorescent dye-UTP), and ALEXA FLUOR™ 546- 14-UTP (orange fluorescent dye- UTP) (Molecular Probes, Inc. Eugene, Oreg.). Any suitable methods for custom synthesis of nucleotides having other fluorophores may be utilized.
[0217] In some embodiments, the base-by-base sequencing comprises using a polymerase that is fluorescently labeled. In some embodiments, the base-by-base sequencing73MOFO-358213026202412023740 comprises using a polymerase-nucleotide conjugate comprising a fluorescently labeled polymerase linked to a nucleotide moiety that is not fluorescently labeled. In some embodiments, the base-by-base sequencing comprises using a multivalent polymer-nucleotide conjugate comprising a polymer core, multiple nucleotide moieties, and one or more fluorescent labels.
[0218] In some embodiments, sequencing is performed by sequencing-by-synthesis (SBS). In some embodiments, a sequencing primer is complementary to sequences at or near the one or more barcode(s). In such embodiments, sequencing-by-synthesis can comprise reverse transcription and / or amplification in order to generate a template sequence from which a primer sequence can bind. Exemplary SBS methods comprise those described for example, but not limited to, US 2007 / 0166705, US 2006 / 0188901, US 7,057,026, US 2006 / 0240439, US 2006 / 0281109, US 2011 / 005986, US 2005 / 0100900, US 9,217,178, US 2009 / 0118128, US 2012 / 0270305, US 2013 / 0260372, and US 2013 / 0079232, all of which are herein incorporated by reference in their entireties.
[0219] In some embodiments, sequencing is performed by sequencing-by-binding (SBB). Various aspects of SBB are described in U.S. Pat. No. 10,655,176 B2, the content of which is herein incorporated by reference in its entirety. In some embodiments, SBB comprises performing repetitive cycles of detecting a stabilized complex that forms at each position along the template nucleic acid to be sequenced (e.g. a ternary complex that includes the primed template nucleic acid, a polymerase, and a cognate nucleotide for the position), under conditions that prevent covalent incorporation of the cognate nucleotide into the primer, and then extending the primer to allow detection of the next position along the template nucleic acid. In the sequencing-by-binding approach, detection of the nucleotide at each position of the template occurs prior to extension of the primer to the next position. Generally, the methodology is used to distinguish the four different nucleotide types that can be present at positions along a nucleic acid template by uniquely labelling each type of ternary complex (e.g., different types of ternary complexes differing in the type of nucleotide it contains) or by separately delivering the reagents needed to form each type of ternary complex. In some instances, the labeling may comprise fluorescence labeling of, e.g., the cognate nucleotide or the polymerase that participate in the ternary complex.
[0220] In some embodiments, sequencing is performed by sequencing-by-avidity (SBA). Some aspects of SBA approaches are described in U.S. Pat. No. 10,768,173 B2, the74MOFO-358213026202412023740 content of which is herein incorporated by reference in its entirety. In some embodiments, SBA comprises detecting a multivalent binding complex formed between a fluorescently-labeled polymer-nucleotide conjugate, and a one or more primed nucleic acid sequences (e.g., RNA sequences). Fluorescence imaging is used to detect the bound complex and thereby determine the identity of the N+l nucleotide in the RNA (where the primer extension strand is N nucleotides in length). Following the imaging step, the multivalent binding complex is disrupted and washed away, the correct blocked nucleotide is incorporated into the primer extension strand, and the sequencing cycle is repeated.
[0221] In some embodiments, sequencing is performed using single molecule sequencing by ligation. Such techniques utilize DNA ligase to incorporate oligonucleotides and identify the incorporation of such oligonucleotides. The oligonucleotides typically have different labels that are correlated with the identity of a particular nucleotide in a sequence to which the oligonucleotides hybridize. Aspects and features involved in sequencing by ligation are described, for example, in Shendure et al. Science (2005), 309: 1728-1732, and in US 5,599,675; US 5,750,341; US 6,969,488; US 6,172,218; US and 6,306,597.
[0222] In some embodiments, real-time monitoring of DNA polymerase activity can be used during sequencing. For example, nucleotide incorporations can be detected through fluorescence resonance energy transfer (FRET), as described for example in Levene et al., Science (2003), 299, 682-686, Lundquist et al., Opt. Lett. (2008), 33, 1026-1028, and Korlach et al., Proc. Natl. Acad. Sci. USA (2008), 105, 1176-1181.B. Single Cell Analysis
[0223] In some embodiments, an extended oligonucleotide is processed and used for single cell analysis. In some embodiments, an RNA template as described in Section II is used for single cell analysis. In some embodiments, an assay is performed for single cell whole transcriptome gene expression with multiplexing capabilities to profile hundreds to a million cells. In some instances, the process for generating extended oligonucleotides using RNA templates are used with systems and methods described herein that provide for the compartmentalization, depositing, or partitioning of one or more particles (e.g., biological particles, macromolecular constituents of biological particles, beads, reagents, etc.) into discrete compartments or partitions (referred to interchangeably herein as partitions), where each partition maintains separation of its own contents from the contents of other partitions. In some75MOFO-358213026202412023740 instances, RNA templates or the generated extended oligonucleotides are used to generate a plurality of barcoded nucleic acid molecules (e.g., barcoded extended oligonucleotide or a derivative thereof) comprising a sequence corresponding to the RNA template and a barcode sequence or reverse complement thereof.
[0224] Provided herein are methods for barcoding an extended oligonucleotide. In some instances, the extended oligonucleotide or RNA template used to generate the barcoded nucleic acid molecule is in a cell or nucleus. In some cases, the cell or nucleus is partitioned (e.g., in a droplet or well). In some cases, the cell or nucleus is partitioned with a support (e.g., bead or particle) comprising a nucleic acid barcode molecule. The nucleic acid barcode molecule may comprise identifying information, e.g., a barcode sequence. In some instances, a nucleic acid barcode molecule or a sequence thereof is attached to the extended oligonucleotide or a derivative thereof, thereby generating a barcoded nucleic acid molecule.
[0225] In some instances, a barcoded nucleic acid molecule is generated comprising the barcode sequence and a sequence of the extended oligonucleotide or complement thereof. In some instances, a barcoded nucleic acid molecule is generated comprising the barcode sequence and a sequence of the RNA template as described in Section II. In some instances, a nucleic acid barcode molecule or a portion thereof binds to the extended oligonucleotide. In some instances, the extended oligonucleotide or a derivative thereof is ligated to the nucleic acid barcode molecule. In some cases, one or more nucleic acid reactions may be performed to generate the barcoded nucleic acid molecule. For example, the one or more nucleic acid reactions may comprise a ligation reaction, an extension reaction, an amplification reaction, etc. In some aspects, the nucleic acid barcode molecule may comprise additional functional sequences, e.g., a unique molecular identifier (UMI) sequence, an adaptor sequence, a sequencer specific flow cell attachment sequence (such as an P5, P7, or partial P5 or P7 sequence), a primer or primer binding sequence, a sequencing primer, a partial sequencing primer sequence, or primer binding sequence (such as an Rl, R2, or partial R1 or R2 sequence), etc., or any combinations thereof).
[0226] In some embodiments, an adaptase enzyme is contacted with the generated extended oligonucleotide or a derivative thereof. In some embodiments, the adaptase enzyme is an adapter ligation reagent. In some embodiments, an adaptase enzyme is used to append an adapter to a single-stranded DNA (e.g., the extended oligonucleotide or a derivative thereof). In some embodiments, an adaptase enzyme is used to append a sequencing adapter to the single-76MOFO-358213026202412023740 stranded DNA (e.g., the extended oligonucleotide or a derivative thereof). In some embodiments, an adaptase enzyme is used to append one or more sequences for preparing a sequencing library. In some embodiments, the method comprises using an adaptase enzyme to introduce a 5’ handle and / or a 3’ sequencing handle to the extended oligonucleotide or a derivative thereof.
[0227] In some instances, the barcoded nucleic acid molecule or a derivative thereof (e.g., a oligonucleotide or polynucleotide generated from the barcoded nucleic acid molecule) is subjected to sequencing. In some instances, the barcoded nucleic acid molecule is amplified prior to sequencing, and the amplification product is sequenced. In some instances, the sequencing comprises next-generation sequencing, or any other suitable sequencing method.
[0228] In some aspects, the nucleic acid barcode molecule is coupled to a support. In some aspects, the support is a bead, such as a gel bead or a glass bead. In some aspects, the nucleic acid barcode molecule is coupled to the support via a labile moiety, which may be thermolabile, photocleavable, or enzymatically cleavable. In some embodiments, a single cell barcoding reaction is performed using a plurality of nucleic acid barcode molecules covalently coupled to a particle (e.g., a bead). In some embodiments, a single cell barcoding reaction is performed using barcode carrying beads. In some embodiments, a single cell barcoding reaction is performed by partitioning cells or nuclei of a biological sample.
[0229] In some embodiments, a nucleic acid barcode molecule comprises one or more barcode sequences. In some embodiments, a plurality of nucleic acid barcode molecules are coupled to a particle (e.g., a bead). In some embodiments, the one or more barcode sequences include sequences that are the same for all or a portion of the nucleic acid molecules coupled to a given bead and / or sequences that are different across all (or a portion of the) nucleic acid molecules coupled to the given bead. The nucleic acid molecule may be incorporated into the bead.
[0230] In some aspects, nucleic acid barcode molecules comprise one or more functional sequences for coupling to another nucleic acid molecule (e.g., an extended oligonucleotide). Such functional sequences can include, e.g., a template switch oligonucleotide (TSO) sequence, a primer sequence (e.g., a poly T sequence, or a nucleic acid primer sequence complementary to a target nucleic acid sequence and / or for amplifying a target nucleic acid sequence, a random primer, and a primer sequence for messenger RNA). In some instances, an oligonucleotide described in Section II is a nucleic acid barcode molecule comprising a77MOFO-358213026202412023740 functional sequence for coupling to an RNA template. In some instances, an oligonucleotide comprises a random sequence or sequence complementary to an mRNA.
[0231] In some cases, the nucleic acid barcode molecule further comprises a unique molecular identifier (UMI). In some cases, the barcoded nucleic acid molecule comprises one or more functional sequences, for example, for attachment to a sequencing flow cell, such as, for example, a P5 sequence (or a portion thereof) for Illumina® sequencing. In some cases, the barcoded nucleic acid molecule or derivative thereof comprises another functional sequence, such as, for example, a P7 sequence (or a portion thereof) for attachment to a sequencing flow cell for Illumina sequencing. In some cases, a barcoded nucleic acid molecule can comprise an R1 primer sequence for Illumina sequencing. In some cases, a barcoded nucleic acid molecule comprises an R2 primer sequence for Illumina sequencing. In some cases, a functional sequence comprises a partial sequence, such as a partial barcode sequence, partial anchoring sequence, partial sequencing primer sequence (e.g., partial R1 sequence, partial R2 sequence, etc.), a partial sequence configured to attach to the flow cell of a sequencer (e.g., partial P5 sequence, partial P7 sequence, etc.), or a partial sequence of any other type of sequence described elsewhere herein. A partial sequence may contain a contiguous or continuous portion or segment, but not all, of a full sequence, for example. In some cases, a downstream procedure may extend the partial sequence, or derivative thereof, to achieve a full sequence of the partial sequence, or derivative thereof.
[0232] Examples of such nucleic acid molecules (e.g., oligonucleotides, polynucleotides, etc.) and uses thereof, as may be used with compositions, devices, methods and systems of the present disclosure, are provided in U.S. Patent Pub. Nos. 2014 / 0378345 and 2015 / 0376609, each of which is entirely incorporated herein by reference.
[0233] In some embodiments, a nucleic acid barcode molecule is coupled to a bead by a releasable linkage, such as, for example, a disulfide linker. The same bead may be coupled (e.g., via releasable linkage) to one or more other nucleic acid barcode molecules. In some instances, a nucleic acid barcode molecule comprises a barcode. As noted elsewhere herein, the structure of the barcode may comprise a number of sequence elements. In some embodiments, a nucleic acid barcode molecule comprises a functional sequence that may be used in subsequent processing. For example, the functional sequence comprises one or more of a sequencer specific flow cell attachment sequence (e.g., a P5 sequence for Illumina® sequencing systems) and a78MOFO-358213026202412023740 sequencing primer sequence (e.g., a R1 primer for Illumina® sequencing systems), or partial sequence(s) thereof. In some embodiments, a nucleic acid barcode molecule comprises a barcode sequence for use in barcoding the sample and identifying the cell of origin of the analytes (e.g., DNA, RNA, protein, etc.). In some cases, the barcode sequence is bead-specific such that the barcode sequence is common to all nucleic acid barcode molecules coupled to the same bead. Alternatively, or in addition, the barcode sequence can be partition-specific such that the barcode sequence is common to all nucleic acid barcode molecules coupled to one or more beads that are partitioned into the same partition. In some instances, a nucleic acid barcode molecule comprises a sequence complementary to an extended oligonucleotide generated as described in Section II. In some instances, a nucleic acid barcode molecule comprises a sequence complementary to an RNA template as described in Section II. In some instances, a nucleic acid barcode molecule comprises a sequence complementary to a cleaved RNA template as described in Section II.
[0234] In some instances, a nucleic acid barcode molecule comprises a unique molecular identifying sequence (e.g., unique molecular identifier (UMI)). In some cases, the unique molecular identifying sequence may comprise from about 5 to about 8 nucleotides. In some cases, the unique molecular identifying sequence comprises 5, 6, 7, or 8 nucleotides. Alternatively, the unique molecular identifying sequence may compress less than about 5 or more than about 8 nucleotides. In some instances, the unique molecular identifying sequence comprises 1, 2, 3, or 4 nucleotides. In some instances, the unique molecular identifying sequence comprises 9, 10, 11, 12, 13, 14, 15, or more nucleotides. The unique molecular identifying sequence may be a unique sequence that varies across individual nucleic acid barcode molecules coupled to a single bead. In some cases, the unique molecular identifying sequence may be a random sequence (e.g., such as a random N-mer sequence). For example, the UMI may provide a unique identifier of the starting analyte (e.g., mRNA) molecule that was captured, in order to allow quantitation of the number of original expressed RNA molecules.
[0235] In some embodiments, an RNA template or a generated extended oligonucleotide is co-partitioned along with a barcode bearing bead. In some instances, a nucleic acid barcode molecule is released from the bead in the partition. In some instances, a released nucleic acid barcode molecule hybridizes to a sequence of a generated extended oligonucleotide. In some instances, a barcoded extended oligonucleotide can be amplified, cleaned up and79MOFO-358213026202412023740 sequenced to identify the sequence of associated RNA template, as well as to sequence the barcode segment and the UMI segment. In some instances, further processing may be performed, in the partitions or outside the partitions (e.g., in bulk). For instance, additional adapter sequences may be added to the barcoded nucleic acid molecules, or other nucleic acid reactions (e.g., amplification, nucleic acid extension) may be performed. The beads or products thereof (e.g., barcoded nucleic acid molecules) may be collected from the partitions, and / or pooled together and subsequently subjected to clean up and further characterization (e.g., sequencing).
[0236] The operations described herein may be performed at any useful or convenient step. For instance, the beads comprising nucleic acid barcode molecules may be introduced into a partition (e.g., well or droplet) prior to, during, or following introduction of a sample into the partition. In some instances, extended oligonucleotides or derivatives thereof are subjected to barcoding, which may occur on the bead (in cases where the nucleic acid molecules remain coupled to the bead) or following release of the nucleic acid barcode molecules into the partition. In some instances, beads from various partitions may be collected, pooled, and subjected to further processing (e.g., reverse transcription, adapter attachment, amplification, clean up, sequencing). In other instances, the processing may occur in the partition. For example, conditions sufficient for barcoding, adapter attachment, reverse transcription, or other nucleic acid processing operations may be provided in the partition and performed prior to clean up and sequencing.
[0237] In some instances, a bead may comprise a capture sequence or binding sequence configured to bind to a corresponding capture sequence or binding sequence. In some instances, a bead may comprise a plurality of different capture sequences or binding sequences configured to bind to different respective corresponding capture sequences or binding sequences. For example, a bead may comprise a first subset of one or more capture sequences each configured to bind to a first corresponding capture sequence, a second subset of one or more capture sequences each configured to bind to a second corresponding capture sequence, a third subset of one or more capture sequences each configured to bind to a third corresponding capture sequence, and etc. A bead may comprise any number of different capture sequences. In some instances, a bead may comprise at least 2, 3, 4, 5, 6, 7, 8, 9, 10 or more different capture sequences or binding sequences configured to bind to different respective capture sequences or80MOFO-358213026202412023740 binding sequences, respectively. Alternatively, or in addition, a bead may comprise at most about 10, 9, 8, 7, 6, 5, 4, 3, or 2 different capture sequences or binding sequences configured to bind to different respective capture sequences or binding sequences.
[0238] FIG. 4 illustrates an example of a barcode carrying bead for processing an extended oligonucleotide. In some instances, a nucleic acid barcode molecule is coupled to a bead by a releasable linkage such as, for example, a disulfide linker. In some instances, the nucleic acid barcode molecule comprises a capture sequence. In some instance, the nucleic acid barcode molecule comprises a barcode (e.g., bead-specific sequence common to bead, partitionspecific sequence common to partition, etc.), and a unique molecular identifier (e.g., unique sequence within different molecules attached to the bead), or partial sequences thereof.
[0239] Barcodes can be releasably, cleavably or reversibly attached to the beads such that barcodes can be released or be releasable through cleavage of a linkage between the barcode molecule and the bead, or released through degradation of the underlying bead itself, allowing the barcodes to be accessed or be accessible by other reagents, or both. In non-limiting examples, cleavage may be achieved through reduction of di-sulfide bonds, use of restriction enzymes, photo-activated cleavage, or cleavage via other types of stimuli (e.g., chemical, thermal, pH, enzymatic, etc.) and / or reactions, such as described elsewhere herein. Releasable barcodes may sometimes be referred to as being activatable, in that they are available for reaction once released. Thus, for example, an activatable barcode may be activated by releasing the barcode from a bead (or other suitable type of partition described herein). Other activatable configurations are also envisioned in the context of the described methods and systems.
[0240] As will be appreciated from the above disclosure, the degradation of a bead may refer to the disassociation of a bound or entrained species from a bead, both with and without structurally degrading the physical bead itself. For example, the degradation of the bead may involve cleavage of a cleavable linkage via one or more species and / or methods described elsewhere herein. In another example, entrained species may be released from beads through osmotic pressure differences due to, for example, changing chemical environments. See, e.g., WO20 14210353, which is hereby incorporated by reference in its entirety.
[0241] A degradable bead may be introduced into a partition, such as a droplet of an emulsion or a well, such that the bead degrades within the partition and any associated species (e.g., oligonucleotides) are released within the droplet when the appropriate stimulus is applied.81MOFO-358213026202412023740The free species (e.g., oligonucleotides, nucleic acid molecules) may interact with other reagents contained in the partition. See, e.g., WO2014210353, which is hereby incorporated by reference in its entirety.
[0242] In some cases, a species (e.g., oligonucleotide molecules comprising barcodes) that are attached to a solid support (e.g., a bead) may comprise a U-excising element that allows the species to release from the bead. In some cases, the U-excising element may comprise a single- stranded DNA (ssDNA) sequence that contains at least one uracil. The species may be attached to a solid support via the ssDNA sequence containing the at least one uracil. The species may be released by a combination of uracil-DNA glycosylase (e.g., to remove the uracil) and an endonuclease (e.g., to induce an ssDNA break). If the endonuclease generates a 5’ phosphate group from the cleavage, then additional enzyme treatment may be included in downstream processing to eliminate the phosphate group, e.g., prior to ligation of additional sequencing handle elements, e.g., Illumina full P5 sequence, partial P5 sequence, full R1 sequence, and / or partial R1 sequence.
[0243] The barcodes that are releasable as described herein may sometimes be referred to as being activatable, in that they are available for reaction once released. Thus, for example, an activatable barcode may be activated by releasing the barcode from a bead (or other suitable type of partition described herein). Other activatable configurations are also envisioned in the context of the described methods and systems.
[0244] The nucleic acid barcode sequences can include from about 6 to about 20 or more nucleotides within the sequence of the nucleic acid molecules (e.g., oligonucleotides). The nucleic acid barcode sequences can include from about 6 to about 20, 30, 40, 50, 60, 70, 80, 90, 100 or more nucleotides. In some cases, the length of a barcode sequence may be about 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 nucleotides or longer. In some cases, the length of a barcode sequence may be at least about 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 nucleotides or longer. In some cases, the length of a barcode sequence may be at most about 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 nucleotides or shorter. These nucleotides may be completely contiguous, e.g., in a single stretch of adjacent nucleotides, or they may be separated into two or more separate subsequences that are separated by 1 or more nucleotides. In some cases, separated barcode subsequences can be from about 4 to about 16 nucleotides in length. In some cases, the barcode subsequence may be about 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 1682MOFO-358213026202412023740 nucleotides or longer. In some cases, the barcode subsequence may be at least about 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16 nucleotides or longer. In some cases, the barcode subsequence may be at most about 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16 nucleotides or shorter.
[0245] The co-partitioned nucleic acid molecules can also comprise other functional sequences useful in the processing of the nucleic acids from the co-partitioned biological particles. These sequences include, e.g., targeted or random / universal amplification primer sequences for amplifying nucleic acids (e.g., mRNA, the genomic DNA) from the individual biological particles within the partitions while attaching the associated barcode sequences, sequencing primers or primer recognition sites, hybridization or probing sequences, e.g., for identification of presence of the sequences or for pulling down barcoded nucleic acids, or any of a number of other potential functional sequences (e.g., restriction sites, transposition sites). Other mechanisms of co-partitioning oligonucleotides may also be employed, including, e.g., coalescence of two or more droplets, where one droplet contains oligonucleotides, or microdispensing of oligonucleotides (e.g., attached to a bead) into partitions, e.g., droplets within microfluidic systems.
[0246] In an example, beads are provided that include large numbers of the above described nucleic acid barcode molecules releasably attached to the beads, where all or at least a subset of the nucleic acid barcode molecules attached to a particular bead will include a common nucleic acid barcode sequence, but where a large number of diverse barcode sequences are represented across the population of beads used. In some embodiments, hydrogel beads, e.g., comprising polyacrylamide polymer matrices, are used as a solid support and delivery vehicle for the nucleic acid barcode molecules into the partitions, as they are capable of carrying large numbers of nucleic acid barcode molecules, and may be configured to release those nucleic acid molecules upon exposure to a particular stimulus, as described elsewhere herein. In some cases, the population of beads provides a diverse barcode sequence library that includes at least about 1,000 different barcode sequences, at least about 5,000 different barcode sequences, at least about 10,000 different barcode sequences, at least about 50,000 different barcode sequences, at least about 100,000 different barcode sequences, at least about 1,000,000 different barcode sequences, at least about 5,000,000 different barcode sequences, or at least about 10,000,000 different barcode sequences, or more. In some cases, the population of beads provides a diverse barcode sequence library that includes about 1,000 to about 10,000 different barcode sequences,83MOFO-358213026202412023740 about 5,000 to about 50,000 different barcode sequences, about 10,000 to about 100,000 different barcode sequences, about 50,000 to about 1,000,000 different barcode sequences, or about 100,000 to about 10,000,000 different barcode sequences.
[0247] Additionally, beads can be provided with large numbers of nucleic acid (e.g., oligonucleotide) molecules attached. In particular, the number of molecules of nucleic acid molecules including the barcode sequence on an individual bead can be at least about 1,000 nucleic acid molecules, at least about 5,000 nucleic acid molecules, at least about 10,000 nucleic acid molecules, at least about 50,000 nucleic acid molecules, at least about 100,000 nucleic acid molecules, at least about 500,000 nucleic acids, at least about 1,000,000 nucleic acid molecules, at least about 5,000,000 nucleic acid molecules, at least about 10,000,000 nucleic acid molecules, at least about 50,000,000 nucleic acid molecules, at least about 100,000,000 nucleic acid molecules, at least about 250,000,000 nucleic acid molecules and in some cases at least about 1 billion nucleic acid molecules, or more. In some embodiments, the number of nucleic acid molecules including the barcode sequence on an individual bead is between about 1,000 to about 10,000 nucleic acid molecules, about 5,000 to about 50,000 nucleic acid molecules, about 10,000 to about 100,000 nucleic acid molecules, about 50,000 to about 1,000,000 nucleic acid molecules, about 100,000 to about 10,000,000 nucleic acid molecules, about 1,000,000 to about 1 billion nucleic acid molecules.
[0248] Nucleic acid molecules of a given bead can include identical (or common) barcode sequences, different barcode sequences, or a combination of both. Nucleic acid molecules of a given bead can include multiple sets of nucleic acid molecules. Nucleic acid molecules of a given set can include identical barcode sequences. The identical barcode sequences can be different from barcode sequences of nucleic acid molecules of another set.
[0249] Moreover, when the population of beads is partitioned, the resulting population of partitions can also include a diverse barcode library that includes at least about 1,000 different barcode sequences, at least about 5,000 different barcode sequences, at least about 10,000 different barcode sequences, at least at least about 50,000 different barcode sequences, at least about 100,000 different barcode sequences, at least about 1,000,000 different barcode sequences, at least about 5,000,000 different barcode sequences, or at least about 10,000,000 different barcode sequences. Additionally, each partition of the population can include at least about 1,000 nucleic acid barcode molecules, at least about 5,000 nucleic acid84MOFO-358213026202412023740 barcode molecules, at least about 10,000 nucleic acid barcode molecules, at least about 50,000 nucleic acid barcode molecules, at least about 100,000 nucleic acid barcode molecules, at least about 500,000 nucleic acids, at least about 1,000,000 nucleic acid barcode molecules, at least about 5,000,000 nucleic acid barcode molecules, at least about 10,000,000 nucleic acid barcode molecules, at least about 50,000,000 nucleic acid barcode molecules, at least about 100,000,000 nucleic acid barcode molecules, at least about 250,000,000 nucleic acid barcode molecules and in some cases at least about 1 billion nucleic acid barcode molecules.
[0250] In some cases, the resulting population of partitions provides a diverse barcode sequence library that includes about 1,000 to about 10,000 different barcode sequences, about 5,000 to about 50,000 different barcode sequences, about 10,000 to about 100,000 different barcode sequences, about 50,000 to about 1,000,000 different barcode sequences, or about 100,000 to about 10,000,000 different barcode sequences. Additionally, each partition of the population can include between about 1,000 to about 10,000 nucleic acid barcode molecules, about 5,000 to about 50,000 nucleic acid barcode molecules, about 10,000 to about 100,000 nucleic acid barcode molecules, about 50,000 to about 1,000,000 nucleic acid barcode molecules, about 100,000 to about 10,000,000 nucleic acid barcode molecules, about 1,000,000 to about 1 billion nucleic acid barcode molecules.
[0251] In some cases, it may be desirable to incorporate multiple different barcodes within a given partition, either attached to a single or multiple beads within the partition. For example, in some cases, a mixed, but known set of barcode sequences may provide greater assurance of identification in the subsequent processing, e.g., by providing a stronger address or attribution of the barcodes to a given partition, as a duplicate or independent confirmation of the output from a given partition.
[0252] In some instances, the nucleic acid molecules (e.g., oligonucleotides) are releasable from the beads upon the application of a particular stimulus to the beads. In some cases, the stimulus may be a photo- stimulus, e.g., through cleavage of a photo-labile linkage that releases the nucleic acid molecules. In other cases, a thermal stimulus may be used, where elevation of the temperature of the beads environment will result in cleavage of a linkage or other release of the nucleic acid molecules from the beads. In still other cases, a chemical stimulus can be used that cleaves a linkage of the nucleic acid molecules to the beads, or otherwise results in release of the nucleic acid molecules from the beads. In one case, such85MOFO-358213026202412023740 compositions include the polyacrylamide matrices described above for encapsulation of biological particles and may be degraded for release of the attached nucleic acid molecules through exposure to a reducing agent, such as DTT. In some embodiments, extension of the oligonucleotides occurs prior to its release from the bead. In some embodiments, extension of the oligonucleotides occurs after its release from the bead.
[0253] In some embodiments, a barcoding reaction is performed for appending a barcode or a complement thereof to an extended oligonucleotide or the RNA template as described in Section II. In some embodiments, a nucleic acid barcode molecule is co-partitioned with an extended oligonucleotide or the RNA template. In some instances, nucleic acid barcode molecule is attached to a support (e.g., a bead, such as a gel bead), such as those described elsewhere herein. For example, nucleic acid barcode molecule is attached to support via a releasable linkage (e.g., comprising a labile bond), such as those described elsewhere herein. In some aspects, the nucleic acid barcode molecule comprises a functional sequence and optionally comprise other additional sequences, for example, a barcode sequence (e.g., common cell barcode, partition- specific barcode, or other functional sequences described elsewhere herein), and / or a UMI sequence. In some aspects, the nucleic acid barcode molecule comprises a capture sequence that is complementary to another nucleic acid sequence (e.g., an extended oligonucleotide, RNA template or derivative thereof), such that it hybridizes to a particular sequence, e.g., capture handle sequence.
[0254] In some embodiments, the capture sequence is configured to hybridize to an RNA template. In some embodiments, the capture sequence comprises a poly-T sequence and is used to hybridize to an RNA template. In some embodiments, the capture sequence comprises a randomer and is used to hybridize to an RNA template. In some embodiments, a nucleic acid barcode molecule comprises a capture sequence complementary to a sequence of a cleaved RNA template. In some instances, a nucleic acid barcode molecule comprises a capture sequence complementary to a sequence of a generated extended oligonucleotide or derivative thereof (as shown in FIG. 4). In some instances, the capture sequence comprises a sequence corresponding to a specific RNA template. In some instances, a capture sequence comprises a common capture sequence for hybridizing to different RNA templates or extended oligonucleotides or derivatives thereof. In some instances, a nucleic acid extension reaction is performed, thereby generating a86MOFO-358213026202412023740 barcoded nucleic acid molecule comprising a capture sequence, functional sequence, barcode sequence, any other functional sequence, and a sequence corresponding to the RNA template.
[0255] In another example, a capture sequence is complementary to an adapter sequence that has been appended to an extended oligonucleotide.C. Spatial Array
[0256] In some embodiments, an extended oligonucleotide is processed using a spatial array. In some embodiments, an RNA template as described in Section II is captured on a spatial array. In some embodiments, an assay is performed for spatial analysis using a spatial array. In some instances, the process for generating extended oligonucleotides using RNA templates are used with a spatial array to generate barcoded nucleic acid molecules. In some instances, RNA templates or the generated extended oligonucleotides are used to generate a plurality of barcoded nucleic acid molecules (e.g., barcoded extended oligonucleotide or a derivative thereof) comprising a sequence corresponding to the RNA template and a spatial barcode or a complement thereof.
[0257] In some embodiments, the detecting of an extended oligonucleotide or RNA template used to generate the extended oligonucleotide comprises detecting a location of a spatial array of oligonucleotides (e.g., capture probes), in which spatial information from a biological sample is preserved. In some embodiments, an extended oligonucleotide or RNA template is transferred to an array for spatial assay. In some instances, a generated barcoded nucleic acid molecule is analyzed by performing NGS sequencing to determine one or more sequences of the oligonucleotides captured on the array.
[0258] In one aspect, provided herein are methods, compositions, apparatus, and systems for spatial analysis of a biological sample, for example, a spatial array-based analysis. Non-limiting aspects of spatial analysis methodologies are described in U.S. Pat. Pub. No. 10,308,982; U.S. Pat. Pub. No. 9,879,313; U.S. Pat. Pub. No. 9,868,979; Liu et al., bioRxiv 788992, 2020; U.S. Pat. Pub. No. 10,774,372; U.S. Pat. Pub. No. 10,774,374; WO 2018 / 091676; U.S. Pat. Pub. No. 10,030,261; U.S. Pat. Pub. No. 9,593,365; U.S. Patent No. 10,002,316; U.S. Patent No. 9,727,810; U.S. Pat. Pub. No. 10,640,816; Rodriques et al., Science 363(6434): 1463-1467, 2019; US 11,447,807; Lee et al., Nat. Protoc. 10(3):442-458, 2015; U.S. Pat. Pub. No. 10,179,932; US 11,085,072; U.S. Pat. Pub. No. 10,138,509; Trejo et al., PLoS ONE 14(2) :e0212031, 2019; U.S. Patent Application Publication No. 2018 / 0245142; Chen et al.,87MOFO-358213026202412023740Science 348(6233):aaa6090, 2015; Gao et al., BMC Biol. 15:50, 2017; WO 2017 / 144338; US 2018 / 0372736; US 2022 / 0290228; US 11,597,965; WO 2011 / 094669; U.S. Patent No. 7,709,198; U.S. Patent No. 8,604,182; U.S. Patent No. 8,951,726; U.S. Patent No. 9,783,841; U.S. Patent No. 10,041,949; WO 2016 / 057552; US 2021 / 0238665; U.S. Pat. Pub. No. 10,370,698; US 10,724,078; U.S. Pat. Pub. No. 10,364,457; U.S. Pat. Pub. No. 10,317,321; US 2021 / 0395796; US 2020 / 0239946; U.S. Patent No. 10,059,990; US 11,505,819; US 11,104,936; and Gupta et al., Nature Biotechnol. 36:1197-1202, 2018, all of which are herein incorporated by reference in their entireties, and can be used herein in any combination. Further non-limiting aspects of spatial analysis methodologies are described herein.
[0259] In some embodiments, the spatial assays disclosed herein comprise capturing an extended oligonucleotide or a derivative thereof. In some embodiments, the spatial assays disclosed herein comprise capturing the RNA template used to generate the extended oligonucleotide. In some aspects, the biological sample is on the substrate (e.g., a cover slip with sufficient strength comprising the capture array). In some embodiments, the biological sample is on a second substrate. The biological sample may be positioned between the first substrate (e.g., substrate comprising the capture array) and the second substrate (e.g., slide comprising the biological sample) such that the capture agents are allowed to capture the RNA template, the extended oligonucleotide or derivatives thereof. In some embodiments, the biological sample is processed to release the RNA template, the extended oligonucleotide or derivatives thereof.
[0260] In some embodiments, a method disclosed herein comprises transferring the RNA template, the extended oligonucleotide or a derivative thereof from a biological sample to an array of features on a substrate, each of which is associated with a unique spatial location on the array. Each feature may comprise a plurality of capture agents (e.g., capture probes) capable of capturing one or more nucleic acid molecules, and each of the capture agents of the same feature may comprise a spatial barcode corresponding to a unique spatial location of the feature on the array. In some embodiments, the method comprises capturing the RNA template, the extended oligonucleotide or a derivative thereof by a capture agent (e.g., capture probe). The capture probe comprises a capture domain that binds to a sequence in the RNA template, the extended oligonucleotide or a derivative thereof. One or more reactions (e.g., extension, and / or ligation) are performed to generate a spatially labeled polynucleotide sequence comprising (i) a sequence of the captured RNA template, extended oligonucleotide or a complement thereof, and88MOFO-358213026202412023740(ii) a sequence of the spatial barcode (e.g., spatial barcode of the capture probe) or complement thereof. Subsequent analysis of the transferred analytes includes determining the identity of the analytes and the spatial location of each analyte within the sample. The spatially labeled polynucleotide or a portion thereof may be removed from the substrate (e.g., capture array) for sequencing using any suitable nucleic acid sequencing techniques, including next-generation sequencing (NGS). In some embodiments, the sequence of the spatially labeled polynucleotide is determined to detect the spatial barcode and the RNA template. All or part of the sequence of the generated spatially labeled polynucleotide may be determined. The spatial location of each analyte (e.g., RNA template) within the sample is determined based on the feature to which each analyte is bound in the array, and the feature’s relative spatial location within the array.
[0261] In some embodiments, a method disclosed herein comprises associating a spatial barcode with one or more analytes (e.g., RNA template), such that the spatial barcode identifies the one or more analytes, and / or contents of the one or more cells, as associated with a particular spatial location.
[0262] In some embodiments, a method disclosed herein comprises contacting RNA templates, extended oligonucleotides or derivatives thereof with a spatially-barcoded array populated with capture probes. In some instances, the RNA template, extended oligonucleotide or a derivative thereof comprises a sequence that is configured to interact with a capture probe on the spatially-barcoded array. Once the RNA template, extended oligonucleotide or a derivative thereof is captured (e.g., via hybridization or splint ligation to the capture probe), the sample is optionally removed from the array and the captured molecules are processed and analyzed in order to obtain spatially-resolved analyte information.
[0263] In some embodiments, a method disclosed herein comprises delivering or driving spatially-barcoded nucleic acid molecules (e.g., capture probes) towards and / or into or onto a sample. In some embodiments, a method disclosed herein comprises cleaving spatially- barcoded nucleic acid molecules (e.g., capture probes) from an array and driving the cleaved nucleic acid molecules towards and / or into or onto a sample. Alternatively, the sample may be permeabilized and fixed / crosslinked to restrict mobility of one or more target analytes, while allowing spatially-barcoded capture probes to migrate towards and / or into or onto the sample. Once the spatially-barcoded capture probe is associated with a particular analyte (e.g., RNA template, extended oligonucleotide or a derivative thereof), the sample can be optionally89MOFO-358213026202412023740 removed for analysis. The sample can be optionally dissociated before analysis. Once the tagged analyte or cell is associated with the spatially-barcoded capture probe, the capture probes can be analyzed to obtain spatially-resolved information about the tagged analyte or cell.
[0264] Examples of workflows for sample preparation, permeabilization, DNA generation (e.g., first strand cDNA generation and second strand generation), DNA amplification (e.g., cDNA amplification) and quality control, and spatial gene expression library construction are disclosed for example in US 2021 / 0317524, US 2021 / 0332424, US 2021 / 0317524, US 2021 / 0324457, US 2021 / 0332425, all of which are incorporated herein by reference in their entireties.
[0265] A capture probe herein can comprise any molecule capable of capturing (directly or indirectly) an analyte of interest in a biological sample (e.g., RNA template, extended oligonucleotide or a derivative thereof). In some embodiments, the capture probe is a nucleic acid. In some embodiments, the capture probe includes a barcode (e.g., a spatial barcode and / or a unique molecular identifier (UMI)) and a capture domain. In some embodiments, the capture probe is the oligonucleotide as described in Section II.
[0266] In some embodiments, analytes in a biological sample are pre-processed prior to interaction with a capture probe. For example, prior to interaction with capture probes, polymerization reactions catalyzed by a polymerase (e.g., DNA polymerase or reverse transcriptase) are performed in the biological sample. In some embodiments, a primer for the polymerization reaction includes a functional group that enhances hybridization with the capture probe. The capture probes can include appropriate capture domains to capture biological analytes of interest (e.g., RNA template, extended oligonucleotide or a derivative thereof as described in Section II).
[0267] In some embodiments, biological analytes are pre-processed for library generation via next generation sequencing. For example, analytes can be pre-processed by addition of a modification (e.g., ligation of sequences that allow interaction with capture probes). In some embodiments, analytes are fragmented using fragmentation techniques. Fragmentation can be followed by a modification of the analyte. For example, a modification can be the addition through ligation of an adapter sequence that allows hybridization with the capture probe. In some embodiments, where the analyte of interest is RNA (e.g., mRNA), poly(A) tailing is performed. Addition of a poly(A) tail to RNA that does not contain a poly(A) tail can facilitate90MOFO-358213026202412023740 hybridization with a capture probe that includes a capture domain with a functional amount of poly(dT) sequence.
[0268] In some embodiments, prior to interaction with capture probes, ligation reactions catalyzed by a ligase are performed in the biological sample. In some embodiments, ligation is performed by chemical ligation.
[0269] In some embodiments, prior to interaction with capture probes, target- specific reactions are performed in the biological sample. Examples of target specific reactions include, but are not limited to, ligation of target specific adaptors, probes and / or other oligonucleotides, target specific amplification using primers specific to one or more analytes (e.g., RNA template, extended oligonucleotide or a derivative thereof). In some embodiments, a capture probe includes capture domains targeted to target- specific products (e.g., amplification or ligation).
[0270] In some embodiments, the capture probes may comprise one or more cleavable capture probes, wherein the cleaved capture probe can enter into a non-permeabilized cell and bind to target analytes within the sample. The capture probe may contain a cleavage domain, a cell penetrating peptide, a reporter molecule, and a disulfide bond (-S-S-). In some cases, the capture probe may also include a spatial barcode and a capture domain. i. Capture Domain
[0271] In some embodiments, each capture agent (e.g., a capture probe) comprises at least one capture domain, which may comprise an oligonucleotide that binds specifically to a desired analyte. In some embodiments, a capture domain can be used to capture or detect a desired analyte (e.g., RNA template, extended oligonucleotide or a derivative thereof).
[0272] In some embodiments, the capture domain comprises a functional nucleic acid sequence configured to interact with one or more analytes (e.g.,, RNA templates, extended oligonucleotides or derivatives thereof). In some embodiments, the functional nucleic acid sequence can include an N-mer sequence (e.g., a randomer sequence), which N-mer sequences are configured to interact with a plurality of nucleic acid molecules. In some embodiments, the functional sequence comprises a poly(T) sequence, which poly(T) sequences are configured to interact with messenger RNA (mRNA) molecules via the poly (A) tail of an mRNA transcript.
[0273] In some embodiments, capture probes include ribonucleotides and / or deoxyribonucleotides as well as synthetic nucleotide residues that are capable of participating in Watson-Crick type or analogous base pair interactions. In some embodiments, the capture91MOFO-358213026202412023740 domain is capable of priming a reverse transcription reaction to generate cDNA that is complementary to the captured RNA template. In some embodiments, the capture domain can template a ligation reaction between the captured nucleic acid (e.g., RNA template, extended oligonucleotide or a derivative thereof) and a surface probe that is directly or indirectly immobilized on the substrate. In some embodiments, the capture domain can be ligated to one strand of the captured nucleic acid. For example, SplintR® ligase along with RNA or DNA sequences (e.g., degenerate RNA) can be used to ligate a single-stranded DNA or RNA to the capture domain. In some embodiments, ligases with RNA-templated ligase activity, e.g., SplintR® ligase, T4 RNA ligase 2 or KOD ligase, can be used to ligate a single- stranded DNA or RNA to the capture domain. In some embodiments, a capture domain includes a splint oligonucleotide. In some embodiments, a capture domain captures a splint oligonucleotide.
[0274] In some embodiments, the capture domain is located at the 3’ end of the capture probe and includes a free 3’ end that can be extended, e.g. by template dependent polymerization, to form an extended capture probe as described herein. In some embodiments, the capture domain includes a nucleotide sequence that is capable of hybridizing to nucleic acid, e.g. RNA template, extended oligonucleotide or a derivative thereof, present in the cells of the tissue sample contacted with the array. In some embodiments, the capture domain can be selected or designed to bind selectively or specifically to a target nucleic acid (e.g., RNA template, extended oligonucleotide or a derivative thereof). For example, the capture domain can be selected or designed to capture mRNA by way of hybridization to the mRNA poly(A) tail. Thus, in some embodiments, the capture domain includes a poly(T) DNA oligonucleotide, e.g., a series of consecutive deoxythymidine residues linked by phosphodiester bonds, which is capable of hybridizing to the poly(A) tail of mRNA. In some embodiments, the capture domain can include nucleotides that are functionally or structurally analogous to a poly(T) tail. For example, a poly(U) oligonucleotide or an oligonucleotide included of deoxy thymidine analogues. In some embodiments, the capture domain includes at least 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides. In some embodiments, the capture domain includes at least 25, 30, or 35 nucleotides.
[0275] In some embodiments, random sequences, e.g., random hexamers or similar sequences, can be used to form all or a part of the capture domain. For example, random sequences can be used in conjunction with poly(T) (or poly(T) analogue) sequences. Thus, where92MOFO-358213026202412023740 a capture domain includes a poly(T) (or a “poly(T)-like”) oligonucleotide, it can also include a random oligonucleotide sequence (e.g., “poly(T)-random sequence” probe). This can, for example, be located 5’ or 3’ of the poly(T) sequence, e.g. at the 3’ end of the capture domain. The poly(T)-random sequence probe can facilitate the capture of the mRNA poly(A) tail. In some embodiments, the capture domain can be an entirely random sequence. In some embodiments, degenerate capture domains can be used.
[0276] In some embodiments, a pool of two or more capture probes form a mixture, where the capture domain of one or more capture probes includes a poly(T) sequence and the capture domain of one or more capture probes includes random sequences. In some embodiments, a pool of two or more capture probes form a mixture where the capture domain of one or more capture probes includes poly(T)-like sequence and the capture domain of one or more capture probes includes random sequences. In some embodiments, a pool of two or more capture probes form a mixture where the capture domain of one or more capture probes includes a poly(T)-random sequences and the capture domain of one or more capture probes includes random sequences. In some embodiments, probes with degenerate capture domains can be added to any of the preceding combinations listed herein. In some embodiments, probes with degenerate capture domains can be substituted for one of the probes in each of the pairs described herein.
[0277] In some instances, a capture domain is based on a particular gene sequence or particular motif sequence or common / conserved sequence that it is designed to capture (e.g., a sequence-specific capture domain). Thus, in some embodiments, the capture domain is capable of binding selectively to a desired sub-type or subset of nucleic acid. In some embodiments, the capture domain is capable of binding selectively to one or more RNA template, extended oligonucleotide or a derivative thereof disclosed herein. In some instances, the capture domain is configured to bind to a sequence 3’ to a region of interest as described in Section II.
[0278] In some embodiments, a capture domain includes an “anchor” or “anchoring sequence”, which is a sequence of nucleotides that is designed to ensure that the capture domain hybridizes to the intended biological analyte. In some embodiments, an anchor sequence includes a sequence of nucleotides, including a 1-mer, 2-mer, 3-mer or longer sequence. In some embodiments, the short sequence is random. For example, a capture domain including a poly(T) sequence can be designed to capture an mRNA. In such embodiments, an anchoring sequence93MOFO-358213026202412023740 can include a random 3-mer (e.g., GGG) that helps ensure that the poly(T) capture domain hybridizes to an mRNA. In some embodiments, an anchoring sequence can be VN, N, or NN. Alternatively, the sequence can be designed using a specific sequence of nucleotides. In some embodiments, the anchor sequence is at the 3’ end of the capture domain. In some embodiments, the anchor sequence is at the 5’ end of the capture domain. ii. Cleavage Domain
[0279] Each capture probe can optionally include at least one cleavage domain. The cleavage domain represents the portion of the probe that is used to reversibly attach the probe to an array feature, as will be described further below. Further, one or more segments or regions of the capture probe can optionally be released from the array feature by cleavage of the cleavage domain. As an example spatial barcodes and / or universal molecular identifiers (UMIs) can be released by cleavage of the cleavage domain.
[0280] In some embodiments, the cleavage domain linking the capture probe to a feature is a disulfide bond. A reducing agent can be added to break the disulfide bonds, resulting in release of the capture probe from the feature. As another example, heating can also result in degradation of the cleavage domain and release of the attached capture probe from the array feature. In some embodiments, laser radiation is used to heat and degrade cleavage domains of capture probes at specific locations. In some embodiments, the cleavage domain is a photosensitive chemical bond (e.g., a chemical bond that dissociates when exposed to light such as ultraviolet light).
[0281] Other examples of cleavage domains include labile chemical bonds such as, but not limited to, ester linkages (e.g., cleavable with an acid, a base, or hydroxylamine), a vicinal diol linkage (e.g., cleavable via sodium periodate), a Diels- Alder linkage (e.g., cleavable via heat), a sulfone linkage (e.g., cleavable via a base), a silyl ether linkage (e.g., cleavable via an acid), a glycosidic linkage (e.g., cleavable via an amylase), a peptide linkage (e.g., cleavable via a protease), or a phosphodiester linkage (e.g., cleavable via a nuclease (e.g., DNAase)).
[0282] In some embodiments, the cleavage domain includes a sequence that is recognized by one or more enzymes capable of cleaving a nucleic acid molecule, e.g., capable of breaking the phosphodiester linkage between two or more nucleotides. A bond can be cleavable via other nucleic acid molecule targeting enzymes, such as restriction enzymes (e.g., restriction endonucleases). For example, the cleavage domain can include a restriction endonuclease94MOFO-358213026202412023740(restriction enzyme) recognition sequence. Restriction enzymes cut double- stranded or single stranded DNA at specific recognition nucleotide sequences known as restriction sites. In some embodiments, a rare-cutting restriction enzyme, e.g., enzymes with a long recognition site (at least 8 base pairs in length), is used to reduce the possibility of cleaving elsewhere in the capture probe.
[0283] In some embodiments, the cleavage domain includes a poly(U) sequence which can be cleaved by a mixture of Uracil DNA glycosylase (UDG) and the DNA glycosylase- lyase Endonuclease VIII, commercially known as the USER™ enzyme. Releasable capture probes can be available for reaction once released. Thus, for example, an activatable capture probe can be activated by releasing the capture probes from a feature.
[0284] In some embodiments, where the capture probe is attached indirectly to a substrate, e.g., via a surface probe, the cleavage domain includes one or more mismatch nucleotides, so that the complementary parts of the surface probe and the capture probe are not 100% complementary (for example, the number of mismatched base pairs can one, two, or three base pairs). Such a mismatch is recognized, e.g., by the MutY and T7 endonuclease I enzymes, which results in cleavage of the nucleic acid molecule at the position of the mismatch.
[0285] In some embodiments, where the capture probe is attached to a feature indirectly, e.g., via a surface probe, the cleavage domain includes a nickase recognition site or sequence. Nickases are endonucleases which cleave only a single strand of a DNA duplex. Thus, the cleavage domain can include a nickase recognition site close to the 5’ end of the surface probe (and / or the 5’ end of the capture probe) such that cleavage of the surface probe or capture probe destabilizes the duplex between the surface probe and capture probe thereby releasing the capture probe) from the feature.
[0286] Nickase enzymes can also be used in some embodiments where the capture probe is attached to the feature directly. For example, the substrate can be contacted with a nucleic acid molecule that hybridizes to the cleavage domain of the capture probe to provide or reconstitute a nickase recognition site, e.g., a cleavage helper probe. Thus, contact with a nickase enzyme will result in cleavage of the cleavage domain thereby releasing the capture probe from the feature. Such cleavage helper probes can also be used to provide or reconstitute cleavage recognition sites for other cleavage enzymes, e.g., restriction enzymes.95MOFO-358213026202412023740
[0287] Some nickases introduce single-stranded nicks only at particular sites on a DNA molecule, by binding to and recognizing a particular nucleotide recognition sequence. A number of naturally-occurring nickases have been discovered, of which at present the sequence recognition properties have been determined for at least four. Nickases are described in U.S. Patent No. 6,867,028, which is incorporated herein by reference in its entirety. In general, any suitable nickase can be used to bind to a complementary nickase recognition site of a cleavage domain. Following use, the nickase enzyme can be removed from the assay or inactivated following release of the capture probes to prevent unwanted cleavage of the capture probes.
[0288] In some embodiments, a cleavage domain is absent from the capture probe. Examples of substrates with attached capture probes lacking a cleavage domain are described for example in Macosko et al., (2015) Cell 161, 1202-1214, the entire contents of which are incorporated herein by reference.
[0289] In some embodiments, the region of the capture probe corresponding to the cleavage domain can be used for some other function. For example, an additional region for nucleic acid extension or amplification can be included where the cleavage domain would normally be positioned. In such embodiments, the region can supplement the functional domain or even exist as an additional functional domain. In some embodiments, the cleavage domain is present but its use is optional. iii. Functional Domain
[0290] Each capture probe can optionally include at least one functional domain. Each functional domain typically includes a functional nucleotide sequence for a downstream analytical step in the overall analysis procedure. In some embodiments, a functional domain is appended to an extended oligonucleotide or a derivative thereof, e.g., by performing one or more reactions (e.g., extension, and / or ligation).
[0291] In some cases, an extended oligonucleotide or derivative thereof captured by the capture probe can comprise one or more functional sequences. For example, a functional sequence can comprise a sequence for attachment to a sequencing flow cell, such as, for example, a P5 sequence for Illumina® sequencing. In some cases, an extended oligonucleotide or derivative thereof captured by the capture probe can comprise another functional sequence, such as, for example, a P7 sequence for attachment to a sequencing flow cell for Illumina sequencing. In some cases, the functional sequence can comprise a barcode sequence or multiple96MOFO-358213026202412023740 barcode sequences. In some cases, the functional sequence can comprise a unique molecular identifier (UMI). In some cases, the functional sequence can comprise a primer sequence (e.g., an R1 primer sequence for Illumina sequencing, an R2 primer sequence for Illumina sequencing, etc.). In some cases, a functional sequence can comprise a partial sequence, such as a partial barcode sequence, partial anchoring sequence, partial sequencing primer sequence (e.g., partial R1 sequence, partial R2 sequence, etc.), a partial sequence configured to attach to the flow cell of a sequencer (e.g., partial P5 sequence, partial P7 sequence, etc.), or a partial sequence of any other type of sequence described elsewhere herein. A partial sequence may contain a contiguous or continuous portion or segment, but not all, of a full sequence, for example. In some cases, a downstream procedure may extend the partial sequence, or derivative thereof, to achieve a full sequence of the partial sequence, or derivative thereof. Examples of such capture probes and uses thereof are described in U.S. Patent Publication Nos. 2014 / 0378345 and 2015 / 0376609, the entire contents of each of which are incorporated herein by reference. The functional domains can be selected for compatibility with a variety of different sequencing systems, e.g., 454 Sequencing, Ion Torrent Proton or PGM, Illumina X10, etc., or other platforms from Illumina, BGI, Qiagen, Thermo-Fisher, PacBio, and Roche, and the requirements thereof. iv. Spatial Barcode
[0292] As discussed above, the capture probe can include one or more spatial barcodes (e.g., two or more, three or more, four or more, five or more) spatial barcodes. A “spatial barcode” is a contiguous nucleic acid segment or two or more non-contiguous nucleic acid segments that function as a label or identifier that conveys or is capable of conveying spatial information. In some embodiments, a capture probe includes a spatial barcode that possesses a spatial aspect, where the barcode is associated with a particular location within an array or a particular location on a substrate. Exemplary spatial barcodes are described in US Patent No. 10,030,261, which is incorporated herein by reference. In some embodiments, the oligonucleotide (e.g., as described in Section II) comprises a spatial barcode. In some embodiments, the extended oligonucleotide (e.g., as described in Section II) or a derivative thereof comprises a spatial barcode.
[0293] A spatial barcode can be part of a capture probe. A spatial barcode can be unique. In some embodiments where the spatial barcode is unique, the spatial barcode functions97MOFO-358213026202412023740 both as a spatial barcode and as a unique molecular identifier (UMI), associated with one particular capture probe.
[0294] Spatial barcodes can have a variety of different formats. For example, spatial barcodes can include polynucleotide spatial barcodes; random nucleic acid and / or amino acid sequences; and synthetic nucleic acid and / or amino acid sequences. In some embodiments, a spatial barcode is attached to an analyte in a reversible or irreversible manner. In some embodiments, a spatial barcode allows for identification and / or quantification of individual sequencing-reads. In some embodiments, a spatial barcode is a used as a barcode for which fluorescently labeled oligonucleotide probes hybridize to the spatial barcode.
[0295] In some embodiments, the spatial barcode is a nucleic acid sequence that does not substantially hybridize to nucleic acid molecules such as mRNAs in a biological sample. In some embodiments, the spatial barcode has less than 80% sequence identity (e.g., less than 70%, 60%, 50%, or less than 40% sequence identity) to the nucleic acid molecules such as mRNAs across a substantial part (e.g., 80% or more) of the nucleic acid molecules in the biological sample.
[0296] The spatial barcode sequences can include from about 6 to about 20 or more nucleotides within the sequence of the capture probes. In some embodiments, the length of a spatial barcode sequence can be about 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 nucleotides or longer. In some embodiments, the length of a spatial barcode sequence can be at least about 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 nucleotides or longer. In some embodiments, the length of a spatial barcode sequence is at most about 6, 7, 8, 9, 10, 11, 12, 13,14, 15, 16, 17, 18, 19, 20 nucleotides or shorter. These nucleotides can be completely contiguous, e.g., in a single stretch of adjacent nucleotides, or they can be separated into two or more separate subsequences that are separated by 1 or more nucleotides. Separated spatial barcode subsequences can be from about 4 to about 16 nucleotides in length. In some embodiments, the spatial barcode subsequence can be about 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14,15, 16 nucleotides or longer. In some embodiments, the spatial barcode subsequence can be at least about 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16 nucleotides or longer. In some embodiments, the spatial barcode subsequence can be at most about 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16 nucleotides or shorter.98MOFO-358213026202412023740
[0297] For multiple capture probes that are attached to a common array feature, the one or more spatial barcode sequences of the multiple capture probes can include sequences that are the same for all capture probes coupled to the feature, and / or sequences that are different across all capture probes coupled to the feature. In some embodiments, a plurality of capture probes attached to a common array feature possesses the same spatial barcode.
[0298] In some embodiments, capture probes attached to a single array feature comprises identical (or common) spatial barcode sequences, different spatial barcode sequences, or a combination of both. Capture probes attached to a feature can include multiple sets of capture probes. Capture probes of a given set can include identical spatial barcode sequences. The identical spatial barcode sequences can be different from spatial barcode sequences of capture probes of another set.
[0299] The plurality of capture probes can include spatial barcode sequences that are associated with specific locations on a spatial array. For example, a first plurality of capture probes can be associated with a first region, based on a spatial barcode sequence common to the capture probes within the first region, and a second plurality of capture probes can be associated with a second region, based on a spatial barcode sequence common to the capture probes within the second region. The second region may or may not be associated with the first region. Additional pluralities of capture probes can be associated with spatial barcode sequences common to the capture probes within other regions. In some embodiments, the spatial barcode sequences can be the same across a plurality of capture probe molecules.
[0300] In some embodiments, multiple different spatial barcodes are incorporated into a single arrayed capture probe. For example, a mixed but known set of spatial barcode sequences can provide a stronger address or attribution of the spatial barcodes to a given spot or location, by providing duplicate or independent confirmation of the identity of the location. In some embodiments, the multiple spatial barcodes represent increasing specificity of the location of the particular array point. v. Unique Molecular Identifier
[0301] The capture probe can include one or more (e.g., two or more, three or more, four or more, five or more) Unique Molecular Identifiers (UMIs). A unique molecular identifier is a contiguous nucleic acid segment or two or more non-contiguous nucleic acid segments that99MOFO-358213026202412023740 function as a label or identifier for a particular analyte, or for a capture probe that binds a particular analyte (e.g., via the capture domain).
[0302] A UMI can be unique. A UMI can include one or more specific polynucleotides sequences, one or more random nucleic acid and / or amino acid sequences, and / or one or more synthetic nucleic acid and / or amino acid sequences.
[0303] In some embodiments, the UMI is a nucleic acid sequence that does not substantially hybridize to analyte nucleic acid molecules in a biological sample. In some embodiments, the UMI has less than 80% sequence identity (e.g., less than 70%, 60%, 50%, or less than 40% sequence identity) to the nucleic acid sequences across a substantial part (e.g., 80% or more) of the nucleic acid molecules in the biological sample.
[0304] The UMI can include from about 6 to about 20 or more nucleotides within the sequence of the capture probes. In some embodiments, the length of a UMI sequence can be about 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 nucleotides or longer. In some embodiments, the length of a UMI sequence can be at least about 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 nucleotides or longer. In some embodiments, the length of a UMI sequence is at most about 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 nucleotides or shorter. These nucleotides can be completely contiguous, e.g., in a single stretch of adjacent nucleotides, or they can be separated into two or more separate subsequences that are separated by 1 or more nucleotides. Separated UMI subsequences can be from about 4 to about 16 nucleotides in length. In some embodiments, the UMI subsequence can be about 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16 nucleotides or longer. In some embodiments, the UMI subsequence can be at least about 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16 nucleotides or longer. In some embodiments, the UMI subsequence can be at most about 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16 nucleotides or shorter.
[0305] In some embodiments, a UMI is attached to an analyte in a reversible or irreversible manner. In some embodiments, a UMI allows for identification and / or quantification of individual sequencing-reads. In some embodiments, a UMI is a used as a fluorescent barcode for which fluorescently labeled oligonucleotide probes hybridize to the UMI. vi. Other Aspects of Capture Probes
[0306] For capture probes that are attached to an array feature, an individual array feature can include one or more capture probes. In some embodiments, an individual array feature includes hundreds or thousands of capture probes. In some embodiments, the capture100MOFO-358213026202412023740 probes are associated with a particular individual feature, where the individual feature contains a capture probe including a spatial barcode unique to a defined region or location on the array.
[0307] In some embodiments, a particular feature can contain capture probes including more than one spatial barcode (e.g., one capture probe at a particular feature can include a spatial barcode that is different than the spatial barcode included in another capture probe at the same particular feature, while both capture probes include a second, common spatial barcode), where each spatial barcode corresponds to a particular defined region or location on the array. For example, multiple spatial barcode sequences associated with one particular feature on an array can provide a stronger address or attribution to a given location by providing duplicate or independent confirmation of the location. In some embodiments, the multiple spatial barcodes represent increasing specificity of the location of the particular array point. In a nonlimiting example, a particular array point can be coded with two different spatial barcodes, where each spatial barcode identifies a particular defined region within the array, and an array point possessing both spatial barcodes identifies the sub-region where two defined regions overlap, e.g., such as the overlapping portion of a Venn diagram.
[0308] In another non-limiting example, a particular array point can be coded with three different spatial barcodes, where the first spatial barcode identifies a first region within the array, the second spatial barcode identifies a second region, where the second region is a subregion entirely within the first region, and the third spatial barcode identifies a third region, where the third region is a subregion entirely within the first and second subregions.
[0309] In some embodiments, capture probes attached to array features are released from the array features for sequencing. Alternatively, in some embodiments, capture probes remain attached to the array features, and the probes are sequenced while remaining attached to the array features. Further aspects of the sequencing of capture probes are described in subsequent sections of this disclosure.
[0310] In some embodiments, an array feature can include different types of capture probes attached to the feature. For example, the array feature can include a first type of capture probe with a capture domain designed to bind to one type of analyte, and a second type of capture probe with a capture domain designed to bind to a second type of analyte. In general, array features can include one or more (e.g., two or more, three or more, four or more, five or101MOFO-358213026202412023740 more, six or more, eight or more, ten or more, 12 or more, 15 or more, 20 or more, 30 or more, 50 or more) different types of capture probes attached to a single array feature.
[0311] In some embodiments, the capture probe is nucleic acid. In some embodiments, the capture probe is attached to the array feature via its 5’ end. In some embodiments, the capture probe includes from the 5’ to 3’ end: one or more barcodes (e.g., a spatial barcode and / or a UMI) and one or more capture domains. In some embodiments, the capture probe includes from the 5’ to 3’ end: one barcode (e.g., a spatial barcode or a UMI) and one capture domain. In some embodiments, the capture probe includes from the 5’ to 3’ end: a cleavage domain, a functional domain, one or more barcodes (e.g., a spatial barcode and / or a UMI), and a capture domain. In some embodiments, the capture probe includes from the 5’ to 3’ end: a cleavage domain, a functional domain, one or more barcodes (e.g., a spatial barcode and / or a UMI), a second functional domain, and a capture domain. In some embodiments, the capture probe includes from the 5’ to 3’ end: a cleavage domain, a functional domain, a spatial barcode, a UMI, and a capture domain. In some embodiments, the capture probe does not include a spatial barcode. In some embodiments, the capture probe does not include a UMI. In some embodiments, the capture probe includes a sequence for initiating a sequencing reaction.
[0312] In some embodiments, the capture probe is immobilized on a feature via its 3’ end. In some instances, the capture probe comprises: an adapter sequence - a barcode (e.g., a spatial barcode) - an optional unique molecular identifier (UMI) sequence - a capture domain. In some embodiments, the capture probe includes from the 3’ to 5’ end: one or more barcodes (e.g., a spatial barcode and / or a UMI) and one or more capture domains. In some embodiments, the capture probe includes from the 3’ to 5’ end: one barcode (e.g., a spatial barcode or a UMI) and one capture domain. In some embodiments, the capture probe includes from the 3’ to 5’ end: a cleavage domain, a functional domain, one or more barcodes (e.g., a spatial barcode and / or a UMI), and a capture domain. In some embodiments, the capture probe includes from the 3’ to 5’ end: a cleavage domain, a functional domain, a spatial barcode, a UMI, and a capture domain.
[0313] In some embodiments, a capture probe includes an in situ synthesized oligonucleotide. In some embodiments, the in situ synthesized oligonucleotide includes one or more constant sequences, one or more of which serves as a priming sequence (e.g., a primer for amplifying target nucleic acids). In some embodiments, a constant sequence is a cleavable sequence. In some embodiments, the in situ synthesized oligonucleotide includes a barcode102MOFO-358213026202412023740 sequence, e.g., a variable barcode sequence. In some embodiments, the in situ synthesized oligonucleotide is attached to a feature of an array.
[0314] In some embodiments, a capture probe is a product of two or more oligonucleotide sequences, e.g., two or more oligonucleotide sequences that are ligated together. In some embodiments, one of the oligonucleotide sequences is an in situ synthesized oligonucleotide.
[0315] In some embodiments, the capture probe includes a splint oligonucleotide. Two or more oligonucleotides can be ligated together using a splint oligonucleotide and any variety of suitable ligases described herein (e.g., a SplintR® ligase).
[0316] In some embodiments, one of the oligonucleotides includes: a constant sequence (e.g., a sequence complementary to a portion of a splint oligonucleotide), a degenerate sequence, and a capture domain (e.g., as described herein). In some embodiments, the capture probe is generated by having an enzyme add polynucleotides at the end of an oligonucleotide sequence. The capture probe can include a degenerate sequence, which can function as a unique molecular identifier.
[0317] A capture probe can include a degenerate sequence, which is a sequence in which some positions of a nucleotide sequence contain a number of possible bases. A degenerate sequence can be a degenerate nucleotide sequence including about or at least 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 45, or 50 nucleotides. In some embodiments, a nucleotide sequence contains 1, 2, 3, 4, 5, 6, 7, 8, 9, 0, 10, 15, 20, 25, or more degenerate positions within the nucleotide sequence. In some embodiments, the degenerate sequence is used as a UMI.
[0318] In some embodiments, a capture probe includes a restriction endonuclease recognition sequence or a sequence of nucleotides cleavable by specific enzyme activities. For example, uracil sequences can be cleaved by specific enzyme activity. As another example, other modified bases (e.g., modified by methylation) can be recognized and cleaved by specific endonucleases. The capture probes can be subjected to an enzymatic cleavage, which removes the blocking domain and any of the additional nucleotides that are added to the 3’ end of the capture probe during the modification process. The removal of the blocking domain reveals and / or restores the free 3’ end of the capture domain of the capture probe. In some embodiments,103MOFO-358213026202412023740 additional nucleotides can be removed to reveal and / or restore the 3’ end of the capture domain of the capture probe.
[0319] In some embodiments, a blocking domain can be incorporated into the capture probe when it is synthesized, or after its synthesis. The terminal nucleotide of the capture domain is a reversible terminator nucleotide (e.g., 3’-O-blocked reversible terminator and 3 ’-unblocked reversible terminator), and can be included in the capture probe during or after probe synthesis, vii. Extended Capture Probes
[0320] An “extended capture probe” is a capture probe with an enlarged nucleic acid sequence. For example, where the capture probe includes nucleic acid, an “extended 3’ end” indicates that further nucleotides were added to the most 3’ nucleotide of the capture probe to extend the length of the capture probe, for example, by standard polymerization reactions utilized to extend nucleic acid molecules including templated polymerization catalyzed by a polymerase (e.g., a DNA polymerase or reverse transcriptase).
[0321] In some embodiments, extending the capture probe includes generating cDNA from the captured (hybridized) RNA template (e.g., as described in Section II). In some instances, the extended oligonucleotide is generated by extending the capture probe. This process involves synthesis of a complementary strand of the hybridized nucleic acid, e.g., generating cDNA based on the captured RNA template (the RNA hybridized to the capture domain of the capture probe). Thus, in an initial step of extending the capture probe, e.g., the cDNA generation, the captured (hybridized) nucleic acid, e.g., RNA, acts as a template for the extension, e.g., reverse transcription, step. In some instances, a cleaved RNA template is captured by the capture probe is extended.
[0322] In some embodiments, the capture probe is extended using reverse transcription. For example, reverse transcription includes synthesizing cDNA (complementary or copy DNA) from RNA, e.g., (messenger RNA), using a reverse transcriptase. In some embodiments, reverse transcription is performed while the tissue is still in place, generating an analyte library, where the analyte library includes the spatial barcodes from the adjacent capture probes. In some embodiments, the capture probe is extended using one or more DNA polymerases.
[0323] In some embodiments, the capture domain of the capture probe includes a primer for producing the complementary strand of the nucleic acid hybridized to the capture104MOFO-358213026202412023740 probe, e.g., a primer for DNA polymerase and / or reverse transcription. The nucleic acid, e.g., DNA and / or cDNA, molecules generated by the extension reaction incorporate the sequence of the capture probe. The extension of the capture probe, e.g., a DNA polymerase and / or reverse transcription reaction, can be performed using a variety of suitable enzymes and protocols.
[0324] For example, the capture sequences in the RNA template, the extended oligonucleotide or a derivative thereof can be captured onto the capture array slide. In some embodiments, capture probes with capture domains can bind to capture sequences on an RNA template, an extended oligonucleotide or complements thereof disclosed herein. In some embodiments, one or more reactions (e.g., extension, and / or ligation) are performed to generate a spatially labeled polynucleotide sequence comprising a sequence of the RNA template, the extended oligonucleotide or complement thereof and a sequence of the spatial barcode or complement thereof.
[0325] In some embodiments, a nucleic acid barcode molecule of a spatial array comprises a capture sequence complementary to a sequence of a cleaved RNA template. In some embodiments, a nucleic acid barcode molecule of a spatial array comprises a capture sequence complementary to a sequence of a generated extended oligonucleotide or derivative thereof (as shown in FIG. 4). In some instances, a nucleic acid barcode molecule comprises a capture sequence complementary to a sequence of a generated extended oligonucleotide or derivative thereof. In some instances, the capture sequence comprises a sequence corresponding to a specific RNA template. In some instances, a capture sequence comprises a common capture sequence for hybridizing to different RNA templates or extended oligonucleotides or derivatives thereof. In some instances, a nucleic acid extension reaction is performed, thereby generate a spatially labeled polynucleotide sequence comprising a sequence of the RNA template or complement thereof and a sequence of the spatial barcode or complement thereof, and optionally a functional sequence.
[0326] In some embodiments, the 3’ end of the extended oligonucleotide is modified. For example, a linker or adaptor can be ligated to the 3’ end of the extended probes. This can be achieved using single stranded ligation enzymes such as T4 RNA ligase or Circligase™ (available from Epicentre Biotechnologies, Madison, WI). In some embodiments, template switching oligonucleotides are used to extend cDNA in order to generate a full-length cDNA (or as close to a full-length cDNA as possible). In some embodiments, an adaptase enzyme is used105MOFO-358213026202412023740 to modify the generated extended oligonucleotide or a derivative thereof. In some embodiments, a second strand synthesis helper probe (a partially double stranded DNA molecule capable of hybridizing to the 3’ end of the extended oligonucleotide), can be ligated to the 3’ end using a double stranded ligation enzyme such as T4 DNA ligase. Any suitable enzymes appropriate for the ligation step may be used and include, e.g., Tth DNA ligase, Taq DNA ligase, Thermococcus sp. (strain 9°N) DNA ligase (9°N™ DNA ligase, New England Biolabs), Ampligase™ (available from Epicentre Biotechnologies, Madison, WI), and SplintR® (available from New England Biolabs, Ipswich, MA).
[0327] In some embodiments, double- stranded extended capture probes are treated to remove any unextended capture probes prior to amplification and / or analysis, e.g. sequence analysis. This can be achieved by a variety of methods, e.g., using an enzyme to degrade the unextended probes, such as an exonuclease enzyme, or purification columns.
[0328] In some embodiments, extended capture probes are amplified to yield quantities that are sufficient for analysis, e.g., via DNA sequencing. In some embodiments, the first strand of the extended capture probes (e.g., DNA and / or cDNA molecules) acts as a template for the amplification reaction (e.g., a polymerase chain reaction).
[0329] In some embodiments, the amplification reaction incorporates an affinity group onto the extended capture probe (e.g., RNA-cDNA hybrid) using a primer including the affinity group. In some embodiments, the primer includes an affinity group and the extended capture probes includes the affinity group. The affinity group can correspond to any of the affinity groups described previously.
[0330] In some embodiments, the extended capture probes including the affinity group can be coupled to an array feature specific for the affinity group. In some embodiments, the array feature includes avidin or streptavidin and the affinity group includes biotin. In some embodiments, the array feature includes maltose and the affinity group includes maltose-binding protein. In some embodiments, the array feature includes maltose-binding protein and the affinity group includes maltose. In some embodiments, amplifying the extended capture probes can function to release the extended probes from the array feature, insofar as copies of the extended probes are not attached to the array feature.
[0331] In some embodiments, the extended capture probe or complement or amplicon thereof is released from an array feature. The step of releasing the extended capture probe or106MOFO-358213026202412023740 complement or amplicon thereof from an array feature can be achieved in a number of ways. In some embodiments, an extended capture probe or a complement thereof is released from the feature by nucleic acid cleavage and / or by denaturation (e.g., by heating to denature a doublestranded molecule). viii. Analysis of Captured Analytes
[0332] A wide variety of different sequencing methods can be used to analyze spatially barcoded oligonucleotides (e.g., barcoded extended oligonucleotides). In general, sequenced polynucleotides can be, for example, nucleic acid molecules such as deoxyribonucleic acid (DNA) or ribonucleic acid (RNA), including variants or derivatives thereof (e.g., single stranded DNA or DNA / RNA hybrids, and nucleic acid molecules with a nucleotide analog). In some embodiments, the spatially labeled polynucleotide (e.g., spatially barcoded extended oligonucleotide or a derivative thereof) comprises a sequence of the spatial barcode or complement thereof.
[0333] Sequencing of polynucleotides can be performed by various commercial systems. More generally, sequencing can be performed using nucleic acid amplification, polymerase chain reaction (PCR) (e.g., digital PCR and droplet digital PCR (ddPCR), quantitative PCR, real time PCR, multiplex PCR, PCR-based singleplex methods, emulsion PCR), and / or isothermal amplification.
[0334] Other examples of methods for sequencing genetic material include, but are not limited to, DNA hybridization methods (e.g., Southern blotting), restriction enzyme digestion methods, Sanger sequencing methods, next-generation sequencing methods (e.g., singlemolecule real-time sequencing, nanopore sequencing, and Polony sequencing), ligation methods, and microarray methods. Additional examples of sequencing methods that can be used include targeted sequencing, single molecule real-time sequencing, exon sequencing, electron microscopy-based sequencing, panel sequencing, transistor-mediated sequencing, direct sequencing, random shotgun sequencing, Sanger dideoxy termination sequencing, wholegenome sequencing, sequencing by hybridization, pyrosequencing, capillary electrophoresis, gel electrophoresis, duplex sequencing, cycle sequencing, single-base extension sequencing, solidphase sequencing, high-throughput sequencing, massively parallel signature sequencing, coamplification at lower denaturation temperature-PCR (COLD-PCR), sequencing by reversible dye terminator, paired-end sequencing, near-term sequencing, exonuclease sequencing,107MOFO-358213026202412023740 sequencing by ligation, short-read sequencing, single-molecule sequencing, sequencing-by- synthesis, real-time sequencing, reverse-terminator sequencing, nanopore sequencing, 454 sequencing, Solexa Genome Analyzer sequencing, SOLiD™ sequencing, MS-PET sequencing, and any combinations thereof.
[0335] In some embodiments, direct sequencing of one or more captured analytes is performed by sequencing-by- synthesis (SBS). In some embodiments, a sequencing primer is complementary to a sequence in one or more of the domains of a capture probe (e.g., functional domain). In such embodiments, sequencing-by- synthesis can include reverse transcription and / or amplification in order to generate a template sequence (e.g., functional domain) from which a primer sequence can bind.
[0336] SBS can involve hybridizing an appropriate primer, sometimes referred to as a sequencing primer, with the nucleic acid template to be sequenced, extending the primer, and detecting the nucleotides used to extend the primer. Preferably, the nucleic acid used to extend the primer is detected before a further nucleotide is added to the growing nucleic acid chain, thus allowing base-by-base nucleic acid sequencing. The detection of incorporated nucleotides is facilitated by including one or more labelled nucleotides in the primer extension reaction. To allow the hybridization of an appropriate sequencing primer to the nucleic acid template to be sequenced, the nucleic acid template should normally be in a single stranded form. If the nucleic acid templates making up the nucleic acid spots are present in a double stranded form these can be processed to provide single stranded nucleic acid templates using any suitable methods, for example by denaturation, cleavage etc. The sequencing primers which are hybridized to the nucleic acid template and used for primer extension are preferably short oligonucleotides, for example, 15 to 25 nucleotides in length. The sequencing primers can be provided in solution or in an immobilized form. Once the sequencing primer has been annealed to the nucleic acid template to be sequenced by subjecting the nucleic acid template and sequencing primer to appropriate conditions, primer extension is carried out, for example using a nucleic acid polymerase and a supply of nucleotides, at least some of which are provided in a labelled form, and conditions suitable for primer extension if a suitable nucleotide is provided. In some embodiments, a nucleic acid molecule or a complement thereof captured on the spatial array is released for sequencing. In some embodiments, the released nucleic acid molecule or complement thereof comprises the variant sequence. In some embodiments, the released nucleic108MOFO-358213026202412023740 acid molecule or complement thereof is a spatially labeled polynucleotide comprising (i) a sequence corresponding to the RNA template or complement thereof and (ii) a sequence of the spatial barcode or complement thereof.III. Samples and AnalytesA. Samples
[0337] In some aspects, a sample disclosed herein is or is derived from any biological sample. Methods and compositions disclosed herein may be used for analyzing a biological sample, which may be obtained from a subject using any of a variety of techniques including, but not limited to, biopsy, surgery, and laser capture microscopy (LCM), and generally includes cells and / or other biological material from the subject. In addition to the subjects described above, a biological sample can be obtained from a prokaryote such as a bacterium, an archaea, a virus, or a viroid. A biological sample can also be obtained from non-mammalian organisms (e.g., a plant, an insect, an arachnid, a nematode, a fungus, or an amphibian). A biological sample can also be obtained from a eukaryote, such as a tissue sample, a patient derived organoid (PDO) or patient derived xenograft (PDX). A biological sample from an organism may comprise one or more other organisms or components therefrom. For example, a mammalian tissue section may comprise a prion, a viroid, a virus, a bacterium, a fungus, or components from other organisms, in addition to mammalian cells and non-cellular tissue components. Subjects from which biological samples can be obtained can be healthy or asymptomatic individuals, individuals that have or are suspected of having a disease (e.g., a patient with a disease such as cancer) or a predisposition to a disease, and / or individuals in need of therapy or suspected of needing therapy.
[0338] The biological sample can include any number of macromolecules, for example, cellular macromolecules and organelles (e.g., mitochondria and nuclei). In some aspects, the biological sample includes nucleic acids (such as DNA or RNA), proteins / polypeptides, carbohydrates, and / or lipids. In some embodiments, the biological sample is obtained as a tissue sample, such as a tissue section, biopsy, a core biopsy, needle aspirate, or fine needle aspirate. In some embodiments, the biological sample is or comprise a cell pellet or a section of a cell pellet. In some embodiments, the biological sample is or comprise a cell block or a section of a cell block. In some aspects, a biological sample is a fluid sample, such as a blood sample, urine sample, or saliva sample. The sample can be a skin sample, a colon sample, a cheek swab, a histology sample, a histopathology sample, a plasma or serum sample, a tumor 109MOFO-358213026202412023740 sample, living cells, cultured cells, a clinical sample such as, for example, whole blood or blood- derived products, blood cells, or cultured tissues or cells, including cell suspensions. In some embodiments, the biological sample comprises cells which are deposited on a surface. In some embodiments, the RNA templates described herein are in a cell or tissue sample.
[0339] In some aspects, a biological sample is derived from a homogeneous culture or population of the subjects or organisms mentioned herein or alternatively from a collection of several different organisms. In some aspects, a biological sample comprises one or more diseased cells. A diseased cell can have altered metabolic properties, gene expression, protein expression, and / or morphologic features. Examples of diseases include inflammatory disorders, metabolic disorders, nervous system disorders, and cancer. Cancer cells can be derived from solid tumors, hematological malignancies, cell lines, or obtained as circulating tumor cells. Biological samples can also include fetal cells and immune cells.
[0340] In some embodiments, a substrate herein can be any support (e.g., a solid support) that is insoluble in aqueous liquid and which allows for positioning of biological samples, analytes, features, and / or reagents on the support. In some embodiments, a biological sample is attached to a substrate (e.g., a slide). Attachment of the biological sample can be irreversible or reversible, depending upon the nature of the sample and subsequent steps in the analytical method. In certain embodiments, the sample is attached to the substrate reversibly by applying a suitable polymer coating to the substrate, and contacting the sample to the polymer coating. The sample can then be detached from the substrate, e.g., using an organic solvent that at least partially dissolves the polymer coating. Hydrogels are examples of polymers that are suitable for this purpose. In some embodiments, the substrate is coated or functionalized with one or more substances to facilitate attachment of the sample to the substrate. Suitable substances that can be used to coat or functionalize the substrate include, but are not limited to, lectins, poly-lysine, antibodies, and polysaccharides.
[0341] A variety of steps can be performed to prepare or process a biological sample for and / or during an assay. Except where indicated otherwise, the preparative or processing steps described below can generally be combined in any manner and in any order to appropriately prepare or process a particular sample for and / or analysis.110MOFO-358213026202412023740(i) Preparation
[0342] In some aspects, a biological sample is harvested from a subject (e.g., via surgical biopsy, whole subject sectioning) or grown in vitro on a growth substrate or culture dish as a population of cells, and prepared for analysis as a tissue slice or tissue section. Grown samples may be sufficiently thin for analysis without further processing steps. Alternatively, grown samples, and samples obtained via biopsy or sectioning, can be prepared as thin tissue sections using a mechanical cutting apparatus such as a vibrating blade microtome. As another alternative, in some embodiments, a thin tissue section can be prepared by applying a touch imprint of a biological sample to a suitable substrate material.
[0343] The thickness of the tissue section can be a fraction of (e.g., less than 0.9, 0.8, 0.7, 0.6, 0.5, 0.4, 0.3, 0.2, or 0.1) the maximum cross-sectional dimension of a cell. However, tissue sections having a thickness that is larger than the maximum cross-section cell dimension can also be used. For example, cryostat sections can be used, which can be, e.g., 10-20 pm thick. More generally, the thickness of a tissue section typically depends on the method used to prepare the section and the physical characteristics of the tissue, and therefore sections having a wide variety of different thicknesses can be prepared and used. For example, the thickness of the tissue section can be at least 0.1, 0.2, 0.3, 0.4, 0.5, 0.7, 1.0, 1.5, 2, 3, 4, 5, 6, 7, 8, 9, 10, 12, 13, 14, 15, 20, 30, 40, or 50 pm. Thicker sections can also be used if desired or convenient, e.g., at least 70, 80, 90, or 100 pm or more. Typically, the thickness of a tissue section is between 1-100 pm, 1-50 pm, 1-30 pm, 1-25 pm, 1-20 pm, 1-15 pm, 1-10 pm, 2-8 pm, 3-7 pm, or 4-6 pm, but as mentioned above, sections with thicknesses larger or smaller than these ranges can also be analysed.
[0344] In some examples, multiple sections are obtained from a single biological sample. For example, multiple tissue sections can be obtained from a surgical biopsy sample by performing serial sectioning of the biopsy sample using a sectioning blade. Spatial information among the serial sections can be preserved in this manner, and the sections can be analysed successively to obtain three-dimensional information about the biological sample.
[0345] In some embodiments, the biological sample (e.g., a tissue section as described above) is prepared by deep freezing at a temperature suitable to maintain or preserve the integrity (e.g., the physical characteristics) of the tissue structure. The frozen tissue sample can be sectioned, e.g., thinly sliced, onto a substrate surface using any number of suitable111MOFO-358213026202412023740 methods. For example, a tissue sample can be prepared using a chilled microtome (e.g., a cryostat) set at a temperature suitable to maintain both the structural integrity of the tissue sample and the chemical properties of the nucleic acids in the sample. Such a temperature can be, e.g., less than -15°C, less than -20°C, or less than -25°C.
[0346] In some embodiments, the biological sample is prepared using formalinfixation and paraffin-embedding (FFPE), which are established methods. In some embodiments, cell suspensions and other non-tissue samples can be prepared using formalin-fixation and paraffin-embedding. Following fixation of the sample and embedding in a paraffin or resin block, the sample can be sectioned as described above. Prior to analysis, the paraffin-embedding material can be removed from the tissue section (e.g., deparaffinization) by incubating the tissue section in an appropriate solvent (e.g., xylene) followed by a rinse (e.g., 99.5% ethanol for 2 minutes, 96% ethanol for 2 minutes, and 70% ethanol for 2 minutes). In some embodiments, the biological sample (e.g., FFPE sample) is permeable after deparaffinization. In some embodiments, processing of the biological sample, such as de-waxing, allows the biological sample to b...
Claims
202412023740CLAIMSWhat is claimed is:
1. A method of nucleic acid processing, comprising:(a) hybridizing an oligonucleotide to a ribonucleic acid (RNA) template;(b) extending the oligonucleotide using a reverse transcriptase and a plurality of free nucleotides to generate an extended oligonucleotide, wherein the plurality of free nucleotides comprises a terminating nucleotide that prevents further extension of the extended oligonucleotide;(c) digesting the RNA template; and(d) using the extended oligonucleotide to generate a circularized template.
2. The method of claim 1, wherein the terminating nucleotide is a nucleotide lacking a 3' hydroxyl group in the deoxyribose sugar.
3. The method of claim 2, wherein the terminating nucleotide is a dideoxyribonucleotide (ddNTP).
4. The method of claim 2 or claim 3, wherein the plurality of free nucleotides comprises two, three or four different ddNTPs selected from the group consisting of ddGTP, ddATP, ddTTP and ddCTP.
5. The method of claim 1, wherein the terminating nucleotide comprises a terminating group or a terminating modification.
6. The method of claim 5, wherein the terminating modification is a reverse linkage.
7. The method of claim 6, wherein the terminating nucleotide comprises a 3’ inverted dT.
8. The method of any one of claims 1-7, wherein the terminating nucleotide is present in an amount lower than an amount of a corresponding free nucleotide in the plurality of free nucleotides.145MOFO-3582130262024120237409. The method of any one of claims 1-7, wherein the terminating nucleotide is present in an amount greater than an amount of a corresponding free nucleotide in the plurality of free nucleotides.
10. The method of claim 8, wherein the amount of the terminating nucleotide is at least 2- fold lower than the amount of the corresponding free nucleotide in the plurality of free nucleotides.
11. The method of claim 9, wherein the amount of the terminating nucleotide is at least 2- fold greater than the amount of the corresponding free nucleotide in the plurality of free nucleotides.
12. The method of any one of claims 1-7, wherein the plurality of free nucleotides comprises a cleavable nucleotide.
13. The method of claim 12, wherein the cleavable nucleotide is a uracil.
14. The method of claim 12, wherein the cleavable nucleotide is an inosine.
15. The method of claim 12, wherein the cleavable nucleotide is a ribonucleotide and the extended oligonucleotide is DNA.
16. The method of claim 13, further comprising using Uracil-Specific Excision Reagent (USER) to cleave the extended oligonucleotide before the using in (d).
17. The method of claim 14, further comprising using an endonuclease to cleave the extended oligonucleotide before the using in (d).
18. The method of claim 1, wherein the terminating nucleotide is a reversible terminator nucleotide.
19. The method of claim 18, wherein the reversible terminator nucleotide comprises an azidomethyl group, an amino group, a nitrobenzyl group, an allyl group, a carbonate, a functional photocleavable ether, a methyl group, or a cyanoethyl group.146MOFO-35821302620241202374020. The method of claim 19, wherein the reversible terminator nucleotide comprises an azidomethyl group, and amino group, a nitrobenzyl group, or an allyl group.
21. The method of claim 20, wherein the reversible terminator nucleotide is a 3'-O-blocked reversible terminator nucleotide.
22. The method of any one of claims 5-21, wherein the plurality of free nucleotides comprises two, three or four different ddNTPs comprising the terminating group or the terminating modification.
23. The method of any one of claims 1-22, further comprising before the using in (d), removing the terminating nucleotide from the extended oligonucleotide.
24. The method of any one of claims 1-23, further comprising before the using in (d), cleaving the terminating nucleotide from the extended oligonucleotide.
25. The method of any one of claims 5-23, further comprising before the using in (d), removing the terminating group or the terminating modification from the terminating nucleotide incorporated into the extended oligonucleotide.
26. The method of claim 25, wherein removing the terminating group or the terminating modification from the terminating nucleotide is performed by cleaving a linker in the terminating nucleotide.
27. The method of any one of claims 1-26, wherein the plurality of free nucleotides comprise all four canonical bases: adenine, thymine, guanine and cytosine; and at least two different bases of terminating nucleotides.
28. The method of any one of claims 1-27, further comprising before the using in (d), using a kinase to phosphorylate a 5’ end of the extended oligonucleotide.
29. The method of any one of claims 1-28, wherein the using in (d) comprises ligating a 5’ end and a 3’ end of the extended oligonucleotide to generate the circularized template.147MOFO-35821302620241202374030. The method of any one of claims 1-29, wherein the extended oligonucleotide is used as a template to generate a copy of the extended oligonucleotide and a 5’ end and a 3’ end of the copy of the extended oligonucleotide is ligated to generate the circularized template.
31. A method, comprising:(a) hybridizing an oligonucleotide to an RNA template;(b) extending the oligonucleotide using a reverse transcriptase and a plurality of free nucleotides to generate an extended oligonucleotide, wherein the plurality of free nucleotides comprises a ribonucleotide, an inosine, and / or a uracil;(c) digesting the RNA template and cleaving the extended oligonucleotide at the ribonucleotide, the inosine, and / or the uracil that is incorporated into the extended oligonucleotide, thereby generating a cleaved extended sequence; and(d) using the cleaved extended sequence to generate a circularized template.
32. The method of claim 31, wherein a 5’ end and a 3’ end of the cleaved extended sequence is ligated to generate the circularized template.
33. The method of claim 31 or claim 32, wherein the inosine is incorporated into the extended oligonucleotide and the digesting in (c) comprises using an endonuclease to cleave the extended oligonucleotide.
34. The method of claim 31, wherein the uracil is incorporated into the extended oligonucleotide and the digesting in (c) comprises using Uracil-Specific Excision Reagent (USER) to cleave the extended oligonucleotide.
35. The method of any one of claims 31-34, wherein the plurality of free nucleotides comprise all four canonical bases: adenine, thymine, guanine and cytosine; and at least two different ribonucleotides.
36. The method of any one of claims 31-34, wherein the plurality of free nucleotides comprise the inosine and all four canonical bases: adenine, thymine, guanine and cytosine.
37. The method of any one of claims 31-34, wherein the plurality of free nucleotides comprise the uracil and all four canonical bases: adenine, thymine, guanine and cytosine.148MOFO-35821302620241202374038. The method of any one of claims 1-37, further comprising before the using in (d), using a kinase to phosphorylate a 5’ end of the extended oligonucleotide.
39. The method of any one of claims 1-38, wherein the using in (d) comprises ligating a 5’ end and a 3’ end of the extended oligonucleotide to generate the circularized template.
40. The method of any one of claims 31-39, wherein RNase H is used to perform the digesting and / or cleaving.
41. The method of any one of claims 1-40, wherein the using in (d) comprises performing an enzymatic ligation.
42. The method of any one of claims 1-40, wherein the using in (d) comprises performing a chemical ligation.
43. The method of claim 41, wherein the enzymatic ligation is performed using CircLigase.
44. The method of any one of claims 1-43, wherein the using in (d) comprises performing a ligation using a splint.
45. The method of any one of claims 1-44, wherein the circularized template does not comprise a barcode sequence.
46. The method of any one of claims 1-45, wherein the oligonucleotide comprises a primer binding sequence.
47. The method of any one of claims 1-46, further comprising:(e) performing rolling circle amplification of the circularized template to generate a rolling circle amplification product (RCP); and(f) detecting the RCP.
48. The method of claim 47, wherein the performing rolling circle amplification in (e) comprises binding a primer to a primer binding sequence of the circularized template, and extending the primer to generate an amplification product comprising multiple copies of the RNA template or a complement thereof.149MOFO-35821302620241202374049. The method of claim 47 or claim 48, comprising using a polymerase having stranddisplacement activity to generate the RCP.
50. The method of claim 49, wherein the polymerase is a Phi29 polymerase.
51. The method of any one of claims 1-50, wherein the extended oligonucleotide has a length of 50-200 nucleotides.
52. The method of any one of claims 1-51, wherein the extended oligonucleotide has a length of 70-100 nucleotides.
53. The method of any one of claims 1-52, wherein the method is performed in a biological sample.
54. The method of claim 53, wherein the biological sample is a cell or tissue sample.
55. The method of any one of claims 47-54, wherein detecting the RCP comprises detecting a sequence of the RCP using sequencing by ligation, sequencing by synthesis, sequencing by binding, sequencing by avidity, or a combination thereof.
56. The method of claim 47, wherein the detecting in (f) comprises binding a sequencing primer to the RCP.
57. The method of claim 56, wherein the sequencing primer binds to the primer binding sequence or a complement thereof.
58. The method of any one of claims 53-57, wherein the method comprises imaging the biological sample to detect the RCP in situ in the biological sample or a matrix embedding the biological sample.
59. The method of claim 55, wherein the sequence of the RCP or a complement thereof is sequenced using in situ sequencing by synthesis (SBS) in the biological sample.
60. The method of claim 59, wherein the sequence of the RCP or a complement thereof is detected using single nucleotide sequencing by synthesis.150MOFO-35821302620241202374061. The method of claim 55, wherein the sequence of the RCP or a complement thereof is sequenced using in situ sequencing by ligation (SBL) in the biological sample.
62. The method of claim 55, wherein the sequence of the RCP or a complement thereof is sequenced using in situ sequencing by binding (SBB) or sequencing-by-avidity (SB A) in the biological sample.
63. The method of claim 59, comprising contacting a sequencing primer bound to the RCP with (i) a polymerase and (ii) a first plurality of nucleotide molecules to form a complex comprising a 3’ terminus of the sequencing primer, the RCP, the polymerase, and a nucleotide molecule of the first plurality of nucleotide molecules and detecting presence of the nucleotide molecules in the complex to identify a complementary nucleotide in the RCP.
64. The method of claim 59, wherein a sequencing primer comprises a 3’ terminal nucleotide that is reversibly blocked.
65. The method of claim 64, wherein the method further comprises removing the complex, unblocking the reversibly blocked 3’ terminal nucleotide molecule and contacting the sequencing primer bound to the RCP with a polymerase and a second plurality of nucleotide molecules.
66. The method of any one of claims 59-65, further comprising repeating contacting a sequencing primer with an additional plurality of nucleotide molecules to identify additional complementary nucleotides of the sequence in the RCP for at least 2, 5, 10, 20, or 30 additional cycles.
67. The method of any one of claims 53-66, wherein the RNA template is attached directly or indirectly to the biological sample or to a matrix embedding the biological sample.
68. The method of any one of claims 53-67, wherein the RNA template is crosslinked in the biological sample or in a matrix embedding the biological sample.
69. A method of nucleic acid processing, comprising:(a) hybridizing an oligonucleotide to an RNA template in a biological sample;(b) extending the oligonucleotide using a reverse transcriptase and a plurality of free nucleotides to generate an extended oligonucleotide, wherein the plurality of free nucleotides151MOFO-358213026202412023740 comprises a reversible terminator nucleotide comprising a terminating group and incorporation of the reversible terminator nucleotide stops further extension of the extended oligonucleotide;(c) removing the terminating group from the reversible terminator nucleotide incorporated into the extended oligonucleotide;(d) digesting the RNA template;(e) ligating the 5’ and 3’ ends of the extended oligonucleotide to generate a circularized template;(f) performing rolling circle amplification of the circularized template to generate a rolling circle amplification product (RCP); and(g) detecting a sequence of the RCP in the biological sample using in situ sequencing by synthesis (SBS) in the biological sample, thereby detecting a sequence of the RCP or a complement thereof at a location in the biological sample.
70. The method of claim 69, wherein the digesting in (d) is performed prior to the removing in (c).
71. The method of claim 69, wherein the removing in (c) is performed prior to the digesting in (d).
72. The method of any one of claims 53-71, wherein the biological sample is a cell or tissue sample comprising cells or cellular components.
73. The method of any one of claims 53-72, wherein the biological sample is a tissue section.
74. The method of any one of claims 53-73, wherein the biological sample is a formalin- fixed, paraffin-embedded (FFPE) sample, a frozen tissue sample, or a fresh tissue sample.
75. The method of any one of claims 53-74, wherein the biological sample is fixed and / or permeabilized.
76. The method of any one of claims 53-75, wherein the biological sample is crosslinked and / or embedded in a matrix.
77. The method of claim 76, wherein the matrix comprises a hydrogel.
78. The method of any one of claims 53-77, wherein the biological sample is cleared.152MOFO-35821302620241202374079. A system, comprising:(i) an oligonucleotide;(ii) a plurality of free nucleotides that comprise different nucleic acid bases, wherein the plurality of free nucleotides comprises a terminating nucleotide that prevents further extension; and(iii) a reverse transcriptase for extending the oligonucleotide to generate an extended oligonucleotide.
80. The system of claim 79, further comprising one or more reagents for circularizing the extended oligonucleotide.
81. The system of claim 80, further comprising a polymerase having strand-displacement activity.
82. The system of claim 81, wherein the polymerase is a Phi29 polymerase.
83. The system of claim 82, wherein the oligonucleotide comprises a primer binding sequence.
84. The system of any one of claims 79-83, comprising a plurality of a sequencing primer, a plurality of detectably labeled nucleotides, and a polymerase.
85. The system of any one of claims 79-84, wherein the terminating nucleotide is a reversible terminator nucleotide.
86. The system of claim 85, wherein the reversible terminator nucleotide comprises an azidomethyl group, an amino group, a nitrobenzyl group, an allyl group, a carbonate, a functional photocleavable ether, a methyl group, or a cyanoethyl group.
87. The system of claim 86, wherein the reversible terminator nucleotide comprises an azidomethyl group, and amino group, a nitrobenzyl group, or an allyl group.
88. The system of claim 87, wherein the reversible terminator nucleotide is a 3'-O-blocked reversible terminator nucleotide.153MOFO-35821302620241202374089. The system of claim 84, wherein the oligonucleotide comprises a primer binding sequence configured to hybridize to the sequencing primer.
90. A method of nucleic acid processing, comprising: in a biological sample, hybridizing a nucleic acid strand to a cleavage region in a ribonucleic acid (RNA) template; cleaving the RNA template in the cleavage region; hybridizing an oligonucleotide to the cleaved RNA template; and extending the oligonucleotide using a reverse transcriptase and a plurality of free nucleotides to generate an extended oligonucleotide comprising a sequence complementary to the cleaved RNA template.
91. The method of claim 90, further comprising detecting a sequence of the extended oligonucleotide or derivative thereof.
92. The method of claim 90 or claim 91, wherein cleaving the RNA template comprises contacting the biological sample with an RNase H.
93. The method of any one of claims 90-92, wherein the cleavage region is 5’ to a region of interest in the RNA template.
94. The method of any one of claims 90-93, further comprising using the extended oligonucleotide to generate a circularized template.
95. The method of any one of claims 92-94, wherein the biological sample is contacted with the nucleic acid strand and with the RNase H simultaneously.
96. The method of claim 94 or claim 95, further comprising using a kinase to phosphorylate a 5’ end of the extended oligonucleotide prior to generating the circularized template.
97. The method of any one of claims 94-96, wherein generating the circularized template comprises ligating a 5’ end and a 3’ end of the extended oligonucleotide to generate the circularized template.154MOFO-35821302620241202374098. The method of any one of claims 94-96, wherein the extended oligonucleotide is used as a template to generate a copy of the extended oligonucleotide and a 5’ end and a 3’ end of the copy of the extended oligonucleotide is ligated to generate the circularized template.
99. The method of any one of claims 92-98, wherein the RNase H comprises an RNase Hl and / or an RNase H2.
100. The method of any one of claims 92-99, wherein contacting the biological sample with the RNase H comprises contacting the biological sample with between about 0.5 enzyme units (U) and about 50 U of the RNase H.
101. The method of any one of claims 90-100, wherein the cleavage region is about 10 to about 22 nucleotides in length, or about 15 to about 20 nucleotides in length.
102. The method of any one of claims 94-101, wherein the circularized template is generated by an enzymatic ligation.
103. The method of any one of claims 94-101, wherein the circularized template is generated by a chemical ligation.
104. The method of claim 102, wherein the enzymatic ligation is performed using CircLigase.
105. The method of any one of claims 94-104, wherein the circularized template is generated by performing a ligation using a splint.
106. The method of any one of claims 94-105, wherein the circularized template does not comprise a barcode sequence.
107. The method of any one of claims 90-106, wherein the oligonucleotide comprises a primer binding sequence.
108. The method of any one of claims 90-107, wherein the oligonucleotide is complementary to a sequence 3’ to a region of interest in the RNA template.155MOFO-358213026202412023740109. The method of any one of claims 90-107, wherein the oligonucleotide comprises a poly-T sequence or a randomer.
110. The method of any one of claims 94-109, further comprising: performing rolling circle amplification (RCA) of the circularized template to generate a rolling circle amplification product (RCP).
111. The method of claim 110, wherein performing RCA comprises binding a primer to the primer binding sequence of the circularized template, and extending the primer to generate an amplification product comprising multiple copies of the RNA template or a complement thereof.
112. The method of claim 110 or claim 111, comprising using a polymerase having stranddisplacement activity to generate the RCP.
113. The method of claim 112, wherein the polymerase is a Phi29 polymerase.
114. The method of any one of claims 90-113, wherein the extended oligonucleotide has a length of 50-200 nucleotides.
115. The method of any one of claims 90-114, wherein the biological sample is a cell or tissue sample.
116. The method of any one of claims 110-115, wherein detecting the sequence of the extended oligonucleotide or derivative thereof comprises detecting a sequence of the RCP using sequencing by ligation, sequencing by synthesis, sequencing by binding, sequencing by avidity, or a combination thereof.
117. The method of claim 116, wherein detecting the sequence of the RCP comprises binding a sequencing primer to the RCP.
118. The method of claim 117, wherein the sequencing primer binds to the primer binding sequence or a complement thereof.
119. The method of any one of claims 110-118, wherein the method comprises imaging the biological sample to detect a sequence of the RCP in situ in the biological sample or a matrix embedding the biological sample.156MOFO-358213026202412023740120. The method of any one of claims 116-119, wherein the sequence of the RCP comprises a region of interest or a complement of a sequence in the region of interest.
121. The method of any one of claims 116-120, wherein the sequence of the RCP or a complement thereof is sequenced using in situ sequencing by synthesis (SBS) in the biological sample.
122. The method of any one of claims 116-120, wherein the sequence of the RCP or a complement thereof is sequenced using in situ sequencing by avidity (SBA) in the biological sample.
123. The method of any one of claims 116-120, wherein the sequence of the RCP or a complement thereof is sequenced using in situ sequencing by ligation (SBL) in the biological sample.
124. The method of any one of claims 116-120, wherein the sequence of the RCP or a complement thereof is sequenced using in situ sequencing by binding (SBB) in the biological sample.
125. The method of any one of claims 117-122, comprising contacting the sequencing primer bound to the RCP with (i) a polymerase and (ii) a first plurality of nucleotide molecules to form a complex comprising a 3’ terminus of the sequencing primer, the RCP, the polymerase, and a nucleotide molecule of the first plurality of nucleotide molecules and detecting presence of the nucleotide molecules in the complex to identify a complementary nucleotide in the RCP.
126. The method of any one of claims 117-122, wherein the sequencing primer comprises a 3’ terminal nucleotide that is reversibly blocked.
127. The method of claim 126, wherein the method further comprises removing the complex, unblocking the reversibly blocked 3’ terminal nucleotide molecule and contacting the sequencing primer bound to the RCP with a polymerase and a second plurality of nucleotide molecules.
128. The method of any one of claims 117-122, further comprising repeating contacting the sequencing primer with an additional plurality of nucleotide molecules to identify additional157MOFO-358213026202412023740 complementary nucleotides of the sequence in the R.CP for at least 2, at least 5, at least 10, at least 20, or at least 30 additional cycles.
129. The method of any one of claims 90-128, wherein the RNA template is attached directly or indirectly to the biological sample or to a matrix embedding the biological sample.
130. The method of any one of claims 90-128, wherein the RNA template is crosslinked in the biological sample or in a matrix embedding the biological sample.
131. The method of any one of claims 90-130, further comprising digesting the cleaved RNA template after the extended oligonucleotide is generated.
132. The method of claim 131, wherein additional RNase H is used to digest the cleaved RNA template.
133. The method of any one of claims 90-93 and 129-132, wherein the extended oligonucleotide is a barcoded nucleic acid molecule.
134. The method of any one of claims 90-93 and 129-132, wherein the extended oligonucleotide is used to generate a barcoded nucleic acid molecule.
135. The method of any one of claims 90-93 and 129-134, wherein the oligonucleotide or a derivative of the extended oligonucleotide comprises a capture sequence.
136. The method of claim 135, wherein the capture sequence is complementary to a capture domain of a capture probe.
137. The method of claim 136, wherein the capture probe is immobilized on a substrate.
138. The method of any one of claims 90-93 and 129-136, wherein the oligonucleotide is immobilized on a substrate.
139. The method of any one of claims 136-138, wherein the capture probe comprises a spatial barcode.
140. The method of claim 138 or claim 139, wherein the substrate is a bead.158MOFO-358213026202412023740141. The method of claim 140, wherein the bead and the RNA template is in a partition.
142. The method of any one of claims 90-141, wherein the extended oligonucleotide is partitioned in a droplet or a well.
143. The method of any one of claims 133-142, further comprising processing the barcoded nucleic acid molecule to append a functional sequence.
144. The method of any one of claims 133-143, wherein detecting the sequence of the extended oligonucleotide or a derivative thereof comprises sequencing the barcoded nucleic acid molecule or a derivative thereof.
145. A kit, comprising: a nucleic acid strand, wherein the nucleic acid strand hybridizes to a cleavage region in a ribonucleic acid (RNA) template; an RNase H for cleaving the RNA template; a reverse transcriptase for generating an extended oligonucleotide; and an oligonucleotide configured to bind to the RNA template.
146. The kit of claim 145, further comprising a ligase for generating a circularized template and a polymerase for performing rolling circle amplification (RCA).
147. The kit of claim 145 or claim 146, further comprising a kinase.
148. The kit of any one of claims 145-147, further comprising a plurality of free nucleotides that comprise different nucleic acid bases.
149. The kit of any one of claims 146-148, wherein the ligase for generating the circularized template is a CircLigase.
150. The kit of any one of claims 145-149, wherein the RNase H comprises an RNase Hl and / or an RNAse H2.159MOFO-358213026202412023740151. The kit of any one of claims 145-150, wherein the nucleic acid strand is about 10 to about 22 nucleotides in length, or about 15 to about 20 nucleotides in length.
152. The kit of any one of claims 145-151, wherein the oligonucleotide comprises a primer binding sequence.
153. The kit of any one of claims 145-152, wherein the oligonucleotide is complementary to a sequence 3’ to a region of interest in the RNA template.
154. The kit of any one of claims 145-152, wherein the oligonucleotide comprises a poly-T sequence or a randomer.
155. The kit of any one of claims 146-154, wherein the polymerase for performing rolling circle amplification is a Phi29 polymerase.
156. The kit of any one of claims 145-155, comprising a plurality of a sequencing primer, a plurality of detectably labeled nucleotides, and a polymerase.
157. A system, comprising: a biological sample comprising a ribonucleic acid (RNA) template; a nucleic acid strand, wherein the nucleic acid strand hybridizes to a cleavage region in the RNA template; an RNase H for cleaving the RNA template; a reverse transcriptase for generating an extended oligonucleotide; and an oligonucleotide configured to bind to the RNA template.
158. The system of claim 157, further comprising a ligase for generating a circularized template; and a polymerase for performing rolling circle amplification (RCA).
159. The system of claim 157 or 158, wherein the biological sample is a cell or tissue sample provided on a solid support.
160. The system of any one of claims 157-159, further comprising a kinase.
161. The system of any one of claims 157-160, further comprising a plurality of free nucleotides that comprise different nucleic acid bases.160MOFO-358213026202412023740162. The system of any one of claims 158-161, wherein the ligase for generating the circularized template is a CircLigase.
163. The system of any one of claims 157-162, wherein the RNase H comprises an RNase Hl and / or an RNase H2.
164. The system of any one of claims 157-163, wherein the nucleic acid strand is about 10 to about 22 nucleotides in length, or about 15 to about 20 nucleotides in length.
165. The system of any one of claims 157-164, wherein the oligonucleotide comprises a primer binding sequence.
166. The system of any one of claims 157-165, wherein the oligonucleotide is complementary to a sequence 3’ to the region of interest in the RNA template.
167. The system of any one of claims 157-165, wherein the oligonucleotide comprises a poly- T sequence or a randomer.
168. The system of any one of claims 158-167, wherein the polymerase for performing rolling circle amplification is a Phi29 polymerase.
169. The system of any one of claims 157-168, comprising reagents for performing sequencing by ligation, sequencing by synthesis, sequencing by binding, sequencing by avidity, or a combination thereof.
170. The system of any one of claims 157-169, comprising a plurality of a sequencing primer, a plurality of detectably labeled nucleotides, and a polymerase.
171. The system of any one of claims 157-170, comprising an optical detection system configured to detect a sequence of or associated with the RNA template or a complement thereof.
172. A method of nucleic acid processing, comprising:(a) hybridizing an oligonucleotide to a ribonucleic acid (RNA) template;(b) extending the oligonucleotide using a reverse transcriptase and a plurality of free nucleotides to generate an extended oligonucleotide, wherein the plurality of free nucleotides161MOFO-358213026202412023740 lacks nucleotides of at least one of four canonical bases: adenine (A), thymine (T), guanine (G), uracil (U), and cytosine (C);(c) digesting the RNA template; and(d) using the extended oligonucleotide to generate a circularized template.
173. The method of claim 172, wherein the plurality of free nucleotides lacks nucleotides of at least two of four canonical bases: adenine (A), thymine (T), guanine (G), uracil (U), and cytosine (C).
174. The method of claim 172, wherein the plurality of free nucleotides lacks nucleotides of three of four canonical bases: adenine (A), thymine (T), guanine (G), uracil (U), and cytosine (C).
175. The method of any one of claims 172-174, further comprising using an additional plurality of free nucleotides in the extending of (b), wherein the additional plurality of free nucleotides comprises nucleotides of the canonical base lacking in the plurality of free nucleotides.
176. A method of nucleic acid processing, comprising:(a) contacting a biological sample with a nucleic acid comprising one or more catalytic domains, wherein the nucleic acid hybridizes to a ribonucleic acid (RNA) template, and cleaves the RNA template;(b) hybridizing an oligonucleotide to the RNA template cleaved by the nucleic acid;(c) extending the oligonucleotide using a reverse transcriptase and a plurality of free nucleotides to generate an extended oligonucleotide;(d) using the extended oligonucleotide to generate a circularized template.
177. The method of any one of claims 172-176, wherein the using in (d) comprises ligating the 5’ and 3’ ends of the extended oligonucleotide to generate the circularized template.
178. The method of any one of claims 172-177, further comprising performing rolling circle amplification of the circularized template to generate a rolling circle amplification product (RCP).162MOFO-358213026202412023740179. The method of claim 178, further comprising detecting a sequence of the RCP in the biological sample.
180. The method of claim 179, wherein detecting the sequence of the RCP comprises performing in situ sequencing by synthesis (SBS) in the biological sample.
181. The method of claim 179 and 180, wherein detecting the sequence of the RCP comprises imaging the biological sample to detect the RCP in situ in the biological sample.
182. The method of any one of claims 1-78, 94-144 and 172-181, wherein the circularized template comprises at least a portion of the extended oligonucleotide.
183. The method of any one of claims 1-78, 94-144 and 172-181, wherein the circularized template comprises a sequence of the extended oligonucleotide or a sequence complementary to the extended oligonucleotide.
184. The method of any one of claims 172-183, wherein the biological sample is a cell or tissue sample comprising cells or cellular components.
185. The method of any one of claims 172-183, wherein the biological sample is crosslinked and / or embedded in a matrix.163MOFO-358213026