Compositions and Methods for In Situ Array Determination

By employing base-by-base sequencing and strategic signal coding, the method addresses the issue of optical crowding in in situ detection, enabling more effective and scalable detection and decoding of analytes in situ.

JP2025516589APending Publication Date: 2025-05-3010X GENOMICS INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024566330
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-05-11
Filing Date
2023-05-11
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The plex scalability of in situ detection methods is limited by optical crowding, where large or high-intensity signals overlap with and mask other signals, while small or low-intensity signals may not reach the detection threshold, impairing signal detection and decoding quality.

Method used

The method involves base-by-base sequencing in situ in a cell or tissue sample using sequencing primers hybridized to identifier sequences, with nucleotides added in periodic steps and signals detected to generate signal code sequences, allowing for the decoding of analytes by reducing optical crowding through strategic signal coding.

Benefits of technology

This approach enhances the ability to detect and decode multiple analytes simultaneously by reducing optical crowding, improving signal detection quality, and minimizing the need for complex probe pools, thereby increasing the plex scalability of in situ assays.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025516589000001_ABST
    Figure 2025516589000001_ABST
Patent Text Reader

Abstract

Methods and compositions are described for performing in situ base-by-base sequencing in a cell or tissue sample that minimizes optical crowding. In some embodiments, a sequencing primer is hybridized to a priming site 3' to an identifier sequence (e.g., a barcode sequence) in the sample, such that the sequencing primer can be extended by a polymerase in a base-by-base manner using the identifier sequence as a template. The sample can be contacted with nucleotides in a series of periodic nucleotide incorporation or binding steps, and a signal indicating an incorporation or binding event is detected to generate a signal code sequence comprising a series of signal codes detected in successive cycles (corresponding to a signal (on signal), absence of a signal (off signal), or a combination thereof). The corresponding analyte can be detected and localized using decoding of the identifier sequence based at least in part on the signal code sequence.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Cross - Reference to Related Applications This application claims the benefit of priority of U.S. Provisional Patent Application No. 63 / 340,730, filed on May 11, 2022, the content of which is hereby incorporated by reference in its entirety.

[0002] Field The present disclosure generally relates to methods and compositions for in - situ detection of analytes in a sample.

Background Art

[0003] Background Genome, transcriptome, and proteome profiling of cell and tissue samples using microscopy imaging can resolve multiple analytes of interest simultaneously, thereby providing valuable information regarding the abundance and localization of analytes in - situ. Thus, in - situ assays are, for example, important tools for understanding the molecular basis of cell identity and developing treatments for diseases. There is a need for new and improved methods for in - situ assays. Methods and compositions addressing such needs and other needs are provided herein.

Summary of the Invention

[0004] Summary The in situ detection method's plex scalability can be limited by optical crowding. For example, large-sized features (e.g., rolling circle amplification (RCA) products, nucleic acid probes, or nucleic acid complexes) and / or signals associated with high intensity, especially when the signals arise from features in close proximity, may overlap with and / or mask other signals, while small-sized features and / or signals associated with low intensity may not reach the detection threshold. In either case, for example, the signal detection and decoding quality of identifier sequences related to target analytes in a multiplex assay can be impaired. In addition, certain existing methods require complex pools of oligonucleotide probes in addition to the cost and time for detection and decoding.

[0005] In some aspects, the present disclosure relates to methods and compositions for base-by-base sequencing in situ in a cell or tissue sample. In some embodiments, the sequencing primer is hybridized to a priming site 3' to an identifier sequence (e.g., a barcode sequence) in the sample, such that the sequencing primer can be extended by a polymerase in a base-by-base manner using the sequence of the identifier sequence as a template. The cell or tissue sample can be contacted with nucleotides in a series of periodic nucleotide incorporation or nucleotide binding steps, and a signal indicating the incorporation or binding event is detected in each cycle to generate a signal code sequence comprising a series of signal codes corresponding to the signal (on-signal), the absence of a signal (off-signal), or a combination of on and off signals detected in successive cycles. Signal code sequences for a plurality of different identifier sequences are generated at positions in the sample and can be used to decode the corresponding analytes.

[0006] In some embodiments, provided herein is a method for analyzing a biological sample, the method comprising: a) contacting the biological sample with a first probe and a second probe, wherein the biological sample is a cell or tissue sample, the biological sample contains a first analyte and a second analyte at a first position and a second position in the biological sample respectively, the first probe and the second probe bind directly or indirectly to the first analyte and the second analyte respectively, the first probe or its product contains i) a first priming site for a first sequencing primer, and ii) a first identifier sequence related to the first analyte, the second probe or its product contains i) a second priming site for a second sequencing primer, and ii) a second identifier sequence related to the second analyte; b) using the first and second sequencing primers to perform base-by-base sequencing of the first and second identifier sequences, for example, sequencing by synthesis (SBS) or sequencing by binding (SBB), thereby generating a first signal code sequence and a second signal code sequence, each containing a signal code corresponding to a signal (on-signal), absence of a signal (off-signal), or a combination thereof, detected in consecutive cycles at the first position and the second position respectively, wherein in one or more of the consecutive cycles, an on-signal is detected at the first position and an off-signal is detected at the second position; c) detecting the first and second identifier sequences in the biological sample based at least in part on the first and second signal code sequences.

[0007] In some embodiments, the first and second analytes are the same. In some embodiments, the first and second analytes are different. In any of the embodiments herein, the first and second identifier sequences may be different. In any of the embodiments herein, the first and second identifier sequences may contain an analyte sequence or its complement.

[0008] In any of the embodiments of this specification, the first and second identifier sequences may include a barcode sequence or its complement. In any of the embodiments of this specification, the barcode sequence may be assigned to the first and second analytes, respectively. In some embodiments, assigning the barcode sequence is based on a decision rule designed to minimize the maximum predicted density of the on-signal detected in each of one or more of the consecutive cycles. In some embodiments, the decision rule for assigning the first barcode sequence to the first analyte and the second barcode sequence to the second analyte includes an assignment based on expression data for the first and second analytes. In some embodiments, assigning the first barcode sequence to the first analyte and the second barcode sequence to the second analyte includes an assignment based on expression data for the first and second analytes in the clustered cell types. In some embodiments, the clustered cell types represent the distribution of cell types found in the biological sample. In some embodiments, the expression data for the first and second analytes at least partially overlap. In some embodiments, the expression data for the first and second analytes includes bulk gene expression data, bulk protein expression data, spatial gene expression data, spatial protein expression data, single cell gene expression data, single cell protein expression data, or any combination thereof.

[0009] In any of the embodiments of this specification, the method may include assigning a first barcode array to a first analyte and a second barcode array to a second analyte. In some embodiments, the nucleotides in the first barcode array detected in a particular cycle correspond to a signal code including an on-signal, and the corresponding nucleotides in the second barcode array detected in the particular cycle correspond to a signal code including an off-signal. In some embodiments, the nucleotides in the first barcode array detected in a particular cycle correspond only to the on-signal, and the corresponding nucleotides in the second barcode array detected in the particular cycle correspond only to the off-signal. In any of the embodiments of this specification, one or more pairs of corresponding nucleotides in the first and second barcode arrays detected in the same cycle may be selected to reduce optical crowding of the signals detected in the cycle.

[0010] In any of the embodiments of this specification, base-by-base sequencing may be performed by contacting a biological sample with nucleotides in successive cycles, and in each cycle, a complex is formed, the complex including: i) a first or second sequencing primer or an extension product thereof that hybridizes to the first or second priming site, respectively; ii) a polymerase; and iii) cognate nucleotides that base pair with the nucleotides in the first or second identifier sequence, and a signal (on-signal) associated with the cognate nucleotides and / or polymerase in the complex and / or the absence of a signal (off-signal) is detected at a specific location in the biological sample, and the on-signal, off-signal, or a combination thereof corresponds to the cognate nucleotides in the first or second identifier sequence and the bases in the corresponding nucleotides.

[0011] In any of the embodiments of this specification, 25% or more of the nucleotides in the first and / or second identifier sequences can be assigned to correspond to an off signal. In any of the embodiments of this specification, the first and / or second identifier sequences can be designed such that more than 25% of the nucleotides therein correspond to an off signal. In any of the embodiments of this specification, 30% or more of the nucleotides in the first and / or second identifier sequences can be assigned to correspond to an off signal. In any of the embodiments of this specification, the first and / or second identifier sequences can be designed such that more than 30% of the nucleotides therein correspond to an off signal. In any of the embodiments of this specification, the first and / or second identifier sequences can be designed such that more than 35% of the nucleotides therein correspond to an off signal. In any of the embodiments of this specification, 40% or more of the nucleotides in the first and / or second identifier sequences can be assigned to correspond to an off signal. In any of the embodiments of this specification, the first and / or second identifier sequences can be designed such that more than 45% of the nucleotides therein correspond to an off signal. In any of the embodiments of this specification, 50% or more of the nucleotides in the first and / or second identifier sequences can be assigned to correspond to an off signal. In any of the embodiments of this specification, the first and / or second identifier sequences can be designed such that more than 55% of the nucleotides therein correspond to an off signal. In any of the embodiments of this specification, 60% or more of the nucleotides in the first and / or second identifier sequences can be assigned to correspond to an off signal. In any of the embodiments of this specification, the first and / or second identifier sequences can be designed such that more than 65% of the nucleotides therein correspond to an off signal. In any of the embodiments of this specification, 70% or more of the nucleotides in the first and / or second identifier sequences can be assigned to correspond to an off signal. In any of the embodiments of this specification, 75% or more of the nucleotides in the first and / or second identifier sequences can be assigned to correspond to an off signal. In any of the embodiments of this specification, 80% or more of the nucleotides in the first and / or second identifier sequences can be assigned to correspond to an off signal.In any of the embodiments of the present specification, 85% or more of the nucleotides in the first and / or second identifier sequences can be assigned to correspond to off-signals. In any of the embodiments of the present specification, 90% or more of the nucleotides in the first and / or second identifier sequences can be assigned to correspond to off-signals. In any of the embodiments of the present specification, 95% or more of the nucleotides in the first and / or second identifier sequences can be assigned to correspond to off-signals.

[0012] In any of the embodiments of the present specification, a plurality of different identifier sequences can be detected in a biological sample, and each different identifier sequence can be detected at one or more positions in the biological sample. In some embodiments, the plurality of different identifier sequences are a plurality of unique identifier sequences, for example, each uniquely corresponding to an analyte. In any of the embodiments of the present specification, 50% or more of the different identifier sequences can each include 50% or more of the nucleotides in the identifier sequence corresponding to the off-signal. In any of the embodiments of the present specification, 80% or more of the different identifier sequences can each include 80% or more of the nucleotides in the identifier sequence corresponding to the off-signal.

[0013] In any of the embodiments of the present specification, each signal code can correspond to a signal of a first color, a signal of a second color, a signal of a third color, or the absence of a signal, and the first, second, and third colors are different. In any of the embodiments of the present specification, each signal code can correspond to a signal of a first color, a signal of a second color, a combination of signals of the first and second colors, or the absence of a signal, and the first and second colors are different. In any of the embodiments of the present specification, each signal code can correspond to a combination of a signal (on-signal) and / or the absence of a signal (off-signal), and the combination of on and / or off-signals is detected in two or more imaging steps.

[0014] In any of the embodiments herein, it may further include detecting first and second analytes in a biological sample based on detecting first and second identifier sequences. In any of the embodiments herein, the first identifier sequence may be the sequence of the first analyte or its complement. In any of the embodiments herein, the second identifier sequence may be the sequence of the second analyte or its complement. In any of the embodiments herein, the first identifier sequence may be the first barcode sequence or its complement that corresponds to, is related to, and / or identifies the first analyte. In any of the embodiments herein, the second identifier sequence may be the second barcode sequence or its complement that corresponds to, is related to, and / or identifies the second analyte.

[0015] In any of the embodiments herein, the first barcode sequence may be used to identify the first analyte. In any of the embodiments herein, the second barcode sequence may be used to identify the second analyte.

[0016] In any of the embodiments herein, the first probe may be provided among a first plurality of probes that bind directly or indirectly to the first analyte. In some embodiments, the first plurality of probes collectively include a first combination of barcode sequences that identify the first analyte. In any of the embodiments herein, the second probe may be provided among a second plurality of probes that bind directly or indirectly to the second analyte. In some embodiments, the second plurality of probes collectively include a second combination of barcode sequences that identify the second analyte.

[0017] In any of the embodiments herein, the first and second analytes may include nucleic acid sequences. In any of the embodiments herein, the first identifier sequence may include the sequence of the first analyte or its complement, and the second identifier sequence may include the sequence of the second analyte or its complement.

[0018] In any of the embodiments of this specification, base-by-base sequencing may involve using a fluorescently labeled polymerase and one or more unlabeled nucleotides. In any of the embodiments of this specification, base-by-base sequencing may involve using a polymerase-nucleotide conjugate that includes a fluorescently labeled polymerase linked to an unlabeled nucleotide moiety. In any of the embodiments of this specification, base-by-base sequencing may involve using a multivalent polymer-nucleotide conjugate that includes a polymer core, a plurality of nucleotide moieties, and one or more fluorescent labels.

[0019] In some embodiments, during base-by-base sequencing, cognate nucleotides are not incorporated by the polymerase into the first or second sequencing primer or their extension products. In some embodiments, the incorporation of cognate nucleotides into the first or second sequencing primer or their extension products by the polymerase is reduced or inhibited.

[0020] In any of the embodiments of this specification, base-by-base sequencing may involve contacting a biological sample with a nucleotide mix that includes fluorescently labeled nucleotides and unlabeled nucleotides. In some embodiments, during base-by-base sequencing, cognate nucleotides are incorporated by the polymerase into the first or second sequencing primer or their extension products, and the cognate nucleotides are either fluorescently labeled or unlabeled.

[0021] In any of the embodiments of this specification, base-by-base sequencing may include contacting a biological sample with a first nucleotide mix, wherein the nucleotide containing the first base is not detectably labeled while the nucleotides containing bases other than the first base are each labeled with one or more detectable labels, and contacting the biological sample with a subsequent nucleotide mix, wherein the nucleotide containing the subsequent base is not detectably labeled while the nucleotides containing bases other than the subsequent base are each labeled with one or more detectable labels, and the subsequent base is the same as the first base, and optionally, the first base and the subsequent base are A, T, C, or G.

[0022] In any of the embodiments of this specification, base-by-base sequencing may include contacting a biological sample with a first nucleotide mix, wherein the nucleotide containing the first base is not detectably labeled while the nucleotides containing bases other than the first base are each labeled with one or more detectable labels, and contacting the biological sample with a subsequent nucleotide mix, wherein the nucleotide containing the subsequent base is not detectably labeled while the nucleotides containing bases other than the subsequent base are each labeled with one or more detectable labels, and the subsequent base is different from the first base.

[0023] In any of the embodiments of this specification, a biological sample can contact two or more of the following nucleotide mixes in consecutive cycles in any order: nucleotide mix 1 in which nucleotides containing G are not detectably labeled while nucleotides containing A, C, or T are detectably labeled; nucleotide mix 2 in which nucleotides containing T are not detectably labeled while nucleotides containing A, C, or G are detectably labeled; nucleotide mix 3 in which nucleotides containing C are not detectably labeled while nucleotides containing A, G, or T are detectably labeled; and nucleotide mix 4 in which nucleotides containing A are not detectably labeled while nucleotides containing G, C, or T are detectably labeled.

[0024] In any of the embodiments of this specification, each nucleotide mix, independent of one another, can contact the biological sample in one or more cycles, and the cycles can be consecutive or non - consecutive. In any of the embodiments of this specification, independently of one another, each nucleotide mix can include: detectably labeled nucleotides having three different - colored fluorescent labels, one color for each of the three bases, for example, red for A, blue for G, green for T, and no detectable label for C; detectably labeled nucleotides having two different - colored fluorescent labels, one color for each of two of the three bases, and nucleotides containing the remaining base are labeled with both colors, for example, red and green for A, red for G, green for T, and no detectable label for C; or detectably labeled nucleotides having the same - colored fluorescent label, the fluorescent label on the nucleotide containing one of the three bases is configured to be cleaved, and the nucleotide containing another one of the three bases is configured to be labeled with a fluorescent label.

[0025] In any of the embodiments of this specification, a biological sample may contact two or more of the following nucleotide mixes in consecutive cycles in any order: nucleotide mix 1 in which nucleotides containing G or A are not detectably labeled while nucleotides containing C or T are detectably labeled; nucleotide mix 2 in which nucleotides containing G or T are not detectably labeled while nucleotides containing C or A are detectably labeled; nucleotide mix 3 in which nucleotides containing G or C are not detectably labeled while nucleotides containing A or T are detectably labeled; nucleotide mix 4 in which nucleotides containing C or A are not detectably labeled while nucleotides containing G or T are detectably labeled; nucleotide mix 5 in which nucleotides containing C or T are not detectably labeled while nucleotides containing G or A are detectably labeled; and nucleotide mix 6 in which nucleotides containing A or T are not detectably labeled while nucleotides containing G or C are detectably labeled.

[0026] In any of the embodiments of this specification, the first priming site and the second priming site may be different. In any of the embodiments of this specification, the method includes: b1) hybridizing a first sequencing primer to the first priming site and performing base-by-base sequencing to generate an extension product of the first sequencing primer and a first signal code sequence; b2) removing, cleaving, or blocking the extension product of the first sequencing primer in b1); and b3) hybridizing a second sequencing primer to the second priming site and performing base-by-base sequencing (e.g., SBS or SBB) to generate an extension product of the second sequencing primer and a second signal code sequence.

[0027] In any of the embodiments of this specification, the probes or their products for the first plurality of analytes may share a common first priming site, and the probes or their products for the second plurality of analytes may share a common second priming site. In any of the embodiments of this specification, the second plurality of analytes may include two or more different analytes that are different from two or more different analytes of the first plurality of analytes.

[0028] In any of the embodiments of this specification, a biological sample can be contacted with a plurality of probes configured to directly or indirectly bind to different analytes, and each probe or its product may include a combination of different priming sites. In any of the embodiments of this specification, the first probe or its product may include a first combination of different priming sites that includes a first priming site. In any of the embodiments of this specification, the second probe or its product may include a second combination of different priming sites that includes a second priming site.

[0029] In any of the embodiments of this specification, a biological sample can be contacted with a third probe that directly or indirectly binds to a third analyte. In any of the embodiments of this specification, the third probe or its product may include a third combination of different priming sites that includes a first priming site, a second priming site, and / or a third priming site.

[0030] In any of the embodiments of this specification, any two or more of the first combination, the second combination, and the third combination may share one or more common priming sites.

[0031] In any of the embodiments of this specification, the method includes: b’) contacting a biological sample with a first sequencing primer for base-by-base sequencing, thereby hybridizing the first sequencing primer to a first priming site in a first probe or its product and in one or more other probes or their products, and generating an extension product of the first sequencing primer; b’’) removing, cleaving, or blocking the extension product of the first sequencing primer in b’); and b’’’) contacting the biological sample with a second sequencing primer for base-by-base sequencing, thereby hybridizing the second sequencing primer to a second priming site in a second probe or its product and in one or more other probes or their products, and generating an extension product of the second sequencing primer.

[0032] In any of the embodiments of this specification, the base-by-base sequencing in b’) can be performed by contacting the biological sample with nucleotides in successive cycles, detecting a signal associated with nucleotide incorporation or binding for each successive cycle, and generating a signal code sequence for a first plurality of analytes.

[0033] In any of the embodiments of this specification, the base-by-base sequencing in b’’’) can be performed by contacting the biological sample with nucleotides in successive cycles, detecting a signal associated with nucleotide incorporation or binding for each successive cycle, and generating a signal code sequence for a second plurality of analytes. In some embodiments, the first plurality of analytes and the second plurality of analytes include one or more common analytes. In some embodiments, the first plurality of analytes and the second plurality of analytes do not include a common analyte.

[0034] In any of the embodiments of this specification, each analyte can independently be a nucleic acid analyte or a non-nucleic acid analyte. In any of the embodiments of this specification, each probe can independently be i) a primary probe that binds directly to its corresponding analyte, or ii) a probe that binds directly or indirectly to the primary probe. In some embodiments, the primary probe and the probe that binds directly or indirectly to the primary probe are independently selected from the group consisting of: a probe comprising a 3' or 5' overhang, optionally wherein the 3' or 5' overhang comprises one or more barcode sequences; a probe comprising a 3' overhang and a 5' overhang, optionally wherein the 3' overhang and the 5' overhang each independently comprise one or more barcode sequences; a circular probe; a circularizable probe or probe set; a probe or probe set comprising a split hybridization region configured to hybridize to a splint, optionally wherein the split hybridization region comprises one or more barcode sequences; and combinations thereof.

[0035] In any of the embodiments of this specification, the product of each probe can include a rolling circle amplification (RCA) product generated in situ in a biological sample. In any of the embodiments of this specification, base-by-base sequencing can be performed in situ in a biological sample.

[0036] In some aspects, a method of analyzing a biological sample is disclosed herein, the method comprising: a) contacting a biological sample with a first sequencing primer and a second sequencing primer, wherein the biological sample is a cell or tissue sample, the biological sample contains a first nucleic acid and a second nucleic acid at a first position and a second position in the biological sample respectively, the first nucleic acid contains i) a first priming site complementary to the first sequencing primer, and ii) a first identifier sequence, and the second nucleic acid contains i) a second priming site complementary to the second sequencing primer, and ii) a second identifier sequence; b) performing base-by-base sequencing of the first identifier sequence using the first sequencing primer hybridized to the first priming site to generate a first signal code sequence containing signal codes detected in consecutive cycles at the first position, wherein the signal code corresponds to a signal (on signal), the absence of a signal (off signal), or a combination thereof; c) then, performing base-by-base sequencing of the second identifier sequence using the second sequencing primer hybridized to the second priming site to generate a second signal code sequence containing signal codes detected in additional consecutive cycles at the second position, wherein at one or both of the first and second positions, an off signal is detected in at least one or more of the consecutive cycles in b) and the additional consecutive cycles in c); and d) detecting the first and second identifier sequences in the biological sample at the first and second positions respectively, based at least in part on the first and second signal code sequences.

[0037] In some embodiments, a method of analyzing a biological sample is disclosed herein, the method comprising: a) contacting the biological sample with a first probe and a second probe, wherein the biological sample is a cell or tissue sample, the biological sample contains a first analyte and a second analyte at a first position and a second position in the biological sample respectively, the first probe and the second probe bind directly or indirectly to the first analyte and the second analyte respectively, the first probe or its product contains i) a first priming site for a first sequencing primer, and ii) a first identifier sequence associated with the first analyte, and the second probe or its product contains i) a second priming site for a second sequencing primer, and ii) a second identifier sequence associated with the second analyte; b) using the first sequencing primer to perform base-by-base sequencing (e.g., using at least one dark base) of the first identifier sequence to generate a first signal code sequence containing the signal code detected in consecutive cycles at the first position, wherein the signal code corresponds to a signal (on-signal), the absence of a signal (off-signal), or a combination thereof; c) then using the second sequencing primer to perform base-by-base sequencing (e.g., using at least one dark base) of the second identifier sequence to generate a second signal code sequence containing the signal code detected in additional consecutive cycles at the second position, wherein an off-signal is detected in at least one or more of the consecutive cycles in b) and the additional consecutive cycles in c); d) detecting the first and second identifier sequences in the biological sample at the first and second positions respectively, based at least in part on the first and second signal code sequences. In some embodiments, an off-signal is detected at one or both of the first position and the second position in at least one or more of the consecutive cycles in b) and the additional consecutive cycles in c).

[0038] In some embodiments, in at least one or more of the successive cycles using the first sequencing primer and the additional successive cycles using the second sequencing primer, the on-signal is not detected at both the first and second positions in the same base-by-base sequencing cycle. In some embodiments, in at least one or more of the successive cycles using the first sequencing primer and in at least one or more of the additional successive cycles using the second sequencing primer, the on-signal is not detected at both the first and second positions in the same base-by-base sequencing cycle. In some embodiments, in two or more of the successive cycles using the first sequencing primer and in two or more of the additional successive cycles using the second sequencing primer, the on-signal is not detected at both the first and second positions in the same base-by-base sequencing cycle. In some embodiments, in three or more of the successive cycles using the first sequencing primer and in three or more of the additional successive cycles using the second sequencing primer, the on-signal is not detected at both the first and second positions in the same base-by-base sequencing cycle. When the first and second positions are in close proximity to each other, the methods disclosed herein reduce optical crowding during base-by-base in situ sequencing in a biological sample.

[0039] In any of the embodiments of the present specification, an off-signal can be generated by performing base-by-base sequencing using one, two, or three nucleotides (e.g., any one, two, or three of A, T / U, C, and G) that are not detectably labeled while other nucleotides are detectably labeled. Different nucleotides or combinations of different nucleotides can be used in a particular cycle as compared to one or more other cycles. In some embodiments, the nucleotides that are not detectably labeled can be natural nucleotides or derivatives thereof that do not contain an exogenous label (e.g., a fluorescent dye or any other label) or a chemical modification, e.g., naturally occurring nucleotides. Nucleotides that are detectably labeled can be conjugated to a detectable label (e.g., a fluorescent dye or any other label) covalently (e.g., via a bond or a linker) or non-covalently (e.g., via a binding pair such as biotin or a derivative or analog thereof and streptavidin or a derivative or analog thereof). In other examples, nucleotides that are detectably labeled can be conjugated to a moiety (e.g., an antigen-binding molecule such as an antigen or an antibody) that enables the detection of nucleotides using an agent that specifically binds to the moiety and generates a detectable signal, and the detectable signal generated by the agent can be used to detect the incorporation of the nucleotides. The moiety can be conjugated to the nucleotide via a cleavable linker.

[0040] In any of the embodiments of the present specification, an off-signal can be generated by performing base-by-base sequencing by excluding one, two, or three nucleotides (e.g., any one, two, or three of A, T / U, C, and G) from a nucleotide mix that contacts a sample in a particular sequencing cycle while the nucleotide mix contains only other nucleotides that are detectably labeled. Different nucleotides or combinations of different nucleotides can be used (or excluded from the nucleotide mix) in a particular cycle as compared to one or more other cycles.

[0041] In any of the embodiments of the present specification, an off-signal can be generated by performing base-by-base sequencing by detecting one, two, or three nucleotides (e.g., any one, two, or three of A, T / U, C, and G) that contact the sample in a specific sequencing cycle while detecting only other nucleotides. Different nucleotides or combinations of different nucleotides can be detected (or not detected) in a specific cycle as compared to one or more other cycles. For example, when nucleotides labeled with an antigen or an antibody are used in a sequencing cycle, not detecting the signal associated with the nucleotide can be achieved by excluding the corresponding antibody or antigen labeled with a detectable label from the signal detection in that sequencing cycle.

[0042] In any of the embodiments of the present specification, the number of "on" cycles / bits must be such that each codeword (e.g., corresponding to a signal code sequence) has and can be selected such that the codeword can be designed accordingly. In any of the embodiments of the present specification, the number of "on" cycles / bits can be from about 3 to about 8 for the codeword. In any of the embodiments of the present specification, the number of "on" cycles / bits can be 4, 5, 6, or 7 for the codeword. In any of the embodiments of the present specification, the constraints on the codeword can be used to facilitate the identification of the correct identifier sequence (e.g., a barcode sequence such as that corresponding to an RCP). In any of the embodiments of the present specification, each codeword can have at least X (X is 3 or 4) colored "on" bits to distinguish it from a background fluorescence source that tends to remain a single color.

[0043] In any of the embodiments of this specification, the first and second sequencing primers can contact the sample either simultaneously or sequentially in any order. In any of the embodiments of this specification, the first and second sequencing primers can contact the sample either simultaneously or sequentially in any order. In any of the embodiments of this specification, the first and second identifier sequences can be in DNA (e.g., genomic DNA), RNA (e.g., mRNA), DNA (e.g., genomic DNA) or RNA (e.g., mRNA), directly or indirectly bound probes, or in products such as DNA (e.g., genomic DNA), RNA (e.g., mRNA), or rolling circle amplification products of the probes.

[0044] In any of the embodiments of this specification, the first and second identifier sequences can be the same. In any of the embodiments of this specification, the first and second identifier sequences can be different.

[0045] In any of the embodiments of this specification, the first and second identifier sequences may include barcode sequences or their complements respectively assigned to the first and second analytes. In some embodiments, based on a decision rule designed to minimize the maximum predicted density of on-signals detected in each of one or more of the consecutive cycles, the first priming site and / or the first barcode sequence may be assigned to the first analyte, and the second priming site and / or the second barcode sequence may be assigned to the second analyte. In some embodiments, the assignment may include an assignment based on expression data for the first and second analytes. In some embodiments, the assignment may include an assignment based on expression data for the first and second analytes in clustered cell types. In some embodiments, the clustered cell types may represent the distribution of cell types found in a biological sample. In some embodiments, the expression data for the first and second analytes may at least partially overlap. In some embodiments, the expression data for the first and second analytes may include bulk gene expression data, bulk protein expression data, spatial gene expression data, spatial protein expression data, single-cell gene expression data, single-cell protein expression data, or any combination thereof.

[0046] In any of the embodiments of this specification, the nucleotides in the first barcode sequence or the second barcode sequence detected in a particular cycle may correspond to a signal code including an on-signal. In any of the embodiments of this specification, the nucleotides in the first barcode sequence or the second barcode sequence detected in a particular cycle may correspond to a signal code including an off-signal. In some embodiments, the nucleotides in the first barcode sequence or the second barcode sequence detected in a particular cycle may correspond only to an on-signal. In some embodiments, the nucleotides in the first barcode sequence or the second barcode sequence detected in a particular cycle may correspond only to an off-signal.

[0047] Also disclosed herein is a method for decoding identifier sequences in a biological sample while minimizing optical crowding. The method comprises: a) contacting a biological sample with a first probe and a second probe, wherein the biological sample is a cell or tissue sample, the biological sample contains a first analyte and a second analyte at a first position and a second position in the biological sample, respectively, the first probe and the second probe are directly or indirectly bound to the first analyte and the second analyte, respectively, the first probe or its product contains i) a first priming site for a first sequencing primer and ii) a first identifier sequence associated with the first analyte, and the second probe or its product contains i) a second priming site for a second sequencing primer and ii) a second identifier sequence associated with the second analyte; b) performing base-by-base sequencing of the first and second identifier sequences using the first and second sequencing primers, thereby generating a first signal code sequence and a second signal code sequence, each containing a signal code corresponding to a signal (on-signal), absence of a signal (off-signal), or a combination thereof detected in successive sequencing cycles at the first position and the second position, respectively, and wherein the base-by-base sequencing comprises contacting the biological sample in each successive cycle with a mixture of nucleotides containing a polymerase and at least one nucleotide that is not detectably labeled; and c) detecting the first and second identifier sequences in the biological sample based at least in part on the first and second signal code sequences.

[0048] In some embodiments, the first and second analytes can be the same or different. In some embodiments, the first and second identifier sequences are different.

[0049] In any of the embodiments of this specification, the first and second identifier sequences may include the analyte sequence or its complement. In any of the embodiments of this specification, the first and second identifier sequences include barcode sequences or their complements respectively assigned to the first and second analytes. In any of the embodiments of this specification, the method may include assigning a first barcode sequence to a first analyte and a second barcode sequence to a second analyte. In some embodiments, assigning the barcode sequences may be based on a decision rule designed to minimize the maximum predicted density of the on-signal detected in each of one or more of the consecutive cycles. In some embodiments, assigning a first barcode sequence to a first analyte and a second barcode sequence to a second analyte may be based on the expression data for the first and second analytes. In some embodiments, assigning a first barcode sequence to a first analyte and a second barcode sequence to a second analyte may include an assignment based on the expression data for the first and second analytes in the clustered cell types. In some embodiments, the clustered cell types represent the distribution of cell types found in the biological sample. In some embodiments, the expression data for the first and second analytes at least partially overlap. In some embodiments, the expression data for the first and second analytes includes bulk gene expression data, bulk protein expression data, spatial gene expression data, spatial protein expression data, single-cell gene expression data, single-cell protein expression data, or any combination thereof.

[0050] In any of the embodiments of the present specification, the nucleotide mixture may include at least two nucleotides that are not detectably labeled. In some embodiments, the biological sample is contacted with two or more of the following nucleotide mixtures in successive sequencing cycles in any order: nucleotide mixture 1 in which the nucleotide containing G is not detectably labeled while the nucleotides containing A, C, or T are detectably labeled; nucleotide mixture 2 in which the nucleotide containing T is not detectably labeled while the nucleotides containing A, C, or G are detectably labeled; nucleotide mixture 3 in which the nucleotide containing C is not detectably labeled while the nucleotides containing A, G, or T are detectably labeled; and nucleotide mixture 4 in which the nucleotide containing A is not detectably labeled while the nucleotides containing G, C, or T are detectably labeled. In some embodiments, independently of each other, each nucleotide mixture is contacted with the biological sample in one or more cycles, and the cycles are consecutive or non-consecutive. In some embodiments, independently of each other, in each nucleotide mixture, the detectably labeled nucleotides include: i) three different-colored fluorescent labels, one color for each of the three bases; ii) two different-colored fluorescent labels, one color for each of two of the three bases, and the nucleotides containing the remaining base are labeled with both colors; or iii) fluorescent labels of the same color, where the fluorescent label on the nucleotide containing one of the three bases is configured to be cleaved, and the nucleotide containing another one of the three bases is configured to be labeled with a fluorescent label.

[0051] In some embodiments, the biological sample can be contacted with two or more of the following nucleotide mixes in successive cycles in any order: nucleotide mix 1 in which nucleotides containing G or A are not detectably labeled while nucleotides containing C or T are detectably labeled; nucleotide mix 2 in which nucleotides containing G or T are not detectably labeled while nucleotides containing C or A are detectably labeled; nucleotide mix 3 in which nucleotides containing G or C are not detectably labeled while nucleotides containing A or T are detectably labeled; nucleotide mix 4 in which nucleotides containing C or A are not detectably labeled while nucleotides containing G or T are detectably labeled; nucleotide mix 5 in which nucleotides containing C or T are not detectably labeled while nucleotides containing G or A are detectably labeled; and nucleotide mix 6 in which nucleotides containing A or T are not detectably labeled while nucleotides containing G or C are detectably labeled.

Brief Description of the Drawings

[0052] The drawings illustrate certain features and advantages of the present disclosure. These embodiments are not intended to limit the scope of the appended claims in any way.

[0053]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

DETAILED DESCRIPTION OF THE INVENTION

[0054] DETAILED DESCRIPTION All publications, including patent documents, scientific papers, and databases, referred to in this application are hereby incorporated by reference in their entirety for all purposes to the same extent as if each individual publication were incorporated by reference separately. If the definitions set forth herein conflict with or are contrary to the definitions set forth in patents, applications, published applications, and other publications incorporated by reference herein, the definitions set forth herein shall control over the definitions incorporated by reference herein.

[0055] The section headings used herein are for organizational purposes only and are not to be construed as limiting the subject matter described.

[0056] I. SUMMARY The plex scalability of in situ detection methods can be limited by optical crowding. For example, a signal having a large size and / or high intensity can overlap with and / or mask other signals, especially when the signals are close, while a signal having a small size and / or low intensity may not reach the detection threshold. In either case, the quality of signal detection and decoding can be impaired. In addition, certain existing methods require complex pools of oligonucleotide probes, in addition to the cost and time for detection and decoding.

[0057] For example, in some embodiments, single-base sequencing chemistries (e.g., sequencing by synthesis (SBS), sequencing by binding (SBB), etc.) provide fast reaction times and can avoid the need for complex pools of oligonucleotide probes and sequential hybridizations of probe pools for decoding identifier sequences associated with target analytes in, for example, multiplexed assays. However, these per-base sequencing methods can have limited plex scalability for in situ analysis due to optical crowding during signal detection and decoding steps. In particular, it can be difficult to detect a large number of analytes (e.g., genes and / or their transcripts) in a cell or tissue sample in parallel in situ.

[0058] Signal crowding can occur when a large number of signals are detected. Using a conventional base-by-base sequencing method in which each nucleotide in a sequencing cycle generates an optical signal (e.g., a spot in an image obtained using fluorescence microscopy), the sample can become congested with adjacent (e.g., somewhat overlapping) signal spots, which can make it difficult to resolve individual spots. Thus, spatial overlap can limit the ability to multiplex in microscopy-based nucleic acid sequencing assays. In some embodiments, signal crowding can occur if one or more of the detected signals are significantly stronger than other signals (e.g., have a significantly greater amplitude or intensity). For example, in the same microscopic field, one or more fluorescent spots may be significantly stronger than other spots including adjacent spots. If too many signal spots are present in the sample, or if the amplitude of a signal is significantly greater than that of another signal, it can be difficult to accurately and reliably detect all of the signals in the same field of view and / or the same detection channel (e.g., the same fluorescence channel). In some cases, signal crowding can mask and / or omit weaker signals (e.g., lower amplitude) or overlapping signals, which can ultimately lead to loss of information from the analyte in the sample. In such situations, the effective dynamic range of detection can be reduced.

[0059] In some embodiments, methods and compositions are provided herein that can be used to prevent and / or address issues associated with optical crowding during in situ base-by-base sequencing.

[0060] In some embodiments, methods for analyzing a cell or tissue sample are provided herein, where identifier sequences for various analytes (e.g., the sequence of a nucleic acid analyte, or a barcode sequence in an analyte-targeting probe) are sequenced in situ, thereby detecting the corresponding analyte at one or more positions in the cell or tissue sample. In some embodiments, to detect identifier sequences for a plurality of analytes sequenced in situ, signals corresponding to at least two different identifier sequences are shifted between cycles. In some embodiments, the identifier sequences are decoded in nucleotide-by-nucleotide sequencing cycles using SBS or SBB or any other nucleotide-by-nucleotide sequencing chemistry (e.g., sequencing by affinity). In some embodiments, for many (e.g., most) of the nucleotide-by-nucleotide sequencing cycles, only signals associated with a limited number of analytes are detected in a particular cycle, while signals associated with many (e.g., most) other analytes are dark in that particular cycle. For example, an analyte may be dark in a particular cycle if the incorporation (e.g., in SBS) or binding (e.g., in SBB) of a cognate nucleotide that base pairs with a nucleotide in the identifier sequence for that analyte does not produce a detectable signal. In some embodiments, the cognate nucleotides contain bases and are not detectably labeled (e.g., produce an off-signal), while nucleotides containing one or more other different bases are detectably labeled (e.g., each produces an on-signal).

[0061] In some embodiments, using the methods disclosed herein, an on-signal is detected at a first position in a cell or tissue sample and an off-signal is detected at a second position in the cell or tissue sample. In some embodiments, the first and second positions do not overlap. In some embodiments, the first and second positions at least partially overlap. In some embodiments, the identifier sequences are such that only a subset of the nucleotides in the identifier sequences that are detected in the same base-by-base sequencing cycle produce a detectable signal, thereby limiting optical crowding of the signals detected in that cycle.

[0062] Methods are provided herein that involve the use of one or more polynucleotides (e.g., circularizable probes such as padlock probes) for analyzing one or more analytes (e.g., one or more messenger RNAs) present in a biological sample such as a cell or tissue sample. Also provided are probes, sets of probes, compositions, kits, systems, and devices for use according to the provided methods. In some aspects, the provided methods and systems can be applied to sequence, detect, image, quantify, and / or determine the presence of one or more analytes (e.g., target nucleic acids) or portions thereof. In some aspects, the provided methods and systems can be applied to simultaneously reduce optical crowding of the signals generated in a sample while sequencing multiple analytes, or identifier sequences associated therewith, in parallel in situ.

[0063] In some embodiments, an exemplary workflow for analyzing a biological sample includes contacting the biological sample with a first probe and a second probe, where the biological sample is a cell or tissue sample, the biological sample includes a first analyte and a second analyte at a first position and a second position in the biological sample, and the first probe and the second probe bind directly or indirectly to the first analyte and the second analyte, respectively. In some embodiments, the first probe and the second probe are subjected to ligation, circularization, and amplification to form a first product and a second product. In some aspects, the amplification is rolling circle amplification, and the products formed are a first rolling circle amplification product (RCP) and a second rolling circle amplification product. In some embodiments, the product of each probe is a rolling circle amplification (RCA) product generated in situ in the biological sample. In some embodiments, the first probe or its product includes i) a first priming site for a first sequencing primer, and ii) a first identifier sequence associated with the first analyte, and the second probe or its product includes i) a second priming site for a second sequencing primer, and ii) a second identifier sequence associated with the second analyte. In some embodiments, the first identifier sequence is a first barcode sequence or its complement, and the second identifier sequence is a second barcode sequence or its complement.

[0064] In some embodiments, an exemplary workflow for analyzing a biological sample includes performing base-by-base sequencing of first and second identifier sequences using first and second sequencing primers, and a periodic series of nucleotide incorporation or nucleotide binding steps, thereby generating a first signal code sequence and a second signal code sequence, each including a series of signal codes corresponding respectively to a signal (on-signal), absence of a signal (off-signal), or a combination thereof, detected in successive cycles at a first position and a second position, respectively, wherein in one or more of the successive cycles, the on-signal is detected at the first position and the off-signal is detected at the second position. In some embodiments, the method includes detecting the first and second identifier sequences in the biological sample based on the first and second signal code sequences. In some embodiments, the method includes detecting an amplification product or a portion thereof (e.g., an RCA product) based on the first and second signal code sequences. In some aspects, the first identifier sequence (e.g., a barcode sequence) identifies a first analyte and / or the second barcode sequence identifies a second analyte.

[0065] In some embodiments, an exemplary workflow for analyzing a biological sample comprises hybridizing a first sequencing primer to a first priming site and performing base-by-base sequencing to generate an extension product of the first sequencing primer and a first signal code sequence; removing (e.g., stripping), cleaving, or blocking the extension product of the first sequencing primer so as to prevent it from interfering with the base-by-base sequencing (e.g., by using terminator nucleotides in the last cycle); hybridizing a second sequencing primer to a second priming site and performing base-by-base sequencing to generate an extension product of the second sequencing primer and a second signal code sequence. In some embodiments, the first priming site and the second priming site are different. In some embodiments, probes or amplification products thereof for a first plurality of analytes share a common first priming site, and probes or amplification products thereof for a second plurality of analytes share a common second priming site. The second plurality of analytes may include two or more different analytes that are different from two or more different analytes of the first plurality of analytes.

[0066] In some embodiments, an exemplary workflow for analyzing a biological sample includes contacting the biological sample with a plurality of probes configured to each directly or indirectly bind to a different analyte. In some embodiments, each probe, or its rolling circle amplification product, includes a combination of two or more different priming sites. In some embodiments, the first probe or its product includes a first combination of two or more different priming sites including a first priming site, and / or the second probe or its product includes a second combination of two or more different priming sites including a second priming site. In some embodiments, the first combination of two or more different priming sites and the second combination of two or more different priming sites may share one or more common priming sites. In some embodiments, an exemplary workflow for analyzing a biological sample includes contacting the biological sample with a first sequencing primer and performing base-by-base sequencing using a periodic series of nucleotide incorporations or bindings, respectively, thereby generating an extension product of the first sequencing primer and a detectable signal coding sequence for a first plurality of analytes, removing, cleaving, or blocking the extension product of the first sequencing primer, contacting the biological sample with a second sequencing primer, and performing base-by-base sequencing using a periodic series of nucleotide incorporations or bindings, respectively, thereby generating an extension product of the second sequencing primer and a detectable signal coding sequence for a second plurality of analytes different from the first plurality of analytes.

[0067] In some embodiments, an exemplary workflow for analyzing a biological sample includes contacting the biological sample with a first probe and a second probe, where the biological sample is a cell or tissue sample, the biological sample includes a first analyte and a second analyte at a first position and a second position in the biological sample, the first probe is provided among a first plurality of probes that bind directly or indirectly to the first analyte, the second probe is provided among a second plurality of probes that bind directly or indirectly to the second analyte, the first plurality of probes collectively include a first combination of barcode sequences, and the second plurality of probes collectively include a second combination of barcode sequences. In some embodiments, the first and second pluralities of probes include first and second combinations of priming sites for the binding of a plurality of first and second sequencing primers. In some embodiments, the exemplary method uses a plurality of first and second sequencing primers and a periodic series of nucleotide incorporation or nucleotide binding steps, respectively, to perform base-by-base sequencing of the first and second combinations of barcode sequences, thereby generating a first signal code sequence and a second signal code sequence, each including a series of signal codes corresponding to a signal (on-signal), absence of a signal (off-signal), or a combination thereof, respectively detected in consecutive cycles at the first position and the second position, wherein in one or more of the consecutive cycles, the on-signal is detected at the first position and the off-signal is detected at the second position. In some embodiments, the exemplary method includes detecting the first and second combinations of barcode sequences in the biological sample based on the first and second signal code sequences. In some embodiments, the first combination of barcode sequences identifies the first analyte and the second combination of barcode sequences identifies the second analyte.

[0068] II. Identifier Sequence In some embodiments, methods and compositions for analyzing multiple analytes in a sample are provided herein by detecting identifier sequences for the analytes in the sample. In some embodiments, the identifier sequences are present in or derived from analytes in the sample, such as DNA or RNA analytes. For example, the identifier sequences can be part of a DNA or RNA analyte sequence. In some embodiments, the identifier sequences can be an analyte sequence (e.g., an arm of a padlock probe, or a gap filling sequence) or its complement. In some embodiments, an analyte can include two or more different identifier sequences. In some embodiments, an analyte can include two or more copies of the same identifier sequence. In some embodiments, different identifier sequences, and / or copies of the same identifier sequence, can be directly linked by phosphodiester bonds. In some embodiments, different identifier sequences, and / or copies of the same identifier sequence, can be separated from each other by one or more nucleotide residues. In some embodiments, different identifier sequences, and / or copies of the same identifier sequence, can be partially overlapping.

[0069] In some embodiments, the identifier array is present in a labeling agent that includes the identifier array, such as a nucleic acid probe, or an antibody conjugated to a reporter oligonucleotide that includes the identifier array. In some embodiments, the labeling agent can include a binding moiety that directly or indirectly interacts (e.g., binds and / or reacts) with an analyte (e.g., an endogenous analyte in a sample). In some embodiments, the labeling agent can include a reporter oligonucleotide that indicates an analyte or a portion thereof that interacts with the binding moiety. For example, the reporter oligonucleotide can include (as an identifier array) a barcode sequence that enables identification of the binding moiety and the corresponding analyte. In some cases, the sample contacted by the labeling agent can further contact a nucleic acid probe that hybridizes to the reporter oligonucleotide of the labeling agent. In some embodiments, the labeling agent includes one or more barcode sequences, such as a barcode sequence corresponding to and / or associated with an analyte binding moiety. In some embodiments, the barcode sequence is associated with or otherwise identifies the analyte binding moiety. In some embodiments, by identifying the associated barcode sequence to identify the analyte binding moiety, the analyte to which the analyte binding moiety binds can be identified. In some embodiments, the barcode sequence can be a nucleic acid sequence of a given length and / or a sequence that is associated with, corresponds to, and / or identifies the analyte binding moiety. The barcode sequence can generally include any of the various aspects of the barcode sequences described herein, e.g., in Section II-B.

[0070] In some embodiments, the identifier array is present in the product of a DNA or RNA analyte in a sample. For example, the identifier array can be present in amplification products such as hybridization products, ligation products, extension products (e.g., by DNA or RNA polymerase), replication products, transcription / reverse transcription products, and / or rolling circle amplification (RCA) products of a DNA or RNA analyte in a sample.

[0071] In some embodiments, the identifier sequences are present in the products of the labeling agents. In some embodiments, the identifier sequences are present in the products of the nucleic acid probes. In some embodiments, the identifier sequences are present in the products of the reporter oligonucleotides of the labeling agents. For example, the identifier sequences can be present in amplification products such as hybridization products, ligation products, extension products (e.g., by DNA or RNA polymerase), replication products, transcription / reverse transcription products, and / or rolling circle amplification (RCA) products of nucleic acid probes or reporter oligonucleotides in a sample.

[0072] In some embodiments, the products of the DNA or RNA analytes, nucleic acid probes, and / or reporter oligonucleotides in a sample can contain two or more different identifier sequences. In some embodiments, the products can contain two or more copies of the same identifier sequence. In some embodiments, the different identifier sequences, and / or copies of the same identifier sequence, can be in the same molecule or different molecules (e.g., molecules that form a complex such as a branched structure via hybridization). In some embodiments, the different identifier sequences, and / or copies of the same identifier sequence, can be directly linked by phosphodiester bonds. In some embodiments, the different identifier sequences, and / or copies of the same identifier sequence, can be separated from each other by one or more nucleotide residues. In some embodiments, the different identifier sequences, and / or copies of the same identifier sequence, can be partially overlapping. The products can be generated in the sample (e.g., in situ), or at least a portion of the products can be generated outside of the sample and then contacted with the sample. The products can be generated enzymatically and / or non-enzymatically. Exemplary products include, for example, RCA products as described in Section II-C, hybridization chain reaction (HCR) products, linear oligonucleotide hybridization chain reaction (LO-HCR) products, branched DNA reaction (bDNA) products, primer exchange reaction (PER) products, or products generated using any combination of these enzymatic and / or non-enzymatic reactions, but are not limited thereto.

[0073] The identifier sequences of the present specification can be contiguous sequences or can include two or more sequences on the same molecule or on separate molecules. When in the same molecule, the two or more sequences can be separated by one or more nucleotide residues.

[0074] The identifier sequences of the present specification can be of any suitable length. In some embodiments, the identifier sequence is from about 1 to about 500 nucleotides in length. In some embodiments, the identifier sequence is about 5, about 10, about 15, about 20, about 25, about 30, about 35, about 40, about 45, about 50, about 55, about 60, about 65, about 70, about 75, about 80, about 85, about 90, about 95, about 100, about 105, about 110, about 115, about 120, about 125, about 130, about 135, about 140, about 145, about 150, about 155, about 160, about 165, about 170, about 175, about 180, about 185, about 190, about 195, or about 200 nucleotides in length or of any integer (or range of integers) of nucleotides between the values shown.

[0075] In some embodiments, at a particular location in a cell or tissue sample, a molecule or complex includes multiple copies of an identifier sequence that is related to, corresponds to, and / or identifies an analyte at that particular location in the sample. In some embodiments, the multiple copies of the identifier sequence are detected using base-by-base sequencing and are used to identify the signal, the analyte. In some embodiments, the molecule or complex includes at least two, at least five, at least ten, at least fifteen, at least twenty, at least twenty-five, or more different identifier sequences that are related to, correspond to, and / or identify the analyte. In some embodiments, the molecule or complex includes at least two, at least five, at least ten, at least twenty-five, at least fifty, at least one hundred, at least two hundred and fifty, at least five hundred, at least one thousand, at least two thousand five hundred, at least five thousand, or more copies of each of one or more identifier sequences. Exemplary molecules and complexes that include one or more different identifier sequences are described in Section II-C.

[0076] In some embodiments, multiple analytes are detected in a cell or tissue sample, and each analyte can be identified using an identifier sequence or a combination of identifier sequences. In some embodiments, the number of different analytes detected in the sample is at least 5, at least 10, at least 25, at least 50, at least 100, at least 250, at least 500, at least 1,000, at least 2,500, at least 5,000, or more. In some embodiments, the number of different identifier sequences used to identify the multiple analytes in the sample is at least 5, at least 10, at least 25, at least 50, at least 100, at least 250, at least 500, at least 1,000, at least 2,500, at least 5,000, or more.

[0077] A. Identifier Sequences from Analytes In some embodiments, the identifier sequences herein include analyte sequences, sequences derived from analytes, or their complements. In some embodiments, the analyte includes a nucleic acid sequence, and the identifier sequence includes the nucleic acid sequence in the analyte or the complement of the nucleic acid sequence.

[0078] In some embodiments, the identifier sequence comprises the sequence of viral nucleic acid or cellular nucleic acid. In some embodiments, the identifier sequence comprises the sequence of viral DNA or RNA. In some embodiments, the identifier sequence is in or derived from a virus or viral particle (e.g., a cell or tissue sample). In some embodiments, the identifier sequence comprises the sequence of cellular DNA or RNA. In some embodiments, the cellular DNA or RNA is from or derived from a prokaryotic cell (e.g., in a tissue sample). In some embodiments, the cellular DNA or RNA is from or derived from a eukaryotic cell (e.g., in a tissue sample). In some embodiments, the identifier sequence comprises the sequence of a nucleic acid molecule in or derived from the nucleus, mitochondria, or chloroplast. In some embodiments, the identifier sequence comprises the sequence of genomic DNA, cellular RNA, or cDNA. In some embodiments, the identifier sequence comprises the sequence of coding RNA and / or non-coding RNA. In some embodiments, the identifier sequence comprises the sequence of messenger RNA (mRNA) including nascent RNA, pre-mRNA, primary transcript RNA, and processed RNA such as capped mRNA (e.g., with a 5’ 7-methylguanosine cap), polyadenylated mRNA (with a poly-A tail at the 3’ end), and spliced mRNA with one or more introns removed. In some embodiments, the identifier sequence comprises the sequence of non-capped mRNA, non-polyadenylated mRNA, or non-spliced mRNA. In some embodiments, the identifier sequence comprises the sequence of a transcript of another nucleic acid molecule (e.g., RNA such as DNA or viral RNA) present in a cell or tissue sample.In some embodiments, the identifier sequence comprises the sequence of a non-coding RNA, and examples of non-coding RNAs (ncRNAs) that are not translated into proteins include transfer RNA (tRNA) and ribosomal RNA (rRNA), as well as microRNA (miRNA), small interfering RNA (siRNA), Piwi-interacting RNA (piRNA), small nucleolar RNA (snoRNA), small nuclear RNA (snRNA), extracellular RNA (exRNA), small Cajal body-specific RNA (scaRNA), and other small non-coding RNAs, as well as long non-coding RNAs such as Xist and HOTAIR. In some embodiments, the identifier sequence comprises the sequence of a small RNA (e.g., less than 200 nucleobases in length) or a large RNA (e.g., an RNA greater than 200 nucleobases in length). Examples of small RNAs include 5.8S ribosomal RNA (rRNA), 5S rRNA, tRNA, miRNA, siRNA, snoRNA, piRNA, small RNAs derived from tRNA (tsRNA), and small RNAs derived from small rDNA (srRNA). The RNA can include double-stranded RNA or single-stranded RNA. The RNA can be circular RNA. The RNA can be bacterial rRNA (e.g., 16s rRNA or 23s rRNA). In some embodiments, the identifier sequence comprises, for example, a sequence spanning an exon-exon junction in a splicing RNA. In some embodiments, the identifier sequence comprises, for example, a sequence spanning an intron-exon or exon-intron junction in DNA or non-splicing RNA. In some embodiments, the identifier sequence can comprise the sequence of a reverse transcript of any of the RNAs disclosed herein. In some embodiments, the identifier sequence can comprise cDNA of any of the RNAs disclosed herein, or the complement of the cDNA.

[0079] In some embodiments, the nucleic acid analyte or its complement is circularized, for example, using template-independent ligation and used as a template for rolling circle amplification (RCA). For example, cDNA produced from reverse transcription of mRNA in a cell or tissue sample can be directly circularized using a single-stranded DNA ligase (e.g., CircLigase™) and used as a template for RCA. In some cases, prior knowledge of the mRNA is not required and a direct ligation approach can be used to sample the entire transcriptome in situ. The RCA product contains the complementary sequence of the cDNA (i.e., the sequence of the mRNA that was reverse transcribed to generate the cDNA), which can be an identifier sequence sequenced using the methods disclosed herein to detect the corresponding mRNA at its location in the cell or tissue sample.

[0080] In some embodiments, the cyclizable probe or probe set hybridizes to a nucleic acid analyte (e.g., cDNA or mRNA) and is ligated to form a circular probe using the nucleic acid analyte as a template, with or without gap filling prior to ligation. In some embodiments, the cyclizable probe or probe set includes 3' and 5' hybridization regions (e.g., 3' arm and 5' arm) that hybridize to a nucleic acid analyte (e.g., cDNA or mRNA). In some embodiments, upon hybridization to the nucleic acid analyte, the 3' and 5' terminal nucleotides of the cyclizable probe or probe set are configured to be ligated after gap filling by a polymerase (e.g., an enzyme having DNA polymerase or reverse transcriptase activity). In some embodiments, the 3' and 5' hybridization regions of the cyclizable probe or probe set are complementary to sequences in a nucleic acid analyte (e.g., cDNA or mRNA) that are not directly linked by a phosphodiester bond, and a gap of one or more nucleotides is formed between the two hybridization regions. In some embodiments, the gap is about 1 to about 500 nucleotides in length. In some embodiments, the gap is about 5, about 10, about 15, about 20, about 25, about 30, about 35, about 40, about 45, about 50, about 55, about 60, about 65, about 70, about 75, about 80, about 85, about 90, about 95, about 100, about 105, about 110, about 115, about 120, about 125, about 130, about 135, about 140, about 145, about 150, about 155, about 160, about 165, about 170, about 175, about 180, about 185, about 190, about 195, or about 200 nucleotides in length, or of any integer (or range of integers) of nucleotides between the values shown. In some embodiments, the gap is about 200 to about 300, about 300 to about 400, or about 400 to about 500 nucleotides in length.In some embodiments, one or more gaps between the 3' and 5' ends of the cyclizable probe or probe set are formed upon hybridization to a nucleic acid analyte (e.g., cDNA or mRNA). For example, two, three, or more gaps can be filled using the nucleic acid analyte as a template.

[0081] In some embodiments, the nucleic acid analyte is DNA (e.g., cDNA), and the gaps are filled by an enzyme having DNA polymerase activity. In some embodiments, the nucleic acid analyte is RNA (e.g., mRNA), and the gaps are filled by an enzyme having reverse transcriptase activity. In some embodiments, gap filling by the enzyme copies the sequence of the nucleic acid analyte (e.g., cDNA or mRNA) into the circularized probe formed from the cyclizable probe or probe set. In some embodiments, because a known sequence flanks the gap, the methods disclosed herein can be used to read out the nucleic acid analyte sequence or its complement (e.g., a "cell barcode") as an identifier sequence of the nucleic acid analyte. For example, a known sequence 3' to the identifier sequence from the nucleic acid analyte can provide a priming site for binding by a sequencing primer disclosed herein, and the identifier sequence can be sequenced base by base.

[0082] In some embodiments, one or more barcode arrays can be constructed within the backbone of a cyclizable probe (e.g., a padlock probe) to distinguish, for example, cyclizable probes that target the same nucleic acid analyte or different nucleic acid analytes (e.g., cDNA or mRNA). In some embodiments, in addition to an identifier sequence (e.g., a "cell barcode"), one or more barcode sequences from a probe or probe set can also be read using the base-by-base sequencing methods disclosed herein. In some embodiments, in a circularized probe, the barcode sequence is adjacent to an identifier sequence (e.g., a sequence complementary to the sequence of a nucleic acid analyte) generated using gap filling, and both the barcode sequence and the identifier sequence can be sequenced in situ, for example, by base-by-base sequencing of the corresponding sequences in the RCA product of the circularized probe. The adjacent barcode sequence and identifier sequence can be sequenced using the same sequencing primer or using separate sequencing primers. In some embodiments, in a circularized probe, the 3' or 5' hybridization region sequence of the cyclizable probe is between the adjacent barcode sequence and the identifier sequence, and the 3' or 5' hybridization region sequence (a known sequence) can be used as a priming site or can be sequenced as a control.

[0083] Exemplary methods for generating circular probes and RCA products that contain identifier sequences derived from nucleic acid analytes (e.g., genomic DNA, mRNA, or cDNA) include those described in Chen et al., Efficient in situ barcode sequencing using padlock probe-based BaristaSeq, Nucleic Acid Research (2018) 46(4):e22, Lee et al., Fluorescent in situ sequencing (FISSEQ) of RNA for gene expression profiling in intact cells and tissues, Nature Protocols (2015) 10:442-458, Lee et al., Highly multiplexed subcellular RNA sequencing in situ, Science (2014) 343:1360-1363, US10,138,509, US10,179,932, US10,494,662, US11,078,520, and US11,085,072, but are not limited thereto, and all of these are incorporated herein by reference.

[0084] B. Barcode sequences as identifier sequences In some embodiments, provided herein are methods that include contacting a cell or tissue sample with a plurality of probes, each of which binds directly or indirectly to a different analyte in the sample. In some embodiments, each probe can include: i) a hybridization sequence (or an analyte-binding moiety such as an antibody or antibody fragment) that binds directly or indirectly to its corresponding analyte; and ii) one or more identifier sequences that are related to, correspond to, and / or identify the corresponding analyte. In some embodiments, the one or more identifier sequences are barcode sequences that are not derived from the corresponding analyte but are assigned to the corresponding analyte. In some cases, the assignment of the identifier sequence to the corresponding analyte provides a particular advantage over the sequence of the endogenous analyte by enabling the design of an identifier sequence that includes a desired number and / or type of nucleotides (e.g., A, T, C, or G) such that a majority (e.g., 70% or more) can be related to off-signals. In some cases, using a barcode sequence instead of the sequence of the endogenous analyte enables the design of a plurality of barcode sequences that collectively reduce the optical crowding of signals detected in successive decoding cycles (e.g., by having a sufficient number of "off" bases among the bases analyzed in the same cycle across a plurality of barcode sequences), facilitating unambiguous analyte identification.

[0085] In some embodiments, in barcode sequencing methods, the barcode sequence is detected for the identification of other molecules that contain nucleic acid molecules (DNA or RNA) that are longer than the barcode sequence itself, as opposed to direct sequencing of longer nucleic acid molecules. In some embodiments, given a sequencing read of N bases, an N-mer barcode sequence has a complexity of 4 N and can require a much shorter sequencing read for molecule identification compared to non-barcode sequencing methods such as direct sequencing. For example, 1024 molecular species can be discriminated with a 5-nucleotide barcode sequence (4 5While it can be identified using, e.g., a 1024-nucleotide barcode, an 8-nucleotide barcode can be used to identify up to 65,536 molecular species, which is more than the total number of distinct genes in the human genome. In some embodiments, it is the barcode sequences (e.g., contained in a probe or RCA product), rather than endogenous sequences that may be efficient reads in terms of information per cycle of sequencing, that are detected. In some embodiments, since the barcode sequences are pre-determined, they can also be designed to feature error detection and correction mechanisms. See, e.g., U.S. Patent Publication No. 2019 / 0055594 and U.S. Patent Publication No. 2021 / 0164039, which are hereby incorporated by reference in their entirety.

[0086] In some aspects, an analyte can be associated with, corresponding to, and / or identified using one or more barcode sequences, e.g., at least 2, 3, 4, 5, 6, 7, 8, 9, 10, or more different barcode sequences. Assignment of barcode sequences to analytes can be performed as described in Section III.

[0087] In some embodiments, a single barcode array may uniquely identify one of a plurality of different analytes. In some embodiments, two or more different barcode arrays may be assigned to the same analyte, and any of the barcode arrays may uniquely identify one of the plurality of different analytes. In some embodiments, a single barcode array is provided as a combination of barcode arrays, and the combination may uniquely identify one of the plurality of different analytes. The combination of barcode arrays may be provided by two or more probe molecules. In some embodiments, the number of distinct barcode arrays in a population of nucleic acid probes is less than the number of distinct target analytes (e.g., nucleic acid analytes and / or protein analytes) of the nucleic acid probes, yet the distinct target analytes can still be uniquely identified from one another, for example, by encoding probes with different combinations of barcode arrays. However, it is not necessary to use all possible combinations of a given set of barcode arrays. For example, each probe may contain 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, or more barcode arrays. In some embodiments, for example, a population of nucleic acid probes may each contain the same number of barcode arrays, but in other cases, different numbers of barcode arrays may be present on different probes.

[0088] As an illustrative example, a first probe may contain a first target binding sequence (or a first target binding moiety), a first barcode array, and a second barcode array, and a second different probe may contain a second target binding sequence or target binding moiety (different from the first target binding sequence (or target binding moiety) in the first probe), and may contain the same first barcode array as in the first probe, but may contain a third barcode array instead of the second barcode array. Thereby, such probes can be distinguished by determining the combinations of the various barcode arrays that are present at a given location in the sample or associated with a given probe at a given location.

[0089] The barcode array can be attached to the analyte or another moiety or structure in a reversible or irreversible manner. In some embodiments, the barcode array can include about 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, or more than 30 nucleotides.

[0090] In some embodiments, a probe such as a nucleic acid probe (or a combination of nucleic acid probes configured to target the same analyte) can include 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15 or more, 20 or more, 32 or more, 40 or more, or 50 or more different barcode arrays. In some embodiments, the nucleic acid probe or combination of nucleic acid probes can include 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15 or more, 20 or more, 32 or more, 40 or more, or 50 or more copies of a particular barcode array. The different barcode arrays, or copies of the same barcode array, can be positioned anywhere within or between the nucleic acid probes. If two or more barcode arrays or two or more copies are present, the barcode arrays or copies can be positioned adjacent to each other and / or other sequences can be interspersed. In some embodiments, two or more of the barcode arrays or copies can also at least partially overlap. In some embodiments, two or more of the barcode arrays or copies within the same probe do not overlap. In some embodiments, any two or more or all of the barcode arrays or copies within the same probe can be separated from each other by at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more nucleotides, such as by at least a phosphodiester bond (e.g., they can be immediately adjacent to each other but do not overlap).

[0091] The barcode array can be of any length. In some embodiments, the barcode arrays can independently have the same or different lengths, such as at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 50 nucleotides in length. In some embodiments, the individual barcode arrays can be 24 or less, 16 or less, 15 or less, 14 or less, 13 or less, 12 or less, 10 or less, 9 or less, or 8 or less nucleotides in length. Any combination of these is possible. For example, the barcode array can be 5 to 10 nucleotides, 8 to 15 nucleotides, etc.

[0092] In some embodiments, the barcode array can include two or more sub-barcode arrays that function together as a single barcode array. For example, the polynucleotide can include two or more polynucleotide sequences (e.g., sub-barcode arrays) separated by one or more non-barcode sequences. In some embodiments, the barcode array can also provide a platform for targeting functional groups such as oligonucleotides, oligonucleotide-antibody conjugates, oligonucleotide-streptavidin conjugates, modified oligonucleotides, affinity purification, detectable moieties, enzymes, enzymes for detection assays, or other functional groups, and / or for the detection and discrimination of polynucleotides that contain or are directly or indirectly linked to the barcode array.

[0093] The barcode sequences can be arbitrary or random. In certain cases, for example, the barcode sequences are selected to reduce or minimize homology with other components in the sample so that the barcode sequences themselves do not bind to or hybridize with other nucleic acids suspected of being in the cell or other sample. In some embodiments, between a particular barcode sequence and another sequence (e.g., a cellular nucleic acid sequence in the sample, or another barcode sequence in a probe added to the sample), the homology can be less than 10%, less than 8%, less than 7%, less than 6%, less than 5%, less than 4%, less than 3%, less than 2%, or less than 1%. In some embodiments, the homology can be 20, 19, 18, 17, 16, 15, 14, 13, 12, 11, 10, 9, 8, 7, 6, 5, 4, 3, or less than 2 bases, and in some embodiments, the bases are consecutive bases.

[0094] In some embodiments, the nucleic acid probes and barcode sequences in the probes disclosed herein can be made using only 2 or only 3 of the 4 bases, such as omitting all "G"s and / or all "C"s in the probe. Sequences lacking either "G" or "C" can form very little secondary structure and, in certain embodiments, can contribute to more uniform and faster hybridization.

[0095] In some embodiments, the set of signals detected during a periodic in situ decoding process can be used to generate a signal code sequence corresponding to the barcode sequence assigned to the analyte, and the signal code sequence can be compared to the codewords in the codebook to identify a match. In some cases, the codewords in the codebook can correspond to a physical barcode that includes a series of on and off bits corresponding to a series of on and off signals detected in an image acquired in successive cycles during the decoding process. The presence and location of one or more target analytes can then be inferred from the detected signal code sequence and the detected locations of the corresponding on and off signals for each signal code sequence at specific locations in a series of images of the sample. In some aspects, the signal codes of the signal code sequence can include, for example, a signal (on signal), the absence of a signal (off signal), or combinations thereof.

[0096] A codebook that includes codewords (corresponding to an identifier sequence or a “physical” nucleic acid barcode sequence) can be designed to meet a set of specific design criteria. For example, the codebook can be designed to ensure that it includes a specified number (or minimum number) of unique codewords / barcode sequences (e.g., corresponding to a specified number (or minimum number) of target analytes to be identified). In some cases, for example, the codebook includes at least 2, at least 5, at least 10, at least 20, at least 40, at least 60, at least 80, at least 100, at least 200, at least 400, at least 600, at least 800, at least 1,000, at least 2,000, at least 4,000, at least 6,000, at least 8,000, at least 10,000, at least 20,000, at least 40,000, at least 60,000, at least 80,000, at least 100,000, at least 200,000, at least 400,000, at least 600,000, at least 800,000, at least 1,000,000, at least 2×10 6 at least 3×106 , at least 4×10 6 , at least 5×10 6 , at least 6×10 6 , at least 7×10 6 , at least 8×10 6 , at least 9×10 6 , at least 10 7 , at least 10 8 , at least 10 9 , or more than 10 9 unique codeword / barcode arrays. In some cases, the codebook may include any number of unique codeword / barcode arrays within the range of values of this paragraph.

[0097] In some cases, the codeword / barcode arrays within a given codebook may be designed to meet a specified pairwise edit distance (e.g., a specified minimum pairwise edit distance) to enable barcode error detection and correction. "Edit distance" is a numerical value that quantifies how different two strings (e.g., text strings) are from each other by counting the minimum number of edit operations required to transform one string into the other. Examples of edit distance metrics include, but are not limited to, Hamming distance, Levenshtein distance, longest common subsequence (LCS) distance, etc. For example, the Levenshtein distance between two strings is the minimum number of single-character edits (e.g., insertions, deletions, or substitutions) required to transform one string into the other. The longest common subsequence (LCS) distance is an edit distance where the only permitted edit operations are insertions and deletions, and each of them has a unit cost assigned. The Hamming distance between two strings of equal length (e.g., substitution is the only permitted edit operation) is the number of positions in the two strings where the corresponding symbols are different.

[0098] The Hamming distance and / or the Levenshtein distance (when the error penalty assigned to the difference between two strings is an integer value) allows for a natural interpretation of error correction, and a minimum pairwise barcode distance of 2k + 1 allows for the correction of up to k errors. In some cases, the designed barcodes of a given codebook may need to have a minimum pairwise edit distance (e.g., minimum pairwise Hamming distance, minimum pairwise Levenshtein distance, or minimum pairwise LCS distance) to guarantee the error correction ability to correct at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, or at least 10 barcode errors (e.g., errors that occur during the decoding process).

[0099] In some cases, it is desirable to design the codebook so that optical crowding in the image used for decoding is minimized. For example, in some cases, the set of codewords / barcodes can be designed such that the on-signal (or combined weight) of the set of codewords / barcodes is distributed more or less evenly over multiple decoding cycles. In some cases, the codewords / barcodes within the codebook can be designed such that at most 1 bit is on during any given decoding cycle. For example, the codewords / barcodes within a codebook designed for use with a 4 "color" (e.g., 4 detection channels) decoding device can include a block of 4 bits (1 for each detection channel) for each decoding cycle, with at most 1 bit on at any given cycle. As used herein, "color" can include any color used in fluorescence microscopy, as well as a dark or off state (e.g., no detectable signal).

[0100] In some embodiments, the nucleotides within the barcode array corresponding to the codewords can be selected to reduce optical crowding of the signals detected in the cycle. For example, in multiple barcode arrays, the corresponding nucleotides detected in the same base sequencing cycle can be selected to reduce optical crowding of the signals detected in the cycle, and the multiple barcode arrays can be designed such that optical crowding in multiple (e.g., most) cycles is limited. In some embodiments, a signal code (e.g., a signal of a first color, a signal of a second color, a signal of a third color, or the absence of a signal) representing a nucleotide in a first barcode array is an off-signal for each cycle that is an on-signal, and a signal code representing the corresponding nucleotide in a second barcode array is an off-signal. In some embodiments, a dark state or an off state can be included in one or more codewords / barcode arrays, and a nucleotide mix for periodic base-by-base sequencing is used to detect the dark state or the off state. In some embodiments, the assignment of barcode arrays to each of the plurality of analytes is performed using any suitable scheme (e.g., as described in Section III.A).

[0101] C. Molecules or Complexes Containing Identifier Sequences An identifier sequence containing an analyte and from the barcode array assigned to the analyte is present in various molecules or complexes thereof and can be detected using in situ base-by-base sequencing. Exemplary molecules and complexes are described herein.

[0102] (i) Labeling Agents In the methods and systems described herein, one or more labeling agents that can be conjugated to or otherwise coupled to one or more features can be used to characterize an analyte, a cell, and / or a cell feature. In some cases, the cell feature includes a cell surface feature. Analytes can include, but are not limited to, proteins, receptors, antigens, surface proteins, transmembrane proteins, clusters of differentiation proteins, protein channels, protein pumps, carrier proteins, phospholipids, glycoproteins, glycolipids, cell-cell interaction protein complexes, antigen presenting complexes, major histocompatibility complexes, engineered T cell receptors, T cell receptors, B cell receptors, chimeric antigen receptors, gap junctions, adherens junctions, or any combination thereof. In some cases, the cell feature can include intracellular analytes such as proteins, protein modifications (e.g., phosphorylation state or other post-translational modifications), nuclear proteins, nuclear membrane proteins, or any combination thereof. In some embodiments, the method includes one or more post-fixing (also referred to as post-fixation) steps after contacting the sample with one or more labeling agents.

[0103] In some embodiments, the labeling agent includes an analyte-binding moiety that can bind to an analyte (e.g., a biological analyte, such as a macromolecular component). The binding moiety can include a protein, a peptide, an antibody (or an epitope-binding fragment thereof), a lipophilic moiety (such as cholesterol), a cell surface receptor-binding molecule, a receptor ligand, a small molecule, a bispecific antibody, a bispecific T cell engager, a T cell receptor engager, a B cell receptor engager, a prodrug, an aptamer, a monobody, an affimer, a darpin, and a protein scaffold, or any combination thereof, but is not limited thereto. The binding moiety can be attached directly or indirectly to a reporter oligonucleotide that indicates the analyte or feature to which the binding moiety binds. For example, the reporter oligonucleotide can include a barcode sequence that enables identification of the binding moiety and the corresponding analyte. For example, a labeling agent that is specific for one type of analyte or cell feature can have a first reporter oligonucleotide coupled thereto, while a labeling agent that is specific for a different analyte or cell feature can have a different reporter oligonucleotide coupled thereto. Exemplary labeling agents, reporter oligonucleotides, and methods of use are described, for example, in U.S. Patent No. 10,550,429, U.S. Patent Publication No. 2019 / 0177800, and U.S. Patent Publication No. 2019 / 0367969, each of which is incorporated herein by reference in its entirety.

[0104] In some embodiments, the analyte-binding moiety includes one or more nucleic acid moieties. The one or more nucleic acid moieties can specifically bind to a target analyte, such as a target nucleic acid via nucleic acid hybridization. In some embodiments, the analyte-binding moiety includes one or more antibodies or epitope-binding fragments thereof. The antibody or epitope-binding fragment containing the analyte-binding moiety can specifically bind to the target analyte. In some embodiments, the analyte is a protein (e.g., a protein on the surface of a biological sample (e.g., a cell) or an intracellular protein).

[0105] In some embodiments, a plurality of analyte labels comprising a plurality of analyte binding moieties bind to a plurality of analytes present in a biological sample. In some embodiments, the plurality of analytes comprises a single species of analyte (e.g., a single species of polynucleotide or polypeptide). In some embodiments where the plurality of analytes comprises a single species of analyte, the analyte binding moieties of the plurality of analyte labels are the same. In some embodiments where the plurality of analytes comprises a single species of analyte, the analyte binding moieties of the plurality of analyte labels are different (e.g., members of the plurality of analyte labels can have two or more species of analyte binding moieties, and each of the two or more species of analyte binding moieties binds to the single species of analyte at, for example, different binding sites). In some embodiments, the plurality of analytes comprises a plurality of different species of analyte (e.g., a plurality of different species of polynucleotide or polypeptide).

[0106] In other cases, for example, to facilitate sample multiplexing, different subsets of labels that are specific for a particular analyte or cell feature can be used. For example, a first subset of labels comprises a binding moiety (e.g., a nucleic acid or an antibody or a lipophilic moiety) coupled to a first reporter oligonucleotide, and a second subset of labels comprises a binding moiety coupled to a second reporter oligonucleotide that is different from the first reporter oligonucleotide.

[0107] In some aspects, these reporter oligonucleotides can comprise nucleic acid barcode sequences that enable identification of the binding moiety to which the reporter oligonucleotide is coupled. The selection of oligonucleotides as reporters can provide the advantage of being able to provide significant diversity from a sequence perspective while being attachable to most biomolecules such as nucleic acids and antibodies, and detectable using, for example, the in situ detection techniques described herein.

[0108] Attachment (coupling) of the reporter oligonucleotide to the binding moiety can be achieved through any of a variety of direct or indirect, covalent or non-covalent associations or attachments. For example, the oligonucleotide can be covalently attached to a portion of the binding moiety (a protein, such as an antibody or antibody fragment, etc.) using chemical conjugation techniques (e.g., the Lightning-Link® protein labeling kit available from Innova Biosciences), and also using other non-covalent mechanisms such as biotinylated antibodies and oligonucleotides (or beads containing one or more biotinylated linkers coupled to the oligonucleotide) using, for example, an avidin or streptavidin linker. Antibody and oligonucleotide biotinylation techniques are available. See, for example, Fang, et al., “Fluoride-Cleavable Biotinylation Phosphoramidite for 5’-end-Labelling and Affinity Purification of Synthetic Oligonucleotides,” Nucleic Acids Res. Jan. 15, 2003; 31(2):708-715, which is hereby incorporated by reference in its entirety for all purposes. Similarly, protein and peptide biotinylation techniques can be used. See, for example, U.S. Patent No. 6,265,552, which is hereby incorporated by reference in its entirety for all purposes. Further, click reaction chemistry can be used to couple the reporter oligonucleotide to the binding moiety. Commercially available kits such as those from Thunderlink and Abcam can be used to couple the reporter oligonucleotide to the binding moiety as needed. In another example, the binding moiety is indirectly coupled to a reporter oligonucleotide that contains a barcode sequence that identifies the binding moiety (e.g., via hybridization). For example, the binding moiety can be directly coupled (e.g., covalently) to a hybridization oligonucleotide that contains a sequence that hybridizes to the sequence of the reporter oligonucleotide.Hybridization of a hybridization oligonucleotide to a reporter oligonucleotide couples the binding moiety to the reporter oligonucleotide. In some embodiments, the reporter oligonucleotide is releasable from the binding moiety, such as upon application of a stimulus. For example, the reporter oligonucleotide can be attached to the binding moiety through a labile bond (e.g., chemically labile, photocleavable, thermally labile, etc.).

[0109] In some embodiments, multiple different species of an analyte (e.g., polynucleotide or polypeptide) from a biological sample can then be related to one or more physical properties of the biological sample. For example, multiple different species of the analyte can be related to the location of the analyte in the biological sample. Such information (e.g., proteomics information when the analyte binding moiety recognizes a polypeptide) can be used in relation to other spatial information (e.g., DNA sequence information, transcriptome information (e.g., sequence of transcripts), or both, such as genetic information from the biological sample). For example, cell surface proteins of a cell can be related to one or more physical properties of the cell (e.g., cell shape, size, activity, or type). The one or more physical properties can be characterized by imaging the cell. The cell can be bound by an analyte labeling agent comprising an analyte binding moiety that binds to the cell surface protein and an analyte binding moiety barcode that identifies the analyte binding moiety. The results of protein analysis in a sample (e.g., tissue sample or cell) can be related to DNA and / or RNA analysis in the sample.

[0110] (ii) Nucleic acid molecules and complexes In some embodiments, the binding of one or more nucleic acid molecules (e.g., probes) to an analyte in a biological sample can be direct or indirect. In some embodiments, the one or more nucleic acid molecules include primary probes that bind directly to their corresponding analytes. In other embodiments, the one or more nucleic acid molecules include one or more probes that bind directly or indirectly to the primary probes. Identifier sequences can be present in nucleic acid molecules and complexes that include one or more nucleic acid analytes and / or one or more nucleic acid probes that bind directly or indirectly to a nucleic acid analyte. One or more identifier sequences in the nucleic acid molecules and complexes can be located downstream (5') of one or more priming sites for the binding of one or more sequencing primers. One or more identifier sequences can be related to an analyte to which a primary probe binds, and sequencing of the identifier sequence can enable identification of the analyte.

[0111] For example, single molecule fluorescence in situ hybridization (smFISH) can be used to determine expression levels by detecting cellular nucleic acid analytes such as mRNA. In smFISH, typically a set of 30-50 oligonucleotides, each about 20 nucleotides in length, can hybridize to a complementary mRNA target. The individual transcripts are then visualized and quantified as diffraction-limited spots using widefield epifluorescence microscopy. In some embodiments, the smFISH probes can carry an overhang sequence 10-30 nucleotides in length, instead of a directly conjugated fluorescent label, which does not hybridize to the mRNA target and can be detected by hybridization to a probe that is detectably (e.g., fluorescently) labeled. The overhang sequence in the smFISH probe can include one or more identifier sequences, e.g., barcode sequences. Thus, the smFISH probe and the mRNA target can form a nucleic acid complex that includes a plurality of barcode sequences that can be sequenced in situ base-by-base using the methods disclosed herein.

[0112] In some embodiments, probes (e.g., first and / or second nucleic acid probes) that are introduced into cells or otherwise used to contact a biological sample such as a tissue sample are disclosed herein. The probe can typically include any of a variety of entities that can hybridize to a nucleic acid by Watson-Crick base pairing, such as DNA, RNA, LNA, PNA, etc. The probe typically contains a targeting sequence or hybridization region that can bind directly or indirectly to at least a portion of a nucleic acid (e.g., a target nucleic acid or a probe). For example, the probes described herein can bind to a specific target nucleic acid (e.g., mRNA, or other nucleic acids contemplated herein). In some embodiments, the probe includes one or more priming sites for the binding of sequencing primers and one or more downstream identifier sequences that can be sequenced using the sequencing primers.

[0113] Any of the probes described herein can be a linear probe. In some embodiments, the linear probe can include an analyte binding sequence (sometimes also referred to as a target recognition sequence). In some embodiments, the linear probe can include sequences that do not hybridize to the target nucleic acid or target probe, such as 5' overhangs and / or 3' overhangs. In some embodiments, the sequences (e.g., 5' overhangs, 3' overhangs) are non-hybridizing to the target nucleic acid or target probe but can hybridize to each other and / or one or more other probes to form hybridization complexes, such as those in a hybridization chain reaction (HCR), branched DNA reaction, etc. The hybridization complex can include one or more priming sites for the binding of sequencing primers and one or more downstream identifier sequences that can be sequenced using the sequencing primers.

[0114] The analyte-binding sequence of the probe can be positioned anywhere within the probe. For example, the analyte-binding sequence of a probe that binds to an analyte can be 5' or 3' relative to an identifier sequence (e.g., a barcode sequence) in the probe. In some embodiments, the analyte-binding sequence of the probe can be 5' or 3' relative to a priming site or a portion thereof in the probe. In some embodiments, the analyte-binding sequence can include a sequence that is substantially complementary to a portion of the analyte. In some embodiments, the portion can be at least 50%, at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% complementary. The analyte-binding sequence of the probe can be determined with reference to a target analyte (e.g., cellular RNA) present or suspected to be present in the sample. In some embodiments, two or more analyte-binding sequences can be used to identify a particular analyte that includes or is related to the target nucleic acid. The two or more analyte-binding sequences can be in the same probe or in different probes. For example, a plurality of probes that can bind (e.g., hybridize) continuously and / or simultaneously to different regions of the same target analyte or to different target analytes can be used. In some embodiments, a first probe is provided in a first plurality of probes that bind directly or indirectly to a first analyte. In some embodiments, a second probe is provided in a second plurality of probes that bind directly or indirectly to a second analyte. The first and second analytes can be the same or different, and the first plurality of probes can be contacted with the sample, followed by detecting a signal associated with the first plurality of probes, and the sample can be contacted with the second plurality of probes and a signal associated therewith can be detected.

[0115] In some embodiments, the identifier array (e.g., barcode array) can be on a single overhang of a probe provided herein (e.g., on the overhang of an L-shaped nucleic acid probe, or on one overhang of a U-shaped probe). In other embodiments, the barcode array can be positioned on two overhang regions of the probe (e.g., both the 5’ and 3’ overhangs of a U-shaped probe). In some embodiments, the barcode sequence of an individual probe is associated with and uniquely identifies an analyte. In other embodiments, a combination of barcode sequences of multiple probes that hybridize to the same target analyte uniquely identifies the target analyte. For example, a first probe is provided among a first plurality of probes that bind directly or indirectly to a first analyte. In some aspects, the first plurality of probes collectively include a first combination of barcode sequences, and the first combination of barcode sequences identifies the first analyte. In some embodiments, a second probe is provided among a second plurality of probes that bind directly or indirectly to a second analyte. In some aspects, the second plurality of probes collectively include a second combination of barcode sequences, and the second combination of barcode sequences identifies the second analyte. In some embodiments, splitting the barcode sequence that identifies a target analyte among multiple probes (e.g., multiple primary probes that do the following) can reduce the length requirement of the probe (e.g., the first probe).

[0116] In some embodiments, the identifier array (e.g., barcode array) can be in a primary probe, such as a primary probe that hybridizes directly to a cellular nucleic acid molecule, such as mRNA, or in a product of the primary probe (e.g., RCA product). In some embodiments, the primary probe can include a 3' or 5' overhang upon hybridization to the cellular nucleic acid molecule. In some embodiments, the 3' or 5' overhang includes one or more barcode arrays. In some embodiments, the primary probe can include both a 3' overhang and a 5' overhang upon hybridization to the cellular nucleic acid molecule. In some embodiments, the 3' overhang and the 5' overhang each independently include one or more barcode arrays. In some embodiments, the primary probe can be a circular primary probe. In some embodiments, the circular primary probe includes one or more barcode arrays outside the analyte binding sequence of the probe. In some embodiments, the primary probe can be a circularizable primary probe or probe set. In some embodiments, the circularizable primary probe or probe set includes one or more barcode arrays outside the analyte binding sequence of the probe or probe set. In some embodiments, the primary probe can include a split hybridization region configured to hybridize to a splint. In some embodiments, the split hybridization region includes one or more barcode arrays. In some embodiments, the primary probe includes one or more barcode arrays outside the analyte binding sequence and the split hybridization region of the primary probe.

[0117] In some embodiments, the identifier array (e.g., barcode array) can be in a probe that binds indirectly to a cellular nucleic acid molecule such as mRNA, or in a product of the probe (e.g., RCA product). In some embodiments, the probe can be an intermediate probe that crosslinks the binding of a probe (e.g., a primary probe) or its product and another probe or its product. In some embodiments, the probe can be a detection probe that hybridizes to a primary probe or intermediate probe or their products but does not bind directly or indirectly to any other probe. The detection probe can further include a priming site for sequencing the identifier array in the detection probe using the methods disclosed herein.

[0118] In some embodiments, an intermediate probe or a detection probe may include a 3' or 5' overhang upon hybridization to another probe. In some embodiments, the 3' or 5' overhang includes one or more barcode sequences. In some embodiments, an intermediate probe or a detection probe may include both a 3' overhang and a 5' overhang upon hybridization to another probe. In some embodiments, the 3' overhang and the 5' overhang each independently include one or more barcode sequences. In some embodiments, the intermediate probe or the detection probe may be a circular probe. In some embodiments, the circular probe includes one or more barcode sequences outside the target binding sequence of the probe. In some embodiments, the intermediate probe or the detection probe may be a circularizable probe or a probe set. In some embodiments, the circularizable probe or the probe set includes one or more barcode sequences outside the target binding sequence of the probe or the probe set. In some embodiments, the intermediate probe or the detection probe may include a split hybridization region configured to hybridize to a splint. In some embodiments, the split hybridization region may include a priming site, and the splint or a portion thereof may be used as a sequencing primer to sequence an identifier sequence in the intermediate probe or the detection probe using the methods disclosed herein. Exemplary probes including a split hybridization region are described in US2022 / 0049302, which is hereby incorporated by reference in its entirety, but are not limited thereto. In some embodiments, the split hybridization region includes one or more barcode sequences. In some embodiments, the intermediate probe or the detection probe includes one or more barcode sequences outside the target binding sequence of the probe and the split hybridization region.

[0119] In some embodiments, the identifier array (e.g., barcode array) can be in a rolling circle amplification (RCA) product molecule, a complex comprising initiators and amplifiers for hybridization chain reaction (HCR), a complex comprising initiators and amplifiers for linear oligonucleotide hybridization chain reaction (LO-HCR), a primer exchange reaction (PER) product molecule, a complex comprising pre-amplifiers and amplifiers for branched DNA (bDNA), or a complex comprising any two or more of the aforementioned molecules and complexes. For example, a bDNA complex or an HCR complex can assemble on an RCA product. See, e.g., US2021 / 0198727, which is incorporated herein by reference in its entirety. The priming site can be provided 3' to each identifier array in the molecule or complex for sequencing the identifier array in situ using SBS, SBB, or any other base-by-base sequencing method.

[0120] In some embodiments, a molecule or complex comprising an identifier array (e.g., barcode array) and a priming site for sequencing the identifier array can be generated using targeted assembly of branched structures (e.g., bDNA or branched assays using locked nucleic acids (LNAs)), programmed in situ proliferation of concatemers by enzymatic rolling circle amplification (RCA) (e.g., as described in US2019 / 0055594, which is incorporated herein by reference), HCR or the like, assembly of phase-linked linear DNA structures using successive rounds of chemical ligation (clampFISH), hairpin-mediated concatemerization (e.g., as described in US2020 / 0362398, which is incorporated herein by reference), primer exchange reactions such as signal amplification by exchange reaction (SABER) or SABER with DNA exchange (Exchange-SABER). In some embodiments, non-enzymatic methods can be used.

[0121] In some embodiments, a complex comprising an identifier array (e.g., a barcode array) and a priming site hybridizes directly or indirectly (via one or more oligonucleotides) to the sequence of a nucleic acid analyte, for example, as shown in FIG. 2, and includes an amplification agent, a probe that directly or indirectly targets the nucleic acid analyte, or a product of the nucleic acid analyte or the probe. In some embodiments, the assembly includes one or more amplification agents, each of which includes an amplification agent repeat sequence. In some aspects, one or more of the amplification agents are labeled. In some aspects, one or more of the amplification agents are unlabeled. For exemplary complexes, see, for example, US2020 / 0399689 and US2022 / 0064697, which are hereby incorporated by reference in their entirety.

[0122] In some embodiments, the complex comprising an identifier array (e.g., a barcode array) and a priming site comprises, for example, as shown in FIG. 2, an HCR complex. HCR is an enzyme-free nucleic acid amplification based on the triggered strand of hybridization of nucleic acid molecules starting from HCR monomers that hybridize to each other to form a nicked nucleic acid polymer. This polymer is the product of the HCR reaction that is ultimately detected to indicate the presence of the target analyte. HCR is described in detail in Dirks and Pierce, 2004, PNAS, 101(43), 15275-15278, as well as US7,632,641 and US7,721,721 (see also US2006 / 00234261, Chemeris et al, 2008 Doklady Biochemistry and Biophysics, 419, 53-55, Niu et al, 2010, 46, 3089-3091, Choi et al, 2010, Nat. Biotechnol. 28(11), 1208-1212, and Song et al, 2012, Analyst, 137, 1396-1401). HCR monomers typically contain a hairpin or other metastable nucleic acid structure. In the simplest form of HCR, two different types of stable hairpin monomers, herein referred to as the first and second HCR monomers, undergo a chain reaction of hybridization events upon introduction of an "initiator" nucleic acid molecule to form a long nicked double-stranded DNA molecule. An HCR monomer has a hairpin structure that includes a double-stranded stem region, a loop region connecting the two strands of the stem region, and a single-stranded region at one end of the double-stranded stem region. The single-stranded region that is exposed when the monomer is in the hairpin structure (and thus available for hybridization to another molecule, e.g., an initiator or another HCR monomer) can be known as the "toehold region" (or "input domain"). Each of the first HCR monomers further includes a sequence that is complementary to a sequence within the exposed toehold region of the second HCR monomer. This complementary sequence in the first HCR monomer can be known as the "interaction region" (or "output domain").Similarly, each of the second HCR monomers includes an interaction region (output domain), e.g., an array that is complementary to the exposed toehold region (input domain) of the first HCR monomer. In the absence of the HCR initiator, these interaction regions are protected by the secondary structure (e.g., they are not exposed), and thus the hairpin monomers are stable or kinetically trapped (also referred to as "quasi-stable"), and the first and second sets of HCR monomers cannot hybridize to each other and thus remain as monomers (e.g., preventing the system from rapidly equilibrating). However, when the initiator is introduced, it can hybridize to and enter the exposed toehold region of the first HCR monomer, opening it up. This exposes the interaction region of the first HCR monomer (e.g., an array complementary to the toehold region of the second HCR monomer), enabling it to hybridize to and enter the second HCR monomer in the toehold region. This hybridization and entry then opens up the second HCR monomer, exposing its interaction region (complementary to the toehold region of the first HCR monomer), enabling it to hybridize to and enter another first HCR monomer. The reaction continues in this way until all of the HCR monomers are consumed (e.g., all of the HCR monomers are incorporated into the polymer chain). Ultimately, this chain reaction leads to the formation of a nicked strand with alternating units of the first and second monomer species. Thus, the presence of the HCR initiator is required to initiate the HCR reaction by hybridization to and entry into the first HCR monomer. The first and second HCR monomers are designed to hybridize to each other and can thus be defined as cognate to each other. They are also cognate to a given HCR initiator sequence. HCR monomers that interact (hybridize) with each other can be described as a set of HCR monomers or as an HCR monomer, or hairpin, system.

[0123] The HCR reaction can be carried out using three or more species or types of HCR monomers. For example, a system containing three HCR monomers can be used. In such a system, each first HCR monomer can include an interaction region that binds to the toehold region of a second HCR monomer, each second HCR can include an interaction region that binds to the toehold region of a third HCR monomer, and each third HCR monomer can include an interaction region that binds to the toehold region of the first HCR monomer. Then, the HCR polymerization reaction will proceed as described above, except that the resulting product is a polymer having repeating units of the first, second, and third monomers in sequence. Corresponding systems having a greater number of sets of HCR monomers can be used.

[0124] In some embodiments, the HCR product (e.g., a nicked strand of alternating units of monomer species) can include multiple overhangs, each containing one or more identifier sequences (e.g., barcode sequences) and a priming site for sequencing the identifier sequences in situ. Thus, the HCR as used herein does not require the use of detectably labeled HCR monomers. Rather, one or more HCR monomers can include one or more identifier sequences (e.g., barcode sequences), and detection of the HCR product includes sequencing the identifier sequences in situ.

[0125] In some embodiments, similar to the HCR reaction using hairpin monomers, a linear oligohybridization chain reaction (LO-HCR) can also generate a complex comprising an identifier sequence (e.g., a barcode sequence) and a priming site for sequencing the identifier sequence. In some embodiments, methods for detecting an analyte in a sample are provided herein, (i) performing a linear oligohybridization chain reaction (LO-HCR), wherein the initiator contacts a plurality of LO-HCR monomers of at least a first and a second species to generate a polymeric LO-HCR product hybridized to a target nucleic acid molecule, the first species comprising a first hybridization region complementary to the initiator and a second hybridization region complementary to the second species, the first and second species being linear single-stranded nucleic acid molecules, the initiator being provided in one or more moieties that hybridize directly or indirectly to or are included within the target nucleic acid molecule, and (ii) detecting the polymeric product, thereby detecting the analyte. In some embodiments, the first species and / or the second species may not comprise a hairpin structure. In some embodiments, the plurality of LO-HCR monomers may not comprise a metastable secondary structure. In some embodiments, the LO-HCR polymer may not comprise a branched structure. In some embodiments, performing the linear oligohybridization chain reaction comprises contacting the target nucleic acid molecule with the initiator to provide an initiator hybridized to the target nucleic acid molecule. Exemplary methods and compositions for LO-HCR are described in US2021 / 0198723, which is hereby incorporated by reference in its entirety.

[0126] In some embodiments, the polymeric products in LO-HCR can include multiple overhangs, each of which includes one or more identifier sequences (e.g., barcode sequences) and priming sites for sequencing the identifier sequences in situ. Thus, LO-HCR as used herein does not require the use of detectably labeled LO-HCR monomers; rather, one or more LO-HCR monomers can include one or more identifier sequences (e.g., barcode sequences), and detection of the LO-HCR products includes sequencing the identifier sequences in situ.

[0127] In some embodiments, molecules (e.g., concatemer molecules) that include an identifier sequence (e.g., barcode sequence) and a priming site are generated by a primer exchange reaction (PER). In various embodiments, a primer having a domain at its 3’ end binds to a catalytic hairpin and is extended with a new domain by a strand-displacing polymerase. In various embodiments, the strand-displacing polymerase is Bst polymerase. In various embodiments, the catalytic hairpin includes a stopper that releases the strand-displacing polymerase. In various embodiments, branch migration displaces the extended primer, which can then dissociate. In various embodiments, the primer undergoes iterative cycles to form a concatemer molecule; see, e.g., US2019 / 0106733, which is incorporated herein by reference for exemplary molecules and PER reaction components. In various embodiments, the concatemer molecule includes multiple copies of one or more identifier sequences (e.g., barcode sequences) and priming sites for sequencing the identifier sequences. Thus, instead of hybridizing multiple labeled oligonucleotide probes to the concatemer molecule, the concatemer molecule can be detected by sequencing the identifier sequences in situ.

[0128] (iii) Rolling circle amplification (RCA) products In some embodiments, molecules (e.g., concatemer molecules) that include an identifier sequence (e.g., a barcode sequence) and a priming site are generated by RCA of a circular nucleic acid molecule, as shown, for example, in FIG. 2. As described in Section II-A, a nucleic acid analyte (e.g., cDNA) can be circularized to generate a circular nucleic acid molecule, and the RCA product includes an identifier sequence derived from the nucleic acid analyte. In some embodiments, any suitable cyclizable probe or probe set disclosed herein can be circularized, with or without gap filling prior to circularization, using, for example, a nucleic acid analyte or a probe as a template. The RCA product can include a barcode sequence or its complement that can be sequenced in situ.

[0129] In some aspects, the probes disclosed herein (e.g., primary probes, intermediate probes, detection probes, etc.) are ligated to form a circular construct (e.g., a circular probe). In some embodiments, the circular construct is formed using template primer extension followed by ligation. In some embodiments, the ligation probe is generated using an analyte as a template. In some embodiments, the circular construct is formed by providing an insert between the ends that are ligated. In some embodiments, the circular construct is formed using any combination of the foregoing. In some embodiments, the ligation is DNA template ligation. In some embodiments, the ligation is RNA template ligation. In some embodiments, a sprint is provided as a template for ligation.

[0130] The nature of the ligation reaction depends on the structural components of the probe used. In some embodiments, the 3' and 5' ends of a cyclizable probe or probe set can be ligated using an analyte (e.g., RNA) as a template. In some embodiments, the 3' and 5' ends are ligated without gap filling prior to ligation. In some embodiments, gap filling precedes the ligation of the 3' and 5' ends. The gap can be 1, 2, 3, 4, 5, or more nucleotides. In some embodiments, the ligation can include enzymatic ligation, chemical ligation, template-dependent ligation, and / or template-independent ligation. In any one of the embodiments herein, the ligation can include using a ligase having RNA template DNA ligase activity and / or RNA template RNA ligase activity. In some embodiments, enzymatic ligation involves the use of a ligase (e.g., an RNA ligase, a DNA ligase). The ligase includes any one of the ligases described in ATP-dependent double-stranded polynucleotide ligase, NAD-i-dependent double-stranded DNA or RNA ligase, and single-stranded polynucleotide ligase, e.g., EC 6.5.1.1 (ATP-dependent ligase), EC 6.5.1.2 (NAD+-dependent ligase), EC 6.5.1.3 (RNA ligase). Specific examples of ligases include bacterial ligases such as E. coli DNA ligase, Tth DNA ligase, Thermococcus species (strain 9°N) DNA ligase (9°N™ DNA ligase, New England Biolabs), Taq DNA ligase, Ampligase™ (Epicentre Biotechnologies), and phage ligases such as T3 DNA ligase, T4 DNA ligase, and T7 DNA ligase, and variants thereof. In any one of the embodiments herein, the ligation can include using a ligase selected from the group consisting of Chlorella virus DNA ligase (PBCV DNA ligase), T4 RNA ligase, T4 DNA ligase, and single-stranded DNA (ssDNA) ligase.In any one of the embodiments of this specification, ligation may involve using PBCV-1 DNA ligase or its variants or derivatives, and / or T4 RNA ligase 2 (T4 Rnl2) or its variants or derivatives. In some embodiments, the ligase is T4 RNA ligase. In some embodiments, the ligase includes splintR ligase. In some embodiments, the ligase is a single-stranded DNA ligase. In some embodiments, the ligase is T4 DNA ligase. In some embodiments, the ligase is a ligase having DNA splint DNA ligase activity. In some embodiments, the ligase is a ligase having RNA splint DNA ligase activity.

[0131] In some aspects, high-fidelity ligases such as thermostable DNA ligases (e.g., Taq DNA ligase) are used. Thermostable DNA ligases are active at high temperatures and allow for further discrimination by incubating the ligation at a temperature close to the melting temperature of the DNA strands. High-fidelity ligation can be achieved through the inherent selectivity of the ligase active site and a balanced combination of conditions to reduce the incidence of annealed mismatched dsDNA.

[0132] In some embodiments, a removing step is performed to remove molecules that have not specifically hybridized. In some embodiments, a removing step is performed to remove unligated probes. In some embodiments, the removing step is performed after ligation and before amplification. A washing step can be performed at any point during the process to remove non-specifically bound probes, unligated probes, etc. In some embodiments, the circularized probe remains specifically hybridized to the analyte after the removing step.

[0133] In some cases, primer oligonucleotides are added for amplification. In some cases, the primer oligonucleotides are added with a cyclizable probe or probe set. In some cases, the primer oligonucleotides are added before or after the cyclizable probe or probe set contacts the sample. In some cases, the primer oligonucleotides for amplification of the circularized nucleic acid molecule can include a sequence complementary to a nucleic acid (e.g., cDNA or mRNA), as well as a sequence complementary to the cyclizable probe that hybridizes to the nucleic acid. In some embodiments, a washing step is performed to remove any unbound probes, primers, etc. In some embodiments, the washing is a stringent wash.

[0134] The primer oligonucleotides for amplification of the circularized nucleic acid molecule can include a single-stranded nucleic acid sequence having a 3' end that can be used as a substrate for a nucleic acid polymerase in a nucleic acid extension reaction. The primer oligonucleotides can include both RNA nucleotides and DNA nucleotides (e.g., in a random or designed pattern). The primer oligonucleotides can also include other natural or synthetic nucleotides described herein that can have additional functionality. The primer oligonucleotides can be from about 6 bases to about 100 bases, such as about 25 bases.

[0135] In some cases, in the presence of appropriate dNTP precursors and other cofactors, addition of DNA polymerase causes the amplification primers to be extended by replication of multiple copies of the template. The amplification step may utilize isothermal amplification or non-isothermal amplification. In some embodiments, after formation of the hybridization complex and any subsequent cyclization (e.g., ligation of padlock probes, etc.), the circular nucleic acid molecule is amplified by rolling circle amplification to produce an RCA product (e.g., an amplicon) containing multiple copies of the sequence of the circular nucleic acid molecule. See, for example, Baner et al, Nucleic Acids Research, 26:5073-5078, 1998, Lizardi et al, Nature Genetics 19:226, 1998, Mohsen et al., Acc Chem Res. 2016 November 15;49(11):2540-2550, Schweitzer et al. Proc. Natl Acad. Sci. USA 97:10113-119, 2000, Faruqi et al, BMC Genomics 2:4, 2000, Nallur et al, Nucl. Acids Res. 29:e118, 2001, Dean et al. Genome Res. 11:1095-1099, 2001, Schweitzer et al, Nature Biotech. 20:359-365, 2002, U.S. Patent Nos. 6,054,274, 6,291,187, 6,323,009, 6,344,329, and 6,368,801, all of which are incorporated herein by reference.

[0136] In some embodiments, the rolling circle amplification product is generated using a polymerase selected from the group consisting of Phi29 DNA polymerase, Phi29-like DNA polymerase, M2 DNA polymerase, B103 DNA polymerase, GA-1 DNA polymerase, phi-PRD1 polymerase, Vent DNA polymerase, Deep Vent DNA polymerase, Vent(exo-)DNA polymerase, KlenTaq DNA polymerase, DNA polymerase I, the Klenow fragment of DNA polymerase I, DNA polymerase III, T3 DNA polymerase, T4 DNA polymerase, T5 DNA polymerase, T7 DNA polymerase, Bst polymerase, rBST DNA polymerase, N29 DNA polymerase, TopoTaq DNA polymerase, T7 RNA polymerase, SP6 RNA polymerase, T3 RNA polymerase, and variants or derivatives thereof. In some embodiments, the polymerase is Phi29 DNA polymerase.

[0137] In some embodiments, the polymerase comprises a modified recombinant Phi29-type polymerase. In some embodiments, the polymerase comprises a modified recombinant Phi29, B103, GA-1, PZA, Phi15, BS32, M2Y, Nf, G1, Cp-1, PRD1, PZE, SF5, Cp-5, Cp-7, PR4, PR5, PR722, or L17 polymerase. In some embodiments, the polymerase comprises a modified recombinant DNA polymerase having at least one amino acid substitution or combination of substitutions compared to wild-type Phi29 polymerase. Exemplary polymerases are described in U.S. Patent Nos. 8,257,954, 8,133,672, 8,343,746, 8,658,365, 8,921,086, and 9,279,155, all of which are incorporated herein by reference. In some embodiments, the polymerase is not directly or indirectly immobilized on a substrate such as beads or a planar substrate (e.g., a glass slide) prior to contacting the sample, although the sample may be immobilized on the substrate.

[0138] In some embodiments, the amplification is carried out at a temperature of 20°C to 60°C or about 20°C to about 60°C. In some embodiments, the amplification is carried out at a temperature of 30°C to 40°C or about 30°C to about 40°C. In some aspects, the amplification is carried out at a temperature of 25°C to 50°C or about 25°C to about 50°C, such as 25°C, 27°C, 29°C, 31°C, 33°C, 35°C, 37°C, 39°C, 41°C, 43°C, 45°C, 47°C, or 49°C, or about 25°C, 27°C, 29°C, 31°C, 33°C, 35°C, 37°C, 39°C, 41°C, 43°C, 45°C, 47°C, or 49°C.

[0139] In any one of the embodiments herein, the RCA product can be generated in situ in a biological sample. In any one of the embodiments herein, the product can be generated using linear RCA, branched RCA, dendritic RCA, or any combination thereof.

[0140] In some embodiments, the RCA product comprises multiple copies of one or more identifier sequences or their complementary sequences. In some embodiments, the RCA product comprises multiple copies of one or more priming sites or their complementary sequences. In some aspects, one or more identifier sequences or their complementary sequences are located downstream of one or more priming sequences or their complementary sequences.

[0141] In some embodiments, during the amplification step, modified nucleotides can be added to the reaction to incorporate the modified nucleotides into the amplification product (e.g., nanoball). Examples of modified nucleotides include amine-modified nucleotides. For example, in some embodiments for anchoring or cross-linking the generated amplification product (e.g., nanoball) to a scaffold, to a cellular structure, and / or to other amplification products (e.g., other nanoballs). In some embodiments, the amplification product contains modified nucleotides such as amine-modified nucleotides. In some embodiments, the amine-modified nucleotide reacts with an N-hydroxysuccinimide moiety of acrylic acid. Other examples of amine-modified nucleotides include, but are not limited to, 5-aminoallyl-dUTP moiety modification, 5-propynylamino-dCTP moiety modification, N 6 -6-aminohexyl-dATP moiety modification, or 7-deaza-7-propynylamino-dATP moiety modification. In some embodiments, the modified nucleotides include azide and / or alkyne base modifications, dibenzylcyclooctyl (DBCO) modifications, vinyl modifications, base modifications such as trans-cyclooctene (TCO).

[0142] In some embodiments, the primer extension reaction mixture can contain deoxynucleoside triphosphates (dNTPs) or derivatives, variants, or analogs thereof. In some embodiments, the primer extension reaction mixture can contain catalytic cofactors of the polymerase. In any of the preceding embodiments, the primer extension reaction mixture can contain catalytic dication such as Mg 2+ and / or Mn 2+ etc.

[0143] In some embodiments, amplification products (e.g., RCA products) can be anchored to a polymer matrix. The amplification products are generally immobilized within the matrix at the location of the nucleic acid being amplified, thereby creating localized colonies of amplicons. The amplification products can be immobilized within the matrix by steric factors. The amplification products can also be immobilized within the matrix by covalent or non-covalent bonds. In this way, the amplification products can be considered to be attached to the matrix. By being immobilized to the matrix, such as by covalent bonding or cross-linking, the size and spatial relationships of the original amplicons are maintained. By being immobilized to the matrix, such as by covalent bonding or cross-linking, the amplification products are resistant to movement or disintegration under mechanical stress.

[0144] In some embodiments, amplification products (e.g., RCA products) are copolymerized and / or covalently bonded to the surrounding matrix, thereby retaining their spatial relationships and any information unique to them. In some embodiments, the RCA products are generated from DNA or RNA within cells embedded in the matrix. In some embodiments, the RCA products are also functionalized to form covalent bonds to the matrix that preserves their spatial information intracellularly, thereby providing an intracellular localization distribution pattern. In some embodiments, the methods provided involve embedding the RCA products in the presence of hydrogel subunits to form one or more hydrogel-embedded amplification products. In some embodiments, the hydrogel-histochemistry described involves covalently bonding nucleic acids to hydrogels synthesized in situ for tissue clearing, enzyme diffusion, and multi-cycle sequencing or probe hybridization, which existing hydrogel histochemistry methods cannot do. In some embodiments, amine-modified nucleotides are included in the amplification step (e.g., RCA) to enable embedding of the amplification products in the tissue-hydrogel environment, are functionalized with acrylamide moieties using N-hydroxysuccinimide ester of acrylic acid, and copolymerized with acrylamide monomers to form a hydrogel.

[0145] III. In Situ Sequencing of Identifier Arrays In some embodiments, methods are provided herein for analyzing a cell or tissue sample, the method comprising contacting the sample with a probe that binds directly or indirectly to an analyte at a position in the sample, the probe or its product (e.g., an amplification product such as an RCA product) comprising a priming site and an identifier sequence that is present in, related to, corresponding to, and / or identifying the analyte. In some embodiments, the method further comprises contacting the sample with a sequencing primer configured to hybridize to the priming site and be ligated by a polymerase and nucleotides for base-by-base sequencing of the identifier sequence or a portion thereof in the probe or its product.

[0146] In some embodiments, the method includes contacting a biological sample with nucleotides in successive cycles (e.g., a nucleotide mix comprising nucleotides containing different bases for each cycle), a complex being formed in each cycle, the complex comprising i) a sequencing primer or an extension product thereof hybridized to a probe or a product thereof, ii) a polymerase, and iii) cognate nucleotides base-pairing with nucleotides in the identifier sequence. In some embodiments, in each cycle, the method includes detecting that a signal (on-signal) and / or the absence of a signal (off-signal) associated with the cognate nucleotides and / or polymerase is detected at the position, the on-signal, off-signal, or a combination thereof corresponding to the cognate nucleotides in the identifier sequence and the bases in the corresponding nucleotides. In some embodiments, the method includes generating a signal code sequence comprising a signal code corresponding to the on-signal, off-signal, or a combination thereof in successive cycles at the position, thereby detecting the identifier sequence at the position in the biological sample. Since the identifier sequence can be present in, associated with, corresponding to, and / or identifying the analyte, detection of the identifier sequence at the position can be used to identify the analyte at the position. For example, if the identifier sequence uniquely identifies an analyte out of a plurality of analytes, detection of the identifier sequence at the position identifies the analyte at the position. In some embodiments, a combination of the analyte and / or an identifier sequence in the probe or a product thereof identifies the analyte, and each identifier sequence can be sequenced to decode the combination and identify the analyte.

[0147] The present disclosure provides a method for detecting multiple analytes in situ in a cell sample or a tissue sample. Probes (e.g., a first and a second probe) designed to reduce optical crowding of signals in a biological sample are provided herein. In some embodiments, the first and second probes contact first and second analytes (e.g., nucleic acid molecules such as mRNA) in the biological sample (FIG. 1(101)). In some embodiments, the first and second probes directly or indirectly bind to the first and second analytes, respectively, at first and second positions in the sample. In some aspects, the first and second probes are primary probes that directly bind to their corresponding analytes at positions in the sample. In some aspects, the first and second probes directly or indirectly bind to the first and second primary probes in the sample. In some embodiments, the first and second probes can be circularizable probes or probe sets. In some embodiments, the first and second probes can be linear probes or probe sets. In some embodiments, the first and second probes are amplified in situ to generate first and second products of the first and second probes (e.g., first and second rolling circle amplification (RCA) products). In some embodiments, each of the first and second probes, or their products, includes i) a priming site for a sequencing primer, and ii) an identifier sequence related to the corresponding analyte in the sample. For example, the first probe or its product can include, at a first position in the sample, for example, i) a first priming site for a first sequencing primer, and ii) a first identifier sequence related to the first analyte. The priming site for a sequencing primer is a site for initiating a sequencing reaction such as a synthesis-based sequencing (SBS) or a binding-based sequencing (SBB) reaction. In some embodiments, the first and second identifier sequences are related to the first and second analytes in the sample. Decoding of the first and second identifier sequences can enable identification of the first and second analytes in the sample, and their respective positions.In some embodiments, the identifier array is a barcode array that is related to the analyte or corresponds to the identity of the analyte, and decoding the barcode array enables identification of the analyte in addition to revealing the spatial location of the analyte in the biological sample.

[0148] In some embodiments, the method includes performing SBS or SBB of the first and second identifier arrays (FIG. 1(102)). In some embodiments, the SBS or SBB reaction is performed using first and second sequencing primers. The first and second sequencing primers bind to first and second priming sites on the first and second probes or their products, respectively. In some embodiments, the method includes performing a cyclic series of nucleotide incorporation steps, such as incorporation of A, T, C, or G in SBS. In some embodiments, the method includes performing a cyclic series of nucleotide binding steps, such as binding of A, T, C, or G in SBB.

[0149] In some embodiments, the method includes detecting a signal, such as an on-signal, or the absence of a signal, such as an off-signal, during the sequencing step (FIG. 1(103)). In some embodiments, the method includes detecting first and second on-signals and / or off-signals at first and second positions in the sample. In some embodiments, the signal detected in a particular cycle corresponds to a signal code for the cycle. In some embodiments, the signal code corresponds to a signal of a first color, a signal of a second color, a signal of a third color, or the absence of a signal. In some aspects, the first, second, and third colors are different. In some aspects, the first, second, and third colors are the same. In some aspects, the SBS or SBB reaction is repeated in one or more consecutive cycles to generate a series of signal codes, each signal code corresponding to a signal (on-signal) or the absence of a signal (off-signal), or a combination of on-signals and / or off-signals (FIG. 1(104)).

[0150] In some embodiments, a series of signal codes are detected at different positions (e.g., a first and a second position) in a sample. In some embodiments, the nucleotides in the first barcode array detected in a particular cycle correspond to a signal code that includes an on-signal, and the corresponding nucleotides in the second barcode array detected in the particular cycle correspond to a signal code that includes an off-signal. In some embodiments, one or more pairs of corresponding nucleotides in the first and second barcode arrays detected in the same cycle are selected to reduce optical crowding of the signals detected in the cycle.

[0151] In some aspects, each series of signal codes includes a signal code array. For example, a series of signal codes detected in consecutive cycles at a first position includes a first signal code array. A series of signal codes detected in consecutive cycles at a second position includes a second signal code array. In some aspects, the first and second signal code arrays are used to decode identifier arrays on the first and second probes or products thereof. In some embodiments, the method includes detecting first and second identifier arrays based on the first and second signal code arrays (FIG. 1(105)). In some aspects, the identifier array on the probe or product thereof is a barcode array. In some aspects, the barcode array on the probe or product thereof is assigned to an analyte bound to the probe. In some aspects, in situ detection of the barcode array by SBS or SBB enables identification of analytes in parallel while simultaneously reducing optical crowding of the signals during multiple decoding cycles.

[0152] FIG. 2 shows an exemplary diagram of sequencing by synthesis of an identifier array on a probe or product thereof (e.g., an RCA product). An exemplary hybridization complex including the identifier array is also shown.

[0153] The probe or product may include: i) a priming site for a sequencing primer, and ii) an identifier sequence. In some embodiments, the priming site is a site for hybridization of a sequencing primer for initiation of sequencing. In some embodiments, the identifier sequence is a barcode sequence. In some embodiments, the identifier sequence or barcode sequence is associated with a corresponding analyte in the sample. Decoding of the identifier sequence or barcode sequence on the probe or its product may enable identification of the analyte bound to the probe. In some embodiments, the SBS reaction is performed by contacting a biological sample with a nucleotide mixture in successive cycles (Figure 2). In some embodiments, in each cycle, a complex is formed that includes: i) a sequencing primer or its extension product hybridized to the priming site, ii) a polymerase, and iii) cognate nucleotides that base pair with nucleotides in the identifier sequence. In some aspects, the sequencing primer binds to the priming site on the probe or its product and generates an extension product as the sequencing reaction proceeds. In some embodiments, the sequencing reaction includes performing a periodic series of nucleotide incorporation steps. Cognate nucleotides A, T, C, or G bind to their corresponding nucleotides in the identifier sequence and are incorporated into the sequencing primer or extension product. In some embodiments, the cognate nucleotides are incorporated into the sequencing primer or extension product by a polymerase. In some embodiments, the nucleotide mixture that contacts the biological sample includes a mixture of fluorescently labeled nucleotides and non-fluorescently labeled nucleotides. In some aspects, the nucleotide mixture includes three fluorescently labeled nucleotides and one non-fluorescently labeled nucleotide. In some aspects, the nucleotide mixture includes two fluorescently labeled nucleotides and two non-fluorescently labeled nucleotides. For example, as shown in Figure 2, nucleotides T, A, and G are labeled with a fluorescent moiety while nucleotide C is unlabeled.In some embodiments, signals (on-signals) and / or the absence of signals (off-signals) associated with cognate nucleotides are detected at specific locations in a biological sample. In some embodiments, an on-signal, an off-signal, or a combination thereof corresponds to a base in a cognate nucleotide and a corresponding nucleotide in an identifier sequence. For example, an on-signal can be detected as a result of incorporation of a labeled nucleotide T, A, or G, while an off-signal can be detected as a result of incorporation into a sequencing primer or extension product of an unlabeled nucleotide C. In some aspects, the identifier sequence is detected based on a signal code sequence generated by SBS. In some aspects, detection of the identifier sequence shown in FIG. 2 enables identification of an analyte while simultaneously reducing optical crowding of signals.

[0154] A. Dark Bases and Nucleotide Mixes In some embodiments, provided herein are methods that include determining, in situ, the sequence of an identifier sequence in a cell or tissue sample, the identifier sequence including bases that do not give rise to a detectable signal in a sequencing cycle for each corresponding base, e.g., these bases are “dark.” In any particular identifier sequence, the dark bases can be the same (e.g., all G) or different (the first dark base is G and the second dark base is C). In any particular identifier sequence, any two or more dark bases can be contiguous (nucleotide residues are directly linked by phosphodiester bonds) or non-contiguous (e.g., nucleotide residues are separated by one or more nucleotide residues, at least one of which includes a non-dark base).

[0155] In some embodiments, multiple different identifier arrays are sequenced in situ in a cell or tissue sample using a base-by-base sequencing method, such as SBS or SBB, and each of the multiple different identifier arrays includes multiple dark bases such that optical crowding is limited in many (e.g., most) of the base-by-base sequencing cycles to facilitate accurate and efficient signal detection and decoding.

[0156] In some embodiments, a method of analyzing a biological sample is provided herein, the method comprising: a) contacting a biological sample with a first probe and a second probe, wherein the biological sample is a cell or tissue sample, the biological sample includes a first analyte and a second analyte at a first position and a second position in the biological sample, the first probe and the second probe are directly or indirectly bound to the first analyte and the second analyte, respectively, and the first probe or its product includes i) a first priming site for a first sequencing primer, and ii) a first identifier array associated with the first analyte, and the second probe or its product includes i) a second priming site for a second sequencing primer, and ii) a second identifier array associated with the second analyte; b) performing base-by-base sequencing of the first and second identifier arrays using the first and second sequencing primers, thereby generating a first signal code sequence and a second signal code sequence, each including a signal code corresponding to a signal (on signal), absence of a signal (off signal), or a combination thereof, respectively, detected in consecutive cycles at the first position and the second position, and in one or more of the consecutive cycles, an on signal is detected at the first position and an off signal is detected at the second position; and c) detecting the first and second identifier arrays in the biological sample based on the first and second signal code sequences.

[0157] In some embodiments, the first and second identifier sequences are different and are each associated with different first and second analytes. In some embodiments, the first and second identifier sequences are different and are associated with the same analyte. In some embodiments, the first and second identifier sequences may include a barcode sequence or its complement. In some embodiments, the first barcode sequence and the second barcode sequence may be assigned to the first and second analytes, respectively.

[0158] In some embodiments, the assignment of codewords (corresponding to physical barcodes or identifier arrays) to each of a plurality of analytes is performed using any suitable scheme that minimizes optical crowding in the fluorescence images used to decode the identifier arrays. Methods for designing codebooks and assigning codewords to analytes to minimize optical crowding are described in US63 / 317,842, filed Mar. 8, 2022, and International Patent Application No. PCT / US2023 / 063866, filed Mar. 7, 2023, both entitled "in situ Code Design Methods for Minimizing Optical Crowding", the contents of which are hereby incorporated by reference in their entirety. In some aspects, the assignment of an identifier array (corresponding to a codeword) to each of a plurality of analytes is optimized for each of the analytes. In some embodiments, for example, the assignment of an identifier array to each of a plurality of analytes is performed based on expression data for the plurality of analytes. For example, the expression data can include whole transcriptomes clustered according to cell type, single cell reference gene expression data. In some embodiments, codewords / identifier arrays containing the maximum number of off-bits can be sequentially assigned to analytes (e.g., genes) proceeding in descending order of the number of off-bits in the codeword and the highest expression levels of the analytes (e.g., genes) across all cell types. In some embodiments, specifying a codeword (corresponding to an identifier array) for a gene involves calculating a metric (e.g., worst density or maximum predicted density) that would be achieved by assigning any of the currently available codewords to the current gene. Next, the codeword that results in the lowest worst density is assigned to the particular analyte. In some cases, the worst density is the expression density of the (cell type, code bit) pair having the highest total expression density of the genes that are on at that code bit.

[0159] In some cases, the codewords (corresponding to the identifier arrays) can be assigned to a plurality of analytes according to a decision rule (e.g., a minimax decision rule) designed to minimize the maximum predicted density of on-signals over a series of images acquired over one or more detection channels (e.g., 1, 2, 3, 4, or more than 4 detection channels) during a plurality of array determination cycles. For example, in some cases, the codewords of the codebook are for a series of images acquired during a plurality of array determination cycles (e.g., for 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or more than 20 array determination cycles) to minimize, for each image, the maximum predicted density of on-signals (corresponding to the on-bits of the target analyte-related codewords) detected (e.g., different images of the biological sample are acquired or received for 1, 2, 3, 4, or more than 4 determination channels in each array determination cycle) according to a minimax decision rule, for a plurality of analytes (e.g., at least 5, at least 10, at least 20, at least 40, at least 60, at least 80, at least 100, at least 200, at least 400, at least 600, at least 800, at least 1,000, at least 2,000, at least 4,000, at least 6,000, at least 8,000, at least 10,000, at least 20,000, at least 40,000, at least 60,000, at least 80,000, at least 100,000, at least 200,000, at least 400,000, at least 600,000, at least 800,000, at least 1,000,000, at least 2×10 6 , at least 3×10 6 , at least 4×10 6 , at least 5×10 6 , at least 6×10 6 , at least 7×10 6 , at least 8×10 6 , at least 9×10 6 , at least 10 7 , at least 10 8 , at least 109 or 10 9 can be assigned to (more than 10 analytes).

[0160] In some cases, the codewords in the codebook can be assigned to multiple analytes according to a decision rule that ensures that the on-bits are distributed more or less evenly across the multiple sequencing cycles and detection channels used for sequencing. For example, in some cases, the codewords of the codebook are such that the total number of on-signals detected in a given image for a given sequencing cycle is within ±5%, ±10%, ±15%, ±20%, or ±25% of the average number of on-signals detected per image (e.g., different images of a biological sample are acquired or received for 1, 2, 3, 4, or more than 4 determination channels in each sequencing cycle) for a series of images acquired over multiple sequencing cycles (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or more than 20 sequencing cycles) used to sequence multiple analytes (e.g., at least 5, at least 10, at least 20, at least 40, at least 60, at least 80, at least 100, at least 200, at least 400, at least 600, at least 800, at least 1,000, at least 2,000, at least 4,000, at least 6,000, at least 8,000, at least 10,000, at least 20,000, at least 40,000, at least 60,000, at least 80,000, at least 100,000, at least 200,000, at least 400,000, at least 600,000, at least 800,000, at least 1,000,000, at least 2×10 6 at least 3×10 6 at least 4×10 6 at least 5×10 6 at least 6×10 6 at least 7×10 6 at least 8×10 6 at least 9×106 、at least 10 7 、at least 10 8 、at least 10 9 、or 10 9 that can be assigned to (more than 10 analytes).

[0161] In some cases, the codewords in the codebook ensure that the number of target analytes visible in a given image (e.g., the number of target analytes for which the corresponding codeword has an on-bit in a given image) is distributed more or less evenly across the multiple sequencing cycles and detection channels used for sequencing. For example, in some cases, the codewords of the codebook ensure that the number of target analytes visible (e.g., having the corresponding codeword with an on-bit) in a given image for a given sequencing cycle is within ±5%, ±10%, ±15%, ±20%, or ±25% of the average number of target analytes detected per image (e.g., different images of the biological sample are acquired or received for 1, 2, 3, 4, or more than 5 determination channels in each cycle) for a series of images acquired over the multiple sequencing cycles (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or more than 20 cycles) used to sequence a plurality of analytes (e.g., at least 5, at least 10, at least 20, at least 40, at least 60, at least 80, at least 100, at least 200, at least 400, at least 600, at least 800, at least 1,000, at least 2,000, at least 4,000, at least 6,000, at least 8,000, at least 10,000, at least 20,000, at least 40,000, at least 60,000, at least 80,000, at least 100,000, at least 200,000, at least 400,000, at least 600,000, at least 800,000, at least 1,000,000, at least 2×10 6 、at least 3×106 , at least 4×10 6 , at least 5×10 6 , at least 6×10 6 , at least 7×10 6 , at least 8×10 6 , at least 9×10 6 , at least 10 7 , at least 10 8 , at least 10 9 , or 10 9 can be assigned to (more than analyte).

[0162] In some cases, the codewords are designed to use a minimax decision rule (e.g., designed to minimize the maximum predicted density of on-signals across a series of images), an average on-signal decision rule (e.g., ensuring that the total number of on-signals detected in a given image for a given array determination cycle is within ±5%, ±10%, ±15%, ±20%, or ±25% of the average number of on-signals detected per image), an average target number decision rule (e.g., ensuring that the number of target analytes visible (e.g., having the corresponding codeword with an on-bit) in a given image for a given array determination cycle is within ±5%, ±10%, ±15%, ±20%, or ±25% of the average number of target analytes detected per image), or any combination thereof, to be assigned to the analyte.

[0163] In some cases, the determination rule (or determination process) can be implemented in an iterative fashion. For example, in some cases, one or more codewords can be ranked according to the codeword weight (e.g., the total number of on bits within a given codeword), one or more analytes can be ranked according to the predicted density, and one or more ranked codewords can be assigned to one or more ranked analytes using an iterative process that is repeated for each of the one or more analytes in decreasing order of the maximum predicted density. The iterative process includes calculating the predicted density of the on signal for all combinations of the remaining unassigned codewords and analytes across a series of images, selecting a codeword from the remaining unassigned codewords that minimizes the predicted density of the on signal across a series of images, and assigning the selected codeword to an analyte. In some cases, the iterative process can further include reviewing the previous assignment of codewords to analytes and changing the selected codeword for the current analyte to minimize the predicted density of the on signal across a series of images for the analyte to which the codeword was previously assigned.

[0164] In some embodiments, the optimized assignment of codewords to analytes is such that the codewords are assigned to the corresponding analytes according to decision rules based on prior knowledge of the abundance or distribution of the target analytes in a given biological sample, e.g., expression data for the analytes, and the weight of the codewords corresponding to highly expressed analytes is reduced. In some cases, the codewords are assigned to the corresponding analytes according to decision rules based on single-cell expression data for the target analytes in, e.g., clustered cell types, and the weight of the codewords corresponding to highly expressed target analytes can be reduced, where the clustered cell types represent the distribution of cell types found in the biological sample. In some cases, the expression data for one or more target analytes includes bulk gene expression data, bulk protein expression data, spatial gene expression data, spatial protein expression data, single-cell gene expression data, single-cell protein expression data, or any combination thereof.

[0165] For example, assume that single-cell expression data (e.g., single-cell gene expression data or single-cell protein expression data) is available for a biological sample of interest (e.g., a tissue sample of interest), and the expression data is clustered according to cell type clusters, each having an average gene or protein expression profile, and the clustered cell types represent the distribution of cell types found in the biological sample. In some cases, the clustered single-cell expression data can provide the best prior information about the expression profiles likely to be observed in in situ experiments, and the density of labeled features is likely to mimic the expression profiles.

[0166] Given a codebook and the assignment of analytes (e.g., gene transcripts) to the codewords within the codebook, the single-cell type expression data can be used to determine the expected density of labeled spots observed for each cell type in each sequencing cycle and detection channel of the sequencing process.

[0167] In some embodiments, as described herein, the codewords for each gene transcript are designed to distribute the number (or density) of on-features across multiple sequencing cycles and detection channels, and to explicitly avoid this situation and reduce the occurrence of overcrowded cycles by including the absence of a signal (off-signal). As described above, in some cases, the codewords can be selected / assigned according to a decision rule (e.g., a minimax decision rule) that minimizes the maximum predicted density of on-signals in the images acquired during the periodic sequencing process. In some cases, the codewords can be selected / assigned according to a decision rule (e.g., an average on-signal decision rule) that ensures that the total number of on-signals detected in the image for a given sequencing cycle is within ±5%, ±10%, ±15%, ±20%, or ±25% of the average number of on-signals detected per image. In some cases, the codewords can be selected / assigned according to a decision rule (e.g., an average target number decision rule) that ensures that the number of target analytes visible (e.g., having the corresponding codeword with an on-bit) in a given image for a given sequencing cycle is within ±5%, ±10%, ±15%, ±20%, or ±25% of the average number of target analytes detected per image. In any of these cases, the decision rule can further include an assignment based on expression data.

[0168] In some cases, for example, the codewords may be ranked according to the codeword weight, the analytes may be ranked according to the maximum expression level across the clustered cell types, and the ranked codewords may be assigned to the ranked analytes using an iterative process that is repeated for each of the analytes in decreasing order of the maximum expression level. The iterative process includes calculating the predicted density of the on-signal for all combinations of the remaining unassigned codewords and analytes across a series of images, selecting from the remaining unassigned codewords the codeword that minimizes the predicted density of the on-signal across a series of images, and assigning the selected codeword to the analyte. In some cases, the iterative process may further include reviewing the previous assignment of codewords to analytes and changing the selected codeword for the current analyte to minimize the predicted density of the on-signal across a series of images for the analyte to which the codeword was previously assigned.

[0169] In some cases, for example, the codewords may be ranked according to the codeword weight (i.e., the total number of on-bits within a given codeword), and the detected analytes may be ranked according to their corresponding single-cell expression data or predicted density in the sample. In some cases, the lowest-ranked codewords may then be assigned to the highest-ranked analytes. In some cases, an algorithm for assigning codewords, for example, to gene transcripts, may be developed, and the assignment algorithm may be optimized to minimize optical crowding and distribute the total number of on-bits or density across multiple sequencing cycles and detection channels.

[0170] In some embodiments, the nucleotides in the first identifier array (e.g., barcode array) detected in a particular cycle correspond to a signal code including an on signal, and the corresponding nucleotides in the second identifier array (e.g., barcode array) detected in the particular cycle correspond to a signal code including an off signal. In some embodiments, one or more pairs of corresponding nucleotides in the first and second barcode arrays detected in the same cycle can be selected to reduce optical crowding of the signal detected in the cycle.

[0171] In some embodiments, the first and / or second identifier arrays, such as barcode arrays (or their corresponding codewords), can be designed such that about or at least 30% of the nucleotides therein correspond to off signals, e.g., such that they are dark bases. In some embodiments, about 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, or more of the nucleotides in the first and / or second identifier arrays (e.g., barcode arrays) are dark bases.

[0172] In some embodiments, the signals detected in each cycle of base-by-base sequencing correspond to a block of bits (e.g., a signal code) within a codeword corresponding to an identifier sequence (e.g., a first and / or second identifier sequence). The identifier sequence (e.g., a barcode sequence) can be assigned based on the codeword design and the base-dye labeling scheme during use (the latter can be the same for each decoding (imaging) cycle or different for different decoding cycles (e.g., different dark bases (corresponding to off-signals) can be used in each cycle). For example, the signals detected in each cycle of base-by-base sequencing can correspond to a 3-bit within a codeword (corresponding to an identifier sequence), or a portion thereof, and each group of 3 bits can be mapped to a particular nucleotide. The barcode sequence can be derived from the codeword design, and the base / dye assignment scheme during use can be the same in each cycle or the assignment of dark bases can be adjusted in each cycle. Table 1 provides an example of a group of 3 bits within a codeword (corresponding to an identifier sequence) that can be mapped to nucleotides using the exemplary patterns provided.

[0173]

Table 1

[0174] In some cases, the identifier is made up of 3-bit blocks drawn from 4 options. The codebook design may involve generating candidate codewords (e.g., identifier arrays) having an appropriate number of on-bits (e.g., 4, 5, or 6), and then successively identifying new valid codewords that are sufficiently separated from existing codewords generated within the codebook. As long as the codeword has a correctly valid 3-bit block and the correct number of on-bits (corresponding to the dye / signal), the on-bits can appear anywhere in the identifier. In some aspects, a plurality of “ON” cycles / bits that each codeword must have are determined, and the codewords can be designed accordingly. For example, each codeword can use 3, 4, 5, 6, 7, or 8 “on” cycles / bits. In some examples, each codeword uses 5 “ON” cycles / bits. In some cases, there may be certain constraints on the codewords that make it easier to identify “actual” features, for example, by requiring that each codeword have “on” bits that emit light in at least a certain number of colors. For example, each codeword may need to have “on” bits that emit light in at least 2, 3, or 4 colors. In some cases, this makes it possible to distinguish the signal from background fluorescence sources that tend to remain monochromatic. In some cases, the barcode array is derived from a codeword assignment scheme (e.g., optimization of codeword assignment), and the base and corresponding dye assignment scheme are used to designate nucleotides as “dark”. In some cases, a nucleotide can be designated as dark for a given cycle, or the dark nucleotides can move around.

[0175] In some cases, the codewords are made up of 3-bit blocks drawn from four options for the dye label (e.g., as shown in Table 1), and each bit within the block corresponds to a different detection (color) channel. The codebook design may involve generating candidate codewords (e.g., identifier arrays) having an appropriate number of on-bits (e.g., 2, 3, 4, 5, or 6), and then successively identifying newly valid codewords that are sufficiently separated from the existing codewords generated within the codebook (e.g., according to an edit distance such as the Hamming distance). The on-bits can appear anywhere within the codeword as long as the codeword has the correct number of 3-bit blocks (or codeword bits) and the correct number of on-codeword bits (corresponding to the detection of on-signals in a given sequencing cycle). In some embodiments, multiple on-codeword bits can be specified such that each codeword must have a specified number of on-codeword bits (corresponding to the number of sequencing cycles in which an on-signal is detected), and the codewords can be designed accordingly. For example, in some cases, each codeword can use 1, 2, 3, 4, 5, 6, 7, or 8 on-bits. In some examples, each codeword uses 5 on-bits. In some cases, there can be certain constraints on the codeword design that make it easier to identify "real" features, for example, by requiring that each codeword includes on-bits that emit light in at least a certain number of colors. For example, each codeword may need to have on-bits that emit light in at least 2, 3, or 4 colors. In some cases, this makes it possible to distinguish the signal from background fluorescence sources that tend to remain monochromatic.

[0176] In some cases, an identifier array, such as a barcode array, is derived from a codeword assignment scheme (e.g., by optimizing codeword assignment to minimize optical crowding), and bases and corresponding dye-labeling schemes can be used to designate a given nucleotide as "dark." In some cases, a nucleotide can be designated as dark for a given sequencing cycle, or the designated dark nucleotides can be different for different cycles.

[0177] In some embodiments, a plurality of different identifier arrays (or barcode arrays) can be detected in a biological sample, and each different identifier array can be detected at one or more positions in the biological sample. In some embodiments, at least 25%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, or more of the nucleotides of each of the different identifier arrays can be dark bases. In some embodiments, at least 25%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, or more of the nucleotides of at least 40% of the different identifier arrays can be dark bases. In some embodiments, at least 25%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, or more of the nucleotides of at least 50% of the different identifier arrays can each be dark bases. In some embodiments, at least 25%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, or more of the nucleotides of at least 60% of the different identifier arrays can be dark bases. In some embodiments, at least 25%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, or more of the nucleotides of at least 70% of the different identifier arrays can each be dark bases. In some embodiments, at least 25%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, or more of the nucleotides of at least 80% of the different identifier arrays can each be dark bases. In some embodiments, at least 25%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, or more of the nucleotides of at least 90% of the different identifier arrays can each be dark bases. In some embodiments, at least 25%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, or more of the nucleotides of at least 95% of the different identifier arrays can each be dark bases.

[0178] In some embodiments, each signal code can correspond to a signal of a first color, a signal of a second color, a signal of a third color, or the absence of a signal (e.g., corresponding to a dark base), and the first, second, and third colors are different. The dark base can be any one or two of A, T, C, and G. The dark bases in the identifier sequence can be "detected" by using a nucleotide mix in which the cognate nucleotides containing bases complementary to the dark bases are not detectably labeled. In some embodiments, the biological sample is contacted with a first nucleotide mix in which the nucleotides containing the first base are not detectably labeled while the nucleotides containing bases other than the first base are each labeled with one or more detectable labels, and the biological sample is then contacted with subsequent nucleotide mixes in which the nucleotides containing the subsequent bases are not detectably labeled while the nucleotides containing bases other than the subsequent bases are each labeled with one or more detectable labels.

[0179] In some embodiments, the same nucleotide mix can be used in multiple sequencing cycles, e.g., the subsequent bases can be the same as the first base. Any two or more of the multiple sequencing cycles can be consecutive or non - consecutive. For example, two or more cycles can be separated by one or more cycles, with or without using the same nucleotide mix as in the two or more cycles.

[0180] In some cases, for example, G in an identifier array (e.g., a barcode array) can be designated as a dark base, and a nucleotide mix containing unlabeled C nucleotides can be used in two or more or all of the sequencing cycles, while the nucleotides containing A, T, or G in the nucleotide mix are detectably labeled. In some embodiments, the nucleotides assigned as dark bases can be alternating in different cycles. In some cases, the assignment of different nucleotides as dark bases can avoid long homopolymer runs of the same base in the identifier array. For example, G in the identifier array can be designated as a dark base in cycles 1, 5, 9, T in the identifier array can be designated as a dark base in cycles 2, 6, 10, C in the identifier array can be designated as a dark base in cycles 3, 7, 11, and A in the identifier array can be designated as a dark base in cycles 4, 8, 12. In some embodiments, two nucleotides assigned as dark bases can be alternating in different cycles. In some embodiments, by assigning different dark bases in different cycles (e.g., the dark bases can be G in cycles 1, 5, 9, T in cycles 2, 6, 10, C in cycles 3, 7, 11, and A in cycles 4, 8, 12), long homopolymer runs of the same base in the barcode array can be avoided.

[0181] In some cases, the nucleotide mix contains unlabeled C nucleotides, as well as A, T, and G nucleotides labeled with fluorophores of three different colors, e.g., red for T, blue for A, and green for G, such that A, T, and C in the identifier array are each associated with red, blue, and green signals (all on-signals), while the dark base G in the identifier array is associated with an off-signal. Two exemplary identifier (e.g., barcode) arrays are shown in Table 2B below, along with the signals detected using three-color chemistry (dark is indicated by "-").

[0182]

Table 2B

[0183] In some cases, the nucleotide mix includes unlabeled C nucleotides, and two different colors, e.g., red for T and green for A (each nucleotide containing A is labeled with a red fluorophore and a green fluorophore), and A, T, and G nucleotides labeled with a green fluorophore for G, such that A, T, and C in the identifier sequence are associated with red, red and green, and green signals (all on-signals), respectively, while the dark base G in the identifier sequence is associated with an off-signal. Table 2A below shows the signals detected using two-color chemistry to sequence two exemplary barcode sequences.

[0184]

Table 2A

[0185] In some cases, the nucleotide mix includes unlabeled C nucleotides and A, T, and G nucleotides labeled with a single-color fluorophore. For example, each sequencing cycle can include two chemical steps and two imaging steps. The first chemical step exposes the sample to a mixture of nucleotides having fluorescently labeled A and T nucleotides, where the C and G nucleotides in the mixture are unlabeled, but the C nucleotides include a functional group or moiety for attaching a fluorescent label thereto. During the first imaging step, the signal or its absence at multiple positions in the sample is detected (Image 1). The second chemical step removes the fluorescent label from incorporated A nucleotides and adds a fluorescent label to incorporated C nucleotides. In both chemical steps, the G nucleotides are unlabeled. The signal or its absence at multiple positions in the sample is detected again (Image 2). The combination of Image 1 and Image 2 is processed to identify which base is incorporated at each position. For example, the signal code (Image 1 / Image 2) in a sequencing cycle can be on / off for T, off / on for G, on / on for A, and off / off for C in an identifier sequence (e.g., a barcode sequence). In some embodiments, the nucleotides in the first barcode sequence detected in a particular cycle correspond to an on signal (e.g., on / off for T, off / on for G, on / on for A), and the corresponding nucleotides in the second barcode sequence detected in the particular cycle correspond to only an off signal (e.g., off / off for C).

[0186] In some embodiments, different nucleotide mixes can be used in multiple sequencing cycles. For example, a first nucleotide mix can include nucleotides that contain a first base that is not detectably labeled, while the nucleotides that contain bases other than the first base in the first nucleotide mix are each labeled with one or more detectable labels (e.g., one, two, or three different colors), and a subsequent nucleotide mix can include nucleotides that contain a subsequent base that is not detectably labeled, while the nucleotides that contain bases other than the subsequent base in the second nucleotide mix are each labeled with one or more detectable labels (e.g., one, two, or three different colors), and the subsequent base is different from the first base.

[0187] In some embodiments, provided herein is a method for analyzing a biological sample, comprising: a) contacting the biological sample with a probe that directly or indirectly binds to an analyte at a position in the biological sample, wherein the probe or its product comprises a priming site and a barcode sequence, and a sequencing primer hybridizes to the priming site; b) contacting the biological sample with a first nucleotide mix comprising nucleotides containing different bases, wherein in the first nucleotide mix, the nucleotide containing the first base is not detectably labeled, the nucleotides containing one or more other bases are detectably labeled, a complex is formed, the complex comprising: i) a sequencing primer hybridized to the probe or its product, ii) a polymerase, and iii) cognate nucleotides that base pair with the first nucleotide in the barcode sequence, and a signal (on-signal) associated with the cognate nucleotides and / or the absence of a signal (off-signal) is detected at the position, and the on-signal, off-signal, or a combination thereof corresponds to the base in the cognate nucleotides and the first nucleotide in the barcode sequence; c) contacting the biological sample with a subsequent nucleotide mix comprising nucleotides containing different bases, wherein in the second nucleotide mix, the nucleotide containing the subsequent base is not detectably labeled, the nucleotides containing one or more other bases are detectably labeled, the second base is different from the first base, a complex is formed, the complex comprising: i) an extension product of the sequencing primer hybridized to the probe or its product, ii) a polymerase, and iii) cognate nucleotides that base pair with the subsequent nucleotide in the barcode sequence, and a signal (on-signal) associated with the cognate nucleotides and / or the absence of a signal (off-signal) is detected at the position, and the on-signal, off-signal, or a combination thereof corresponds to the base in the cognate nucleotides and the subsequent nucleotide in the barcode sequence; d) generating a signal code sequence comprising a signal code corresponding to the on-signal, off-signal, or a combination thereof in steps b) and c) at the position, wherebyDetecting a barcode array at a position in a biological sample, and,

[0188] In some embodiments, the biological sample can be contacted with two or more of the following nucleotide mixes in successive cycles in any order: Nucleotide mix 1, in which nucleotides containing G are not detectably labeled while nucleotides containing A, C, or T are detectably labeled (e.g., using 1, 2, or 3-color chemistry); Nucleotide mix 2, in which nucleotides containing T are not detectably labeled while nucleotides containing A, C, or G are detectably labeled (e.g., using 1, 2, or 3-color chemistry); Nucleotide mix 3, in which nucleotides containing C are not detectably labeled while nucleotides containing A, G, or T are detectably labeled (e.g., using 1, 2, or 3-color chemistry); and Nucleotide mix 4, in which nucleotides containing A are not detectably labeled while nucleotides containing G, C, or T are detectably labeled (e.g., using 1, 2, or 3-color chemistry).

[0189] In some embodiments, any of the nucleotide mixes (e.g., nucleotide mixes 1-4) can be used in two or more successive base-by-base sequencing cycles where cycles using different nucleotide mixes precede and / or follow. An exemplary order can be nucleotide mix 4 - nucleotide mix 2 - nucleotide mix 2 - nucleotide mix 1 - nucleotide mix 1 - nucleotide mix 1 - nucleotide mix 3.

[0190] B. Multiple sequencing primers for different identifier arrays In some embodiments, methods are provided herein that include determining in situ the sequences of a plurality of different identifier sequences (e.g., barcode sequences) in a cell or tissue sample, wherein a first subset of the different identifier sequences is sequenced using a first sequencing primer and a second subset of the different identifier sequences is sequenced using a second sequencing primer. In some aspects, the identifier sequences sequenced using different sequencing primers may include bases that do not generate a detectable signal in the corresponding base-by-base sequencing cycle, e.g., these bases are “dark” as described in Section III.A. In some embodiments, the first and second sequencing primers are different, e.g., they bind to different priming sites and when the first subset of the different identifier sequences is sequenced using the first sequencing primer, the signal associated with the second subset of the different identifier sequences is not detected. Similarly, when the second subset of the different identifier sequences is sequenced using the second sequencing primer, the signal associated with the first subset of the different identifier sequences is not detected. Thus, by using multiple sequencing primers across different identifier sequences and detecting the signals separately using each sequencing primer, optical crowding of the signals can be improved.

[0191] In some embodiments, a method for analyzing a biological sample is provided herein, comprising: a) contacting the biological sample with a plurality of probes, each of which binds directly or indirectly to an analyte at a position in the biological sample, wherein the first probe or its product comprises a first priming site and a first barcode sequence, the second probe or its product comprises a second priming site and a second barcode sequence, and the first priming site and the second priming site are different; b) hybridizing a first sequencing primer to the first priming site; c) contacting the biological sample with nucleotides in successive cycles, wherein in each cycle, a complex is formed at a first position, the complex comprising: i) the first sequencing primer or its extension product hybridized to the first probe or its product, ii) a polymerase, and iii) cognate nucleotides that base pair with nucleotides in the first barcode sequence, and a signal (on-signal) associated with the cognate nucleotides and / or the absence of a signal (off-signal) is detected at the first position, and the on-signal, off-signal, or a combination thereof corresponds to the base in the cognate nucleotides and the corresponding nucleotide in the first barcode sequence; d) removing or blocking the first sequencing primer or its extension product; e) hybridizing a second sequencing primer to the second priming site; f) contacting the biological sample with nucleotides in successive cycles, wherein in each cycle, a complex is formed at a second position, the complex comprising: i) the second sequencing primer or its extension product hybridized to the second probe or its product, ii) a polymerase, and iii) cognate nucleotides that base pair with nucleotides in the second barcode sequence, and a signal (on-signal) associated with the cognate nucleotides and / or the absence of a signal (off-signal) is detected at the second position, and the on-signal, off-signal, or a combination thereof corresponds to the base in the cognate nucleotides and the corresponding nucleotide in the second barcode sequence; g) a first signal code sequence comprising a signal code corresponding to the signal in the successive cycles at the first position.and generating a second signal code array including signal codes corresponding to signals in the continuous cycles at the second position, thereby detecting a first barcode array at the first position and a second barcode array at the second position in the biological sample.

[0192] In some embodiments, the first probe and the second probe bind directly or indirectly to the same analyte. In some embodiments, the first probe and the second probe bind directly or indirectly to different analytes. In some embodiments, the first barcode array and the second barcode array are of the same sequence, but the 3' priming sites are different. Thus, the same barcode array can be sequenced using different sequencing primers. In some embodiments, the first barcode array and the second barcode array are different and are sequenced using different sequencing primers. The first and second barcode arrays may be related to, correspond to, and / or identify the same or different analytes. In some cases, a plurality of different probes include the same first priming site (e.g., primer binding sequence), but each of the plurality of different probes includes a different barcode array. In some cases, a first plurality of different probes includes the same first priming site (e.g., primer binding sequence), but each of the first plurality of different probes includes a different barcode array, and a second plurality of different probes includes the same second priming site (e.g., primer binding sequence), but each of the second plurality of different probes includes a different barcode array.

[0193] In some embodiments, the method can include hybridizing a first sequencing primer to a first priming site and performing base-by-base sequencing (e.g., SBS or SBB) to generate an extension product of the first sequencing primer. As shown in FIG. 3, the first sequencing primer hybridizes to a first priming site in a molecule or complex (e.g., a probe or its product) that includes a first set of identifier sequences (a "block"), e.g., barcode sequences for genes 1-3 (gene block 1 in the drawing), and can be used to sequence the first set of identifier sequences without detecting signals associated with bases in a second set of identifier sequences (a "block"), e.g., barcode sequences for genes 4-6 (gene block 2 in the drawing). In some embodiments, the extension product of the first sequencing primer is prevented from generating signals for base-by-base sequencing such as SBS and SBB. In some embodiments, the method includes removing, cleaving, or blocking the extension product of the first sequencing primer. In some embodiments, the method further includes hybridizing a second sequencing primer to a second priming site and performing base-by-base sequencing (e.g., SBS or SBB) to generate an extension product of the second sequencing primer. As shown in FIG. 3, the second sequencing primer hybridizes to a second priming site in a molecule or complex (e.g., a probe or its product) that includes a second set of identifier sequences (e.g., the barcode sequence in gene block 2), and can be used to sequence the second set of identifier sequences without detecting signals associated with bases in a first set of identifier sequences (e.g., the barcode sequence in gene block 1). In some embodiments, the extension product of the second sequencing primer is prevented from generating signals for base-by-base sequencing such as SBS and SBB. In some embodiments, the method includes removing, cleaving, or blocking the extension product of the second sequencing primer.

[0194] In some embodiments, the method further comprises hybridizing a third sequencing primer to a third priming site and performing base-by-base sequencing (e.g., SBS or SBB) to generate an extension product of the third sequencing primer. In some embodiments, the third sequencing primer has a different sequence than the first and second sequencing primers. The third sequencing primer hybridizes to a third priming site in a molecule or complex (e.g., a probe or its product) that includes a third set of identifier sequences and can be used to sequence the third set of identifier sequences without detecting a signal related to a base in the first or second set of identifier sequences. For example, the third sequencing primer can be used to sequence a barcode sequence for a gene that is different from genes 1-6 shown in FIG. 3.

[0195] In some embodiments, probes or their products for a first plurality of analytes can share a common first priming site, and probes or their products for a second plurality of analytes can share a common second priming site. In some embodiments, the second plurality of analytes can include two or more different analytes that are different from two or more different analytes of the first plurality of analytes. For example, as shown in FIG. 3, probes or their products for gene block 1 (including genes 1-3) share a common first priming site (the "sequencing primer 1" binding site), and probes or their products for gene block 2 (including genes 4-6) share a common second priming site (the "sequencing primer 2" binding site), and at least one, two, or all of the genes in gene block 1 are different from those in gene block 2.

[0196] In some embodiments, a first sequencing primer (e.g., "Sequencing Primer 1" in FIG. 3) and a second sequencing primer (e.g., "Sequencing Primer 2" in FIG. 3) can be pre-mixed or can be separate but contacted with the biological sample simultaneously, followed by multiple decoding cycles. In some embodiments, one of the sequencing primers can be selectively blocked from generating a signal in base-by-base sequencing while sequencing from the other sequencing primer is performed. When sequencing from one of the sequencing primers is complete, the blocked sequencing primer can be unblocked and used to sequence the identifier sequence.

[0197] In some embodiments, the plurality of identifier sequences includes a plurality of subsets, each subset being sequenced in a separate series of sequencing cycles using different sequencing primer / priming sites, thereby reducing optical crowding compared to a method in which all of the plurality of identifier sequences are sequenced in the same series of sequencing cycles.

[0198] In some cases, barcode sequences corresponding to two analytes (e.g., genes) can be sequenced using two different sequencing primers, and the barcode sequences can be shared between the two genes. In some cases, the barcode sequences can be reused for different analytes.

[0199] C. Combinations of Sequencing Primers for the Same Identifier Sequence In some embodiments, the same identifier sequence can be sequenced using two or more different sequencing primers. For example, a molecule or complex can include two different priming sites, as shown, for example, in FIG. 4, and each of the molecules or complexes for genes 1-5 can include two different priming sites for sequencing an identifier sequence corresponding to an analyte (e.g., a transcript of the corresponding gene). The molecule or complex can be any of those disclosed herein, such as an RCA product or a probe complex as shown in FIG. 2. In each molecule or complex, both of the two different priming sites can be 3′ of the identifier sequence (e.g., a barcode sequence), and one of the priming sites can be 3′ or 5′ of the other priming site. In some embodiments, the different priming sites can partially overlap. In some embodiments, in each molecule or complex, one of the priming sites can be 3′ of one copy of the identifier sequence, and the other priming site can be 3′ of another copy of the identifier sequence. In some aspects, the identifier sequences sequenced using different priming sites can include bases that do not produce a detectable signal in the corresponding base-by-base sequencing cycle, e.g., these bases are “dark” as described in Section III.A.

[0200] In some embodiments, a biological sample can be contacted with a plurality of probes configured to each directly or indirectly bind a different analyte, and each probe or its product can include a combination of different priming sites. In some embodiments, the combination includes 2, 3, 4, 5, or more different priming sites. As an example, FIG. 4 shows that a probe or its product for each gene includes two different priming sites, e.g., a “sequencing primer 1” binding site and a “sequencing primer 2” binding site for gene 1, and a “sequencing primer 1” binding site and a “sequencing primer 3” binding site for gene 2, etc.

[0201] In some embodiments, the first probe or its product can include a first combination of different priming sites that includes a first priming site. For example, the barcode sequence for gene 1 can be determined using "sequencing primer 1" and "sequencing primer 2" as shown in FIG. 4. In some embodiments, the second probe or its product can include a second combination of different priming sites that includes a second priming site. For example, the barcode sequence for gene 2 can be determined using "sequencing primer 1" and "sequencing primer 3" as shown in FIG. 4. In some embodiments, the biological sample can be contacted with a third probe that binds directly or indirectly to a third analyte. In some embodiments, the third probe or its product can include a third combination of different priming sites that includes a first priming site, a second priming site, and / or a third priming site. In some embodiments, any two or more of the first combination, the second combination, and the third combination can share one or more common priming sites. For example, the barcode sequence for gene 3 can be determined using "sequencing primer 1" (shared with genes 1 and 2), and "sequencing primer 4" (not shared with genes 1 and 2).

[0202] In some embodiments, the method comprises contacting a biological sample with a first sequencing primer for base-by-base sequencing, thereby hybridizing the first sequencing primer to a first priming site in a first probe or its product and in one or more other probes or their products, and generating an extension product of the first sequencing primer. For example, as shown in FIG. 4, "Sequencing Primer 1" can hybridize to a priming site in a probe for genes 1-3 or its product, and an identifier sequence (e.g., a barcode sequence) for genes 1-3 is sequenced. In some embodiments, the method further comprises removing, cleaving, or blocking the extension product of the first sequencing primer so as to prevent generation of a signal for base-by-base sequencing using, for example, SBS or SBB. In some embodiments, the method further comprises contacting the biological sample with a second sequencing primer for base-by-base sequencing, thereby hybridizing the second sequencing primer to a second priming site in a second probe or its product and in one or more other probes or their products, and generating an extension product of the second sequencing primer. For example, as shown in FIG. 4, "Sequencing Primer 2" can hybridize to a priming site in a probe for genes 1, 4, and 5 or its product, and an identifier sequence (e.g., a barcode sequence) for these genes is sequenced. Similarly, "Sequencing Primer 3" can hybridize to a priming site in a probe for genes 2, 4, and 6 or its product, and an identifier sequence (e.g., a barcode sequence) for these genes is sequenced.

[0203] In some embodiments, the method may include contacting a biological sample with a first sequencing primer and a second sequencing primer, thereby hybridizing both sequencing primers to corresponding priming sites in a first probe or its product, a second probe or its product, and optionally one or more other probes or their products, and generating an extension product of the first sequencing primer and an extension product of the second sequencing primer. In some embodiments, the first and second sequencing primers may be pre-mixed or may be separate but contacted with the biological sample simultaneously, followed by a plurality of decoding cycles. In some embodiments, the extension products of the first sequencing primer and the second sequencing primer may be removed (e.g., stripped), cleaved, or blocked, or otherwise prevented from generating a signal for base-by-base sequencing, followed by hybridization of a different mix of sequencing primers (e.g., a third sequencing primer and a fourth sequencing primer). For example, as shown in FIG. 4, "Sequencing Primer 1" and "Sequencing Primer 2" may hybridize to priming sites in a probe or its product, where "Sequencing Primer 1" is for Genes 1-3 and "Sequencing Primer 2" is for Genes 1, 4, and 5. For an identifier sequence (e.g., a barcode sequence) for Gene 1 for which sequencing can be performed from both sequencing primers, the signal may be distinguishable, for example, as "ACG" from "Sequencing Primer 1" and "CGA" from "Sequencing Primer 2". This can be achieved, for example, if the identifier sequence includes a plurality of sub-sequences separated by priming sites, such as 3'-"Sequencing Primer 1" binding site-"ACG"-"Sequencing Primer 2" binding site-"CGA"-5'.

[0204] In some embodiments, multiple cycles of nucleotide incorporation or ligation may be performed. In some embodiments, the method comprises contacting a biological sample with nucleotides in successive cycles, detecting a signal associated with nucleotide incorporation or ligation for each successive cycle, and generating a signal code sequence for a plurality of analytes, e.g., a first plurality of analytes and a second plurality of analytes. In some embodiments, the first plurality of analytes and the second plurality of analytes include one or more common analytes. In some embodiments, the first plurality of analytes and the second plurality of analytes do not include a common analyte.

[0205] In some embodiments, by using combinations of sequencing primers for the same identifier sequence, different portions of the identifier sequence can be sequenced in separate sets of sequencing cycles using different sequencing primers or combinations of different sequencing primers, which helps reduce optical crowding compared to methods where all of the plurality of identifier sequences are sequenced in the same set of sequencing cycles. In some embodiments, the method is not constrained by a block diagonal code. For example, using combinations of sequencing primers can allow for an increase in complexity for the identifier sequence.

[0206] In some embodiments, any identifier sequence or portion thereof disclosed herein can be sequenced multiple times using the same sequencing primer or one or more different sequencing primers, using the same nucleotide mix or one or more different nucleotide mixes, and / or using the same polymerase or one or more different polymerases (including, e.g., the same polymerase with different labels or modifications).

[0207] In some embodiments, after the identifier array or a portion thereof has been sequenced, for example, in round N, the extension product of the round 1 sequencing primer (or a portion thereof) is removed from the template strand to enable hybridization of the sequencing primer for round (N + 1). In some embodiments, the round (N + 1) sequencing primer and the round N sequencing primer have the same sequence, such that the identifier array or a portion thereof can be sequenced again. In some embodiments, in round (N + 1), the identifier array or a portion thereof is sequenced using the same nucleotide mix as in round N. In some embodiments, in rounds (N + 1) and N, the identifier array or a portion thereof is sequenced using different nucleotide mixes, for example, as described in Section III-A. In some embodiments, the round (N + 1) sequencing primer and the round N sequencing primer have different sequences, and the nucleotide mixes in rounds (N + 1) and N can be the same or different. In some embodiments, both the round (N + 1) priming site and the round N priming site are 3' of the identifier array, and the round (N + 1) priming site can be 3' or 5' of the round N priming site. In some embodiments, the round (N + 1) priming site and the round N priming site can partially overlap.

[0208] IV. Detection and Analysis Generally, in synthetic sequencing methods, a first population of detectably labeled nucleotides (e.g., dNTPs) is introduced to contact template nucleotides hybridized to a sequencing primer, and a first detectably labeled nucleotide (e.g., an A, T, C, or G nucleotide) is incorporated by polymerase to extend the sequencing primer in the 5' to 3' direction using a complementary nucleotide (the first nucleotide residue) in the template nucleotides as a template. Then, a signal from the first detectably labeled nucleotide can be detected. The first population of nucleotides can be introduced continuously, but in order to incorporate a second detectably labeled nucleotide into the extended sequencing primer, the nucleotides in the first population of nucleotides not incorporated into the sequencing primer are generally removed (e.g., by washing), and a second population of detectably labeled nucleotides is introduced into the reaction. Then, a second detectably labeled nucleotide (e.g., an A, T, C, or G nucleotide) is incorporated by the same or a different polymerase to extend the already extended sequencing primer in the 5' to 3' direction using a complementary nucleotide (the second nucleotide residue) in the template nucleotides as a template. Thus, in some embodiments, cycles of introducing and removing detectably labeled nucleotides are performed.

[0209] In some cases, for the in situ analysis of analytes in a biological sample, such as target nucleic acids in cells of intact tissue, cycles of introducing and removing detectably labeled nucleotides are performed in a temporally sequential manner. In some embodiments, methods are provided herein for detecting detectably labeled nucleotides, thereby generating a signal signature associated with the labeled oligonucleotide. In some cases, the signal signature corresponds to one of a plurality of analytes. In some cases, the methods described herein are based in part on the development of multiplexed biological assays and readouts, first contacting a sample with a plurality of nucleic acid probes that allow the probes to bind directly or indirectly to the target analyte, which can then be optically detected (e.g., by sequencing) in a temporally sequential manner. In some embodiments, the nucleic acid probes are added exogenously and their sequences can be controlled and designed, thereby allowing the use of identifier sequences optimized for each analyte (e.g., as compared to directly sequencing endogenous molecules). In some embodiments, methods are provided herein involving multiplexed biological assays and sequencing readouts that include optically detecting labeled oligonucleotides in a temporally sequential manner. In some cases, the positions of the analyte, probe, and / or its product can be maintained in the sample through multiple sequencing cycles, such that fluorescent spots corresponding to the analyte, probe, or its product remain in fixed positions during multiple rounds and can be aligned to read a string of signals associated with each target analyte. The string of signals observed at a given position (e.g., on or off signals in each round) can be compared to a codebook that includes identifier sequences assigned to a plurality of analytes.

[0210] In some embodiments, the methods disclosed herein include using one or more nucleotides or analogs thereof, including natural nucleotides, nucleotide analogs, or modified nucleotides (e.g., labeled with one or more detectable labels). In some embodiments, nucleotide analogs include a nitrogenous base, a 5-carbon sugar, and a phosphate group, and any component of the nucleotide can be modified and / or substituted. In some embodiments, the methods disclosed herein may include using one or more non-incorporable nucleotides. Non-incorporable nucleotides can be modified to be incorporable at any point during the sequencing method.

[0211] Examples of nucleotide analogs include, but are not limited to, alpha-phosphate modified nucleotides, alpha-beta nucleotide analogs, beta-phosphate modified nucleotides, beta-gamma nucleotide analogs, gamma-phosphate modified nucleotides, caged nucleotides, or ddNTPs. Examples of nucleotide analogs are described in U.S. Patent No. 8,071,755, which is incorporated herein by reference in its entirety.

[0212] In some embodiments, the methods disclosed herein may include using a terminator that reversibly prevents nucleotide incorporation at the 3' end of the primer. One type of reversible terminator is a 3'-O-blocked reversible terminator. Here, the terminator moiety is linked to the oxygen atom at the 3'-OH terminus of the 5-carbon sugar of the nucleotide. For example, U.S. Patent Nos. 7,544,794 and 8,034,923 (the disclosures of these patents are incorporated by reference) describe 3'-ONH 2Describes reversible terminator dNTPs having a 3'-OH group replaced by . Another type of reversible terminator is a 3'-unblocked reversible terminator, where the terminator moiety is linked to the nitrogenous base of the nucleotide. For example, U.S. Patent No. 8,808,989 (the disclosure of which is incorporated by reference) discloses specific examples of base-modified reversible terminator nucleotides that can be used in connection with the methods described herein. Other reversible terminators that can be similarly used in connection with the methods described herein include those described in U.S. Patent Nos. 7,956,171, 8,071,755, and 9,399,798, which are incorporated herein by reference.

[0213] In some embodiments, the methods disclosed herein may include using nucleotide analogs having a terminator moiety that irreversibly prevents nucleotide incorporation at the 3'-end of the primer. Irreversible nucleotide analogs include 2',3'-dideoxynucleotides, ddNTPs (ddGTP, ddATP, ddTTP, ddCTP). Dideoxynucleotides lack the 3'-OH group of dNTPs that is essential for polymerase-mediated synthesis.

[0214] In some embodiments, the methods disclosed herein may include using non-incorporable nucleotides that include a blocking moiety that inhibits or prevents a nucleotide from forming a covalent bond to a second nucleotide (the 3'-OH of the primer) during the incorporation step of the nucleic acid polymerization reaction. The blocking moiety can be removed from the nucleotide to allow nucleotide incorporation.

[0215] In some embodiments, the methods disclosed herein may involve using one, two, three, four, or more nucleotide analogs present in the SBS reaction. In some embodiments, the nucleotide analogs are substituted, diluted, or sequestered during the incorporation step. In some embodiments, the nucleotide analogs are replaced with natural nucleotides. In some embodiments, the nucleotide analogs are modified during the incorporation step. The modified nucleotide analogs may be similar to or the same as natural nucleotides.

[0216] In some embodiments, the methods disclosed herein may involve using nucleotide analogs that have a binding affinity for polymerases different from that of natural nucleotides. In some embodiments, the nucleotide analogs have an interaction with adjacent bases different from that of natural nucleotides. Nucleotide analogs and / or non-incorporable nucleotides may base pair with complementary bases of the template nucleic acid.

[0217] In some embodiments, one or more nucleotides may be labeled with a distinguishable and / or detectable tag or label. The tags may be distinguishable by their differences in fluorescence, Raman spectrum, charge, mass, refractive index, luminescence, length, or any other measurable property. The tags may be attached at one or more different positions on the nucleotide as long as the fidelity of binding to the polymerase-nucleic acid complex is sufficiently maintained to allow correct identification of the complementary bases on the template nucleic acid. In some embodiments, the tag is attached to the nucleobase of the nucleotide. Alternatively, the tag is attached to the gamma phosphate position of the nucleotide.

[0218] Detectable labels may be suitable for small-scale detection and / or may be suitable for high-throughput screening. Thus, suitable detectable labels include, but are not limited to, radioisotopes, fluorophores, chemiluminescent compounds, bioluminescent compounds, and dyes. Detectable labels can be detected qualitatively (e.g., optically or spectrally), or they can be quantified. Qualitative detection generally includes detection methods in which the existence or presence of a detectable label is confirmed, while quantifiable detection generally includes detection methods having quantifiable (e.g., numerically reportable) values such as intensity, duration, polarization, and / or other characteristics. In some embodiments, the detectable label is attached to another moiety, such as a nucleotide or nucleotide analog, and may include a fluorescent label, a colorimetric label, or a chemiluminescent label.

[0219] In some embodiments, the detectable label can be attached to another moiety, such as a nucleotide or nucleotide analog. In some embodiments, the detectable label is a fluorophore. For example, the fluorophore can be from the group including: 7-AAD (7-aminoactinomycin D), acridine orange (+DNA), acridine orange (+RNA), Alexa Fluor® 350, Alexa Fluor® 430, Alexa Fluor® 488, Alexa Fluor® 532, Alexa Fluor® 546, Alexa Fluor® 555, Alexa Fluor® 568, Alexa Fluor® 594, Alexa Fluor® 633, Alexa Fluor® 647, Alexa Fluor® 660, Alexa Fluor® 680, Alexa Fluor® 700, Alexa Fluor® 750, allophycocyanin (APC), AMCA / AMCA-X, 7-aminoactinomycin D (7-AAD), 7-amino-4-methylcoumarin, 6-aminoquinoline, aniline blue, ANS, APC-Cy7, ATTO-TAG™ CBQCA, ATTO-TAG™ FQ, Auramine O-Feulgen, BCECF (high pH), BFP (blue fluorescent protein), BFP / GFP FRET, BOBO™-1 / BO-PRO™-1, BOBO™-3 / BO-PRO™-3, BODIPY® FL, BODIPY® TMR, BODIPY® TR-X, BODIPY® 530 / 550, BODIPY® 558 / 568, BODIPY® 564 / 570, BODIPY® 581 / 591, BODIPY® 630 / 650-X, BODIPY® 650-665-X, BTC, calcein, calcein blue, Calcium Crimson™, Calcium Green-1™, CalciumOrange (trademark), Calcofluor (registered trademark) White, 5-carboxyfluorescein (5-FAM), 5-carboxynaphthofluorescein, 6-carboxyrhodamine 6G, 5-carboxytetramethylrhodamine (5-TAMRA), carboxy-X-rhodamine (5-ROX), Cascade Blue (registered trademark), Cascade Yellow (trademark), CCF2 (GeneBLAzer (trademark)), CFP (cyan fluorescent protein), CFP / YFP FRET, chromomycin A3, Cl-NERF (low pH), CPM, 6-CR 6G, CTC formazan, Cy2 (registered trademark), Cy3 (registered trademark), Cy3.5 (registered trademark), Cy5 (registered trademark), Cy5.5 (registered trademark), Cy7 (registered trademark), Cychrome (PE-Cy5), dansylamine, dansylcadaverine, dansyl chloride, DAPI, dapoxyl, DCFH, DHR, DiA (4-Di-16-ASP), DiD (DilC18(5)), DIDS, Dil (DilC18(3)), DiO (DiOC18(3)), DiR (DilC18(7)), Di-4 ANEPPS, Di-8 ANEPPS, DM-NERF (4.5~6.5pH), DsRed (red fluorescent protein), EBFP, ECFP, EGFP, ELF (registered trademark)-97 alcohol, eosin, erythrosin, ethidium bromide, ethidium homodimer-1 (EthD-1), europium(III) chloride, 5-FAM (5-carboxyfluorescein), fast blue, fluorescein-dT phosphoramidite, FITC, Fluo-3, Fluo-4, FluorX (registered trademark), Fluoro-Gold (trademark) (high pH), Fluoro-Gold (trademark) (low pH), Fluoro-Jade, FM (registered trademark)1-43, Fura-2 (high calcium), Fura-2 / BCECF, Fura Red (trademark) (high calcium), Fura Red (trademark) / Fluo-3, GeneBLAzer (trademark) (CCF2), GFP red shift (rsGFP), GFP wild type, GFP / BFP FRET, GFP / DsRed FRET, Hoechst 33342 & 33258, 7-hydroxy-4-methylcoumarin (pH9), 1,5IAEDANS, Indo-1 (high calcium), Indo-1 (low calcium), indodicarbocyanine, indotricarbocyanine, JC-1, 6-JOE, JOJO(trademark)-1 / JO-PRO(trademark)-1, LDS 751(+DNA), LDS 751(+RNA), LOLO(trademark)-1 / LO-PRO(trademark)-1, lucifer yellow, LysoSensor(trademark) Blue(pH5), LysoSensor(trademark) Green(pH5), LysoSensor(trademark) Yellow / Blue(pH4.2), LysoTracker(registered trademark) Green, LysoTracker(registered trademark) Red, LysoTracker(registered trademark) Yellow, Mag-Fura-2, Mag-Indo-1, Magnesium Green(trademark), Marina Blue(registered trademark), 4-methylumbelliferone, mitramycin, MitoTracker(registered trademark) Green, MitoTracker(registered trademark) Orange, MitoTracker(registered trademark) Red, NBD(amine), Nile red, Oregon Green(registered trademark) 488, Oregon Green(registered trademark) 500, Oregon Green(registered trademark) 514, Pacific blue, PBF1, PE(R-phycoerythrin), PE-Cy5, PE-Cy7, PE-Texas Red, PerCP(peridinin chlorophyll protein), PerCP-Cy5.5(TruRed), PharRed(APC-Cy7), C-phycocyanin, R-phycocyanin, R-phycoerythrin(PE), PI(propidium iodide), PKH26, PKH67, POPO(trademark)-1 / PO-PRO(trademark)-1, POPO(trademark)-3 / PO-PRO(trademark)-3, propidium iodide(PI), PyMPO, pyrene, pyronin Y, Quantum Red(PE-Cy5), quinacrine mustard, R670(PE-Cy5), Red 613(PE-Texas Red), red fluorescent protein(DsRed), resorufin, RH 414, Rhod-2, rhodamine B, Rhodamine Green(trademark), RhodamineRed (Trademark), rhodamine phalloidin, rhodamine 110, rhodamine 123, 5-ROX (carboxy-X-rhodamine), S65A, S65C, S65L, S65T, SBFI, SITS, SNAFL®-1 (high pH), SNAFL®-2, SNARF®-1 (high pH), SNARF®-1 (low pH), Sodium Green (Trademark), SpectrumAqua®, SpectrumGreen® #1, SpectrumGreen® #2, SpectrumOrange®, SpectrumRed®, SYTO® 11, SYTO® 13, SYTO® 17, SYTO® 45, SYTOX® Blue, SYTOX® Green, SYTOX® Orange, 5-TAMRA (5-carboxytetramethylrhodamine), tetramethylrhodamine (TRITC), Texas Red® / Texas Red®-X, Texas Red®-X (NHS ester), thiazolylcarbocyanine, thiazole orange, TOTO®-1 / TO-PRO®-1, TOTO®-3 / TO-PRO®-3, TO-PRO®-5, tri-color (PE-Cy5), TRITC (tetramethylrhodamine), TruRed (PerCP-Cy5.5), WW 781, X-rhodamine (XRITC), Y66F, Y66H, Y66W, YFP (yellow fluorescent protein), YOYO®-1 / YO-PRO®-1, YOYO®-3 / YO-PRO®-3, 6-FAM (fluorescein), 6-FAM (NHS ester), 6-FAM (azide), HEX, TAMRA (NHS ester), Yakima Yellow, MAX, TET, TEX615, ATTO 488, ATTO 532, ATTO 542, ATTO 550, ATTO 565, ATTO Rho101, ATTO 590, ATTO 633, ATTO 647N, TYE 563, TYE 665, TYE 705, 5’ IRDye® 700, 5’ IRDye® 800, 5’IRDye® 800CW (NHS ester), WellRED D4 dye, WellRED D3 dye, WellRED D2 dye, Lightcycler® 640 (NHS ester), and Dy 750 (NHS ester).

[0220] A detectable label can be directly detectable by itself (e.g., a radioisotope label or a fluorescent label), or in the case of an enzyme label, can be indirectly detectable, for example, by catalyzing a chemical change in a substrate compound or composition, where the substrate compound or composition is directly detectable. The label can emit a signal or change a signal delivered to the label so that the presence or absence of the label can be detected. In some cases, the coupling can be through a cleavable linker that can be photocleavable (e.g., cleavable under ultraviolet light), chemically cleavable (e.g., through a reducing agent such as dithiothreitol (DTT), tris(2-carboxyethyl)phosphine (TCEP), etc.), or enzymatically cleavable (e.g., through an esterase, lipase, peptidase, or protease).

[0221] In addition to the synthetic sequencing approaches described herein, in some cases, nucleic acid sequencing can be performed using alternative sequencing biochemistries. Examples include, but are not limited to, the "sequencing by binding" (SBB) approach described in WO2017 / 117235, U.S. Patent Nos. 9,951,385 and 10,655,176, and the "sequencing by avidity" (SBA) approach described in U.S. Patent Nos. 10,768,173 and 10,982,280, all of which are incorporated herein by reference. The sequencing chemistry described in US2020 / 0370113, entitled "Polymerase-nucleotide conjugates for sequencing by trapping" and incorporated herein by reference, can also be used.

[0222] In some embodiments, the SBB approach of the present specification performs iterative cycles of detecting a stabilized complex formed at each position along a template under conditions that prevent covalent incorporation of cognate nucleotides into a primer (e.g., a ternary complex comprising a primed template (tethered to a sample support structure), a polymerase, and a cognate nucleotide for the position), and then extending the primer to enable detection of the next position along the template (see, e.g., U.S. Patent Nos. 9,951,385 and 10,655,176). In a binding-based sequencing approach, detection of a nucleotide at each position of the template occurs prior to extension of the primer to the next position. Generally, the methodology is used to distinguish the four different nucleotide types that may be present at positions along a nucleic acid template by uniquely labeling each type of ternary complex (i.e., different types of ternary complexes that differ in the type of nucleotide they contain), or by separately delivering the reagents required to form each type of ternary complex. In some cases, the label may comprise, for example, a fluorescent label of a cognate nucleotide or polymerase involved in the ternary complex.

[0223] In some cases, for example, a method for sequencing a nucleic acid molecule using a ligation-based array determination approach may include: (a) forming a mixture under ternary complex stabilization conditions, the mixture comprising a primed template nucleic acid, a polymerase, and nucleotide homologs of a first, second, and third base type in the template; (b) examining the mixture to determine whether a ternary complex has formed; (c) identifying the next correct nucleotide for the primed template nucleic acid molecule, wherein if a ternary complex is detected in step (b), the next correct nucleotide is identified as a homolog of the first, second, or third base type, and based on the absence of ternary complex formation in step (b), the next correct nucleotide is considered to be a nucleotide homolog of a fourth base type; (d) after step (b), adding the next correct nucleotide to the primer of the primed template nucleic acid, thereby generating an extended primer; and (e) repeating steps (a) to (d) for the primed template nucleic acid comprising the extended primer. In some cases, the mixture formed in step (a) may include nucleotide homologs of a first base type in the template. In some cases, the mixture formed in step (a) may include nucleotide homologs of first and second base types in the template.

[0224] The "Sequence By Affinity" (or SBA) approach relies on the increase in affinity (or functional affinity) resulting from the formation of a complex containing multiple individual non-covalent interactions (see, for example, U.S. Patent Nos. 10,768,173 and 10,982,280). The Sequence By Affinity approach is based on the detection of a multivalent binding complex formed between a fluorescently labeled polymer-nucleotide conjugate, a polymerase, and a plurality of primed target nucleic acid molecules tethered to a sample support structure, thereby enabling the separation of the detection / base calling step from the nucleotide incorporation step. Fluorescent imaging is used to detect the bound complex and thereby determine the identity of the N+1 nucleotide in the target nucleic acid sequence (the primer extension strand is of length N nucleotides).

[0225] In some cases, for example, nucleic acid sequencing using the Sequence By Affinity approach may include: (a) providing a composition comprising (i) two or more copies of a target nucleic acid sequence, (ii) two or more primer nucleic acid molecules complementary to one or more regions of the target nucleic acid sequence, and (iii) two or more polymerase molecules; (b) contacting the composition with a polymer-nucleotide conjugate under conditions sufficient to allow a multivalent binding complex to form between the polymer-nucleotide conjugate and the two or more copies of the target nucleic acid sequence in the composition, wherein the polymer-nucleotide conjugate comprises two or more nucleotide moieties; and (c) detecting the multivalent binding complex (e.g., by fluorescence imaging) and thereby determining the identity of the nucleotides in the target nucleic acid sequence. After the imaging step, the multivalent binding complex is disrupted, washed away, and the correct nucleotide (e.g., a blocked nucleotide) is incorporated into the primer extension strand (e.g., after unblocking a previously incorporated nucleotide), and the cycle is repeated.

[0226] Any suitable enzyme having polymerase activity can be used in the sequencing reactions described herein. Exemplary polymerases include, but are not limited to, bacterial DNA polymerases, eukaryotic DNA polymerases, archaeal DNA polymerases, viral DNA polymerases, and phage DNA polymerases. Bacterial DNA polymerases include Escherichia coli DNA polymerases I, II, and III, IV, and V, the Klenow fragment of E. coli DNA polymerase, Clostridium stercorarium (Cst) DNA polymerase, Clostridium thermocellum (Cth) DNA polymerase, and Sulfolobus solfataricus (Sso) DNA polymerase. Eukaryotic DNA polymerases include DNA polymerases α, β, γ, δ, ε, η, ζ, λ, σ, μ, and κ, as well as Rev1 polymerase (terminal deoxynucleotidyl transferase) and terminal deoxynucleotidyl transferase (TdT). Viral DNA polymerases include T4 DNA polymerase, phi-29 DNA polymerase, GA-1, phi-29-like DNA polymerase, PZA DNA polymerase, phi-15 DNA polymerase, Cpl DNA polymerase, Cp7 DNA polymerase, T7 DNA polymerase, and T4 polymerase. Other DNA polymerases include Thermus aquaticus (Taq) DNA polymerase, Thermus filiformis (Tfi) DNA polymerase, Thermococcus zilligi (Tzi) DNA polymerase, Thermus thermophilus (Tth) DNA polymerase, Thermus flavusu (Tfl) DNA polymerase, Pyrococcus woesei (Pwo) DNA polymerase, Pyrococcus furiosus (PyrococcusPyrococcus furiosus (Pfu) DNA polymerase and Turbo Pfu DNA polymerase, Thermococcus litoralis (Tli) DNA polymerase, Pyrococcus GB-D polymerase, Thermotoga maritima (Tma) DNA polymerase, Bacillus stearothermophilus (Bst) DNA polymerase, Pyrococcus Kodakaraensis (KOD) DNA polymerase, Pfx DNA polymerase, Thermococcus sp. JDF-3 (JDF-3) DNA polymerase, Thermococcus gorgonarius (Tgo) DNA polymerase, Thermococcus acidophilium DNA polymerase; Sulfolobus acidocaldarius DNA polymerase; Thermococcus sp. go N-7 DNA polymerase; Pyrodictium occultum DNA polymerase; Methanococcus voltae DNA polymerase; Methanococcus thermoautotrophicum DNA polymerase; Methanococcus jannaschii DNA polymerase; Desulfurococcus sp. TOK DNA polymerase (D.Tok Pol); Pyrococcus abyssi DNA polymerase; Pyrococcus horikoshii DNA polymerase; Pyrococcus islandicum DNA polymerase; Thermococcus fumicolansDNA polymerases; the DNA polymerase of Aeropyrum pernix; and thermostable and / or thermophilic DNA polymerases such as the DNA polymerase isolated from the heterodimeric DNA polymerase DP1 / DP2. Engineered and modified polymerases are also useful in connection with the disclosed technology. For example, a modified form of the extremely thermophilic marine archaeon Thermococcus species 9°N (e.g., Terminator DNA polymerase from New England BioLabs Inc., Ipswich, Mass.) can be used. Still other useful DNA polymerases, including 3PDX polymerase, are disclosed in U.S. Patent No. 8,703,461, the disclosure of which is incorporated herein by reference in its entirety. Additional examples include viral RNA polymerases such as T7 RNA polymerase, T3 polymerase, SP6 polymerase, and K11 polymerase; eukaryotic RNA polymerases such as RNA polymerase I, RNA polymerase II, RNA polymerase III, RNA polymerase IV, and RNA polymerase V, archaeal RNA polymerase, HIV-1 reverse transcriptase from human immunodeficiency virus type 1 (PDB 1HMV), HIV-2 reverse transcriptase from human immunodeficiency virus type 2, M-MLV reverse transcriptase from Moloney murine leukemia virus, AMV reverse transcriptase from avian myeloblastosis virus, and telomerase reverse transcriptase that maintains the telomeres of eukaryotic chromosomes.

[0227] Fluorescence detection in tissue samples can often be hampered by the presence of strong background fluorescence. "Autofluorescence" is a common term used to distinguish background fluorescence (which can arise from various sources including aldehyde fixation, extracellular matrix components, red blood cells, lipofuscin, etc.) from the desired immunofluorescence from fluorescently labeled antibodies or probes. Tissue autofluorescence can pose difficulties in distinguishing signals from fluorescent antibodies or probes from common background. In some embodiments, the methods disclosed herein utilize one or more agents to reduce tissue autofluorescence, e.g., Autofluorescence Eliminator (Sigma / EMD Millipore), TrueBlack Lipofuscin Autofluorescence Quencher (Biotium), MaxBlock Autofluorescence Reducing Reagent Kit (MaxVision Biosciences), and / or very strong black dyes (e.g., Sudan black or equivalent dark chromophores).

[0228] Examples of fluorescent labels and nucleotides and / or polynucleotides conjugated to such fluorescent labels include those described, for example, in Hoagland, Handbook of Fluorescent Probes and Research Chemicals, Ninth Edition (Molecular Probes, Inc., Eugene, 2002), Keller and Manak, DNA Probes, 2nd Edition (Stockton Press, New York, 1993), Eckstein, editor, Oligonucleotides and Analogues: A Practical Approach (IRL Press, Oxford, 1991), and Wetmur, Critical Reviews in Biochemistry and Molecular Biology, 26:227-259 (1991). In some embodiments, exemplary techniques and methodologies applicable to the provided embodiments include those described, for example, in US4,757,141, US5,151,507, and US5,091,519. In some embodiments, one or more fluorescent dyes are used as labels for labeled target sequences as described, for example, in US5,188,934 (4,7-dichlorofluorescein dye), US5,366,860 (spectrally resolvable rhodamine dye), US5,847,162 (4,7-dichlororhodamine dye), US4,318,846 (ether-substituted fluorescein dye), US5,800,996 (energy transfer dye), US5,066,580 (xanthine dye), and US5,688,648 (energy transfer dye). Labeling can also be carried out using quantum dots described in US6,322,901, US6,576,291, US6,423,551, US6,251,303, US6,319,426, US6,426,513, US6,444,143, US5,990,479, US6,207,392, US2002 / 0045045, and US2003 / 0017264.As used herein, the term "fluorescent label" includes a signaling moiety that conveys information through the fluorescence absorption and / or emission characteristics of one or more molecules. Exemplary fluorescence characteristics include fluorescence intensity, fluorescence lifetime, emission spectral characteristics, and energy transfer.

[0229] In some embodiments, detection (including imaging) is performed using any of a number of different types of microscopy, such as confocal microscopy, two-photon microscopy, light field microscopy, tissue-clearing expansion microscopy, and / or CLARITY (trademark)-optimized light sheet microscopy (COLM).

[0230] In some embodiments, fluorescence microscopy is used for detection and imaging of samples. In some embodiments, a fluorescence microscope is an optical microscope that uses fluorescence and phosphorescence instead of, or in addition to, reflection and absorption to study the properties of organic or inorganic substances. In fluorescence microscopy, the sample is irradiated with light of a wavelength that excites fluorescence in the sample. Then, light that emits fluorescence, which is usually at a longer wavelength than the irradiation, is imaged through the microscope objective lens. In this technique, two filters can be used: an irradiation (or excitation) filter that ensures that the irradiation is substantially monochromatic and of the correct wavelength, and a second emission (or barrier) filter that ensures that neither the excitation light source nor the irradiation reaches the detector. Alternatively, both of these functions can be achieved by a single dichroic filter. "Fluorescence microscope" includes any microscope that uses fluorescence to generate an image, whether in a simpler configuration such as an epi-fluorescence microscope or a more complex design such as a confocal microscope that uses optical sectioning to obtain better resolution of the fluorescence image.

[0231] In some embodiments, confocal microscopy is used for sample detection and imaging. Confocal microscopy uses point illumination and a pinhole within an optically conjugate plane in front of the detector to remove out-of-focus signals. Since only the light generated by fluorescence very close to the focal plane can be detected, the optical resolution of the image, particularly in the depth direction of the sample, is much better than that of wide-field microscopes. However, since most of the light from the fluorescence of the sample is blocked by the pinhole, this increase in resolution comes at the expense of a decrease in signal intensity, and thus long exposures are often required. Since only one point in the sample is illuminated at a time, 2D or 3D imaging needs to scan across a regular raster (i.e., a rectangular pattern of parallel scan lines) in the specimen. The achievable thickness of the focal plane is defined mainly by the wavelength of the light used divided by the numerical aperture of the objective lens, but also by the optical properties of the specimen. The ability to perform very thin optical sectioning makes these types of microscopes particularly good for 3D imaging and surface profiling of samples. CLARITY (trademark) Optimized Light Sheet Microscopy (COLM) provides an alternative microscopy method for rapid 3D imaging of large cleared samples. COLM collates large immunostained tissues, enabling an increase in acquisition speed and resulting in higher quality of the generated data.

[0232] Other types of microscopy that can be used include brightfield microscopy, oblique illumination microscopy, darkfield microscopy, phase contrast, differential interference contrast (DIC) microscopy, interference reflection microscopy (also known as reflection interference contrast or RIC), single plane illumination microscopy (SPIM), super-resolution microscopy, laser microscopy, electron microscopy (EM), transmission electron microscopy (TEM), scanning electron microscopy (SEM), reflection electron microscopy (REM), scanning transmission electron microscopy (STEM), and low-voltage electron microscopy (LVEM), scanning probe microscopy (SPM), atomic force microscopy (ATM), ballistic electron emission microscopy (BEEM), chemical force microscopy (CFM), conductive atomic force microscopy (C-AFM), electrochemical scanning tunneling microscopy (ECSTM), electrostatic force microscopy (EFM), fluidic force microscopy (FluidFM), force modulation microscopy (FMM), feature-oriented scanning probe microscopy (FOSPM), Kelvin probe force microscopy (KPFM), magnetic force microscopy (MFM), magnetic resonance force microscopy (MRFM), near-field scanning optical microscopy (NSOM) (or SNOM, scanning near-field optical microscopy, SNOM), piezoresponse force microscopy (PFM), PSTM, photon scanning tunneling microscopy (PSTM), PTMS, photothermal conversion spectroscopy / microscopy (PTMS), SCM, scanning capacitance microscopy (SCM), SECM, scanning electrochemical microscopy (SECM), SGM, scanning gate microscopy (SGM), SHPM, scanning Hall probe microscopy (SHPM), SICM, scanning ion conductance microscopy (SICM), SPSM spin-polarized scanning tunneling microscopy (SPSM), SSRM, scanning spreading resistance microscopy (SSRM), SThM, scanning thermal microscopy (SThM), STM, scanning tunneling microscopy (STM), STP, scanning tunneling potentiometry (STP), SVM, scanning voltage microscopy (SVM), and synchrotron X-ray scanning tunneling microscopy (SXSTM), as well as label-free tissue expansion microscopy (exM).

[0233] In some embodiments, the methods herein include subjecting a sample to expansion microscopy and techniques. Expansion enables spatially resolving individual targets (e.g., mRNAs or RNA transcripts) that are densely packed within cells in a high-throughput manner. Expansion microscopy techniques are known in the art and can be carried out as described in US2016 / 0116384 and Chen et al., Science, 347, 543 (2015), each of which is incorporated herein by reference in its entirety. In some embodiments, the method does not include subjecting a sample to expansion microscopy. In some embodiments, the method does not include dissociating cells from a sample such as a tissue or cellular microenvironment. In some embodiments, the method does not include lysing the sample or cells therein. In some embodiments, the method does not include embedding the sample or molecules from the sample in an exogenous matrix.

[0234] V. Sample and Sample Processing The methods and compositions disclosed herein can be obtained from a subject using any of a variety of techniques including, but not limited to, biopsy, surgery, and laser capture microscopy (LCM), and can generally be used to analyze a biological sample comprising cells and / or other biological materials from the subject. Biological samples can also be obtained from eukaryotes such as tissue samples, patient-derived organoids (PDOs), or patient-derived xenografts (PDXs). Biological samples derived from an organism can contain one or more other organisms or components derived therefrom. For example, a mammalian tissue section can contain components derived from prions, viroids, viruses, bacteria, fungi, or other organisms in addition to mammalian cells and acellular tissue components.

[0235] Subjects from which biological samples can be obtained can be healthy or asymptomatic individuals, individuals having or suspected of having a disease (e.g., a patient having a disease such as cancer), or individuals predisposed to a disease, and / or individuals in need of or suspected of needing therapy.

[0236] In some embodiments, the biological sample corresponds to cells (e.g., from a cell culture, tissue sample, or cells deposited on a surface). In a cell sample having a plurality of cells, the individual cells can be naturally non-aggregated. For example, the cells can be from a cell suspension (e.g., a body fluid such as blood), and / or dissociated or non-aggregated cells from a tissue or tissue section. The number of cells in the biological sample can vary. Some biological samples contain a large number of cells, such as a blood sample, while other biological samples contain a smaller or only a few cells, or may only be suspected of containing cells, such as plasma, serum, urine, saliva, synovial fluid, amniotic fluid, tears, lymph fluid, liquor, cerebrospinal fluid, etc.

[0237] In some embodiments, the cell-containing biological sample includes a body fluid, or a cell-containing sample derived from a body fluid, such as a cell-containing sample derived from whole blood, a sample derived from blood such as plasma or serum, buffy coat, urine, sputum, tears, lymph fluid, sweat, liquor, cerebrospinal fluid, ascites, milk, feces, bronchoalveolar lavage fluid, saliva, amniotic fluid, nasal secretions, vaginal secretions, semen / seminal fluid, wound secretions, cell cultures, and swab samples, or any cell-containing sample derived from the aforementioned samples. In some embodiments, the cell-containing biological sample can be a body fluid, a body secretion or a body excretion, such as lymph fluid, blood, buffy coat, plasma, or serum. In some embodiments, the cell-containing biological sample can be a circulating body fluid such as blood or lymph fluid, such as peripheral blood obtained from a mammal such as a human.

[0238] A biological sample can contain any number of macromolecules, such as cellular macromolecules and organelles (e.g., mitochondria and nuclei). The biological sample can be obtained as a tissue sample, such as a tissue section, biopsy, core biopsy, needle aspirate, or fine needle aspirate. The sample can be a liquid sample, such as a blood sample, urine sample, or saliva sample. The sample can be a skin sample, colon sample, buccal swab, histological sample, histopathological sample, plasma or serum sample, tumor sample, live cells, cultured cells, e.g., clinical samples such as whole blood or blood-derived products, blood cells, or cultured tissue or cells containing a cell suspension. In some embodiments, the biological sample can contain cells deposited on a surface. In some embodiments, the biological sample can contain transcripts of antigen receptor molecules.

[0239] The biological sample can be derived from a homogeneous culture or population of the subject or organism referred to herein, or alternatively, from a collection of several different organisms in, e.g., a community or ecosystem.

[0240] The biological sample can contain one or more diseased cells. The diseased cells can have altered metabolic properties, gene expression, protein expression, and / or morphological features. Examples of diseases include inflammatory disorders, metabolic disorders, nervous system disorders, and cancer. Cancer cells can be derived from solid tumors, hematological malignancies, cell lines, or obtained as circulating tumor cells. The biological sample can also contain fetal cells and immune cells.

[0241] The biological sample can contain an analyte (e.g., protein, RNA, and / or DNA) embedded in a 3D matrix. In some embodiments, amplicons (e.g., rolling circle amplification products) derived from or related to the analyte (e.g., protein, RNA, and / or DNA) can be embedded in the 3D matrix. In some embodiments, the 3D matrix can contain a network of natural and / or synthetic molecules chemically and / or enzymatically linked, e.g., by crosslinking. In some embodiments, the 3D matrix can contain a synthetic polymer. In some embodiments, the 3D matrix contains a hydrogel.

[0242] In some embodiments, the substrate herein can be any support that is insoluble in aqueous liquids and enables the positioning of a biological sample, an analyte, a feature, and / or a reagent (e.g., a probe) on the support. In some embodiments, the biological sample can adhere to the substrate. The adhesion of the biological sample can be irreversible or reversible, depending on the nature of the sample and the subsequent steps of the analysis method. In certain embodiments, by applying a suitable polymer coating to the substrate and contacting the sample with the polymer coating, the sample can adhere reversibly to the substrate. Then, for example, using an organic solvent that at least partially dissolves the polymer coating, the sample can be separated from the substrate. A hydrogel is an example of a polymer suitable for this purpose.

[0243] In some embodiments, the substrate can be coated or functionalized with one or more substances to facilitate the adhesion of the sample to the substrate. Suitable substances that can be used to coat or functionalize the substrate include, but are not limited to, lectin, poly-lysine, antibodies, and polysaccharides.

[0244] Various steps can be used for the assay and / or to prepare or process a biological sample during the assay. Unless otherwise indicated, the preparation or processing steps described below can generally be combined in any manner and in any order and / or can appropriately prepare or process a specific sample for analysis.

[0245] (i) Tissue sectioning Biological samples can be collected from a subject (e.g., via surgical biopsy, whole subject sectioning), or grown in vitro as a cell population on a growth substrate or culture dish and prepared for analysis as a tissue slice or section. The grown sample can be thin enough for analysis without further processing steps. Alternatively, the grown sample, as well as samples obtained via biopsy or sectioning, can be prepared as thin tissue sections using a mechanical cutting device such as a vibrating blade microtome. As another alternative, in some embodiments, thin tissue sections can be prepared by applying a touch imprint of the biological sample to a suitable substrate material.

[0246] The thickness of the tissue section can be a fraction of the maximum cross-sectional dimension of the cell (e.g., less than 0.9, 0.8, 0.7, 0.6, 0.5, 0.4, 0.3, 0.2, or 0.1). However, tissue sections having a thickness greater than the maximum cross-sectional cell dimension can also be used. For example, cryostat sections can be used, which can be, for example, 10 - 20 μm thick.

[0247] More generally, the thickness of the tissue section typically depends on the method used to prepare the section and the physical characteristics of the tissue, and thus sections having a wide variety of different thicknesses can be prepared and used. For example, the thickness of the tissue section can be at least 0.1, 0.2, 0.3, 0.4, 0.5, 0.7, 1.0, 1.5, 2, 3, 4, 5, 6, 7, 8, 9, 10, 12, 13, 14, 15, 20, 30, 40, or 50 μm. Thicker sections, for example, at least 70, 80, 90, or 100 μm or more, can also be used if desired or convenient. Typically, the thickness of the tissue section is 1 - 100 μm, 1 - 50 μm, 1 - 30 μm, 1 - 25 μm, 1 - 20 μm, 1 - 15 μm, 1 - 10 μm, 2 - 8 μm, 3 - 7 μm, or 4 - 6 μm, but as noted above, sections having a thickness greater than or less than these ranges can also be analyzed.

[0248] Multiple sections can also be obtained from a single biological sample. For example, multiple tissue sections can be obtained from a surgical biopsy sample by performing serial sectioning of the biopsy sample using a blade for sectioning. The spatial information between serial sections can be thus preserved, and the sections can be serially analyzed to obtain three-dimensional information regarding the biological sample.

[0249] (ii) Freezing In some embodiments, a biological sample (e.g., the tissue sections described above) can be prepared by cryogenic freezing at a temperature suitable for maintaining or preserving the integrity of the tissue structure (e.g., physical characteristics). The frozen tissue sample can be sectioned, e.g., thinly sliced, on a substrate surface using any number of suitable methods. For example, the tissue sample can be prepared using a cryomicrotome (e.g., a cryostat) set to a temperature suitable for maintaining both the structural integrity of the tissue sample and the chemical properties of the nucleic acids in the sample. Such a temperature can be, for example, less than -15°C, less than -20°C, or less than -25°C.

[0250] (iii) Fixation and Post-Fixation In some embodiments, a biological sample can be prepared using formalin fixation and paraffin embedding (FFPE), which is an established method. In some embodiments, cell suspensions and other non-tissue samples can be prepared using formalin fixation and paraffin embedding. After fixation of the sample and embedding in a paraffin or resin block, the sample can be sectioned as described above. Prior to analysis, the paraffin-embedded material can be removed (e.g., deparaffinized) from the tissue section by incubating the tissue section in a suitable solvent (e.g., xylene), followed by rinsing (e.g., 99.5% ethanol for 2 minutes, 96% ethanol for 2 minutes, and 70% ethanol for 2 minutes).

[0251] As an alternative to the above-described formalin fixation, the biological sample can be fixed in any of a variety of other fixatives to preserve the biological structure of the sample prior to analysis. For example, the sample can be fixed via immersion in ethanol, methanol, acetone, paraformaldehyde (PFA)-Triton, and combinations thereof.

[0252] In some embodiments, acetone fixation is used with fresh frozen samples, which can include, but are not limited to, cortical tissue, mouse olfactory bulb, human brain tumor, human postmortem brain, and breast cancer samples. If acetone fixation is performed, the pre-permeabilization step (described below) may not be performed. Alternatively, acetone fixation can be performed in conjunction with the permeabilization step.

[0253] In some embodiments, the methods provided herein include one or more post-fixing (also referred to as postfixation) steps. In some embodiments, the one or more post-fixing steps are performed after contacting the sample with one or more probes, such as polynucleotides disclosed herein, e.g., circular probes or padlock probes. In some embodiments, the one or more post-fixing steps are performed after a hybridization complex comprising the probe and the target has been formed in the sample. In some embodiments, the one or more post-fixing steps are performed prior to a ligation reaction disclosed herein, such as ligation to circularize a padlock probe.

[0254] In some embodiments, the one or more post-fixing steps are performed after contacting the sample with a binder or labeler (e.g., an antibody or an antigen-binding fragment thereof) for non-nucleic acid analytes such as protein analytes. The labeler can include a nucleic acid molecule (e.g., a reporter oligonucleotide) that contains a sequence corresponding to the labeler and thus corresponding to (e.g., uniquely identifying) the analyte. In some embodiments, the labeler can include a reporter oligonucleotide that contains one or more barcode sequences.

[0255] The post-fixation step can be carried out using any suitable fixation reagent disclosed herein, such as 3% (w / v) paraformaldehyde in DEPC-PBS.

[0256] (iv) Embedding As an alternative to the above paraffin embedding, the biological sample can be embedded in any of a variety of other embedding materials to provide a structural substrate to the sample prior to sectioning and other handling steps. In some cases, the embedding material can be removed, for example, prior to analysis of tissue sections obtained from the sample. Suitable embedding materials include, but are not limited to, wax, resins (such as methacrylate resins), epoxies, and agar.

[0257] In some embodiments, the biological sample can be embedded in a matrix (such as a hydrogel matrix). Embedding the sample in this way typically involves contacting the biological sample with the hydrogel such that the biological sample is surrounded by the hydrogel. For example, the sample can be embedded by contacting the sample with a suitable polymeric material and activating the polymeric material to form a hydrogel. In some embodiments, the hydrogel is formed such that the hydrogel is internalized within the biological sample.

[0258] In some embodiments, the biological sample is immobilized in the hydrogel via crosslinking of the polymeric material forming the hydrogel. Crosslinking can be carried out chemically and / or photochemically, or alternatively by any other hydrogel-forming method.

[0259] The composition and application of the hydrogel-matrix to a biological sample typically depend on the nature and preparation of the biological sample (e.g., sectioned, non-sectioned, type of fixation). As an example, when the biological sample is a tissue section, the hydrogel-matrix can include a monomer solution and an ammonium persulfate (APS) initiator / tetramethylethylenediamine (TEMED) accelerator solution. As another example, when the biological sample consists of cells (e.g., cultured cells or cells dissociated from a tissue sample), the cells can be incubated with the monomer solution and the APS / TEMED solution. For cells, the hydrogel-matrix gel is formed within a compartment including, but not limited to, devices used to culture, maintain, or transport the cells. For example, the hydrogel-matrix can be formed using the monomer solution and APS / TEMED added to the compartment to a depth in the range of about 0.1 μm to about 2 mm.

[0260] Additional methods and embodiments of hydrogel embedding of biological samples are described, for example, in Chen et al., Science 347(6221):543-548, 2015, the entire content of which is incorporated herein by reference.

[0261] (v) Staining and immunohistochemistry (IHC) To facilitate visualization, a biological sample can be stained using a wide variety of stains and staining techniques. In some embodiments, for example, the sample can be stained using any number of dyes and / or immunohistochemical reagents. One or more staining steps can be performed to prepare or process the biological sample for the assays described herein, or can be performed during and / or after the assay. In some embodiments, the sample can be contacted with one or more nucleic acid stains, membrane stains (e.g., cell membrane or nuclear membrane), cytological stains, or combinations thereof. In some examples, the stain can be specific for proteins, phospholipids, DNA (e.g., dsDNA, ssDNA), RNA, organelles, or cellular compartments. The sample can be contacted with one or more labeled antibodies (e.g., a primary antibody specific for an analyte of interest and a labeled secondary antibody specific for the primary antibody). In some embodiments, cells in the sample can be segmented using one or more captured images of the stained sample.

[0262] In some embodiments, staining is performed using lipophilic dyes. In some examples, staining is performed using lipophilic carbocyanine or aminostyryl dyes, or analogs thereof (e.g., DiI, DiO, DiR, DiD). Other cell membrane stains can include FM and RH dyes or immunohistochemical reagents specific for cell membrane proteins. In some examples, stains can include, but are not limited to, acridine orange, acid fuchsin, Bismarck brown, carmine, Coomassie blue, cresyl violet, DAPI, eosin, ethidium bromide, acid fuchsine, hematoxylin, Hecht stain, iodine, methyl green, methylene blue, neutral red, Nile blue, Nile red, osmium tetroxide, ruthenium red, propidium iodide, rhodamine (e.g., rhodamine B), or safranin, or derivatives thereof. In some embodiments, the sample can be stained with hematoxylin and eosin (H&E).

[0263] The sample can be stained using hematoxylin and eosin (H&E) staining techniques, Papanicolaou staining techniques, Masson trichrome staining techniques, silver staining techniques, Sudan staining techniques, and / or periodic acid Schiff (PAS) staining techniques. PAS staining is typically performed after formalin or acetone fixation. In some embodiments, the sample can be stained using Romanowsky staining, including Wright staining, Giemsa staining, Can-Grunwald staining, Leishman staining, and Giemsa staining.

[0264] In some embodiments, the biological sample can be de-stained. The method for de-staining or discoloring the biological sample generally depends on the nature of the staining applied to the sample. For example, in some embodiments, one or more immunofluorescent stains are applied to the sample via antibody coupling. Such stains can be removed using techniques such as treatment with reducing agents and detergent washes, chaotropic salt treatment, treatment with antigen retrieval solutions, and cleavage of disulfide linkages via treatment with acidic glycine buffer. Methods for multiplex staining and de-staining are described, for example, in Bolognesi et al., J. Histochem. Cytochem. 2017;65(8):431-444, Lin et al., Nat Commun. 2015;6:8390, Pirici et al., J. Histochem. Cytochem. 2009;57:567-75, and Glass et al., J. Histochem. Cytochem. 2009;57:899-905, the entire contents of each of which are incorporated herein by reference.

[0265] (vi) Isometric expansion In some embodiments, a biological sample embedded in a matrix (e.g., a hydrogel) can expand isometrically. As described in Chen et al., Science 347(6221):543-548, 2015, an isometric expansion method that can be used is hydration, which is a preparation step in expansion microscopy.

[0266] Isometric expansion can be carried out by fixing one or more components of a biological sample in a gel and subsequently performing gel formation, proteolysis, and swelling. In some embodiments, an analyte in the sample, a product of the analyte, and / or a probe related to the analyte in the sample can be immobilized in a matrix (e.g., a hydrogel). Isometric expansion of the biological sample can occur before the biological sample is immobilized on a substrate or after the biological sample has been immobilized on the substrate. In some embodiments, the isometrically expanded biological sample can be removed from the substrate before contacting the substrate with the probes disclosed herein.

[0267] Generally, the steps used to perform isometric expansion of a biological sample can depend on the characteristics of the sample (e.g., thickness of a tissue section, fixation, crosslinking), and / or the analyte of interest (e.g., different conditions for immobilizing RNA, DNA, and proteins in a gel).

[0268] In some embodiments, proteins in a biological sample are immobilized in a swellable gel such as a polyelectrolyte gel. Antibodies can be directed against the proteins either before, after, or in conjunction with immobilization in the swellable gel. DNA and / or RNA in the biological sample can also be immobilized in a swellable gel via a suitable linker. Examples of such linkers include 6-((acryloyl)amino)hexanoic acid (Acryloyl-X SE) (available from ThermoFisher, Waltham, MA), Label-IT Amine (available from MirusBio, Madison, WI), and Label X (e.g., as described in Chen et al., Nat. Methods 13:679-684, 2016, the entire content of which is incorporated herein by reference), but are not limited thereto.

[0269] Isometric expansion of the sample can increase the spatial resolution of subsequent analysis of the sample. The increase in resolution in spatial profiling can be determined by comparison of the isometrically expanded sample to an unexpanded sample.

[0270] In some embodiments, the biological sample is isometrically expanded to a size that is at least 2-fold, 2.1-fold, 2.2-fold, 2.3-fold, 2.4-fold, 2.5-fold, 2.6-fold, 2.7-fold, 2.8-fold, 2.9-fold, 3-fold, 3.1-fold, 3.2-fold, 3.3-fold, 3.4-fold, 3.5-fold, 3.6-fold, 3.7-fold, 3.8-fold, 3.9-fold, 4-fold, 4.1-fold, 4.2-fold, 4.3-fold, 4.4-fold, 4.5-fold, 4.6-fold, 4.7-fold, 4.8-fold, or 4.9-fold the size of its non-expanded size. In some embodiments, the sample is isometrically expanded to at least 2-fold and less than 20-fold the size of its non-expanded size.

[0271] (vii) Crosslinking and de-crosslinking In some embodiments, the biological sample is reversibly crosslinked before or during the in situ assay. In some aspects, the analyte, the polynucleotide of the analyte and / or the amplification product (e.g., amplicon), or a probe bound thereto can be immobilized in the polymer matrix. For example, the polymer matrix can be a hydrogel. In some embodiments, one or more of the polynucleotide probes and / or its amplification products can be modified to contain a functional group that can be used as an anchor site for attaching the polynucleotide probe and / or the amplification product (e.g., amplicon) to the polymer matrix. In some embodiments, a modified probe containing oligo dT can be used to bind to the mRNA molecule of interest, followed by reversible crosslinking of the mRNA molecule.

[0272] The hydrogel can include a polymeric polymer gel including a network. Within the network, some of the polymer chains can be optionally crosslinked, but crosslinking does not necessarily occur.

[0273] In some embodiments, the hydrogel can include a hydrogel subunit, including, but not limited to, acrylamide, bis-acrylamide, polyacrylamide and its derivatives, poly(ethylene glycol) and its derivatives (e.g., PEG-acrylate (PEG-DA), PEG-RGD), gelatin-methacryloyl (GelMA), methacrylated hyaluronic acid (MeHA), polyaliphatic polyurethane, polyether polyurethane, polyester polyurethane, polyethylene copolymer, polyamide, polyvinyl alcohol, polypropylene glycol, polytetramethylene oxide, polyvinyl pyrrolidone, polyacrylamide, poly(hydroxyethyl acrylate), and poly(hydroxyethyl methacrylate), collagen, hyaluronic acid, chitosan, dextran, agarose, gelatin, alginate, protein polymer, methylcellulose, etc., and combinations thereof.

[0274] In some embodiments, the hydrogel includes a hybrid material, e.g., the hydrogel material includes elements of both synthetic polymers and natural polymers. Examples of suitable hydrogels are described, for example, in U.S. Patent Nos. 6,391,937, 9,512,422, and 9,889,422, and U.S. Patent Application Publication Nos. 2017 / 0253918, 2018 / 0052081, and 2010 / 0055733, the entire contents of each of which are incorporated herein by reference.

[0275] In some embodiments, the hydrogel can form a substrate. In some embodiments, the substrate includes the hydrogel and one or more second materials. In some embodiments, the hydrogel is disposed on top of one or more second materials. For example, the hydrogel can be pre-formed and then disposed on, under, or within any other configuration having one or more second materials. In some embodiments, hydrogel formation occurs after contact with one or more second materials during substrate formation. Hydrogel formation can also occur within structures (e.g., wells, ridges, protrusions, and / or markings) located on the substrate.

[0276] In some embodiments, hydrogel formation on the substrate occurs before, simultaneously with, or after the probe is provided to the sample. For example, hydrogel formation can be performed on a substrate that already contains the probe.

[0277] In some embodiments, hydrogel formation occurs within the biological sample. In some embodiments, the biological sample (e.g., tissue section) is embedded in the hydrogel. In some embodiments, the hydrogel subunits are injected into the biological sample and the polymerization of the hydrogel is initiated by an external or internal stimulus.

[0278] In embodiments where the hydrogel is formed within the biological sample, functionalization chemistries can be used. In some embodiments, the functionalization chemistries include hydrogel-tissue chemistry (HTC). Any hydrogel-tissue backbone (e.g., synthetic or natural) suitable for HTC can be used to immobilize the biopolymer and regulate the functionalization. Non-limiting examples of methods using HTC backbone variants include CLARITY, PACT, ExM, SWITCH, and ePACT. In some embodiments, the hydrogel formation within the biological sample is permanent. For example, the biopolymer can be permanently adhered to the hydrogel, allowing for multiple rounds of interrogation. In some embodiments, the hydrogel formation within the biological sample is reversible.

[0279] In some embodiments, additional reagents are added to the hydrogel subunits before, simultaneously with, and / or after polymerization. For example, additional reagents can include, but are not limited to, oligonucleotides (e.g., probes), endonucleases to fragment DNA, fragmentation buffer for DNA, DNA polymerase enzymes, dNTPs used to amplify nucleic acids and attach to amplified fragments with barcodes. Other enzymes including, but not limited to, RNA polymerase, ligase, proteinase K, and DNAse can be used. Additional reagents can also include reverse transcriptase, including enzymes having terminal transferase activity, primers, and switch oligonucleotides. In some embodiments, optical labels are added to the hydrogel subunits before, simultaneously with, and / or after polymerization.

[0280] In some embodiments, HTC reagents are added to the hydrogel before, simultaneously with, and / or after polymerization. In some embodiments, cell labeling agents are added to the hydrogel before, simultaneously with, and / or after polymerization. In some embodiments, cell membrane permeabilizing agents are added to the hydrogel before, simultaneously with, and / or after polymerization.

[0281] The hydrogel embedded within the biological sample can be removed using any suitable method. For example, an electrophoretic tissue removal method can be used to remove biopolymers from the hydrogel-embedded sample. In some embodiments, the hydrogel-embedded sample is stored in a medium (e.g., an encapsulant, methylcellulose, or other semi-solid medium) before or after removal of the hydrogel.

[0282] In some embodiments, the methods disclosed herein include depolymerizing a reversibly cross-linked biological sample. The depolymerization need not be complete. In some embodiments, only a portion of the cross-linked molecules in the reversibly cross-linked biological sample are depolymerized and become mobile.

[0283] (viii) Tissue permeabilization and treatment In some embodiments, the biological sample can be permeabilized to facilitate the movement of species (such as probes) into the sample. If the sample is not sufficiently permeabilized, the amount of species (such as probes) in the sample may be too low to allow for proper analysis. Conversely, if the tissue sample is too permeable, the relative spatial relationships of the analytes within the tissue sample may be lost. Thus, a balance is desired between permeabilizing the tissue sample enough to obtain good signal strength while maintaining the spatial resolution of the analyte distribution in the sample.

[0284] Generally, a biological sample can be permeabilized by exposing the sample to one or more permeabilizing agents. Suitable agents for this purpose include, but are not limited to, organic solvents (such as acetone, ethanol, and methanol), crosslinking agents (such as paraformaldehyde), surfactants (such as saponin, Triton X-100™, or Tween-20™), and enzymes (such as trypsin, protease). In some embodiments, the biological sample can be incubated with a cell permeabilizing agent to facilitate permeabilization of the sample. Additional methods for sample permeabilization are described, for example, in Jamur et al., Method Mol. Biol. 588:63-66, 2010, the entire contents of which are incorporated herein by reference. Any suitable method for sample permeabilization can generally be used in connection with the samples described herein.

[0285] In some embodiments, the biological sample can be permeabilized by adding one or more lysis reagents to the sample. Examples of suitable lysis agents include, but are not limited to, bioactive reagents such as lysozyme, achromopeptidase, lysostaphin, labiase, kitase, lithicase, and various other commercially available lysis enzymes that are used for the lysis of different cell types such as gram-positive or gram-negative bacteria, plants, yeast, mammals, etc.

[0286] Other lysing agents can be added to the biological sample additionally or alternatively to facilitate permeabilization. For example, a surfactant-based lysis solution can be used to lyse the sample cells. The lysis solution can include, for example, ionic surfactants such as sarcosyl and sodium dodecyl sulfate (SDS). More generally, chemical lysing agents can include, but are not limited to, organic solvents, chelating agents, detergents, surfactants, and chaotropic agents.

[0287] In some embodiments, the biological sample can be permeabilized by non-chemical permeabilization methods. Non-chemical permeabilization methods that can be used include physical lysis techniques such as electroporation, mechanical permeabilization methods (e.g., bead beating using a homogenizer and grinding balls to mechanically disrupt the sample tissue structure), sonoporation (e.g., sonication), and thermal lysis techniques such as heating to induce thermal permeabilization of the sample, but are not limited thereto.

[0288] Additional reagents can be added to the biological sample to perform various functions prior to analysis of the sample. In some embodiments, a DNase and RNase inactivator or inhibitor such as proteinase K, and / or a chelating agent such as EDTA can be added to the sample. For example, the methods disclosed herein can include steps to enhance the accessibility of nucleic acids for binding, such as a denaturation step to open intracellular DNA for hybridization by a probe. For example, proteinase K treatment can be used to release DNA bound by proteins.

[0289] (ix) Selective enrichment of RNA species or cDNA species In some embodiments where the analyte is RNA or cDNA, one or more RNA or cDNA analyte species of interest can be selectively enriched. For example, one or more RNA or cDNA species of interest can be selected by adding one or more oligonucleotides to the sample. In some embodiments, the additional oligonucleotides are sequences used to prime a reaction by an enzyme (e.g., polymerase). For example, one or more primer sequences having sequence complementarity to one or more RNA or cDNA of interest can be used to amplify the one or more RNA or cDNA of interest, thereby selectively enriching these RNA or cDNA.

[0290] In some aspects, when two or more analytes are analyzed, first and second probes that are specific for (e.g., specifically hybridize to) each RNA or cDNA analyte are used. For example, in some embodiments of the methods provided herein, template ligation is used to detect gene expression in a biological sample. An analyte of interest (such as a protein) that is bound by a labeling agent or binding agent (e.g., an antibody or an epitope-binding fragment thereof), where the binding agent is conjugated to or otherwise associated with a reporter oligonucleotide that contains a reporter sequence that identifies the binding agent, can be targeted for analysis. The probe can hybridize to the reporter oligonucleotide and be ligated in a template ligation reaction to produce a product for analysis. In some embodiments, the gap between the probe oligonucleotides can be first filled prior to ligation using, for example, Mu polymerase, DNA polymerase, RNA polymerase, reverse transcriptase, VENT polymerase, Taq polymerase, and / or any combination, derivative, and variant thereof (e.g., engineered mutants). In some embodiments, the assay can further include amplification (e.g., by multiplex PCR) of the template ligation product.

[0291] In some embodiments, the analyte can be further concentrated for in situ readout by immobilization at a location in the biological sample. In non-limiting examples, the analyte can include one or more fragments that are specific to a location in the biological sample.

[0292] Alternatively, one or more species of RNA can be narrowed down (e.g., removed) using any of a variety of methods. For example, probes can be administered to a sample that selectively hybridize to ribosomal RNA (rRNA), thereby reducing the pool and concentration of rRNA in the sample. Additionally and alternatively, double-stranded specific nuclease (DSN) treatment can remove rRNA (see, e.g., Archer, et al, Selective and flexible depletion of problematic sequences from RNA-seq libraries at the cDNA stage, BMC Genomics, 15 401, (2014), the entire content of which is incorporated herein by reference). Further, hydroxyapatite chromatography can remove abundant species (e.g., rRNA) (see, e.g., Vandernoot, V.A., cDNA normalization by hydroxyapatite chromatography to enrich transcriptome diversity in RNA-seq applications, Biotechniques, 53(6)373-80, (2012), the entire content of which is incorporated herein by reference).

[0293] The biological sample can contain one or more analytes of interest. A method for performing a multiplexed assay for analyzing two or more different analytes in a single biological sample is provided.

[0294] VI. Compositions, Kits, and Systems For example, a kit is provided herein that includes one or more oligonucleotides, such as any of those described in Sections I-V, and instructions for carrying out the methods provided herein. In some embodiments, the kit further includes one or more reagents (e.g., nucleotide mixes, probes, etc.) for carrying out the methods provided herein. In some embodiments, the kit further includes one or more reagents necessary for one or more steps including hybridization, ligation, extension, amplification, detection, and / or sample preparation as described herein. In some embodiments, the kit further includes an enzyme such as a ligase and / or polymerase as described herein. In some embodiments, the kit includes a polymerase for performing primer extension and incorporating nucleotides, for example. In some embodiments, the kit may contain reagents for forming a functionalized matrix (e.g., a hydrogel) such as any suitable functional moiety. In some examples, buffers and reagents for tethering probes and products (e.g., RCA products) to the functionalized matrix are also provided. The various components of the kit may be present in separate containers or certain compatible components may be pre-combined in a single container. In some embodiments, the kit further contains instructions for carrying out the methods provided using the components of the kit.

[0295] In some embodiments, the kit may contain the reagents and / or consumables necessary to perform one or more steps of the provided methods. In some embodiments, the kit contains reagents for fixing, embedding, and / or permeabilizing a biological sample. In some embodiments, the kit contains reagents such as enzymes and buffers for ligation and / or amplification, such as ligase and / or polymerase. In some aspects, the kit may also include any of the reagents described herein, such as a wash buffer and a ligation buffer. In some embodiments, the kit contains reagents for detection and / or sequencing, such as detectably labeled nucleotides, polymerase, or conjugate. In some embodiments, the kit optionally contains other components such as nucleic acid primers, enzymes and reagents, buffers, nucleotides, modified nucleotides, reagents for additional assays, etc.

[0296] VII. Microfluidic Devices for Analysis of Biological Samples As described herein, a device having integrated optical and fluidic modules (a "microfluidic device" or "microfluidic system") for detecting target molecules (e.g., nucleic acids, proteins, antibodies, etc.) in a biological sample (e.g., one or more cell or tissue samples) is provided herein. In the microfluidic device, the fluidic module is configured to deliver one or more reagents (e.g., detectably labeled nucleotides, polymerase, or conjugate) to the biological sample and / or remove used reagents therefrom. Further, the optical module is configured to irradiate the biological sample with light having one or more spectral emission profiles (over a range of wavelengths) and then capture one or more images of the emitted light signal from the biological sample during one or more sequencing cycles (e.g., as described in Section III). In some embodiments, the in situ assays (e.g., sequencing by synthesis) disclosed herein may be performed using an automated instrument or system, such as the microfluidic device or system disclosed herein.

[0297] In various embodiments, the captured images can be processed in real time and / or at a later time to determine the presence of one or more target molecules in the biological sample and the three-dimensional position information associated with each detected target molecule. Additionally, the optofluidic device includes a sample module configured to receive (and optionally immobilize) one or more biological samples. In some cases, the sample module includes an X-Y stage configured to move the biological sample along the X-Y plane (e.g., perpendicular to the objective lens of the optical module).

[0298] In various embodiments, the optofluidic device is configured to analyze one or more target molecules (e.g., any of the analytes described in Section III) at their native positions (i.e., in-situ) within the biological sample. For example, the optofluidic device can be an in-situ analysis system used to analyze a biological sample and detect target molecules (e.g., analytes) including, but not limited to, DNA, RNA, proteins, antibodies, and / or the like.

[0299] An optofluidic device that can be used for in-situ target molecule detection via base-by-base sequencing (e.g., sequencing of an identifier sequence such as a barcode sequence) and / or other imaging or target molecule detection techniques. That is, for example, the optofluidic device can include a fluid module that includes the fluids required to establish the experimental conditions necessary for a probe of the target molecule in the sample. Additionally, such an optofluidic device can also include a sample module configured to receive the sample and an optical module that includes an imaging system for irradiating (e.g., exciting one or more fluorescently labeled nucleotides within the sample) and / or imaging the optical signal received from the sample. The in-situ analysis system can also include other auxiliary modules configured to facilitate the operation of the optofluidic device, such as, but not limited to, a cooling system, an operation calibration system, and the like.

[0300] FIG. 5 shows an exemplary analysis workflow of a biological sample 510 (e.g., a cell or tissue sample) using an optofluidic device or system 500 according to various embodiments. In various embodiments, the sample 510 can be a biological sample (e.g., a tissue) containing molecules such as DNA, RNA, proteins, antibodies, etc. For example, the sample 510 can be a sectioned tissue that has been processed to access its RNA for probe hybridization and sequencing as described herein (e.g., in Section III).

[0301] In various embodiments, the sample 510 can be placed within an optofluidic device or system 500 for analysis and detection of molecules in the sample 510. In various embodiments, the optofluidic device or system 500 can be a system configured to facilitate experimental conditions that aid in the detection of target molecules. For example, the optofluidic device or system 500 can include a fluid module 540, an optical module 550, a sample module 560, and an auxiliary module 570, which are operated by a system controller 530 to create experimental conditions for base-by-base sequencing of nucleic acid molecules in the sample 510 and to facilitate imaging of the sample (e.g., by the imaging system of the optical module 550). In various embodiments, the various modules of the optofluidic device or system 500 can be separate components that communicate with each other or at least some of them can be integrated together.

[0302] In various embodiments, the sample module 550 can be configured to receive the sample 510 within the optofluidic device or system 500. For example, the sample module 560 can include a sample interface module (SIM) configured to receive a sample device (e.g., a cassette) on which the sample 510 can be deposited. That is, the sample 510 (e.g., a sectioned tissue) can be placed within the optofluidic device or system 500 by depositing the sample 510 on a sample device that is then inserted into the SIM of the sample module 560. In some cases, the sample module 560 can also include an X-Y stage to which the SIM is attached. The X-Y stage can be configured to move the SIM attached thereto (e.g., and thus the sample device containing the sample 510 inserted therein) vertically along a two-dimensional (2D) plane of the optofluidic device or system 500.

[0303] The experimental conditions that assist in the detection of molecules in the sample 510 can depend on the target molecule detection technique used by the optofluidic device or system 500. For example, in various embodiments, the optofluidic device or system 500 can be a system configured to detect molecules in the sample 510 (e.g., by detecting nucleotides incorporated into an extension sequencing primer using an identifier sequence as a template).

[0304] In various embodiments, the fluid module 540 can include one or more components that can be used to store a reagent and to transport the reagent to and from a sample device containing a sample 510. For example, the fluid module 540 can include a reservoir configured to store a reagent and a waste container configured to collect the reagent (e.g., and other waste) after use by the optofluidic device or system 500 for analyzing and detecting molecules of the sample 510. Additionally, the fluid module 540 can also include pumps, tubes, pipettes, etc. configured to facilitate the transport of the reagent to the sample device (e.g., sample 510). For example, the fluid module 540 can include a pump (a "reagent pump") configured to pump a wash / stripping reagent into the sample device for use in washing / stripping the sample 510 (e.g., and other washing functions such as washing the objective lens of the imaging system of the optical module 550).

[0305] In various embodiments, the auxiliary module 570 can be a cooling system of the optofluidic device or system 500, and the cooling system can include a network of coolant-carrying tubes configured to transport coolant to various modules of the optofluidic device or system 500 to regulate the temperature of those modules. In such cases, the fluid module 540 includes a coolant reservoir for storing coolant and a pump (e.g., a "coolant pump") for generating a differential pressure, thereby allowing the coolant to flow from the reservoir through the coolant-carrying tubes to various modules of the optofluidic device or system 500. In some cases, the fluid module 540 can include a return coolant reservoir configured to receive and store the heated coolant that has absorbed heat released by various modules of the optofluidic device or system 500 and then flows back to the return coolant reservoir. In such cases, the fluid module 540 can also include a cooling fan configured to send air (e.g., cold air and / or ambient air) into the return coolant reservoir to cool the heated coolant stored therein. In some cases, the fluid module 540 can also include a cooling fan configured to directly send air into components of the optofluidic device or system 500 to cool those components. For example, the fluid module 540 can include a cooling fan configured to direct cold air or ambient air into the system controller 530 to cool it.

[0306] As described above, the optofluidic device or system 500 can include an optical module 550 that includes various optical components of the optofluidic device or system 500, such as, but not limited to, a camera, an illumination module (e.g., an LED), an objective lens, and / or the like. The optical module 550 can include a fluorescence imaging system configured to image that fluorescence emitted by detectably labeled nucleotides after the detectable label has been excited by light from the illumination module of the optical module 550 and incorporated into the extended sequencing primers in the sample 510.

[0307] In some cases, the optical module 550 may also include an optical frame to which the camera, the illumination module, and / or the X-Y stage of the sample module 560 can be attached.

[0308] In various embodiments, the system controller 530 may be configured to control the operation of the optofluidic device or system 500 (e.g., and the operation of one or more of its modules). In some cases, the system controller 530 may take various forms, including a processor, a single computer (or computer system), or multiple computers communicating with each other. In various embodiments, the system controller 530 may be communicatively coupled to data storage, a set of input devices, a display system, or a combination thereof. In some cases, some or all of these components may be considered to be part of the system controller 530, or otherwise integrated with it, or may be separate components communicating with each other, or may be integrated together. In other examples, the system controller 530 may be, or communicate with, a cloud computing platform.

[0309] In various embodiments, the optofluidic device or system 500 can analyze a sample 510 and generate an output 590 that includes an indication of the presence of a target molecule in the sample 510. For example, with respect to the exemplary embodiments described above where the optofluidic device or system 500 uses hybridization techniques for detecting molecules, the optofluidic device or system 500 can subject the sample 510 to successive sequencing cycles, and during the same sequencing cycle, the sample is imaged to detect signals associated with nucleotide binding and / or incorporation events at at least some positions in the sample 510 (e.g., positions where one or more genes in gene block 1 shown in FIG. 3 are present), as well as the absence of signals at other positions in the sample (e.g., positions where one or more genes in gene block 2 shown in FIG. 3 are present) (e.g., the absence of a signal can result from the use of unmodified nucleotides as dark bases or the absence of an agent capable of detecting incorporated nucleotides). In such cases, the output 590 can include an optical signature (e.g., a codeword) specific to each identifier sequence (e.g., a barcode sequence) that enables identification of the target molecule.

[0310] VIII. Terms Unless otherwise defined, all specialized terms, notations, and other technical and scientific terms or terminology used herein are intended to have the same meaning as commonly understood by one of ordinary skill in the art to which the claimed subject matter pertains. In some cases, terms having commonly understood meanings are defined herein for clarity and / or for ready reference, and the inclusion of such definitions herein should not necessarily be construed as representing a substantial difference from what is commonly understood in the art.

[0311] As used interchangeably herein, the terms "polynucleotide", "polynucleotide", and "nucleic acid molecule" refer to polymeric forms of nucleotides of any length, either ribonucleotides or deoxyribonucleotides. Thus, the term includes single-stranded, double-stranded, or multi-stranded DNA or RNA, genomic DNA, cDNA, DNA-RNA hybrids, or polymers containing purine and pyrimidine bases, or other natural, chemically or biochemically modified, non-natural, or derivatized nucleotide bases, but is not limited thereto. The backbone of the polynucleotide can include sugar and phosphate groups (such as can be found in RNA or DNA typically), or modified or substituted sugar or phosphate groups.

[0312] As used herein, "hybridization" can refer to the process by which two single-stranded polynucleotides non-covalently associate to form a stable double-stranded polynucleotide. In one aspect, the resulting double-stranded polynucleotide can be a "hybrid" or "duplex". "Hybridization conditions" typically include a salt concentration of less than approximately 1 M, often less than about 500 mM, and may be less than about 200 mM. A "hybridization buffer" includes a buffered salt solution such as 5% SSPE or other such buffers. The hybridization temperature can be as low as about 5°C, but is typically greater than 22°C, more typically greater than about 30°C, and typically greater than 37°C. Hybridization is often carried out under stringent conditions, i.e., conditions under which the sequence hybridizes to its target sequence but not to other non-complementary sequences. Stringent conditions are sequence-dependent and vary in different circumstances. For example, longer fragments may require a higher hybridization temperature for a particular hybridization than shorter fragments. Other factors, including the base composition and length of the complementary strands, the presence of organic solvents, and the degree of base mismatches, can affect the stringency of hybridization, so the combination of parameters is more important than the absolute measure of any one parameter alone. Generally, stringent conditions are selected to be about 5°C lower than the T m for a particular sequence at a defined ionic strength and pH. The melting temperature T m can be the temperature at which a population of double-stranded nucleic acid molecules is half-dissociated into single strands. As shown by standard references, a simple estimate of the T m value is given by the equation T mIt can be calculated by = 81.5 + 0.41(%G + C) (see, for example, Anderson and Young, Quantitative Filter Hybridization, Nucleic Acid Hybridization (1985)). Other references (e.g., Allawi and SantaLucia, Jr., Biochemistry, 36: 10581-94 (1997)) include alternative computational methods that take into account structural and environmental features, as well as sequence features, for the calculation of T m including alternative computational methods that take into account structural and environmental features, as well as sequence features, for the calculation of T

[0313] Generally, the stability of a hybrid is a function of ion concentration and temperature. Typically, the hybridization reaction is carried out under conditions of lower stringency, followed by washes at various, but higher, stringencies. Exemplary stringent conditions include a salt concentration of at least 0.01 M to 1 M or less of sodium ions (or other salts) at a pH of about 7.0 to about 8.3 and a temperature of at least 25 °C. For example, conditions of 5×SSPE (750 mM NaCl, 50 mM sodium phosphate, 5 mM EDTA at pH 7.4) and a temperature of about 30 °C are suitable for allele-specific hybridization, but the suitable temperature depends on the length and / or GC content of the hybridized region. In one aspect, the "stringency of hybridization" when determining the proportion of mismatches can be as follows: 1) high stringency: 0.1×SSPE, 0.1% SDS, 65 °C, 2) moderate stringency: 0.2×SSPE, 0.1% SDS, 50 °C (also referred to as moderate stringency), and 3) low stringency: 1.0×SSPE, 0.1% SDS, 50 °C. It is understood that equivalent stringencies can be achieved using alternative buffers, salts, and temperatures. For example, moderately stringent hybridization can refer to conditions that allow a nucleic acid molecule, such as a probe, to bind to a complementary nucleic acid molecule. Hybridized nucleic acid molecules generally have at least 60% identity, including any of 70%, 75%, 80%, 85%, 90%, or 95% identity. Moderately stringent conditions can be equivalent to hybridization in 50% formamide, 5× Denhardt's solution, 5×SSPE, 0.2% SDS at 42 °C, followed by washing in 0.2×SSPE, 0.2% SDS at 42 °C. High stringency conditions can be provided, for example, by hybridization in 50% formamide, 5× Denhardt's solution, 5×SSPE, 0.2% SDS at 42 °C, followed by washing in 0.1×SSPE and 0.1% SDS at 65 °C.Low stringency hybridization can refer to conditions equivalent to hybridization in 10% formamide, 5× Denhardt's solution, 6× SSPE, 0.2% SDS at 22° C., followed by washing in 1× SSPE, 0.2% SDS at 37° C. Denhardt's solution contains 1% Ficoll, 1% polyvinylpyrrolidone, and 1% bovine serum albumin (BSA). 20× SSPE (sodium chloride, sodium phosphate, ethylenediaminetetraacetic acid (EDTA)) contains 3M sodium chloride, 0.2M sodium phosphate, and 0.025M EDTA. Other suitable moderate stringency and high stringency hybridization buffers and conditions are described, for example, in Sambrook et al., Molecular Cloning: A Laboratory Manual, 2nd ed., Cold Spring Harbor Press, Plainview, N.Y. (1989), and Ausubel et al., Short Protocols in Molecular Biology, 4th ed., John Wiley & Sons (1999).

[0314] Alternatively, substantial complementarity exists when an RNA or DNA strand hybridizes to its complement under selective hybridization conditions. Typically, selective hybridization occurs when there is at least about 65% complementarity, preferably at least about 75%, more preferably at least about 90% complementarity over a range of at least 14 to 25 nucleotides. See M. Kanehisa, Nucleic Acids Res. 12:203 (1984).

[0315] As used herein, a "primer" can be either a natural or synthetic oligonucleotide, which, when forming a double strand with a polynucleotide template, acts as an initiation point for nucleic acid synthesis and can extend along the template from its 3'-end so that an extended double strand is formed. The sequence of nucleotides added during the extension process is determined by the sequence of the template polynucleotide. A primer is usually extended by DNA polymerase.

[0316] "Ligation" can refer to the formation of a covalent bond or linkage between the ends of two or more nucleic acids, such as oligonucleotides and / or polynucleotides, in a template-driven reaction. The nature of the bond or linkage can vary widely and ligation can be carried out enzymatically or chemically. As used herein, ligation is usually carried out enzymatically to form a phosphodiester linkage between the 5'-carbon terminal nucleotide of one oligonucleotide and the 3'-carbon of another nucleotide.

[0317] "Sequencing", "sequence determination", etc. mean the determination of information related to the nucleotide base sequence of a nucleic acid. Such information may include the identification or determination of partial as well as complete sequence information of the nucleic acid. The sequence information can be determined with varying degrees of statistical certainty or reliability. In one aspect, the term includes the determination of the identity and order of multiple consecutive nucleotides in a nucleic acid. "High-throughput digital sequencing" or "next-generation sequencing" refers to determining many (typically thousands to billions) of nucleic acid sequences in an essentially parallel fashion, i.e., DNA templates are prepared for sequencing in a bulk process rather than one at a time, and many sequences are preferably read out in parallel, or alternatively, using a super-high-throughput sequential process that can itself be parallelized. Such methods include, but are not limited to, pyrosequencing (e.g., commercialized by 454 Life Sciences, Inc., Branford, Conn.), sequencing by ligation (e.g., SOLiD™ technology, commercialized by Life Technologies, Inc., Carlsbad, Calif.), sequencing by synthesis using modified nucleotides (TruSeq™ and HiSeq™ technologies by Illumina, ...

Claims

1. A method for analyzing a biological sample, comprising: a) contacting the biological sample with a first probe and a second probe, wherein the biological sample is a cell or tissue sample, wherein the biological sample contains a first analyte and a second analyte at a first position and a second position in the biological sample, respectively, wherein the first probe and the second probe are directly or indirectly bound to the first analyte and the second analyte, respectively, wherein the first probe or its product contains i) a first priming site for a first sequencing primer, and ii) a first identifier sequence associated with the first analyte, and wherein the second probe or its product contains i) a second priming site for a second sequencing primer, and ii) a second identifier sequence associated with the second analyte, and b) performing base-by-base sequencing of the first and second identifier sequences using the first and second sequencing primers, thereby generating a first signal code sequence and a second signal code sequence, each containing a signal code corresponding to a signal (on-signal), absence of a signal (off-signal), or a combination thereof, respectively, detected in consecutive cycles at the first position and the second position, wherein in one or more of the consecutive cycles, an on-signal is detected at the first position and an off-signal is detected at the second position, and c) detecting the first and second identifier sequences in the biological sample based at least on the first and second signal code sequences.

2. The method according to claim 1, wherein the first and second analytes are the same or different.

3. The method according to claim 1 or 2, wherein the first and second identifier sequences are different.

4. The method according to any one of claims 1 to 3, wherein the first and second identifier sequences contain an analyte sequence or its complement.

5. The method according to any one of claims 1 to 4, wherein the first and second identifier sequences contain a barcode sequence or its complement assigned to the first and second analytes, respectively.

6. The method according to claim 5, comprising assigning a first barcode sequence to the first analyte and a second barcode sequence to the second analyte.

7. The method according to claim 6, wherein assigning the barcode array is based on a decision rule designed to minimize the maximum predicted density of the on-signals detected in each of the one or more of the continuous cycles.

8. The method according to claim 6, wherein assigning the first barcode array to the first analyte and the second barcode array to the second analyte includes an assignment based on expression data for the first analyte and the second analyte.

9. The method according to claim 8, wherein assigning the first barcode array to the first analyte and the second barcode array to the second analyte includes an assignment based on expression data for the first analyte and the second analyte in the clustered cell type.

10. The method according to claim 9, wherein the clustered cell type represents the distribution of cell types found in the biological sample.

11. The method according to any one of claims 8 to 10, wherein the expression data for the first analyte and the second analyte at least partially overlap.

12. The method according to any one of claims 8 to 11, wherein the expression data for the first analyte and the second analyte includes bulk gene expression data, bulk protein expression data, spatial gene expression data, spatial protein expression data, single cell gene expression data, single cell protein expression data, or any combination thereof.

13. The method according to any one of claims 6 to 12, wherein the nucleotide in the first barcode array detected in a particular cycle corresponds to a signal code including an on-signal, and the corresponding nucleotide in the second barcode array detected in the particular cycle corresponds to a signal code including an off-signal.

14. The method according to claim 13, wherein the nucleotide in the first barcode array detected in the particular cycle corresponds only to an on-signal, and the corresponding nucleotide in the second barcode array detected in the particular cycle corresponds only to an off-signal.

15. The method according to any one of claims 6 to 14, wherein one or more pairs of corresponding nucleotides in the first and second barcode sequences detected in the same cycle are selected to reduce optical crowding of the signals detected in the cycle.

16. The base-by-base sequencing is performed by contacting the biological sample with nucleotides in successive cycles, and in each cycle, a complex is formed, the complex comprising: i) the first or second sequencing primer, or an extension product thereof, hybridized to the first or second priming site, respectively; ii) a polymerase; and iii) cognate nucleotides that base pair with nucleotides in the first or second identifier sequence, and a signal (on-signal) and / or absence of a signal (off-signal) associated with the cognate nucleotides and / or the polymerase in the complex is detected at a specific position in the biological sample, and the on-signal, the off-signal, or a combination thereof corresponds to the cognate nucleotides in the first or second identifier sequence and bases in the corresponding nucleotides. The method according to any one of claims 1 to 15.

17. The method according to any one of claims 1 to 16, wherein 25% or more of the nucleotides in the first and / or second identifier sequence correspond to an off-signal.

18. The method according to claim 17, wherein 40% or more, 45% or more, 50% or more, 55% or more, 60% or more, 65% or more, 70% or more, 75% or more, 80% or more, 85% or more, 90% or more, or 95% or more of the nucleotides of the first and / or second identifier sequence are assigned to correspond to an off-signal.

19. The method according to any one of claims 1 to 18, wherein a plurality of different identifier sequences are detected in the biological sample, and each different identifier sequence is detected at one or more positions in the biological sample.

20. The method according to claim 19, wherein 50% or more of the different identifier sequences each contain 50% or more of the nucleotides in the identifier sequence that correspond to an off-signal.

21. The method according to claim 19 or 20, wherein 80% or more of the different identifier sequences each contain 80% or more of the nucleotides in the identifier sequence that correspond to an off-signal.

22. The method according to any one of claims 1 to 21, wherein the signal code corresponds to a signal of a first color, a signal of a second color, a signal of a third color, or the absence of a signal, and the first, second, and third colors are different.

23. The method according to any one of claims 1 to 21, wherein the signal code corresponds to a signal of a first color, a signal of a second color, a combination of the signals of the first and second colors, or the absence of a signal, and the first and second colors are different.

24. The method according to any one of claims 1 to 23, wherein the signal code corresponds to a combination of a signal (on-signal) and / or the absence of a signal (off-signal), and the combination of on and / or off signals is detected in two or more imaging steps.

25. The method according to any one of claims 1 to 24, further comprising detecting the first and second analytes in the biological sample based on detecting the first and second identifier sequences.

26. The method according to any one of claims 1 to 25, wherein the first identifier sequence is a first barcode sequence or its complement, and the second identifier sequence is a second barcode sequence or its complement.

27. The first barcode sequence identifies the first analyte and / or The second barcode sequence identifies the second analyte, The method according to claim 26.

28. The first probe is provided among a first plurality of probes that directly or indirectly bind to the first analyte, the first plurality of probes collectively include a barcode sequence of a first combination, and the barcode sequence of the first combination identifies the first analyte and / or The second probe is provided among a second plurality of probes that directly or indirectly bind to the second analyte, the second plurality of probes collectively include a barcode sequence of a second combination, and the barcode sequence of the second combination identifies the second analyte, The method according to claim 26 or 27.

29. The method according to any one of claims 1 to 28, wherein the first and second analytes include nucleic acid sequences, the first identifier sequence is the sequence of the first analyte or its complement, and the second identifier sequence is the sequence of the second analyte or its complement.

30. The base-by-base determination comprises using a fluorescently labeled polymerase and one or more unlabeled nucleotides, using a polymerase-nucleotide conjugate comprising a fluorescently labeled polymerase linked to an unlabeled nucleotide moiety, or using a multivalent polymer-nucleotide conjugate comprising a polymer core, a plurality of nucleotide moieties, and one or more fluorescent labels The method according to any one of claims 1 to 29.

31. The method according to claim 30, wherein cognate nucleotides are not incorporated into the first or second sequencing primer or an extension product thereof by the polymerase.

32. The method according to claim 30, wherein incorporation of cognate nucleotides into the first or second sequencing primer or an extension product thereof by the polymerase is reduced or inhibited.

33. The method according to any one of claims 1 to 32, wherein the base-by-base determination comprises contacting the biological sample with a nucleotide mix comprising fluorescently labeled nucleotides and unlabeled nucleotides.

34. The method according to claim 33, wherein cognate nucleotides are incorporated into the first or second sequencing primer or an extension product thereof by the polymerase, and the cognate nucleotides are either fluorescently labeled or unlabeled.

35. The base-by-base determination comprises contacting the biological sample with a first nucleotide mix, wherein in the first nucleotide mix, nucleotides comprising a first base are not detectably labeled, while nucleotides comprising bases other than the first base are each labeled with one or more detectable labels, and contacting the biological sample with a subsequent nucleotide mix, wherein in the subsequent nucleotide mix, nucleotides comprising a subsequent base are not detectably labeled, while nucleotides comprising bases other than the subsequent base are each labeled with one or more detectable labels, comprising wherein the subsequent base is the same as the first base, and optionally, the first and subsequent bases are A, T, C, or G. The method according to claim 33 or 34.

36. The base-by-base determination comprises contacting the biological sample with a first nucleotide mix, wherein in the first nucleotide mix, nucleotides containing a first base are not detectably labeled, while other nucleotides in the first nucleotide mix are each labeled with one or more detectable labels, and contacting the biological sample with a subsequent nucleotide mix, wherein in the subsequent nucleotide mix, nucleotides containing a subsequent base are not detectably labeled, while other nucleotides in the subsequent nucleotide mix are each labeled with one or more detectable labels, comprising wherein the subsequent base is different from the first base, The method according to claim 33 or 34.

37. The method according to claim 36, wherein the biological sample contacts two or more of the following nucleotide mixes in consecutive cycles in any order: Nucleotide mix 1, wherein nucleotides containing G are not detectably labeled, while nucleotides containing A, C, or T are detectably labeled, Nucleotide mix 2, wherein nucleotides containing T are not detectably labeled, while nucleotides containing A, C, or G are detectably labeled, Nucleotide mix 3, wherein nucleotides containing C are not detectably labeled, while nucleotides containing A, G, or T are detectably labeled, and Nucleotide mix 4, wherein nucleotides containing A are not detectably labeled, while nucleotides containing G, C, or T are detectably labeled.

38. The method according to claim 37, wherein, independently of each other, each nucleotide mix contacts the biological sample in one or more cycles, and the cycles are consecutive or non - consecutive.

39. Independently of each other, in each nucleotide mix, the detectably labeled nucleotides are i) fluorescent labels of three different colors, one color for each of the three bases, ii) fluorescent labels of two different colors, one color for each of two of the three bases, and nucleotides containing the remaining base are labeled with both colors, or iii) a fluorescent label of the same color, wherein the fluorescent label on a nucleotide containing one of the three bases is configured to be cleaved, and a nucleotide containing another one of the three bases is configured to be labeled with the fluorescent label The method according to claim 37 or 38, comprising the fluorescent label **Claim 40** The method according to claim 36, wherein the biological sample contacts two or more of the following nucleotide mixtures in consecutive cycles in any order: Nucleotide mixture 1, in which a nucleotide containing G or A is not detectably labeled while a nucleotide containing C or T is detectably labeled Nucleotide mixture 2, in which a nucleotide containing G or T is not detectably labeled while a nucleotide containing C or A is detectably labeled Nucleotide mixture 3, in which a nucleotide containing G or C is not detectably labeled while a nucleotide containing A or T is detectably labeled Nucleotide mixture 4, in which a nucleotide containing C or A is not detectably labeled while a nucleotide containing G or T is detectably labeled Nucleotide mixture 5, in which a nucleotide containing C or T is not detectably labeled while a nucleotide containing G or A is detectably labeled, and Nucleotide mixture 6, in which a nucleotide containing A or T is not detectably labeled while a nucleotide containing G or C is detectably labeled **Claim 41** The first priming site and the second priming site are different, and the method comprises b1) hybridizing the first sequencing primer to the first priming site and performing base-by-base sequencing to generate an extension product of the first sequencing primer and the first signal code sequence; b2) removing, cleaving, or blocking the extension product of the first sequencing primer in b1); b3) hybridizing the second sequencing primer to the second priming site and performing base-by-base sequencing to generate an extension product of the second sequencing primer and the second signal code sequence The method according to any one of claims 1 to 40, comprising the steps above **Claim 42** The method according to claim 41, wherein the probes or products thereof for the first plurality of analytes share a common first priming site, and the probes or products thereof for the second plurality of analytes share a common second priming site.

43. The method according to claim 42, wherein the second plurality of analytes includes two or more different analytes that are different from two or more different analytes among the first plurality of analytes.

44. In a), the biological sample is contacted with a plurality of probes configured to bind directly or indirectly to different analytes, and each probe or product thereof includes a combination of different priming sites. The method according to any one of claims 1 to 43.

45. The first probe or product thereof includes a first combination of different priming sites including the first priming site, and / or the second probe or product thereof includes a second combination of different priming sites including the second priming site. The method according to claim 44.

46. In a), the biological sample is contacted with a third probe that binds directly or indirectly to a third analyte, the third probe or product thereof includes a third combination of different priming sites including the first priming site, the second priming site, and / or a third priming site. The method according to claim 45.

47. The method according to claim 46, wherein any two or more of the first combination, the second combination, and the third combination share one or more common priming sites.

48. b') contacting the biological sample with the first sequencing primer for base-by-base sequencing, thereby hybridizing the first sequencing primer to the first priming site in the first probe or product thereof and in one or more other probes or products thereof, and generating an extension product of the first sequencing primer; b'') removing, cleaving, or blocking the extension product of the first sequencing primer in b'); and b''') contacting the biological sample with the second sequencing primer for base-by-base sequencing, thereby hybridizing the second sequencing primer to the second priming site in the second probe or its product and in one or more other probes or their products, and generating an extension product of the second sequencing primer The method according to any one of claims 44 to 47, comprising: **Claim 49** wherein the base-by-base sequencing in b') is contacting the biological sample with nucleotides in successive cycles, detecting a signal associated with nucleotide incorporation or binding for each successive cycle, and generating a signal code sequence for a first plurality of analytes The method according to claim 48, wherein the method is carried out by: **Claim 50** wherein the base-by-base sequencing in b''') is contacting the biological sample with nucleotides in successive cycles, detecting a signal associated with nucleotide incorporation or binding for each successive cycle, and generating a signal code sequence for a second plurality of analytes The method according to claim 48 or 49, wherein the method is carried out by: **Claim 51** The method according to claim 50, wherein the first plurality of analytes and the second plurality of analytes comprise one or more common analytes. **Claim 52** The method according to claim 50, wherein the first plurality of analytes and the second plurality of analytes do not comprise a common analyte. **Claim 53** The method according to any one of claims 1 to 52, wherein each analyte is independently a nucleic acid analyte or a non-nucleic acid analyte. **Claim 54** The method according to any one of claims 1 to 53, wherein each probe is independently i) a primary probe that binds directly to its corresponding analyte, or ii) a probe that binds directly or indirectly to the primary probe. **Claim 55** The primary probe and the probe that binds directly or indirectly to the primary probe are A probe comprising a 3' or 5' overhang, optionally wherein the 3' or 5' overhang comprises one or more barcode sequences; a probe comprising a 3' overhang and a 5' overhang, optionally wherein the 3' overhang and the 5' overhang each independently comprise one or more barcode sequences; a circular probe; a circularizable probe or probe set; a probe or probe set comprising a split hybridization region configured to hybridize to a splint, optionally wherein the split hybridization region comprises one or more barcode sequences; and combinations thereof The method according to claim 54, independently selected from the group consisting of

56. The method according to any one of claims 1 to 55, wherein the product of each probe is a rolling circle amplification (RCA) product generated in situ in the biological sample or an aggregate of branched structures formed in situ in the biological sample

57. The method according to any one of claims 1 to 56, wherein in (b), the sequencing of each base is performed in situ in the biological sample

58. A method for analyzing a biological sample, comprising: a) contacting the biological sample with a first probe and a second probe, wherein the biological sample is a cell or tissue sample, wherein the biological sample contains a first analyte and a second analyte at a first position and a second position in the biological sample, respectively, wherein the first probe and the second probe are directly or indirectly bound to the first analyte and the second analyte, respectively, wherein the first probe or its product comprises i) a first priming site for a first sequencing primer, and ii) a first identifier sequence associated with the first analyte, and wherein the second probe or its product comprises i) a second priming site for a second sequencing primer, and ii) a second identifier sequence associated with the second analyte such that b) performing base-by-base sequencing of the first identifier array using the first sequencing primer to generate a first signal code array including the signal codes detected in consecutive cycles at the first position, wherein the signal codes correspond to a signal (on-signal), the absence of a signal (off-signal), or a combination thereof; c) subsequently performing base-by-base sequencing of the second identifier array using the second sequencing primer to generate a second signal code array including the signal codes detected in additional consecutive cycles at the second position; wherein an off-signal is detected in at least one or more of the consecutive cycles in b) and the additional consecutive cycles in c); and d) detecting the first and second identifier arrays in the biological sample at the first and second positions, respectively, based at least in part on the first and second signal code arrays. **Claim 59** The method according to claim 58, wherein the first and second identifier arrays are the same. **Claim 60** The method according to claim 58, wherein the first and second identifier arrays are different. **Claim 61** The method according to any one of claims 58 to 60, wherein the first and second identifier arrays include barcode sequences or their complements respectively assigned to the first and second analytes. **Claim 62** Based on a determination rule designed to minimize the maximum predicted density of on-signals detected in each of the one or more of the consecutive cycles, the first priming site and / or the first barcode sequence is assigned to the first analyte, and the second priming site and / or the second barcode sequence is assigned to the second analyte. The method according to claim 61. **Claim 63** The method according to claim 61, wherein the assigning includes an assignment based on expression data for the first and second analytes. **Claim 64** The method according to claim 63, wherein the assigning includes an assignment based on expression data for the first and second analytes in a clustered cell type. **Claim 65** The method according to claim 64, wherein the clustered cell type represents the distribution of cell types found in the biological sample.

66. The method according to any one of claims 63 to 65, wherein the expression data of the first analyte and the second analyte at least partially overlap.

67. The method according to any one of claims 63 to 66, wherein the expression data for the first analyte and the second analyte comprises bulk gene expression data, bulk protein expression data, spatial gene expression data, spatial protein expression data, single cell gene expression data, single cell protein expression data, or any combination thereof.

68. The method according to any one of claims 58 to 67, wherein the nucleotides in the first barcode sequence or the second barcode sequence detected in a particular cycle correspond to a signal code comprising an on-signal.

69. The method according to any one of claims 58 to 68, wherein the nucleotides in the first barcode sequence or the second barcode sequence detected in a particular cycle correspond to a signal code comprising an off-signal.

70. The method according to claim 68 or claim 69, wherein the nucleotides in the first barcode sequence or the second barcode sequence detected in the particular cycle correspond only to an on-signal.

71. The method according to claim 69 or claim 70, wherein the nucleotides in the first barcode sequence or the second barcode sequence detected in the particular cycle correspond only to an off-signal.

72. A method for decoding an identifier sequence in a biological sample while minimizing optical crowding, comprising: a) contacting the biological sample with a first probe and a second probe, wherein the biological sample is a cell or tissue sample, wherein the biological sample contains a first analyte and a second analyte at a first position and a second position in the biological sample, respectively, and wherein the first probe and the second probe bind directly or indirectly to the first analyte and the second analyte, respectively, wherein the first probe or its product comprises i) a first priming site for a first sequencing primer, and ii) a first identifier sequence associated with the first analyte, wherein the second probe or its product comprises i) a second priming site for a second sequencing primer, and ii) a second identifier sequence associated with the second analyte. a b) performing base-by-base sequencing of the first and second identifier sequences using the first and second sequencing primers, thereby generating a first signal code sequence and a second signal code sequence, each including a signal code corresponding to a signal (on-signal), absence of signal (off-signal), or combination thereof detected in successive sequencing cycles at the first and second positions, respectively; wherein the base-by-base sequencing includes contacting the biological sample in each successive cycle with a mixture of polymerase and nucleotides, wherein i) at least one nucleotide (e.g., A, T / U, C, or G) is not detectably labeled, ii) at least one nucleotide (e.g., A, T / U, C, or G) is excluded from the mixture, or iii) at least one nucleotide (e.g., A, T / U, C, or G) is not detected in c); a, and c) detecting the first and second identifier sequences in the biological sample based on at least the first and second signal code sequences. Claim 73 The method of claim 72, wherein the first and second analytes are the same or different. Claim 74 The method of claim 72 or 73, wherein the first and second identifier sequences are different. Claim 75 The method according to any one of claims 72 to 74, wherein the first and second identifier sequences include an analyte sequence or a complement thereof. Claim 76 The method according to any one of claims 72 to 75, wherein the first and second identifier sequences include barcode sequences or complements thereof assigned to the first and second analytes, respectively. Claim 77 The method of claim 76, including assigning a first barcode sequence to the first analyte and a second barcode sequence to the second analyte. Claim 78 The method of claim 77, wherein the assignment of the barcode sequences is based on a decision rule designed to minimize the maximum predicted density of on-signals detected in each of the one or more of the successive cycles. Claim 79 The method of claim 77, wherein the determination rule for assigning the first barcode array to the first analyte and the second barcode array to the second analyte includes an assignment based on expression data for the first analyte and the second analyte.

80. The method of claim 79, wherein assigning the first barcode array to the first analyte and the second barcode array to the second analyte includes an assignment based on expression data for the first analyte and the second analyte in a clustered cell type.

81. The method of claim 80, wherein the clustered cell type represents the distribution of cell types found in the biological sample.

82. The method according to any one of claims 79 - 81, wherein the expression data for the first analyte and the second analyte at least partially overlap.

83. The method according to any one of claims 79 - 82, wherein the expression data for the first analyte and the second analyte includes bulk gene expression data, bulk protein expression data, spatial gene expression data, spatial protein expression data, single - cell gene expression data, single - cell protein expression data, or any combination thereof.

84. The method according to any one of claims 72 - 83, wherein the nucleotide mixture includes at least two nucleotides that are not detectably labeled.

85. The method according to any one of claims 72 - 83, wherein the biological sample contacts two or more of the following nucleotide mixtures in any order in successive sequencing cycles: Nucleotide mix 1, in which the nucleotide containing G is not detectably labeled while the nucleotide containing A, C, or T is detectably labeled; Nucleotide mix 2, in which the nucleotide containing T is not detectably labeled while the nucleotide containing A, C, or G is detectably labeled; Nucleotide mix 3, in which the nucleotide containing C is not detectably labeled while the nucleotide containing A, G, or T is detectably labeled; and Nucleotide mix 4, in which the nucleotide containing A is not detectably labeled while the nucleotide containing G, C, or T is detectably labeled.

86. The method according to claim 85, wherein each nucleotide mix independently of the others is contacted with the biological sample in one or more cycles, the cycles being consecutive or non-consecutive.

87. Independently of each other, in each nucleotide mix, the detectably labeled nucleotide is i) a fluorescent label that is three different colored fluorescent labels, one color for each of the three bases; ii) a fluorescent label that is two different colored fluorescent labels, one color for each of two of the three bases, and the nucleotide containing the remaining base is labeled with both colors; or iii) a fluorescent label that is the same color, wherein the fluorescent label on the nucleotide containing one of the three bases is configured to be cleaved, and the nucleotide containing another one of the three bases is configured to be labeled with the fluorescent label. The method according to claim 85 or 86, comprising.

88. The method according to any one of claims 72 to 83, wherein the biological sample is contacted with two or more of the following nucleotide mixes in consecutive cycles in any order: Nucleotide mix 1, wherein the nucleotide containing G or A is not detectably labeled, while the nucleotide containing C or T is detectably labeled; Nucleotide mix 2, wherein the nucleotide containing G or T is not detectably labeled, while the nucleotide containing C or A is detectably labeled; Nucleotide mix 3, wherein the nucleotide containing G or C is not detectably labeled, while the nucleotide containing A or T is detectably labeled; and Nucleotide mix 4, wherein the nucleotide containing C or A is not detectably labeled, while the nucleotide containing G or T is detectably labeled; Nucleotide mix 5, wherein the nucleotide containing C or T is not detectably labeled, while the nucleotide containing G or A is detectably labeled; and Nucleotide mix 6, wherein the nucleotide containing A or T is not detectably labeled, while the nucleotide containing G or C is detectably labeled.