Systems and methods for nucleic acid mismatch error detection

By preserving and analyzing both strands of the double-stranded template nucleic acid molecule during nucleic acid sequencing, the problem of information loss during amplification is solved, improving the accuracy of sequencing and the reliability of SNP detection, especially in the presence of base mismatches.

CN121488034APending Publication Date: 2026-02-06ULTIMA GENOMICS INC
View PDF 20 Cites 0 Cited by

Patent Information

Application Number
CN202480022349.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-05-10
Filing Date
2024-01-26
Publication Date
2026-02-06

AI Technical Summary

Technical Problem

In existing nucleic acid sequencing technologies, the two strands of a double-stranded template nucleic acid molecule are prone to losing information during amplification, especially when base mismatches exist. This leads to a high error rate in SNP detection and makes it difficult to efficiently preserve and detect the two strands of the double-stranded template nucleic acid molecule.

Method used

A system and method are provided to identify single nucleotide variants (SNVs) or single nucleotide polymorphisms (SNPs) by preserving both strands of a double-stranded template nucleic acid molecule during amplification and determining locus inconsistencies through sequencing signal analysis. This includes probing both strands of the nucleic acid molecule using sequencing primers with the same or different sequences, generating and aligning sequencing reads, and determining the proportion and presence of the strands.

Benefits of technology

It improves the accuracy of nucleic acid sequencing, enabling the identification and exclusion of SNPs, detection of minimal residual disease and tumor fraction, and achieving efficient information retention and detection of double-stranded template nucleic acid molecules.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121488034A_ABST
    Figure CN121488034A_ABST
Patent Text Reader

Abstract

A double-stranded template nucleic acid molecule may include an adaptor that includes a mismatched moiety. The mismatch moiety may include a first mismatch sequence in a first chain and a second mismatch sequence in a second chain that is not complementary to the first mismatch sequence. When the double-stranded template nucleic acid molecule is amplified to produce a cluster of amplified strands and the cluster is sequenced to produce a sequenced read, the double-stranded template nucleic acid molecule is amplified to produce the sequenced read. A portion of the sequencing read corresponding to the mismatch portion can be analyzed to determine whether the cluster of amplified strands originates only from one strand or both strands of the double-stranded template nucleic acid molecule. A sequencing signal between two chain derivatives in the cluster, a base recognition, or an inconsistency at one or more loci in a sequencing read can be used to correct for sequencing errors and improve sequencing accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-references

[0002] This application claims U.S. Provisional Patent Application No. 63 / 441,727, filed January 27, 2023; U.S. Provisional Patent Application No. 63 / 452,696, filed March 17, 2023; and U.S. Provisional Patent Application No. 63 / 465,346, filed May 10, 2023, each of which is incorporated herein by reference for all purposes. Background Technology

[0003] Biological sample handling has a variety of applications in molecular biology and medicine (e.g., diagnostics). For example, nucleic acid sequencing can provide information that can be used to diagnose certain conditions in a subject and, in some cases, to tailor treatment plans. Sequencing is widely used in molecular biology applications, including vector design, gene therapy, vaccine design, and industrial strain design and validation. Biological sample handling can involve fluid systems and / or detection systems.

[0004] Despite advancements in sequencing technology, significant effort is still required to analyze samples with high throughput and efficiency. Summary of the Invention

[0005] Template nucleic acid molecules input into sequencing libraries can be double-stranded. In other cases, template nucleic acid molecules may exist in double-stranded form and / or be converted to double-stranded form at various time points during the sequencing process (such as during one or more operations of library preparation, amplification, or enrichment). Prior to sequencing, template nucleic acid molecules may undergo amplification, such as to produce clusters of template nucleic acid molecules (e.g., on sites on a vector such as beads or on a surface)—during sequencing, the cumulative signal detected from clusters of amplified molecules can be significantly stronger than the signal from a single molecule. However, some amplification protocols can amplify only one of the two strands of the double-stranded template nucleic acid molecule, thereby discarding information from the other strand (e.g., sequence information). While in some cases the two strands of a double-stranded template nucleic acid molecule may be perfectly inverse complementary sequences, in which case discarding one strand does not result in loss of information, in other cases the two strands may contain base mismatch sites. In this case, discarding one strand results in the loss of valuable information from the template nucleic acid molecule, including potential alternative base markers in the sequence and the fact that a base mismatch was originally present. As used herein, a base mismatch, also referred to as a mispair, generally refers to the occurrence of non-complementary base pairings within the double-stranded portion of a nucleic acid molecule. In some instances, PCR-free library DNA can carry sites with a relatively high frequency of base mismatches. In other amplification protocols, it is possible to amplify both strands of a double-stranded template nucleic acid molecule to produce a “mixture” of strand derivatives in the amplified clusters. However, the signal detected from such a “mixture” of amplified clusters may need to be limited by understanding the percentage of amplified clusters originating from one strand relative to the percentage originating from the other. Therefore, it can be beneficial to retain both strands of the template nucleic acid molecule during amplification. Such retention can be particularly advantageous in detecting or excluding single nucleotide polymorphisms (SNPs) and in improving the error rate of SNP detection. SNPs are genetic variations in a subject's DNA that are particularly prone to false detection. When both strands of a double-stranded template nucleic acid molecule are retained for amplification, it is important that a significant amount of both strands are actually amplified and presented in the amplified cluster.

[0006] This invention provides systems and methods that at least address the problems described above. The invention provides systems and methods for preserving both strands of a double-stranded template nucleic acid molecule during amplification, and systems and methods for quantitatively measuring the proportion of clusters of amplified molecules derived from said two strands. This invention also provides systems and methods for amplifying both strands of a double-stranded template nucleic acid without dissociating said two strands.

[0007] On one hand, the present invention provides a method for high-accuracy sequencing, the method comprising: (a) providing an amplified cluster of a plurality of nucleic acid molecules derived from a double-stranded nucleic acid molecule comprising a first strand and a second strand, wherein a first subset of the plurality of nucleic acid molecules each comprises a copy of a first sequence as at least a portion of a first strand sequence, and wherein a second subset of the plurality of nucleic acid molecules each comprises a second sequence comprising a reverse complementary copy of a second sequence as at least a portion of a second strand sequence; (b) collecting sequencing signals from the amplified cluster to determine inconsistencies between the first and second sequences at a locus; and (c) excluding the locus from single nucleotide variant (SNV) or single nucleotide polymorphism (SNP) recognition at least in part based on the inconsistency.

[0008] In some embodiments, in (b) the sequencing signal from the amplified cluster is collected by hybridizing sequencing primers with the plurality of nucleic acid molecules and simultaneously probing both the first subset and the second subset with the same mixture of nucleotides.

[0009] In some implementations, the sequencing primers that hybridize with the first subset and the second subset comprise the same sequence.

[0010] In some implementations, the sequencing primers that hybridize with the first subset and the second subset comprise different sequences.

[0011] In some implementations, in (b), the same nucleotide mixture probes for the same locus between the first sequence and the second sequence.

[0012] In some embodiments, the length of the same locus is one base position.

[0013] In some embodiments, the length of the same locus is at least two base positions.

[0014] In some embodiments, the homonucleotide mixture probes for the same locus between the first sequence and the second sequence.

[0015] In some embodiments, the length of the same locus is one base position.

[0016] In some embodiments, the length of the same locus is at least two base positions.

[0017] In some implementations, (b) includes generating sequencing data simultaneously based on the first subset and the second subset.

[0018] In some embodiments, in (b) the sequencing signal from the amplified cluster is collected by: (i) hybridizing a first set of sequencing primers with a first subset of the plurality of nucleic acid molecules and probing the first subset with a first set of nucleotide mixtures; and (ii) hybridizing a second set of sequencing primers with a second subset of the plurality of nucleic acid molecules and probing the second subset with a second set of nucleotide mixtures different from the first set of nucleotide mixtures.

[0019] In some implementations, the first set of sequencing primers and the second set of sequencing primers comprise different sequences.

[0020] In some implementations, the first set of sequencing primers and the second set of sequencing primers have the same sequence.

[0021] In some implementations, (b) includes generating sequencing data at different time points based on the first subset and the second subset.

[0022] In some embodiments, the method further includes generating a single sequencing read from the amplified cluster, wherein the single sequencing read is generated by probing both a first subset and a second subset of the plurality of nucleic acid molecules.

[0023] In some embodiments, the method further includes generating two sequencing reads from the amplified cluster: a first sequencing read generated by probing a first subset of the plurality of nucleic acid molecules and a second sequencing read generated by probing a second subset of the plurality of nucleic acid molecules.

[0024] In some embodiments, the method further includes generating at least two candidate base recognitions at the locus where the inconsistency occurs.

[0025] In some embodiments, the method further includes comparing the at least two candidate base identifications with the locus in a reference sequence, and selecting one of the at least two candidate base identifications or a base at the locus in the reference sequence for a shared read.

[0026] In some embodiments, the method further includes using the sequencing signal to generate sequencing reads for the amplified cluster.

[0027] In some embodiments, the method further includes aligning the sequencing read with the reference sequence.

[0028] In some embodiments, the method further includes identifying one or more SNVs or SNPs for the sample or subject from which the double-stranded nucleic acid molecule originates.

[0029] In some embodiments, the method further includes detecting minimal residual disease (MRD), tumor score, or circulating tumor score in the sample or the subject based on the identified one or more SNVs or SNPs.

[0030] In some embodiments, each of the first subsets further includes a first chain identification element, and each of the second subsets further includes a second chain identification element that is different from the first chain identification element.

[0031] In some implementations, the first chain identification element and the second chain identification element are different sequences.

[0032] In some implementations, the different sequences are different homopolymer sequences.

[0033] In some implementations, the different sequences are different heteropolymer sequences.

[0034] In some implementations, the different sequences have different lengths.

[0035] In some implementations, the different sequences have the same length.

[0036] In some embodiments, the different sequences each have a single base length.

[0037] In some embodiments, each of the different sequences has a length of at least one base.

[0038] In some embodiments, each of the different sequences has a length of at least 3 bases.

[0039] In some embodiments, each of the different sequences has a length of at least 5 bases.

[0040] In some implementations, the first and second chain recognition elements are not nucleic acid sequences.

[0041] In some embodiments, the method further includes detecting the presence of one or both of the first chain identification element and the second chain identification element in the amplified cluster.

[0042] In some embodiments, the detection includes sequencing the plurality of nucleic acid molecules to identify the sequence of the first strand recognition element, the second strand recognition element, or both.

[0043] In some embodiments, the detection includes sequencing the plurality of nucleic acid molecules to identify inconsistencies in the sequences of the plurality of nucleic acid molecules, including a common portion of the first or second strand recognition element.

[0044] In some embodiments, the detection includes hybridizing a labeled oligonucleotide probe with a first chain recognition element, a second chain recognition element, or both, and detecting a signal from the labeled oligonucleotide probe.

[0045] In some embodiments, the method further includes determining the ratio of the first subset to the second subset of the plurality of nucleic acid molecules.

[0046] In some implementations, the ratio is determined by processing the signal strength collected from probing the first chain identification element, the second chain identification element, or both.

[0047] In some embodiments, the method further includes generating sequencing reads for the amplified clusters, wherein the sequencing reads are generated at least in part based on the sequencing signal and the ratio.

[0048] In some implementations, the amplified cluster is fixed to an individually addressable location on the substrate.

[0049] In some implementations, the substrate includes at least 1,000,000 individually addressable locations.

[0050] In some implementations, the substrate includes at least 1,000,000,000 individually addressable locations.

[0051] In some implementations, the substrate includes at least 5,000,000,000 individually addressable locations.

[0052] In some implementations, the substrate includes at least 10,000,000,000 individually addressable locations.

[0053] In some implementations, the substrate includes at least 20,000,000,000 individually addressable locations.

[0054] In some implementations, the substrate is substantially planar.

[0055] In some implementations, the substrate is textured or patterned.

[0056] In some implementations, the substrate is unpatterned.

[0057] In some embodiments, the substrate includes an aminosilane layer that immobilizes the amplified clusters.

[0058] In some embodiments, the substrate includes a surface primer layer that immobilizes the amplified cluster.

[0059] In some embodiments, the plurality of nucleic acid molecules are coupled to beads, which are fixed to individually addressable locations on the substrate.

[0060] In some embodiments, the substrate is rotated during sequencing of the plurality of nucleic acid molecules.

[0061] In some embodiments, the plurality of nucleic acid molecules are single-stranded molecules.

[0062] On the other hand, a method for high-accuracy sequencing is provided, the method comprising: (a) providing an amplified cluster of a plurality of nucleic acid molecules derived from a double-stranded nucleic acid molecule, the double-stranded nucleic acid molecule comprising a first strand and a second strand, wherein a first subset of the plurality of nucleic acid molecules each comprises a first strand recognition element and a first sequence as a copy of at least a portion of the first strand sequence, and wherein a second subset of the plurality of nucleic acid molecules each comprises a second strand recognition element and a second sequence as a reverse complementary copy of at least a portion of the second strand sequence, wherein the first strand recognition element and the second strand recognition element are different; and (b) detecting the presence of the first strand recognition element, the second strand recognition element, or both in the amplified cluster.

[0063] In some embodiments, the plurality of nucleic acid molecules are immobilized onto a vector.

[0064] In some embodiments, the method further includes amplifying the double-stranded nucleic acid molecules to generate the plurality of nucleic acid molecules.

[0065] In some embodiments, the amplification includes emulsion PCR (ePCR), emulsion recombinase polymerase amplification (eRPA), PCR, RPA, rolling circle amplification (RCA), multiple displacement amplification (MDA), bridging amplification, or a combination thereof.

[0066] In some embodiments, the method further includes linking a strand recognition adaptor comprising a pair of mismatched sequences to the double-stranded nucleic acid molecule prior to the amplification.

[0067] In some implementations, double-stranded nucleic acid molecules are amplified by: (1) attaching a hairpin adapter to each end to generate a dumbbell-shaped molecule and subjecting the dumbbell-shaped molecule to rolling circle amplification (RCA) to generate a first amplification product; and (2) cleaving or digesting the portion of the first amplification product corresponding to the hairpin adapter to generate a plurality of copy molecules, each of the plurality of copy molecules comprising a copy of the first strand and a copy of the second strand.

[0068] In some embodiments, double-stranded nucleic acid molecules are amplified by: (1) attaching a hairpin adaptor to each end to generate a dumbbell-shaped molecule, and subjecting the dumbbell-shaped molecule to rolling circle amplification (RCA) to generate a first amplification product by contacting the dumbbell-shaped molecule with a plurality of random primers and dNTPs including dUTPs but not dTTPs; (2) hybridizing a second primer to the first amplification product in the presence of dNTPs including dTTPs but not dUTPs to generate a second amplification product; (3) degrading the first amplification product based on uracil residues to separate the second amplification product; and (4) cleaving or digesting the portion of the second amplification product corresponding to the hairpin adaptor to generate a plurality of copy molecules, each of the plurality of copy molecules including a copy of the first strand and a copy of the second strand.

[0069] In some implementations, double-stranded nucleic acid molecules are amplified by: (1) attaching a hairpin adapter to each end to generate a dumbbell-shaped molecule and subjecting the dumbbell-shaped molecule to rolling circle amplification (RCA) to generate a first amplification product; and (2) hybridizing a second primer with the first amplification product under repressor strand substitution conditions to generate a plurality of second amplification products, each of the plurality of second amplification products comprising a single copy of the first strand and the second strand.

[0070] In some embodiments, the method further includes linking a strand recognition adaptor comprising a pair of mismatched sequences to each of the plurality of copy molecules to generate the plurality of nucleic acid molecules.

[0071] In some embodiments, at least one of the hairpin connectors includes a chain-identifying connector that includes a pair of mismatched sequences, and each of the plurality of copy molecules includes a copy of the pair of mismatched sequences.

[0072] On the other hand, a method for detecting amplified strands on a vector is provided, the method comprising: (a) conjugating a double-stranded template molecule to a vector to produce a template-conjugated vector, wherein the double-stranded template molecule includes a first strand and a second strand, wherein the double-stranded template molecule includes an adapter, the adapter including a mismatch portion, wherein the mismatch portion includes a first mismatch sequence in the first strand and a second mismatch sequence in the second strand that is not complementary to the first mismatch sequence; (b) amplifying the double-stranded template molecule of the template-conjugated vector to produce an amplified vector, the amplified vector including a plurality of amplified strands conjugated thereto; (c) sequencing the amplified vector to produce sequencing reads; and (d) determining, at least in part, the percentage of amplified strands originating from the first strand among the plurality of amplified strands in the amplified vector based on portions of the sequencing reads corresponding to the mismatch portions.

[0073] On the other hand, a method for detecting amplified strands on a vector is provided, the method comprising: (a) conjugating a double-stranded template molecule to a vector to produce a template-conjugated vector, wherein the double-stranded template molecule includes a first strand and a second strand, wherein the double-stranded template molecule includes an adapter, the adapter including a mismatch portion, wherein the mismatch portion includes a first homopolymer sequence in the first strand and a second homopolymer sequence in the second strand that is not complementary to the first homopolymer sequence; (b) subjecting the template-conjugated vector to amplification of the double-stranded template molecule to produce an amplified vector, the amplified vector including a plurality of amplified strands conjugated thereto; (c) sequencing the amplified vector to produce sequencing reads; and (d) determining, at least in part, based on portions of the sequencing reads corresponding to the mismatch portion, the percentage of amplified strands in the plurality of amplified strands in the amplified vector originating from the first strand.

[0074] On the other hand, a kit is provided, the kit comprising: a double-stranded linker, the double-stranded linker including a mismatch portion, wherein the double-stranded linker includes a first chain and a second chain, wherein the mismatch portion includes a first mismatch sequence in the first chain and includes a second mismatch sequence in the second chain that is not complementary to the first mismatch sequence.

[0075] In some embodiments, the kit further includes a plurality of double-stranded linkers, the plurality of double-stranded linkers including the mismatch portion, wherein the mismatch portion of the plurality of double-stranded linkers is identical.

[0076] On the other hand, a kit is provided comprising: a double-stranded linker including a mismatch portion, wherein the double-stranded linker includes a first chain and a second chain, wherein the mismatch portion includes a first homopolymer sequence in the first chain and a second homopolymer sequence in the second chain that is not complementary to the first homopolymer sequence.

[0077] In some embodiments, the kit further includes a plurality of double-stranded linkers, the plurality of double-stranded linkers including the mismatch portion, wherein the mismatch portion of the plurality of double-stranded linkers is identical.

[0078] On the other hand, a composition is provided comprising: a double-stranded connector, the double-stranded connector including a mismatch portion, wherein the double-stranded connector includes a first chain and a second chain, wherein the mismatch portion includes a first mismatch sequence in the first chain and includes a second mismatch sequence in the second chain that is not complementary to the first mismatch sequence.

[0079] In some embodiments, the composition further comprises a plurality of double-stranded connectors, each of the plurality of double-stranded connectors including the mismatch portion, wherein the mismatch portion of the plurality of double-stranded connectors is identical.

[0080] In some embodiments, the composition further comprises a plurality of template molecules, wherein the plurality of template molecules include a plurality of double-stranded template insert molecules connected to the plurality of double-stranded linkers.

[0081] In some embodiments, the composition further comprises a template molecule, wherein the template molecule includes a double-stranded template insert molecule connected to the double-stranded linker.

[0082] In some embodiments, the composition further comprises a carrier.

[0083] In some embodiments, the carrier is connected to the template molecule.

[0084] In some embodiments, the composition further comprises a carrier.

[0085] On the other hand, a composition is provided comprising: a double-stranded linker including a mismatch portion, wherein the double-stranded linker includes a first chain and a second chain, wherein the mismatch portion includes a first homopolymer sequence in the first chain and a second homopolymer sequence in the second chain that is not complementary to the first homopolymer sequence.

[0086] In some embodiments, the composition further comprises a plurality of double-stranded connectors, each of the plurality of double-stranded connectors including the mismatch portion, wherein the mismatch portion of the plurality of double-stranded connectors is identical.

[0087] On the other hand, a method for enrichment is provided, the method comprising: (a) providing a plurality of balancing vectors, the plurality of balancing vectors comprising a plurality of amplified strands, wherein the plurality of amplified strands are derived from a plurality of double-stranded template nucleic acid molecules, each of the plurality of double-stranded template nucleic acid molecules comprising a mismatch portion, wherein the mismatch portion comprises a first strand and a second strand, the first strand comprising a first mismatch sequence, and the second strand comprising a second mismatch sequence that is not complementary to the first mismatch sequence, such that each of the plurality of amplified strands comprises, at a corresponding portion corresponding to the mismatch portion, an inverse complementary sequence of the first mismatch sequence or the second mismatch sequence; (b) contacting a plurality of first enrichment molecules with the plurality of amplified strands, wherein the plurality of first enrichment molecules comprises a first capture sequence complementary to the first mismatch sequence to produce a first set of enriched balancing vectors; and (c) contacting a plurality of second enrichment molecules with the first set of enriched balancing vectors, wherein the plurality of second enrichment molecules comprises a second capture sequence comprising the second mismatch sequence to produce a second set of enriched balancing vectors.

[0088] On the other hand, a method for enrichment is provided, the method comprising: (a) providing a plurality of balancing vectors, the plurality of balancing vectors comprising a plurality of amplified strands, wherein the plurality of amplified strands are derived from a plurality of double-stranded template nucleic acid molecules, each of the plurality of double-stranded template nucleic acid molecules comprising a mismatch portion, wherein the mismatch portion comprises a first strand and a second strand, the first strand comprising a first mismatch sequence, and the second strand comprising a second mismatch sequence that is not complementary to the first mismatch sequence, such that each of the plurality of amplified strands comprises, at a corresponding portion corresponding to the mismatch portion, an inverse complementary sequence of the first mismatch sequence or the second mismatch sequence; (b) contacting a plurality of first enrichment molecules with the plurality of amplified strands, wherein the plurality of first enrichment molecules comprises a first capture sequence comprising the second mismatch sequence to produce a first set of enriched balancing vectors comprising the second mismatch sequence; and (c) contacting a plurality of second enrichment molecules with the first set of enriched balancing vectors, wherein the plurality of second enrichment molecules comprises a second capture sequence that is complementary to the first mismatch sequence to produce a second set of enriched balancing vectors.

[0089] On the other hand, a method for enrichment is provided, the method comprising: (a) providing a plurality of balancing vectors, the plurality of balancing vectors comprising a plurality of amplified strands, wherein the plurality of amplified strands are derived from a plurality of double-stranded template nucleic acid molecules, each of the plurality of double-stranded template nucleic acid molecules comprising a mismatch portion, wherein the mismatch portion comprises a first strand and a second strand, the first strand comprising a first mismatch sequence, and the second strand comprising a second mismatch sequence that is not complementary to the first mismatch sequence, such that each of the plurality of amplified strands comprises, at a corresponding portion corresponding to the mismatch portion, an inverse complementary sequence of the first mismatch sequence or the second mismatch sequence; (b) contacting a plurality of first enrichment molecules with the plurality of amplified strands, wherein the plurality of first enrichment molecules comprises a first capture sequence complementary to the first mismatch sequence to capture a subset of balancing vectors; and (c) removing at least a portion of the subset of balancing vectors from the plurality of balancing vectors to produce a first set of enriched balancing vectors.

[0090] On the other hand, a method for enrichment is provided, the method comprising: (a) providing a plurality of balancing vectors, the plurality of balancing vectors comprising a plurality of amplified strands, wherein the plurality of amplified strands are derived from a plurality of double-stranded template nucleic acid molecules, each of the plurality of double-stranded template nucleic acid molecules comprising a mismatch portion, wherein the mismatch portion comprises a first strand and a second strand, the first strand comprising a first mismatch sequence, and the second strand comprising a second mismatch sequence that is not complementary to the first mismatch sequence, such that each of the plurality of amplified strands comprises, at a corresponding portion corresponding to the mismatch portion, an inverse complementary sequence of the first mismatch sequence or the second mismatch sequence; (b) contacting a plurality of first enrichment molecules with the plurality of amplified strands, wherein the plurality of first enrichment molecules comprises a first capture sequence comprising the second mismatch sequence to capture a subset of balancing vectors; and (c) removing at least a portion of the subset of balancing vectors from the plurality of balancing vectors to produce a first set of enriched balancing vectors.

[0091] On the other hand, a method is provided for generating an amplified vector having a predetermined forward-reverse strand ratio, the method comprising: contacting: (i) a template-linked vector, wherein the template-linked vector comprises a vector linked to a double-stranded template molecule, wherein the double-stranded template molecule comprises a first strand and a second strand, wherein the double-stranded template molecule comprises an adapter, the adapter comprising a mismatch portion, wherein the mismatch portion comprises a first mismatch sequence in the first strand and a second mismatch sequence in the second strand that is not complementary to the first mismatch sequence; and (ii) a plurality of forward amplification primers, the plurality of forward amplification primers being at a first predetermined concentration or comprising a first predetermined concentration of the template molecule. (iii) a plurality of forward amplification primers having a first annealing temperature, wherein each of the plurality of forward amplification primers includes an inverse complementary sequence of the first mismatched sequence to hybridize with a first amplified strand derived from the first strand, and (iii) a plurality of reverse amplification primers having a second predetermined concentration or including primers having a second annealing temperature with an inverse complementary sequence of the second mismatched sequence, wherein each of the plurality of reverse amplification primers includes the second mismatched sequence to hybridize with a second amplified strand derived from the second strand, to produce the amplified vector comprising a plurality of amplified strands having the predetermined forward-reverse strand ratio.

[0092] On the other hand, a method for detecting strands on a vector is provided, the method comprising: (a) providing a plurality of balanced vectors immobilized to a substrate, wherein the plurality of balanced vectors comprise a plurality of amplified strands, wherein the plurality of amplified strands are derived from a plurality of double-stranded template nucleic acid molecules, each of the plurality of double-stranded template nucleic acid molecules comprising a mismatch portion, wherein the mismatch portion comprises a first strand and a second strand, the first strand comprising a first mismatch sequence, and the second strand comprising a second mismatch sequence that is not complementary to the first mismatch sequence, such that each of the plurality of amplified strands comprises, at a corresponding portion corresponding to the mismatch portion, an inverse complementary sequence of the first mismatch sequence or the second mismatch sequence; (b) contacting a plurality of first-type probes with the plurality of amplified strands, wherein each of the plurality of first-type probes comprises a first capture sequence complementary to the first mismatch sequence and a first detectable portion; and (c) detecting a first signal, the first signal being derived from the first detectable portion of a subset of the plurality of first-type probes bound to a subset of the plurality of amplified strands.

[0093] In some embodiments, the method further includes: (d) contacting a plurality of second-type probes with the plurality of amplified strands, wherein each of the plurality of second-type probes includes a second capture sequence comprising the second mismatch sequence and a second detectable portion; and (e) detecting a second signal from the second detectable portion of a subset of the plurality of second-type probes combined with a second subset of the plurality of amplified strands.

[0094] On the other hand, an amplification method is provided, the method comprising: (a) providing a double-stranded target molecule and a plurality of adaptors, wherein the double-stranded target molecule includes a first strand and a second strand, and at least one of the plurality of adaptors includes a hairpin sequence; (b) exposing the double-stranded target molecule and the plurality of adaptors to conditions sufficient to connect the adaptors to each end of the double-stranded target molecule, thereby producing a double-stranded template-adaptor molecule, wherein at least one connected adaptor includes a hairpin sequence; and (c) subjecting the double-stranded template-adaptor molecule to amplification to produce a plurality of copies of the double-stranded template-adaptor molecule, wherein each copy of the double-stranded template-adaptor molecule includes a copy of a sequence of the first strand and a copy of a sequence of the second strand.

[0095] In some embodiments, the amplification includes rolling circle amplification (RCA), and the plurality of copies of the double-stranded template-adaptor molecule are linked to each other.

[0096] In some implementations, the amplification includes PCR.

[0097] In some embodiments, the amplification includes loop-mediated isothermal amplification (LAMP).

[0098] In some embodiments, the at least one linker further includes a first region and a second region that do not have sequence complementarity, wherein the first region and the second region are located away from the double-stranded template molecule.

[0099] In some embodiments, the double-stranded template-linker molecule is connected to a vector, wherein the vector includes a plurality of primers, wherein a first subset of the primers is sequence complementary to the first region, and a second subset of the primers is sequence complementary to the second region.

[0100] On the other hand, a sequencing method is provided, the method comprising: (a) providing a double-stranded template molecule, the double-stranded template molecule comprising a first strand and a second strand having sequence complementarity with each other and at least one adapter region including a single-stranded hairpin region; (b) annealing primers with the single-stranded hairpin region; (c) extending the primers to generate a partially single-stranded template molecule, the partially single-stranded template molecule comprising a double-stranded region and a single-stranded region; (d) processing the partially single-stranded template molecule to generate a single-stranded template molecule; and (e) sequencing the single-stranded template molecule.

[0101] In some embodiments, processing (d) the template molecule of the partial single strand includes filtering based on the sequence of the single-stranded region; and sequencing (e) includes targeted sequencing.

[0102] In some embodiments: processing (d) of the template molecule of the partial single strand includes methylation conversion of the single-stranded region; and sequencing (e) includes methylation sequencing.

[0103] On the other hand, a method for sequencing is provided, the method comprising: (a) providing a balanced construct comprising a mixture of a forward strand and a reverse strand, wherein, respectively, (i) the forward strand comprises a first sequence identical to or as the reverse complementary sequence of the first strand of a double-stranded template molecule of a sample, and (ii) the reverse strand comprises a second sequence as the reverse complementary sequence of or identical to the methylated sequence of the second strand of the double-stranded template molecule of the sample; and (b) sequencing the forward strand and the reverse strand by: i. hybridizing primers to the forward strand and the reverse strand, respectively; ii. extending the primers with nucleotides from a nucleotide stream provided according to a repeating stream sequence, wherein the nucleotide stream comprises nucleotides of a single typical base type, wherein the repeating stream sequence comprises three consecutive stream sequences of thymine base stream, cytosine base stream and thymine base stream; and iii. detecting a stream signal indicating the incorporation or absence of a nucleotide by passing the primer after each corresponding nucleotide stream.

[0104] In some embodiments, the method further includes using the flow signal detected in (b)(iii) to determine the methylation state of the double-stranded template molecule.

[0105] In some embodiments, the forward chain includes a first chain recognition element comprising a first homopolymer sequence, and the reverse chain includes a second chain recognition element comprising a second homopolymer sequence, wherein the first homopolymer sequence and the second homopolymer sequence comprise different bases.

[0106] In some embodiments, the method further includes using a subset of the flow signals detected in (b)(iii) corresponding to the first chain identification element and the second chain identification element to determine the forward-reverse ratio of the plurality of copies of the forward chain to the plurality of copies of the reverse chain on the balanced construct.

[0107] In some embodiments, the method further includes determining the methylation state of the double-stranded template molecule based at least in part on the forward-reverse ratio.

[0108] Another aspect of this disclosure provides a system comprising one or more computer processors and computer memory coupled thereto. The computer memory includes machine-executable code that, when executed by the one or more computer processors, performs any of the methods described above or elsewhere herein. Another aspect of this disclosure provides a non-transitory computer-readable medium comprising machine-executable code that, when executed by one or more computer processors, performs any of the methods described above or elsewhere herein.

[0109] Further aspects and advantages of this disclosure will become apparent to those skilled in the art from the following detailed description, in which only illustrative embodiments of the disclosure are shown and described. It should be understood that this disclosure is capable of other and different embodiments, and that certain details thereof can be modified in various obvious ways without departing from this disclosure. Therefore, the drawings and description are to be regarded in an illustrative rather than restrictive manner.

[0110] By incorporating references

[0111] All publications, patents, and patent applications mentioned in this specification are incorporated herein by reference to the extent that each individual publication, patent, or patent application is specifically and individually indicated as incorporated by reference. In the event that any publication or patent or patent application incorporated by reference conflicts with any disclosure included in this specification, this specification is intended to substitute for and / or give precedence to any such conflicting material. Attached Figure Description

[0112] The novel features of the invention are set forth in the appended claims. The features and advantages of the invention can be better understood by referring to the following detailed description and the accompanying drawings (also referred to herein as “Figure” and “FIG.”), which illustrate illustrative embodiments utilizing the principles of the invention:

[0113] Figure 1An example workflow for processing samples for sequencing is shown.

[0114] Figure 2 Examples of individually addressable locations distributed on the substrate, as described in this paper, are shown.

[0115] Figure 3A-3G Different examples of cross-sectional surface profiles of the substrate as described in this paper are shown.

[0116] Figure 4 An example coating of a substrate with a hexagonal lattice bead as described herein is shown.

[0117] Figures 5A-5B Example systems and methods for loading samples or reagents onto a substrate, as described herein, are demonstrated.

[0118] Figure 6 A computerized system for sequencing nucleic acid molecules was demonstrated.

[0119] Figures 7A-7C The multiplexing station in the sequencing system is demonstrated.

[0120] Figure 8 Computer systems that are programmed or otherwise configured to implement the methods provided herein are shown.

[0121] Figure 9A The workflow for extension and amplification on a vector is demonstrated, in which one template strand is not amplified.

[0122] Figure 9B The workflow for extension and amplification on a vector is demonstrated, in which both template strands are amplified.

[0123] Figure 9C This demonstrates an alternative workflow for extension and amplification on a vector, in which both template strands are amplified.

[0124] Figures 10A-10B This demonstrates an alternative workflow for extension and amplification on a vector, in which both template strands are amplified and an adapter including mismatched portions is used.

[0125] Figure 10C-10D This demonstrates an alternative workflow for extension and amplification on a vector, in which both template strands are amplified and an adapter including mismatched portions is used.

[0126] Figure 11 The frequency distribution diagrams of beads with different forward and reverse chain ratios from different samples are shown.

[0127] Figure 12AThe graph shows the relationship between (number of reads) and (average mass of 20bp reads after cyclic jump replacement) for beads with different forward and reverse percentages for the first sample.

[0128] Figure 12B The graph shows the relationship between (number of reads) and (average mass of 20bp reads after cyclic jump replacement) for beads with different forward and reverse percentages for the second sample.

[0129] Figure 13 A two-step enrichment scheme according to an embodiment of the present disclosure is demonstrated.

[0130] Figure 14 The vectors are shown in different example configurations, including template nucleic acid molecules, pre-amplification, and strand recognition elements.

[0131] Figures 15A-15C A rolling circle amplification (RCA) workflow according to an embodiment of the present disclosure is shown, in which both template strands are amplified (e.g., before being ligated to a vector).

[0132] Figure 15D This demonstrates an alternative workflow for amplifying double-stranded template molecules while maintaining forward-reverse strand association (e.g., amplification without losing the double strand).

[0133] Figure 15E The workflow for dumbbell-based amplification using suppression chain replacement conditions is demonstrated.

[0134] Figure 15F The workflow for dumbbell-based amplification using random primers is demonstrated.

[0135] Figure 15G This paper demonstrates potential sources of error during library preparation when performing flat-end shearing and connecting hairpin connectors.

[0136] Figures 16A-16C Showing Figures 15A-15D The document presents alternatives for each stage of the workflow. Figure 16A and 16B An alternative connector sub-constructor was demonstrated. Figure 16C The locations of alternative restriction sites and the resulting product molecules are shown.

[0137] Figure 16D A sequencing method according to an embodiment of the present disclosure is shown.

[0138] Figure 17 Another workflow for amplification according to an embodiment of the present disclosure is shown, in which both template strands are amplified.

[0139] Figure 18A and18B A bridged PCR amplification workflow according to an embodiment of the present disclosure is shown, wherein both template strands are amplified.

[0140] Figure 19 A method for forming a partially double-stranded template molecule according to embodiments of the present disclosure is shown.

[0141] Figure 20 An example workflow for using unique molecular identifiers in double-stranded template molecules is shown.

[0142] Figure 21A An example workflow for generating double-stranded molecules is shown, the double-stranded molecules comprising sequences corresponding to the two strands of a double-stranded insert molecule.

[0143] Figure 21B An additional example workflow for generating double-stranded molecules using bending proteins is shown, the double-stranded molecules comprising sequences corresponding to the two strands of a double-stranded insert molecule.

[0144] Figure 22 A method for generating double-stranded template molecules that are partially converted for use in methylation sequencing is demonstrated.

[0145] Figure 23 Different methods for generating balanced constructs using methylation sequencing of partially transformed molecules are demonstrated.

[0146] Figure 24 This demonstrates an example of how sequencing information from both strands of a template nucleic acid molecule can distinguish between real mutations and artificial (fake) mutations in a sample.

[0147] Figure 25A An example streaming sequencing method that can be used to generate the sequencing data described in this paper is demonstrated.

[0148] Figure 25B An example flowchart is shown.

[0149] Figure 25C Two example sequences (e.g., extended sequencing primer sequences) that differ at at least one locus are shown. Sequence 1 includes TATGGTCATCGA (SEQ ID NO:389), and sequence 2 includes TATGGTCGTCGA (SEQ ID NO:390).

[0150] Figure 25D Sequencing data for two different extended sequences using a single streaming sequence are shown: TATGGTCATCGA (SEQ ID NO:389) and TATGGTCGTCGA (SEQ ID NO:390).

[0151] Figure 25E Sequencing data for two different extended sequences using TACG streaming ordering are presented: TATGGTCATCGA (SEQ ID NO:389) and TATGGTCGTCGA (SEQ ID NO:390).

[0152] Figure 25F Sequencing data for two different extended sequences using AGCT stream sequencing are presented: TATGGTCATCGA (SEQ ID NO:389) and TATGGTCGTCGA (SEQ ID NO:390).

[0153] Figure 26A A comparison of the SNV spectra of patient-matched FF and FFPE samples is presented.

[0154] Figure 26B The SNVQ distributions of standard cfDNA libraries and ppmSeq cfDNA libraries are shown. ppmSeq libraries are divided into libraries that collect information from both strands (mixed) and libraries that primarily collect information from a single strand (non-mixed).

[0155] Figure 26C The correlation between circulating tumor fractions measured for each spectrum in eight patient-specific spectra using standard library preparation methods and ppmSeq library preparation methods is demonstrated.

[0156] Figure 26D This displays the estimated tumor fractions for standard cfDNA preparation (left) and ppmSeq (right) in matched cfDNA samples and control cfDNA samples. Each column corresponds to a patient's SNV profile, with one matched cfDNA sample and nine negative control samples.

[0157] Figure 27 The error distribution for different mutation types and different measurement types is shown.

[0158] Figure 28 An example of a template molecule linked by a chain-recognized linker is shown. Detailed Implementation

[0159] While various embodiments of the invention have been shown and described herein, it will be apparent to those skilled in the art that these embodiments are provided by way of example only. Various changes, modifications, and substitutions will occur to those skilled in the art without departing from the invention. It should be understood that various alternatives to the embodiments of the invention described herein may be employed.

[0160] As used herein, unless the context clearly indicates otherwise, the singular forms “a / an” and “the” include plural indicators.

[0161] Where a range of values ​​is provided, it should be understood that every intermediate value between the upper and lower limits of the range, as well as any other stated or intermediate values ​​within the range, are covered within the scope of this disclosure. Where the range includes an upper or lower limit, the range excluding either of those included limits is also included in this disclosure.

[0162] As used herein, the term "biological sample" generally refers to any sample derived from a subject or specimen. A biological sample can be a fluid, tissue, collection of cells (e.g., a cheek swab), hair sample, or fecal sample. Fluids can be blood (e.g., whole blood), saliva, urine, or sweat. Tissues can be derived from organs (e.g., liver, lung, or thyroid gland) or large amounts of cellular material, such as tumors. A biological sample can be a cellular sample or a cell-free sample. Examples of biological samples include nucleic acid molecules, amino acids, peptides, proteins, carbohydrates, fats, or viruses. In one instance, a biological sample is a nucleic acid sample comprising one or more nucleic acid molecules, such as deoxyribonucleic acid (DNA) and / or ribonucleic acid (RNA). The nucleic acid sample may include cell-free nucleic acid molecules, such as cell-free DNA or cell-free RNA. Further, samples can be extracted from a variety of animal fluids containing cell-free sequences, including but not limited to blood, serum, plasma, vitreous humor, sputum, urine, tears, sweat, saliva, semen, mucosal excretions, mucus, cerebrospinal fluid, amniotic fluid, lymph, etc. Cell-free polynucleotides can be of fetal origin (via fluid obtained from a pregnant subject) or can be derived from the subject's own tissue. Biological samples can also refer to samples engineered to mimic one or more characteristics (e.g., nucleic acid sequence characteristics, such as sequence identity, length, GC content, etc.) of samples derived from a subject or sample.

[0163] As used herein, the term "subject" generally refers to an individual from whom a biological sample was obtained. A subject can be a mammal or a non-mammal. A subject can be a human, a non-human mammal, an animal, an ape, a monkey, a chimpanzee, a reptile, an amphibian, a bird, or a plant. A subject can be a patient. A subject may exhibit symptoms of a disease. A subject may be asymptomatic. A subject may be receiving treatment. A subject may not be receiving treatment. A subject may have or be suspected of having a disease such as cancer (e.g., breast cancer, colorectal cancer, brain cancer, leukemia, lung cancer, skin cancer, liver cancer, pancreatic cancer, lymphoma, esophageal cancer, cervical cancer, etc.) or an infectious disease. Subjects may have or be suspected of having genetic disorders such as achondroplasia, alpha-1 antitrypsin deficiency, antiphospholipid syndrome, autism, autosomal dominant polycystic kidney disease, Charcot-Marie-tooth syndrome, Cri-du-chat syndrome, Crohn's disease, cystic fibrosis, Dercum disease, Down syndrome, Duane syndrome, Duchenne muscular dystrophy, factor V Leidenthrombophilia, familial hypercholesterolemia, familial Mediterranean fever, fragile X syndrome, Gaucher disease, hemochromatosis, hemophilia, holoencephaly, Huntington's disease, Klinefelter syndrome, Marfan syndrome, myotonic dystrophy, neurofibromatosis, and Noonan syndrome. Osteogenesis imperfecta, Parkinson's disease, phenylketonuria, Poland anomaly, porphyria, progeria, retinitis pigmentosa, severe complex immunodeficiency, sickle cell disease, spinal muscular atrophy, Tay-Sachs disease, thalassemia, trimethylaminuria, Turner syndrome, palatal-cardiac syndrome, WAGR syndrome, or Wilson's disease.

[0164] As used herein, the term "analyte" generally refers to a subject being analyzed, or a subject being analyzed directly or indirectly during the process, whether or not it is a subject being analyzed. Analytes may be synthetic. Analytes may be derived from and / or derived from samples, such as biological samples. In some instances, analytes are or include molecules, macromolecules (e.g., nucleic acids, carbohydrates, proteins, lipids, etc.), nucleic acids, carbohydrates, lipids, antibodies, antibody fragments, antigens, peptides, polypeptides, proteins, macromolecular groups (e.g., glycoproteins, proteoglycans, ribozymes, liposomes, etc.), cells, tissues, biological particles or organisms, or any engineered copy or variant thereof, or any combination thereof. As used herein, the term "processing an analyte" generally refers to one or more stages of interaction with one or more samples. Processing an analyte may include chemical, biochemical, enzymatic, hybridization, polymerization, physical, any other reaction, or a combination thereof, in the presence of or on the analyte. Processing an analyte may include physical and / or chemical manipulation of the analyte. For example, processing analytes may include detecting chemical or physical changes, adding or removing materials, atoms or molecules, molecular confirmation, detecting the presence of fluorescent labels, detecting Forster resonance energy transfer (FRET) interactions, or inferring the absence of fluorescence.

[0165] As used herein, the terms “nucleic acid,” “nucleic acid molecule,” “nucleic acid sequence,” “nucleic acid fragment,” “oligonucleotide,” and “polynucleotide” generally refer to polynucleotides that can have bases of varying lengths, including, for example, deoxyribonucleotides, deoxyribonucleic acid (DNA), ribonucleotides, or ribonucleic acid (RNA) or analogues thereof. Nucleic acids can be single-stranded. Nucleic acids can be double-stranded. Nucleic acids can be partially double-stranded, such as having at least one double-stranded region and at least one single-stranded region. Partially double-stranded nucleic acids may have one or more overhanging regions. As used herein, “overhanging” generally refers to a single-stranded portion of a nucleic acid that extends from or is connected to the double-stranded portion of the same nucleic acid molecule, and the single-stranded portion is located at the 3' or 5' end of the same nucleic acid molecule. Non-restrictive examples of nucleic acids include DNA, RNA, genomic DNA, or synthetic DNA / RNA, or coding or non-coding regions of genes or gene fragments, loci (loci / locus) defined by linkage analysis, exons, introns, messenger RNA (mRNA), transfer RNA, ribosomal RNA (rRNA), short interfering RNA (siRNA), short hairpin RNA (shRNA), microRNA (miRNA), ribozymes, cDNA, recombinant nucleic acids, branched nucleic acids, plasmids, vectors, isolated DNA of any sequence, and isolated RNA of any sequence. Nucleic acids can be at least about 10 bases (“bases”), 20 bases, 30 bases, 40 bases, 50 bases, 100 bases, 200 bases, 300 bases, 400 bases, 500 bases, 1 kilobits (kb), 2kb, 3kb, 4kb, 5kb, 10kb, 20kb, 30kb, 40kb, 50kb, 100kb, 200kb, 300kb, 400kb, 500kb, 1 megabase (Mb), 10Mb, 100Mb, 1 gigabase or more. Nucleic acids can include a sequence of four natural nucleotide bases: adenine (A); cytosine (C); guanine (G); and thymine (T) (or, when the nucleic acid is RNA, uracil (U) instead of thymine (T)). Nucleic acids may include one or more non-standard nucleotides, nucleotide analogs, and / or modified nucleotides.

[0166] As used herein, the term "nucleotide" generally refers to any nucleotide or nucleotide analogue. Nucleotides can be naturally occurring or non-natural. Nucleotides can be modified, synthetic, or engineered. Nucleotides can include typical or atypical bases. Nucleotides can include substituted bases. Nucleotides can include modified polyphosphate chains (e.g., triphosphates coupled to a fluorophore). Nucleotides can include tags. Nucleotides can be terminated (e.g., reversibly terminated). Non-standard nucleotides, nucleotide analogs, and / or modified analogs may include, but are not limited to, diaminopurine, 5-fluorouracil, 5-bromouracil, 5-chlorouracil, 5-iodouracil, hypoxanthine, xanthine, 4-acetylcytosine, 5-(carboxyhydroxymethyl)uracil, 5-carboxymethylaminomethyl-2-thiouracil, 5-carboxymethylaminomethyluracil, dihydrouracil, beta-D-galactosylqueuosine, inosine, N6-isopentene adenine, 1-methylguanine, 1-methylinosine, 2,2-dimethylguanine, 2-methyladenine, 2-methylguanine, 3-methylcytosine, 5-methylcytosine, N6-adenine, 7-methylguanine, 5-methylaminomethyluracil, and 5-methoxyaminomethyluracil. 2-thiouracil, β-D-mannosylqueuosine, 5'-methoxycarboxymethyluracil, 5-methoxyuracil, 2-methylthio-D46-isopentene adenine, uracil-5-oxyacetic acid (v), huaitoxyglycin, pseudouracil, queuosine, 2-thiocytosine, 5-methyl-2-thiouracil, 2-thiouracil, 4-thiouracil, 5-methyluracil, methyl uracil-5-oxyacetic acid, uracil-5-oxyacetic acid (v), 5-methyl-2-thiouracil, 3-(3-amino-3-N-2-carboxypropyl)uracil, (acp3)w, 2,6-diaminopurine, ethynyl nucleotide bases, 1-propynyl nucleotide bases, azide nucleotide bases, selenophosphate nucleic acids, etc. In some cases, the phosphate moiety of a nucleotide may include modifications, including modifications to the triphosphate moiety. Other non-limiting examples of modifications include longer phosphate chains (e.g., phosphate chains having 4, 5, 6, 7, 8, 9, 10, or more phosphate moieties), modifications with thiol moieties (e.g., α-thiotriphosphate and β-thiotriphosphate), or modifications with selenium moieties (e.g., selenophosphate nucleic acids). Nucleic acids may also be modified at the base moiety (e.g., at one or more atoms that are normally available to form hydrogen bonds with complementary nucleotides and / or at one or more atoms that are normally not available to form hydrogen bonds with complementary nucleotides), the sugar moiety, or the phosphate backbone.Nucleic acids may also include amine-modified groups, such as aminoallyl-dUTP (aa-dUTP) and aminohexylacrylamide-dCTP (aha-dCTP), to achieve covalent linkage with amine-reactive moieties (such as N-hydroxysuccinimide ester (NHS)). Alternatives to standard DNA or RNA base pairs in the oligonucleotides of this disclosure can provide higher bit density per cubic mm², greater safety (resistance to accidental or intentional synthesis of natural toxins), easier recognition in photoprogrammed polymerases, or lower secondary structure. Nucleotides can also react with or bind to detectable moieties for nucleotide detection.

[0167] As used herein, the term “sequencing” generally refers to the process of generating or identifying a sequence of a biomolecule, such as a nucleic acid molecule. The sequence can be a nucleic acid sequence comprising a sequence of nucleic acid bases. As used herein, the term “template nucleic acid” generally refers to the nucleic acid to be sequenced. The template nucleic acid can be an analyte or associated with an analyte. For example, the analyte can be mRNA, and the template nucleic acid is mRNA or cDNA derived from mRNA, or other derivatives thereof. In another instance, the analyte can be a protein, and the template nucleic acid is an oligonucleotide conjugated to an antibody that binds to the protein, or a derivative thereof. Examples of sequencing include, for example, single-molecule sequencing or synthetic sequencing. Sequencing can include generating a sequencing signal and / or sequencing reads. Sequencing can be performed on a template nucleic acid immobilized on a vector, such as a flow cell, substrate, and / or one or more beads. In some cases, the template nucleic acid can be amplified to generate colonies of nucleic acid molecules linked to the vector to produce an amplified sequencing signal. In one example, (i) the template nucleic acid undergoes a nucleic acid reaction, such as amplification, to generate a clonal population of nucleic acids linked to beads, which are immobilized to a substrate; (ii) the amplified sequencing signal from the immobilized beads is detected from the substrate surface during or after one or more nucleotide streams; and (iii) the sequencing signal is processed to generate sequencing reads. Multiple beads, each comprising a different nucleic acid colony, may be immobilized at different locations on the substrate surface, and multiple sequencing signals from different immobilized beads at different locations may be processed simultaneously or substantially simultaneously once the substrate surface is detected to generate multiple sequencing reads. In some sequencing methods, the nucleotide stream includes non-terminated nucleotides. In some sequencing methods, the nucleotide stream includes terminated nucleotides.

[0168] As used herein, the term “nucleotide stream” generally refers to instances where nucleotide-containing reagents are provided to the sequencing reaction space at different times. As used herein, the term “stream” generally refers to a nucleotide stream when not limited by another reagent. For example, providing two streams can refer to (i) providing a nucleotide-containing reagent (e.g., a solution containing A bases) to the sequencing reaction space at a first time point, and (ii) providing a nucleotide-containing reagent (e.g., a solution containing G bases) to the sequencing reaction space at a second time point different from the first time point. A “sequencing reaction space” can be any reaction environment comprising template nucleic acids. For example, a sequencing reaction space can be or include a substrate surface comprising template nucleic acids immobilized thereon; a substrate surface containing beads immobilized thereon comprising template nucleic acids immobilized thereon; or any reaction chamber or surface comprising template nucleic acids, which may or may not be immobilized. Nucleotide streams can have any number of base types (e.g., A, T, G, C; or U), such as 1, 2, 3, or 4 typical base types. As used in this article, "stream sequence" generally refers to the sequence of nucleotides used to sequence template nucleic acids. The stream sequence can be represented as a one-dimensional matrix or linear array of bases, corresponding to the identity of the nucleotide streams provided to the sequencing reaction space and arranged in chronological order:

[0169] (For example, [ATGCATGCATGATGATGAT GC ATGC]).

[0170] This one-dimensional matrix or linear array of bases in a flow sequence may also be referred to herein as a “flow space.” A flow sequence can have any number of nucleotide flows. As used herein, a “flow position” generally refers to the sequential position of a given nucleotide flow entering the flow space (e.g., elements in a one-dimensional matrix or linear array). As used herein, a “flow cycle” generally refers to the sequence of nucleotide flows of a subgroup of consecutive nucleotide flows within a flow sequence. A flow cycle can be represented as a one-dimensional matrix or linear array of base sequences that correspond to the identity of the nucleotide flows provided within a subgroup of consecutive flows and are arranged in chronological order (e.g., [ATGC], [AATTGGCC], [AT], [A / TA / G], [AA], [A], [ATG], etc.). A flow cycle can have any number of nucleotide flows. A given flow cycle can be repeated once or more, either consecutively or discontinuously, in the flow sequence. Therefore, as used herein, the term “flow cycle sequence” generally refers to the ordering of flow cycles within a flow sequence and can be expressed in units of flow cycles. For example, if [ATGC] is identified as the first flow cycle and [ATG] is identified as the second flow cycle, the flow sequence [AT GC AT GC ATGATGATG AT GCATGC] can be described as a flow cycle sequence having [first flow cycle; first flow cycle; second flow cycle; second flow cycle; second flow cycle; first flow cycle; first flow cycle]. Alternatively or additionally, the flow cycle sequence can be described as [cycle 1, cycle 2, cycle 3, cycle 4, cycle 5, cycle 6], where cycle 1 is the first flow cycle, cycle 2 is the first flow cycle, cycle 3 is the second flow cycle, and so on.

[0171] As used herein, the term "locus" generally refers to a base position in a sequence. In some cases, a locus may refer to a specific position in a reference sequence, such as a reference genome, or in a natural sample or source. A locus may or may not be a base pair. A locus can have any length, such as 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more bases or base pairs. A locus can refer to any continuous sequence region, such as a homopolymer or heteropolymer portion of a sequence. The same locus described relative to two or more sequences, chains, and / or molecules may refer to a common specific base position or common specific position relative to a natural sample or reference sequence.

[0172] The terms “amplifying,” “amplification,” and “nucleic acid amplification” are used interchangeably and generally refer to the production of one or more copies of a nucleic acid or template. For example, “amplification” of DNA generally refers to the production of one or more copies of a DNA molecule. Nucleic acid amplification can be linear, exponential, or a combination thereof. Amplification can be emulsion-based or non-emulsion-based. Non-limiting examples of nucleic acid amplification methods include reverse transcription, primer extension, polymerase chain reaction (PCR), ligase chain reaction (LCR), helicase-dependent amplification, asymmetric amplification, rolling circle amplification (RCA), recombinase polymerase reaction (RPA), loop-mediated isothermal amplification (LAMP), sequence-based amplification (NASBA), autonomous sequence replication (3SR), and multiple substitution amplification (MDA). When using PCR, any form of PCR can be used. Non-limiting examples include real-time PCR, allele-specific PCR, assembly PCR, asymmetric PCR, digital PCR, emulsion PCR (ePCR or emPCR), emulsion RPA (eRPA), dial-out PCR, helicase-dependent PCR, nested PCR, hot-start PCR, reverse PCR, methylation-specific PCR, small primer PCR, multiplex PCR, nested PCR, overlap-extension PCR, thermal asymmetric staggered PCR, and falling PCR. Amplification can be performed in a reaction mixture containing various components that participate in or promote amplification (e.g., primers, template, nucleotides, polymerase, buffer components, cofactors, etc.). In some cases, the reaction mixture includes a buffer that allows its components to be independently incorporated into nucleotides. Non-limiting examples include magnesium ion, manganese ion, and isocitrate buffers. Further examples of such buffers are described in Tabor, S. et al., CCPNAS, 1989, 86, 4076-4080, and U.S. Patent Nos. 5,409,811 and 5,674,716, each of which is incorporated herein by reference in its entirety.Useful methods for clonal amplification from single molecules include rolling circle amplification (RCA) (Lizardi et al., Nat. Genet. 19:225-232 (1998), which is incorporated herein by reference in its entirety), bridge PCR (Adams and Kron, Method for Performing Amplification of Nucleic Acid with Two Primers Bound to a Single Solid Support, Mosaic Technologies, Inc. (Winter Hill, Mass.); Whitehead Institute for Biomedical Research, Cambridge, Mass., (1997); Adessi et al., Nucl. Acids Res. 28:E87 (2000); Pemov et al., Nucl. Acids Res. 33:e11 (2005); or U.S. Patent No. 5,641,658, each of which is incorporated herein by reference in its entirety), polymerase colony generation (polony) generation (Mitra et al., Proc. Natl. Acad. Sci. USA 100:5926-5931 (2003); Mitra et al., Anal. Biochem. 320:55-65 (2003), each of which is incorporated herein by reference) and clonal amplification on beads using emulsions (Dressman et al., Proc. Natl. Acad. Sci. USA 100:8817-8822 (2003), which is incorporated herein by reference) or ligation with bead-based linker libraries (Brenner et al., Nat. Biotechnol. 18:630-634 (2000); Brenner et al., Proc. Natl. Acad. Sci. USA 97:1665-1670 (2000)); Reinartz et al., Brief Funct. Genomic Proteomic 1:95-104 (2002), each of which is incorporated herein by reference. Amplification products from nucleic acids may be identical or substantially identical. Nucleic acid colonies generated by amplification may have identical or substantially identical sequences.

[0173] As used herein, when referring to two or more nucleic acid or polypeptide sequences, the term "identical" or "percentage of identity" means two or more sequences that, when compared and aligned to obtain maximum correspondence, have the same or alternatively a specific percentage of the same amino acid residues or nucleotides, as determined by one or more of the following sequence comparison algorithms: Needleman-Wunsch (see, for example, Needleman, Saul B.; and Wunsch, Christian D. (1970). "A general method applicable to the search for similarities in the amino acid sequence of two proteins", Journal of Molecular Biology 48(3):443-53); Smith-Waterman (see, for example, Smith, Temple F.; and Waterman, Michael S., "Identification of common molecular subsequences." (1981) Journal of Molecular Biology). Biology 147:195-197); or BLAST (Basic Local Alignment Search Tool); see, for example, Altschul SF, Gish W, Miller W, Myers EW, Lipman DJ, “Basic local alignment search tool” (1990) J Mol Biol 215(3):403-410), each of which is incorporated herein by reference in its entirety. As used herein, when referring to two or more nucleic acid or polypeptide sequences, the terms “substantially identical” or “substantially identical” means two or more sequences or subsequences (such as biologically active fragments) that, when compared and aligned to obtain maximum correspondence, have at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% nucleotide or amino acid residue identity, as determined by sequence comparison algorithms or by visual inspection. Broadly consistent sequences are generally considered homologous, regardless of their actual pedigree. In some implementations, "fundamental consistency" exists in regions of the sequences being compared.In some embodiments, substantial identity exists in regions of at least 25 residues, at least 50 residues, at least 100 residues, at least 150 residues, at least 200 residues, or greater than 200 residues. In some embodiments, the compared sequences are substantially identical across their full length. Generally, substantially identical nucleic acid or protein sequences include less than 100% nucleotide or amino acid residue identity, and such sequences are generally considered "identical."

[0174] As used herein, the term "detector" generally refers to a device capable of detecting signals, including signals indicating the presence or absence of one or more incorporated nucleotides or fluorescently labeled signals. Detectors can detect multiple signals simultaneously or substantially simultaneously. Detectors can detect signals in real time during, substantially during, or after a biological reaction (such as sequencing reactions, e.g., sequencing during primer extension reactions). In some cases, detectors may include optical and / or electronic components capable of detecting signals. Non-limiting examples of detection methods used by detectors include optical detection, spectroscopic detection, electrostatic detection, electrochemical detection, acoustic detection, magnetic detection, etc. Optical detection methods include, but are not limited to, light absorption, ultraviolet-visible (UV-vis) light absorption, infrared light absorption, light scattering, Rayleigh scattering, Raman scattering, surface-enhanced Raman scattering, Mie scattering, fluorescence, luminescence, and phosphorescence. Spectroscopic detection methods include, but are not limited to, mass spectrometry, nuclear magnetic resonance (NMR) spectroscopy, and infrared spectroscopy. Electrostatic detection methods include, but are not limited to, gel-based techniques, such as gel electrophoresis. Electrochemical detection methods include, but are not limited to, electrochemical detection of amplified products after separation by high-performance liquid chromatography (HPLC). The detector can be a continuous-area scanning detector. For example, the detector may include an imaging array sensor capable of continuous integration over the scanning area, wherein the scanning is electronically synchronized with an image of an object in relative motion. Continuous-area scanning detectors may include time-delay and integration (TDI) charge-coupled devices (CCDs), hybrid TDIs, complementary metal-oxide-semiconductor (CMOS) pseudo-TDI devices, or TDI line-scan cameras.

[0175] As used herein, the terms 'tagment', 'tagmentation', and 'tagmented' can refer to the use of a transposase (e.g., Tn5 transposase) to link an oligonucleotide to a target molecule (e.g., a double-stranded target molecule). For example, when a Tn5 transposase is loaded with a desired adaptor, tagging can be used to link the adaptor to the target molecule. Descriptions of tagging methods are provided, for example, in: Picelli. 2014. Tn5 transposase and tagmentation procedures formassively scaled sequencing projects, Genome Res. 24(12), 2033–2040. In some cases, tagging is an alternative method for linking multiple double-stranded molecules together.

[0176] Sample processing methods

[0177] This article describes apparatus, systems, methods, compositions, and kits for processing samples, such as for preparing samples for sequencing, for sequencing samples, and / or for analyzing sequencing data. Figure 1 An example sequencing workflow 100 according to the apparatus, system, method, composition and kit of this disclosure is shown.

[0178] Vectors and / or template nucleic acids (101) can be prepared and / or provided to be compatible with downstream sequencing operations (e.g., 107). Vectors (e.g., beads) can be used to help facilitate sequencing of template nucleic acids on a substrate. Vectors can help immobilize template nucleic acids onto the substrate, such as when the template nucleic acid is coupled to the vector, and the vector is subsequently immobilized to the substrate. Vectors can further act as binding entities to retain molecules of the template nucleic acid colonies (e.g., including copies of the same or substantially the same sequence as the template nucleic acid) together for any downstream processing, such as for sequencing operations. This can be particularly useful in distinguishing colonies from other colonies (e.g., on other vectors) and in generating amplified sequencing signal for the template nucleic acid sequence.

[0179] The prepared and / or provided vector may include oligonucleotides comprising one or more functional nucleic acid sequences. For example, the vector may include a capture sequence configured to capture or conjugate a template nucleic acid (or a treated template nucleic acid). For example, the vector may include a capture sequence, primer sequence, barcode sequence, sample index sequence, unique molecular identifier (UMI), flow cell adaptor sequence, adaptor sequence, binding sequence for any molecule (e.g., splint, primer, template nucleic acid, capture sequence, etc.), or any other functional sequence or any combination thereof for downstream applications. The oligonucleotide may be single-stranded, double-stranded, or partially double-stranded.

[0180] The carrier may include one or more capturing entities, wherein the capturing entities are configured to be captured by the capturing entity. The capturing entity may be coupled to an oligonucleotide conjugated with the carrier. The capturing entity may be coupled to the carrier. For example, the capturing entity may include streptavidin (SA) when the capturing portion includes biotin. In another instance, when the capturing entity includes a capturing sequence (e.g., a capturing oligonucleotide complementary to a complementary capturing sequence), the capturing entity may include a complementary capturing sequence. In another instance, the capturing entity may include a device, system, or apparatus configured to apply a magnetic field when the capturing entity includes magnetic particles. In another instance, the capturing entity may include a device, system, or apparatus configured to apply an electric field when the capturing entity includes charged particles. In some cases, the capturing entity may include one or more other mechanisms configured to capture the capturing entity. The capturing entities and the capturing entities may bind to, be coupled to, hybridize with, or otherwise associate with each other. Association may include the formation of covalent bonds, non-covalent bonds, and / or releasable bonds (e.g., cleavable bonds that are cleavable upon application of a stimulus). In some cases, association may not form any bonds. For example, association can increase the physical proximity (or decrease the physical distance) between capturing entities. In some cases, a single capturing entity can associate with a single capturing entity. Alternatively, a single capturing entity can associate with multiple capturing entities. The capturing entity can be linked to a nucleotide. Chemically modified bases, including biotin, azides, cyclooctyne, tetrazolium, and thiol, as well as many other bases, are suitable as capturing entities. The capturing entity / capturing entity pair can be any combination. Such pairs can include, but are not limited to, biotin / streptavidin, azide / cyclooctyne, and thiol / maleimide. It should be understood that any of these pairs can be used as a capturing entity or a capturing entity. In some cases, the capturing entity can include a secondary capturing entity, for example, for subsequent capture by a secondary capturing entity. Secondary capturing entities and secondary capturing entities can include any one or more of the capturing mechanisms described elsewhere herein (e.g., biotin and streptavidin, complementary capturing sequences, etc.). In some cases, the secondary trapping entity may include magnetic particles (e.g., magnetic beads), and the secondary trapping entity may include a magnetic system (e.g., a magnet, device, system, or apparatus configured to apply a magnetic field). In some cases, the secondary trapping entity may include charged particles (e.g., charged beads carrying a charge), and the secondary trapping entity may include an electrical system (e.g., a magnet, device, system, or apparatus configured to apply an electric field).

[0181] The vector may include one or more cleavable portions. A cleavable portion may be a part of or linked to an oligonucleotide conjugated to the vector. The cleavable portion may be conjugated to the vector. The cleavable portion may include any useful cleavable or excisable part that can be used to cleave the oligonucleotide (or a portion thereof) from the vector. For example, the cleavable portion may include uracil, ribonucleotides, or other modified nucleotides that can be excised or cleaved by an enzyme (e.g., UDG, RNase, endonuclease, exonuclease, etc.). The cleavable portion may include a base-free site or analogues of a base-free site (e.g., dSpacer), dideoxyribose. The cleavable portion may include spacers, such as C3 spacers, hexanediol, triethylene glycol spacers (e.g., spacer 9), hexaethylene glycol spacers (e.g., spacer 18), or combinations thereof or similar. The cleavable portion may include an optically cleavable portion. The cleavable portion may include modified nucleotides, such as methylated nucleotides. Modified nucleotides may be enzyme-specifically recognized (e.g., methylated nucleotides may be recognized by MspJI). The cleavable portion can be enzymatically cleaved (e.g., using enzymes such as UDG, RNase, APE1, MspJI, etc.). Alternatively or additionally, the cleavable portion can be cleaved using one or more stimuli (e.g., light stimulation, chemical stimulation, heat stimulation, etc.).

[0182] In some instances, a single vector comprises copies of oligonucleotides from a single species that are identical or substantially identical to each other. In some instances, a single vector comprises copies of oligonucleotides from at least two species (e.g., including different sequences). For example, a single vector may include a first subset of oligonucleotides configured to capture a first adaptor sequence of a template nucleic acid and a second subset of oligonucleotides configured to capture a second adaptor sequence of a template nucleic acid.

[0183] In some instances, a population of vectors of a single species can be prepared and / or provided, wherein all vectors within a single species are identical (e.g., having the same oligonucleotide composition (e.g., sequence) etc.). In some instances, a population of vectors of multiple species can be prepared and / or provided. For example, a population of vectors can be prepared to include multiple unique vector species, wherein each unique vector species includes primer sequences unique to said vector species. When a template nucleic acid is ligated to a vector, only template nucleic acid including a given adaptor sequence compatible with (e.g., at least partially complementary to) a given primer sequence can be ligated to a given vector of a vector species including the given primer sequence. In another instance, a population of vectors can be prepared such that each unique vector species includes multiple primer sequences (e.g., a pair of primer sequences) unique to said vector species. In some embodiments, the systems and methods disclosed herein may include a population of vectors comprising two, three, four, five, six, seven, eight, nine, ten, or more unique vector species. Each unique vector species may include a unique primer sequence that allows selective interaction between the corresponding vector species and its intended binding partner (e.g., a complementary nucleic acid sequence within the adaptor region of the template nucleic acid or an intermediate primer sequence that can subsequently bind to a complementary nucleic acid sequence within the adaptor region of the sample nucleic acid). Populations of multiple species vectors can be prepared by first preparing different populations of all different single-species vectors, and then mixing these different populations of single-species vectors to produce a final population of multiple species vectors. The concentrations of the different vector species within the final mixture can be adjusted accordingly. Apparatus, systems, methods, compositions, and kits for preparing and using vector species are further described in detail in U.S. Patent Publication No. 20220042072A1 and International Patent Publication No. WO 2022040557A2, each of which is incorporated herein by reference in its entirety for all purposes.

[0184] Template nucleic acids may include insert sequences derived from biological samples. In some cases, the insert sequence may be derived from a larger nucleic acid (e.g., an endogenous nucleic acid) or its reverse complementary sequence in the biological sample (e.g., by fragmentation, transposition, and / or replication of the larger nucleic acid). Template nucleic acids may be derived from any nucleic acid in the biological sample and produced by any number of nucleic acid processing operations, such as, but not limited to, fragmentation, degradation or digestion, transposition, ligation, reverse transcription, extension, etc. The prepared and / or provided template nucleic acid may include one or more functional nucleic acid sequences. In some cases, the one or more functional nucleic acid sequences may be placed at one end of the insert sequence. In some cases, the one or more functional nucleic acid sequences may be isolated and placed at both ends of the insert sequence, for example, to clamp the insert sequence. In some cases, nucleic acid molecules including the insert sequence or its complementary sequence may be linked to one or more adaptor oligonucleotides including such functional nucleic acid sequences. In some cases, nucleic acid molecules including the insert sequence or its complementary sequence may hybridize with primers including such functional nucleic acid sequences and extend to produce template nucleic acids including such functional nucleic acid sequences. In some cases, nucleic acid molecules, including insert sequences or their complementary sequences, can hybridize with primers comprising one or more functional nucleic acid sequences and extend to produce intermediate molecules, and the intermediate molecules can hybridize with primers comprising additional functional nucleic acid sequences and extend, etc., to perform any number of extension reactions to produce template nucleic acids comprising one or more functional nucleic acid sequences. For example, the template nucleic acid may include an adaptor sequence configured to be captured by a capture sequence on an oligonucleotide coupled to a vector. For example, the template nucleic acid may include a capture sequence, primer sequence, barcode sequence, sample index sequence, unique molecular identifier (UMI), flow cell adaptor sequence, adaptor sequence, binding sequence for any molecule (e.g., splint, primer, template nucleic acid, capture sequence, etc.), or any other functional sequence or any combination thereof that can be used for downstream operations. The template nucleic acid may be single-stranded, double-stranded, or partially double-stranded.

[0185] The template nucleic acid may include one or more capture entities described elsewhere herein. In some cases, in the workflow, only the vector includes a capture entity and the template nucleic acid does not. In other cases, in the workflow, only the template nucleic acid includes a capture entity and the vector does not. In other cases, both the template nucleic acid and the vector include capture entities. In other cases, neither the vector nor the template nucleic acid includes capture entities.

[0186] The template nucleic acid may include one or more cleavable portions described elsewhere herein. In some cases, in the workflow, only the vector includes a cleavable portion and the template nucleic acid does not. In other cases, in the workflow, only the template nucleic acid includes a cleavable portion and the vector does not. In other cases, both the template nucleic acid and the vector include cleavable portions. In other cases, neither the vector nor the template nucleic acid includes cleavable portions. For example, the cleavable portion may be strategically placed based on the desired downstream amplification workflow.

[0187] In some instances, libraries of insert sequences are processed to provide a population of template sequences with the same conformation, such as positions of the same sequence and / or one or more functional sequences. For example, the population of template sequences may include multiple nucleic acid molecules, each including the same first adaptor sequence linked to the same end. In some instances, libraries of insert sequences are processed to provide a population of template sequences with different conformations (such as positions of different sequences and / or one or more functional sequences). For example, the population of template sequences may include a first subset of nucleic acid molecules and a second subset of nucleic acid molecules, each of the first subset including the same first adaptor sequence at a first end, and each of the second subset including the same second adaptor sequence at a second end, wherein the second adaptor sequence is different from the first adaptor sequence. In some cases, populations of template sequences with different conformations (e.g., different adaptor sequences) may be used in conjunction with populations of vectors from multiple species, such as to reduce polyclonal problems during downstream amplification. A population of template nucleic acids with multiple configurations can be prepared by first preparing different populations of template nucleic acids with all different single configurations, and then mixing these different populations of template nucleic acids with single configurations to produce a final population of template nucleic acids with multiple configurations. The concentrations of template nucleic acids with different configurations in the final mixture can be adjusted accordingly.

[0188] Optionally, the vector and / or template nucleic acid can be pre-enriched (102). For example, vectors comprising different oligonucleotide sequences are isolated from a mixture comprising vectors not having different oligonucleotide sequences. Alternatively, a population of vectors can be provided to comprise substantially homogeneous vectors, wherein each vector comprises the same surface primer molecule immobilized thereon. For example, template nucleic acids comprising different conformations (e.g., comprising a specific adaptor sequence) are isolated from a mixture comprising template nucleic acids not having different conformations. Alternatively, a population of template nucleic acids can be provided to comprise substantially homogeneous conformations. In some cases, capture entities on the vector and / or template nucleic acid are used for pre-enrichment.

[0189] After preparing the vector and template nucleic acid, they can be ligated (103). The template nucleic acid can be coupled to the vector by any method that induces stable association between the template nucleic acid and the vector. For example, the template nucleic acid can hybridize with an oligonucleotide on the vector. In another instance, the template nucleic acid can hybridize with one or more intermediate molecules, such as splints, bridges, and / or primer molecules, which hybridize with an oligonucleotide on the vector. Alternatively or additionally, the template nucleic acid can be ligated to one or more nucleic acids on or coupled to the vector. Alternatively or additionally, the template nucleic acid can be fragmented in solution or with a nucleic acid coupled to the vector. Alternatively or additionally, the template nucleic acid can hybridize with an oligonucleotide on the vector, which includes a primer sequence and is subsequently extended from the primer sequence. Once ligated, multiple vector-template complexes can be generated.

[0190] Optionally, the vector-template complex can be pre-enriched (104), wherein the vector-template complex is isolated from a mixture comprising a vector and / or template nucleic acid that are not connected to each other. In some cases, capture entities on the vector and / or template nucleic acid are used for pre-enrichment.

[0191] After the template nucleic acid molecule is ligated to the vector, the template nucleic acid can undergo an amplification reaction (105) to produce multiple amplification products immobilized to the vector. For example, such amplification reactions can include polymerase chain reaction (PCR) or any other amplification method described herein, including but not limited to emulsion PCR (ePCR or emPCR), isothermal amplification (e.g., recombinase polymerase amplification (RPA), emulsion RPA (eRPA)), bridging amplification, template walking, etc. In some cases, the amplification reaction can occur while the vector is immobilized to the substrate. In other cases, the amplification reaction can occur outside the substrate, such as in solution, or on different surfaces or platforms. In some cases, the amplification reaction can occur in separate reaction volumes, such as within multiple droplets of an emulsion during emulsion PCR (ePCR or emPCR) or emulsion RPA (eRPA), or in wells. Emulsion PCR methods are further described in detail in U.S. Patent Publication No. 20220042072A1 and International Patent Publication No. WO 2022040557A2, each of which is incorporated herein by reference in its entirety for all purposes. Emulsion RPA can be performed in separated droplets within an emulsion. For example, in eRPA, the droplets can include beads (e.g., single beads), a template (which may or may not be pre-ligated to the beads) (e.g., a single template), and RPA reagents (e.g., recombinases, single-stranded binding (SSB) proteins, polymerases, primers, dNTPs, etc.).

[0192] Following amplification, the vector (e.g., comprising a template nucleic acid) may undergo post-amplification processing (106). Typically, after amplification, the resulting mixture may comprise a mixture of positive vectors (e.g., those comprising a template nucleic acid molecule) and negative vectors (e.g., those not linked to a template nucleic acid molecule). An enrichment procedure can separate positive vectors from the mixture. Example enrichment methods for amplified vectors are described in U.S. Patent No. 10,900,078 and U.S. Patent Publication No. 20210079464A1 and International Patent Publication No. WO 2022040557A2, each of which is incorporated herein by reference in its entirety. For example, a substrate enrichment procedure may immobilize only positive vectors onto a substrate surface to separate positive vectors. In some cases, positive vectors may be immobilized at desired locations (e.g., individually addressable locations) on the substrate surface, distinguished from undesirable locations (e.g., spacers between individually addressable locations). In some cases, positive and / or negative vectors can be processed to selectively remove unamplified surface primers (on the vector), such that the resulting positive vector retains the template nucleic acid molecule, and the resulting negative vector is stripped of the unamplified surface primers. Subsequently, the template nucleic acid on the positive vector can be used, for example, to enrich the positive vector by capturing the template nucleic acid.

[0193] Following amplification and post-processing, the template nucleic acid can undergo sequencing (107). The template nucleic acid can be sequenced while ligated to a vector. Alternatively, the template nucleic acid molecule may be vector-free when performing sequencing and / or analysis. In some cases, the template nucleic acid can be sequenced while ligated to a vector immobilized to a substrate. Examples of substrate-based sample processing systems are described elsewhere in this document. Any sequencing method described elsewhere in this document may be used. In some cases, sequencing by synthesis (SBS) is performed.

[0194] In one embodiment (Example A), the SBS method includes flowing nucleotide reagents according to a repeating sequence of streams comprising a 4-base stream (e.g., [A / T / G / C]), wherein each nucleotide is reversibly terminated (e.g., dideoxynucleotide), and wherein each base is labeled with a different dye (producing a different light signal). For each stream, other sequencing reagents, such as sequencing primers, polymerase, buffers, etc., are present to provide sufficient conditions for the incorporation of reversibly terminated, labeled nucleotides into the growth strand hybridizing with the template nucleic acid. After each stream, the incorporation event or the absence of each base can be detected by probing the different dyes in the four channels. After an incorporation event in a stream (where at most one nucleotide is incorporated into each growth strand due to the termination state), termination can be reversed (e.g., cleaving the termination portion) to allow subsequent stepwise incorporation events in subsequent streams. After each or one detection event, the label can be removed (e.g., cleaved) to reduce signal noise for the next detection. In another embodiment (Example B), the SBS method includes flowing nucleotide reagents according to a repeating flow sequence comprising four single-base streams (e.g., [AT GC]), wherein each nucleotide is reversibly terminated and each base is labeled with the same dye (producing an optical signal of the same frequency). For each stream, other sequencing reagents, such as sequencing primers, polymerase, buffers, etc., are present to provide sufficient conditions for the reversibly terminated, labeled nucleotide incorporation into the growth chain hybridizing with the template nucleic acid. After each stream, an incorporation event or a deficiency of a specific base in the stream can be detected by probing the wavelength of the dye. After an incorporation event in a stream (where at most one nucleotide is incorporated into each growth chain due to the termination state), termination can be reversed (e.g., cleaving the termination portion) to allow subsequent stepwise incorporation events in subsequent streams. After each or one detection event, the label can be removed (e.g., cleaved) to reduce signal noise for the next detection. In another embodiment (Example C), the SBS method includes flowing nucleotide reagents according to a repeating flow sequence comprising four single-base streams (e.g., [AT GC]), wherein each nucleotide is not terminated and each base is labeled with the same dye (producing an optical signal of the same frequency). For each stream, other sequencing reagents, such as sequencing primers, polymerase, buffers, etc., are present to provide sufficient conditions for the incorporation of labeled nucleotides into the growth chain that hybridizes with the template nucleic acid. After each stream, an incorporation event of a specific base or a lack of a specific base in the stream can be detected by probing the wavelength of the dye. Because the nucleotides are not terminated, multiple nucleotides can be incorporated during a single stream if the growth chain extends through the homopolymer region (e.g., poly-T region, etc.) of the template nucleic acid.After each or one detection event, the markers can be removed (e.g., cleavage dye) to reduce signal noise for the next detection. In another embodiment (Example D), the SBS method includes flowing nucleotide reagents according to a repeating sequence of flow cycles comprising four single base streams (e.g., [AT GC]), wherein each nucleotide is not terminated, and only a portion of the bases in each stream (e.g., less than 90%, 80%, 70%, 60%, 50%, 40%, 30%, 20%, 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, 1%, etc.) is marked with the same dye (producing an optical signal of the same frequency). For each stream, other sequencing reagents, such as sequencing primers, polymerase, buffers, etc., are present to provide sufficient conditions for the incorporation of nucleotides into the growth chain that hybridizes with the template nucleic acid. After each stream, the incorporation event of a specific base or the absence of a specific base in the stream can be detected by probing the wavelength of the dye. Because the nucleotides are not terminated, multiple nucleotides can be incorporated during a single flow if the growth chain extends through the homopolymer region (e.g., poly-T region, etc.) of the template nucleic acid. After each or more detection events, the label (e.g., cleavage dye) can be removed to reduce signal noise for the next detection. In another embodiment (Example E), the SBS method includes flowing nucleotide reagents according to a repeating flow sequence comprising eight single-base flows, wherein each of the four typical base types flows twice consecutively within the flow cycle (e.g., [AAT TG GCC]), wherein each nucleotide is not terminated, and wherein only a portion of the bases in each other flow in the flow cycle (e.g., less than 90%, 80%, 70%, 60%, 50%, 40%, 30%, 20%, 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, 1%, etc.) is labeled with the same dye (producing an optical signal of the same frequency), and nucleotides in alternating other flows are unlabeled. For each stream, additional sequencing reagents, such as sequencing primers, polymerase, buffers, etc., are available to provide sufficient conditions for the incorporation of nucleotides into the growth chain that hybridizes with the template nucleic acid. After one or two streams for each typical base type, an incorporation event or absence of a specific base in the stream can be detected by probing the wavelength of the dye. Because the nucleotides are not terminated, multiple nucleotides can be incorporated during a single stream if the growth chain extends across a homopolymer region (e.g., a poly-T region) of the template nucleic acid. A first stream of a typical base type (e.g., A) is followed by a second stream of the same typical base type (e.g., A), which can help facilitate completion of the incorporation reaction across each growth chain, such as to reduce phase fixation problems. After each or more detection events, the marker can be removed (e.g., cleavage dye) to reduce signal noise for the next detection.

[0195] Labeled nucleotides can include dyes, fluorophores, or quantum dots.

[0196] It should be understood that, except as listed in Example AE, the termination state on the nucleotides, the type of label (e.g., the type of dye or other detectable portion), the proportion of labeled nucleotides in the stream, the type of nucleotide bases in each stream, the type of nucleotide bases in each stream cycle, and / or the sequence of streams in the stream cycle and / or the combination of stream sequences can vary depending on the SBS method.

[0197] Following sequencing, the collected and / or generated sequencing signals can undergo data analysis (108). Sequencing signals can be processed to generate base recognition and / or sequencing reads. In some cases, sequencing reads can be processed to generate diagnostic data for a biological sample or for a subject from which the biological sample is derived. In some cases, as described elsewhere herein, sequencing reads can be processed to determine whether the amplified cluster originates from only one or both strands of the double-stranded template nucleic acid molecule and / or to determine the ratio or percentage of amplified strands from each strand of the double-stranded template nucleic acid molecule.

[0198] Although about Figure 1 The sequencing workflow 100 describes the use of vectors to bind to template molecules, but it should be understood that different vectors can be effectively replaced by using spatially different locations on one or more surfaces, and said one or more surfaces need not be the surface of a single vector (e.g., a bead). For example, a first spatially different location on a surface can directly immobilize a first colony of a first template nucleic acid, and a second spatially different location on the same surface (or different surfaces) can directly immobilize a second colony of a second template nucleic acid to distinguish it from the first colony. In some cases, the surface including the spatially different locations can be the surface of a substrate on which the sample is sequenced, thereby simplifying the amplification-sequencing workflow.

[0199] It should be understood that in some cases, the different operations described in sequencing workflow 100 may be performed in a different order. It should be understood that in some cases, one or more operations described in sequencing workflow 100 may be omitted or replaced by other similar operations. It should be understood that in some cases, one or more additional operations described in sequencing workflow 100 may be performed.

[0200] The different operations described in sequencing workflow 100 can be performed with the aid of the open substrate system described in this article.

[0201] Open base system

[0202] This document describes apparatuses, systems, and methods for processing samples using open substrate or open flow cell geometry. As used herein, the term "open substrate" generally refers to a substrate on which any point on the active surface can be physically accessed from a direction perpendicular to the substrate. The apparatuses, systems, and methods described herein can be used to facilitate any application or process involving a reaction or interaction between two objects, such as between an analyte and a reagent or between two reagents. For example, the reaction or interaction can be chemical (e.g., polymerase reaction) or physical (e.g., displacement). The apparatuses, systems, and methods described herein can benefit from greater efficiency, such as faster reagent delivery and lower reagent volumes required per surface area. The apparatuses, systems, and methods described herein avoid the contamination problems common in microfluidic channel flow cells fed from multi-port valves, which can be sources of one reagent carried to another. The apparatuses, systems, and methods can benefit from shorter completion times, use of fewer resources (e.g., various reagents), and / or reduced system costs. As described herein, open substrate or flow cell geometries can be used to process any analyte from any sample, such as, but not limited to, nucleic acid molecules, protein molecules, antibodies, antigens, cells, and / or organisms. As described herein, open substrate or flow cell geometries can be used in any application or process, such as, but not limited to, sequencing by synthesis, ligation sequencing, amplification, proteomics, single-cell processing, barcoding technology, and sample preparation.

[0203] A sample processing system may include a substrate and means and systems for performing one or more operations on or on the substrate. The sample processing system may allow for efficient dispensing of reagents onto the substrate. Sample processing may allow for efficient imaging of one or more analytes or their corresponding signals on the substrate. The sample processing system may include an imaging system comprising a detector. Substrates and detectors that may be used in the sample processing system are described in further detail in U.S. Patent Publications 20200326327A1, 20210354126A1, and 20210079464A1, each of which is incorporated herein by reference in its entirety for all purposes.

[0204] base

[0205] The substrate may be a solid-phase substrate. The substrate may include, wholly or partially, one or more of the following: rubber; glass; silicon; metals such as aluminum, copper, titanium, chromium, or steel; ceramics such as titanium oxide or silicon nitride; plastics such as polyethylene (PE), low-density polyethylene (LDPE), high-density polyethylene (HDPE), polypropylene (PP), polystyrene (PS), high-impact polystyrene (HIPS), polyvinyl chloride (PVC), polyvinylidene chloride (PVDC), acrylonitrile butadiene styrene (ABS), polyacetylene, polyamide, polycarbonate, polyester, polyurethane, polyepoxide, polymethyl methacrylate (PMMA), polytetrafluoroethylene (PTFE), phenolic resin (PF), melamine-formaldehyde (MF), urea-formaldehyde (UF), polyetheretherketone (PEEK), polyetherimide (PEI), polyimide, polylactic acid (PLA), furan, silicone, polysulfone, any mixture of the foregoing materials, or any other suitable material. The substrate may be completely or partially coated with one or more of the following: metals, such as aluminum, copper, silver or gold; oxides, such as silicon oxide (SiO2). x O y (where x and y can take any possible values), photoresist such as SU8, surface coating such as aminosilane or hydrogel, polyacrylic acid, polyacrylamide dextran, polyethylene glycol (PEG), or any mixture of the foregoing materials or any other suitable coating. The substrate may comprise multiple layers of the same or different types of materials. The substrate may be completely or partially opaque to visible light. The substrate may be completely or partially transparent to visible light. The surface of the substrate may be modified to include active chemical groups such as amines, esters, hydroxyl groups, epoxides, etc., or combinations thereof. The surface of the substrate may be modified to include any binders or linkers described herein. In some cases, such binders, linkers, active chemical groups, etc., may be added to the substrate as additional layers or coatings.

[0206] The substrate can be in the general form of a cylinder, a cylindrical shell or disk, a rectangular prism, or any other geometry. The thickness of the substrate (e.g., the minimum dimension) can be at least 100 micrometers (μm), at least 200 μm, at least 500 μm, at least 1 mm, at least 2 millimeters (mm), at least 5 mm, at least 10 mm, or greater. The first lateral dimension of the substrate (such as the width of a substrate in the general form of a rectangular prism, or the radius or diameter of a substrate in the general form of a cylinder) and / or the second lateral dimension (such as the length of a substrate in the general form of a rectangular prism) can be at least 1 mm, at least 2 mm, at least 5 mm, at least 10 mm, at least 20 mm, at least 50 mm, at least 100 mm, at least 200 mm, at least 500 mm, at least 1,000 mm, or greater.

[0207] One or more surfaces of the substrate may be exposed to and accessible from the surrounding open environment. For example, an array may be exposed to and accessible from such an open environment. In some cases, as described elsewhere herein, the surrounding open environment may be controlled and / or confined to a larger controlled environment.

[0208] The substrate may include multiple individually addressable locations. Individually addressable locations may include physically accessible locations for manipulation. Manipulation may include, for example, placement, extraction, reagent dispensing, inoculation, heating, cooling, or stirring. Manipulation can be achieved through, for example, local microfluidic, pipette, optical, laser, acoustic, magnetic, and / or electromagnetic interactions with the analyte or its surrounding environment. Individually addressable locations may include digitally accessible locations. For example, each individually addressable location may be located, identified, and / or accessed electronically or digitally for indexing, mapping, sensing, associating with devices (e.g., detectors, processors, dispensers, etc.), or otherwise processed.

[0209] The plurality of individually addressable locations can be arranged in an array on the substrate randomly or according to any pattern. Figure 2Different bases (top views) are shown, including different arrangements of individually addressable locations 201. Inset A shows a substantially rectangular base with a regular linear array, inset B shows a substantially annular base with a regular linear array, and inset C shows a base of arbitrary shape with an irregular array. The base can have any number of individually addressable locations, for example, at least 1, at least 2, at least 5, at least 10, at least 20, at least 50, at least 100, at least 200, at least 500, at least 1,000, at least 2,000, at least 5,000, at least 10,000, at least 20,000, at least 50,000, at least 100,000, at least 200,000, at least 500,000, at least 1,000,000, at least 2,000,000, at least 5,000,000, at least 10,000. The substrate may have at least 20,000,000, at least 50,000,000, at least 100,000,000, at least 200,000,000, at least 500,000,000, at least 1,000,000,000, at least 2,000,000,000, at least 5,000,000,000, at least 10,000,000,000, at least 20,000,000,000, at least 50,000,000,000, at least 100,000,000,000 or more individually addressable locations. The substrate may have multiple individually addressable locations, which are within the range defined by any two of the foregoing values.

[0210] Each individually addressable location can have a circular, recessed, raised, rectangular, or any other general shape or form (e.g., polygonal, non-polygonal). Multiple individually addressable locations can have the same shape or form or different shapes or forms. Individually addressable locations can have any size. In some cases, the area of ​​an individually addressable location can be approximately 0.1 square micrometers (μm). 2 ), approximately 0.2μm 2 Approximately 0.25μm 2 Approximately 0.3μm 2 Approximately 0.4μm 2 Approximately 0.5μm 2 Approximately 0.6μm 2 Approximately 0.7μm 2 Approximately 0.8μm 2 Approximately 0.9μm 2 Approximately 1μm 2 Approximately 1.1 μm 2 Approximately 1.2μm 2Approximately 1.25μm 2 Approximately 1.3μm 2 Approximately 1.4μm 2 Approximately 1.5μm 2 Approximately 1.6μm 2 Approximately 1.7μm 2 Approximately 1.75μm 2 Approximately 1.8μm 2 Approximately 1.9μm 2 Approximately 2μm 2 Approximately 2.25μm 2 Approximately 2.5μm 2 Approximately 2.75μm 2 Approximately 3μm 2 Approximately 3.25μm 2 Approximately 3.5μm 2 Approximately 3.75μm 2 Approximately 4μm 2 Approximately 4.25μm 2 Approximately 4.5μm 2 Approximately 4.75μm 2 Approximately 5μm 2 Approximately 5.5μm 2 Approximately 6μm 2 Or larger. The area of ​​an individually addressable location can be within the range defined by any two of the aforementioned values. The area of ​​an individually addressable location can be less than approximately 0.1 μm. 2 or greater than approximately 6 μm 2 .

[0211] Individually addressable locations can be distributed on the substrate, with the spacing determined by the distance between the center of a first location and the center of the nearest or adjacent individually addressable location. The locations can be spaced apart at the following intervals: approximately 0.1 micrometers (μm), approximately 0.2 μm, approximately 0.25 μm, approximately 0.3 μm, approximately 0.4 μm, approximately 0.5 μm, approximately 0.6 μm, approximately 0.7 μm, approximately 0.8 μm, approximately 0.9 μm, approximately 1 μm, approximately 1.1 μm, approximately 1.2 μm, approximately 1.25 μm, approximately 1.3 μm, approximately 1.4 μm, approximately 1.5 μm, approximately 1.6 μm, approximately 1.7 μm, approximately 1.75 μm, and approximately 1. The possible values ​​are approximately 8 μm, 1.9 μm, 2 μm, 2.25 μm, 2.5 μm, 2.75 μm, 3 μm, 3.25 μm, 3.5 μm, 3.75 μm, 4 μm, 4.25 μm, 4.5 μm, 4.75 μm, 5 μm, 5.5 μm, 6 μm, 6.5 μm, 7 μm, 7.5 μm, 8 μm, 8.5 μm, 9 μm, 9.5 μm, or 10 μm. In some cases, the location can be positioned at a spacing within a range defined by any two of the aforementioned values. The location can be positioned at a spacing less than approximately 0.1 μm or greater than approximately 10 μm. In some cases, the spacing between two individually addressable locations can be determined as a function of the size of the loaded object (e.g., beads). For example, if the loaded object is a bead with the largest diameter, the spacing can be at least approximately the largest diameter of the loaded object.

[0212] Each of the plurality of individually addressable sites, or each subset of such sites, can immobilize an analyte (e.g., nucleic acid molecules, protein molecules, carbohydrate molecules, etc.) or a reagent (e.g., nucleic acid molecules, probe molecules, barcode molecules, antibody molecules, primer molecules, beads, etc.). In some cases, the analyte or reagent can be immobilized to the individually addressable site using a carrier (e.g., beads). In one example, beads are immobilized to the individually addressable site, and the analyte or reagent is immobilized onto the beads. In some cases, multiple analytes or reagents can be immobilized to the individually addressable site, as with a carrier. The substrate can immobilize multiple analytes or reagents across multiple individually addressable sites. The multiple analytes or reagents can be of the same type (e.g., nucleic acid molecules) or can be a combination of different types of analytes or reagents (e.g., nucleic acid molecules, protein molecules, etc.). In one instance, a first bead comprising a first cluster of nucleic acid molecules, each comprising a first template sequence, is fixed to a first individually addressable location, and a second bead comprising a second cluster of nucleic acid molecules, each comprising a second template sequence, is fixed to a second individually addressable location.

[0213] The substrate may include more than one type of individually addressable sites, which are arranged in an array on the substrate randomly or according to any pattern. In some cases, different types of individually addressable sites may have different chemical, physical, and / or biological properties (e.g., hydrophobicity, charge, color, morphology, size, dimension, geometry, etc.). For example, a first type of individually addressable site may be used with a first type of bioanalyte but not with a second type of bioanalyte, and a second type of individually addressable site may be used with a second type of bioanalyte but not with a first type of bioanalyte.

[0214] In some cases, individually addressable locations may include different surface chemical compositions. Different surface chemical compositions can distinguish different addressable locations. Different surface chemical compositions can differentiate individually addressable locations on a substrate from surrounding locations. For example, a first location type may include a first surface chemical composition, and a second location type may lack the first surface chemical composition. In another example, a first location type may include a first surface chemical composition, and a second location type may include a different second surface chemical composition. The first location type may have a first affinity for an object (e.g., beads containing nucleic acid molecules (e.g., amplicons) immobilized thereon), and the second location type may have a different second affinity for the same object due to different surface chemical compositions. In other examples, a first location type including a first surface chemical composition may have an affinity for a first sample type (e.g., beads containing nucleic acid molecules (e.g., amplicons) immobilized thereon) and exclude a second sample type (e.g., beads lacking nucleic acid molecules (e.g., amplicons) immobilized thereon). The first and second location types may or may not be arranged alternately on the surface. For example, a first position type or region type may include a positively charged surface chemical component, and a second position type or region type may include a negatively charged surface chemical component. In another example, a first position type or region type may include a hydrophobic surface chemical component, and a second position type or region type may include a hydrophilic surface chemical component. In another example, a first position type includes a binding agent as described elsewhere herein, and a second position type does not include a binding agent or includes a different binding agent. In some cases, the surface chemical component may include an amine. In some cases, the surface chemical component may include a silane (e.g., tetramethylsilane). In some cases, the surface chemical component may include hexamethyldisilazane (HMDS). In some cases, the surface chemical component may include (3-aminopropyl)triethoxysilane (APTMS). In some cases, the surface chemical component may include a surface primer molecule or any oligonucleotide molecule having any degree of affinity for another molecule. In one example, the substrate includes a plurality of individually addressable sites, each defined by an APTMS, which are positively charged and have an affinity for amplified beads exhibiting a negative charge (e.g., beads comprising nucleic acid molecules (e.g., amplicons) immobilized thereon). Sites surrounding the plurality of individually addressable sites may include HMDS that repel amplified beads.

[0215] In some cases, individually addressable locations can be indexed, such as in spatial indexing. Data collected over multiple time periods corresponding to indexed locations can be linked to the same indexed location. In some cases, during iterations of a synthetic sequencing stream, sequencing signal data collected from an indexed location is linked to the indexed location to generate sequencing reads for an analyte fixed at the indexed location. In some implementations, individually addressable locations are indexed by partitioning a portion of a surface, such as by etching or grooving the surface, using dyes or inks, depositing topographic markers, depositing samples (e.g., control nucleic acid samples), depositing reference objects (e.g., reference beads that consistently emit detectable signals during detection), etc., and such partitions can be used to index individually addressable locations. As will be understood, a combination of positive and negative partitions (lack of partitions) can be used to index individually addressable locations. In some implementations, each of the individually addressable locations is indexed. In some implementations, a subset of individually addressable locations is indexed. In some implementations, individually addressable locations are not indexed, and different regions of the substrate are indexed.

[0216] The substrate may include a planar or substantially planar surface. A substantially planar surface may refer to a flatness on the micrometer scale (e.g., the extent of unevenness on a planar surface does not exceed the micrometer scale) or a flatness on the nanometer scale (e.g., the extent of unevenness on a planar surface does not exceed the nanometer scale). Alternatively, a substantially planar surface may refer to a flatness smaller than the nanometer scale or larger than the micrometer scale (e.g., the millimeter scale). Alternatively or additionally, the surface of the substrate may be textured or patterned. For example, the substrate may include grooves, valleys, mounds, and / or pillars. The substrate may define one or more cavities (e.g., micrometer-scale cavities or nanometer-scale cavities). The substrate may define one or more channels. The substrate may have a regular texture and / or pattern across its surface. For example, the substrate may have a regular geometry above or below a surface reference level (e.g., wedges, cuboids, cylinders, spheres, hemispheres, etc.). Alternatively, the substrate may have an irregular texture and / or pattern across its surface. In some cases, the texture of the substrate may include a structure with a maximum dimension of up to about 100%, 90%, 80%, 70%, 60%, 50%, 40%, 30%, 20%, 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, 1%, 0.1%, 0.01%, 0.001%, 0.0001%, or 0.00001% of the total thickness of the substrate or its layers. In some cases, the texture and / or pattern of the substrate may define at least a portion of individually addressable locations on the substrate. The textured and / or patterned substrate may be substantially planar. Figure 3A-3G Different examples of cross-sectional surface profiles of the substrate are shown. Figure 3AA cross-sectional surface profile of a substrate with a completely flat surface is shown. Figure 3B A cross-sectional surface profile of a substrate with a hemispherical groove or recess is shown. Figure 3C The cross-sectional surface profile of a base having columns or alternatively, or both, holes is shown. Figure 3D A cross-sectional surface profile of a substrate with a coating is shown. Figure 3E A cross-sectional surface profile of a substrate with spherical particles is shown. Figure 3F Showing Figure 3B The cross-sectional surface profile, in which the first type of binder is inoculated or associated with the corresponding groove. Figure 3G Showing Figure 3B The cross-sectional surface profile, in which a second type of binder is inoculated or associated with a corresponding groove.

[0217] The binder can be configured to immobilize an analyte or reagent to individually addressable sites. In some cases, the surface chemical composition of the individually addressable sites may include one or more binders. In some cases, multiple individually addressable sites may be coated with a binder. In some cases, at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% of the total number of individually addressable sites or the surface area of ​​the substrate are coated with a binder. The binder may be integral with the array. The binder may be added to the array. For example, the binder may be added to the array as one or more coatings on the array. The substrate may include orders of magnitude of at least about 10, 100, 10 3 10 4 10 5 10 6 10 7 10 8 10 9 10 10 10 11 Or more binders. Alternatively or additionally, the substrate may comprise orders of magnitude up to about 10. 11 10 10 10 9 10 8 10 7 10 6 10 5 10 4 10 3 100, 10 or less binder.

[0218] Binders can immobilize analytes or reagents through nonspecific interactions, such as one or more of the following: hydrophilic interactions, hydrophobic interactions, electrostatic interactions, physical interactions (e.g., adhesion to a column or sedimentation within a well), etc. Alternatively or additionally, binders can immobilize analytes or reagents through specific interactions. For example, in the case where the analyte or reagent is a nucleic acid molecule, the binder may include an oligonucleotide adaptor configured to bind to the nucleic acid molecule. In other instances, binders may include one or more of the following: antibodies, oligonucleotides, nucleic acid molecules, adaptors, affinity-binding proteins, lipids, carbohydrates, etc. Binders can immobilize analytes or reagents through any possible combination of interactions. For example, binders can immobilize nucleic acid molecules through combinations of physical and chemical interactions, combinations of protein and nucleic acid interactions, etc. In some cases, a single binder can bind a single analyte (e.g., a nucleic acid molecule) or a single reagent. In some cases, a single binder can bind multiple analytes (e.g., multiple nucleic acid molecules) or multiple reagents. In some cases, multiple binders can bind a single analyte or a single reagent. While the examples described herein depict the interaction of binders with nucleic acid molecules, binders can immobilize other molecules (such as proteins), other particles, cells, viruses, other organisms, etc. Similarly, while the examples described herein depict the interaction of binders with samples or analytes, binders can also similarly immobilize reagents. In some cases, the substrate may include multiple types of binders, for example, to bind different types of analytes or reagents. For example, a first type of binder (e.g., oligonucleotide) is configured to bind a first type of analyte (e.g., nucleic acid molecule) or reagent, and a second type of binder (e.g., antibody) is configured to bind a second type of analyte (e.g., protein) or reagent. In another example, a first type of binder (e.g., a first type of oligonucleotide molecule) is configured to bind a first type of nucleic acid molecule, and a second type of binder (e.g., a second type of oligonucleotide molecule) is configured to bind a second type of nucleic acid molecule. For example, the substrate may be configured to bind different types of analytes or reagents to certain portions or specific locations on the substrate by having different types of binders in certain portions or specific locations on the substrate.

[0219] The substrate can rotate about an axis. The axis of rotation may or may not be an axis passing through the center of the substrate. In some cases, the systems, apparatus, and devices described herein may further include an automatic or manual rotation unit configured to rotate the substrate. The rotation unit may include a motor and / or a rotor to rotate the substrate. For example, the substrate may be fixed to a chuck (such as a vacuum chuck). The substrate can rotate at the following speeds: at least 1 revolution per minute (rpm), at least 2 rpm, at least 5 rpm, at least 10 rpm, at least 20 rpm, at least 50 rpm, at least 100 rpm, at least 200 rpm, at least 500 rpm, at least 1,000 rpm, at least 2,000 rpm, at least 5,000 rpm, at least 10,000 rpm, or higher. Alternatively or additionally, the substrate may rotate at speeds of up to approximately the following: 10,000 rpm, 5,000 rpm, 2,000 rpm, 1,000 rpm, 500 rpm, 200 rpm, 100 rpm, 50 rpm, 20 rpm, 10 rpm, 5 rpm, 2 rpm, 1 rpm, or lower. The substrate may be configured to rotate at speeds within a range defined by any two of the foregoing values. The substrate may be configured to rotate at different speeds during different operations described herein. The substrate may be configured to rotate at speeds varying according to a time-dependent function (such as a ramp, sine wave, pulse, or other function or combination of functions). The time-varying function may be periodic or non-periodic.

[0220] During rotation, analytes or reagents can be fixed to the substrate. Analytes or reagents can be dispensed onto the substrate before or during rotation. When the substrate rotates at a relatively high speed, high-speed coating across the substrate can be achieved during rotation by tangential inertia guiding unconstrained rotating reagents in a partially radial direction (i.e., away from the axis of rotation); this phenomenon is often referred to as centrifugal force. In some cases, the substrate can rotate at a relatively low speed so that reagents dispensed to one location do not move to another location due to rotation, or move only minimally due to rotation, allowing for controlled dispensing of reagents to the desired location. For controlled dispensing, the substrate can be rotated at the following frequencies: no more than 60 rpm, no more than 50 rpm, no more than 40 rpm, no more than 30 rpm, no more than 25 rpm, no more than 20 rpm, no more than 15 rpm, no more than 14 rpm, no more than 13 rpm, no more than 12 rpm, no more than 11 rpm, no more than 10 rpm, no more than 9 rpm, no more than 8 rpm, no more than 7 rpm, no more than 6 rpm, no more than 5 rpm, no more than 4 rpm, no more than 3 rpm, no more than 2 rpm, or no more than 1 rpm. In some cases, the rotation frequency can be within the range defined by any two of the aforementioned values. In some cases, during controlled dispensing, the substrate can be rotated at a frequency of approximately 5 rpm. The substrate rotation speed can be adjusted according to appropriate operations (e.g., high speed for spin coating, high speed for washing the substrate, low speed for sample loading, low speed for detection, etc.).

[0221] In some cases, the substrate can move in any vector or direction. For example, such movement can be nonlinear (e.g., rotation about an axis), linear, or a mixture of linear and nonlinear motion. In some cases, the systems, apparatuses, and devices described herein may further include motion units configured to move the substrate. Motion units can include any mechanical components such as motors, rotors, actuators, linear platforms, rollers, pulleys, etc., to move the substrate. During any such movement, an analyte or reagent can be fixed to the substrate. The analyte or reagent can be dispensed onto the substrate before, during, or after the movement of the substrate.

[0222] Reagents were loaded onto an open substrate.

[0223] The surface of the substrate may be in fluid communication with at least one fluid nozzle (of the fluid channel). The surface may be in fluid communication with the fluid nozzle through a non-solid gap (e.g., an air gap). In some cases, the surface may additionally be in fluid communication with at least one fluid outlet. The surface may be in fluid communication with the fluid outlet through an air gap. The nozzle may be configured to direct the solution into the array. The outlet may be configured to receive the solution from the substrate surface. One or more dispensing nozzles may be used to direct the solution into the surface. For example, at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or more dispensing nozzles may be used to direct the solution into the array. Multiple nozzles within the range defined by any two of the foregoing values ​​may be used to direct the solution into the array. In some cases, different reagents (e.g., different types of nucleotide solutions, different probes, washing solutions, etc.) may be dispensed through different nozzles, such as to prevent contamination. Each nozzle may be connected to a dedicated fluid line or fluid valve, which may further prevent contamination. A type of reagent can be dispensed via one or more nozzles. The one or more nozzles may be directed to be located at or near the center of the substrate. Alternatively, the one or more nozzles may be directed to be located at or near a location on the substrate other than the center of the substrate. Alternatively or in combination, one or more nozzles may be directed closer to the center of the substrate than one or more other nozzles. For example, one or more nozzles for dispensing a detergent may be directed closer to the center of the substrate than one or more nozzles for dispensing an active reagent. The one or more nozzles may be arranged at different radii from the center of the substrate. Two or more nozzles may be operated in combination to deliver fluid to the substrate more efficiently. One or more nozzles may be configured to deliver fluid to the substrate in the form of a jet, spray (or other dispersed fluid), and / or droplets. One or more nozzles may be operated to atomize the fluid before delivery to the substrate. For example, the fluid may be delivered as aerosol particles.

[0224] In some cases, the solution can be dispensed onto the substrate while it is stationary; the substrate can then undergo rotation (or other motion) after solution dispensing. Alternatively, the substrate can undergo rotation (or other motion) before solution dispensing; the solution can then be dispensed onto the substrate while it is being rotated (or otherwise moved). In some cases, the rotation of the substrate may generate centrifugal forces (or inertial forces directed away from the axis) on the solution, causing the solution to flow radially outward across the array. In this way, the rotation of the substrate can guide the solution across the array. Continuous rotation of the substrate over a period of time can dispense a fluid film of nearly constant thickness across the array.

[0225] One or more conditions, such as the rotational speed of the substrate, the acceleration of the substrate (e.g., rate of change of velocity), the viscosity of the solution, the angle of solution dispensing (e.g., contact angle of the reagent flow), the radial coordinates of the solution dispensing (e.g., center, eccentric, etc.), the temperature of the substrate, the temperature of the solution, and other factors, can be adjusted and / or optimized to achieve desired wetting on the substrate and / or to achieve desired film thickness on the substrate, such as to promote uniform coating of the substrate. For example, one or more conditions can be applied to obtain film thicknesses of at least 10 nanometers (nm), 20 nm, 50 nm, 100 nm, 200 nm, 500 nm, 1 micrometer (μm), 2 μm, 5 μm, 10 μm, 20 μm, 50 μm, 100 μm, 200 μm, 500 μm, 1 millimeter (mm), or greater. Alternatively or additionally, one or more conditions can be applied to obtain film thicknesses of up to 10 nanometers (nm), 20 nm, 50 nm, 100 nm, 200 nm, 500 nm, 1 micrometer (μm), 2 μm, 5 μm, 10 μm, 20 μm, 50 μm, 100 μm, 200 μm, 500 μm, 1 millimeter (mm), or smaller. One or more conditions can be applied to obtain film thicknesses within the range defined by any two of the foregoing values. Film thickness can be measured or monitored using various techniques, such as thin-film spectroscopy with a thin-film spectrometer (e.g., fiber optic spectrometer). In some cases, surfactants can be added to the solution or to the surface to promote uniform coating or improve sample loading efficiency. Alternatively or in combination, mechanical, electrical, physical, or other mechanisms can be used to adjust the thickness of the solution. For example, the solution can be dispensed onto a substrate and subsequently leveled using, for example, a physical scraper (e.g., a blade) to obtain the desired uniform thickness across the substrate.

[0226] Reagents can be dispensed to multiple sites on a substrate through various mechanisms, and / or multiple reagents can be dispensed to a single site on the substrate. The reagent dispensing mechanisms disclosed herein can be applied to sample dispensing. For example, reagents may include samples. As described herein with reference to reagents or samples, the term "loaded onto a substrate" can refer to dispensing a reagent or sample onto the surface of a substrate according to any reagent dispensing mechanism described herein.

[0227] In some cases, dispensing can be achieved through relative movement between the substrate and the dispenser (e.g., a nozzle). For example, a reagent may be dispensed onto the substrate at a first location and then travel to a second location different from the first location due to forces (e.g., centrifugal, centripetal, inertial, etc.) generated by the movement of the substrate (e.g., rotational movement of the substrate, linear movement of the substrate, combinations thereof). In another example, the reagent may be dispensed to a reference location, and the substrate may be moved relative to the reference location such that the reagent is dispensed to multiple locations on the substrate. In yet another example, the dispenser may be moved relative to the substrate to dispense the reagent at different locations, such as before, during, or after dispensing. In one example, the reagent is 'coated' onto the substrate by moving the dispenser and / or the substrate relative to each other along a desired path on the substrate. Open substrate geometry allows for flexible and controlled dispensing of reagents to desired locations on the substrate. In some cases, dispensing can be achieved without relative movement between the substrate and the dispenser. For example, multiple dispensers may be used to dispense reagents to different locations, and / or to dispense multiple reagents to a single location, or combinations thereof (e.g., to dispense multiple reagents to multiple locations).

[0228] In another example, an external force (e.g., involving pressure differential, physical force, magnetic force, electrical force, etc.), such as wind, a field generating device, or a physical device, can be applied to one or more surfaces of the substrate to direct the reagent to different locations across the substrate. In another example, the method for dispensing the reagent may include vibration. In such examples, the reagent may be distributed or dispensed into a single area or multiple areas on the substrate (or substrate surface). The substrate (or its surface) may then be subjected to vibration, which may cause the reagent to diffuse to different locations across the substrate (or surface). Alternatively or in combination, the method may include dispensing the reagent onto the substrate using mechanical, electrical, physical, or other mechanisms. For example, a solution may be dispensed onto the substrate, and a physical scraper (e.g., a scraper) may be used to disperse the dispensed material or reagent to different locations and / or across the substrate to achieve the desired thickness or uniformity. Advantageously, this flexible dispensing can be achieved without contaminating the reagent.

[0229] In some cases, where a volume of reagent is dispensed onto the substrate at a first location and then travels to a second location different from the first location, the volume of reagent can travel along one or more paths, such that the one or more paths are coated with reagent. In some cases, such one or more paths can cover a desired surface area of ​​the substrate (e.g., the entire surface area, a portion of the surface area, etc.). In some cases, two or more reagents can be mixed on the surface of the substrate, such as by dispensing at the same location and / or by guiding the first reagent to encounter the other reagent. In some cases, the reagent mixture formed on the substrate can be homogeneous or substantially homogeneous. The reagent mixture can be formed at the first location on the substrate before dispersing the reagent mixture to other locations on the substrate (e.g., at locations where it encounters other reagents or analytes).

[0230] In some embodiments, one or more solutions can be delivered directly to the reaction site without significantly displacing the solutions from the delivery site. Methods for delivering solutions directly to the reaction site may include aerosol delivery of the solution, application of the solution using an applicator, curtain coating of the solution, slit coating, dispensing the solution from a translational dispensing probe, dispensing the solution from a dispensing probe array, immersing a substrate in the solution, or contacting the substrate with a sheet comprising the solution.

[0231] Aerosol delivery can include delivering the solution to the substrate in aerosol form by guiding the solution to the substrate using a pressure nozzle or an ultrasonic nozzle. Applying the solution using an applicator can include contacting the substrate with the applicator containing the solution and translating the applicator relative to the substrate. For example, applying the solution using an applicator can include coating the substrate. The solution can be applied in a patterned manner by translating the applicator, rotating the substrate, translating the substrate, or a combination thereof. Curtain coating can include dispensing the solution from a dispensing probe in a continuous flow (e.g., a curtain or plate) to the substrate and translating the dispensing probe relative to the substrate. The solution can be curtain coated in a patterned manner by translating the dispensing probe, rotating the substrate, translating the substrate, or a combination thereof. Slit coating can include dispensing the solution from a dispensing probe positioned near the substrate, such that the solution forms a meniscus between the substrate and the dispensing probe, and translating the dispensing probe relative to the substrate. The solution can be slit coated in a patterned manner by translating the dispensing probe, rotating the substrate, translating the substrate, or a combination thereof. Dispensing solution from a translational dispensing probe can include translating the dispensing probe relative to the substrate in a patterned manner (e.g., a spiral pattern, a ring pattern, a linear pattern, a stripe pattern, a crosshair pattern, or a diagonal pattern). Dispensing solution from an array of dispensing probes can include dispensing solution from an array of nozzles (e.g., spray heads) positioned above the substrate, such that the solution is dispensed substantially simultaneously across a region of the substrate. Immersing the substrate in the solution can include immersing the substrate in a reservoir containing the solution. In some embodiments, the reservoir can be a shallow reservoir to reduce the volume of solution required to coat the substrate. Contacting the substrate with a sheet containing the solution can include contacting the substrate with a sheet of material permeated with the solution (e.g., a porous sheet or a fibrous sheet). The solution can be transferred onto the substrate. In some embodiments, the sheet of material can be a disposable sheet. In some embodiments, the sheet of material can be a reusable sheet. In some embodiments, it can be used... Figure 5B The method shown in the figure dispenses a solution onto a substrate, wherein the solution jet can be dispensed from a nozzle onto a rotating substrate. The nozzle can be radially translated relative to the rotating substrate, thereby dispensing the solution onto the substrate in a helical pattern.

[0232] One or more solutions or reagents may be delivered to a substrate using any of the delivery methods disclosed herein. In some embodiments, two or more solutions or reagents are delivered to the substrate using the same or different delivery methods. In some embodiments, two or more solutions are delivered to the substrate such that, for each region of the substrate in contact with the one or more solutions or reagents, the time between contact with one solution or reagent and contact with a subsequent solution or reagent is substantially the same. In some embodiments, the solutions or reagents may be delivered as a single mixture. In some embodiments, the solutions or reagents may be dispensed in two or more component solutions. For example, each component of the two or more component solutions may be dispensed from a different nozzle. The different nozzles may dispense the two or more component solutions substantially simultaneously into substantially the same region of the substrate, such that a homogeneous solution is formed on the substrate. In some embodiments, the dispensing of each of the two or more components may be time-separated. The dispensing of each component may be performed using the same or different delivery methods. In some embodiments, direct delivery of solutions or reagents may be combined with spin coating.

[0233] The solution can be incubated on the substrate for any desired duration (e.g., minutes, hours, etc.). In some embodiments, the solution can be incubated on the substrate while maintaining a fluid layer on the surface. One or more of the following can be adjusted: chamber temperature, chamber humidity, substrate rotation, or fluid composition, such that the fluid layer remains constant during incubation. In some cases, during incubation, the substrate can be rotated at a rotation frequency not exceeding: 60 rpm, 50 rpm, 40 rpm, 30 rpm, 25 rpm, 20 rpm, 15 rpm, 14 rpm, 13 rpm, 12 rpm, 11 rpm, 10 rpm, 9 rpm, 8 rpm, 7 rpm, 6 rpm, 5 rpm, 4 rpm, 3 rpm, 2 rpm, 1 rpm, or lower. In some cases, during incubation, the substrate can be rotated at a rotation frequency of approximately 5 rpm.

[0234] The substrate or its surface may include other features that facilitate the retention of a solution or reagent on the substrate, or contribute to the uniformity of solution or reagent thickness on the substrate. In some cases, the surface may include raised edges (e.g., margins) that can be used to retain the solution on the surface. The surface may include margins near the outer edge of the surface, thereby reducing the amount of solution flowing over the outer edge.

[0235] The dispensed solution may include any sample or analyte disclosed herein. The dispensed solution may include any reagent disclosed herein. In some cases, the solution may be a reaction mixture comprising various components. In some cases, the solution may be a component of the final mixture (e.g., to be mixed after dispensing). In non-limiting examples, the solution may include samples, analytes, carriers, beads, probes, nucleotides, oligonucleotides, labels (e.g., dyes), terminators (e.g., blocking groups), other components that aid, accelerate, or decelerate the reaction (e.g., enzymes, catalysts, buffers, saline solutions, chelating agents, reducing agents, other reagents, etc.), washing solutions, cleaving agents, combinations thereof, deionized water, and other reagents and buffers.

[0236] In some cases, the sample can be diluted to control the approximate occupancy of individually addressable sites. In some cases, the sample may include beads, as described elsewhere in this document, such as beads comprising nucleic acid colonies bound thereto. In some cases, beads of orders of magnitude of at least about 10, 100, 1000, 10,000, 100,000, 1,000,000, 100,000,000, 1,000,000,000, 10,000,000,000, 10,000,000,000, 100,000,000,000 or more may be loaded onto the substrate, such as to anchor them to as many individually addressable sites as possible. Alternatively or additionally, beads in quantities ranging from approximately 100,000,000,000, 10,000,000,000, 1,000,000,000, 100,000,000, 10,000,000, 1,000,000, 100,000, 10,000, 10,000, 100, 100, or 10 can be mounted on a substrate, such as to be fixed to as many individually addressable locations as possible. In some cases, beads can be distinguished from each other using characteristics such as color, reflectivity, anisotropy, brightness, fluorescence, etc. In some cases, as described elsewhere herein, different beads may include different tags (e.g., nucleic acid sequences) coupled to them. For example, beads may include oligonucleotide molecules that include tags for identifying beads in a plurality of beads. Figure 4Images of a portion of the substrate surface are shown after a sample comprising beads has been loaded onto a substrate patterned with a substantially hexagonal lattice and having individually addressable locations. The right inset shows a reduced-size image of a portion of the surface, and the left inset shows a magnified image of a segment of that portion of the surface. In some cases, after sample loading, "bead occupancy" can generally refer to the number of individually addressable locations of a certain type comprising at least one bead out of the total number of individually addressable locations of the same type. Bead "landing efficiency" can generally refer to the number of beads bonded to the surface out of the total number of beads distributed on the surface.

[0237] In some cases, it can be based on Figures 5A-5B One or more systems and methods are shown to dispense beads onto a substrate. For example... Figure 5A As shown, a solution comprising beads can be dispensed from a dispensing probe 501 (e.g., a nozzle) onto a substrate 503 (e.g., a wafer) to form a layer 505. The dispensing probe can be positioned at a height (“Z”) above the substrate. In the illustrated example, the beads are held in layer 505 by electrostatic retention and can be anchored to the substrate at respective individually addressable locations. A group of beads in the solution can each comprise a population of amplified products (e.g., nucleic acid molecules) anchored thereon, the amplified products accumulating a negative charge on the beads, which has an affinity for positive charges. Furthermore, the beads can include negatively charged reagents. The substrate comprises alternating surface chemical compositions with distinguishable positions, wherein a first position type comprises positively charged APTMS that has an affinity for the negative charge of amplified beads (e.g., beads comprising amplified products immobilized thereon, and distinct from negatively charged beads that do not include said amplified products) or other beads comprising a negative charge, and a second position type comprises HMDS that has a lower affinity for and / or repels amplified beads or other beads comprising a negative charge. Within layer 505, beads can successfully settle at a first position of the first position type (as in 507). In the illustrated example, the position size is 1 micrometer, the spacing between different positions of the same position type (e.g., the first position type) is 2 micrometers, and the layer depth is 15 micrometers. Figure 5B This demonstrates reagents (e.g., beads) dispensed along pathways on the open surface of a substrate. Figure 5B As shown, the reagent solution can be dispensed from a dispensing probe (e.g., a nozzle). The reagent can be dispensed onto the surface in any desired pattern or path. This can be achieved by moving one or both of the substrate and the dispensing nozzle. The substrate and the dispensing probe can be moved relative to each other in any configuration to achieve any pattern (e.g., linear pattern, substantially spiral pattern, etc.).

[0238] In some cases, after the solution has been contacted with the substrate, a subset or all of the solution can be recovered. Recovery may include collecting, filtering, and reusing a subset or all of the solution. Filtration may be molecular filtration.

[0239] Detection

[0240] An optical system including a detector can be configured to detect one or more signals from a detection region on a substrate before, during, or after reagent dispensing to generate an output. During a single detection event, signals from multiple individually addressable locations can be detected. Signals from the same individually addressable location can be detected in multiple scenarios.

[0241] Once a reaction occurs between the probe and the analyte in solution, a detectable signal, such as an optical signal (e.g., a fluorescence signal), can be generated. For example, the signal may originate from the probe and / or the analyte. The detectable signal can indicate the reaction or interaction between the probe and the analyte. The detectable signal can be a non-optical signal. For example, the detectable signal can be an electronic signal. The detectable signal can be detected by a detector (e.g., one or more sensors). For example, in optical detection schemes described elsewhere herein, an optical signal can be detected by one or more optical detectors. The signal can be detected during substrate rotation. The signal can be detected after rotation is terminated. The signal can be detected when the analyte comes into contact with the solution fluid. The signal can be detected after washing the solution. In some cases, the signal may be muted after detection, such as by cutting a mark from the probe and / or the analyte, and / or altering the probe and / or the analyte. Such cutting and / or alteration may be influenced by one or more stimuli, such as exposure to chemicals, enzymes, light (e.g., ultraviolet light), or temperature changes (e.g., heat). In some cases, a signal may become undetectable by disabling or changing the mode of one or more sensors (e.g., the detection wavelength) or terminating or reversing the excitation of the signal. In some cases, signal detection may include capturing an image or generating a digital output (e.g., between different images).

[0242] The following operations can be repeated any number of times: (i) directing a solution to the substrate, and (ii) detecting a signal indicating the reaction between one or more probes in the solution and an analyte immobilized on the substrate. Such operations can be repeated iteratively. For example, the same analyte immobilized at a given location in an array can interact with multiple solutions in multiple repeated cycles. For each iteration, additional signals detected can provide incremental or final data about the analyte during processing. For example, in the case where the analyte is a nucleic acid molecule and the processing is sequencing, additional signals detected in each iteration can indicate bases in the nucleic acid sequence of the nucleic acid molecule. In some cases, multiple solutions can be provided to the substrate without intervention in detection events. In some cases, multiple detection events can be performed after a single solution flow. In some cases, wash solutions, cleavage solutions (e.g., including cleavage agents), and / or other solutions can be directed to the substrate between each operation, between each cycle, or within a certain number of cycles.

[0243] Optical systems can be configured to perform continuous area scanning of a substrate during rotational motion. As used herein, the term "continuous area scanning (CAS)" generally refers to a method of imaging a relatively moving object by repeatedly, electronically or computationally advancing (timing or triggering) an array of sensors at a speed that compensates for the object's motion in the detection plane (focal plane). CAS can produce images with a scan dimension larger than the field of the optical system. TDI scanning can be an example of CAS, where a clock needs to shift the photoelectric charge on the area sensor during signal integration. For a TDI sensor, the charge can be shifted one row at each clock step, with the last row being read out and digitized. Other modalities can achieve similar functionality through high-speed area imaging and the overlay of digital data to synthesize continuous or progressively continuous scans.

[0244] An optical system may include one or more sensors. Sensors can detect an image optically projected from a sample. An optical system may include one or more optical elements. Optical elements may be, for example, lenses, prisms, mirrors, waveplates, filters, attenuators, gratings, apertures, beam splitters, diffusers, polarizers, depolarizers, retroreflectors, spatial light modulators, or any other optical element. The system may include any number of sensors. In some cases, the sensor is any detector as described herein. In some instances, the sensor may include an image sensor, a CCD camera, a CMOS camera, a TDI camera (e.g., a TDI line scan camera), a pseudo-TDI fast frame rate sensor, or a CMOS TDI or hybrid camera. The optical system may further include any light source. In some cases, where multiple sensors are present, different sensors may image the same or different regions of the rotating substrate, and in some cases, simultaneously. Each of the multiple sensors may be timed at a rate suitable for imaging a region of the rotating substrate by the sensor, the rate may be based on the distance of the region from the center of the rotating substrate or the tangential velocity of the region. In some cases, multiple scanning heads can operate in parallel along different imaging paths (e.g., staggered spiral scanning, nested spiral scanning, staggered circular scanning, nested circular scanning). A scanning head may include one or more detector elements, such as a camera (e.g., a TDI line scan camera), an illumination source (e.g., as described herein), and one or more optical elements (e.g., as described herein).

[0245] The system may further include a controller. The controller may be operatively coupled to the one or more sensors. The controller may be programmed to process optical signals from each zone of the rotating substrate. For example, the controller may be programmed to process optical signals from each zone at independent clocks during rotational motion. Independent timing may be based at least in part on the distance of each zone's projection onto the axis and / or the tangential velocity of the rotational motion. Independent timing may be based at least in part on the angular velocity of the rotational motion. While a single controller has been described, multiple controllers may be configured to perform the operations described herein individually or collectively.

[0246] In some cases, the optical system may include an immersion objective. The immersion objective may contact an immersion fluid that is in contact with an open substrate. The immersion fluid may include any suitable immersion medium for imaging (e.g., water, aqueous solution, organic solution). In some cases, a housing may partially or completely surround the sample-facing end of the optical imaging objective. The housing may be configured to contain fluid. The housing may not be in contact with the substrate; for example, the gap between the housing and the substrate may be filled by the fluid contained within the housing (e.g., the housing may retain the fluid through surface tension). In some cases, an electric field may be used to modulate the hydrophobicity of one or more surfaces of the housing to retain at least a portion of the fluid in contact with the immersion objective and the open substrate.

[0247] Figure 6 A computerized system 600 for sequencing nucleic acid molecules is illustrated. The system may include a substrate 610, such as any substrate described herein. The system may further include a fluid flow unit 611. The fluid flow unit may include any elements associated with the fluid flow described herein. The fluid flow unit may be configured to guide a solution comprising the various nucleotides described herein to a substrate array before or during rotation of the substrate. The fluid flow unit may be configured to guide a wash solution, comprising the wash solution described herein, to the substrate array before or during rotation of the substrate. In some cases, the fluid flow unit may include a pump, compressor, and / or actuator to guide fluid from a first position to a second position. The fluid flow unit may be configured to guide any solution to the substrate 610. The fluid flow system may be configured to collect any solution from the substrate 610. The system may further include a detector 670, such as any detector described herein. The detector may be in sensing communication with the substrate surface.

[0248] The system may further include one or more processors 620. The one or more processors may be programmed individually or collectively to implement any of the methods described herein. For example, the one or more processors may be programmed individually or collectively to implement any or all operations of the methods disclosed herein. Specifically, the one or more processors may be programmed individually or collectively to perform the following: (i) guiding a fluid flow unit to guide a solution comprising the plurality of nucleotides through an array during or prior to substrate rotation; (ii) subjecting the nucleic acid molecule to a primer extension reaction under conditions sufficient to incorporate at least one nucleotide of the plurality of nucleotides into a growth chain complementary to the nucleic acid molecule; and (iii) using a detector to detect a signal indicating the incorporation of the at least one nucleotide, thereby sequencing the nucleic acid molecule.

[0249] High throughput

[0250] The open substrate system disclosed herein may include a barrier system configured to maintain a fluid barrier between a sample handling environment and an external environment. The barrier system is further described in detail in U.S. Patent Publication No. 20210354126A1, which is incorporated herein by reference in its entirety. The sample environment system may include a sample handling environment defined by a chamber and a cover plate, wherein the cover plate does not contact the chamber. The gap between the cover plate and the chamber may include a fluid barrier. The fluid barrier may include fluid (e.g., air) from the sample handling environment and / or the external environment, and may have a pressure lower than that of the sample environment, the external environment, or both. The fluid in the fluid barrier may be in coherent motion or in a uniform motion.

[0251] The sample processing environment may include a substrate, such as any substrate described elsewhere herein. As described elsewhere herein, any operations performed on or with the substrate can be performed within the sample processing environment while maintaining a fluid barrier. For example, the substrate may be rotated within the sample processing environment during various operations. In another example, when the substrate is in the sample processing environment, fluid can be directed to the substrate via a fluid processor (e.g., a nozzle) that penetrates the cover plate into the sample processing environment. In another example, when the substrate is in the sample processing environment, a detector can image the substrate via a detector that penetrates the cover plate into the sample processing environment. Advantageously, the fluid barrier helps maintain the temperature and / or relative humidity or its range within the sample processing environment during various processing operations.

[0252] The system or any component thereof described herein may be environmentally controlled. For example, the system may be maintained at a specified temperature or humidity. For operation, the system (or any component thereof) may be maintained at a temperature of at least 20 degrees Celsius (°C), 25°C, 30°C, 35°C, 40°C, 45°C, 50°C, 55°C, 60°C, 65°C, 70°C, 75°C, 80°C, 85°C, 90°C, 95°C, 100°C, or higher. Alternatively or additionally, for operation, the system (or any component thereof) may be maintained at a temperature of at most 100°C, 95°C, 90°C, 85°C, 80°C, 75°C, 70°C, 65°C, 60°C, 55°C, 50°C, 45°C, 40°C, 35°C, 30°C, 25°C, 20°C, or lower. Different components of the system may be maintained at different temperatures or different temperature ranges, as described herein. Components of the system may be set at temperatures above the dew point to prevent condensation. The system components can be set at temperatures below the dew point to collect condensation. In one example, the sample processing environment, including the substrate as described elsewhere herein, can be environmentally controlled from the external environment. The sample processing environment can be further divided into separate zones maintained at different local temperatures and / or relative humidityes, such as a first zone in contact with or near the surface of the substrate, and a second zone in contact with or near the top portion (e.g., a cover) of the sample processing environment. For example, the local environment of the first zone can be maintained at a first set of temperatures and a first set of humidityes configured to prevent or minimize the evaporation of one or more reagents on the surface of the substrate, and the local environment of the second zone can be maintained at a second set of temperatures and a second set of humidityes configured to enhance or limit condensation. The first set of temperatures can be the lowest temperature within the sample processing environment, and the second set of temperatures can be the highest temperature within the sample processing environment.

[0253] In some cases, different environmental conditions in different zones can be achieved by controlling the temperature of the outer shell. In some cases, different environmental conditions in different zones can be achieved by controlling the temperature of selected portions or the entire container. In some cases, different environmental conditions in different zones can be achieved by controlling the temperature of selected portions or the entire substrate. In some cases, different environmental conditions in different zones can be achieved by controlling the temperature of the reagent dispensed onto the substrate. Any combination thereof can be used to control the environmental conditions in different zones. Heat transfer can be achieved by any method, including, for example, conduction, convection, and radiation.

[0254] While the examples described herein provide examples of relative rotational motion of the substrate and / or detector system, the substrate and / or detector system may alternatively or additionally experience relative non-rotational motion, such as relative linear motion, relative nonlinear motion (e.g., bending, bowing, angular, etc.), and any other type of relative motion.

[0255] In some cases, the open substrate is retained in the same or nearly the same physical location during the processing of the analyte and the subsequent detection of the signal associated with the processed analyte.

[0256] In some cases, different operations are performed on or using an open substrate at different stations. The different stations can be located in different physical locations. For example, a first station can be located above, below, near, or opposite a second station. In some cases, the different stations can be housed within an integrated enclosure. Alternatively, the different stations can be housed separately. In some cases, the different stations may be separated by barriers, such as retractable barriers (e.g., sliding doors). One or more different stations or portions thereof in the system may be affected by different physical conditions, such as different temperatures, pressures, or atmospheric compositions. In an example, a processing station may include a first atmospheric environment containing a first set of conditions and a second atmospheric environment containing a second set of conditions. Barrier systems can be used to maintain the different physical conditions of one or more different stations or portions thereof in the system, as described elsewhere herein.

[0257] An open substrate can be transitioned between different stations by transporting a sample handling environment, including an open substrate (such as the sample handling environment described with respect to a barrier system). One or more mechanical components or mechanisms, such as robotic arms, lifting mechanisms, actuators, tracks, etc., or other mechanisms, can be used to transport the sample handling environment.

[0258] Environmental units (e.g., humidifiers, heaters, heat exchangers, compressors, etc.) can be configured to regulate one or more operating conditions in each station. In some cases, each station can be regulated by an independent environmental unit. In others, a single environmental unit can regulate multiple stations. In still others, multiple environmental units can regulate different stations individually or collaboratively. Environmental units can regulate operating conditions using active or passive methods. For example, heating or cooling elements can be used to control temperature. Humidity can be controlled using humidifiers or dehumidifiers. In some cases, a portion of a particular station, such as the sample handling environment, can be further controlled from other portions of that station. Different portions can have different local temperatures, pressures, and / or humidity levels.

[0259] In one example, reagent delivery and / or dispersion may be performed at a first station with first operating conditions, and the detection process may be performed at a second station with second operating conditions different from the first operating conditions. During the delivery and / or dispersion process, the first station may be located at a first physical location where the open substrate is accessible to the fluid processing unit, and the second station may be located at a second physical location where the open substrate is accessible to the detector system.

[0260] One or more modular sample environment systems (each with its own barrier system) can be used between different stations. In some cases, the systems described herein can be scaled up to include two or more of the same station type. For example, a sequencing system may include multiple processing stations and / or detection stations. Figures 7A-7C System 300 demonstrates the reuse of two modular sample environment systems within a three-station system. Figure 7B In this system, a first chemical station (e.g., 320a) can operate (e.g., dispense reagents, e.g., to incorporate nucleotides for synthesis sequencing) on ​​a first substrate (e.g., 311) in a first sample environment system (e.g., 305a) at least via a first operating unit (e.g., fluid dispenser 309a), while substantially simultaneously, a detection station (e.g., 320b) can operate (e.g., scan) on a second substrate in a second sample environment system (e.g., 305b) at least via a second operating unit (e.g., detector 301), while substantially simultaneously, a second chemical station (e.g., 320c) is idle. The idle station may not be able to operate on the substrate. The idle station (e.g., 320c) can be charged, reloaded, replaced, cleaned, washed (e.g., to rinse reagents), calibrated, reset, kept active (e.g., powered on), and / or otherwise maintained during idle time. After the operating cycle is complete, the sample environment system can be repositioned, such as... Figure 7C As shown, a second substrate in a second sample environment system (e.g., 305b) is repositioned from a detection station (e.g., 320b) to a second chemistry station (e.g., 320c) for operation by the second chemistry workstation (e.g., dispensing reagents, for example, to incorporate nucleotides for synthesis sequencing), and a first substrate in a first sample environment system (e.g., 305a) is repositioned from a first chemistry station (e.g., 320a) to a detection station (e.g., 320b) for operation by the detection station (e.g., scanning). An operation cycle is considered complete when all operations at each active parallel station are completed. During repositioning, different sample environment systems may be physically moved (e.g., along the same or dedicated track, e.g., track 307) to different stations and / or different stations may be physically moved to different sample environment systems. One or more components of a station, such as modular plates 303a, 303b, 303c defining a particular station, may be physically moved to allow a sample environment system to leave, enter, or pass through the station. During substrate processing at the station, the environment of the sample environment zone (e.g., 315) of the sample environment system (e.g., 305a) can be controlled and / or adjusted according to the station's requirements. After the next operating cycle is completed, the sample environment system can be repositioned, such as returning to... Figure 7B The configuration, and this repositioning can be repeated each time an operation cycle is completed (e.g., in...). Figure 7B and 7C Between configurations), until the substrate processing required is completed. In this illustrative repositioning scheme, by providing the detection station with alternating different sample environment systems for each continuous operating cycle, the detection station can remain active for all operating cycles (e.g., there is no idle time not operating on the substrate). Advantageously, the use of the detection station is optimized. Based on different processing or equipment requirements, the operator can choose to run two chemical stations (e.g., 320a, 320c) substantially simultaneously, while the detection station (e.g., 320b) remains idle, such as... Figure 7A As shown.

[0261] Advantageously, different operations within the system can be reused with high flexibility and control. For example, as described herein, one or more processing stations can operate in parallel with one or more detection stations on different substrates in different modular sample environment systems to reduce or eliminate hysteresis between different operation sequences (e.g., chemical reaction first, then detection). Modular sample environment systems can be shifted accordingly between different stations to optimize efficient use of the equipment (e.g., so that the detection station is in operation almost 100% of the time). In some instances, at least one, two, three, four, five, six, seven, eight, nine, or more modules or stations of a sequencing system can be reused. For example, two or more modules can each perform their intended function simultaneously or according to the methods described elsewhere herein. Such instances can include two-station reuse, such as optical and chemical stations as described herein. Another instance can include reuse of three or more stations and processing stages. For example, the method can include staggered chemical stages using shared scanning stations. The scanning station can be a high-speed scanning station. Various sequences and configurations can be used to reuse modules or stations.

[0262] The nucleic acid sequencing system and optical system (or any of its components) described in this article can be combined in various architectures.

[0263] Retain and use both the forward and reverse strands for sequencing.

[0264] This document provides apparatus, systems, methods, compositions, and kits that (i) preserve both strands of a double-stranded template nucleic acid molecule during amplification and / or (ii) are capable of identifying or quantifying the proportion of clusters of amplified molecules derived from the respective two strands. This document provides apparatus, systems, methods, compositions, and kits that allow for the simultaneous sequencing of materials derived from both strands of a double-stranded template nucleic acid molecule. Such apparatus, systems, methods, compositions, and kits may alternatively or in addition to those described above. Figure 1This applies to one or more operations 101-108 outside of the sequencing workflow described in 100. Such devices, systems, methods, compositions, and kits can be used in conjunction with the sample processing systems and methods described herein or components thereof (e.g., substrates, detectors, reagent dispensing, serial scanning, etc.).

[0265] The template nucleic acid molecule input into the sequencing library can be double-stranded. In other cases, the template nucleic acid molecule may exist in double-stranded form and / or be converted to double-stranded form at various time points during the sequencing process (such as during one or more operations of library preparation, amplification, and / or enrichment).

[0266] Prior to sequencing, the template nucleic acid molecule can undergo amplification, such as generating colonies of multiple copies of the template nucleic acid molecule in clusters (e.g., on vectors such as beads or surface sites). Amplification can produce multiple different molecules, each comprising a copy of the template nucleic acid molecule (or its strand), or single molecules (e.g., tandem molecules) comprising multiple copies of the template nucleic acid molecule (or its strand)—either type of amplification product may be referred to herein as a colony or cluster. However, some amplification protocols can amplify only one of the two strands of a double-stranded template nucleic acid molecule, thereby discarding information (e.g., sequence information) from the other strand. While in some cases the two strands of a double-stranded template nucleic acid molecule may be perfectly inverse complementary sequences, in which case discarding one strand does not result in loss of information, in other cases the two strands may contain sites of base mismatch, in which case discarding one strand results in the loss of valuable information from the template nucleic acid molecule, including potential surrogate bases at the site in the sequence and the fact that the site of the base mismatch was originally present. For example, PCR-free DNA may include a high percentage of damaged bases that carry base mismatches between the two strands.

[0267] Sequencing errors can be quantified as the error rate out of the total number of sequenced bases. For example, if the human genome (approximately 3.4 billion bases) is sequenced to a depth of 30x in whole-genome sequencing, and the sequencing error rate is 1e-5, this means that 3.4e9 x 30 x 1e-5 = approximately 1e6, or 1 million bases, are incorrectly sequenced. The method presented in this paper can significantly reduce the sequencing error rate, such as by an order of magnitude.

[0268] Figure 24This document illustrates an example of how preserving the sequencing information from both strands of a template nucleic acid molecule can allow for the differentiation between genuine mutations and artificial (pseudo) mutations in a sample. Excluding artificial mutations can reduce sequencing error rates. Single nucleotide variants (SNVs) in a sample can be identified by sequencing the sample to produce sequencing reads, aligning the sequencing reads to a reference sequence, and comparing the bases identified at each locus in the sequencing reads with the bases in the reference sequence. However, as described elsewhere in this document, sample processing can introduce non-sample-inherent artificial errors (e.g., artificial base mismatches) that may be incorrectly identified as genuine SNVs after sequencing, thereby increasing the sequencing error rate and reducing overall accuracy. Such artificial errors typically introduce only one of the two strands in the sample. One example is the spontaneous (or stimulated) deamination of C residues to U residues. This deamination can be 'repaired' in living cells by various enzymes (e.g., by uracil-DNA glycosylase, AP endonuclease, polymerase, etc.) to remove uracil and fill the gap with the correct base (C). However, outside of cells, such as in the laboratory, the deamination site will continue to be sequenced downstream if it is not otherwise interpreted or corrected.

[0269] exist Figure 24In the diagram, a solid circle on a line represents a base site in the strand where a real mutation (e.g., a C->T mutation) has occurred (compared to the "reference"), and a hollow circle on a line represents a base site in the strand where an artificial mutation (e.g., a C->T mutation) has occurred. In the illustration, the post-library preparation sample molecule includes two sites: a real mutation occurring in both strands (at the corresponding base site), and an artificial mutation occurring in only one strand. For clarity, solid and dashed lines each represent nucleic acid strands that are reverse complementary sequences to each other (except for the artificial base mismatch site). The post-library sample can be processed for sequencing, for example, undergoing amplification. As shown in the left-hand inset after processing, when information from only one strand (e.g., dashed lines) is retained for sequencing, as on a bead, both the real mutation and the artificial base mismatch site may be identified as variants in the reference, making it impossible to distinguish between the real mutation and the artificial base mismatch. In contrast, as shown in the right-hand inset after processing, when information from both strands (e.g., dashed lines, solid lines) is retained for sequencing, as on a bead, it is possible to distinguish between real mutations and artificial base mismatches by recognizing: (1) the two strand derivatives being sequenced, (2) confirming base identification in the sequencing read when there is relatively high consistency in base identification for variants at the locus (e.g., indicated by high confidence or sequencing quality score, etc.), and identifying the locus as a site of a real mutation when compared to a reference, and (3) rejecting base identification for variants in the sequencing read when there is relatively high inconsistency in base identification for variants at the locus (e.g., indicated by reduced or diluted sequencing signal, low confidence or sequencing quality score, etc.), and therefore the locus cannot be identified as a site of a real mutation when compared to a reference. In some cases, it is recognized that (1) the two-strand derivatives being sequenced can be achieved by using a processing (e.g., amplification) scheme that preserves or guarantees the preservation of information about the two strands in the sequenceable product, and / or by labeling each strand type (e.g., positive and negative strands) with strand identification elements (e.g., sequences) and identifying the strand identification elements before, during or after sequencing to confirm the presence of information about the two-strand derivatives.

[0270] Therefore, preserving information from both strands of the template nucleic acid molecule during amplification can be beneficial. Such preservation can be particularly advantageous in detecting or excluding single nucleotide variants (SNVs) and / or single nucleotide polymorphisms (SNPs) and in improving the error rate of SNV and / or SNP detection. SNPs are genetic variations in a subject's DNA that are particularly prone to false detection. When information from both strands of the double-stranded template nucleic acid molecule is preserved for amplification, it is important that a significant amount of information from both strands is actually amplified and represented in the amplified cluster, as if only 0.01% of the material originated from one strand and 99.99% from the other. Any signal detected from the cluster can then be effectively attributed to only the other strand.

[0271] Chain identification mismatch connector for forward-reverse amplified chain detection

[0272] Figure 9A This demonstrates a first-stage workflow for extension and amplification on a vector, where one template strand is not amplified and therefore its information is not retained in the downstream product. Vector 901 and template nucleic acid molecule 903 can be provided. Vector 901 may include multiple surface primers, such as 902 (for clarity, ...). Figure 9A (Only one is shown in the image), the plurality of surface primers 902 may be identical and / or include a common primer sequence. The template nucleic acid molecule 903 may include a first adaptor 903a attached to one end and a second adaptor 903b attached to the other end. The first adaptor may be a partially double-stranded adaptor, which includes a protruding end 904 configured to bind to the surface primer 902. One or more strands of the first adaptor may include one or more cleavable portions (denoted as "U"), such as uracil residues. Figure 9A In the illustrated example, the chain with the protruding end includes a cleavable portion. The second linker 903b may be a double-stranded linker. The second linker may include a trapping portion 905, such as biotin, and one or more cleavable portions 906 (denoted as "U"), such as uracil residues, said trapping portion being disposed at 5' of the cleavable portion. Figure 9AIn the example shown, the overhang of the first adaptor and the capture portion of the second adaptor are linked to different strands of the template insert of the template nucleic acid molecule 903. The template nucleic acid molecule 903 can be linked to the vector 901 by annealing the overhang 904 with the surface primer 902 and connecting it to the notch (951) in the top strand 907a. Pre-enrichment can be performed before or after cleavage, wherein the capture portion is used to separate the vector-template assembly from the negative vector (not bound to the template). In the example, the capture portion is biotin, and streptavidin-magnetic beads are used to capture biotin. The U-based cleavage site (952) can then be cleaved using a mixture of uracil-DNA glycosylase (UDG) and endonuclease VIII or a mixture of USER enzymes. UDG catalyzes the hydrolysis of the N-glycosidic bond of deoxyuridine to release uracil, and endonuclease VIII cleaves the DNA phosphodiester backbone at AP, thereby forming a 1-nucleotide DNA gap with 5' and 3' phosphate ends. In other words, using the UDG / Endo VIII enzyme combination produces a blocking 3' end of the bottom strand 907b of the template. Before, during, or after the initiation of the extension and / or amplification reaction (953), the blocked bottom strand 907b can be separated from the top strand 907a, which is covalently bound to the vector 901 (e.g., the surface of a bead or substrate). Primer 910 can anneal with the top strand 907a, and the primer extension reaction can be initiated using the top strand 907a as a template. The template and / or the copied strand can then be extended and / or amplified, such that multiple surface primers on the vector 901 are extended to copy the top strand 907a. The extension and / or amplification reaction can be performed in partitions (e.g., wells, emulsions) or in batches. Figure 9A In the workflow, one of the template strands fails to extend due to its blocked end, and therefore the resulting amplified vector includes a derivative (e.g., a copy) of only one strand (top strand 907a). This implies certain errors (such as incorrect SNPs, e.g., Figures 9A-9BArtificial SNPs (SNPs) at G / T loci introduced during deamination are shown to be retained (or their source lost) during the extension and / or amplification phases because the vector retains information from only the first strand (e.g., only 'G' at the G / T locus). These errors can be sequencing or processing errors that contaminate the original template sequence (e.g., from the sample) (e.g., chimeric byproducts, etc.). Errors can be base mismatch errors, such as SNPs, indels, or artificial base mismatch errors. It should be understood that alternatively or additionally, alternative enzymes performing the same or similar functions as described herein can be used (e.g., any enzyme that cleaves the DNA phosphodiester backbone at AP to produce a 1-nucleotide DNA gap with 5' and 3' phosphate ends). In another instance, the cleavable portion may include ribonucleotides, and the enzyme may include ribonuclease HII (RNase HII), wherein a polymerase is used to fill the downstream portion during cleavage.

[0273] Figure 9B This demonstrates a second workflow for extension and amplification on a vector, where both template strands are amplified. Vector 901 and template nucleic acid molecule 903 can be provided. Vector 901 may include multiple surface primers, such as 902 (for clarity, ...). Figure 9B (Only one is shown in the image), the plurality of surface primers 902 may be identical and / or include a common primer sequence. The template nucleic acid molecule 903 may include a first adaptor 903a attached to one end and a second adaptor 903b attached to the other end. The first adaptor may be a partially double-stranded adaptor, which includes a protruding end 904 configured to bind to the surface primer 902. Figure 9A Unlike other connectors, the first connector 903a may not include any cleavable portion. The second connector 903b may be a double-stranded connector. The second connector may include a capture portion 905, such as biotin, and one or more cleavable portions 906 (denoted as "U"), such as uracil residues, said capture portion being disposed at 5' of the cleavable portion. Figure 9BIn the illustrated example, the protruding end of the first adaptor and the capture portion of the second adaptor are connected to different strands of the template insert of the template nucleic acid molecule 903. The template nucleic acid molecule 903 can be connected to the vector 901 (e.g., the surface of a bead or substrate) by annealing the protruding end 904 with the surface primer 902 and connecting it to a notch (951) in the top strand 907a. Pre-enrichment (952) can be performed before or after dicing, where the capture portion is used to separate the vector-template assembly from the negative vector (not bound to the template). In this example, the capture portion is biotin, and streptavidin-magnetic beads are used to capture the biotin. The U-based dicing site can then be diced using a mixture of uracil-DNA glycosylase (UDG) and endonuclease VIII or a mixture of USER enzymes. Before, during, or after the extension and / or amplification reaction (953), the bottom chain 907b can be separated from the top chain 907a, which is covalently bound to the carrier 901, and annealed to another surface primer on the carrier 901 (e.g., Figure 9B (As shown in the last small figure). Primer 910 can be annealed with the top strand 907a, and primer extension reactions can be initiated using the top strand 907a as a template. Separately, the bottom strand 907b can be used as a template to extend another surface primer annealed with said bottom strand. The template strand and / or copies thereof can then be extended and / or amplified, such that multiple surface primers on vector 901 are extended to copy the top strand 907a and the bottom strand 907b. Extension and / or amplification reactions can be performed in partitions (e.g., wells, emulsions) or in batches. Figure 9B In the workflow, with Figure 9A The workflow differs; both template strands are extended, resulting in amplified vectors comprising derivatives (e.g., copies) of both strands (top strand 907a, bottom strand 907b). This implies that certain sequencing information, such as artificial base mismatches (e.g., ...), may be lost. Figures 9A-9B An artificial SNP of the G / T locus introduced during the deamination reaction is shown to be retained on the vector via a first-strand copy (e.g., 'G' of the G / T locus) and a second-strand copy (e.g., 'T' of the G / T locus). It should be understood that alternatively or additionally, alternative enzymes performing the same or similar functions as described herein (e.g., any enzyme that cleaves the DNA phosphodiester backbone at the AP site by hydrolysis, leaving a 1-nucleotide gap with 3'-hydroxyl and 5'-deoxyribose phosphate (dRP) ends) can be used. In another example, the cleavable portion may comprise a ribonucleotide, and the enzyme may comprise an RNase HII, wherein a polymerase is used to fill the downstream portion during cleavage.

[0274] Figure 9C This demonstrates a third workflow for extension and amplification on a vector, where the adaptor of the template nucleic acid molecule and... Figure 9A The linkers in the workflow are the same, but both template strands are amplified. Vector 901 and template nucleic acid molecule 903 can be linked together, and regarding... Figure 9A Pre-enrichment is performed as described. A mixture of uracil-DNA glycosylase (UDG) and apurine / pyrimidine endonuclease 1 (APE1) can be used to cleave U-based cleavage sites (961). UDG catalyzes the hydrolysis of the N-glycosidic bond of deoxyuridine to release uracil, and APE1 cleaves the DNA phosphodiester backbone at the AP site via hydrolysis, leaving a 1-nucleotide gap with a 3'-hydroxyl and 5'-deoxyribose phosphate (dRP) terminus. That is, using the UDG / APE1 enzyme combination produces an extendable 3' end of the template's bottom strand 908b. Taq or similar polymerases and dNTPs can be added to extend the 3'-hydroxyl site (962). The bottom strand 908b can be extended using the top strand 908a bound to a vector 901 (e.g., the surface of a bead or substrate) as a template to produce an extended bottom strand including a sequence complementary to the surface primer 902. Then, the extended bottom chain 908b can be separated from the top chain 908a and annealed with another surface primer on vector 901. Primer 910 can be annealed with the top chain 908a, and the primer extension reaction can be initiated using the top chain 908a as a template (963). Separately, the extended bottom chain can be used as a template to extend another surface primer annealed with the extended bottom chain (963). The template chain and / or copies thereof can be extended and / or amplified such that multiple surface primers on vector 901 are extended to copy the top chain 908a and the bottom chain 908b. The extension and / or amplification reactions can be performed in partitions (e.g., wells, emulsions) or in batches. In some cases, only the first-stage extension reaction (e.g., extending the bottom chain 908b to produce the extended bottom chain) can be performed outside the partition, and the next-stage extension reaction can be performed within the partition such that within the partition, the vector, including one copy of the top chain and one copy of the bottom chain, undergoes amplification. Figure 9C In the workflow, with Figure 9A The workflow differs; both template strands are extended, resulting in amplified vectors comprising derivatives (e.g., copies) of both strands (top strand 908a, bottom strand 908b). This implies that certain sequencing information, such as artificial base mismatches (e.g., ...), may be lost. Figures 9A-9CAn artificial SNP of the G / T locus introduced during the deamination reaction is shown to be retained on the vector via a first-strand copy (e.g., 'G' of the G / T locus) and a second-strand copy (e.g., 'T' of the G / T locus). It should be understood that alternatively or additionally, alternative enzymes performing the same or similar functions as described herein (e.g., any enzyme that cleaves the DNA phosphodiester backbone at the AP site by hydrolysis, leaving a 1-nucleotide gap with 3'-hydroxyl and 5'-deoxyribose phosphate (dRP) ends) can be used. In another example, the cleavable portion may comprise a ribonucleotide, and the enzyme may comprise an RNase HII, wherein a polymerase is used to fill the downstream portion during cleavage.

[0275] The amplified vector obtained from the workflow, wherein the amplified strands linked to the vector can be derived from two strands, which can be pseudo-polyclonal beads, wherein the first set of strands covalently linked to the vector is a copy of the top strand 908a, and the second set of strands covalently linked to the vector is a copy of the bottom strand 908b. The ratio of the top strand and / or top strand copy in all extended strands on the vector can be at least about 0.001, 0.01, 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 0.99, or 0.999. The ratio of the top strand and / or top strand copy in all extended strands on the vector can be at most about 0.001, 0.01, 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 0.99, or 0.999. The ratio of bottom chain and / or bottom chain copies among all extended chains on the vector can be at least about 0.001, 0.01, 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 0.99, or 0.999. The ratio of bottom chain and / or bottom chain copies among all extended chains on the vector can be at most about 0.001, 0.01, 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 0.99, or 0.999. In some cases, the ratio of top chain copies to bottom chain copies can be approximately 5:5.

[0276] Figures 10A-10B This demonstrates a fourth workflow for extension and amplification on a vector, where both template strands are amplified using an adaptor that includes a mismatched portion. Vector 1070 and template nucleic acid molecule 1007 are available. Vector 1070 may include multiple surface primers, such as 1072 (for clarity, ...). Figure 10A(Only two are shown in the figure), the plurality of surface primers 1072 may be the same and / or include a common primer sequence. Template nucleic acid molecule 1007 may include insert molecule 1003, a first adaptor 1001 (1051) attached at one end of insert molecule and a second adaptor 1005 (1051) attached at the other end of insert molecule.

[0277] The first connector 1001 may be a partially double-stranded connector, including a protruding end configured to bind to the surface primer 1072. The first connector 1001 may not include any cleavable portion. In some cases, the first connector 1001 may correspond to... Figure 9B The first connector 903a. The second connector 1005 may be a double-chain connector. The second connector 1005 may include a mismatch portion, wherein the two chains have a sequence mismatch (having a non-complementary sequence) such that they do not anneal in the mismatch portion. In some cases, the second connector may include a mismatch portion between two double-chain portions to form the following structure: (double-chain portion) - (mismatch portion) - (double-chain portion). When the mismatch portion is located between two double-chain portions, the mismatch portion may be referred to as a 'loop' structure. In some cases, the second connector may include a mismatch portion at one end of the connector to form the following structure: (double-chain portion) - (mismatch portion). In some cases, the second connector may include more than one mismatch portion. In some cases, the second connector may include a protruding end (single-chain portion) at one end. See also Figure 14 Non-limiting example configurations of template nucleic acid molecules including a second adaptor are provided. As described herein, at least a partially double-stranded adaptor including a mismatch portion may be referred to herein as a strand recognition adaptor.

[0278] In some cases, the two chains in the mismatch portion may each comprise a homopolymer sequence that is not complementary to the other. For example, the top / bottom chain pair in the mismatch portion may be selected from the following homopolymer sequence pairs: (poly(T) / poly(G); (poly(T) / poly(C); (poly(T) / poly(T); (poly(G) / poly(G); (poly(G) / poly(T); (poly(G) / poly(A)); (poly(C) / poly(C); (poly(C) / poly(T); (poly(C) / poly(A)); (poly(A) / poly(G); (poly(A) / poly(C)); and (poly(A) / poly(A)). In some cases, the two chains in the mismatch portion may each comprise any sequence that is not complementary to the other. In some cases, the two chains in the mismatch portion may have the same length or different lengths (e.g., pentomer / pentomer pairs or pentomer / octamer pairs). In some cases, the mismatch portion may comprise a single base mismatch or a raised loop (e.g., an insertion in one chain). The mismatch sequence in the mismatch portion may act as a chain identification element. The second adaptor may include a capture portion 1075, such as biotin, and one or more cleavable portions (denoted as "U"), such as uracil residues, said capture portion being positioned at the 5' of the cleavable portion. The first and / or second adaptor (e.g., a strand recognition adaptor) may include any other functional sequence, such as a capture sequence, primer sequence, amplification primer sequence, sequencing primer sequence, barcode sequence, sample index sequence, unique molecular identifier (UMI), flow cell adaptor sequence, adaptor sequence, binding sequence for any molecule (e.g., splint, primer, template nucleic acid, capture sequence, etc.), or any other functional sequence or any combination thereof that may be used for downstream operations. Figures 10A-10B In the example shown, the protruding end of the first adapter and the capture portion of the second adapter are connected to different strands of the template nucleic acid molecule 1007.

[0279] The template nucleic acid molecule 1007 can be linked to the vector 1070 (1052) by annealing the protruding end of the first adaptor to a surface primer (e.g., 1072) and connecting it to the gap in the top strand 1078a. Pre-enrichment (1053) can be performed before or after cleavage, where a capture portion 1075 is used to separate the vector-template assembly from the negative vector (not bound to the template). In this example, the capture portion is biotin, and streptavidin-magnetic beads are used to capture the biotin. The U-based cleavage site can then be cleaved using a mixture of uracil-DNA glycosylase (UDG) and endonuclease VIII or a mixture of USER enzymes. Before, during, or after the extension and / or amplification reaction (1054), the bottom strand 1078b can be separated from the top strand 1078a, which is covalently bound to the vector 1070, and annealed to another surface primer on the vector 1070. Primer 1080 can be annealed with the top strand 1078a, and primer extension reactions can be initiated using the top strand 1078a as a template. Separately, the bottom strand 1078b can be used as a template to extend another surface primer annealed with said bottom strand. The template strand and / or copies thereof can then be extended and / or amplified, such that multiple surface primers on vector 1070 are extended to copy the top strand 1078a and the bottom strand 1078b. Extension and / or amplification reactions can be performed in partitions (e.g., wells, emulsions) or in batches.

[0280] Figure 10C-10D This demonstrates a fifth workflow for extension and amplification on a vector, where both template strands are amplified using an adaptor that includes a mismatched portion. Vector 1070 and template nucleic acid molecule 1007 are available. Vector 1070 may include multiple surface primers, such as 1072 (for clarity, ...). Figure 10A (Only two are shown in the figure), the plurality of surface primers 1072 may be the same and / or include a common primer sequence. Template nucleic acid molecule 1007 may include insert molecule 1003, a first adaptor 1001 (1051) attached at one end of insert molecule and a second adaptor 1005 (1051) attached at the other end of insert molecule.

[0281] The first adapter may be a partially double-stranded adapter, including a protruding end configured to bind to surface primer 1072. The first adapter 1001 may include a cleavable portion, denoted as "U", such as a uracil residue. In some cases, the first adapter 1001 may correspond to... Figure 9A The first connector 903a. The second connector 1005 can be a double-stranded connector including the mismatch portion, as per [reference to...]. Figures 10A-10BAs described. The second adaptor may include a capture portion 1075, such as biotin, and one or more cleavable portions (denoted as "U"), such as uracil residues, said capture portion being disposed at 5' of the cleavable portion. Figure 10C-10D In the example shown, the protruding end of the first adapter and the capture portion of the second adapter are connected to different strands of the template nucleic acid molecule 1007.

[0282] The template nucleic acid molecule 1007 can be linked to the vector 1070 (1052) by annealing the protruding end of the first adaptor to a surface primer (e.g., 1072) and connecting it to the gap in the top strand 1078a. Pre-enrichment (1053) can be performed before or after cleavage, where a capture portion 1075 is used to separate the vector-template assembly from the negative vector (not bound to the template). In this example, the capture portion is biotin, and streptavidin-magnetic beads are used to capture the biotin. U-based cleavage sites can be cleaved using a mixture of uracil-DNA glycosylase (UDG) and purine-free / pyrimidine-free endonuclease 1 (APE1). Using the UDG / APE1 enzyme combination may produce an extendable 3' end of the bottom strand 1078b. Taq or similar polymerases and dNTPs can be added to extend the 3'-hydroxyl site (1054). The bottom chain 1078b can be extended using the top chain 1078a bound to the vector 1070 as a template to produce an extended bottom chain comprising a sequence complementary to the surface primer 1072. Before, during, or after the extension and / or amplification reaction (1055) begins, the extended bottom chain 1078b can be separated from the top chain 1078a covalently bound to the vector 1070 (e.g., the surface of a bead or substrate) and annealed with another surface primer on the vector 1070. Primer 1080 can be annealed with the top chain 1078a, and the primer extension reaction can be initiated using the top chain 1078a as a template. Separately, the extended bottom chain 1078b can be used as a template to extend another surface primer annealed with the extended bottom chain. The template chain and / or copies thereof can then be extended and / or amplified such that multiple surface primers on the vector 1070 are extended to copy the top chain 1078a and the bottom chain 1078b. Extension and / or amplification reactions can be performed in partitions or in batches (e.g., wells, emulsions).

[0283] After amplification, if a template nucleic acid molecule containing an adaptor with a mismatch is used, as according to Figures 10A-10B and Figure 10C-10DThe workflow allows the vector to include at least two types of amplified chains. The amplified chain on the vector, derived from the top chain 1078a, may include a mismatch sequence (e.g., a homopolymer sequence) in the top chain of the mismatch portion of the second connector 1005, and the amplified chain on the vector, derived from the bottom chain 1078b, may include a complementary sequence to the mismatch sequence (e.g., a homopolymer sequence) in the bottom chain of the mismatch portion of the second connector 1005. That is, the two sets of amplified chains on the vector include different mismatch sequences (e.g., homopolymer sequences) corresponding to the mismatch portion of the second connector. Figures 10A-10B and Figure 10C-10D In the example shown, the first linker includes a polyA homopolymer sequence in the top chain of the mismatched portion and a polyC homopolymer sequence in the bottom chain of the mismatched portion. The resulting amplified beads include (i) a first type of amplified chain (derived from the top chain) which includes a polyA homopolymer sequence from the top chain of the mismatched portion, and (ii) a second type of amplified chain (derived from the bottom chain) which includes a polyG homopolymer sequence as a complementary sequence to the polyC homopolymer sequence from the bottom chain of the mismatched portion.

[0284] It should be understood that any of these workflows may omit the pre-enrichment and / or enrichment operations, and therefore may be implemented even if the template connector does not include the capture portion and / or the slicable portion adjacent to the capture portion.

[0285] Sequencing reads of amplified clusters (e.g., on amplified beads, on a substrate, etc.) can be used to determine whether the amplified cluster comprises only one or both amplified strands derived from the double-stranded template molecule, and / or to determine the percentage or ratio of amplified strands derived from the first (e.g., top or bottom) strand of the double-stranded template molecule to amplified strands derived from the second (e.g., top or bottom) strand. As used herein, the term “forward” with respect to an amplified strand can correspond to an amplified strand derived from the first strand of the double-stranded template (intended to include the first homopolymer sequence of the first strand of the mismatched portion), and the term “reverse” with respect to an amplified strand can correspond to an amplified strand derived from the second strand (not the first) of the double-stranded template (intended to include a homopolymer sequence complementary to the second homopolymer sequence of the second strand of the mismatched portion). It should be understood that the terms “forward” and “reverse” can be reversed relative to the two strands (first and second strands) of the double-stranded template molecule, but generally refer to different strands of the two strands. As used herein, the terms “forward” and “reverse” for template molecules are generally interchangeable with the terms “positive” and “negative” for template molecules.

[0286] The amplified clusters can be sequenced using any sequencing workflow or sequencing method described herein to produce sequencing reads. Sequencing may include stream-based sequencing. Stream-based sequencing methods and errors associated with these methods (such as in the stream space (relative to the base space)) are described in U.S. Patent Publication No. 2020 / 0372971, which is incorporated herein by reference in its entirety for all purposes. Sequencing may include non-terminating sequencing. Sequencing may include reversibly terminated sequencing. Chain recognition elements in mismatched portions can be identified in the stream space or the base space. Mismatched portions of a population of amplified copies (e.g., in colonies or clusters and / or tandems) can be identified by any sequencing method, such as stream-based sequencing, non-terminating sequencing, or reversibly terminated sequencing. For example, in non-terminating sequencing methods, such as when using a single base stream (e.g., A stream -> T stream -> G stream -> C stream -> repeat), sequencing signals read in the stream space can be used to identify mismatched portions. In another instance, in a sequencing termination method, for example, when using a 4-base stream (e.g., A / T / G / C stream -> A / T / G / C stream -> repeat), the sequencing signal read in the base space can be used to identify mismatches. One or more photometric and / or base-identifier algorithms can be used to process the sequencing signal and / or sequencing reads.

[0287] For ease of illustration, when sequencing is performed in a streaming space (e.g., in the case of a non-terminated nucleotide stream), if the mismatched portion in the adaptor includes the following pair of hexameric homopolymer sequences: TTTTTT / CCCCCC,

[0288] For an amplified bead containing 100% forward strand and 0% reverse strand, at the mismatch portion, the sequencing signal collected from the bead can correspond to 6T's at 100% intensity, and the base recognizer or photometric algorithm can identify the following sequence: TTTTTT;

[0289] For amplified beads comprising 0% forward strand and 100% reverse strand, at the mismatch portion, the sequencing signal collected from the bead can correspond to 6G's at 100% intensity, and the base recognizer or photometric algorithm can identify the following sequence: GGGGGG.

[0290] For an amplified bead comprising 50% forward and 50% reverse strands, at the mismatch portion, the sequencing signal collected from the bead can correspond to 6T's at 50% intensity and 6G's at 50% intensity (or 3T's at 100% intensity and 3G's at 100% intensity), and the base recognizer or photometric algorithm can identify the following sequence: TTTGGG or GGGTTT (depending on the stream order);

[0291] For an amplified bead comprising 66% forward strand and 33% reverse strand, at the mismatch portion, the sequencing signal collected from the bead can correspond to 6T's at 66% intensity and 6G's at 33% intensity (or 4T's at 100% intensity and 2G's at 100% intensity), and the base recognizer or photometric algorithm can identify the following sequence: TTTTGG or GGTTTT (depending on the flow order); and

[0292] For an amplified bead comprising 17% forward strand and 83% reverse strand, at the mismatch portion, the sequencing signal collected from the bead can correspond to 6T's at 17% intensity and 6G's at 83% intensity (or 1T's at 100% intensity and 5G's at 100% intensity), and the base recognizer or photometric algorithm can identify the following sequences: TGGGGG or GGGGGT (depending on the stream order).

[0293] In another instance, when sequencing is performed in the base space (e.g., in the case of a reversibly terminated nucleotide stream), the mismatch portion in the adaptor includes the following pair of mismatched sequences: AT / GC.

[0294] For an amplified cluster comprising 100% forward and 0% reverse strands, at the mismatch portion, the sequencing signal collected from the bead can correspond to: [A at 100% intensity; T at 100% intensity] values. In some cases, the base recognizer or photometric algorithm can identify the following sequence: AT from this data. In other cases, the base recognizer or photometric algorithm may not output any recognition for the corresponding portion. In some cases, the sequencing algorithm may output confirmation of data lacking both strand representations and / or % strand type representations, corresponding to: 100% forward and 0% reverse strand representations;

[0295] For an amplified cluster comprising 0% forward and 100% reverse strands, at the mismatch portion, the sequencing signal collected from the bead can correspond to: [C at 100% intensity; G at 100% intensity] values. In some cases, the base recognizer or photometric algorithm can identify the following sequence: CG. In other cases, the base recognizer or photometric algorithm may not output any recognition for the corresponding portion. In some cases, the sequencing algorithm may output confirmation lacking both strand representations and / or % strand type representation data, corresponding to: 0% forward and 100% reverse strand representation;

[0296] For an amplified cluster comprising 50% forward and 50% reverse strands, at the mismatch portion, the sequencing signal collected from the bead can correspond to: [A at 50% intensity and C at 50% intensity; T at 50% intensity and G at 50% intensity]. Based on these signals, in some cases, the base recognizer or photometric algorithm can identify sequences of MK (IUPAC nomenclature, where M represents (A or C) and K represents (G or T)) or NN (IUPAC nomenclature, where N represents (any base)). In other cases, the base recognizer or photometric algorithm may not output any recognition for the corresponding portion. In some cases, the sequencing algorithm can output confirmation of two-strand representation and / or % strand type representation data, corresponding to: 50% forward and 50% reverse strand representation;

[0297] For an amplified cluster comprising 66% forward and 33% reverse strands, at the mismatch portion, the sequencing signal collected from the bead can correspond to: [A at 66% intensity and C at 33% intensity; T at 66% intensity and G at 33% intensity]. Based on these signals, in some cases, the base recognizer or photometric algorithm can identify sequences of AT, MK (IUPAC nomenclature, where M represents (A or C) and K represents (G or T)), or NN (IUPAC nomenclature, where N represents (any base)). In other cases, the base recognizer or photometric algorithm may not output any recognition for the corresponding portion. In some cases, the sequencing algorithm can output confirmation of two-strand representation and / or % strand type representation data, corresponding to: 66% forward and 33% reverse strand representation; and

[0298] For an amplified cluster comprising 17% forward and 83% reverse strands, at the mismatch portion, the sequencing signal collected from the bead can correspond to [A at 17% intensity and C at 83% intensity; T at 17% intensity and G at 83% intensity]. Based on these signals, in some cases, the base recognizer or photometric algorithm can identify sequences of CG, MK (IUPAC nomenclature, where M represents (A or C) and K represents (G or T)), or NN (IUPAC nomenclature, where N represents (any base)). In other cases, the base recognizer or photometric algorithm may not output any recognition for the corresponding portion. In some cases, the sequencing algorithm can output confirmation of two-strand representation and / or %strand type representation data, corresponding to: 17% forward and 83% reverse strand representation.

[0299] Thus, sequencing reads at the mismatched portion of the adaptor can be processed for the pair of sequences (e.g., non-complementary sequences, non-complementary homopolymer sequences, etc.) to determine the forward and reverse strand percentages on the amplified vector. The sequencing algorithm can output confirmation (or lack of confirmation) for both the strand type representation and / or the % strand type representation data (e.g., 17% forward and 83% reverse in a cluster). In some cases, confirmation (or lack of confirmation) for both the strand type representation and / or the % strand type representation data (e.g., 17% forward and 83% reverse in a cluster) can be determined by a photometric-only algorithm (instead of or in addition to a base recognition algorithm). For example, the mismatched portion can be classified as a preamble (e.g., which may additionally include other functional sequences such as calibration sequences, barcode sequences, adaptor sequences, etc.), whose sequencing signal is decrypted by a photometric algorithm, rather than base recognition. In some cases, the % strand type representation data can be used downstream by a base recognizer or a photometric algorithm to identify the sequence of the template insertion portion. In some cases, loci that have identified sequencing signals, base recognition, and / or sequencing read inconsistencies between positive and negative strand derivatives can be excluded from downstream variant recognition, such as SNV or SNP recognition.

[0300] High-accuracy sequencing, high-accuracy SNV or SNP identification, or both, achieved through the amplification and / or sequencing methods described herein can be particularly advantageous for applications such as the detection of residual disease or minimal residual disease (MRD). Detection of circulating tumor DNA (ctDNA) in the blood (e.g., from cell-free DNA (cfDNA) samples) can help identify patients more likely to have cancer recurrence and monitor treatment efficacy based on the detection and quantification of residual ctDNA before, during, and after treatment. Compared to currently available MRD assays that rely on targeted sequencing based on enriched, limited numbers of target mutations or methylation markers, the ability to perform high-accuracy SNV identification, with data point redundancy (from the target) used to tolerate sequencing errors (e.g., SNV identification errors), enables the detection of residual disease or MRD using relatively low-depth whole-genome sequencing (WGS). SNVs or SNPs identified using any of the strand-recognition adaptor compositions, kits, methods, or systems described herein can be used to determine the level of residual disease, MRD, tumor fraction, or circulating tumor fraction in subjects whose samples have been sequenced.

[0301] While one example illustrates a mismatch portion comprising a pair of hexameric homopolymer sequences, the mismatch portion of the linker can include homopolymer sequences of any length, such as 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, or longer. In some cases, the mismatch portion can include homopolymer sequences of at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, or longer. Alternatively or additionally, the mismatch portion may include a homopolymer sequence length of up to 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, or 30 polymers.

[0302] As described elsewhere in this document, the mismatch portion of the linker can include any pair of non-complementary sequences. The sequence of the pair of non-complementary sequences can include a homopolymer sequence. The sequence of the pair of non-complementary sequences may not include a homopolymer sequence. The sequence of the pair of non-complementary sequences can include multiple homopolymer sequences (e.g., AAAGATTT / GGGTCCCC). The pair of non-complementary sequences can include the same length (e.g., 5-mer / 5-mer pair). The pair of non-complementary sequences can include different lengths (e.g., 5-mer / 8-mer pair). The pair of non-complementary sequences can include sequences of any length, such as about, at least about, and / or at most about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 45, 50, or longer. In some cases, portions of the pair of non-complementary sequences may be inversely complementary sequences, but nevertheless, entire segments of the pair may be non-complementary (e.g., in the pair of non-complementary sequences: AATTGCA / CCAACTG, TTG / AAC portions are complementary but other portions are not). In some cases, mismatched portions may include single-base mismatches or protruding loops (e.g., insertions of any number of bases in a single chain).

[0303] In some cases, the template nucleic acid molecule may undergo one or more resynchronization streams after sequencing through the mismatched portion and before sequencing through the inserted sequence portion of the template nucleic acid molecule.

[0304] This document provides a method comprising (1) providing a processed product (e.g., an amplified product) that preserves the pairing association of the two strands of a template nucleic acid, and (2) generating error-corrected sequencing reads of the template nucleic acid by sequencing two sequence portions of the processed product derived from the two strands of the template nucleic acid, respectively, simultaneously, in a temporally different manner, and / or in a spatially different manner. For example, two sequencing primers hybridizing with the processed product may extend simultaneously and synchronously through the two sequence portions (which may or may not be in the same cluster) to generate sequencing data. In another instance, two sequencing primers may extend through the two sequence portions (which may or may not be in the same cluster) at different time points and / or asynchronously to generate sequencing data. In yet another instance, sequencing primers may extend through both of the two sequence portions (e.g., within the same nucleic acid strand of the processed product) to generate sequencing data. In some cases, the pairing association information of the two strands of the template nucleic acid can be preserved by at least two molecules (e.g., one molecule per strand of the template nucleic acid), which can be associated by being fixed to the same spatial location and / or vector. In some cases, the pairing association information of the two strands of a template nucleic acid can be preserved by a single molecule (e.g., including information from both strands), such as by a tandem product or other products that link the two strands of the template nucleic acid or other products derived from the tandem version.

[0305] Advantageously, the methods described herein can simultaneously sequence material derived from two strands of a library molecule and / or generate a single sequencing read representing two strand derivatives of the template nucleic acid. For example, the same nucleotide stream or nucleotide stream set can simultaneously probe material derived from both strands. This simultaneous sequencing of two strand derivatives is different and more efficient than generating two distinct sequencing reads each belonging to a different strand of the library molecule—for example, some paired-end sequencing methods involve generating a first read by sequencing the first strand of the template nucleic acid and generating a second read (different from the first read) by sequencing the second strand of the template nucleic acid (and / or the reverse complementary sequence of the first strand), and then pairing the first and second reads to generate a shared paired read. Advantageously, simultaneous sequencing of two strand derivatives allows for the association and identification of their presence without the need for barcoding, or without sequencing additional bases in the barcoded region—for example, some barcoded duplex methods provide a unique barcode for each molecule and then later associate different strand derivatives using the same barcode; however, such barcoded duplex methods require significant overhead, including large amounts of barcoding reagents and redundant sequencing and data processing of the barcoded region. In some of the methods presented in this paper, the two strand derivatives are already provided in a single amplified cluster, such that the signal read from the amplified cluster (spatial index to the amplified cluster) automatically associates the two strand derivatives.

[0306] This article provides methods for generating amplified products that retain or preserve information about both strands of a library nucleic acid molecule. The amplified products can be generated using any amplification method, such as PCR, ePCR, RPA, eRPA, RCA, MDA, or bridging amplification. Examples of preserving two-strand information using various amplification methods are detailed here. In some cases, the two strands of the library nucleic acid molecule can be ligated into a single strand prior to amplification. The amplified products can include tandem molecules (each comprising multiple copies of the template strand or multiple copies of both template strands) or distinct molecules (each comprising a copy of the template strand or each comprising a copy of each of the two template strands). The amplified products can be generated as clusters or colonies, each derived from a different library molecule. The amplified products can be generated on a vector or while ligated into a vector. The amplified products may be generated in solution. The amplified products can be sequenced while immobilized on a substrate, as described elsewhere in this article.

[0307] This document provides a method for generating a detectable amplified vector comprising amplified strands of two strands derived from a double-stranded template nucleic acid molecule. The method may include (a) conjugating a double-stranded template molecule to a vector to generate a template-conjugated vector, wherein the double-stranded template molecule comprises a first strand and a second strand, wherein the double-stranded template molecule includes an adaptor comprising a mismatch portion, wherein the mismatch portion comprises a first sequence in the first strand and a second sequence in the second strand that is not complementary to the first sequence; and (b) amplifying the double-stranded template molecule of the template-conjugated vector to generate an amplified vector comprising a plurality of amplified strands conjugated thereto. The plurality of amplified strands may include derivatives from both the first and second strands. The method may further include hybridizing a plurality of sequencing primers to the plurality of amplified strands and extending the plurality of sequencing primers to sequence the amplified strands. Sequencing signals collected from individually addressable locations following a sequencing stream may represent the two strand derivatives.

[0308] This document provides a method for detecting amplified strands on a vector, such as the percentage or ratio of amplified strands derived from one or both strands of a double-stranded template nucleic acid molecule. One method may include: (a) conjugating a double-stranded template molecule to a vector to produce a template-conjugated vector, wherein the double-stranded template molecule comprises a first strand and a second strand, wherein the double-stranded template molecule includes an adaptor, the adaptor including a mismatch portion, wherein the mismatch portion includes a first sequence in the first strand and a second sequence in the second strand that is not complementary to the first sequence; (b) amplifying the template-conjugated vector to produce an amplified vector, the amplified vector comprising a plurality of amplified strands conjugated thereto; (c) sequencing the amplified vector to produce sequencing reads; and (d) determining, at least in part, the percentage of amplified strands derived from the first strand among the plurality of amplified strands in the amplified vector based on portions of the sequencing reads corresponding to the mismatch portion. The sequencing described in (c) may include hybridizing a plurality of sequencing primers with the plurality of amplified strands derived from both the first and second strands, and extending the plurality of sequencing primers to sequence the plurality of amplified strands. Sequencing signals collected from individually addressable locations following the sequencing stream may represent the two strand derivatives.

[0309] This document provides a method for generating a detectable amplified vector comprising amplified strands of two strands derived from a double-stranded template nucleic acid molecule. The method may include (a) conjugating a double-stranded template molecule to a vector to generate a template-conjugated vector, wherein the double-stranded template molecule comprises a first strand and a second strand, wherein the double-stranded template molecule includes an adaptor, the adaptor including a mismatch portion, wherein the mismatch portion includes a first homopolymer sequence in the first strand and a second homopolymer sequence in the second strand that is not complementary to the first homopolymer sequence; and (b) amplifying the double-stranded template molecule of the template-conjugated vector to generate an amplified vector comprising a plurality of amplified strands conjugated thereto.

[0310] This document provides a method for detecting amplified strands on a vector, such as the percentage or ratio of amplified strands derived from one or both strands of a double-stranded template nucleic acid molecule. One method may include: (a) conjugating a double-stranded template molecule to a vector to produce a template-conjugated vector, wherein the double-stranded template molecule comprises a first strand and a second strand, wherein the double-stranded template molecule includes an adaptor, the adaptor including a mismatch portion, wherein the mismatch portion includes a first homopolymer sequence in the first strand and a second homopolymer sequence in the second strand that is not complementary to the first homopolymer sequence; (b) subjecting the template-conjugated vector to amplification to produce an amplified vector, the amplified vector comprising a plurality of amplified strands conjugated thereto; (c) sequencing the amplified vector to produce sequencing reads; and (d) determining, at least in part, the percentage of amplified strands derived from the first strand among the plurality of amplified strands in the amplified vector based on portions of the sequencing reads corresponding to the mismatch portion.

[0311] This document provides a kit for generating a detectable amplified vector and detecting the amplified strand on the vector. One kit may include a double-stranded linker comprising a mismatch portion, wherein the double-stranded linker includes a first strand and a second strand, wherein the mismatch portion includes a first sequence in the first strand and a second sequence in the second strand that is not complementary to the first sequence. The kit may include a plurality of double-stranded linkers including the mismatch portion, wherein the mismatch portion of the plurality of double-stranded linkers is identical.

[0312] This document provides a kit for generating detectable amplified vectors and detecting amplified chains on the vectors. One kit may include a double-stranded linker comprising a mismatch portion, wherein the double-stranded linker includes a first chain and a second chain, wherein the mismatch portion includes a first homopolymer sequence in the first chain and a second homopolymer sequence in the second chain that is not complementary to the first homopolymer sequence. The kit may include a plurality of double-stranded linkers including the mismatch portion, wherein the mismatch portion of the plurality of double-stranded linkers is identical.

[0313] This document provides compositions for generating detectable amplified vectors and detecting amplified strands on the vectors. A composition may comprise a double-stranded terminator including a mismatch portion, wherein the double-stranded terminator includes a first strand and a second strand, wherein the mismatch portion includes a first sequence in the first strand and a second sequence in the second strand that is not complementary to the first sequence. The composition may comprise a plurality of double-stranded termins, each of the plurality of double-stranded termins including the mismatch portion, wherein the mismatch portion of the plurality of double-stranded termins is identical. The composition may comprise a plurality of template molecules, wherein the plurality of template molecules include a plurality of double-stranded template insert molecules linked to the plurality of double-stranded termins. The composition may comprise template molecules, wherein the template molecules include double-stranded template insert molecules linked to the double-stranded termins. The composition may comprise a vector. The composition may comprise a plurality of vectors. The vector may be linked to a template molecule.

[0314] This document provides compositions for generating detectable amplified vectors and detecting amplified chains on the vectors. A composition may comprise a double-stranded linker including a mismatch portion, wherein the double-stranded linker includes a first chain and a second chain, wherein the mismatch portion includes a first homopolymer sequence in the first chain and a second homopolymer sequence that is not complementary to the first homopolymer sequence in the second chain. The composition may comprise a plurality of double-stranded linkers, each of the plurality of double-stranded linkers including the mismatch portion, wherein the mismatch portion of the plurality of double-stranded linkers is identical. The composition may comprise a plurality of template molecules, wherein the plurality of template molecules includes a plurality of double-stranded template insert molecules linked to the plurality of double-stranded linkers. The composition may comprise template molecules, wherein the template molecules include double-stranded template insert molecules linked to the double-stranded linkers. The composition may comprise a vector. The composition may comprise a plurality of vectors. The vector may be linked to a template molecule.

[0315] enrichment of balance carriers

[0316] Multiple balanced vectors can be generated using any of the methods described herein. A balanced vector can generally refer to a vector comprising amplified strands, wherein the amplified strands can be derived from the two strands of a double-stranded template nucleic acid molecule, wherein the template nucleic acid molecule includes mismatched portions. A balanced vector can comprise any ratio of a first set of amplified strands and a second set of amplified strands derived from the first and second strands of the template nucleic acid molecule, respectively. In some cases, a balanced vector can comprise an amplified strand derived from only one of the two strands of the template nucleic acid molecule (either forward-facing or reverse-facing only).

[0317] Multiple balanced vectors can be enriched and then loaded onto a substrate for sequencing, for example, to enrich balanced vectors comprising amplified strands of both strands derived from a template nucleic acid molecule (as opposed to balanced vectors comprising only the forward or reverse strand). The portion of the amplified strand corresponding to a mismatch in the template nucleic acid molecule can be used for enrichment. In some cases, the mismatch portion can be designed to include relatively long sequences, such as at least 5-mer, 6-mer, 7-mer, 8-mer, 9-mer, 10-mer, 15-mer, 20-mer, 25-mer, 30-mer, 35-mer, 40-mer, or longer homopolymer sequences, or can be designed to include additional capture sequences to assist in the enrichment method.

[0318] In some cases, high-affinity peptide nucleic acids (PNAs) or locked nucleic acids (LNAs) can be used as enrichment molecules to capture mismatched portions (e.g., hybridize with mismatched portions). Any nucleic acid molecule can be used as an enrichment molecule.

[0319] Enrichment can be performed by using enriching molecules in one or more enrichment steps to capture equilibrium carriers, including amplified chains containing specific sequences, thereby retaining the captured equilibrium carriers and removing uncaptured equilibrium carriers from the mixture.

[0320] Therefore, a method for enrichment may include (a) providing a plurality of balancing vectors comprising a plurality of amplified strands, wherein the plurality of amplified strands are derived from a plurality of double-stranded template nucleic acid molecules, each of the plurality of double-stranded template nucleic acid molecules comprising a mismatch portion, wherein the mismatch portion comprises a first strand and a second strand, the first strand comprising a first mismatch sequence, and the second strand comprising a second mismatch sequence that is not complementary to the first mismatch sequence, such that each of the plurality of amplified strands comprises, at a corresponding portion corresponding to the mismatch portion, an inverse complementary sequence of the first mismatch sequence or the second mismatch sequence; (b) contacting a plurality of first enrichment molecules with the plurality of amplified strands, wherein the plurality of first enrichment molecules comprises a first capture sequence complementary to the first mismatch sequence, to produce a first set of enriched balancing vectors; and (c) contacting a plurality of second enrichment molecules with the first set of enriched balancing vectors, wherein the plurality of second enrichment molecules comprises a second capture sequence comprising the second mismatch sequence, to produce a second set of enriched balancing vectors. The two types of enrichment molecules may be introduced in any order. For example, a method for enrichment may include (a) providing a plurality of balancing vectors comprising a plurality of amplified strands, wherein the plurality of amplified strands are derived from a plurality of double-stranded template nucleic acid molecules, each of the plurality of double-stranded template nucleic acid molecules comprising a mismatch portion, wherein the mismatch portion comprises a first strand and a second strand, the first strand comprising a first mismatch sequence, and the second strand comprising a second mismatch sequence that is not complementary to the first mismatch sequence, such that each of the plurality of amplified strands comprises, at a corresponding portion corresponding to the mismatch portion, an inverse complementary sequence of the first mismatch sequence or the second mismatch sequence; (b) contacting a plurality of first enrichment molecules with the plurality of amplified strands, wherein the plurality of first enrichment molecules comprises a first capture sequence comprising the second mismatch sequence to produce a first set of enriched balancing vectors comprising the second mismatch sequence; and (c) contacting a plurality of second enrichment molecules with the first set of enriched balancing vectors, wherein the plurality of second enrichment molecules comprises a second capture sequence complementary to the first mismatch sequence to produce a second set of enriched balancing vectors.

[0321] When a set of enriched molecules is contacted with a set of balancing carriers, the set of enriched balancing carriers can be pulled down and / or otherwise separated. In some cases, the enriched molecules may include a capturing portion, which is subsequently captured by a capturing portion, according to any capture / capturing portion pair described elsewhere herein. In some cases, the enriched molecules may be immobilized, and / or the capturing portion may be immobilized. Following these methods, in which derivatives of both strands of the mismatched portion are targeted and enriched, a second set of enriched balancing carriers may represent enriched balancing carriers, wherein each carrier comprises amplified strands of both strands derived from the template nucleic acid molecule.

[0322] Figure 13 A two-step enrichment scheme according to embodiments of the present disclosure is illustrated. Multiple balancing vectors 1302 can be contacted with multiple first enrichment molecules 1308 to isolate and / or generate a first set of enriched balancing vectors 1304. The multiple balancing vectors 1302 may include multiple amplified strands, each amplified strand including a portion corresponding to a mismatch portion of a template nucleic acid molecule, the mismatch portion including a first mismatch sequence (e.g., AAA) and a second mismatch sequence (e.g., CCC) that is not complementary to the first mismatch sequence in the first and second strands, respectively. The amplified strand derived from the first strand may include the first mismatch sequence (e.g., AAA), and the amplified strand derived from the second strand may include a complementary sequence (e.g., GGG) of the second mismatch sequence in the corresponding portion corresponding to the mismatch portion. The first enrichment molecule 1308 may include a first capture sequence complementary to one of two sequences in the portion of the amplified strand corresponding to the mismatch portion. In this figure, the first enrichment molecule includes a first capture sequence comprising a second mismatch sequence (e.g., CCC) complementary to a complementary sequence of the second mismatch sequence (e.g., GGG) in the amplified strand. Therefore, the first set of enriched balancing vectors 1304 comprises only balancing vectors comprising the amplified strand derived from the second strand. That is, balancing vectors comprising only the amplified strand derived from the first strand (e.g., forward-only vectors) are removed. In a second step, the first set of enriched balancing vectors 1304 may be contacted with a plurality of second enrichment molecules 1310 to separate and / or generate a second set of enriched balancing vectors 1306. The second enrichment molecule 1310 may include a second capture sequence complementary to another sequence of two sequences in the portion of the amplified strand corresponding to the mismatch portion. In this figure, the second enrichment molecule includes a second capture sequence (e.g., TTT) complementary to the first mismatch sequence (e.g., AAA) in the amplified strand. Therefore, the second set of enriched balancing vectors 1306 comprises balancing vectors only, which include an amplified strand derived from the first strand. That is, balancing vectors comprising an amplified strand derived from the second strand (e.g., reverse-biased vectors only) are removed. After performing these two steps, the second set of enriched balancing vectors comprises balancing vectors only, which include an amplified strand derived from both strands of the template nucleic acid molecule.

[0323] Alternatively or additionally, enrichment can be carried out by using enriching molecules in one or more enrichment steps to capture a balance carrier comprising an amplified chain containing a specific sequence, and to deplete or partially deplete the captured balance carrier from the mixture.

[0324] Therefore, a method for enrichment may include (a) providing a plurality of balancing vectors comprising a plurality of amplified strands, wherein the plurality of amplified strands are derived from a plurality of double-stranded template nucleic acid molecules, each of the plurality of double-stranded template nucleic acid molecules comprising a mismatch portion, wherein the mismatch portion comprises a first strand and a second strand, the first strand comprising a first mismatch sequence, and the second strand comprising a second mismatch sequence that is not complementary to the first mismatch sequence, such that each of the plurality of amplified strands comprises, at a corresponding portion corresponding to the mismatch portion, an inverse complementary sequence of the first mismatch sequence or the second mismatch sequence; (b) contacting a plurality of first enrichment molecules with the plurality of amplified strands, wherein the plurality of first enrichment molecules comprises a first capture sequence complementary to the first mismatch sequence to capture a subset of balancing vectors; and (c) removing at least a portion of the subset of balancing vectors from the plurality of balancing vectors to produce a first set of enriched balancing vectors. In some cases, a method for enrichment may include (a) providing a plurality of balancing vectors comprising a plurality of amplified strands, wherein the plurality of amplified strands are derived from a plurality of double-stranded template nucleic acid molecules, each of the plurality of double-stranded template nucleic acid molecules comprising a mismatch portion, wherein the mismatch portion comprises a first strand and a second strand, the first strand comprising a first mismatch sequence and the second strand comprising a second mismatch sequence that is not complementary to the first mismatch sequence, such that each of the plurality of amplified strands comprises, at a corresponding portion corresponding to the mismatch portion, an inverse complementary sequence of the first mismatch sequence or the second mismatch sequence; (b) contacting a plurality of first enrichment molecules with the plurality of amplified strands, wherein the plurality of first enrichment molecules comprises a first capture sequence comprising the second mismatch sequence to capture a subset of balancing vectors; and (c) removing at least a portion of the subset of balancing vectors from the plurality of balancing vectors to produce a first set of enriched balancing vectors.

[0325] In these partial depletion methods, the first enrichment molecule can be designed to capture only a portion (less than 100%) of the balanced vector, comprising an amplified strand originating from one strand. This portion of the balanced vector can include only a subset of the balanced vector comprising an amplified strand originating from said one strand (e.g., only forward or only reverse balanced vectors). All or a subset of the captured balanced vectors can then be removed from the mixture of balanced vectors to produce a first set of enriched balanced vectors—the first set of enriched balanced vectors can include vectors comprising amplified strands of both strands originating from the template nucleic acid molecule. Alternatively, the first enrichment molecule can be designed to capture substantially all balanced vectors comprising an amplified strand originating from the first strand, and only a subset of the captured balanced vectors can be removed from the mixture of balanced vectors to produce a first set of enriched balanced vectors—the first set of enriched balanced vectors can include vectors comprising amplified strands of both strands originating from the template nucleic acid molecule.

[0326] Any method described herein may further include additional enrichment steps, wherein enrichment molecules are used to capture a subset of an enriched set of balanced vectors, the enrichment molecules comprising capture sequences enriched using corresponding portions of amplified strands corresponding to mismatched portions, and (i) retaining such subsets and removing uncaptured subsets; or (ii) removing at least a portion of such subsets to produce a second set of enriched balanced vectors. The enrichment methods described herein may include any number of steps, e.g., one, two, three, four, five, six, seven, eight, or more, of separating and / or removing subsets of the captured vectors using different sets of enrichment molecules.

[0327] Although Figure 13 The examples illustrate that mismatches include a pair of non-complementary homopolymer sequences (e.g., polyA / polyC), but it should be understood that mismatches can include any non-complementary sequence pair (e.g., non-homogeneous sequences) and can be of any configuration. Template nucleic acid molecules may include alternative or additional chain recognition elements (e.g., mismatches). Figure 14The diagram illustrates different example configurations of a template nucleic acid molecule and a pre-amplified vector having strand recognition elements, with insets showing: (A) including a circular mismatch portion, (B) including a divergent mismatch portion (Y-shaped), (C) including both a circular mismatch portion and a distal divergent mismatch portion (Y-shaped), (D) including two circular mismatch portions, and (E) including two circular mismatch portions and a distal annealed double-stranded portion, the annealed double-stranded portion including a cleavable portion that, upon cleavage (e.g., USER treatment during vector preparation), transforms the distal circular mismatch portion into a divergent mismatch portion (Y-shaped), as shown in inset (C). The template nucleic acid molecule may include any number of mismatch portions, such as 1, 2, 3, 4, 5, or more. The mismatch portions may be positioned closer to the vector relative to the sequencing primer binding site, such that the portions corresponding to the mismatch portions and the portions corresponding to the insert molecule portions are sequenced together. The sequencing reads corresponding to the mismatched portion can be used to determine the presence of one or both strands of the amplified strand derived from the template nucleic acid molecule, and / or to determine the ratio of the forward and / or reverse strands in the total amplified strands of the balancing vector. The mismatched portion can be positioned further away from the vector relative to the sequencing primer binding site, such that the portion corresponding to the mismatched portion is not sequenced. In this case, the mismatched portion can be used for enrichment purposes, as described elsewhere herein. In some cases, where the template nucleic acid molecule includes multiple mismatched portions, including a first mismatched portion closer to the vector and a second mismatched portion farther from the vector, the sequencing primer binding site can be positioned between the first and second mismatched portions, such that the portion corresponding to the first mismatched portion closer to the vector can be sequenced and said portion used for strand ratio detection purposes, and / or the portion corresponding to the second mismatched portion farther from the vector can be not sequenced and said portion used for enrichment purposes, as described herein. In some cases, the portion corresponding to the mismatched portion can be used for both strand detection and enrichment purposes.

[0328] The systems, kits, and compositions disclosed herein may contain any component of any method described herein, such as vectors, amplified vectors, template nucleic acids, template-ligated vectors, adaptors, strand substitution adaptors (e.g., adaptors including mismatched portions), template nucleic acids including mismatched portions, first enrichment molecules, second enrichment molecules, capture portions, capture moieties, different enzymes, different primers, amplification reagents, sequencing reagents, etc.

[0329] Fine-tuning the chain ratio in the balance carrier

[0330] This article provides a method for fine-tuning the forward-reverse strand ratio in a balanced vector. Mismatched portions in the template nucleic acid molecule and / or their corresponding portions in the derived strand can be used as primer binding sites during amplification. In cases where the template nucleic acid molecule includes mismatched portions comprising a first mismatched sequence and a second mismatched sequence, respectively, in the first and second strands, (a) the forward amplification primer may include the reverse complementary sequence of the first mismatched sequence to hybridize with the first amplified strand derived from the first strand and initiate an extension reaction from the first amplified strand, and (b) the reverse amplification primer may include the second mismatched sequence to hybridize with the second amplified strand derived from the second strand and initiate an extension reaction from the second amplified strand. The corresponding concentrations of the forward and reverse amplification primers provided during amplification can be adjusted to fine-tune the forward-reverse strand ratio in the resulting balanced vector. For example, if more forward amplified strand is needed in the balanced vector, the concentration of the forward amplification primer can be increased relative to the concentration of the reverse amplification primer, and vice versa.

[0331] Therefore, a method for generating an amplified vector having a predetermined forward-reverse strand ratio includes contacting: (i) a template-linked vector, wherein the template-linked vector comprises a vector linked to a double-stranded template molecule, wherein the double-stranded template molecule comprises a first strand and a second strand, wherein the double-stranded template molecule comprises an adapter, the adapter comprising a mismatch portion, wherein the mismatch portion comprises a first mismatch sequence in the first strand and a second mismatch sequence in the second strand that is not complementary to the first mismatch sequence; and (ii) a plurality of forward amplification primers at a first predetermined concentration and a plurality of reverse amplification primers at a second predetermined concentration to generate the amplified vector comprising a plurality of amplified strands having a predetermined forward-reverse strand ratio, wherein each of the plurality of forward amplification primers comprises a reverse complementary sequence of the first mismatch sequence to hybridize with a first amplified strand derived from the first strand, and wherein each of the plurality of reverse amplification primers comprises the second mismatch sequence to hybridize with a second amplified strand derived from the second strand.

[0332] Alternatively, or in addition to adjusting the corresponding concentrations of the forward and reverse amplification primers, various other parameters, such as the affinity of the corresponding primers for the amplified strand at the portion corresponding to the mismatch, can be adjusted to fine-tune the forward-reverse strand ratio in the resulting balanced vector. Alternatively, or in addition to adjusting the corresponding parameters of the forward and reverse amplification primers, various parameters of the amplification program can be adjusted to fine-tune the forward-reverse strand ratio. For example, during one or more PCR cycles of the amplification program, the temperature can be adjusted so that the annealing temperature of one type of primer is superior to that of another type. In these cases, the forward and reverse amplification primers can be designed to have different annealing temperatures for the corresponding portions corresponding to the mismatch.

[0333] The systems, kits, and compositions disclosed herein may contain any components of the methods described herein, such as vectors, amplified vectors, template nucleic acids, template-ligated vectors, adaptors, adaptors including mismatched portions, template nucleic acids including mismatched portions, forward amplification primers, reverse amplification primers, amplification reagents, sequencing reagents, etc.

[0334] Independent of sequencing detection chain ratio

[0335] Methods are provided for detecting the forward-reverse strand ratio of a balanced vector independently of and / or in addition to sequencing. These methods can detect the forward-reverse strand ratio of a balanced vector without requiring sequencing or generating sequencing reads. Probes comprising a capture sequence and a detectable portion can be configured to bind their capture sequence to the amplified strand at a location corresponding to a mismatched portion of the template nucleic acid molecule. The probe can be a labeled oligonucleotide probe. In cases where the template nucleic acid molecule includes mismatched portions comprising a first mismatched sequence and a second mismatched sequence, respectively, in a first strand and a second strand, a first type of probe can be configured to detect the amplified strand derived from the first strand, wherein each of the amplified strands includes the first mismatched sequence. The first type of probe may include a first capture sequence comprising the reverse complementary sequence of the first mismatched sequence and a first detectable portion. A second type of probe can be configured to detect the amplified strand derived from the second strand, wherein each of the amplified strands includes the reverse complementary sequence of the second mismatched sequence. The second type of probe may include a second capture sequence comprising the second mismatched sequence and a second detectable portion.

[0336] Only one type of probe can be used. Both types of probes can be used sequentially or simultaneously. The detectable portion of the first type and the detectable portion of the second type can be the same type or different types of detectable portions. The detectable portion can be any label as described herein, such as a fluorescent dye. In some cases, the first type and the second type of detectable portions can be detected at the same frequency or within the same frequency range. In some cases, the first type and the second type of detectable portions can be detected at different frequencies or within different frequency ranges.

[0337] The probe can contact the equilibration vector immobilized to the substrate, wash away unbound probes, and detect a signal from the equilibration vector. After detection, before collecting sequencing signals from the equilibration vector, bound probes can be removed from the equilibration vector (e.g., denatured from the equilibration vector) and / or the detectable portion can be inactivated (e.g., dye cleavage). The signal detected from the probe can be used to determine the relative amounts of the forward and / or reverse strands on the equilibration vector. In some cases, the first type of probe and the second type of probe can contact the equilibration vector simultaneously or at different times. Signals from each probe type can be detected simultaneously, cumulatively, or at different times. Signals from each probe type collected at each location can be processed to determine the forward-to-reverse strand ratio of the equilibration vector immobilized at such loci. In some cases, one type of probe can contact the equilibration vector, and the first signal is detected before the second type of probe contacts the equilibration vector. In some cases, the first type of probe can be removed from the detectable portion and / or the detectable portion can be inactivated before contacting the equilibration vector with the second type of probe and collecting the second signal. In other cases, if the first detectable portion of the first type probe remains present and active when the second type probe is contacted, the second signal can be detected at a different frequency or channel to distinguish it from the first detectable portion, or it can be detected at the same frequency or channel as the first detectable portion, representing the cumulative signal from both the first and second type probes. In this case, the second and first signals can be processed differently to determine the signal collected from only the second type probe. In some cases, signals from both the first and second type probes can be collected after both probes have been contacted with a balancing carrier and detected at different frequencies or channels.

[0338] Therefore, a method for detecting a strand on a vector may include (a) providing a plurality of balanced vectors immobilized to a substrate, wherein the plurality of balanced vectors comprise a plurality of amplified strands, wherein the plurality of amplified strands are derived from a plurality of double-stranded template nucleic acid molecules, each of the plurality of double-stranded template nucleic acid molecules comprising a mismatch portion, wherein the mismatch portion comprises a first strand and a second strand, the first strand comprising a first mismatch sequence, and the second strand comprising a second mismatch sequence that is not complementary to the first mismatch sequence, such that each of the plurality of amplified strands comprises, at a corresponding portion corresponding to the mismatch portion, an inverse complementary sequence of the first mismatch sequence or the second mismatch sequence; (b) contacting a plurality of first-type probes with the plurality of amplified strands, wherein each of the plurality of first-type probes comprises a first capture sequence complementary to the first mismatch sequence and a first detectable portion; and (c) detecting a first signal, the first signal being derived from the first detectable portion of a subset of the plurality of first-type probes bound to a subset of the plurality of amplified strands. The method may further include (d) contacting a plurality of second-type probes with the plurality of amplified strands, wherein each of the plurality of second-type probes includes a second capture sequence comprising the second mismatch sequence and a second detectable portion; and (e) detecting a second signal, the second signal being derived from the second detectable portion of a subset of the plurality of second-type probes that binds to a second subset of the plurality of amplified strands. The systems, kits, and compositions of this disclosure may comprise any component of the methods described herein, such as vectors, amplified vectors, template nucleic acids, template-linked vectors, adaptors, adaptors including mismatch portions, template nucleic acids including mismatch portions, substrates, first probes, second probes, detectable portions, amplification reagents, sequencing reagents, detectors, etc.

[0339] In any method of this disclosure, the template insert molecule or derivative may be attached to a strand recognition adaptor at one end, such that the sequenced product includes a strand recognition element at only one end. Alternatively, in any method of this disclosure, the template insert molecule or derivative may be attached to a strand recognition adaptor at both ends (which may or may not include identical mismatched sequences), such that the sequenced product includes a strand recognition element at both ends. For example, in Figure 10A-10D In the illustrated workflow, adaptor 1001 may also include a mismatch portion (allowing adaptor 1001 to be a strand recognition adaptor in addition to adaptor 1005). In some cases, attaching the strand recognition adaptor to both ends of the template insert molecule allows for strand recognition and / or quantification when read from either end, which may facilitate paired-end sequencing workflows. In some cases, one or two strand recognition sequences can be identified when the entire strand sequence is fully sequenced at sufficient sequencing quality for further validation.

[0340] Other amplification methods for maintaining double-strand association

[0341] In some cases, it is desirable to amplify sample nucleic acids prior to sequencing (e.g., when the sample contains only a small amount of template nucleic acid molecules, for targeted sequencing purposes, etc.). Advantageously, the amplification methods described herein enable the preservation of both strands of a double-stranded template molecule (e.g., before or simultaneously with library preparation). In some cases, this amplification can be performed prior to further processing for sequencing (e.g., PCR, ePCR, RPA, eRPA, or bridging amplification for the purpose of forming colonies on a vector). In some cases, the amplification methods can provide amplified products that maintain double-strand association and have strand recognition elements. In some cases, the amplification methods can provide amplified products that maintain double-strand association but lack strand recognition elements—such amplified products can then be labeled with strand recognition elements (e.g., linked to mismatched adaptors as described herein) and then further amplified to produce amplified clusters comprising the strand with strand recognition elements. This further amplification of the amplified products can generate repeating colonies (optionally immobilized to separate vectors) in different clusters. These repeats can further reduce sequencing errors by providing comparative data points within the sample (e.g., one repeat sample can be compared to another for internal error correction, such as by discarding differentially identified bases or reads between repeat samples). As described elsewhere in this document, error-corrected sequencing can be performed on the amplified clusters based on strand recognition elements. Alternatively or in addition to [the above], Figure 9B-10D In addition to the described workflow, the following methods can be used to maintain double-strand association and preserve information about the two strands of a library molecule.

[0342] Figures 15A-15C A workflow for amplifying double-stranded template molecules while maintaining forward-reverse strand association is demonstrated (e.g., amplification without losing the double strand). The insert molecule 1504 (e.g., sample nucleic acid) is provided with multiple adaptors, wherein at least one type of adaptor includes a mismatch region and acts as a strand recognition adaptor. Figure 14 As shown and described elsewhere in this document, mismatch regions can include single base mismatches, multiple base mismatches, and hairpin regions, etc. Figures 15A-15B The mismatched region in connector 1506 as an inner ring (e.g., an inner hairpin) is shown, but this is for illustrative purposes only, and other configurations may be considered.

[0343] exist Figure 15AThe diagram provides a double-stranded target molecule 1504 and initiators 1502 and 1506. These are linked by 1501 to generate a double-stranded template molecule 1510. The double-stranded template molecule 1510 is further linked by 1503 to additional initiators 1512 and 1514 to generate a double-stranded dual-initiator template molecule 1520. Each link 1501 and 1503 can also alternatively be a tagging reaction.

[0344] exist Figure 15B In the process, double-stranded double-adaptor template molecule 1520 is annealed with primer 1522, wherein primer 1522 anneals with a single-stranded region of one of the adaptors in molecule 1520. Polymerase 1524 extends primer 1522 to produce molecule 1526, which is a copy of molecule 1520. Polymerase 1524 can be a strand displacement polymerase. The extension can be repeated any number of times to amplify double-stranded double-adaptor template molecule 1520. Amplification 1507 can be rolling circle amplification, so the amplification product 1530 is a large nucleic acid molecule containing multiple copies of double-stranded double-adaptor template molecule 1520.

[0345] exist Figure 15C In this process, the multiple copies of the double-stranded double-connector template molecule 1530 undergo conditions sufficient to cleave one or more cleavage sites (sites 1532 and 1534) of 1509. In some cases, cleavage of 1509 involves enzymatic digestion, and cleavage sites 1532 and 1534 include restriction sites. After cleavage / digestion of 1509, the resulting copy molecule 1540 comprises a copy of the double-stranded template molecule 1510. As described herein, these copies retain the double-stranded background of the original target molecule 1504 for further processing purposes. Other methods described herein, such as detecting strand ratios independently of sequencing, can be performed using copy molecule 1540. Copy molecule 1540 can be sequenced to provide sequencing information about the original double-stranded target molecule 1504. In some cases, copy molecule 1540 can be further amplified by any of the amplification methods described herein prior to sequencing. Amplification of copy molecules can generate repeating colonies (optionally immobilized to individual vectors) in different clusters. This repeatability can further reduce sequencing errors by providing comparative data points within the sample (e.g., one repeat sample can be compared with another repeat sample for internal error correction, such as by discarding differentially identified bases or reads between repeat samples). For example, copy molecules can be amplified by ePCR or eRPA to produce amplified vectors.

[0346] Figure 15DAn additional workflow for amplifying a double-stranded template molecule while maintaining forward-reverse strand association (e.g., amplification without losing the double-stranded background) is demonstrated. The insert molecule 1504 (e.g., sample nucleic acid) is circularized (1551) by connecting (e.g., linking) an adaptor comprising a hairpin portion at each end. Primers may anneal to one or both single-stranded hairpin regions of the circularized molecule and extend it via a polymerase. The polymerase may be a strand displacement polymerase. The extension may be repeated any number of times (or for any duration) to amplify (1552) the circularized molecule, as via RCA. The amplification product may be a tandem molecule comprising multiple copies of the double-stranded circularized molecule. The multiple copies may undergo conditions sufficient to cleave (1553) the copies (or portions thereof) from the other copies. For example, cleavage may include enzymatic digestion at or near the hairpin portion or a restriction site. In some cases, the hairpin portion may be cleaved and / or digested. After cleavage / digestion, the resulting copy molecule 1541 comprises a copy retaining the double-stranded background of the original insert molecule 1504. Each of the resulting copy molecules 1541 can be linked to a strand recognition adaptor, including the mismatched portion, as described elsewhere herein, such as linking to each end of the molecule. Other methods described herein, such as detecting strand ratios independently of sequencing, can be performed using copy molecules. Copy molecules can be sequenced to provide sequencing information on the original insert molecule 1504. In some cases, copy molecules can be further amplified prior to sequencing using any of the amplification methods described herein. Amplification of copy molecules can generate repeating colonies (optionally fixed to individual vectors) in different clusters, which can further reduce sequencing errors by providing comparative data points within the sample (e.g., one repeat sample can be compared to another repeat sample for internal error correction, such as by discarding differentially identified bases or reads between repeat samples). For example, copy molecules can be amplified by ePCR or eRPA to produce amplified vectors.

[0347] Figure 15E An alternative workflow is demonstrated for dumbbell-structure-based amplification of double-stranded template molecules using suppressed strand substitution, while maintaining forward-reverse strand association (e.g., amplification without loss of double strand). The insert molecule 1504 (e.g., sample nucleic acid) is circularized (1561) by connecting (e.g., linking) adapters (e.g., adapter 1, adapter 2) that include hairpin portions at each end.

[0348] Procedures 1562-1564 illustrate a first workflow without the use of blocking primers. In this first workflow, a first primer hybridizes to a primer-binding site in one or both of adaptor 1 and adaptor 2 and is extended using a strand displacement polymerase to produce a first tandem product. Second primers each hybridize to a corresponding primer-binding site in the first tandem product and are extended using a strand displacement polymerase to produce a secondary tandem product. A secondary tandem product can replace another second primer that hybridizes with the first tandem product during extension. The composition may contain additional first primers that can bind to and extend the secondary tandem product to produce a tertiary tandem product, which can hybridize with a second primer molecule extended to produce a quaternary tandem product, and so on. After amplification (e.g., RCA, MDA) (1562), amplification products of various sizes can be produced (1563). In some cases, the amplification products may undergo size selection (1564) to isolate products having at least and / or at most the desired copy number. In some cases, the amplification product can be cleaved to produce a single-copy product, such as... Figure 15D The description of 1553 in the text.

[0349] Procedures 1565-1566 illustrate a second workflow in which amplification undergoes suppressed strand substitution. In this second workflow, a first primer may hybridize to a primer binding site in one or both of adaptor 1 and adaptor 2, and be extended using a strand substitution polymerase to produce a first tandem product. In some cases, the hairpin portions of adaptor 1 and adaptor 2 may include different sequences. Second primers may each hybridize to a corresponding primer binding site in the first tandem product and be extended under suppressed strand substitution conditions to produce a secondary amplification product. The secondary amplification product may stop extending when one secondary amplification product reaches the primer binding site where another second primer hybridizes with the first tandem product because strand substitution is suppressed. In some cases, the second primer extension is performed using a non-strand substitution polymerase to prevent strand substitution. In some cases, the second primer includes a reversible crosslinking agent at or near its 5' end for reversible crosslinking at one or more bases (e.g., 3-cyanovinylcarbazole (CNVK), 3-cyanovinylcarbazole modified with D-threonol (CNVD), pyranocarbazole (PCX), pyranocarbazole with D-threonol (PCXD), etc.) to the first tandem product, said crosslinking preventing strand substitution but being reversible, such as by stimulation (e.g., light stimulation), to release the second amplification product; the second primer includes synthetic nucleic acid and / or mimic nucleic acid, such as peptide nucleic acid (PNA), bridged nucleic acid (BNA), or locked nucleic acid (LNA), at or near its 5' end to prevent strand substitution; the second primer includes another blocking agent (e.g., methylated RNA base, etc.) at or near its 5' end to prevent strand substitution; and / or combinations thereof. In some cases, the modification may be included at the portion of the first tandem product corresponding to one of the adaptors, such that the polymerase detaches from the template after reaching this modification (e.g., after generating only one copy), or the modification may be included in one of the adaptors, such that the polymerase detaches from the template after reaching this modification (i.e., after generating only one copy). The template may be re-initiated, such as by nicking the enzyme sequence in the loop. After performing amplification (1565), a second amplification product of substantially the same size (e.g., corresponding to one copy of the template) (1566) may be generated. Advantageously, in the second workflow, the first tandem product (e.g., the long amplicon) remains substantially double-stranded rather than single-stranded during amplification, which prevents the first tandem product from folding itself—this reduces bias due to interference from folded molecules. Furthermore, the second workflow avoids the need to cleave the long tandem product as in the first workflow, which can improve yield. The second amplification product may be separated from the first tandem product, each of which includes a single copy of the insert molecule.

[0350] In the first or second workflow, the resulting copy molecule comprises a copy of Insert 1504 retaining the double-stranded background. In some cases, hairpins may be cut off and / or the hairpin portion may be cut off or digested in the copy molecule. Each of the resulting copy molecules may be connected to a chain-recognition connector including the mismatched portion, as described elsewhere herein, such as by connecting to each end of the molecule (e.g., Figure 15D (In the workflow). Other methods described herein, such as those independent of sequencing detection chain ratios, can be performed using copy molecules. Copy molecules can be sequenced to provide sequencing information on the original insert molecule 1504. In some cases, copy molecules can be further amplified prior to sequencing using any of the amplification methods described herein. Amplification of copy molecules can generate repeating colonies (optionally fixed to individual vectors) in different clusters, which can further reduce sequencing errors by providing comparative data points within the sample (e.g., one repeat sample can be compared with another repeat sample for internal error correction, such as by discarding differentially identified bases or reads between repeat samples). For example, copy molecules can be amplified by ePCR or eRPA to produce amplified vectors.

[0351] Figure 15FAn alternative workflow is demonstrated for dumbbell-structure-based amplification of double-stranded template molecules using random primers, while maintaining forward-reverse strand association (e.g., amplification without loss of double strands). The insert molecule 1504 (e.g., sample nucleic acid) is circularized (1571) by connecting (e.g., linking) adaptors (e.g., adaptor 1, adaptor 2) including hairpin portions at each end. In some cases, random primers, such as random hexamer primers, may contact the circular template, and one or more primers may hybridize to one or more sites in the circular template and be extended using a strand displacement polymerase to produce at least a first tandem product. In some cases, to produce the first tandem product, hybridization and extension with primers including sequences specific to the adaptor sequences of one or both of adaptors 1 and 2 can produce the first tandem product, and random primers may be added and / or present. Random primers, although an example of random hexamers is given here, can have any length (e.g., dimers, trimers, tetramers, pentameres, hexamers, heptameres, octamers, nonameres, etc.). The composition can be provided with dNTPs that include dUTPs instead of dTTPs. Random primers can bind to different sites in the first tandem product and be extended using a chain displacement polymerase to produce secondary amplification products including uracil residues. During amplification (1572), random primers can bind to secondary or higher-order amplification products to produce various amplification products containing uracil residues. After washing, a second set of primers can be introduced into the amplification products, wherein the second primer in the second set includes a sequence that binds to portions of the tandem product corresponding to one or two adaptors. The composition can be provided with dNTPs that include dTTPs instead of dUTPs. The second set of primers can be extended to produce additional amplification products including thymine residues (as opposed to uracil residues). The composition can be heated to enable primer binding. The composition can then undergo enzymatic degradation (1574), such as using a uracil-specific excision reagent (USER) enzyme, to remove amplified material comprising uracil residues (e.g., only one type of amplified strand). In some cases, a non-strand substitution polymerase can be used to extend the second set of primers, in which case the resulting product will be a copy molecule, each comprising a single copy of the insert molecule 1504. In some cases, a strand substitution polymerase can be used to extend the second set of primers, in which case the resulting product can be a tandem polymerase, each comprising multiple copies of the insert molecule 1504. For example, the tandem polymerase may comprise strands containing hairpin copies that can be cleaved and / or the hairpins digested to produce double-stranded copy molecules. It should be understood that substitutes for dUTP (such as other degradable or excitable bases, e.g., ribonucleotides) can be used for the generation of secondary amplification products, or can otherwise be mediated downstream after additional amplification products (e.g., USER substitution).

[0352] The copy molecules may include one or more copies of the inserted 1504 retaining the double-stranded background. In some cases, the hairpin may be cut and / or the hairpin portion may be cut or digested. Each of the resulting copy molecules may be attached to a chain-recognition linker including the mismatched portion, as described elsewhere herein, such as attachment to each end of the molecule (e.g., ...). Figure 15D (In the workflow). Other methods described herein, such as those independent of sequencing detection chain ratios, can be performed using copy molecules. Copy molecules can be sequenced to provide sequencing information on the original insert molecule 1504. In some cases, copy molecules can be further amplified prior to sequencing using any of the amplification methods described herein. Amplification of copy molecules can generate repeating colonies (optionally fixed to individual vectors) in different clusters, which can further reduce sequencing errors by providing comparative data points within the sample (e.g., one repeat sample can be compared with another repeat sample for internal error correction, such as by discarding differentially identified bases or reads between repeat samples). For example, copy molecules can be amplified by ePCR or eRPA to produce amplified vectors.

[0353] exist Figure 15E-15F In any workflow, the initial hairpin connector connected to the insert 1504 to generate the dumbbell loop template may include a chain recognition element. For example, the hairpin connector may include a chain recognition connector, such as... Figure 16A As shown in the diagram. In some cases, the insert may be pre-connected to the chain identification connector at one or both ends, and then connected to the hairpin connector, as per [the diagram / illustration]. Figure 15A The described workflow is advantageous because each chain copy in the resulting copy molecule can be pre-associated with a chain recognition element before cutting and / or digesting the hairpin portion in the tandem. This eliminates the need to reconnect the chain recognition joiners.

[0354] Figure 15G This demonstrates potential sources of error during library preparation when performing flat-end shearing and connecting hairpin connectors, as can be continued... Figure 15A-15F The amplification workflow provided in [the document / platform name]. During shearing, such as after sonication, some templates may include single-stranded overhangs at one or both ends instead of blunt ends, such as [examples of other templates]. Figure 15G The first small figure in the diagram shows that if end repair is incomplete or inefficient, the single-stranded protrusions can fold and self-hybridize to form hairpins on the template molecule, as shown in the first small figure. Figure 15G As shown in the second small figure. After end repair, as shown in the third small figure, and after connecting the hairpin connector, the self-hybrid template can be connected to a hairpin connector on only one end without a self-hybrid hairpin, as shown in the second small figure. Figure 15GThe last small figure in the document illustrates this. When these templates are processed, such as according to the various amplification workflows described in this article, and then sequenced, self-hybridization can lead to the formation of numerous chimeric reads, i.e., complementary strand pairs at the same genomic location. This article provides methods to prevent or reduce this phenomenon.

[0355] In some cases, end repair can be performed at higher temperatures after sonication. Higher temperatures will require more bases to stabilize potential hairpin structures, thus reducing the likelihood of hairpin formation. In some cases, blunting can be performed with exonucleases or endonucleases (such as S1 and mung bean endonucleases), or after pretreatment with exonucleases and endonucleases, to prevent the possibility of hairpin formation. In some cases, hairpin adaptors and / or strand recognition adaptors can be added to the sample DNA without cleavage using transposases (such as Tn5 transposase) in a tagging reaction. This transposase-based hairpin adaptor addition may not require end repair.

[0356] Figure 15A-15F The method shown in the document has many possible variations. For example, such as Figure 16A As shown, in some cases, the target molecule can be linked to two adaptors, namely adaptors 1602 and 1606. Adaptor 1606 can include both a mismatch region (shown here as a ring, but could also be a single base change) and a hairpin region (e.g., a sequence forming a stem-loop structure) (e.g., a combination of adaptors 1506 and 1514). Similarly, another adaptor 1602 can include both a hairpin region and a region capable of further processing (e.g., annealing with a support) (e.g., a combination of adaptors 1502 and 1512). Figure 16B In this configuration, the adaptors 1612 and 1614, which are linked to the target molecule 1604, do not include mismatch regions. In this case, the adaptors including the mismatch regions can then be linked to a copy molecule 1540 obtained from the amplification of the original target molecule 1604, similar to... Figure 15D The workflow within. Figure 16C Possible modifications to the cleavage step 1509 are shown, wherein the cleavage site 1532 is positioned such that all connective regions in the amplified copy molecule 1640 are removed.

[0357] Figure 16D This demonstrates the use of amplified molecule 1530 (e.g., Figure 15BAn exemplary sequencing method used together with (the product generated in the sample). Amplified product 1530 is a large nucleic acid molecule comprising multiple copies of a double-stranded double-connector template molecule 1520. Amplified product 1530 can be ligated to vector 1650 (e.g., beads, planar substrate) for sequencing with primers 1652 and 1654. In some cases, primers 1642 and 1654 can be extended simultaneously. In some cases, primers 1652 and 1654 can be extended sequentially. That is, these primers comprise different sequences and are annealed to different portions of molecule 1530. This can provide sequencing information for paired ends. Any sequencing method provides information for each original strand (e.g., copies of the forward and reverse strands of the original target molecule).

[0358] Figure 17 This demonstrates another workflow for amplifying double-stranded template molecules while maintaining forward-reverse strand association (e.g., amplification without losing the double strand). This workflow is related to... Figure 15A-15F The described workflow differs because it uses only one type of hairpin adaptor, a Y-shaped adaptor instead of a second hairpin adaptor, and the amplification method is not RCA (e.g., it could be PCR, LAMP, etc.). Figure 14 and 16A As shown in -16B, many different adaptor configurations are possible. The double-stranded target molecule 1704 is linked to 1701 (or tagged) to a Y-shaped adaptor 1702 including a mismatch region and an adaptor 1706 having a hairpin region. The template molecule 1710 is provided with a plurality of primers 1712. The plurality of primers 1712 includes at least a first subset 1712b and a second subset 1712a, wherein the first subset 1712b can be annealed with a first region of the adaptor 1702, and the second subset 1712a has sequence complementarity with a different region of the adaptor 1702. The plurality of primers are used to amplify the double-stranded template molecule 1707 to produce multiple copies 1720. The plurality of copies 1720 are exposed to conditions sufficient to cleave one or more cleavage sites in the copies 1720 (e.g., cleavage sites in adaptors 1702 and / or 1706) to provide double-stranded template molecules 1740, wherein these double-stranded template molecules 1740 are copies of the target molecule 1704 and maintain the forward / reverse strand background of the target molecule. Other methods described herein, such as detecting strand ratios independently of sequencing, can be performed using the copy molecule 1740. The copy molecule 1740 can be sequenced to provide sequencing information about the original double-stranded target molecule 1704. In some cases, adaptors including mismatched regions can then be ligated to the copy molecule 1740, similar to... Figure 15D The workflow within the process. In some cases, the copy molecule may be further amplified.

[0359] Figure 18A and18B An alternative workflow for amplifying double-stranded template molecules while maintaining forward-reverse strand association (e.g., amplification without losing the double strand) is demonstrated. The workflow utilizes bridging amplification to provide colonies comprising both forward and reverse strand copies on a vector (e.g., a wafer, beads, etc.). In either case, the double-stranded template molecule may further include mismatched adaptors (e.g., strand recognition adaptors) and / or identification sequences (e.g., UMIs), as described elsewhere herein.

[0360] exist Figure 18A The diagram provides a double-stranded template molecule comprising a first strand 1802a and a second strand 1802b. The double-stranded template molecule further includes a double-stranded linker at each end, wherein each linker comprises a first region 1804a and a second region 1804b that are not chain-complementary to each other. Here, the linkers at each end of the double-stranded template molecule are shown as belonging to the same type. It should be understood that different types of linkers can be used (e.g., in the case where a first type of linker connects to a first end of the template molecule, and a second type of linker connects to the other end of the template molecule). The two strands of the double-stranded template molecule are annealed 1801 with a support, wherein the support comprises primers 1806a and 1806b that are sequence-complementary to each of the linker regions 1804a and 1804b, respectively. Strands 1802a and 1802b are each annealed again with the support 1803 (e.g., each forming a bridge). Using annealed surface primers, each strand is replicated (e.g., surface primers annealed to strand 1802 are extended) to produce original first and second strands 1802 annealed to the vector via surface primers, and strands 1810 covalently linked to the surface primers, where 1810a and 1810b are the inverse complementary sequences of strands 1802a and 1802b, respectively. These strands can be dissociated (e.g., separated into single-stranded molecules). Amplification can be repeated (1809) to produce colonies 1812 comprising copies of each strand 1802 linked to the vector.

[0361] Other methods described in this article, such as sequencing detection strand ratios independent of sequencing, can be performed using colony 1812. Colony 1812 can be sequenced to provide sequencing information about the original strand 1802.

[0362] exist Figure 18B The invention provides a double-stranded template molecule comprising a first strand 1802a and a second strand 1802b. The double-stranded template molecule further comprises a first-type connective tissue, consisting of regions 1804a and 1804b that lack sequence complementarity, and a second-type connective tissue, consisting of a single-stranded hairpin region 1814. Therefore, strands 1802a and 1802b are covalently linked together.

[0363] The double-stranded template molecule is annealed to a support at 1801, wherein the support includes primers that are sequence-complementary to each of the connective regions 1804a and 1804b. The double-stranded template molecule is replicated at 1805 (e.g., by extending surface primers annealed to chain 1802). These chains can be dissociated at 1807 (e.g., separated into single-stranded molecules), resulting in molecule 1820, which is the inverse complementary sequence of chain 1802 and covalently bound to the support, and the original target molecule comprising chain 1802 annealed to the support. Figure 18B This can be repeatedly amplified to produce colonies of individual molecules 1820 covalently linked to the vector. The resulting colonies can be sequenced to provide sequencing information for the original strand 1802.

[0364] Double-stranded molecules used to retain double-stranded information on a single strand.

[0365] In some implementations, such as Figure 20 As described elsewhere in this document, rolling circle amplification can be used to amplify the molecules described in International Patent Publications WO 2023081883A2 and WO 2023164505A2, each of which is incorporated herein by reference in its entirety for all purposes.

[0366] Figure 21A An example workflow for generating a double-stranded molecule is shown, comprising sequences corresponding to the two strands of a double-stranded insert molecule. The single strand can retain information from the two original strands of the double-stranded insert molecule. The double-stranded insert molecule is connected to a linker molecule at each end. In addition to the double-stranded region connected by the insert molecule, each linker molecule has a long single-stranded region capable of hybridizing to form an intramolecular double strand. After an extension reaction, the linker-insertion complex shown at the top yields the long double-stranded molecule shown at the bottom. In the long double-stranded molecule, the original top strand from the insert is now connected to a copy of itself obtained through the extension reaction (the anti-complementary copy of the bottom strand). Similarly, the original bottom strand from the insert is also connected to a copy of itself obtained through the extension reaction (the anti-complementary copy of the top strand). Figure 20 The methods shown, or others described elsewhere in this document, allow one or both chains of a long duplex molecule to be cyclized and subjected to rolling circle amplification.

[0367] Figure 21BAn additional workflow for generating a bistranded molecule is illustrated, the bistranded molecule comprising sequences corresponding to the two strands of a bistranded insert molecule. The single strand can retain information from the two original strands of the bistranded insert molecule. The bistranded insert molecule 2101 is connected to a linker molecule 2102 at each end. The linker molecule includes a bistranded region, a bifurcated (or branched or Y-shaped) mismatched region, and a single-stranded binding region extending from one branch of the bifurcated mismatched region. One end of the bistranded region in the linker molecule 2102 is connected to the insert molecule. The single-stranded binding region of the first linker connected to the first end of the insert can self-hybridize with the single-stranded binding region of the second linker connected to the second end of the insert for intramolecular hybridization to generate a complex 2104. Each of the single-stranded binding regions can be extended to generate the long bistranded molecule shown at the bottom. In the long bistranded molecule, in the first strand 2...

Claims

1. A method for high-accuracy sequencing, the method comprising: (a) Providing an amplified cluster of a plurality of nucleic acid molecules derived from a double-stranded nucleic acid molecule, the double-stranded nucleic acid molecule comprising a first strand and a second strand, wherein a first subset of the plurality of nucleic acid molecules each comprises a copy of a first sequence as at least a portion of a first strand sequence, and wherein a second subset of the plurality of nucleic acid molecules each comprises a second sequence as at least a portion of a reverse complementary copy of a second strand sequence. (b) Collect sequencing signals from the amplified cluster to determine inconsistencies between the first and second sequences at the locus; as well as (c) The locus is excluded from single nucleotide variant (SNV) or single nucleotide polymorphism (SNP) recognition at least in part based on the inconsistency.

2. The method of claim 1, wherein in (b) the sequencing signal from the amplified cluster is collected by hybridizing sequencing primers with the plurality of nucleic acid molecules and simultaneously probing both the first subset and the second subset with the same group of nucleotides.

3. The method of claim 2, wherein the sequencing primers that hybridize with the first subset and the second subset comprise the same sequence.

4. The method of claim 2, wherein the sequencing primers hybridizing with the first subset and the second subset comprise different sequences.

5. The method according to any one of claims 1 to 4, wherein in (b), the same nucleotide mixture probes for the same locus between the first sequence and the second sequence.

6. The method according to claim 5, wherein the length of the same locus is one base position.

7. The method of claim 5, wherein the length of the same locus is at least 2 base positions.

8. The method according to any one of claims 1 to 7, wherein the homonucleotide mixture probes for the same locus between the first sequence and the second sequence.

9. The method of claim 8, wherein the length of the same locus is one base position.

10. The method of claim 8, wherein the length of the same locus is at least 2 base positions.

11. The method according to any one of claims 1 to 10, wherein (b) comprises simultaneously generating sequencing data based on the first subset and the second subset.

12. The method of claim 1, wherein in (b) the sequencing signal from the amplified cluster is collected by: (i) hybridizing a first set of sequencing primers with a first subset of the plurality of nucleic acid molecules and probing the first subset with a first set of nucleotide mixtures, and (ii) hybridizing a second set of sequencing primers with a second subset of the plurality of nucleic acid molecules and probing the second subset with a second set of nucleotide mixtures different from the first set of nucleotide mixtures.

13. The method of claim 11, wherein the first set of sequencing primers and the second set of sequencing primers comprise different sequences.

14. The method of claim 11, wherein the first set of sequencing primers and the second set of sequencing primers have the same sequence.

15. The method according to claims 11 to 14, wherein (b) includes generating sequencing data at different time points based on the first subset and the second subset.

16. The method according to any one of claims 1 to 15, further comprising generating a single sequencing read from the amplified cluster, wherein the single sequencing read is generated by probing both the first subset and the second subset of the plurality of nucleic acid molecules.

17. The method according to any one of claims 1 to 15, further comprising generating two sequencing reads from the amplified cluster, namely a first sequencing read generated by probing a first subset of the plurality of nucleic acid molecules and a second sequencing read generated by probing a second subset of the plurality of nucleic acid molecules.

18. The method according to any one of claims 1 to 17, further comprising generating at least two candidate base recognitions at the locus where the inconsistency occurs.

19. The method of claim 18, further comprising comparing the at least two candidate base identifications with the locus in a reference sequence, and selecting one of the at least two candidate base identifications or a base at the locus in the reference sequence for a shared read.

20. The method according to any one of claims 1 to 19, further comprising using the sequencing signal to generate sequencing reads for the amplified cluster.

21. The method of claim 19, further comprising aligning the sequencing read with the reference sequence.

22. The method of claim 21, further comprising identifying one or more SNVs or SNPs for the sample or subject from which the double-stranded nucleic acid molecule originates.

23. The method according to any one of claims 1 to 22, wherein each of the first subsets further includes a first chain identification element, and wherein each of the second subsets further includes a second chain identification element different from the first chain identification element.

24. The method of claim 23, wherein the first chain identification element and the second chain identification element are different sequences.

25. The method of claim 24, wherein the different sequences are different homopolymer sequences.

26. The method of claim 24, wherein the different sequences are different heteropolymer sequences.

27. The method according to any one of claims 24 to 26, wherein the different sequences have different lengths.

28. The method according to any one of claims 24 to 26, wherein the different sequences have the same length.

29. The method of claim 28, wherein each of the different sequences has a single base length.

30. The method according to any one of claims 24 to 28, wherein each of the different sequences has a length of at least one base.

31. The method according to any one of claims 24 to 30, wherein each of the different sequences has a length of at least 3 bases.

32. The method according to any one of claims 24 to 31, wherein each of the different sequences has a length of at least 5 bases.

33. The method of claim 23, wherein the first strand recognition element and the second strand recognition element are not nucleic acid sequences.

34. The method according to any one of claims 23 to 33, further comprising detecting the presence of one or both of the first chain identification element and the second chain identification element in the amplified cluster.

35. The method of claim 34, wherein the detection comprises sequencing the plurality of nucleic acid molecules to identify the sequence of the first strand recognition element, the second strand recognition element, or both.

36. The method of claim 34, wherein the detection comprises sequencing the plurality of nucleic acid molecules to identify inconsistencies in sequences of the plurality of nucleic acid molecules, including a common portion of the first strand recognition element or the second strand recognition element.

37. The method of claim 34, wherein the detection comprises hybridizing the labeled oligonucleotide probe with the first chain recognition element, the second chain recognition element, or both, and detecting a signal from the labeled oligonucleotide probe.

38. The method of claim 34, further comprising determining the ratio of the first subset to the second subset of the plurality of nucleic acid molecules.

39. The method of claim 38, wherein the ratio is determined by processing the signal strength collected from probing the first chain identification element, the second chain identification element, or both.

40. The method of any one of claims 38 to 39, further comprising generating sequencing reads for the amplified clusters, wherein the sequencing reads are generated at least in part based on the sequencing signal and the ratio.

41. The method according to any one of claims 1 to 39, wherein the amplified cluster is fixed to an individually addressable location on the substrate.

42. The method of claim 41, wherein the substrate comprises at least 1,000,000 individually addressable locations.

43. The method of claim 42, wherein the substrate comprises at least 1,000,000,000 individually addressable locations.

44. The method of claim 43, wherein the substrate comprises at least 5,000,000,000 individually addressable locations.

45. The method of claim 44, wherein the substrate comprises at least 10,000,000,000 individually addressable locations.

46. ​​The method of claim 45, wherein the substrate comprises at least 20,000,000,000 individually addressable locations.

47. The method according to any one of claims 41 to 46, wherein the substrate is substantially planar.

48. The method according to any one of claims 41 to 47, wherein the substrate is textured or patterned.

49. The method according to any one of claims 41 to 47, wherein the substrate is unpatterned.

50. The method according to any one of claims 41 to 49, wherein the substrate comprises an aminosilane layer immobilizing the amplified clusters.

51. The method according to any one of claims 41 to 49, wherein the substrate comprises a surface primer layer immobilizing the amplified cluster.

52. The method according to any one of claims 41 to 51, wherein the plurality of nucleic acid molecules are coupled to beads, the beads being fixed to the individually addressable location on the substrate.

53. The method according to any one of claims 41 to 52, wherein the substrate is rotated during sequencing of the plurality of nucleic acid molecules.

54. The method according to any one of claims 1 to 53, wherein the plurality of nucleic acid molecules are single-stranded molecules.

55. A method for high-accuracy sequencing, the method comprising: (a) Providing an amplified cluster of a plurality of nucleic acid molecules derived from a double-stranded nucleic acid molecule comprising a first strand and a second strand, wherein a first subset of the plurality of nucleic acid molecules each comprises a first strand recognition element and a first sequence as a copy of at least a portion of the first strand sequence, and wherein a second subset of the plurality of nucleic acid molecules each comprises a second strand recognition element and a second sequence as a reverse complementary copy of at least a portion of the second strand sequence, wherein the first strand recognition element and the second strand recognition element are different; as well as (b) Detect the presence of the first chain identification element, the second chain identification element, or both in the amplified cluster.

56. The method of claim 55, wherein the plurality of nucleic acid molecules are immobilized onto the vector.

57. The method according to any one of claims 55 to 56, further comprising subjecting the double-stranded nucleic acid molecules to amplification to generate the plurality of nucleic acid molecules.

58. The method of claim 57, wherein the amplification comprises PCR, emulsion PCR (ePCR), recombinase polymerase amplification (RPA), emulsion RPA (eRPA), rolling circle amplification (RCA), multiple substitution amplification (MDA), bridging amplification, or a combination thereof.

59. The method according to any one of claims 57 to 58, further comprising, prior to the amplification, linking a strand recognition adaptor comprising a pair of mismatched sequences to the double-stranded nucleic acid molecule.

60. The method according to any one of claims 57 to 58, wherein the double-stranded nucleic acid molecule is amplified by: (1) attaching a hairpin adaptor to each end to generate a dumbbell-shaped molecule and subjecting the dumbbell-shaped molecule to rolling circle amplification (RCA) to generate a first amplification product; (2) cleaving or digesting the portion of the first amplification product corresponding to the hairpin adaptor to generate a plurality of copy molecules, each of the plurality of copy molecules comprising a copy of the first strand and a copy of the second strand.

61. The method according to any one of claims 57 to 58, wherein the double-stranded nucleic acid molecule is amplified by: (1) connecting a hairpin adaptor to each end to generate a dumbbell-shaped molecule, and subjecting the dumbbell-shaped molecule to rolling circle amplification (RCA) by contacting the dumbbell-shaped molecule with a plurality of random primers and dNTPs including dUTPs but not dTTPs to generate a first amplification product; (2) hybridizing a second primer to the first amplification product in the presence of dNTPs including dTTPs but not dUTPs to generate a second amplification product; (3) degrading the first amplification product based on uracil residues to separate the second amplification product; and (4) cleaving or digesting the portion of the second amplification product corresponding to the hairpin adaptor to generate a plurality of copy molecules, each of the plurality of copy molecules comprising a copy of the first strand and a copy of the second strand.

62. The method according to any one of claims 57 to 58, wherein the double-stranded nucleic acid molecule is amplified by: (1) connecting a hairpin adaptor to each end to generate a dumbbell-shaped molecule and subjecting the dumbbell-shaped molecule to rolling circle amplification (RCA) to generate a first amplification product; (2) hybridizing a second primer with the first amplification product under repressor strand substitution conditions to generate a plurality of second amplification products, each of the plurality of second amplification products comprising a single copy of the first strand and the second strand.

63. The method according to any one of claims 60 to 61, further comprising linking a strand recognition adaptor comprising a pair of mismatched sequences to each of the plurality of copy molecules to generate the plurality of nucleic acid molecules.

64. The method according to any one of claims 60 to 62, wherein at least one of the hairpin connectors comprises a chain identification connector, the chain identification connector comprising a pair of mismatched sequences, and wherein each of the plurality of copy molecules comprises a copy of the pair of mismatched sequences.

65. A method for detecting an amplified strand on a vector, the method comprising: (a) Connecting a double-stranded template molecule to a vector to generate a template-connected vector, wherein the double-stranded template molecule includes a first strand and a second strand, wherein the double-stranded template molecule includes a linker, wherein the linker includes a mismatch portion, wherein the mismatch portion includes a first mismatch sequence in the first strand and includes a second mismatch sequence that is not complementary to the first mismatch sequence in the second strand; (b) The double-stranded template molecule of the vector to which the template is attached undergoes amplification to produce an amplified vector, the amplified vector comprising a plurality of amplified strands attached thereto. (c) Sequencing the amplified vector to generate sequencing reads; as well as (d) Determine the percentage of amplified strands originating from the first strand in the plurality of amplified strands in the amplified vector, based at least in part on the portion of the sequencing read corresponding to the mismatch portion.

66. A method for detecting an amplified strand on a vector, the method comprising: (a) Connecting a double-stranded template molecule to a carrier to generate a template-connected carrier, wherein the double-stranded template molecule includes a first chain and a second chain, wherein the double-stranded template molecule includes a linker, the linker includes a mismatch portion, wherein the mismatch portion includes a first homopolymer sequence in the first chain and includes a second homopolymer sequence in the second chain that is not complementary to the first homopolymer sequence. (b) The double-stranded template molecule of the vector to which the template is attached undergoes amplification to produce an amplified vector, the amplified vector comprising a plurality of amplified strands attached thereto. (c) Sequencing the amplified vector to generate sequencing reads; as well as (d) Determine the percentage of amplified strands originating from the first strand in the plurality of amplified strands in the amplified vector, based at least in part on the portion of the sequencing read corresponding to the mismatch portion.

67. A reagent kit comprising: A double-chain connector, the double-chain connector including a mismatch portion, wherein the double-chain connector includes a first chain and a second chain, wherein the mismatch portion includes a first mismatch sequence in the first chain and a second mismatch sequence that is not complementary to the first mismatch sequence in the second chain.

68. The kit of claim 67, further comprising a plurality of double-stranded linkers, the plurality of double-stranded linkers including the mismatch portion, wherein the mismatch portion of the plurality of double-stranded linkers is identical.

69. A reagent kit comprising: A double-chain linker, the double-chain linker including a mismatch portion, wherein the double-chain linker includes a first chain and a second chain, wherein the mismatch portion includes a first homopolymer sequence in the first chain and a second homopolymer sequence in the second chain that is not complementary to the first homopolymer sequence.

70. The kit of claim 69, further comprising a plurality of double-stranded linkers, the plurality of double-stranded linkers including the mismatch portion, wherein the mismatch portion of the plurality of double-stranded linkers is identical.

71. A composition comprising: a double-stranded connector, the double-stranded connector including a mismatch portion, wherein the double-stranded connector includes a first chain and a second chain, wherein the mismatch portion includes a first mismatch sequence in the first chain and a second mismatch sequence in the second chain that is not complementary to the first mismatch sequence.

72. The composition of claim 71, further comprising a plurality of double-stranded connectors, each of the plurality of double-stranded connectors including the mismatch portion, wherein the mismatch portion of the plurality of double-stranded connectors is identical.

73. The composition of claim 72, further comprising a plurality of template molecules, wherein the plurality of template molecules comprises a plurality of double-stranded template insert molecules connected to the plurality of double-stranded linkers.

74. The composition of claim 71, further comprising a template molecule, wherein the template molecule comprises a double-stranded template insert molecule connected to the double-stranded linker.

75. The composition according to claim 74, further comprising a carrier.

76. The composition of claim 75, wherein the carrier is connected to the template molecule.

77. The composition according to claim 71, further comprising a carrier.

78. A composition comprising: a double-stranded linker including a mismatch portion, wherein the double-stranded linker includes a first chain and a second chain, wherein the mismatch portion includes a first homopolymer sequence in the first chain and a second homopolymer sequence in the second chain that is not complementary to the first homopolymer sequence.

79. The composition of claim 78, further comprising a plurality of double-stranded connectors, each of the plurality of double-stranded connectors including the mismatch portion, wherein the mismatch portion of the plurality of double-stranded connectors is identical.

80. A method for enrichment, the method comprising: (a) Providing a plurality of balanced vectors, the plurality of balanced vectors comprising a plurality of amplified strands, wherein the plurality of amplified strands are derived from a plurality of double-stranded template nucleic acid molecules, each of the plurality of double-stranded template nucleic acid molecules comprising a mismatch portion, wherein the mismatch portion comprises a first strand and a second strand, the first strand comprising a first mismatch sequence, and the second strand comprising a second mismatch sequence that is not complementary to the first mismatch sequence, such that each of the plurality of amplified strands comprises, at a corresponding portion corresponding to the mismatch portion, an inverse complementary sequence of the first mismatch sequence or the second mismatch sequence; (b) Contacting a plurality of first enrichment molecules with the plurality of amplified strands, wherein the plurality of first enrichment molecules include a first capture sequence complementary to the first mismatch sequence to produce a first set of enriched equilibrium carriers. as well as (c) Contacting a plurality of second enriched molecules with the first set of enriched equilibrium carriers, wherein the plurality of second enriched molecules include a second capture sequence comprising the second mismatch sequence to produce a second set of enriched equilibrium carriers.

81. A method for enrichment, the method comprising: (a) Providing a plurality of balanced vectors, the plurality of balanced vectors comprising a plurality of amplified strands, wherein the plurality of amplified strands are derived from a plurality of double-stranded template nucleic acid molecules, each of the plurality of double-stranded template nucleic acid molecules comprising a mismatch portion, wherein the mismatch portion comprises a first strand and a second strand, the first strand comprising a first mismatch sequence, and the second strand comprising a second mismatch sequence that is not complementary to the first mismatch sequence, such that each of the plurality of amplified strands comprises, at a corresponding portion corresponding to the mismatch portion, an inverse complementary sequence of the first mismatch sequence or the second mismatch sequence; (b) Contacting a plurality of first enrichment molecules with the plurality of amplified chains, wherein the plurality of first enrichment molecules include a first capture sequence comprising the second mismatch sequence to produce a first set of enriched balanced vectors comprising the second mismatch sequence. as well as (c) Contacting a plurality of second enriched molecules with the first set of enriched equilibrium carriers, wherein the plurality of second enriched molecules include a second capture sequence complementary to the first mismatch sequence to produce a second set of enriched equilibrium carriers.

82. A method for enrichment, the method comprising: (a) Providing a plurality of balanced vectors, the plurality of balanced vectors comprising a plurality of amplified strands, wherein the plurality of amplified strands are derived from a plurality of double-stranded template nucleic acid molecules, each of the plurality of double-stranded template nucleic acid molecules comprising a mismatch portion, wherein the mismatch portion comprises a first strand and a second strand, the first strand comprising a first mismatch sequence, and the second strand comprising a second mismatch sequence that is not complementary to the first mismatch sequence, such that each of the plurality of amplified strands comprises, at a corresponding portion corresponding to the mismatch portion, an inverse complementary sequence of the first mismatch sequence or the second mismatch sequence; (b) Contacting a plurality of first enrichment molecules with the plurality of amplified strands, wherein the plurality of first enrichment molecules include a first capture sequence complementary to the first mismatch sequence to capture a subset of the balanced carrier; as well as (c) Remove at least a portion of the subset of the balanced carriers from the plurality of balanced carriers to produce a first set of enriched balanced carriers.

83. A method for enrichment, the method comprising: (a) Providing a plurality of balanced vectors, the plurality of balanced vectors comprising a plurality of amplified strands, wherein the plurality of amplified strands are derived from a plurality of double-stranded template nucleic acid molecules, each of the plurality of double-stranded template nucleic acid molecules comprising a mismatch portion, wherein the mismatch portion comprises a first strand and a second strand, the first strand comprising a first mismatch sequence, and the second strand comprising a second mismatch sequence that is not complementary to the first mismatch sequence, such that each of the plurality of amplified strands comprises, at a corresponding portion corresponding to the mismatch portion, an inverse complementary sequence of the first mismatch sequence or the second mismatch sequence; (b) Contacting a plurality of first enriched molecules with the plurality of amplified strands, wherein the plurality of first enriched molecules include a first capture sequence comprising the second mismatch sequence to capture a subset of the balanced carrier; as well as (c) Remove at least a portion of the subset of the balanced carriers from the plurality of balanced carriers to produce a first set of enriched balanced carriers.

84. A method for generating an amplified vector having a predetermined forward-reverse strand ratio, the method comprising: Make the following items come into contact: (i) a template-linked vector, wherein the template-linked vector comprises a vector linked to a double-stranded template molecule, wherein the double-stranded template molecule comprises a first strand and a second strand, wherein the double-stranded template molecule comprises an adapter, the adapter comprising a mismatch portion, wherein the mismatch portion comprises a first mismatch sequence in the first strand and a second mismatch sequence in the second strand that is not complementary to the first mismatch sequence. (ii) A plurality of forward amplification primers, said plurality of forward amplification primers being at a first predetermined concentration or including primers having a first annealing temperature with the first mismatch sequence, wherein each of said plurality of forward amplification primers includes an inverse complementary sequence of the first mismatch sequence for hybridization with a first amplified strand derived from the first strand, and (iii) A plurality of reverse amplification primers, said plurality of reverse amplification primers being at a second predetermined concentration or comprising primers having a second annealing temperature of a reverse complementary sequence to the second mismatched sequence, wherein each of said plurality of reverse amplification primers comprises the second mismatched sequence for hybridization with a second amplified strand derived from the second strand. To produce the amplified vector comprising multiple amplified chains having the predetermined forward-reverse chain ratio.

85. A method for detecting chains on a carrier, the method comprising: (a) Providing a plurality of balancing vectors immobilized to a substrate, wherein the plurality of balancing vectors comprises a plurality of amplified strands, wherein the plurality of amplified strands are derived from a plurality of double-stranded template nucleic acid molecules, each of the plurality of double-stranded template nucleic acid molecules comprising a mismatch portion, wherein the mismatch portion comprises a first strand and a second strand, the first strand comprising a first mismatch sequence, and the second strand comprising a second mismatch sequence that is not complementary to the first mismatch sequence, such that each of the plurality of amplified strands comprises, at a corresponding portion corresponding to the mismatch portion, an inverse complementary sequence of the first mismatch sequence or the second mismatch sequence; (b) Contacting a plurality of first-type probes with the plurality of amplified strands, wherein each of the plurality of first-type probes includes a first capture sequence complementary to the first mismatch sequence and a first detectable portion; as well as (c) Detecting a first signal, the first signal being derived from the first detectable portion of a subset of the plurality of first-type probes that are combined with a subset of the plurality of amplified strands.

86. The method of claim 85, further comprising: (d) Contacting a plurality of second-type probes with the plurality of amplified strands, wherein each of the plurality of second-type probes includes a second capture sequence comprising the second mismatch sequence and a second detectable portion; as well as (e) Detecting a second signal, the second signal being derived from the second detectable portion of a subset of the plurality of second-type probes that is bound to a second subset of the plurality of amplified strands.

87. A method for amplification, the method comprising: (a) Provides a double-stranded target molecule and a plurality of adaptors, wherein the double-stranded target molecule comprises a first strand and a second strand, and at least one of the plurality of adaptors comprises a hairpin sequence; (b) exposing the double-stranded target molecule and the plurality of adaptors to conditions sufficient to connect the adaptors to each end of the double-stranded target molecule, thereby producing a double-stranded template-adaptor molecule, wherein at least one connected adaptor comprises a hairpin sequence; and (c) The double-stranded template-adaptor molecule is amplified to produce multiple copies of the double-stranded template-adaptor molecule, wherein each copy of the double-stranded template-adaptor molecule includes a copy of the sequence of the first strand and a copy of the sequence of the second strand.

88. The method of claim 87, wherein the amplification comprises rolling circle amplification (RCA), and the plurality of copies of the double-stranded template-adaptor molecule are linked to each other.

89. The method of claim 87, wherein the amplification comprises PCR.

90. The method of claim 87, wherein the amplification comprises loop-mediated isothermal amplification (LAMP).

91. The method of claim 87, wherein the at least one adaptor further comprises a first region and a second region that do not have sequence complementarity, wherein the first region and the second region are located away from the double-stranded template molecule.

92. The method of claim 91, wherein the double-stranded template-linker molecule is connected to the vector, wherein the vector comprises a plurality of primers, wherein a first subset of the primers is sequence complementary to the first region, and a second subset of the primers is sequence complementary to the second region.

93. A sequencing method, the method comprising: (a) Provide a double-stranded template molecule, the double-stranded template molecule comprising a first strand and a second strand having sequence complementarity with each other and at least one linker region including a single-stranded hairpin region; (b) Anneal the primers to the single-stranded hairpin region; (c) Extend the primer to generate a partially single-stranded template molecule, the partially single-stranded template molecule comprising a double-stranded region and a single-stranded region; (d) Processing the partially single-chain template molecule to generate a single-chain template molecule; and (e) Sequencing the single-stranded template molecule.

94. The sequencing method according to claim 93, wherein: Processing (d) the template molecule of the partially single-stranded region includes filtering based on the sequence of the single-stranded region; and Sequencing (e) includes targeted sequencing.

95. The sequencing method according to claim 93, wherein: Treatment (d) of the template molecule of the partially single-chain portion includes methylation conversion of the single-chain region; and Sequencing (e) includes methylation sequencing.

96. A method for sequencing, the method comprising: (a) Providing a balanced construct comprising a mixture of a forward chain and a reverse chain, wherein, respectively, (i) the forward chain comprises a first sequence identical to or serving as the reverse complementary sequence of the first chain of a double-stranded template molecule of the sample, and (ii) the reverse chain comprises a second sequence serving as the reverse complementary sequence of or identical to the methylation-converted sequence of the second chain of the double-stranded template molecule of the sample; and (b) Sequencing of the forward and reverse strands using the following method: i. Hybridize the primers with the forward and reverse strands, respectively. ii. Extend the primer with nucleotides from a nucleotide stream provided according to a repeating sequence, wherein the nucleotide stream comprises nucleotides of a single typical base type, and wherein the repeating sequence comprises three consecutive sequences of thymine bases, cytosine bases, and thymine bases. iii. Detect the flow signal indicating the incorporation or absence of a nucleotide by passing the primer after each corresponding nucleotide flow.

97. The method of claim 96, further comprising using the flow signal detected in (b)(iii) to determine the methylation state of the double-stranded template molecule.

98. The method of claim 96, wherein the forward chain includes a first chain identification element comprising a first homopolymer sequence, and wherein the reverse chain includes a second chain identification element comprising a second homopolymer sequence, wherein the first homopolymer sequence and the second homopolymer sequence comprise different bases.

99. The method of claim 98, further comprising using a subset of the flow signals detected in (b)(iii) corresponding to the first chain identification element and the second chain identification element to determine the forward-reverse ratio of the plurality of copies of the forward chain to the plurality of copies of the reverse chain on the balanced construct.

100. The method of claim 99, further comprising determining the methylation state of the double-stranded template molecule based at least in part on the forward-reverse ratio.

101. The method of claim 22, further comprising detecting minimal residual disease (MRD), tumor fraction, or circulating tumor fraction in the sample or the subject based on the identified one or more SNVs or SNPs.

Citation Information

Patent Citations

  • Methods, devices, and systems for analyte detection and analysis

    US10900078B2

  • Methods and systems for sequence calling

    US11107554B2

  • Methods and systems for analyte detection and analysis

    US20200326327A1

  • Methods for detecting nucleic acid variants

    US20200372971A1

  • Methods for sequencing with single frequency detection

    US20210017593A1