Methods, systems, and compositions for nucleic acid sequencing

By coupling and circularizing nick sites on double-stranded nucleic acid molecules under cell-free conditions, combined with rolling circle amplification technology, the challenge of detecting rare sequence variants in nucleic acid sequencing has been solved, improving the accuracy of detection and the ability to identify early pathological mutations.

CN121472373APending Publication Date: 2026-02-06AXBIO INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511210432.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2018-05-22
Filing Date
2019-05-21
Publication Date
2026-02-06

AI Technical Summary

Technical Problem

Existing nucleic acid sequencing methods are inefficient at detecting rare sequence variants, especially under cell-free conditions where it is difficult to accurately identify the sequence information of double-stranded nucleic acid molecules, which affects the detection and diagnosis of early pathological mutations.

Method used

By providing a double-stranded nucleic acid molecule and a double-stranded linker with a nick site within its strand, coupling and circularization are performed under cell-free conditions, followed by sequencing to generate a growth chain that is sequence-complementary to the nucleic acid molecule, and sequence information is obtained using rolling circle amplification technology.

Benefits of technology

This technology enables efficient and accurate detection of sequence variants of double-stranded nucleic acid molecules under cell-free conditions, improving the detection capability of early pathological mutations and enhancing the accuracy of prenatal diagnosis and tumor cell detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121472373A_ABST
    Figure CN121472373A_ABST
Patent Text Reader

Abstract

The present disclosure describes methods, systems, and compositions for nucleic acid sequencing. The present disclosure provides methods and systems for processing or analyzing nucleic acid molecules. A method for processing or analyzing a double-stranded nucleic acid molecule can include providing a double-stranded nucleic acid molecule and a double-stranded adaptor. The double-stranded adaptor may comprise a nick site within its sense strand or antisense strand. A double-stranded adapter can then be coupled to the double-stranded nucleic acid molecule, and the double-stranded nucleic acid molecule coupled to the double-stranded adapter can be cyclized to produce a cyclized double-stranded nucleic acid molecule.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of Chinese patent application No. 201980049362.0, entitled "Methods, Systems and Compositions for Nucleic Acid Sequencing" (the corresponding PCT application was filed on May 21, 2019, with application number 201980049362.0).

[0002] Cross-referencing

[0003] This application claims the benefit of U.S. Provisional Application No. 62 / 674,706, filed May 22, 2018, which is incorporated herein by reference in its entirety. Background of the Invention

[0004] Nucleic acid sequencing can be used to provide sequence information of nucleic acid samples. Such sequence information may help in the diagnosis or treatment of a subject's condition (e.g., disease). For example, a subject's nucleic acid sequence information can be used to identify, diagnose, or develop treatments for one or more genetic diseases. In another instance, the nucleic acid sequence information of one or more pathogens can guide the treatment of one or more infectious diseases.

[0005] In some cases, methods for nucleic acid sequencing may include creating a nick to produce a circular double-stranded nucleic acid strand and binding a polymerase to the nick. The resulting complex, comprising the polymerase and the circular double-stranded nucleic acid complex, may associate (e.g., conjugate or be adjacent to) a sequencing motif (e.g., nanopore) and may generate a growth strand complementary to at least a portion of the double-stranded nucleic acid (e.g., by rolling circle amplification (RCA)) for sequencing through the sequencing motif. Such methods can be used for whole-genome sequencing or the detection of one or more sequence variants (e.g., mutations) in a nucleic acid library.

[0006] The detection of one or more rare sequence variants (e.g., mutations) can be valuable for healthcare. Detection of rare sequence variants may be important for the early detection of one or more pathological mutations. Detection of one or more cancer-related mutations (e.g., point mutations) in clinical samples can improve the identification of one or more minimal residual diseases in chemotherapy or tumor cell testing in relapsed patients. Furthermore, such mutation detection may be important for assessing exposure to environmental mutagens, monitoring endogenous DNA repair, or studying the accumulation of one or more somatic mutations in aging individuals. Alternatively or additionally, detection of rare sequence variants can enhance prenatal diagnosis and enable the characterization of fetal cells present in maternal blood. Summary of the Invention

[0007] On the one hand, this disclosure provides a method for processing or analyzing double-stranded nucleic acid molecules, comprising: (a) providing (i) the double-stranded nucleic acid molecule and (ii) a double-stranded adapter having a nick site in its sense or antisense strand; (b) coupling the double-stranded adapter to the double-stranded nucleic acid molecule; and (c) cyclizing the double-stranded nucleic acid molecule coupled to the double-stranded adapter to produce a cyclized double-stranded nucleic acid molecule.

[0008] In some embodiments, the double-stranded nucleic acid molecule and the double-stranded adaptor are heterologous to each other. In some embodiments, the double-stranded nucleic acid molecule and the double-stranded adaptor are provided as a cell-free composition. In some embodiments, (b) or (c) is performed under cell-free conditions. In some embodiments, (b) and (c) are performed under cell-free conditions.

[0009] In some embodiments, the coupling includes (i) coupling the sense strand of the double-stranded adapter to the sense strand of the double-stranded nucleic acid molecule, or (ii) coupling the antisense strand of the double-stranded adapter to the antisense strand of the double-stranded nucleic acid molecule. In some embodiments, the coupling includes (i) coupling the sense strand of the double-stranded adapter to the sense strand of the double-stranded nucleic acid molecule, and (ii) coupling the antisense strand of the double-stranded adapter to the antisense strand of the double-stranded nucleic acid molecule.

[0010] In some embodiments, the cleavage site is part of the sense strand of the circularized double-stranded nucleic acid molecule. In some embodiments, the cleavage site is part of the antisense strand of the circularized double-stranded nucleic acid molecule.

[0011] In some embodiments, the method further includes sequencing the double-stranded nucleic acid molecule from the nick site of the double-stranded adaptor. In some embodiments, the sequencing includes (i) extending the double-stranded nucleic acid molecule from the nick site of the double-stranded adaptor to produce a growth strand that is sequence-complementary to at least a portion of the strand of the double-stranded nucleic acid molecule, and (ii) obtaining sequence information of at least a portion of the growth strand. In some embodiments, obtaining the sequence information includes detecting the at least a portion of the growth strand. In some embodiments, the extension reaction includes contacting the double-stranded nucleic acid molecule with the nucleotide coupled to the tag under conditions sufficient to incorporate the nucleotide into the growth strand, and wherein obtaining the sequence information includes detecting the tag. In some embodiments, the method further includes releasing the tag from the nucleotide when incorporating the nucleotide into the growth strand. In some embodiments, the extension reaction is performed without the use of oligonucleotide primers. In some embodiments, the extension reaction includes rolling circle amplification.

[0012] In some embodiments, the sequencing includes (i) cleaving the double-stranded nucleic acid molecule from the cleavage site of the double-stranded connective to cleave at least a portion of the strand of the double-stranded nucleic acid molecule, and (ii) obtaining sequence information of the at least a portion of the strand. In some embodiments, obtaining the sequence information includes detecting the at least a portion of the strand. In some embodiments, the sequencing includes nanopore-based sequencing. In some embodiments, at least a portion of the double-stranded nucleic acid molecule has or is suspected of having one or more sequencing variants compared to at least one reference sequence, and wherein the sequencing is for identifying the presence of the at least a portion of the double-stranded nucleic acid molecule. In some embodiments, the one or more sequencing variants indicate mutations in a gene. In some embodiments, the at least one reference sequence includes a common sequence of at least a portion of the gene.

[0013] In some embodiments, the method further includes, prior to (b), amplifying the double-stranded nucleic acid molecule to produce multiple copies of the double-stranded nucleic acid molecule.

[0014] In some embodiments, the double-stranded nucleic acid molecule contains a recognition sequence, and the method further includes enriching the double-stranded nucleic acid molecules from a random nucleic acid molecule library based at least in part on the recognition sequence. In some embodiments, the enrichment includes generating a selected library of double-stranded nucleic acid molecules, wherein each double-stranded nucleic acid molecule in at least 5% of the selected library contains the recognition sequence. In some embodiments, each double-stranded nucleic acid molecule in at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 95% of the selected library contains the recognition sequence. In some embodiments, the probability of the recognition sequence appearing without any mismatch is 1 x 10^6. 4 At most once per base pair. In some embodiments, the probability of the identified sequence occurring without any mismatch is 1 in every 5 x 10^6 base pairs. 4 7x10 4 1x10 5 1x10 6 1x10 7 1x10 8 1x10 9 1x10 10 1x10 11 Or 1x10 12At most once per base pair. In some embodiments, the recognition sequence comprises at least 5 bases. In some embodiments, the recognition sequence comprises at least 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 35, 40, 45, or 50 bases. In some embodiments, the enrichment comprises (i) binding a recognition portion complementary to the recognition sequence to the double-stranded nucleic acid molecule to form a recognition complex, and (ii) extracting the recognition complex. In some embodiments, the enrichment is performed before (a) or after (b). In some embodiments, the enrichment is performed before (a) and after (b).

[0015] In some embodiments, the double-stranded nucleic acid molecule is derived from or derived from a biological sample of the subject. In some embodiments, the biological sample includes a cell-free biological sample of the subject. In some embodiments, the double-stranded nucleic acid molecule is derived from or derived from a cell-free nucleic acid molecule from the cell-free biological sample. In some embodiments, the cell-free nucleic acid molecule includes circulating tumor nucleic acid molecules or amniotic fluid nucleic acid molecules. In some embodiments, the biological sample includes a tissue sample of the subject. In some embodiments, the double-stranded nucleic acid molecule is derived from or derived from a genomic nucleic acid molecule from the tissue sample. In some embodiments, the tissue sample is derived from: infected tissue, diseased tissue, malignant tissue, calcified tissue, healthy tissue, and combinations thereof. In some embodiments, the tissue sample is derived from malignant tissue comprising a tumor, sarcoma, leukemia, or a derivative thereof.

[0016] In some embodiments, the double-stranded nucleic acid molecule includes DNA, complementary DNA, derivatives thereof, or combinations thereof. In some embodiments, the double-stranded nucleic acid molecule includes RNA.

[0017] In another aspect, this disclosure provides a reaction mixture for processing or analyzing double-stranded nucleic acid molecules, comprising: a composition comprising (i) the double-stranded nucleic acid molecule and (ii) a double-stranded adapter having a nick site within its sense or antisense strand; and at least one enzyme that (i) couples the double-stranded adapter to the double-stranded nucleic acid molecule and (ii) cyclizes the double-stranded nucleic acid molecule coupled with the double-stranded adapter to produce a cyclized double-stranded nucleic acid molecule.

[0018] In some embodiments, the double-stranded nucleic acid molecule and the double-stranded adaptor are heterologous to each other. In some embodiments, the reaction mixture is a cell-free reaction mixture.

[0019] In some embodiments, the at least one enzyme (i) couples the sense strand of the double-stranded adaptor to the sense strand of the double-stranded nucleic acid molecule, or (ii) couples the antisense strand of the double-stranded adaptor to the antisense strand of the double-stranded nucleic acid molecule. In some embodiments, the at least one enzyme (i) couples the sense strand of the double-stranded adaptor to the sense strand of the double-stranded nucleic acid molecule, and (ii) couples the antisense strand of the double-stranded adaptor to the antisense strand of the double-stranded nucleic acid molecule. In some embodiments, the at least one enzyme links the double-stranded adaptor to the double-stranded nucleic acid molecule. In some embodiments, the at least one enzyme includes a ligase, a recombinase, a polymerase, a functional variant thereof, or a combination thereof.

[0020] In some embodiments, the cleavage site is part of the sense strand of the circularized double-stranded nucleic acid molecule. In some embodiments, the cleavage site is part of the antisense strand of the circularized double-stranded nucleic acid molecule.

[0021] In some embodiments, the reaction mixture further comprises at least a second enzyme that performs an extension reaction to produce a growth chain that is sequence complementary to at least a portion of the strands of the double-stranded nucleic acid molecule. In some embodiments, the at least second enzyme produces the growth chain before the double-stranded adapter is coupled to the double-stranded nucleic acid molecule.

[0022] In some embodiments, after the double-stranded adaptor is coupled to the double-stranded nucleic acid molecule, the at least second enzyme performs the extension reaction from the nick site of the double-stranded adaptor to generate the growth chain. In some embodiments, the reaction mixture further comprises at least one nucleotide coupled to a tag, wherein the at least second enzyme incorporates the nucleotide into the growth chain. In some embodiments, when the nucleotide is incorporated into the growth chain, the at least second enzyme releases the tag from the nucleotide. In some embodiments, the at least second enzyme performs the extension reaction without using oligonucleotide primers. In some embodiments, the extension reaction comprises rolling circle amplification. In some embodiments, the at least second enzyme comprises a polymerase.

[0023] In some embodiments, the reaction mixture further comprises at least a third enzyme that performs a cleavage reaction from the cleavage site of the double-stranded linker to cleave at least a portion of the strand of the double-stranded nucleic acid molecule.

[0024] In some embodiments, at least a portion of the double-stranded nucleic acid molecule has, or is suspected of having, one or more variants compared to at least one reference sequence. In some embodiments, the reaction mixture is used to prepare at least one composition for sequencing to identify the presence of the at least a portion of the double-stranded nucleic acid molecule. In some embodiments, the one or more sequencing variants indicate mutations in a gene. In some embodiments, the at least one reference sequence comprises a common sequence of at least a portion of the gene.

[0025] In some embodiments, the double-stranded nucleic acid molecule comprises a recognition sequence. In some embodiments, the reaction mixture further comprises a recognition moiety associated with the recognition sequence to enrich the double-stranded nucleic acid molecules from a random nucleic acid molecule library in the composition, at least partially based on the recognition sequence. In some embodiments, the recognition moiety comprises at least one oligonucleotide complementary to at least the recognition sequence. In some embodiments, the composition comprises a selected library of double-stranded nucleic acid molecules, wherein each double-stranded nucleic acid molecule in at least 5% of the library contains the recognition sequence. In some embodiments, each double-stranded nucleic acid molecule in at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 95% of the selected library contains the recognition sequence. In some embodiments, the probability of the recognition sequence occurring without any mismatch is 1 x 10^6. 4 At most once per base pair. In some embodiments, the probability of the identified sequence occurring without any mismatch is 1 in every 5 x 10^6 base pairs. 4 7x10 4 1x10 5 1x10 6 1x10 7 1x10 8 1x10 9 1x10 10 1x10 11 Or 1x10 12 At most once per base pair. In some embodiments, the identification sequence comprises at least 5 bases. In some embodiments, the identification sequence comprises at least 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 35, 40, 45, or 50 bases.

[0026] In some embodiments, the double-stranded nucleic acid molecule is derived from or derived from a biological sample of the subject. In some embodiments, the biological sample includes a cell-free biological sample of the subject. In some embodiments, the double-stranded nucleic acid molecule is derived from or derived from a cell-free nucleic acid molecule from the cell-free biological sample. In some embodiments, the cell-free nucleic acid molecule includes circulating tumor nucleic acid molecules or amniotic fluid nucleic acid molecules. In some embodiments, the biological sample includes a tissue sample of the subject. In some embodiments, the double-stranded nucleic acid molecule is derived from or derived from a genomic nucleic acid molecule from the tissue sample. In some embodiments, the tissue sample is derived from: infected tissue, diseased tissue, malignant tissue, calcified tissue, healthy tissue, and combinations thereof. In some embodiments, the tissue sample is derived from malignant tissue comprising a tumor, sarcoma, leukemia, or a derivative thereof.

[0027] In some embodiments, the double-stranded nucleic acid molecule includes DNA, complementary DNA, derivatives thereof, or combinations thereof. In some embodiments, the double-stranded nucleic acid molecule includes RNA.

[0028] In various aspects, this disclosure provides libraries of circularized double-stranded nucleic acid molecules comprising (i) double-stranded nucleic acid domains coupled to (ii) double-stranded linker domains containing cleavage sites within the sense or antisense strand of the linker domains, wherein each circularized double-stranded nucleic acid molecule in at least 5% of the library contains a recognition sequence.

[0029] In some embodiments, the nick site is present within the double-stranded connective domain before the double-stranded connective domain is coupled to the double-stranded nucleic acid domain. In some embodiments, the double-stranded nucleic acid domain and the double-stranded connective domain are heterologous to each other. In some embodiments, the library is in a cell-free composition.

[0030] In some embodiments, at least a portion of the circularized double-stranded nucleic acid domain has, or is suspected of having, one or more sequencing variants compared to at least one reference sequence. In some embodiments, the one or more sequencing variants indicate mutations in the gene. In some embodiments, the at least one reference sequence comprises a common sequence of at least a portion of the gene.

[0031] In some embodiments, each circularized double-stranded nucleic acid molecule in at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 95% of the selected library contains the recognition sequence. In some embodiments, the probability of the recognition sequence appearing without any mismatch is 1 x 10^6. 4At most once per base pair. In some embodiments, the probability of the identified sequence occurring without any mismatch is 1 in every 5 x 10^6 base pairs. 4 7x10 4 1x10 5 1x10 6 1x10 7 1x10 8 1x10 9 1x10 10 1x10 11 1x10 12 At most once per base pair. In some embodiments, the identification sequence contains at least 5 specific bases. In some embodiments, the identification sequence contains at least 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 35, 40, 45, or 50 bases.

[0032] In some embodiments, the cleavage site is part of the sense strand of the circularized double-stranded nucleic acid molecule. In some embodiments, the cleavage site is part of the antisense strand of the circularized double-stranded nucleic acid molecule.

[0033] In some embodiments, the circularized double-stranded nucleic acid molecule is derived from or derived from a biological sample of the subject. In some embodiments, the biological sample includes a cell-free biological sample of the subject. In some embodiments, the double-stranded nucleic acid molecule is derived from or derived from a cell-free nucleic acid molecule from the cell-free biological sample. In some embodiments, the cell-free nucleic acid molecule includes circulating tumor nucleic acid molecules or amniotic fluid nucleic acid molecules. In some embodiments, the biological sample includes a tissue sample of the subject. In some embodiments, the double-stranded nucleic acid molecule is derived from or derived from a genomic nucleic acid molecule from the tissue sample. In some embodiments, the tissue sample is derived from: infected tissue, diseased tissue, malignant tissue, calcified tissue, healthy tissue, and combinations thereof. In some embodiments, the tissue sample is derived from malignant tissue comprising a tumor, sarcoma, leukemia, or a derivative thereof.

[0034] In some embodiments, the cyclic double-stranded nucleic acid molecule includes DNA, complementary DNA, derivatives thereof, or combinations thereof. In some embodiments, the cyclic double-stranded nucleic acid molecule includes RNA.

[0035] In various aspects, this disclosure provides methods for processing or analyzing circular nucleic acid molecules, comprising: (a) providing a cell-free composition comprising the circular nucleic acid molecule, the circular nucleic acid molecule comprising (i) a target region and (ii) a nick site at a known distance from the target region; and (b) creating a nick at the nick site of the circular nucleic acid molecule. In some embodiments, (b) is performed under cell-free conditions.

[0036] In some embodiments, the nick site is no more than 100,000 nucleotides away from the target site. In some embodiments, the nick site is no more than 50,000, 10,000, 5,000, 1,000, 500, 400, 300, 200, 100, 90, 80, 70, 60, 50, 40, 30, 25, 20, 15, 10, or 5 nucleotides away from the target site. In some embodiments, the target site comprises up to about 500,000 nucleotides. In some embodiments, the target site comprises up to about 100,000, 50,000, 10,000, 5,000, 1,000, 500, 400, 300, 200, 100, 50, 40, 30, 25, 20, 15, 10, 9, 8, 7, 6, 5, 4, 3, 2, or 1 nucleotide.

[0037] In some embodiments, the nucleic acid molecule comprises a double-stranded nucleic acid molecule. In some embodiments, the cleavage site is part of the sense strand of the circular nucleic acid molecule. In some embodiments, the cleavage site is part of the antisense strand of the circular nucleic acid molecule. In some embodiments, the method further includes determining the cleavage site based at least in part on the position of the target site relative to at least one reference sequence. In some embodiments, the at least one reference sequence comprises a common sequence of at least a portion of a gene. In some embodiments, the cleavage site is endogenous to the circular nucleic acid molecule. In some embodiments, the cleavage site is exogenous to the circular nucleic acid molecule, and the determination includes inserting the exogenous cleavage site into the circular nucleic acid molecule.

[0038] In some embodiments, the nucleic acid molecule further comprises a nicking enzyme binding site specific to the nicking enzyme, and in (b) further comprises providing the nucleic acid molecule with a nicking enzyme under conditions sufficient to cause the nicking enzyme to associate with the nicking enzyme binding site and generate a nick. In some embodiments, the probability of the nicking enzyme binding site appearing without any mismatch is 1 x 10^6 times per ... 4 At most once per base pair. In some embodiments, the probability of the nicking enzyme binding site appearing without any mismatch is 1 in every 5 x 10^6 base pairs. 4 7x104 1x10 5 1x10 6 1x10 7 1x10 8 1x10 9 1x10 10 1x10 11 Or 1x10 12 At most once per base pair. In some embodiments, the nicking enzyme binding site comprises at least 5 bases. In some embodiments, the nicking enzyme binding site comprises at least 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 35, 40, 45, or 50 bases. In some embodiments, the nicking enzyme binding site is endogenous to the nucleic acid molecule. In some embodiments, the nicking enzyme binding site is exogenous to the nucleic acid molecule. In some embodiments, the method further includes inserting the exogenous nicking enzyme binding site into the nucleic acid molecule prior to (b). In some embodiments, the method further includes circularizing the nucleic acid molecule prior to the insertion. In some embodiments, the method further includes circularizing the nucleic acid molecule after the insertion. In some embodiments, the nicking enzyme binding site is no more than 30 nucleotides away from the nicking site. In some embodiments, the nicking enzyme binding site is no more than 25, 20, 19, 18, 17, 16, 15, 14, 13, 12, 11, 10, 9, 8, 7, 6, 5, 4, 3, 2, or 1 nucleotide from the nicking site. In some embodiments, the nicking enzyme binding site comprises the nicking site.

[0039] In some embodiments, the method further includes sequencing the circular nucleic acid molecule. In some embodiments, the sequencing includes (i) extending the circular nucleic acid molecule from the nick to produce a growth strand that is sequence-complementary to at least a portion of the strand of the circular nucleic acid molecule, and (ii) obtaining sequence information of at least a portion of the growth strand. In some embodiments, obtaining the sequence information includes detecting the at least a portion of the growth strand. In some embodiments, the extension reaction includes contacting the circular nucleic acid molecule with the nucleotide coupled to the tag under conditions sufficient to incorporate the nucleotide into the growth strand, and wherein obtaining the sequence information includes detecting the tag. In some embodiments, the method further includes releasing the tag from the nucleotide as it is incorporated into the growth strand. In some embodiments, the extension reaction is performed without the use of oligonucleotide primers. In some embodiments, the extension reaction includes rolling circle amplification.

[0040] In some embodiments, the sequencing includes (i) cleaving the circular nucleic acid molecule through the nick to cut at least a portion of the strand of the double-stranded nucleic acid molecule, and (ii) obtaining sequence information of the at least a portion of the strand. In some embodiments, obtaining the sequence information includes detecting the at least a portion of the strand. In some embodiments, the sequencing includes nanopore-based sequencing. In some embodiments, at least a portion of the target site has or is suspected of having one or more sequencing variants compared to at least one reference sequence, and wherein the sequencing is for identifying the presence of the at least a portion of the target site. In some embodiments, the one or more sequencing variants indicate mutations in a gene. In some embodiments, the at least one reference sequence includes a common sequence of at least a portion of the gene.

[0041] In some embodiments, the method further includes, prior to (a), circularizing at least a linear nucleic acid molecule into a circular nucleic acid molecule. In some embodiments, the at least linear nucleic acid molecule is part of the amplification product. In some embodiments, the method further includes, prior to (b), amplifying the circular nucleic acid molecule to generate multiple copies of the circular nucleic acid molecule.

[0042] In some embodiments, the circular nucleic acid molecule contains a recognition sequence, and the method further includes enriching the circular nucleic acid molecule from a random nucleic acid molecule library at least in part based on the recognition sequence. In some embodiments, the enrichment includes generating a selected library of circular nucleic acid molecules, wherein each circular nucleic acid molecule in at least 5% of the selected library contains the recognition sequence. In some embodiments, each circular nucleic acid molecule in at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 95% of the selected library contains the recognition sequence. In some embodiments, the probability of the recognition sequence appearing without any mismatch is 1 x 10^6. 4 At most once per base pair. In some embodiments, the probability of the identified sequence occurring without any mismatch is 1 in every 5 x 10^6 base pairs. 4 7x10 4 1x10 5 1x10 6 1x10 7 1x10 8 1x10 9 1x10 10 1x10 11 1x10 12At most once per base pair. In some embodiments, the recognition sequence comprises at least 5 bases. In some embodiments, the recognition sequence comprises at least 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 35, 40, 45, or 50 bases. In some embodiments, the enrichment comprises (i) binding a recognition portion complementary to the recognition sequence to the circular nucleic acid molecule to form a recognition complex, and (ii) extracting the recognition complex. In some embodiments, the enrichment is performed before (a) or after (b). In some embodiments, the enrichment is performed before (a) and after (b).

[0043] In some embodiments, the circular nucleic acid molecule is derived from or derived from a biological sample of the subject. In some embodiments, the biological sample includes a cell-free biological sample of the subject. In some embodiments, the circular nucleic acid molecule is derived from or derived from a cell-free nucleic acid molecule from the cell-free biological sample. In some embodiments, the cell-free nucleic acid molecule includes circulating tumor nucleic acid molecules or amniotic fluid nucleic acid molecules. In some embodiments, the biological sample includes a tissue sample of the subject. In some embodiments, the circular nucleic acid molecule is derived from a genomic nucleic acid molecule from the tissue sample. In some embodiments, the tissue sample is derived from: infected tissue, diseased tissue, malignant tissue, calcified tissue, healthy tissue, and combinations thereof. In some embodiments, the tissue sample is derived from malignant tissue comprising a tumor, sarcoma, leukemia, or a derivative thereof.

[0044] In some embodiments, the circular nucleic acid molecule includes DNA, complementary DNA, derivatives thereof, or combinations thereof. In some embodiments, the circular nucleic acid molecule includes RNA.

[0045] In various aspects, this disclosure provides a reaction mixture for processing or analyzing circular nucleic acid molecules, comprising: a cell-free composition containing the circular nucleic acid molecule, the circular nucleic acid molecule including (i) a target site and (ii) a nick site at a known distance from the target site; and at least one enzyme that creates a nick at the nick site of the circular nucleic acid molecule. In some embodiments, the reaction mixture is a cell-free reaction mixture.

[0046] In some embodiments, the nick site is no more than 100,000 nucleotides away from the target site. In some embodiments, the nick site is no more than 50,000, 10,000, 5,000, 1,000, 500, 400, 300, 200, 100, 90, 80, 70, 60, 50, 40, 30, 25, 20, 15, 10, or 5 nucleotides away from the target site.

[0047] In some embodiments, the target site comprises up to about 500,000 nucleotides. In some embodiments, the target site comprises up to about 100,000, 50,000, 10,000, 5,000, 1,000, 500, 400, 300, 200, 100, 50, 40, 30, 25, 20, 15, 10, 9, 8, 7, 6, 5, 4, 3, 2, or 1 nucleotide.

[0048] In some embodiments, the circular nucleic acid molecule comprises a circular double-stranded nucleic acid molecule. In some embodiments, the nick is on the sense strand of the circular double-stranded nucleic acid molecule. In some embodiments, the nick is on the antisense strand of the circular double-stranded nucleic acid molecule. In some embodiments, the nick site is endogenous to the circular nucleic acid molecule. In some embodiments, the nick site is exogenous to the circular nucleic acid molecule.

[0049] In some embodiments, the nucleic acid molecule further comprises an enzyme-binding site specific to the at least one enzyme. In some embodiments, the probability of the enzyme-binding site appearing without any mismatch is in the range of 1 x 10^6. 4 At most once per base pair. In some embodiments, the probability of the enzyme binding site appearing without any mismatch is 1 in every 5 x 10^6 base pairs. 4 7x10 4 1x10 5 1x10 6 1x10 7 1x10 8 1x10 9 1x10 10 1x10 11 1x10 12At most once per base pair. In some embodiments, the enzyme binding site comprises at least 5 bases. In some embodiments, the enzyme binding site comprises at least 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 35, 40, 45, or 50 bases. In some embodiments, the enzyme binding site is endogenous to the circular nucleic acid molecule. In some embodiments, the enzyme binding site is exogenous to the circular nucleic acid molecule. In some embodiments, the enzyme binding site is no more than 30 nucleotides away from the cleavage site. In some embodiments, the enzyme binding site is no more than 25, 20, 19, 18, 17, 16, 15, 14, 13, 12, 11, 10, 9, 8, 7, 6, 5, 4, 3, 2, or 1 nucleotide from the nick site. In some embodiments, the enzyme binding site comprises the nick site.

[0050] In some embodiments, the reaction mixture further comprises at least a second enzyme that performs an extension reaction from the nick to produce a growth chain that is sequence complementary to at least a portion of the circular nucleic acid molecule. In some embodiments, the circular nucleic acid molecule is a circular double-stranded nucleic acid molecule, and wherein the growth chain is sequence complementary to at least a portion of the chain of the circular double-stranded nucleic acid molecule. In some embodiments, the reaction mixture further comprises at least one nucleotide coupled to a tag, wherein the at least second enzyme incorporates the nucleotide into the growth chain. In some embodiments, when the nucleotide is incorporated into the growth chain, the at least second enzyme releases the tag from the nucleotide. In some embodiments, the at least second enzyme performs the extension reaction without using oligonucleotide primers. In some embodiments, the extension reaction comprises rolling circle amplification. In some embodiments, the at least second enzyme comprises a polymerase.

[0051] In some embodiments, the reaction mixture further comprises at least a third enzyme that performs a cleavage reaction from the nick to cleave at least a portion of the circular nucleic acid molecule. In some embodiments, the circular nucleic acid molecule is a circular double-stranded nucleic acid molecule, and wherein the at least third enzyme cleaves at least a portion of the strands of the circular double-stranded nucleic acid molecule.

[0052] In some embodiments, at least a portion of the circular nucleic acid molecule has, or is suspected of having, one or more variants compared to at least one reference sequence. In some embodiments, the reaction mixture is used to prepare at least one composition for sequencing to identify the presence of the at least a portion of the circular nucleic acid molecule. In some embodiments, the one or more sequencing variants indicate mutations in a gene. In some embodiments, the at least one reference sequence comprises a common sequence of at least a portion of the gene.

[0053] In some embodiments, the circular nucleic acid molecule comprises a recognition sequence. In some embodiments, the reaction mixture further comprises a recognition moiety associated with the recognition sequence to enrich the circular nucleic acid molecule from a random nucleic acid molecule library in the composition, at least partially based on the recognition sequence. In some embodiments, the recognition moiety comprises at least one oligonucleotide complementary to at least the recognition sequence. In some embodiments, the composition comprises a selected library of circular nucleic acid molecules, wherein each circular nucleic acid molecule in at least 5% of the library contains the recognition sequence. In some embodiments, each circular nucleic acid molecule in at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 95% of the selected library contains the recognition sequence. In some embodiments, the probability of the recognition sequence appearing without any mismatch is 1 x 10^6 / ... 4 At most once per base pair. In some embodiments, the probability of the identified sequence occurring without any mismatch is 1 in every 5 x 10^6 base pairs. 4 7x10 4 1x10 5 1x10 6 1x10 7 1x10 8 1x10 9 1x10 10 1x10 11 1x10 12 At most once per base pair. In some embodiments, the identification sequence comprises at least 5 bases. In some embodiments, the identification sequence comprises at least 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 35, 40, 45, or 50 bases.

[0054] In some embodiments, the circular nucleic acid molecule is derived from or derived from a biological sample of the subject. In some embodiments, the biological sample includes a cell-free biological sample of the subject. In some embodiments, the circular nucleic acid molecule is derived from or derived from a cell-free nucleic acid molecule from the cell-free biological sample. In some embodiments, the cell-free nucleic acid molecule includes circulating tumor nucleic acid molecules or amniotic fluid nucleic acid molecules. In some embodiments, the biological sample includes a tissue sample of the subject. In some embodiments, the circular nucleic acid molecule is derived from a genomic nucleic acid molecule from the tissue sample. In some embodiments, the tissue sample is derived from: infected tissue, diseased tissue, malignant tissue, calcified tissue, healthy tissue, and combinations thereof. In some embodiments, the tissue sample is derived from malignant tissue comprising a tumor, sarcoma, leukemia, or a derivative thereof.

[0055] In some embodiments, the circular nucleic acid molecule includes DNA, complementary DNA, derivatives thereof, or combinations thereof. In some embodiments, the circular nucleic acid molecule includes RNA.

[0056] In various respects, this disclosure provides cell-free libraries of circular nucleic acid molecules, wherein each individual in at least 5% of the library contains (i) a target site and (ii) a nick at a known distance from the target site.

[0057] In some embodiments, the circular nucleic acid molecules of each individual in at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 95% of the library contain (i) the target site and (ii) the nick at the known distance from the target site. In some embodiments, the circular nucleic acid molecules of the individuals also contain a recognition sequence. In some embodiments, the probability of the recognition sequence appearing without any mismatch is 1 x 10^6 times per ... 4 At most once per base pair. In some embodiments, the probability of the identified sequence occurring without any mismatch is 1 in every 5 x 10^6 base pairs. 4 7x10 4 1x10 5 1x10 6 1x10 7 1x10 8 1x10 9 1x10 10 1x10 11 1x10 12At most once per base pair. In some embodiments, the identification sequence comprises at least 5 bases. In some embodiments, the identification sequence comprises at least 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 35, 40, 45, or 50 bases.

[0058] In some embodiments, (i) the first target site of the first individual nucleic acid molecule and (ii) the second target site of the second individual nucleic acid molecule are the same. In some embodiments, (i) the first known distance between the first nick and the first target site of the first individual nucleic acid molecule and (ii) the second known distance between the second nick and the second target site of the second individual nucleic acid molecule are the same. In some embodiments, (i) the first known distance between the first nick and the first target site of the first individual nucleic acid molecule and (ii) the second known distance between the second nick and the second target site of the second individual nucleic acid molecule are different.

[0059] In some embodiments, (i) the first target site of the first individual nucleic acid molecule and (ii) the second target site of the second individual nucleic acid molecule are different. In some embodiments, (i) the first known distance between the first nick and the first target site of the first individual nucleic acid molecule and (ii) the second known distance between the second nick and the second target site of the second individual nucleic acid molecule are the same. In some embodiments, (i) the first known distance between the first nick and the first target site of the first individual nucleic acid molecule and (ii) the second known distance between the second nick and the second target site of the second individual nucleic acid molecule are different.

[0060] In some embodiments, the nick site is no more than 100,000 nucleotides away from the target site. In some embodiments, the nick site is no more than 50,000, 10,000, 5,000, 1,000, 500, 400, 300, 200, 100, 90, 80, 70, 60, 50, 40, 30, 25, 20, 15, 10, or 5 nucleotides away from the target site.

[0061] In some embodiments, the target site comprises up to about 500,000 nucleotides. In some embodiments, the target site comprises up to about 100,000, 50,000, 10,000, 5,000, 1,000, 500, 400, 300, 200, 100, 50, 40, 30, 25, 20, 15, 10, 9, 8, 7, 6, 5, 4, 3, 2, or 1 nucleotide.

[0062] In some embodiments, at least a portion of the circular nucleic acid molecule has, or is suspected of having, one or more variants compared to at least one reference sequence. In some embodiments, the one or more sequencing variants indicate mutations in the gene. In some embodiments, the at least one reference sequence comprises a common sequence of at least a portion of the gene.

[0063] In some embodiments, the circular nucleic acid molecule comprises a circular double-stranded nucleic acid molecule. In some embodiments, the nick is made on the sense strand of the circular double-stranded nucleic acid molecule. In some embodiments, the nick is made on the antisense strand of the circular double-stranded nucleic acid molecule.

[0064] In some embodiments, the circular nucleic acid molecule is derived from or derived from a biological sample of the subject. In some embodiments, the biological sample includes a cell-free biological sample of the subject. In some embodiments, the circular nucleic acid molecule is derived from or derived from a cell-free nucleic acid molecule from the cell-free biological sample. In some embodiments, the cell-free nucleic acid molecule includes circulating tumor nucleic acid molecules or amniotic fluid nucleic acid molecules. In some embodiments, the biological sample includes a tissue sample of the subject. In some embodiments, the circular nucleic acid molecule is derived from or derived from a genomic nucleic acid molecule from the tissue sample. In some embodiments, the tissue sample is derived from: infected tissue, diseased tissue, malignant tissue, calcified tissue, healthy tissue, and combinations thereof. In some embodiments, the tissue sample is derived from malignant tissue comprising a tumor, sarcoma, leukemia, or a derivative thereof.

[0065] In some embodiments, the circular nucleic acid molecule includes DNA, complementary DNA, derivatives thereof, or combinations thereof. In some embodiments, the circular nucleic acid molecule includes RNA.

[0066] In various aspects, this disclosure provides methods for processing or analyzing nucleic acid molecules, comprising: (a) providing the nucleic acid molecule adjacent to a nanopore and contacting the nucleic acid molecule with a tagged nucleotide under conditions sufficient to incorporate a nucleotide into a nucleic acid chain complementary to at least a portion of the nucleic acid molecule, wherein at least a portion of the tag is located within the nanopore when the nucleotide is incorporated into the nucleic acid chain; (b) detecting one or more signals indicating impedance or impedance changes in the nanopore when at least a portion of the tag is within the nanopore; and (c) using the one or more signals to identify the nucleotide incorporated into the nucleic acid chain.

[0067] In some embodiments, the one or more signals are current or voltage. In some embodiments, the method further includes measuring current or changes therein when at least a portion of the tag is located within the nanopore. In some embodiments, the one or more signals are not tunneling currents. In some embodiments, the current is not a Faraday current. In some embodiments, the nanopore is part of a circuit containing a tunnel junction. In some embodiments, the nanopore includes multiple electrodes, and wherein (c) includes using the multiple electrodes to detect the one or more signals.

[0068] In some embodiments, the nanopores include protein nanopores or solid nanopores.

[0069] In some embodiments, the method further includes, in (a), releasing the tag from the nucleotide upon incorporation into the nucleic acid chain. In some embodiments, the nucleic acid molecule comprises a circular nucleic acid molecule. In some embodiments, the method further includes, prior to (b), performing rolling circle amplification (RCA) on the nucleic acid molecule to generate the nucleic acid chain. In some embodiments, the circular nucleic acid molecule is a circular double-stranded nucleic acid molecule. In some embodiments, the incorporation is performed without the use of oligonucleotide primers. In some embodiments, the provision includes coupling at least one enzyme that performs incorporation into (i) at least a portion of the nanopore or (ii) a membrane having the nanopore. In some embodiments, the membrane is a lipid bilayer. In some embodiments, the membrane is a solid membrane. In some embodiments, the coupling includes conjugating the at least one enzyme to the nanopore or the membrane.

[0070] In various aspects, this disclosure provides a system for processing or analyzing nucleic acid molecules, comprising: a nanopore configured to (i) receive at least a portion of a tag when a tagged nucleotide is incorporated into a nucleic acid chain, wherein the nucleic acid chain is complementary to at least a portion of the nucleic acid molecule; and (ii) detect one or more signals indicative of impedance or impedance changes in the nanopore when the at least a portion of the tag is within the nanopore, wherein the one or more signals can be used to identify the nucleotide incorporated into the nucleic acid chain.

[0071] In some embodiments, the one or more signals are current or voltage. In some embodiments, the nanopore is configured to measure current or changes thereof when at least a portion of the tag is located within the nanopore. In some embodiments, the one or more signals are not tunneling currents. In some embodiments, the current is not a Faraday current. In some embodiments, the nanopore is part of a circuit containing a tunnel junction. In some embodiments, the nanopore includes a plurality of electrodes configured to detect the one or more signals.

[0072] In some embodiments, the nanopores include protein nanopores or solid nanopores.

[0073] In some embodiments, when the nucleotide is incorporated into the nucleic acid chain, at least a portion of the tag is released from the nucleotide. In some embodiments, the system further comprises at least one enzyme configured to perform the incorporation. In some embodiments, the at least one enzyme is coupled to (i) at least a portion of the nanopore or (ii) a membrane having nanopores. In some embodiments, the at least one enzyme is conjugated to (i) at least a portion of the nanopore or (ii) a membrane having nanopores.

[0074] In some embodiments, the membrane is a lipid bilayer. In some embodiments, the membrane is a solid membrane. In some embodiments, the nanopore or membrane is configured to bind at least a portion of the at least one enzyme. In some embodiments, the at least one enzyme is configured to bind to at least a portion of the nanopore or at least a portion of the membrane.

[0075] In various aspects, this disclosure provides a method for sequencing multiple polynucleotides, comprising: circularizing individual polynucleotides to provide multiple cyclic polynucleotides; nicking one strand of the cyclic polynucleotides using a nicking enzyme to provide a nicking site on each of the cyclic polynucleotides; binding a polymerase to the nicking site; and sequencing the cyclic polynucleotides.

[0076] In some embodiments, the polynucleotide comprises a double-stranded polynucleotide. In some embodiments, the polynucleotide comprises a single-stranded polynucleotide, and the method includes adding a primer sequence to the adaptor region of the single-stranded polynucleotide.

[0077] In some embodiments of any of the subject methods, the method further includes amplifying the plurality of polynucleotides prior to cyclizing individual polynucleotides. In some embodiments of any of the subject methods, the polynucleotides comprise DNA, cDNA, ctDNA, or any combination thereof. In some embodiments of any of the subject methods, cyclizing the plurality of polynucleotides comprises reacting the plurality of polynucleotides with a ligase. In some embodiments of any of the subject methods, the cyclic polynucleotide comprises a cyclic double-stranded polynucleotide, and the cleavage comprises a cleavage of the inner strand of the cyclic double-stranded polynucleotide. In some embodiments of any of the subject methods, the cyclic polynucleotide comprises a cyclic double-stranded polynucleotide, and the cleavage comprises a cleavage of the outer strand of the cyclic double-stranded polynucleotide.

[0078] In some embodiments of any of the subject-specific methods, the nicking enzyme comprises an sgRNA-CRISPR Cas9n (Cas9 D10A) nicking enzyme complex. In some embodiments, the sgRNA comprises a nucleotide sequence complementary to the target nucleotide sequence. In some embodiments of any of the subject-specific methods, the nicking enzyme comprises a Cas9n (Cas9 D10A) nicking enzyme.

[0079] In some embodiments of any of the subject methods, the polymerase comprises a adapter. In some embodiments, the method further includes binding the adapter-containing polymerase to the nick site. In some embodiments of any of the subject methods, the polymerase comprises an adapter and a protein nanopore bound to the adapter.

[0080] In some embodiments of any of the subject-matter methods, the method further includes cleaving the polynucleotides to provide a targeted polynucleotide fragment prior to cyclizing the plurality of polynucleotides. In some embodiments of any of the subject-matter methods, the cleavage includes binding a biotinylated sgRNA-CRISPER Cas9n complex to the polynucleotide. In some embodiments, the biotinylated sgRNA comprises a nucleotide sequence complementary to the target polynucleotide sequence. In some embodiments of any of the subject-matter methods, the method further includes enriching the targeted polynucleotide fragment.

[0081] In some embodiments of any of the subject methods, the polymerase exhibits strong chain displacement activity.

[0082] In some embodiments of any of the subject methods, the sequencing includes sequencing the circular polynucleotide more than once while binding to the nanopore. In some embodiments of any of the subject methods, the sequencing includes forward sequencing and reverse sequencing. In some embodiments of any of the subject methods, the polynucleotide comprises a double-stranded polynucleotide, and the uncut polynucleotide chain contains a template for amplification and sequencing. In some embodiments of any of the subject methods, the polynucleotide comprises genomic DNA, cDNA, cell-free DNA, ctDNA, or a combination thereof. In some embodiments of any of the subject methods, the sequencing includes rolling circle amplification and transcription. In some embodiments of any of the subject methods, the sequencing includes whole-genome sequencing. In some embodiments of any of the subject methods, the sequencing includes targeted sequencing. In some embodiments of any of the subject methods, the targeted sequencing includes identifying sequence variants.

[0083] Another aspect of this disclosure provides a non-transitory computer-readable medium containing machine-executable code that, when executed by one or more computer processors, implements any of the methods described above or elsewhere herein.

[0084] Another aspect of this disclosure provides a system comprising one or more computer processors and computer memory coupled thereto. The computer memory contains machine-executable code that, when executed by the one or more computer processors, implements any of the methods described above or elsewhere herein.

[0085] Other aspects and advantages of this disclosure will become apparent to those skilled in the art from the following detailed description of illustrative embodiments shown and described only therein. As will be appreciated, this disclosure is capable of other different implementations, and certain details thereof can be modified in various obvious ways without departing from this disclosure. Therefore, the drawings and detailed descriptions should be considered illustrative in nature and not restrictive.

[0086] Incorporation

[0087] All publications, patents, and patent applications mentioned in this specification are incorporated herein by reference as if specifically and individually indicated to be incorporated by reference for each publication, patent, or patent application. Where any publication, patent, or patent application incorporated by reference conflicts with the disclosure contained in this specification, this specification is intended to supersede and / or give precedence to any such contradictory material. Attached Figure Description

[0088] The novel features of the invention are set forth in the appended claims. A better understanding of the features and advantages of the invention will be obtained by referring to the following detailed description of illustrative embodiments in which the principles of the invention are utilized, along with the accompanying drawings (also referred to herein as “Figures”):

[0089] Figure 1A An example method for providing circular nucleic acids with nicks is illustrated schematically;

[0090] Figure 1B and 1C An example method for isolating or enriching linear or circular nucleic acids containing recognition sites is illustrated schematically.

[0091] Figure 1D An example method is illustrated using one or more uracil-specific enzymes to provide a circular nucleic acid containing a nick at a specific location within the circular nucleic acid;

[0092] Figure 2A and 2B An example method for creating a cut at a known distance from the target site within a circular nucleic acid is illustrated schematically;

[0093] Figure 2C and 2D An example method for isolating or enriching circular nucleic acids containing recognition sites and target sites is illustrated schematically.

[0094] Figure 3 An example method for double-stranded nucleic acid sequencing is illustrated schematically;

[0095] Figure 4A An example method for targeted sequencing using nanopore sequencing is illustrated schematically;

[0096] Figure 4B An example method for genome sequencing using nanopore sequencing is illustrated schematically;

[0097] Figures 5A to 5D An exemplary nanopore sequencing system for obtaining sequence information from one or more nucleic acid samples is schematically illustrated.

[0098] Figure 6 A computer system is shown that is programmed or otherwise configured to implement the methods provided herein.

[0099] Figure 7A An example of a gel electrophoresis image of a sample containing multiple circularized single-stranded nucleic acids is shown. Figure 7B An example of a fluorescence image of an RCA product from a circularized single-stranded nucleic acid is shown;

[0100] Figure 7CAn example of a gel electrophoresis image of a sample containing multiple circularized double-stranded nucleic acids is shown. Figure 7D Examples of fluorescence images of RCA products from circularized double-stranded nucleic acids are shown; and

[0101] Figure 8 Examples of gel electrophoresis images of circular double-stranded nucleic acids complexed with (i) wild-type polymerase and (ii) mutant polymerase are shown. Detailed Implementation

[0102] Although various embodiments of the invention have been shown and described herein, it will be readily understood by those skilled in the art that such embodiments are provided merely by way of example. Many modifications, alterations, and substitutions will occur to those skilled in the art without departing from the invention. It should be understood that various alternatives to the embodiments of the invention described herein may be employed.

[0103] As used in the specification and claims, unless the context clearly indicates otherwise, the singular forms “an,” “a,” and “the” may include plural referents. For example, the term “transmembrane receptor” may include multiple transmembrane receptors.

[0104] As used herein, the terms "about" or "approximately" can refer to an acceptable range of error for a particular value as determined by one of ordinary skill in the art, which will depend in part on how the value is measured or determined, i.e., the limitations of the measurement system. For example, according to practice in the art, "about" can mean within or greater than one standard deviation. Alternatively, "about" can mean a range of up to 20%, up to 10%, up to 5%, or up to 1% of a given value. Or, particularly for biological systems or processes, the term can mean within an order of magnitude of a value, preferably within 5 times, more preferably within 2 times. Where a particular value is described in this application and claims, unless otherwise stated, the term "about" should be assumed to mean within an acceptable range of error for that particular value.

[0105] As used herein, the term "cell" generally refers to a biological cell or cell derivative. A cell can be the basic structural, functional, and / or biological unit of a living organism. A cell can originate from any organism having one or more cells. Some non-limiting examples include: prokaryotic cells, eukaryotic cells, bacterial cells, archaea cells, cells of unicellular eukaryotes, protozoan cells, cells from plants (e.g., cells from plant crops, fruits, vegetables, grains, soybeans, corn, maize, wheat, seeds, tomatoes, rice, cassava, sugarcane, squash, hay, potatoes, cotton, hemp, tobacco, flowering plants, conifers, gymnosperms, ferns, lycophytes, hornworts, bryophytes, mosses), and algal cells (e.g., *Botryococcus braunii*, *Chlamydomonas reinhardtii*, *Nannochloropsis gaditana*, *Chlorella pyrenoidosa*, and *Sargassum patens*). Cells derived from various organisms include: seaweed (e.g., kelp), fungal cells (e.g., yeast cells, mushroom cells), animal cells, cells from invertebrates (e.g., fruit flies, cnidarians, echinoderms, nematodes, etc.), cells from vertebrates (e.g., fish, amphibians, reptiles, birds, mammals), and cells from mammals (e.g., pigs, cattle, goats, sheep, rodents, rats, mice, non-human primates, humans, etc.). Sometimes cells do not originate from natural organisms (e.g., cells can be synthetic, sometimes called artificial cells).

[0106] Where used interchangeably herein, the terms “nucleotide,” “nucleobase,” and “base” generally refer to a base-sugar-phosphate combination. A nucleotide may include synthetic nucleotides. A nucleotide may include synthetic nucleotide analogs. A nucleotide may be a monomeric unit of a nucleic acid sequence (e.g., deoxyribonucleic acid (DNA) and ribonucleic acid (RNA)). The term nucleotide may include ribonucleoside triphosphates (adenosine triphosphate (ATP), uridine triphosphate (UTP), cytosine triphosphate (CTP), guanosine triphosphate (GTP), uridine triphosphate (UTP)) and deoxyribonucleoside triphosphates (such as dATP, dCTP, dITP, dUTP, dGTP, dTTP) or derivatives thereof. Such derivatives may include, for example, [αS]dATP, 7-denitro-dGTP, and 7-denitro-dATP, as well as nucleotide derivatives that confer nuclease resistance on nucleic acid molecules containing them. As used herein, the term nucleotide generally refers to dideoxynucleotide triphosphates (ddNTPs) and their derivatives. Illustrative examples of dideoxynucleotide triphosphates may include, but are not limited to, ddATP, ddCTP, ddGTP, ddITP, and ddTTP. Nucleotides may be unlabeled or detectably labeled. Labeling may also be performed using quantum dots. Detectable labeling may include, for example, radioactive isotopes, fluorescent labels, chemiluminescent labels, bioluminescent labels, and enzyme labels. Fluorescent labeling of nucleotides may include, but is not limited to, fluorescein, 5-carboxyfluorescein (FAM), 2′7′-dimethoxy-4′5-dichloro-6-carboxyfluorescein (JOE), rhodamine, 6-carboxyrhodamine (R6G), N,N,N′,N′-tetramethyl-6-carboxyrhodamine (TAMRA), 6-carboxy-X-rhodamine (ROX), 4-(4′-dimethylaminophenylazo)benzoic acid (DABCYL), cascade blue, Oregon green, Texas red, cyanine, and 5-(2′-aminoethyl)aminonaphthalene-1-sulfonic acid (EDANS).Specific examples of fluorescently labeled nucleotides may include [R6G]dUTP, [TAMRA]dUTP, [R110]dCTP, [R6G]dCTP, [TAMRA]dCTP, [JOE]ddATP, [R6G]ddATP, [FAM]ddCTP, [R110]ddCTP, [TAMRA]ddGTP, [ROX]ddTTP, [dR6G]ddATP, [dR110]ddCTP, [dTAMRA]ddGTP, and [dROX]ddTTP, all available from Perkin Elmer, Foster City, Calif.; and FluoroLink deoxynucleotides, FluoroLink Cy3-dCTP, FluoroLink Cy5-dCTP, FluoroLink FluorX-dCTP, FluoroLink Cy3-dUTP, and FluoroLink, all available from Amersham, Arlington Heights, Ill. Cy5-dUTP; luciferin-15-dATP, luciferin-12-dUTP, tetramethyl-rhodamine-6-dUTP, IR770-9-dATP, luciferin-12-ddUTP, luciferin-12-UTP, and luciferin-15-2'-dATP, available from Boehringer Mannheim, Indianapolis, Ind.; and luciferin-15-2'-dATP, available from Molecular Nucleotides used for chromosome labeling, obtained by Probes, Eugene, and Oreg, include BODIPY-FL-14-UTP, BODIPY-FL-4-UTP, BODIPY-TMR-14-UTP, BODIPY-TMR-14-dUTP, BODIPY-TR-14-UTP, BODIPY-TR-14-dUTP, Cascade Blue-7-UTP, Cascade Blue-7-dUTP, Fluorescein-12-UTP, Fluorescein-12-dUTP, Oregon Green 488-5-dUTP, Rhodamine Green-5-UTP, Rhodamine Green-5-dUTP, Tetramethylrhodamine-6-UTP, Tetramethylrhodamine-6-dUTP, Texas Red-5-UTP, Texas Red-5-dUTP, and Texas Red-12-dUTP. Nucleotides can also be labeled or marked through chemical modifications. The chemically modified single nucleotide can be a biotin-dNTP. Some non-limiting examples of biotinylated dNTPs can include biotin-dATP (e.g., bio-N6-ddATP, biotin-14-dATP), biotin-dCTP (e.g., biotin-11-dCTP, biotin-14-dCTP), and biotin-dUTP (e.g., biotin-11-dUTP, biotin-16-dUTP, biotin-20-dUTP).

[0107] Naturally occurring nucleotides guanine, cytosine, adenine, thymine, and uracil can be abbreviated as G, C, A, T, and U, respectively. Nucleotides can include any subunit that can be incorporated into a growing nucleic acid chain. Such subunits can be A, C, G, T, or U, or any other subunit that is specific to one or more complementary A, C, G, T, or U, or complementary to a purine (i.e., A or G or its variants) or a pyrimidine (i.e., C, T, or U or its variants). Subunits can break down individual nucleic acid bases or base sets (e.g., AA, TA, AT, GC, CG, CT, TC, GT, GT, TG, AC, CA, or their uracil counterparts).

[0108] The terms “polynucleotide,” “oligonucleotide,” “oligomer,” and “nucleic acid” are used interchangeably herein and generally refer to polymeric forms of nucleotides of any length, whether deoxyribonucleotides, ribonucleotides, or analogs, and whether single-stranded, double-stranded, or multi-stranded. Polynucleotides can be exogenous or endogenous to the cell. Polynucleotides can exist in cell-free environments. Polynucleotides can be genes or fragments thereof. Polynucleotides can be DNA. Polynucleotides can be RNA. Polynucleotides can have any three-dimensional structure and can perform any known or unknown function. Polynucleotides can include one or more analogs (e.g., modified backbones, sugars, or nucleobases). Where modifications are present, the nucleotide structure can be modified before or after polymer assembly. Some non-limiting examples of analogues include: 5-bromouracil, peptide nucleic acids, heteronucleic acids, morpholino derivatives, locked nucleic acids, glycol nucleic acids, threonine nucleic acids, dideoxynucleotides, cordycepin, 7-denitro-GTP, fluorophores (e.g., rhodamine or fluorescein linked to sugars), thiol-containing nucleotides, biotin-linked nucleotides, fluorescent base analogues, CpG islands, methyl-7-guanosine, methylated nucleotides, inosine, thiouridine, pseudouridine, dihydrouridine nucleoside, brassinoside, and woyoside. Non-restricted examples of polynucleotides include coding or non-coding regions of genes or gene fragments, loci defined by linkage analysis, exons, introns, messenger RNA (mRNA), transfer RNA (tRNA), ribosomal RNA (rRNA), short interfering RNA (siRNA), short hairpin RNA (shRNA), microRNA (miRNA), ribozymes, complementary DNA (cDNA, such as double-stranded cDNA (dd-cDNA) or single-stranded cDNA (ss-cDNA)), circulating tumor DNA (ctDNA), damaged DNA, recombinant polynucleotides, branched polynucleotides, plasmids, vectors, isolated DNA of any sequence, isolated RNA of any sequence, cell-free polynucleotides (including cell-free DNA (cfDNA) and cell-free RNA (cfRNA)), nucleic acid probes (e.g., fluorescence in situ hybridization (FISH) probes), and primers. The sequence of a nucleotide may be interrupted by non-nucleotide components. Polynucleotides may contain one or more modified nucleotides, such as methylated nucleotides and nucleotide analogs. In the presence of modifications, the nucleotide structure may be modified before or after polymer assembly. The sequence of a nucleotide may be interrupted by non-nucleotide components. Polynucleotides can be further modified after polymerization, such as by conjugation with labeled components.

[0109] As used herein, the term "gene" generally refers to nucleic acids (e.g., DNA, such as genomic DNA and cDNA) that encode RNA transcripts and their corresponding nucleotide sequences. The term "gene," as used herein with respect to genomic DNA, may include intercalated non-coding regions as well as regulatory regions, and may include 5' and 3' ends. In some uses, the term includes the transcribed sequence, including 5' and 3' untranslated regions (5'-UTR and 3'-UTR), exons, and introns. In some genes, the transcribed region will contain an "open reading frame" encoding a polypeptide. In some uses of the term, "gene" includes only the coding sequence necessary to encode a polypeptide (e.g., "open reading frame" or "coding region"). Genes may not encode polypeptides, such as ribosomal RNA (rRNA) genes and transfer RNA (tRNA) genes. The term "gene" can include not only transcribed sequences but also non-transcribed regions, including upstream and downstream regulatory regions, enhancers, and promoters. A gene can be an "endogenous gene" or a natural gene located naturally in the genome of an organism. A gene can be an "exogenous gene" or a non-natural gene. Non-natural genes can be genes that are not normally found in the host organism but are introduced into the host organism through gene transfer (e.g., transgenes). Non-natural genes can be naturally occurring nucleic acid or polypeptide sequences that contain mutations, insertions, and / or deletions (e.g., non-natural sequences).

[0110] As used herein, the term "mutation" generally refers to an alteration in the nucleotide sequence of a normally conserved nucleic acid sequence, resulting in a mutant sequence that differs from the normal (unaltered) or wild-type sequence. Prior to sequencing, the location (e.g., relative to a gene or sample polynucleotide) and sequence of the mutation may be unknown. Alternatively, the location (e.g., relative to a gene or sample polynucleotide) and sequence of the mutation may be known prior to sequencing, in which case sequencing can be performed to detect the presence of the mutation in the sample polynucleotide. Mutations can include base pair substitutions (e.g., single nucleotide substitutions) and frameshift mutations. Frameshift mutations may require the insertion or deletion of one or more nucleotide pairs.

[0111] As used herein, the term "probe" generally refers to a nucleotide or polynucleotide labeled with a marker (e.g., a fluorescent marker) that can be used to detect or identify its corresponding target nucleotide or polynucleotide in a hybridization reaction by hybridization with a corresponding target sequence. Where used interchangeably herein, the terms "nucleotide probe," "nucleotide tag," and "labeled nucleotide" generally refer to a probe having a single nucleotide. Where used interchangeably herein, the terms "polynucleotide probe," "polynucleotide tag," and "labeled polynucleotide" generally refer to a probe having a polynucleotide. A polynucleotide probe may be labeled with at least one marker (e.g., one marker per nucleotide of the polynucleotide probe). The probe may hybridize with one or more target nucleotides or polynucleotides. A polynucleotide probe may be perfectly complementary to one or more target polynucleotides in a sample, or may contain one or more nucleotides that are not complementary (i.e., mismatched) to one or more target polynucleotides in a sample.

[0112] Where used interchangeably herein, the terms “complementary,” “complementary sequence,” “complementary,” and “complementarity” generally refer to a sequence that is completely complementary to and hybridizes with a given sequence. A sequence that hybridizes with a given nucleic acid is called the “complementary sequence” or “reverse complementary sequence” of a given molecule, provided that its base sequence in a given region can bind complementary to the base sequence of its binding partner, such that, for example, AT, AU, GC, and GU base pairs are formed. Typically, a first sequence that hybridizes with a second sequence can hybridize specifically or selectively with the second sequence, such that, during hybridization, hybridization with the second sequence or a group of second sequences is preferred (e.g., more thermodynamically stable under given conditions, such as stringent conditions commonly used in the art) compared to hybridization with a non-target sequence. Typically, hybridizable sequences share a degree of sequence complementarity, such as between 25% and 100%, across their respective lengths, including at least 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, and 100% sequence complementarity. Their respective lengths may include regions having at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50, or 50 or more nucleotides. For purposes such as assessing the percentage of complementarity, sequence identity can be measured using any suitable alignment algorithm, including but not limited to the Needleman-Wunsch algorithm (see, for example, the EMBOSS Needle aligner available at www.ebi.ac.uk / Tools / psa / emboss_needle / nucleotide.html, optionally using the default settings), the BLAST algorithm (see, for example, the BLAST alignment tool available at blast.ncbi.nlm.nih.gov / Blast.cgi, optionally using the default settings), or the Smith-Waterman algorithm (see, for example, the EMBOSS Water aligner available at www.ebi.ac.uk / Tools / psa / emboss_water / nucleotide.html, optionally using the default settings). Optimal alignment can be evaluated using any suitable parameters of the chosen algorithm, including the default parameters.

[0113] Complementarity can be complete or substantial / sufficient. Complete complementarity between two nucleic acids means that the two nucleic acids can form a double helix, in which each base in the double helix binds to a complementary base via Watson-Crick pairing. Substantial or sufficient complementarity means that the sequence in one strand is not completely and / or imperfectly complementary to the sequence in the opposing strand, but under a set of hybridization conditions (e.g., salt concentration and temperature), sufficient binding occurs between the bases on the two strands to form a stable hybrid complex. Such conditions can be predicted by using the sequence and standard mathematical calculations, or by empirically determining the Tm using conventional methods.

[0114] As used herein, the term "hybridization" generally refers to a reaction in which one or more polynucleotides react to form a complex that is stable by hydrogen bonds between the bases of the nucleotide residues. Hydrogen bonds can occur through Watson-Crick base pairing, Hoogstein binding, or in any other sequence-specific manner based on base complementarity. The complex can comprise two strands forming a double-stranded structure, three or more strands forming a multi-stranded complex, a self-hybridizing strand, or any combination thereof. Hybridization reactions can constitute a step in a broader process, such as initiating PCR or the enzymatic cleavage of polynucleotides by endonucleases. A second sequence complementary to the first sequence can be referred to as the "complementary sequence" of the first sequence. The term "hybridizable" applied to polynucleotides generally refers to the ability of the polynucleotide to form a complex that is stable by hydrogen bonds between the bases of the nucleotide residues during the hybridization reaction.

[0115] As used herein, the term "target polynucleotide" generally refers to a polynucleotide in a nucleic acid molecule or population of nucleic acid molecules that has a target sequence. This requires identifying the presence, number, and / or variation of the nucleotide sequence, or one or more of them. The term "target sequence" generally refers to a nucleic acid sequence on a single strand of a nucleic acid. A target sequence can be a portion of a gene, a regulatory sequence, genomic DNA, cDNA, ctDNA, RNA (including mRNA, miRNA, rRNA), or others. A target sequence can be a target sequence derived from a sample or a secondary target, such as a product of an amplification reaction. A target polynucleotide can be a portion of a gene (or a fragment thereof) containing one or more mutations.

[0116] As used herein, the term “target site” generally refers to a polynucleotide sequence containing a target polynucleotide (or target nucleotide). The target polynucleotide (or target nucleotide) of a target site can be one or more sequence variants. Examples of one or more sequence variants may include single nucleotide variations, insertions or deletions of one or more nucleotides (e.g., continuous or discontinuous nucleotides), copy number variations (CNVs) containing one or more repeats of one or more nucleotides (e.g., CNVs with an average size of at least 1, 5, 10, 50, 100, 150, 200 or more kilobases (kb); CNVs with an average size of at most 200, 150, 100, 50, 10, 5, 1 or less kb), and microsatellite instability (MSI). Target sites may include at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 40, 50, 100, 200, 300, 400, 500, 1,000, 5,000, 10,000, 50,000, 100,000, 500,000 or more nucleotides. Target sites may include up to 500,000, 100,000, 50,000, 10,000, 5,000, 1,000, 500, 400, 300, 200, 100, 50, 40, 30, 25, 20, 15, 10, 9, 8, 7, 6, 5, 4, 3, 2 or 1 nucleotide. The target site may contain at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 40, 50, 100, 200, 300, 400, 500, 1,000, 5,000, 10,000, 50,000, 100,000, 500,000, or 500,000 or more nucleotides than the target polynucleotide. The target site may contain up to 500,000, 100,000, 50,000, 10,000, 5,000, 1,000, 500, 400, 300, 200, 100, 50, 40, 30, 25, 20, 15, 10, 9, 8, 7, 6, 5, 4, 3, 2, or 1 nucleotide more than the target polynucleotide. In some instances, the target site may be the target polynucleotide.

[0117] As used herein, the term "strict conditions" generally refers to one or more hybridization conditions under which nucleic acids complementary to the target sequence hybridize primarily with the target sequence and substantially with no hybridization with non-target sequences. Strict conditions can be sequence-dependent and can vary depending on many factors. In some cases, the longer the sequence, the higher the temperature at which the sequence can specifically hybridize with its target sequence.

[0118] As used herein, the term "recognition moiety" generally refers to a molecule (e.g., a small molecule, polynucleotide, protein, its variants, or combinations thereof) capable of interacting with a nucleic acid sequence, i.e., a "recognition sequence" or "recognition site," such as a desired (or target) nucleic acid sequence. The recognition moiety may contain a domain (e.g., a component containing that domain) capable of binding (e.g., hybridizing) to the recognition sequence. Such a domain may contain one or more amino acids, one or more nucleotides, their variants, or combinations thereof. Alternatively or additionally, the recognition moiety may associate (e.g., bind) with a secondary molecule containing such a domain. In some instances, the recognition moiety may contain a nucleic acid molecule capable of hybridizing with the recognition sequence. In some instances, the recognition moiety may contain a component exhibiting specific biological activity, including but not limited to one or more activities of nucleases (e.g., double-stranded nucleases), nickases, transcription activators, transcription repressors, nucleic acid methyltransferases, nucleic acid demethyltransferases, and recombinases. The identification sequence may include at least 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50 or more nucleotides. The identification sequence may include up to 50, 45, 40, 35, 30, 25, 24, 23, 22, 21, 20, 19, 18, 17, 16, 15, 14, 13, 12, 11, 10, 9, 8, 7, 6, 5, 4, 3 or fewer nucleotides.

[0119] The recognition portion can be used to isolate a desired molecule containing a desired nucleic acid sequence from a plurality of molecules (e.g., a plurality of nucleic acid molecules). The recognition portion can be used to enrich a desired molecule containing a desired nucleic acid in a composition or reaction mixture. In some instances, the recognition portion can be captured by a capture system (e.g., magnetic beads) via one or more interactions (e.g., avidin-biotin binding, magnetic binding, etc.). In one instance, the recognition portion may contain biotin, which can be complexed with streptavidin magnetic beads for separation or enrichment.

[0120] Examples of recognition moieties may include CRISPR-associated (Cas) systems (e.g., Cas proteins, including catalytically active or inactive Cas polypeptides); zinc finger nucleases (ZFNs); transcription activator-like effector nucleases (TALENs); large-scale nucleases; RNA-binding proteins (RBPs); Cas RNA-binding proteins; recombinases; flipases; transposases; Argonaute (Ago) proteins (e.g., prokaryotic Argonaute (pAgo), archaea Argonaute (aAgo), and eukaryotic Argonaute (eAgo)); variants thereof; and combinations thereof. The recognition moieties may include polynucleotides (e.g., sequences at least 2, 3, 4, 5, 6, 7, 8, 9, 10, or more nucleotides long) that can be captured by a capture system through one or more interactions. For example, a biotin-tagged polynucleotide sequence may be captured by one or more avidin-functionalized magnetic beads. At least a portion of the polynucleotide may share complementarity with the recognition sequence of the target nucleic acid molecule.

[0121] As used herein, the term "nicking enzyme" generally refers to a molecule (e.g., an enzyme) that cleaves one strand of a double-stranded nucleic acid molecule (i.e., a "nick" in the double-stranded molecule). A nicking enzyme can be a nuclease that cleaves only a single strand of DNA, either due to its natural function or because it has been engineered (e.g., modified by mutation and / or deletion of one or more nucleotides) to cleave only a single strand of DNA. A nicking enzyme can be an enzyme that produces a nick (e.g., a restriction endonuclease, a nicking endonuclease, etc.). A nicking enzyme binds to a nick site in a double-stranded nucleic acid molecule to create a nick (or gap) in one strand of the molecule. The nick can be created within the nick site. Alternatively, the nick can be created near the nick site. In some cases, the nicking enzyme can bind to a nicking enzyme binding site adjacent to the nick site. The length of the nick can be at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more nucleotides. The length of the nick can be up to 10, 9, 8, 7, 6, 5, 4, 3, 2, or 1 nucleotide. Examples of nickases can include the Cas system (e.g., Cas nickases such as Cas9n), N.AlwI, Nb.BbvCl, Nt.BbvCl, Nb.BsmI, Nt.BsmAI, Nt.BspQl, Nb.BsrDI, Nt.BstNBI, Nb.BstsCI, Nt.CviPII, Nb.BpulOI, Nt.Bpu lOI, and Nt,Bst9I, their variants, and combinations thereof. In some instances, the nucleic acid molecule (e.g., a self-complementary double-stranded nucleic acid molecule or a single-stranded nucleic acid molecule) may contain at least one nick site that already includes at least one nick.

[0122] Where used interchangeably herein, the terms “CRISPR-associated system,” “Cas system,” and “Cas complex” generally refer to a two-component nucleoprotein complex having a guide RNA (gRNA) and a Cas polypeptide or protein (e.g., a Cas endonuclease, its catalytic or non-catalytic derivatives, etc.) or other protein with endonuclease activity. The term “CRISPR” refers to clusters of regularly spaced short palindromic repeats and their associated system. At least a portion of the gRNA may be complementary to at least a portion of the target region. The target region may contain a “pre-intercalation region sequence” and a “pre-intercalation region sequence adjacent motif” (PAM), and both domains may be essential for the nuclease activity (e.g., cleavage) of the Cas polypeptide. The pre-intercalation region sequence may be referred to as the target site (or genomic target site). The gRNA may pair (or hybridize) with the opposite strand of the pre-intercalation region sequence (binding site) to guide the Cas polypeptide to the target region. The PAM site generally refers to a short sequence recognized by the Cas polypeptide and, in some cases, may be essential for nuclease (or cleavage) activity. The nucleotide sequence and number of PAM sites may vary depending on the type of Cas enzyme.

[0123] Cas peptides may contain nuclease (or nickase) activity, and gRNA may interact with the Cas peptide to direct its nuclease (or nickase) activity to a desired target region. Alternatively, Cas peptides may be non-catalytic and may not contain nuclease activity. Non-catalytic Cas peptides may be referred to as dead or inactivated Cas (dCas).

[0124] Cas proteins can include proteins from or derived from CRISPR-related type I, II, or III systems, and may possess RNA-directed polynucleotide-binding or nuclease activity. Examples of suitable Cas proteins include Cas3, Cas4, Cas5, Cas5e (or CasD), Cas6, Cas6e, Cas6f, Cas7, Cas8a1, Cas8a2, Cas8b, Cas8c, Cas9 (also known as Csn1 and Csxl2), Cas10, Cas10d, CasF, CasG, CasH, Csy1, Csy2, Csy3, Cse1 (or CasA), Cse2 (or CasB), Cse3 (or CasE), and Cs... e4 (or CasC), Csc1, Csc2, Csa5, Csn2, Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csz1, Csx15, Csf1, Csf2, Csf3, Csf4, Cu1966, their homologs and their modified forms (e.g., catalytic or non-catalytic). In some cases, Cas proteins may include proteins of the CRISPR-associated type V or type VI system or proteins derived from the CRISPR-associated type V or type VI system, such as Cpf1 (or Cas12a), C2c1 (or Cas12b), C2c2, their homologs, and their modified forms (e.g., catalytic or non-catalytic).

[0125] While some examples in this article involve Cas proteins, other proteins with endonuclease activity can be used. For example, such other proteins may not be Cas proteins, but can be configured for use with gRNA.

[0126] Cas peptides or proteins can be engineered to modify nuclease activity into nicking enzyme activity. For example, the aspartic-to-alanine substitution (D10A) in the RuvC I catalytic domain of Cas9 from *S. pyogenes* can convert Cas9 from a two-strand nuclease into a single-strand nicking enzyme, Cas9n. Cas9n nicking enzyme mutants can introduce gRNA-targeted single-strand breaks into the DNA, instead of the double-strand breaks produced by wild-type Cas peptides. Other examples of mutations that make Cas9 a nicking enzyme include H840A, N854A, and N863A.

[0127] As used herein, the term "guide RNA (gRNA)" generally refers to an RNA molecule (e.g., DNA or gene) that can bind to a Cas polypeptide and help target the Cas polypeptide to a specific location within a target nucleic acid region. The complementarity between the gRNA and the specific location within the target nucleic acid region can be at least 50%, 60%, 75%, 80%, 85%, 90%, 95%, 97.5%, 99%, or higher. The guide RNA may comprise a CRISPR RNA (crRNA) segment and a trans-activated crRNA (tracrRNA) segment. As used interchangeably herein, the terms "crRNA" and "crRNA segment" generally refer to an RNA molecule or a portion thereof that includes a polynucleotide-targeting guide sequence, a stem sequence, and an optional 5'-protrusion sequence. As used interchangeably herein, the terms "tracrRNA" and "tracrRNA segment" generally refer to an RNA molecule or a portion thereof that includes a protein-binding segment (e.g., a protein-binding segment capable of interacting with a CRISPR-associated protein such as Cas9). In some cases, the guide RNA can be a single guide RNA (sgRNA), where the crRNA and tracrRNA segments reside within the same RNA molecule. gRNA can contain one or more peptide nucleic acids.

[0128] crRNA can contain at least 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40 or more RNA bases. crRNA can contain up to 40, 35, 30, 25, 24, 23, 22, 21, 20, 19, 18, 17, 16, 15 or fewer RNA bases. The target nucleic acid sequence of Cas system gRNA can contain at least 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40 or more DNA bases. The target nucleic acid sequence of Cas system gRNA can contain up to 40, 35, 30, 25, 24, 23, 22, 21, 20, 19, 18, 17, 16, 15 or fewer DNA bases. crRNA sequences can be selected to target any target sequence. Target sequences can be sequences within the cellular genome. Target sequences can include unique sequences within the target genome.

[0129] As used herein, the term "polymerase" generally refers to an enzyme (e.g., natural or synthetic) capable of catalyzing polymerization reactions. Examples of polymerases may include nucleic acid polymerases (e.g., DNA polymerase or RNA polymerase), transcriptases, and ligases. Polymerases can be polymerization enzymes. The term "DNA polymerase" generally refers to an enzyme capable of catalyzing the polymerization reaction of DNA.

[0130] As used herein, the term "linked polymerase" generally refers to a polymerase, such as a DNA polymerase, that couples to (e.g., fuses to) a linker. The linker may be able to couple to (e.g., bind to or conjugate to) another entity (e.g., a nanopore, such as a protein nanopore or a solid nanopore).

[0131] As used interchangeably herein, the terms “sequence variant” and “sequencing variant” generally refer to any sequence variation relative to one or more reference sequences. Typically, for a given population of individuals for which reference sequences are provided, sequence variants occur less frequently than the reference sequences. For example, a particular bacterial genus may have a common reference sequence for its 16S rRNA gene, but individual species within that genus may have one or more sequence variants within the gene or a portion of the gene that can be used to identify that species within the bacterial population. As a further example, sequences from multiple individuals of the same species or multiple sequencing reads from the same individual can produce a common sequence when optimally aligned, and sequence variants relating to that common sequence can be used to identify mutants in the population that indicate dangerous contamination. Generally, a “common sequence” refers to a nucleotide sequence that reflects the most common base selection at each position in the sequence, in which the associated nucleic acid series has undergone in-depth mathematical and / or sequence analysis, such as optimal alignment according to any of a variety of sequence alignment algorithms. A reference sequence is a single, known reference sequence, such as the genome sequence of a single individual. A reference sequence can be a shared sequence formed by aligning multiple known sequences, such as the genomic sequences of multiple individuals used as a reference population or multiple sequencing reads of a polynucleotide from the same individual. A reference sequence can also be a shared sequence formed by optimally aligning sequences from the sample being analyzed, whereby a sequence variant represents a variation relative to a corresponding sequence in the same sample. Sequence variants can occur at low frequencies in the population (also known as “rare” sequence variants). For example, sequence variants can occur at frequencies less than or equal to 5%, 4%, 3%, 2%, 1.5%, 1%, 0.75%, 0.5%, 0.25%, 0.1%, 0.075%, 0.05%, 0.04%, 0.03%, 0.02%, 0.01%, 0.005%, 0.001%, or lower. Sequence variants can occur at frequencies less than or equal to 0.1%.

[0132] Sequence variants can be any variation relative to a reference sequence. Sequence variants can include alterations, insertions, or deletions of single or multiple nucleotides (such as 2, 3, 4, 5, 6, 7, 8, 9, 10, or more nucleotides). When a sequence variant contains two or more nucleotide differences, these different nucleotides can be continuous or discontinuous. Examples of sequence variant types include single nucleotide polymorphisms (SNPs), deletion / insertion polymorphisms (DIPs), copy number variants (CNVs), short tandem repeats (STRs), simple sequence repeats (SSRs), variable number tandem repeats (VNTRs), amplified fragment length polymorphisms (AFLPs), reverse transcriptase-based insertion polymorphisms, sequence-specific amplification polymorphisms, and epigenetic marker differences that can be detected as sequence variants (e.g., methylation differences).

[0133] As used herein, the term "sequencing" generally refers to the procedure for determining the order in which nucleotides appear in a target nucleotide sequence. Sequencing methods can include high-throughput sequencing, such as next-generation sequencing (NGS). Sequencing can be whole-genome sequencing or targeted sequencing. Sequencing can be single-molecule sequencing or massively parallel sequencing. Next-generation sequencing methods can yield millions of sequences in a single run. In one instance, sequencing can be performed using one or more nanopore sequencing methods, such as synthesis sequencing, ligation sequencing, or cleavage sequencing.

[0134] As used herein, the term "nanopore" generally refers to a pore, channel, or pathway formed or otherwise provided in a membrane. The membrane can be an organic membrane, such as a lipid bilayer, or a synthetic membrane, such as a membrane formed from polymeric materials such as protein nanopores. The membrane can be a solid membrane (e.g., a silicon substrate). Nanopores may be positioned adjacent to or close to electrodes of a sensing circuit (e.g., a complementary metal-oxide-semiconductor (CMOS) or field-effect transistor (FET) circuit). Nanopores may be part of a sensing circuit. Nanopores may have a characteristic width or diameter, for example, from about 0.1 nanometers (nm) to 1000 nm. Nanopores can be biological nanopores, solid nanopores, hybrid biological-solid nanopores, variants thereof, or combinations thereof. Examples of bio-nanopores include, but are not limited to, OmpG from *Escherichia coli*, *Salmonella*, *Shigella*, and *Pseudomonas*, as well as α-hemolysin from *Staphylococcus aureus*, MspA from *Mycobacterium smegmatis*, their functional variants, or combinations thereof. Sequencing may include forward sequencing and / or reverse sequencing. Examples of solid-state nanopores include, but are not limited to, silicon nitride, silicon oxide, graphene, molybdenum sulfide, their functional variants, or combinations thereof. Solid-state nanopores can be fabricated by high-energy beam fabrication, imprinting (e.g., nanoimprinting), laser ablation, chemical etching, plasma etching (e.g., oxygen plasma etching), etc.

[0135] As used herein, the term "nanopore sequencing complex" generally refers to a nanopore linked or coupled to an enzyme, such as a polymerase, which in turn associates with a polymer, such as a polynucleotide template. Nanopore sequencing complexes can be located in membranes, such as lipid bilayers, where their function is to identify polymeric components, such as nucleotides or amino acids.

[0136] Where used interchangeably herein, the terms “nanopore sequencing” and “nanopore-based sequencing” generally refer to methods for determining the sequence of polynucleotides using nanopores. In some cases, the sequence of polynucleotides can be determined in a template-dependent manner. In some cases, the methods, systems, or compositions disclosed herein may not be limited to any particular nanopore sequencing method, system, or apparatus.

[0137] As used herein, the term "barcode" generally refers to a known nucleic acid sequence that allows for the identification of certain characteristics of a polynucleotide associated with the barcode (e.g., a polynucleotide containing at least a portion of the barcode or a polynucleotide complementary to at least a portion of the barcode). In some instances, the characteristics of the polynucleotide to be identified may be the sample from which the polynucleotide originated. The length of the barcode may be at least about 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 or more nucleotides. The length of the barcode may be at most 20, 19, 18, 17, 16, 15, 14, 13, 12, 11, 10, 9, 8, 7, 6, 5, 4, 3 or 2 nucleotides. A barcode associated with a polynucleotide from a first sample may be different (e.g., a different sequence and / or a different length) from a barcode associated with a polynucleotide from a second sample different from the first sample. In this context, identifying the barcode in the corresponding polynucleotide can help identify the sample origin of one or more polynucleotides. Therefore, different samples with different barcodes can be analyzed together (e.g., in batches) (e.g., sequenced), and separation can be performed at least partially based on the barcode during analysis. In some instances, the barcode can be accurately identified even after one or more nucleotides in the barcode sequence have been mutated, inserted, or deleted (e.g., at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more nucleotides). Multiple polynucleotides from the same sample can have the same barcode. Alternatively, multiple polynucleotides from the same sample can have different barcodes. The first barcode can differ from the second barcode by at least three nucleotide positions, such as at least 3, 4, 5, 6, 7, 8, 9, 10, or more nucleotide positions. Multiple barcodes can be represented in a sample library, where each sample contains a polynucleotide containing one or more barcodes that are different from the barcodes contained in polynucleotides derived from other samples in the library. Polynucleotide samples containing one or more barcodes can be grouped according to the barcode sequences to which they are linked, such that all four nucleotide bases A, G, C, and T are approximately uniformly represented in the library along one or more positions of each barcode (such as positions 1, 2, 3, 4, 5, 6, 7, 8, or more, or all, of the barcode). In some instances, the method of the present invention may include identifying the sample from which the target polynucleotide originates based on the barcode sequence linked to the target polynucleotide. The barcode may contain a nucleic acid sequence that, when linked to the target polynucleotide, can act as an identifier of the sample from which the target polynucleotide originates. In one instance, oligonucleotide primers (e.g., amplification primers) may contain one or more barcodes. In another instance, nucleic acid molecules may be coupled (e.g., ligated) to an adaptor nucleic acid (e.g., for circularization), and the adaptor nucleic acid may contain one or more barcodes.

[0138] As used herein, the term "sample" generally refers to any sample that may include one or more components (e.g., nucleic acid molecules) for processing or analysis. A sample can be a biological sample. A sample can be a cell or tissue sample. A sample can be a cell-free sample, such as blood (e.g., whole blood), plasma, serum, sweat, saliva, or urine. A sample can be obtained in vivo or cultured in vitro.

[0139] As used herein, the term "subject" generally refers to the individual or entity from which the sample was derived, such as a vertebrate (e.g., a mammal, such as a human) or an invertebrate. A mammal can be a mouse, monkey, human, farm animal (e.g., a cow, sheep, pig, or chicken), or pet (e.g., a cat or dog). A subject can be a plant. A subject can be a patient. The subject may not have symptoms of a disease (e.g., cancer). Alternatively, the subject may have symptoms of a disease.

[0140] Whenever the terms "at least," "greater than," or "greater than or equal to" precede the first value in a series of two or more values, the terms "at least," "greater than," or "greater than or equal to" apply to each value in the series. For example, greater than or equal to 1, 2, or 3 is equivalent to greater than or equal to 1, greater than or equal to 2, or greater than or equal to 3.

[0141] Whenever the terms "not exceeding," "less than," or "less than or equal to" precede the first value in a series of two or more values, the terms "not exceeding," "less than," or "less than or equal to" apply to each value in that series. For example, less than or equal to 3, 2, or 1 is equivalent to less than or equal to 3, less than or equal to 2, or less than or equal to 1.

[0142] Overview

[0143] Currently available sequencing methods can be used to sequence one or more nucleic acids. However, such methods can be expensive and may not provide sequence information within the time frame necessary for diagnosing or treating subjects (e.g., individuals, patients, etc.) or within the required range (or level) of accuracy.

[0144] Massive parallel sequencing can be used to identify one or more sequence variants within a population (e.g., a complex population). However, massive parallel sequencing using currently available sequencing technologies can be limited by its error rate, which may be greater than the actual frequency of sequence variants in the population. In one instance, currently available high-throughput sequencing methods can exhibit an error rate of approximately 0.1–1 percent (%). In some cases, when the frequency of one or more sequence variants (e.g., one or more rare sequence variants) is low, for example, when the frequency equals or is lower than the error rate, the detection of that sequence variant may have a high false positive rate.

[0145] Methods and compositions for sequencing

[0146] Interchange coupling for sequencing

[0147] On one hand, this disclosure provides methods for processing or analyzing double-stranded nucleic acid molecules. The methods may include providing (i) a double-stranded nucleic acid molecule and (ii) a double-stranded adaptor having a cleavage site within its sense or antisense strand. The methods may include coupling the double-stranded adaptor to the double-stranded nucleic acid molecule. The methods may include cyclizing the double-stranded nucleic acid molecule coupled to the double-stranded adaptor to produce a cyclized double-stranded nucleic acid molecule. The cleavage site may contain a cleavage before coupling between the double-stranded adaptor and the double-stranded nucleic acid molecule. The cleavage may be a break in the sense strand of the double-stranded adaptor. Alternatively, the cleavage may be a break in the antisense strand of the double-stranded adaptor.

[0148] The length of a double-stranded nucleic acid molecule can be at least 2, 3, 4, 5, 10, 50, 100, 500, 1,000, 5,000, 10,000, 50,000, 100,000 or more nucleotides. The length of a double-stranded nucleic acid molecule can be at most 100,000, 50,000, 10,000, 5,000, 1,000, 500, 100, 50, 10 or fewer nucleotides. The length of a double-stranded adapter can be at least 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 45, 50 or more nucleotides. The length of a double-stranded linker can be up to 50, 45, 40, 35, 30, 25, 20, 19, 18, 17, 16, 15, 14, 13, 12, 11, 10, 9, 8, 7, 6, 5, 4, 3 or fewer nucleotides.

[0149] Double-stranded nucleic acid molecules and double-stranded adaptors can be heterologous to each other. They can be the same gene or different parts of different genes. One of the double-stranded nucleic acid molecules and the double-stranded adaptor can be derived from one species, while the other can come from a different species. One of the double-stranded nucleic acid molecules and the double-stranded adaptor can be natural, while the other can be synthetic (e.g., the double-stranded adaptor can be a synthetic molecule). Alternatively, both the double-stranded nucleic acid molecule and the double-stranded adaptor can be synthetic.

[0150] In one instance, the double-stranded nucleic acid molecule may be a fragmented double-stranded nucleic acid molecule from a genomic sample of the subject (e.g., from the subject's cells), and the double-stranded adaptor may be a synthetic adaptor including a nick site. In another instance, the double-stranded nucleic acid molecule may be a cell-free double-stranded nucleic acid molecule from a cell-free biological sample of the subject (e.g., blood, plasma, urine, etc.), and the double-stranded adaptor may be a synthetic adaptor including a nick site.

[0151] Double-stranded nucleic acid molecules and double-stranded adaptors can be provided in the form of cell-free compositions. Cell-free compositions may be substantially free of intact cells. Cell-free compositions may contain cell lysates or extracts. Cell lysates may contain a fluid containing one or more contents of lysed cells. Cell lysates may be crude (i.e., unpurified) or at least partially purified (e.g., to remove cell debris or particles, such as damaged outer cell membranes). Methods for forming cell lysates may include sonication, homogenization, enzymatic lysis using lysozyme, freezing, grinding, and autoclaving. Alternatively, cell-free compositions may be derived from cell-free biological samples.

[0152] The coupling of a double-stranded adaptor to a double-stranded nucleic acid molecule can be performed under cell-free conditions, or the cyclization of the double-stranded nucleic acid molecule coupled with the double-stranded adaptor to a cyclized double-stranded nucleic acid molecule can be performed under cell-free conditions. Alternatively, in the presence of one or more cells, (i) the coupling of the double-stranded adaptor to the double-stranded nucleic acid molecule can be performed, or (ii) the cyclization of the double-stranded nucleic acid molecule coupled with the double-stranded adaptor to a cyclized double-stranded nucleic acid molecule can be performed. In one instance, one or more cells can be configured to express one or more enzymes capable of performing the processes in (i) and / or (ii).

[0153] Coupling may include coupling the sense strand of a double-stranded adaptor to the sense strand of a double-stranded nucleic acid molecule. Alternatively, coupling may include coupling the antisense strand of a double-stranded adaptor to the antisense strand of a double-stranded nucleic acid molecule. In various alternatives, coupling may include (i) coupling the sense strand of a double-stranded adaptor to the sense strand of a double-stranded nucleic acid molecule, and (ii) coupling the antisense strand of a double-stranded adaptor to the antisense strand of a double-stranded nucleic acid molecule. Coupling of two nucleic acid molecules may include ligation (e.g., by an enzyme, such as a ligase), hybridization (e.g., in the absence of an enzyme), or both.

[0154] The cleavage site can be part of the sense strand of a circularized double-stranded nucleic acid molecule. The cleavage site may not be the 5′ end or the 3′ end of the sense strand. The cleavage site may be at least 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, or more nucleotides away from the 5′ or 3′ end of the sense strand. The cleavage site may be at most 30, 25, 20, 15, 10, 5, 4, 3, 2, or 1 nucleotide away from the 5′ or 3′ end of the sense strand. Alternatively or additionally, the cleavage site can be part of the antisense strand of a circularized double-stranded nucleic acid molecule. The cleavage site may not be the 5′ end or the 3′ end of the antisense strand. The cleavage site may be at least 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, or more nucleotides away from the 5′ or 3′ end of the antisense strand. The nick site can be located at a distance of up to 30, 25, 20, 15, 10, 5, 4, 3, 2, or 1 nucleotide from the 5′ or 3′ end of the antisense strand.

[0155] The method may also include sequencing double-stranded nucleic acid molecules from nick sites on double-stranded connectives. Sequencing may be whole-genome sequencing (or whole-genome sequencing) or targeted sequencing. Sequencing may include one or more NGS methods. Sequencing may include nanopore-based sequencing. The nanopore may be a protein nanopore (e.g., α-hemolysin) or a solid nanopore. Alternatively, the nanopore may be a hybrid nanopore comprising at least a portion of a protein nanopore (e.g., α-hemolysin) and at least a portion of a solid nanopore. Nanopore-based sequencing may utilize at least one enzyme (e.g., a polymerase or nuclease) to interact with at least one double-stranded nucleic acid molecule. At least one enzyme may be coupled to the nanopore. At least one enzyme may be fused, conjugated, or bound to a protein nanopore or a membrane containing a nanopore. At least one enzyme may be conjugated or bound to a solid nanopore or a membrane containing a solid nanopore. In some instances, at least one enzyme may have a binding portion capable of binding to a nanopore (or solid nanopore) or a membrane.

[0156] Sequencing may include (i) extending a double-stranded nucleic acid molecule from a nick site on a double-stranded connective to produce a growth strand that is sequence complementary to at least a portion of the strands of the double-stranded nucleic acid molecule, and (ii) obtaining sequence information of at least a portion of the growth strand. The strands of the double-stranded nucleic acid molecule may be its sense strand or its antisense strand. Therefore, the growth strand may exhibit complementarity with at least a portion of the sense strand or at least a portion of the antisense strand. Alternatively, the strands of the double-stranded nucleic acid molecule may be both a sense strand and an antisense strand. Therefore, a first growth strand may exhibit complementarity with at least a portion of the sense strand, and a second growth strand may exhibit complementarity with at least a portion of the antisense strand.

[0157] Extension reactions may include amplifying at least a portion of a double-stranded nucleic acid molecule (e.g., a double-stranded nucleic acid molecule and at least a portion of a double-stranded adapter). Amplification may produce multiple copies (e.g., at least 1, 2, 3, 4, 5, 10, 15, 20 or more copies; up to 20, 15, 10, 5, 4, 3, 2 or 1 copy) of at least a portion of the sense and / or antisense strand of the double-stranded nucleic acid molecule. In some instances, the double-stranded nucleic acid molecule may be a portion of a circular (or circulated) double-stranded nucleic acid molecule, and the extension reaction may be RCA. RCA may generate a growth strand containing one or more copies (e.g., at least 1, 2, 3, 4, 5, 10, 15, 20 or more copies; up to 20, 15, 10, 5, 4, 3, 2 or 1 copy) of at least a portion of the sense or antisense strand of the circular double-stranded nucleic acid molecule.

[0158] Obtaining sequence information may include detecting at least a portion of the growth chain. The extension reaction may include contacting a double-stranded nucleic acid molecule with a nucleotide coupled to a tag under conditions sufficient to incorporate nucleotides into the growth chain. In this case, obtaining sequence information may include detecting the tag. In some cases, the method may also include releasing the tag from the nucleotides during incorporation into the growth chain, and detecting the released tag for sequencing.

[0159] The extension reaction can be performed using oligonucleotide primers. Alternatively, the extension reaction can be performed without the use of oligonucleotide primers. The nick within the double-stranded adapter can act as a binding site for enzymes capable of performing the extension reaction (e.g., polymerases), thus potentially eliminating the need for any oligonucleotide primers.

[0160] Alternatively, sequencing may include (i) cleaving a double-stranded nucleic acid molecule at a cleavage site of a double-stranded connective to cleave at least a portion of the strand of the double-stranded nucleic acid molecule, and (ii) obtaining sequence information of at least a portion of the strand. The strand of the double-stranded nucleic acid molecule to be cleaved may be its sense strand or its antisense strand. Alternatively, the strand of the double-stranded nucleic acid molecule to be cleaved may be both a sense strand and an antisense strand. Subsequently, obtaining the sequence information may include detecting at least a portion of the strand.

[0161] Compared to at least one reference sequence, at least a portion of the double-stranded nucleic acid molecule may have, or be suspected of having, one or more sequencing variants (e.g., one or more mutations). Therefore, sequencing can be performed to identify the presence of at least a portion of the double-stranded nucleic acid molecule. One or more sequencing variants can indicate mutations in the gene. At least one reference sequence may contain a common sequence of at least a portion of the gene. The common sequence and the double-stranded nucleic acid molecule may be derived from the same species or different species. The common sequence may be a representative sequence from a collection of multiple sequences of a gene obtained from multiple samples of the same species (e.g., at least 2, 3, 4, 5, 10, 15, 20, 30, 40, 50 or more samples; up to 50, 40, 30, 20, 15, 10, 4, 3 or 2 samples). In one instance, both the double-stranded nucleic acid molecule and at least one reference sequence may be derived from a human sample, and at least one reference sequence may be a common sequence of at least a portion of a human gene of interest, such as a portion of a gene known to typically not have any mutations.

[0162] This method may include amplifying a double-stranded nucleic acid molecule to produce multiple copies of the double-stranded nucleic acid molecule before coupling a double-stranded adapter to it. Amplification may produce at least 1, 2, 3, 4, 5, 10, 15, 20, 30, 40, 50, 100, 150, 200 or more copies of the double-stranded nucleic acid molecule. Amplification may produce up to 200, 150, 100, 50, 40, 30, 20, 15, 10, 5, 4, 3, 2 or 1 copy of the double-stranded nucleic acid molecule. Thus, this method can utilize multiple double-stranded adapters, for example, the same number or more as the copy number of the double-stranded nucleic acid molecule.

[0163] Double-stranded nucleic acid molecules may contain a recognition sequence. The recognition sequence for double-stranded nucleic acid molecules may be endogenous or exogenous. The recognition sequence may contain at least one natural nucleotide, at least one synthetic nucleotide, or both. The recognition sequence may include at least 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 40, 50, 100 or more nucleotides. The recognition sequence may include up to 100, 50, 40, 30, 25, 24, 23, 22, 21, 20, 19, 18, 17, 16, 15, 14, 13, 12, 11, 10, 9, 8, 7, 6, 5 or fewer nucleotides. The method may also include enriching double-stranded nucleic acid molecules from a random nucleic acid molecule library, at least partially based on the recognition sequence. The library may contain at least 2, 3, 4, 5, 10, 50, 100, 500, 1,000, 5,000, 10,000, 50,000, 100,000, 500,000, 1,000,000 or more random nucleic acid molecules. The library may contain up to 1,000,000, 500,000, 100,000, 50,000, 10,000, 5,000, 1,000, 500, 100, 50, 10, 5, 4, 3 or 2 random nucleic acid molecules. Enrichment may involve isolating double-stranded nucleic acid molecules (containing the recognition sequence) from at least one different nucleic acid molecule that does not contain the recognition sequence. Double-stranded nucleic acid molecules can be isolated from at least 1, 5, 10, 50, 100, 500, 1,000, 5,000, 10,000, 50,000, 100,000, 500,000, 1,000,000 or more different nucleic acid molecules that do not contain a recognition sequence. Double-stranded nucleic acid molecules can also be isolated from up to 1,000,000, 500,000, 100,000, 50,000, 10,000, 5,000, 1,000, 500, 100, 50, 10, 5 or 1 different nucleic acid molecules that do not contain a recognition sequence.

[0164] Prior to amplification, double-stranded nucleic acid molecules in a random nucleic acid library (one of which is a double-stranded nucleic acid molecule containing a recognition sequence) can be enriched. Alternatively or additionally, the random nucleic acid library can be amplified before enrichment of double-stranded nucleic acid molecules.

[0165] Before (i) coupling with at least a double-stranded adaptor and (ii) circularization, double-stranded nucleic acid molecules in a random nucleic acid library (one of which is a double-stranded nucleic acid molecule containing a recognition sequence) can be enriched. Alternatively, the random nucleic acid library containing double-stranded nucleic acid molecules can be coupled with a double-stranded adaptor and then circularized. In different alternatives, the random nucleic acid library containing double-stranded nucleic acid molecules can be (i) coupled with at least a double-stranded adaptor and (ii) circularized before enriching with circularized double-stranded nucleic acid molecules containing the recognition sequence (and any excess of linear double-stranded nucleic acid molecules).

[0166] Enrichment may include generating a library of selected double-stranded nucleic acid molecules (e.g., linear or circular). Each double-stranded nucleic acid molecule in at least 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 60%, 70%, 80%, 90%, 95% or more of the selected library may contain at least a portion of a recognition sequence. Each double-stranded nucleic acid molecule in at most 100%, 95%, 90%, 80%, 70%, 60%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 15%, 10%, 5% or less of the selected library may contain at least a portion of a recognition sequence.

[0167] The probability of a recognized sequence appearing in the absence of any mismatches can be 1 x 10^6. 4 5x10 4 7x10 4 1x10 5 5x10 5 1x10 6 5x10 6 1x10 7 5x10 7 1x10 8 5x10 8 1x10 9 5x10 9 1x10 10 5x10 10 1x10 11 5x10 11 1x10 12 5x10 12 At most once in one or more nucleotides. Alternatively, the probability of a recognition sequence appearing in the absence of any mismatches could be in every 1 x 10^6 nucleotides. 4 5x10 4 7x10 4 1x10 5 5x10 5 1x10 6 5x10 61x10 7 5x10 7 1x10 8 5x10 8 1x10 9 5x10 9 1x10 10 5x10 10 1x10 11 5x10 11 1x10 12 5x10 12 Up to two, three, four, five or more of the nucleotides.

[0168] The identification sequence may contain at least 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 35, 40, 45, 50 or more nucleotides. The identification sequence may contain up to 50, 45, 40, 35, 30, 29, 28, 27, 26, 25, 24, 23, 22, 21, 20, 19, 18, 17, 16, 15, 14, 13, 12, 11, 10, 9, 8, 7, 6, 5 or fewer nucleotides.

[0169] Enrichment may include (i) binding a recognition moiety complementary to the recognition sequence to a double-stranded nucleic acid molecule to form a recognition complex, and (ii) extracting the recognition complex from a random nucleic acid molecule library. The recognition moiety may have at least 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95% or more complementarity to the recognition sequence. The recognition moiety may have at most 100%, 95%, 90%, 80%, 70%, 60%, 50%, 40%, 30% or less complementarity to the recognition sequence.

[0170] One or more double-stranded nucleic acid molecules containing a recognition sequence may be enriched in a random nucleic acid molecule library before (i) providing one or more double-stranded nucleic acid molecules and one or more double-stranded adapters having nick sites, or (ii) after conjugating at least one double-stranded adapter to each of a plurality of double-stranded nucleic acid molecules from a library. Alternatively, one or more double-stranded nucleic acid molecules containing a recognition sequence may be enriched in a random nucleic acid molecule library before (i) providing one or more double-stranded nucleic acid molecules and one or more double-stranded adapters having nick sites, and (ii) after conjugating at least one double-stranded adapter to each of a plurality of double-stranded nucleic acid molecules from a library.

[0171] Alternatively or additionally, the double-stranded adapter may contain at least one recognition sequence. After the double-stranded adapter is coupled (e.g., linked) to a double-stranded nucleic acid molecule, at least one recognition sequence (e.g., identified by at least one recognition portion provided herein) can be used to enrich the double-stranded nucleic acid molecule coupled to the double-stranded adapter. Enrichment may deplete one or more double-stranded nucleic acid molecules not coupled to the double-stranded adapter.

[0172] The recognition sequence of the double-stranded adapter may contain at least one natural nucleotide, at least one synthetic nucleotide, or both. The recognition sequence of the double-stranded adapter may contain at least 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 40, 50, 100 or more nucleotides. The recognition sequence of the double-stranded adapter may contain up to 100, 50, 40, 30, 25, 24, 23, 22, 21, 20, 19, 18, 17, 16, 15, 14, 13, 12, 11, 10, 9, 8, 7, 6, 5 or fewer nucleotides. The method may also include enriching double-stranded nucleic acid molecules coupled to the double-stranded adapter from a random nucleic acid library, at least in part, based on the recognition sequence of the double-stranded adapter. The library may contain at least 2, 3, 4, 5, 10, 50, 100, 500, 1,000, 5,000, 10,000, 50,000, 100,000, 500,000, 1,000,000 or more random nucleic acid molecules. The library may contain up to 1,000,000, 500,000, 100,000, 50,000, 10,000, 5,000, 1,000, 500, 100, 50, 10, 5, 4, 3 or 2 random nucleic acid molecules. Enrichment may involve isolating double-stranded nucleic acid molecules (coupled to a double-stranded linker containing the recognition sequence) from at least one different nucleic acid molecule that does not contain a recognition sequence. Double-stranded nucleic acid molecules can be isolated from at least 1, 5, 10, 50, 100, 500, 1,000, 5,000, 10,000, 50,000, 100,000, 500,000, 1,000,000 or more different nucleic acid molecules that do not contain a recognition sequence. Double-stranded nucleic acid molecules can also be isolated from up to 1,000,000, 500,000, 100,000, 50,000, 10,000, 5,000, 1,000, 500, 100, 50, 10, 5 or 1 different nucleic acid molecules that do not contain a recognition sequence.

[0173] Double-stranded nucleic acid molecules can be derived from or derived from biological samples of the subject. Biological samples can include cell-free biological samples of the subject. Cell-free biological samples can be selected from: blood, plasma, serum, urine, lymph, feces, saliva, semen, amniotic fluid, cerebrospinal fluid, bile, sweat, tears, sputum, synovial fluid, vomitus, and combinations thereof. Double-stranded nucleic acid molecules can be derived from or derived from cell-free nucleic acid molecules derived from cell-free biological samples. Cell-free nucleic acid molecules can include circulating tumor nucleic acid molecules (e.g., ctDNA) or amniotic fluid nucleic acid molecules.

[0174] Biological samples may include tissue samples from the subject. Tissue samples may be derived from: bone, heart, thymus, arteries, blood vessels, lungs, muscles, stomach, intestines, liver, pancreas, spleen, kidneys, gallbladder, thyroid gland, adrenal glands, breast, ovaries, prostate, testes, skin, fat, eyes, brain, and combinations thereof. Double-stranded nucleic acid molecules may be derived from or derived from genomic nucleic acid molecules from tissue samples. Tissue samples may be derived from: infected tissue, diseased tissue, malignant tissue, calcified tissue, healthy tissue, and combinations thereof. Tissue samples may be derived from malignant tissue containing tumors, sarcomas, leukemia, or derivatives thereof.

[0175] Double-stranded nucleic acid molecules can include DNA, cDNA, ctDNA, their derivatives, or combinations thereof. Double-stranded nucleic acid molecules include RNA.

[0176] In another aspect, this disclosure provides a reaction mixture for processing or analyzing double-stranded nucleic acid molecules. The reaction mixture may comprise a composition comprising (i) a double-stranded nucleic acid molecule and (ii) a double-stranded adaptor having a cleavage site within its sense or antisense strand. The reaction mixture may comprise at least one enzyme that (i) couples the double-stranded adaptor to the double-stranded nucleic acid molecule, or (ii) cyclizes the double-stranded nucleic acid molecule coupled to the double-stranded adaptor to produce a cyclized double-stranded nucleic acid molecule. The cleavage site may contain a cleavage prior to coupling between the double-stranded adaptor and the double-stranded nucleic acid molecule. The cleavage may be a break in the sense strand of the double-stranded adaptor. Alternatively, the cleavage may be a break in the antisense strand of the double-stranded adaptor. As provided in this disclosure, the reaction mixture may be used or identified in any of the subject methods for adaptor ligation. The reaction mixture may be used to prepare one or more libraries (e.g., libraries of nucleic acid molecules, enzymes, or combinations thereof) or compositions for one or more sequencing methods. One or more components of the reaction mixture may be used simultaneously in the same reaction. In one instance, the reaction can be carried out in a single reaction vial (e.g., a reaction tube), thereby reducing purification steps and / or sample loss, and / or allowing sequencing with a small amount of nucleic acid sample input. One or more components of the reaction mixture can be used separately in different reactions.

[0177] The reaction mixture may contain at least 1, 2, 3, 4, 5, 10, 50, 100, 500, 1,000, 5,000, 10,000, 50,000, 100,000, 500,000, 1,000,000 or more double-stranded nucleic acid molecules. The reaction mixture may contain up to 1,000,000, 500,000, 100,000, 50,000, 10,000, 5,000, 1,000, 500, 100, 50, 10, 5, 4, 3 or 1 double-stranded nucleic acid molecule. The reaction mixture may contain at least 1, 2, 3, 4, 5, 10, 50, 100, 500, 1,000, 5,000, 10,000, 50,000, 100,000, 500,000, 1,000,000 or more double-stranded linkers. The reaction mixture may contain up to 1,000,000, 500,000, 100,000, 50,000, 10,000, 5,000, 1,000, 500, 100, 50, 10, 5, 4, 3 or 1 double-stranded linker.

[0178] At least one enzyme may be capable of (i) coupling a double-stranded intransitive linker to a double-stranded nucleic acid molecule, and (ii) cyclizing the double-stranded nucleic acid molecule coupled with the double-stranded intransitive linker to produce a cyclized double-stranded nucleic acid molecule. The reaction mixture may contain at least 0.01, 0.1, 1, 10, 100, 1,000, 10,000 or more units (e.g., Weiss units or Modrich-Lehman units). The reaction mixture may contain at most 10,000, 1,000, 100, 10, 1, 0.1, 0.01 or fewer units of at least one enzyme. Alternatively, the reaction mixture may contain an enzyme that (i) couples a double-stranded intransitive linker to a double-stranded nucleic acid molecule, and (ii) an additional enzyme that cyclizes the double-stranded nucleic acid molecule coupled with the double-stranded intransitive linker to produce a cyclized double-stranded nucleic acid molecule. The enzyme and the additional enzyme may be different. The reaction mixture may contain at least 0.01, 0.1, 1, 10, 100, 1,000, 10,000 or more units of enzyme, and at least 0.01, 0.1, 1, 10, 100, 1,000, 10,000 or more units of additional enzyme. The reaction mixture may contain up to 10,000, 1,000, 100, 10, 1, 0.1, 0.01 or fewer units of enzyme, and up to 10,000, 1,000, 100, 10, 1, 0.1, 0.01 or fewer units of additional enzyme.

[0179] The reaction mixture may be a cell-free reaction mixture. A cell-free reaction mixture may be substantially free of intact cells. A cell-free reaction mixture may contain cell lysates or extracts. Cell lysates may contain a fluid containing one or more contents of lysed cells. Cell lysates may be crude (i.e., unpurified) or at least partially purified (e.g., to remove cell debris or particles, such as damaged outer cell membranes). A cell-free reaction mixture may be the product of sonication, homogenization, enzymatic lysis using lysozyme, freezing, grinding, and autoclaving of one or more cells. Alternatively, a cell-free reaction mixture may be derived from a cell-free biological sample.

[0180] At least one enzyme may be able to (i) couple the sense strand of a double-stranded adaptor to the sense strand of a double-stranded nucleic acid molecule, or (ii) couple the antisense strand of a double-stranded adaptor to the antisense strand of a double-stranded nucleic acid molecule. At least one enzyme may be able to (i) couple the sense strand of a double-stranded adaptor to the sense strand of a double-stranded nucleic acid molecule, and (ii) couple the antisense strand of a double-stranded adaptor to the antisense strand of a double-stranded nucleic acid molecule. Alternatively, (i) a first enzyme may be able to couple the sense strand of a double-stranded adaptor to the sense strand of a double-stranded nucleic acid molecule, and (ii) a second enzyme, different from the first enzyme, may be able to couple the antisense strand of a double-stranded adaptor to the antisense strand of a double-stranded nucleic acid molecule. At least one enzyme may include ligases, recombinases, polymerases, functional variants thereof, or combinations thereof. In some instances, at least one enzyme may be able to ligate a double-stranded adaptor to a double-stranded nucleic acid molecule.

[0181] The reaction mixture also contains at least a second enzyme that performs an extension reaction to produce a growth strand that is sequence-complementary to at least a portion of the strands of the double-stranded nucleic acid molecule. The second enzyme may produce the growth strand before the double-stranded adapter is coupled to the double-stranded nucleic acid molecule. In some instances, the second enzyme may produce multiple growth strands to amplify the double-stranded nucleic acid molecule. Amplified copies of the double-stranded nucleic acid molecule may be used in the same reaction mixture. Alternatively, amplified copies of the double-stranded nucleic acid molecule may be split into multiple reaction samples for treatment under the same or different reaction conditions. The second enzyme may include a polymerase. The second enzyme may include polymerases and recombinases, such as recombinase polymerase amplification, which can be used for single-tube isothermal amplification as an alternative to polymerase chain reaction (PCR) amplification.

[0182] After a double-stranded adaptor is coupled to a double-stranded nucleic acid molecule, at least a second enzyme may be able to perform an extension reaction from the nick site (or nick of the nick site) of the double-stranded adaptor to produce a growth chain. The reaction mixture may also contain at least one nucleotide coupled to a tag, wherein at least the second enzyme incorporates the nucleotide into the growth chain. The tag may be a small molecule, nucleotide, polynucleotide, amino acid, polypeptide, polymer, metal and / or ceramic particle, etc. In some instances, each of the nucleotides G, C, A, T, and U may contain a distinct tag that is distinguishable from each other. In some instances, the tag may not be released from the nucleotide when it is incorporated into the growth chain. Alternatively, at least the second enzyme may be able to release the tag from the nucleotide when it is incorporated into the growth chain. At least the second enzyme may perform the extension reaction with the aid of at least one oligonucleotide primer. Alternatively, at least the second enzyme may perform the extension reaction without using an oligonucleotide primer. In one instance, the extension reaction may include an RCA, wherein at least the second enzyme includes a polymerase. The polymerase may bind to the nick of the nick site for use in the RCA.

[0183] The reaction mixture may further contain at least a third enzyme that performs a cleavage reaction from the cleavage site of the double-stranded adapter to cleave at least a portion of the strand of the double-stranded nucleic acid molecule. Starting from the cleavage site, the at least third enzyme can displace and cleave (i) at least a portion of the strand of the double-stranded adapter containing the cleavage site and (ii) at least a portion of the strand of the double-stranded nucleic acid molecule coupled to the strand of the double-stranded adapter. The at least third enzyme may include a nuclease (e.g., an endonuclease, such as a restriction endonuclease).

[0184] As provided in this disclosure, at least a portion of a double-stranded nucleic acid molecule may have, or be suspected of having, one or more variants compared to at least one reference sequence. The reaction mixture can be used to prepare at least one composition for sequencing to identify the presence of at least a portion of the double-stranded nucleic acid molecule.

[0185] As provided in this disclosure, double-stranded nucleic acid molecules may contain a recognition sequence. Therefore, the reaction mixture may further contain a recognition moiety associated with the recognition sequence to enrich double-stranded nucleic acid molecules from a random library of nucleic acid molecules in the composition, at least partially based on the recognition sequence. The recognition moiety may contain at least one oligonucleotide (e.g., at least a portion of a gRNA for Cas system variants) that is complementary to at least the recognition sequence. The oligonucleotide of the recognition moiety may have at least 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95% or more complementarity to the recognition sequence. The oligonucleotide of the recognition moiety may have at most 100%, 95%, 90%, 80%, 70%, 60%, 50%, 40%, 30% or less complementarity to the recognition sequence.

[0186] The reaction mixture may comprise a selected library of double-stranded nucleic acid molecules (e.g., linear or circular). Each double-stranded nucleic acid molecule in at least 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 60%, 70%, 80%, 90%, 95% or more of the selected library may contain at least a portion of a recognition sequence. Each double-stranded nucleic acid molecule in at most 100%, 95%, 90%, 80%, 70%, 60%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 15%, 10%, 5% or less of the selected library may contain at least a portion of a recognition sequence.

[0187] The probability of a recognized sequence appearing in the absence of any mismatches can be 1 x 10^6. 4 5x10 4 7x10 4 1x10 5 5x10 5 1x10 6 5x10 6 1x10 7 5x10 7 1x10 8 5x10 8 1x10 9 5x10 9 1x10 10 5x10 10 1x10 11 5x10 11 1x10 12 5x10 12 At most once in one or more nucleotides (or base pairs). Alternatively, the probability of a recognition sequence appearing without any mismatches could be in every 1 x 10^6 nucleotides. 4 5x10 4 7x10 4 1x10 5 5x10 5 1x10 6 5x10 6 1x10 7 5x10 7 1x10 8 5x10 8 1x10 9 5x10 9 1x10 10 5x10 10 1x10 11 5x10 11 1x10 12 5x1012 Up to two, three, four, five or more of the following nucleotides (or base pairs).

[0188] The recognition sequence may contain at least 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 35, 40, 45, 50 or more nucleotides (or base pairs). The recognition sequence may contain up to 50, 45, 40, 35, 30, 29, 28, 27, 26, 25, 24, 23, 22, 21, 20, 19, 18, 17, 16, 15, 14, 13, 12, 11, 10, 9 or fewer nucleotides (or base pairs).

[0189] In various respects, this disclosure provides libraries of circularized double-stranded nucleic acid molecules. The library may contain (i) a double-stranded nucleic acid domain coupled to (ii) a double-stranded connective domain containing a cleavage site within its sense or antisense strand. Each circularized double-stranded nucleic acid molecule in at least 5% of the library may contain a recognition sequence. The library of circularized double-stranded nucleic acid molecules may be the starting material or product of any of the subject methods or reaction mixtures provided in this disclosure.

[0190] A cleavage site may exist within the double-stranded adaptor domain before coupling the double-stranded adaptor domain to the double-stranded nucleic acid domain. The cleavage site may contain a cleavage before coupling occurs between the double-stranded adaptor and the double-stranded nucleic acid molecule.

[0191] The double-stranded nucleic acid domain and the double-stranded linker domain may be heterologous to each other. Alternatively, the double-stranded nucleic acid domain and the double-stranded linker domain may be non-heterologous to each other. The library may be in a cell-free composition. Alternatively, the library may not be in a cell-free composition.

[0192] Compared to at least one reference sequence, at least a portion of the circularized double-stranded nucleic acid domains in the library may have, or be suspected of having, one or more sequencing variants. One or more sequencing variants may indicate mutations in the gene. At least one reference sequence may contain a common sequence of at least a portion of the gene.

[0193] Each circularized double-stranded nucleic acid molecule in at least 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 60%, 70%, 80%, 90%, 95% or more of the selected library may contain at least a portion of the recognition sequence. Each circularized double-stranded nucleic acid molecule in at most 100%, 95%, 90%, 80%, 70%, 60%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 15%, 10%, 5% or less of the selected library may contain at least a portion of the recognition sequence.

[0194] The probability of a recognized sequence appearing in the absence of any mismatches can be 1 x 10^6. 4 5x10 4 7x10 4 1x10 5 5x10 5 1x10 6 5x10 6 1x10 7 5x10 7 1x10 8 5x10 8 1x10 9 5x10 9 1x10 10 5x10 10 1x10 11 5x10 11 1x10 12 5x10 12 At most once in one or more nucleotides (or base pairs). Alternatively, the probability of a recognition sequence appearing without any mismatches could be in every 1 x 10^6 nucleotides. 4 5x10 4 7x10 4 1x10 5 5x10 5 1x10 6 5x10 6 1x10 7 5x10 7 1x10 8 5x10 8 1x10 9 5x10 9 1x10 10 5x10 10 1x10 11 5x10 11 1x10 12 5x10 12Up to two, three, four, five or more of the following nucleotides (or base pairs).

[0195] The recognition sequence may contain at least 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 35, 40, 45, 50 or more nucleotides (or base pairs). The recognition sequence may contain up to 50, 45, 40, 35, 30, 29, 28, 27, 26, 25, 24, 23, 22, 21, 20, 19, 18, 17, 16, 15, 14, 13, 12, 11, 10, 9 or fewer nucleotides (or base pairs).

[0196] The cleavage site may be part of the sense strand of a circularized double-stranded nucleic acid molecule. Alternatively or additionally, the cleavage site may be part of the antisense strand of a circularized double-stranded nucleic acid molecule.

[0197] Circulated double-stranded nucleic acid molecules can be derived from or derived from biological samples of the subject. Biological samples can include cell-free biological samples of the subject. Circulated double-stranded nucleic acid molecules can be derived from or derived from cell-free nucleic acid molecules derived from cell-free biological samples. Cell-free nucleic acid molecules can include circulating tumor nucleic acid (e.g., ctDNA) molecules or amniotic fluid nucleic acid molecules. Alternatively, biological samples can include tissue samples of the subject. Circulated double-stranded nucleic acid molecules can be derived from or derived from genomic nucleic acid molecules derived from tissue samples. Tissue samples can be derived from: infected tissue, diseased tissue, malignant tissue, calcified tissue, healthy tissue, and combinations thereof. Tissue samples can be derived from malignant tissue containing tumors, sarcomas, leukemia, or derivatives thereof.

[0198] Circular double-stranded nucleic acid molecules can include DNA, cDNA, ctDNA, their derivatives, or combinations thereof. Circular double-stranded nucleic acid molecules include RNA.

[0199] Figure 1AAn example method for providing a nicked circular nucleic acid is schematically illustrated. A first double-stranded nucleic acid molecule 110 and a second double-stranded nucleic acid molecule 120 may be provided. The nucleic acid molecule may be at least a portion of a biological sample or derived from a biological sample. The nucleic acid molecule 120 may be an adaptor (e.g., an adaptor for circularization, polymerization, nuclease activity, etc.). The first double-stranded nucleic acid molecule 110 may contain a recognition sequence 112 and a target site 114. The target site 114 may have or be suspected of having one or more mutations compared to at least a reference sequence (e.g., a shared sequence of a portion of the genes of a species or multiple species). The second double-stranded nucleic acid molecule 112 may include a nick 122 within the nucleic acid molecule 112. The nick 122 may not be located directly at the 5′ or 3′ end of the (i) sense strand or (ii) antisense strand of the nucleic acid molecule 112. Alternatively, the nucleic acid molecule 112 may be characterized by the removal of a phosphate group at the 5′ end of the forward strand or the 5′ end of the reverse strand. In some instances, the phosphate-removed nucleic acid molecule 112 may be produced by DNA synthesis. In other instances, nucleic acid molecules 112 with phosphate groups removed can be generated using PCR primers.

[0200] Reference Figure 1A A first nucleic acid molecule 110 and a second nucleic acid molecule 120 may be coupled (e.g., by ligation and / or hybridization) to produce a coupled double-stranded nucleic acid molecule 130. Coupling may be performed at a constant temperature or at multiple temperatures (e.g., stepwise or gradient or multiple temperatures). Coupling may be performed by one or more enzymes, such as ligases and / or recombinases. Subsequently, the nucleic acid molecule 130 may be circularized (e.g., by ligation and / or hybridization) to form a circular nucleic acid molecule 140. The circular nucleic acid molecule 140 may include at least a portion of the nucleic acid molecule 110 (e.g., recognition sequence 112 and target site 114) and at least a portion of the nucleic acid molecule 120 (e.g., a nick 122). The circular nucleic acid molecule 140 may be a template for sequencing (e.g., nanopore sequencing). In one instance, at least one enzyme (e.g., polymerase or nuclease) can be used to obtain information about the target site 114 of the circular nucleic acid molecule 140 via nanopore sequencing, and at least one enzyme can initiate its activity (e.g., extension reaction or cleavage) at the nick 122 of the circular nucleic acid molecule 140.

[0201] Figure 1B and 1C An example method for isolating or enriching linear or circular nucleic acids containing recognition sites is illustrated schematically. (Refer to...) Figure 1BThe random nucleic acid library 150 may include a first double-stranded nucleic acid molecule 110, which includes a recognition site 112 and a target site 114. The library 150 may also include linear nucleic acid molecules that contain the recognition site 112 but not the target site 114. The library 150 may also include linear nucleic acid molecules that do not contain the recognition site 112. The library 150 may be treated with at least one recognition portion (e.g., a dCas system) 113 to bind to the recognition site 112 and form a recognition complex at the recognition site 112. Any linear nucleic acid molecule containing such a recognition complex can be isolated from one or more nucleic acid molecules that do not contain the recognition complex to produce a purified or enriched library 155. In some instances, at least 5% of the nucleic acid molecules in the library 155 may be characterized as having the recognition sequence 112. In other instances, at least a majority of the nucleic acid molecules in the library 155 may be characterized as having the recognition sequence 112. The recognition portion 113 may be removed (isolated) from the recognition site 112 during purification or enrichment.

[0202] Reference Figure 1C A random circular nucleic acid library 160 may include double-stranded nucleic acid molecules 110 containing a recognition site 112 and a target site 114. Library 160 may also contain other circular nucleic acid molecules that do not contain a recognition site 112. Library 160 may be treated with at least one recognition portion (e.g., a dCas system) 113 to bind to the recognition site 112 and form a recognition complex at the recognition site 112. Any circular nucleic acid molecule containing such a recognition complex can be isolated from one or more circular nucleic acid molecules that do not contain the recognition complex to produce a purified or enriched library 165. In some instances, at least 5% (or most) of the circular nucleic acid molecules in library 165 may be characterized as having the recognition sequence 112. In some instances, at least a majority of the circular nucleic acid molecules in library 165 may be characterized as having the recognition sequence 112. The recognition portion 113 may be removed (isolated) from the recognition site 112 during purification or enrichment.

[0203] Figure 1DAn example method is illustrated using one or more uracil-specific enzymes to provide a circular nucleic acid containing a nick at a specific location within the circular nucleic acid. A double-stranded nucleic acid molecule can be provided. The double-stranded nucleic acid molecule can be a fragmented or complete nucleic acid molecule from a biological sample. The double-stranded nucleic acid molecule may contain two blunt ends. The double-stranded nucleic acid molecule may not contain two blunt ends (e.g., only one blunt end or no blunt ends). In this case, end repair may be required so that the repaired double-stranded nucleic acid molecule can (i) have no protruding ends and (ii) contain a 5′ phosphate group and a 3′ hydroxyl group in both the sense and antisense strands. The blunt ends can be obtained by end-filling with one or more enzymes (e.g., restriction endonucleases and / or exonucleases). 5′ phosphorylation can be achieved by one or more enzymes, such as kinases like T4 polynucleotide kinase. Alternatively or additionally, a non-template deoxyadenosine 5′-monophosphate (dAMP) can be incorporated into the 3′ end of the blunt end of the double-stranded nucleic acid molecule (i.e., dA-tailing). dA tailing can prevent tandem formation (e.g., during one or more ligation steps). dA tailing allows a double-stranded nucleic acid molecule to be linked to one or more adaptors containing a complementary deoxythymidine monophosphate (dTMP or "dT") overhang.

[0204] Reference Figure 1D In process 170, a modified double-stranded nucleic acid molecule may be coupled to one or more adaptors. The adaptor may contain one or more uracil (U) nucleotides (e.g., within one or more strands of the adaptor). The adaptor may contain a dT overhang complementary to the dA tail of the modified double-stranded nucleic acid molecule. Coupling may include joining (e.g., by DNA ligase) the end of the modified double-stranded nucleic acid molecule containing the dA tail to the end of the adaptor containing the dT overhang. In this example, the modified double-stranded nucleic acid molecule may be conjugated to two adaptors, wherein the first adaptor contains one or more uracil residues, and the second adaptor does not contain any uracil residues. Subsequently, the free ends of the first and second adaptors may be coupled (e.g., joined) to produce a circular nucleic acid molecule comprising at least a portion of the original double-stranded nucleic acid molecule, at least a portion of the first adaptor containing one or more uracil residues, and at least a portion of the second adaptor. The circular nucleic acid molecule can then be treated with one or more uracil-specific enzymes (e.g., uracil-specific excision reagents or "USER") to create a single nucleotide gap (i.e., a nick) at the location of a uracil residue. The resulting circular nucleic acid molecule with a nick at a specific site can then be analyzed for sequencing.

[0205] Reference Figure 1DIn process 175, a modified double-stranded nucleic acid molecule may be coupled to one or more adaptors. The adaptor may contain one or more uracil (U) nucleotides (e.g., within one or more strands of the adaptor). The adaptor may contain a dT overhang complementary to the dA tail of the modified double-stranded nucleic acid molecule. The adaptor may also contain a sticky end for hybridization. Coupling may include joining (e.g., by DNA ligase) the end of the modified double-stranded nucleic acid molecule containing the dA tail to the end of the adaptor containing the dT overhang. In this example, the modified double-stranded nucleic acid molecule may be conjugated to two adaptors, wherein the first adaptor contains one or more uracil residues, and the second adaptor does not contain any uracil residues. After coupling, both adaptors may include free sticky ends. Subsequently, the free ends of the first and second adaptors may be coupled to each other (e.g., by hybridization and ligation of the sticky ends) to produce a circular nucleic acid molecule comprising at least a portion of the original double-stranded nucleic acid molecule, at least a portion of the first adaptor containing one or more uracil residues, and at least a portion of the second adaptor. The circular nucleic acid molecule can then be treated with one or more uracil-specific enzymes (e.g., uracil-specific excision reagents) to create a single nucleotide gap (i.e., a nick) at the location of a uracil residue. The resulting circular nucleic acid molecule with a nick at a specific site can then be analyzed for sequencing.

[0206] Other examples of uracil-specific enzymes may include, but are not limited to, uracil-DNA glycosyltransferase (UDG) and / or Afu uracil-DNA glycosyltransferase (Afu UDG), for example, to cleave the N-glycosidic bond of deoxyuridine and create a nick for DNA polymerase to perform DNA elongation. Alternatively, UDG and / or Afu UDG may be combined with one or more repair enzymes specific to depurinyl / depyrimidine sites (e.g., FPG, hOGG1, hNEIL1, etc.).

[0207] Controlling the distance from the incision to the target site for sequencing.

[0208] On one hand, this disclosure provides methods for processing or analyzing circular nucleic acid molecules. The method may include providing a cell-free composition comprising a circular nucleic acid molecule containing (i) a target region and (ii) a nick site at a known distance from the target region. The method may include creating a nick at the nick site of the circular nucleic acid molecule. The target region may be a gene of interest. The target region may be a suspected mutation site or another site adjacent to the suspected mutation site. In one instance, a suspected mutation site for a specific disease may be known, and the sequence of the gene including the suspected mutation site may also be known. Thus, by assigning a nick site to a specific region of a gene characterized by a low probability of mutation, the circular nucleic acid molecule comprises (i) a target region and (ii) a nick site at a known distance from the target region.

[0209] During sequencing, it may be advantageous to position the cut at a known distance from the target site in one or more ways, including but not limited to: (i) increasing the probability of multiple amplifications of the target site during amplification (e.g., RCA), (ii) increasing the probability of at least one amplification of the target site before the activity of the enzyme responsible for amplification (e.g., polymerase) is depleted, and (iii) reducing sequencing errors (e.g., enzyme errors due to enzyme fatigue).

[0210] In some instances, multiple nickases (e.g., multiple different types of nickases or different variants of the same nickase, e.g., Cas nickases with different gRNAs) can be prepared to bind to multiple nick sites on a circular nucleic acid molecule. The binding of multiple nickases to their respective nick sites, any off-target binding, or nick activity can be evaluated. At least one nick site from the multiple nick sites can be selected to produce low off-target binding and / or high nick activity. Thus, by selecting nick sites from multiple nick sites, the distance between the nick site and the target site can be known before processing or analyzing one or more other circular nucleic acid molecules. Nick sites can be selected from at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 30, 40, 50, 100, or more nick sites on a circular nucleic acid molecule. Nick sites can be selected from up to 100, 50, 40, 30, 20, 15, 10, 9, 8, 7, 6, 5, 4, 3, or 2 nick sites on a circular nucleic acid molecule.

[0211] The creation of nicks at nick sites on circular nucleic acid molecules can be performed under cell-free conditions. Cell-free conditions may be substantially free of intact cells. Cell-free conditions may include conditions containing cell lysates or extracts. Cell lysates may contain a fluid containing one or more contents of lysed cells. Cell lysates may be crude (i.e., unpurified) or at least partially purified (e.g., to remove cell debris or particles, such as damaged outer cell membranes). Methods for forming cell lysates may include sonication, homogenization, enzymatic lysis using lysozyme, freezing, grinding, and autoclaving. Alternatively, cell-free conditions may include conditions for cell-free biological samples. Alternatively, the creation of nicks at nick sites on circular nucleic acid molecules can be performed in the presence of one or more cells (e.g., live or dead).

[0212] The cleavage site can be no more than 100,000, 50,000, 10,000, 5,000, 1,000, 500, 400, 300, 200, 100, 90, 80, 70, 60, 50, 40, 30, 25, 20, 15, 10, 5 or fewer nucleotides from the target site. The cleavage site can be at least 5, 10, 15, 20, 25, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 1,000, 5,000, 10,000, 50,000, 100,000 or more nucleotides from the target site.

[0213] Circular nucleic acid molecules may include circular double-stranded nucleic acid molecules. A cleavage site may be part of the sense strand of the circular nucleic acid molecule. Alternatively or additionally, a cleavage site may be part of the antisense strand of the circular nucleic acid molecule. The method may also include determining the cleavage site based at least in part on the position of the target site relative to at least one reference sequence. The at least one reference sequence may comprise a common sequence of at least a portion of a gene. The cleavage site may be endogenous for the circular nucleic acid molecule. Therefore, determining the cleavage site may include selecting an endogenous sequence of the circular nucleic acid molecule. Alternatively, the cleavage site may be exogenous for the circular nucleic acid molecule. Therefore, the determination may include inserting an exogenous cleavage site into the circular nucleic acid molecule. In some instances, an endogenous sequence of the circular nucleic acid molecule may be selected, and then an exogenous cleavage site may be inserted within or near the endogenous sequence of the circular nucleic acid molecule. In this way, the distance between the exogenous cleavage site and the target site can be controlled or known.

[0214] The circular nucleic acid molecule may also contain a nicking enzyme binding site specific to the nicking enzyme. In this case, the method may further include providing the nicking enzyme to the circular nucleic acid molecule under conditions sufficient to cause the nicking enzyme to associate with the nicking enzyme binding site and generate a nick. The probability of the nicking enzyme binding site appearing in the absence of any mismatches may be per 1 x 10^6 nucleotides. 4 5x10 4 7x10 4 1x10 5 5x10 5 1x10 6 5x10 6 1x10 7 5x10 7 1x10 8 5x10 8 1x10 9 5x10 9 1x10 10 5x10 10 1x10 11 5x10 111x10 12 5x10 12 At most once in one or more nucleotides (or base pairs). Alternatively, the probability of a nickase binding site appearing in the absence of any mismatches could be in every 1 x 10^6 nucleotides. 4 5x10 4 7x10 4 1x10 5 5x10 5 1x10 6 5x10 6 1x10 7 5x10 7 1x10 8 5x10 8 1x10 9 5x10 9 1x10 10 5x10 10 1x10 11 5x10 11 1x10 12 5x10 12 Up to two, three, four, five or more of the following nucleotides (or base pairs).

[0215] The nicking enzyme binding site may include at least 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 35, 40, 45, 50 or more nucleotides (or base pairs). The nicking enzyme binding site may also include up to 50, 45, 40, 35, 30, 29, 28, 27, 26, 25, 24, 23, 22, 21, 20, 19, 18, 17, 16, 15, 14, 13, 12, 11, 10, 9 or fewer nucleotides (or base pairs).

[0216] The nicking enzyme binding site can be endogenous for circular nucleic acid molecules. In one example, a linear nucleic acid molecule containing an endogenous nicking enzyme binding site can be circularized to form a circular nucleic acid molecule. Alternatively, the nicking enzyme binding site can be exogenous for circular nucleic acid molecules. In this case, the method may further include inserting the exogenous nicking enzyme binding site into the circular nucleic acid molecule or its starting material (e.g., a linear nucleic acid molecule) prior to nicking. In some examples, the exogenous nicking enzyme binding site may be inserted into the linear nucleic acid molecule prior to circularization into a circular nucleic acid molecule. The linear nucleic acid molecule may contain at least one recognition site as provided herein, and the recognition portion (e.g., a catalytically active recognition portion, such as Cas9) may be used to (i) bind to the recognition site, (ii) cleave the linear nucleic acid molecule at the recognition site, and (iii) facilitate the insertion of the exogenous nicking enzyme binding site into at least one recognition site of the linear nucleic acid molecule or adjacent to at least one recognition site (e.g., by homology-directed repair). The method may further include circularizing the nucleic acid molecule after inserting the exogenous nicking enzyme binding site. In other instances, the method may further include cyclizing the linear nucleic acid molecule into a circular nucleic acid molecule before inserting the exogenous nickase binding site into the circular nucleic acid molecule. After cyclization, the recognition portion (e.g., a catalytically active recognition portion, such as Cas9) can be used to (i) bind to the recognition site of the circular nucleic acid molecule, (ii) cleave the circular nucleic acid molecule at the recognition site, and (iii) facilitate the insertion of the exogenous nickase binding site into at least one recognition site of the circular nucleic acid molecule or adjacent to at least one recognition site (e.g., via homology-directed repair).

[0217] The nicking enzyme binding site can be no more than 30, 25, 20, 19, 18, 17, 16, 15, 14, 13, 12, 11, 10, 9, 8, 7, 6, 5, 4, 3, 2, or 1 nucleotide from the nicking site. The nicking enzyme binding site can be at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, or more nucleotides from the nicking site. In some instances, the nicking enzyme binding site may contain the nicking site. In this case, the nicking site may be within the nicking enzyme binding site.

[0218] The method may also include sequencing circular (or circularized) nucleic acid molecules from a nick in the circular nucleic acid molecule. Sequencing can be used for whole-genome sequencing (or whole-genome sequencing) or targeted sequencing. Sequencing may include one or more NGS methods. Sequencing may include nanopore-based sequencing. The nanopore may be a protein nanopore (e.g., α-hemolysin) or a solid nanopore. Alternatively, the nanopore may be a hybrid nanopore comprising at least a portion of a protein nanopore (e.g., α-hemolysin) and at least a portion of a solid nanopore. Nanopore-based sequencing may utilize at least one enzyme (e.g., a polymerase or nuclease) to interact with at least a circular nucleic acid molecule. At least one enzyme may be coupled to the nanopore. At least one enzyme may be fused, conjugated, or bound to a protein nanopore or a membrane containing a nanopore. At least one enzyme may be conjugated or bound to a solid nanopore or a membrane containing a solid nanopore. In some instances, at least one enzyme may have a binding portion capable of binding to a nanopore (or solid nanopore) or a membrane.

[0219] Sequencing may include sequencing a circular (or circularized) nucleic acid molecule. Sequencing may include extending the circular nucleic acid molecule from a nick to produce a growth strand that is sequence-complementary to at least a portion of the strand of the circular nucleic acid molecule. The method may also include obtaining sequence information of at least a portion of the growth strand. Obtaining sequence information may include detecting at least a portion of the growth strand. The extension reaction may include contacting the circular nucleic acid molecule with a nucleotide coupled to a tag under conditions sufficient to incorporate a nucleotide into the growth strand. Obtaining sequence information may include detecting at least a portion of the tag. When performing sequencing analysis, at least a portion of the tag may be attached to the growth strand. Alternatively, the method may include releasing the tag from the nucleotide as the nucleotide is incorporated into the growth strand, and detecting the released tag for sequencing.

[0220] The extension reaction can be performed using oligonucleotide primers. Alternatively, the extension reaction can be performed without the use of oligonucleotide primers. The nick within the double-stranded adapter can serve as a binding site for enzymes (e.g., polymerases) capable of performing extension reactions (e.g., rolling circle amplification), thus eliminating the need for oligonucleotide primers.

[0221] Alternatively, sequencing may include (i) cleaving a circular nucleic acid molecule through a cleavage reaction to cut at least a portion of the strand of a double-stranded nucleic acid molecule, and (ii) obtaining sequence information of at least a portion of the strand. The strand of the circular nucleic acid molecule may be its sense strand or its antisense strand. Therefore, the growing strand may exhibit complementarity with at least a portion of the sense strand or at least a portion of the antisense strand. Alternatively, the strand of the double-stranded nucleic acid molecule may be a sense strand and an antisense strand. Therefore, a first growing strand may exhibit complementarity with at least a portion of the sense strand, and a second growing strand may exhibit complementarity with at least a portion of the antisense strand.

[0222] The extension reaction may include amplifying at least a portion of a circular nucleic acid molecule (e.g., at least a portion of a nick site and at least a portion of a target site). Amplification (e.g., RCA) may produce multiple copies (e.g., at least 1, 2, 3, 4, 5, 10, 15, 20 or more copies; at most 20, 15, 10, 5, 4, 3, 2 or 1 copy) of at least a portion of the sense and / or antisense strand of the circular nucleic acid molecule. The extension product based on the circular nucleic acid molecule will have a first domain complementary to at least a portion of the nick site and a second domain complementary to at least a portion of the target site, and the distance between the first and second domains may be substantially the same as the known distance between the target site and the nick site in the circular nucleic acid molecule. Subsequently, obtaining sequence information may include detecting at least a portion of the strand.

[0223] Compared to at least one reference sequence, at least a portion of the circular nucleic acid molecule may have, or be suspected of having, one or more sequencing variants (e.g., one or more mutations). Therefore, sequencing can be performed to identify the presence of at least a portion of the circular nucleic acid molecule. One or more sequencing variants can indicate mutations in a gene. At least one reference sequence may contain a common sequence of at least a portion of the gene. The common sequence and the double-stranded nucleic acid molecule may be derived from the same or different species. The common sequence may be a representative sequence of a collection of multiple sequences of a gene obtained from multiple samples of the same species (e.g., at least 2, 3, 4, 5, 10, 15, 20, 30, 40, 50 or more samples; up to 50, 40, 30, 20, 15, 10, 4, 3 or 2 samples). In one instance, both the circular nucleic acid molecule and at least one reference sequence may be derived from a human sample, and at least one reference sequence may be a common sequence of at least a portion of a human gene of interest, such as a portion of a gene known to typically not have any mutations.

[0224] This method may include circularizing at least a linear nucleic acid molecule to produce a circular nucleic acid molecule. In some cases, the method may also include amplifying the linear nucleic acid molecule to produce multiple copies of the linear nucleic acid molecule. Amplification may produce at least 1, 2, 3, 4, 5, 10, 15, 20, 30, 40, 50, 100, 150, 200 or more copies of the linear nucleic acid molecule. Amplification may produce up to 200, 150, 100, 50, 40, 30, 20, 15, 10, 5, 4, 3, 2 or 1 copy of the linear nucleic acid molecule. The method may also include circularizing one or more copies of the linear nucleic acid molecule to produce multiple circular nucleic acid molecules. Circularization may include self-ligation (e.g., by one or more ligases), ligation by an adaptor, hybridization by an adaptor, or a combination thereof.

[0225] Alternatively or additionally, the method may include amplifying a circular nucleic acid molecule to produce multiple copies of the circular nucleic acid molecule. Amplification may produce at least 1, 2, 3, 4, 5, 10, 15, 20, 30, 40, 50, 100, 150, 200 or more copies of the circular nucleic acid molecule. Amplification may produce up to 200, 150, 100, 50, 40, 30, 20, 15, 10, 5, 4, 3, 2 or 1 copy of the circular nucleic acid molecule.

[0226] Circular nucleic acid molecules may contain a recognition sequence. The recognition sequence for circular nucleic acid molecules may be endogenous or exogenous. The recognition sequence may contain at least one natural nucleotide, at least one synthetic nucleotide, or both. The recognition sequence may contain at least 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 40, 50, 100 or more nucleotides. The recognition sequence may contain up to 100, 50, 40, 30, 25, 24, 23, 22, 21, 20, 19, 18, 17, 16, 15, 14, 13, 12, 11, 10, 9, 8, 7, 6, 5 or fewer nucleotides. The method may also include enriching circular nucleic acid molecules from a random nucleic acid molecule library, at least partially based on the recognition sequence. The library may contain at least 2, 3, 4, 5, 10, 50, 100, 500, 1,000, 5,000, 10,000, 50,000, 100,000, 500,000, 1,000,000 or more random nucleic acid molecules. The library may contain up to 1,000,000, 500,000, 100,000, 50,000, 10,000, 5,000, 1,000, 500, 100, 50, 10, 5, 4, 3 or 2 random nucleic acid molecules. Enrichment may involve isolating circular nucleic acid molecules (containing the recognition sequence) from at least one different nucleic acid molecule that does not contain a recognition sequence. Circular nucleic acid molecules can be isolated from at least 1, 5, 10, 50, 100, 500, 1,000, 5,000, 10,000, 50,000, 100,000, 500,000, 1,000,000 or more different nucleic acid molecules that do not contain a recognition sequence. Circular nucleic acid molecules can also be isolated from up to 1,000,000, 500,000, 100,000, 50,000, 10,000, 5,000, 1,000, 500, 100, 50, 10, 5 or 1 different nucleic acid molecules that do not contain a recognition sequence.

[0227] Prior to amplification, linear nucleic acid molecules containing the recognition sequence from a random nucleic acid library (one of which is a linear nucleic acid molecule containing the recognition sequence) can be enriched. Alternatively or additionally, the random nucleic acid library can be amplified before enriching linear nucleic acid molecules containing the recognition sequence. Alternatively, circular nucleic acid molecules from a random nucleic acid library (one of which is a circular nucleic acid molecule containing the recognition sequence) can be enriched before amplification. Alternatively or additionally, the random nucleic acid library can be amplified before enriching circular nucleic acid molecules.

[0228] Enrichment may include generating a selected library of circular nucleic acid molecules. Each circular nucleic acid molecule in at least 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 60%, 70%, 80%, 90%, 95% or more of the selected library may contain at least a portion of a recognition sequence. Each circular nucleic acid molecule in at most 100%, 95%, 90%, 80%, 70%, 60%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 15%, 10%, 5% or less of the selected library may contain at least a portion of a recognition sequence.

[0229] The probability of a recognized sequence appearing in the absence of any mismatches can be 1 x 10^6. 4 5x10 4 7x10 4 1x10 5 5x10 5 1x10 6 5x10 6 1x10 7 5x10 7 1x10 8 5x10 8 1x10 9 5x10 9 1x10 10 5x10 10 1x10 11 5x10 11 1x10 12 5x10 12 At most once in one or more nucleotides (or base pairs). Alternatively, the probability of a recognition sequence appearing without any mismatches could be in every 1 x 10^6 nucleotides. 4 5x10 4 7x10 4 1x10 5 5x10 5 1x10 6 5x10 6 1x10 7 5x107 1x10 8 5x10 8 1x10 9 5x10 9 1x10 10 5x10 10 1x10 11 5x10 11 1x10 12 5x10 12 Up to two, three, four, five or more of the following nucleotides (or base pairs).

[0230] The recognition sequence may contain at least 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 35, 40, 45, 50 or more nucleotides (or base pairs). The recognition sequence may contain up to 50, 45, 40, 35, 30, 29, 28, 27, 26, 25, 24, 23, 22, 21, 20, 19, 18, 17, 16, 15, 14, 13, 12, 11, 10, 9 or fewer nucleotides (or base pairs).

[0231] Enrichment may include (i) binding a recognition moiety complementary to the recognition sequence to a circular nucleic acid molecule to form a recognition complex, and (ii) extracting the recognition complex from a random nucleic acid library. The recognition moiety may have at least 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95% or more complementarity to the recognition sequence. The recognition moiety may have at most 100%, 95%, 90%, 80%, 70%, 60%, 50%, 40%, 30% or less complementarity to the recognition sequence.

[0232] One or more circular nucleic acid molecules containing a recognition sequence from a random nucleic acid molecule library may be enriched, the enrichment being performed (i) prior to providing a cell-free composition containing a circular nucleic acid molecule that includes a target region and a nick site at a known distance from the target region, or (ii) after a nick has been created at the nick site of the circular nucleic acid molecule. Alternatively, one or more circular nucleic acid molecules containing a recognition sequence from a random nucleic acid molecule library may be enriched, the enrichment being performed (i) prior to providing a cell-free composition containing a circular nucleic acid molecule that includes a target region and a nick site at a known distance from the target region, and (ii) after a nick has been created at the nick site of the circular nucleic acid molecule.

[0233] Circular nucleic acid molecules can be derived from or derived from biological samples of the subject. Biological samples can include cell-free biological samples of the subject. Cell-free biological samples can be selected from: blood, plasma, serum, urine, lymph, feces, saliva, semen, amniotic fluid, cerebrospinal fluid, bile, sweat, tears, sputum, synovial fluid, vomitus, and combinations thereof. Circular nucleic acid molecules can be derived from or derived from cell-free nucleic acid molecules derived from cell-free biological samples. Cell-free nucleic acid molecules can include circulating tumor nucleic acid molecules (e.g., ctDNA) or amniotic fluid nucleic acid molecules.

[0234] Biological samples may include tissue samples from the subject. Tissue samples may be derived from: bone, heart, thymus, arteries, blood vessels, lungs, muscles, stomach, intestines, liver, pancreas, spleen, kidneys, gallbladder, thyroid gland, adrenal glands, breast, ovaries, prostate, testes, skin, fat, eyes, brain, and combinations thereof. Circular nucleic acid molecules may originate from or be derived from genomic nucleic acid molecules derived from tissue samples. Tissue samples may be derived from: infected tissue, diseased tissue, malignant tissue, calcified tissue, healthy tissue, and combinations thereof. Tissue samples may originate from malignant tissue containing tumors, sarcomas, leukemia, or derivatives thereof.

[0235] Circular nucleic acid molecules can include DNA, cDNA, ctDNA, their derivatives, or combinations thereof. Circular nucleic acid molecules include RNA.

[0236] On the other hand, this disclosure provides reaction mixtures for processing or analyzing circular nucleic acid molecules. The reaction mixture may contain a cell-free composition comprising circular nucleic acid molecules. The circular nucleic acid molecule may contain (i) a target site and (ii) a nick site at a known distance from the target site. The reaction mixture may contain at least one enzyme that creates a nick at the nick site of the circular nucleic acid molecule. At least one enzyme may comprise a nuclease (e.g., a restriction endonuclease) or a nicking enzyme (e.g., a Cas9n nicking enzyme). At least one enzyme may comprise both a nuclease and a nicking enzyme. As provided in this disclosure, the reaction mixture may be used or identified in any subject method for sequencing at a known nick-to-target distance. The reaction mixture may be used to prepare one or more libraries (e.g., libraries of nucleic acid molecules, enzymes, or combinations thereof) or compositions for one or more sequencing methods. One or more components of the reaction mixture may be used simultaneously in the same reaction. In one example, the reaction may be carried out in a single reaction vial (e.g., a reaction tube), thereby reducing purification steps and / or sample loss, and / or allowing sequencing with a small amount of nucleic acid sample input. One or more components of the reaction mixture may be used separately in different reactions. The reaction mixture may be a cell-free reaction mixture. Alternatively, the reaction mixture may not be a cell-free reaction mixture.

[0237] The cleavage site can be no more than 100,000, 50,000, 10,000, 5,000, 1,000, 500, 400, 300, 200, 100, 90, 80, 70, 60, 50, 40, 30, 25, 20, 15, 10, 5 or fewer nucleotides from the target site. The cleavage site can be at least 5, 10, 15, 20, 25, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 1,000, 5,000, 10,000, 50,000, 100,000 or more nucleotides from the target site.

[0238] Circular nucleic acid molecules may include circular double-stranded nucleic acid molecules. A cleavage site may be part of the sense strand of the circular nucleic acid molecule. Alternatively or additionally, a cleavage site may be part of the antisense strand of the circular nucleic acid molecule. The method may also include determining the cleavage site based at least in part on the position of the target site relative to at least one reference sequence. The at least one reference sequence may comprise a common sequence of at least a portion of a gene. The cleavage site may be endogenous for the circular nucleic acid molecule. Therefore, determining the cleavage site may include selecting an endogenous sequence of the circular nucleic acid molecule. Alternatively, the cleavage site may be exogenous for the circular nucleic acid molecule. Therefore, the determination may include inserting an exogenous cleavage site into the circular nucleic acid molecule. In some instances, an endogenous sequence of the circular nucleic acid molecule may be selected, and then an exogenous cleavage site may be inserted within or near the endogenous sequence of the circular nucleic acid molecule. In this way, the distance between the exogenous cleavage site and the target site can be controlled or known.

[0239] The nucleic acid molecule may also contain an enzyme binding site specific to at least one enzyme. In some instances, the at least one enzyme may exhibit nicking enzyme activity, and the enzyme binding site may be identical to a nicking enzyme binding site, as provided in this disclosure.

[0240] The probability of an enzyme binding site appearing in the absence of any mismatches can be as high as 1 x 10^6. 4 5x10 4 7x10 4 1x10 5 5x10 5 1x10 6 5x10 6 1x10 7 5x10 7 1x10 8 5x10 8 1x10 9 5x10 9 1x10 10 5x10 10 1x1011 5x10 11 1x10 12 5x10 12 At most once in one or more nucleotides (or base pairs). Alternatively, the probability of an enzyme binding site appearing in the absence of any mismatches could be one per 1 x 10^6 nucleotides (or base pairs). 4 5x10 4 7x10 4 1x10 5 5x10 5 1x10 6 5x10 6 1x10 7 5x10 7 1x10 8 5x10 8 1x10 9 5x10 9 1x10 10 5x10 10 1x10 11 5x10 11 1x10 12 5x10 12 Up to two, three, four, five or more of the following nucleotides (or base pairs).

[0241] Enzyme binding sites may contain at least 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 35, 40, 45, or 50 or more nucleotides (or base pairs). Enzyme binding sites may contain up to 50, 45, 40, 35, 30, 29, 28, 27, 26, 25, 24, 23, 22, 21, 20, 19, 18, 17, 16, 15, 14, 13, 12, 11, 10, or 9 or fewer nucleotides (or base pairs).

[0242] The enzyme binding site can be endogenous for circular nucleic acid molecules. In one example, a linear nucleic acid molecule containing an endogenous enzyme binding site can be circularized to form a circular nucleic acid molecule. Alternatively, the enzyme binding site can be exogenous for circular nucleic acid molecules. In this case, the method may further include inserting the exogenous enzyme binding site into the circular nucleic acid molecule or its starting material (e.g., a linear nucleic acid molecule) before creating a nick. In some examples, the exogenous enzyme binding site can be inserted into the linear nucleic acid molecule before it is circularized into a circular nucleic acid molecule. The linear nucleic acid molecule may contain at least one recognition site as provided herein, and the recognition portion (e.g., a catalytically active recognition portion, such as Cas9) can be used to (i) bind to the recognition site, (ii) cleave the linear nucleic acid molecule at the recognition site, and (iii) facilitate the insertion of the exogenous enzyme binding site into at least one recognition site of the linear nucleic acid molecule or adjacent to at least one recognition site (e.g., by homology-directed repair). The method may further include circularizing the linear nucleic acid molecule after inserting the exogenous enzyme binding site. In other instances, the method may further include cyclizing the linear nucleic acid molecule into a circular nucleic acid molecule before inserting the nicking enzyme binding site into the circular nucleic acid molecule. After cyclization, the recognition portion (e.g., a catalytically active recognition portion, such as Cas9) can be used to (i) bind to the recognition site of the circular nucleic acid molecule, (ii) cleave the circular nucleic acid molecule at the recognition site, and (iii) facilitate the insertion of the exogenous enzyme binding site into at least one recognition site of the circular nucleic acid molecule or adjacent to at least one recognition site (e.g., via homology-directed repair).

[0243] The enzyme binding site may be no more than 30, 25, 20, 19, 18, 17, 16, 15, 14, 13, 12, 11, 10, 9, 8, 7, 6, 5, 4, 3, 2, or 1 nucleotide from the cleavage site. The enzyme binding site may be at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, or more nucleotides from the cleavage site. In some instances, the enzyme binding site may contain the cleavage site. In this case, the cleavage site may be within the enzyme binding site.

[0244] The reaction mixture further comprises at least a second enzyme that performs an extension reaction from the nick to produce a growth chain that is sequence complementary to at least a portion of the circular nucleic acid molecule. The circular nucleic acid molecule may be a circular double-stranded nucleic acid molecule, and the growth chain may exhibit sequence complementarity with at least a portion of the chain of the circular double-stranded nucleic acid molecule. The chain of the circular double-stranded nucleic acid molecule may be its sense strand, antisense strand, or both. In some instances, the at least second enzyme may be a polymerase. The reaction mixture may also comprise at least one nucleotide coupled to a tag, wherein the at least second enzyme incorporates the nucleotide into the growth chain. The tag may be a small molecule, nucleotide, polynucleotide, amino acid, polypeptide, polymer, metal and / or ceramic particle, etc. In some instances, each of the nucleotides G, C, A, T, and U may contain a distinct tag that is distinguishable from each other. In some instances, the tag may not be released from the nucleotide when it is incorporated into the growth chain. Alternatively, the at least second enzyme may release the tag from the nucleotide when it is incorporated into the growth chain. The at least second enzyme may perform the extension reaction with the aid of at least one oligonucleotide primer. Alternatively, at least the second enzyme can perform the extension reaction without using an oligonucleotide primer. In one example, the extension reaction may include RCA, wherein at least the second enzyme includes a polymerase. The polymerase may bind to the nick at the nick site for use in RCA.

[0245] The reaction mixture may further contain at least a third enzyme that performs a cleavage reaction from the nick to cleave at least a portion of the circular nucleic acid molecule. Starting from the nick, the at least third enzyme can displace and cleave at least a portion of the chain of the circular nucleic acid molecule. The at least third enzyme may include a nuclease (e.g., an endonuclease, such as a restriction endonuclease).

[0246] As provided in this disclosure, at least a portion of the circular nucleic acid molecule may have, or be suspected of having, one or more variants compared to at least one reference sequence. The reaction mixture can be used to prepare at least one composition for sequencing to identify the presence of at least a portion of the circular nucleic acid molecule.

[0247] As provided in this disclosure, circular nucleic acid molecules may contain a recognition sequence. Therefore, the reaction mixture may further contain a recognition moiety associated with the recognition sequence to enrich circular nucleic acid molecules from a random nucleic acid molecule library in the composition, at least partially based on the recognition sequence. The recognition moiety may contain at least one oligonucleotide (e.g., at least a portion of a gRNA for Cas system variants) that is complementary to at least the recognition sequence. The oligonucleotide of the recognition moiety may have at least 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95% or more complementarity to the recognition sequence. The oligonucleotide of the recognition moiety may have at most 100%, 95%, 90%, 80%, 70%, 60%, 50%, 40%, 30% or less complementarity to the recognition sequence.

[0248] The reaction mixture may comprise a library of selected circular nucleic acid molecules (e.g., single-stranded or double-stranded). Each circular nucleic acid molecule in at least 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 60%, 70%, 80%, 90%, 95% or more of the selected library may contain at least a portion of a recognition sequence. Each circular nucleic acid molecule in at most 100%, 95%, 90%, 80%, 70%, 60%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 15%, 10%, 5% or less of the selected library may contain at least a portion of a recognition sequence.

[0249] The probability of a recognized sequence appearing in the absence of any mismatches can be 1 x 10^6. 4 5x10 4 7x10 4 1x10 5 5x10 5 1x10 6 5x10 6 1x10 7 5x10 7 1x10 8 5x10 8 1x10 9 5x10 9 1x10 10 5x10 10 1x10 11 5x10 11 1x10 12 5x10 12 At most once in one or more nucleotides (or base pairs). Alternatively, the probability of a recognition sequence appearing without any mismatches could be in every 1 x 10^6 nucleotides. 4 5x10 4 7x104 1x10 5 5x10 5 1x10 6 5x10 6 1x10 7 5x10 7 1x10 8 5x10 8 1x10 9 5x10 9 1x10 10 5x10 10 1x10 11 5x10 11 1x10 12 5x10 12 Up to two, three, four, five or more of the following nucleotides (or base pairs).

[0250] The recognition sequence may contain at least 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 35, 40, 45, 50 or more nucleotides (or base pairs). The recognition sequence may contain up to 50, 45, 40, 35, 30, 29, 28, 27, 26, 25, 24, 23, 22, 21, 20, 19, 18, 17, 16, 15, 14, 13, 12, 11, 10, 9 or fewer nucleotides (or base pairs).

[0251] In various respects, this disclosure provides cell-free libraries of circular nucleic acid molecules. Each individual circular nucleic acid molecule in at least 5% of the library may contain (i) a target site and (ii) a nick at a known distance from the target site. The circular nucleic acid molecule library may be the starting material or product of any of the subject methods or reaction mixtures provided in this disclosure.

[0252] Cell-free libraries of circular nucleic acid molecules can be products or byproducts of enriching a random nucleic acid molecule library containing any nucleic acid molecule with a recognition site. In one example, a random circular nucleic acid molecule library may enrich at least one circular nucleic acid molecule containing a recognition site and treat it with at least one enzyme (e.g., a nicking enzyme, such as Cas9n nicking enzyme) to create a nick at a known distance from the target site. In another example, a random circular nucleic acid molecule library may be treated with at least one enzyme to create a nick at a known distance from the target site and enriched with at least one circular nucleic acid molecule containing a recognition site. In various other examples, at least one linear nucleic acid molecule containing a recognition site from a random linear nucleic acid molecule library may be enriched, treated with at least one first enzyme (e.g., a ligase and / or recombinase) to circularize one or more linear nucleic acid molecules, and treated with at least one second enzyme (e.g., a nicking enzyme, such as Cas9n nicking enzyme) to create a nick at a known distance from the target site. In various instances, at least one circular nucleic acid molecule containing a recognition site can be enriched from a random circular nucleic acid molecule library and treated with at least one enzyme (e.g., a recognition portion) to insert a nick site (which may or may not be present before insertion). In various instances, a random circular nucleic acid molecule library can be treated with at least one enzyme (e.g., a recognition portion) to insert a nick site (which may or may not be present before insertion) and enriched with at least one circular nucleic acid molecule containing a recognition site. In various instances, a random linear nucleic acid molecule library can be treated with at least a first enzyme (e.g., a recognition portion) to insert a nick site (which may or may not be present before insertion), treated with at least a second enzyme (e.g., a ligase and / or a recombinase) to circularize one or more linear nucleic acid molecules and enriched with at least one circular nucleic acid molecule containing a recognition site. In different instances, a random linear nucleic acid library may be treated with at least a first enzyme (e.g., a recognition moiety) to insert at a nick site (which may or may not exist prior to insertion), enriching at least one linear nucleic acid molecule containing the recognition site, and then treated with at least a second enzyme (e.g., a ligase and / or a recombinase) to circularize one or more linear nucleic acid molecules. In another different instance, at least one linear nucleic acid molecule containing a recognition site may be enriched from a random linear nucleic acid library, treated with at least a first enzyme (e.g., a recognition moiety) to insert at a nick site (which may or may not exist prior to insertion), and then treated with at least a second enzyme (e.g., a ligase and / or a recombinase) to circularize one or more linear nucleic acid molecules.

[0253] The circular nucleic acid molecules of each individual in at least 5%, 10%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95% or more of the library may contain (i) a target site and (ii) a cut at a known distance from the target site. The circular nucleic acid molecules of each individual in at most 100%, 95%, 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20% or less of the library may contain (i) a target site and (ii) a cut at a known distance from the target site.

[0254] Individual circular nucleic acid molecules may also contain recognition sequences. The probability of a recognition sequence appearing in the absence of any mismatches can be as high as 1 x 10^6. 4 5x10 4 7x10 4 1x10 5 5x10 5 1x10 6 5x10 6 1x10 7 5x10 7 1x10 8 5x10 8 1x10 9 5x10 9 1x10 10 5x10 10 1x10 11 5x10 11 1x10 12 5x10 12 At most once in one or more nucleotides (or base pairs). Alternatively, the probability of a recognition sequence appearing without any mismatches could be in every 1 x 10^6 nucleotides. 4 5x10 4 7x10 4 1x10 5 5x10 5 1x10 6 5x10 6 1x10 7 5x10 7 1x10 8 5x10 8 1x10 9 5x10 9 1x10 10 5x10 10 1x10 11 5x10 111x10 12 5x10 12 Up to two, three, four, five or more of the following nucleotides (or base pairs).

[0255] The recognition sequence may contain at least 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 35, 40, 45, 50 or more nucleotides (or base pairs). The recognition sequence may contain up to 50, 45, 40, 35, 30, 29, 28, 27, 26, 25, 24, 23, 22, 21, 20, 19, 18, 17, 16, 15, 14, 13, 12, 11, 10, 9 or fewer nucleotides (or base pairs).

[0256] Cell-free libraries may contain at least a first somatic nucleic acid molecule and a second somatic nucleic acid molecule. In some instances, (i) the first target site of the first somatic nucleic acid molecule and (ii) the second target site of the second somatic nucleic acid molecule may be the same. In this case, (1) the first known distance between the first nick of the first somatic nucleic acid molecule and the first target site and (2) the second known distance between the second nick of the second somatic nucleic acid molecule and the second target site may be the same. Alternatively, (1) the first known distance between the first nick of the first somatic nucleic acid molecule and the first target site and (2) the second known distance between the second nick of the second somatic nucleic acid molecule and the second target site may be different. In some instances, (i) the first target site of the first somatic nucleic acid molecule and (ii) the second target site of the second somatic nucleic acid molecule may be different. In this case, (1) the first known distance between the first nick of the first somatic nucleic acid molecule and the first target site and (2) the second known distance between the second nick of the second somatic nucleic acid molecule and the second target site may be the same. Alternatively, (1) the first known distance between the first nick of the first somatic nucleic acid molecule and the first target site and (2) the second known distance between the second nick of the second somatic nucleic acid molecule and the second target site can be different.

[0257] The cleavage site can be no more than 100,000, 50,000, 10,000, 5,000, 1,000, 500, 400, 300, 200, 100, 90, 80, 70, 60, 50, 40, 30, 25, 20, 15, 10, 5 or fewer nucleotides from the target site. The cleavage site can be at least 5, 10, 15, 20, 25, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 1,000, 5,000, 10,000, 50,000, 100,000 or more nucleotides from the target site.

[0258] The selected library may include circular double-stranded nucleic acid molecules. The cleavage can be on the sense strand or the antisense strand of the circular double-stranded nucleic acid molecule.

[0259] Figure 2A and 2B An example method for creating a nick at a known distance from the target site within a circular nucleic acid is illustrated schematically. (Refer to...) Figure 2A A circular nucleic acid 210a may be provided. The circular nucleic acid 210a may contain a target site 214 and a nicking enzyme binding site 212 specific to a nicking enzyme 220 (e.g., Cas9n). The target site 214 may have or be suspected of having one or more mutations compared to at least a reference sequence (e.g., a shared sequence of a portion of a gene from a species or multiple species). The nicking enzyme binding site 212 may be located at a known distance 216 from the target site 214 (e.g., the number of nucleotides between the nicking enzyme binding site 212 and the target site 214 is known). The circular nucleic acid 210a may also contain a recognition sequence 218 that can be specifically recognized by a recognition portion 230 (e.g., Cas or dCas). The circular nucleic acid 210a may be treated with a nicking enzyme 220. The nicking enzyme 220 may bind to the nicking enzyme binding site 212 and create a nick, thereby forming a circular nucleic acid 210b with a nick 222 located at, near, or within the nicking enzyme binding site 212. When nick 222 is formed, nickase 220 can be removed or isolated (e.g., automatically). Nickase binding site 212 can be endogenous or exogenous for circular nucleic acid 210a.

[0260] Reference Figure 2B A circular nucleic acid 220a may be provided. Circular nucleic acid 220a may include a target site 214 as provided herein and a recognition sequence 218, which may be specifically recognized by a recognition portion 230 (e.g., Cas). Circular nucleic acid 220a may be treated with the recognition portion 230 to create a break at, near, or within the recognition site 218. After the break is created, a nicking enzyme binding site 212 may be inserted (e.g., by homology-directed repair) into the break, and the circular nucleic acid may be blocked to form a circular nucleic acid molecule 220b. Subsequently, a nicking enzyme 220 may bind to the nicking enzyme binding site 212 and create a nick, thereby forming a circular nucleic acid molecule 220c with a nick 222 at, near, or within the nicking enzyme binding site 212. During the creation of the nick 222, the nicking enzyme 220 may be removed or isolated (e.g., automatically isolated).

[0261] Figure 2C and 2DAn example method for isolating or enriching circular nucleic acids containing recognition and target sites is illustrated schematically. (Refer to...) Figure 2C The random nucleic acid library may include a circular nucleic acid molecule 210a containing a target site 214, a nickase binding site 212 at a known distance 216 from the target site 214, and a recognition site 218. The random nucleic acid library may be treated with at least a recognition portion 150 (e.g., dCas) to form a recognition complex. In some instances, the recognition portion 150 may be conjugated to magnetic beads. Alternatively, the recognition portion 150 may contain one or more biotin molecules, which may subsequently be coupled to magnetic beads presenting avidin via an avidin-biotin interaction. The resulting recognition complex can be pulled out (separated) from other nucleic acid molecules lacking the recognition site 218 by magnetic bead separation. Similarly, the random nucleic acid library containing circular nucleic acid molecules 210b may be enriched (e.g., by using the recognition portion 150).

[0262] Reference Figure 2D The random nucleic acid library may contain circular nucleic acid molecules 220a, which contain a target site 214 and a recognition site 218. Figure 2C Using similar methods provided (e.g., by using a recognition portion), the library can be enriched with circular nucleic acid molecules 220a, or circular nucleic acid molecules 220a can be isolated from the library. Alternatively, the random nucleic acid molecule library may contain circular nucleic acid molecules 220b, which contain a target site 214, a nick site 212 at a known distance 216 from the target site 214, and a recognition site 218. Figure 2C Using similar methods provided (e.g., by using a recognition portion), the library can be enriched with circular nucleic acid molecules 220b, or circular nucleic acid molecules 220b can be isolated from the library. In different alternatives, the random nucleic acid molecule library may contain circular nucleic acid molecules 220c, which include a target site 214, a nick 222 at a known distance 216 from the target site 214, and a recognition site 218. Figure 2C Similar methods provided in the literature (e.g., by using an identification component) can enrich the library with circular nucleic acid molecule 220c, or allow circular nucleic acid molecule 220c to be isolated from the library.

[0263] Figure 3An example method for sequencing double-stranded nucleic acids is illustrated schematically. Double-stranded nucleic acid molecules can be provided. The double-stranded nucleic acid molecules can be fragmented or whole nucleic acid molecules from biological samples. The double-stranded nucleic acid molecules may contain two blunt ends. Alternatively, the double-stranded nucleic acid molecules may not contain two blunt ends (e.g., only one blunt end or no blunt ends). In this case, end repair may be required so that the repaired double-stranded nucleic acid molecules can (i) have no protruding ends, or (ii) contain a 5′ phosphate group and a 3′ hydroxyl group in both the sense and antisense strands, as shown in... Figure 1D Provided in [the document / source]. See reference [the document / source]. Figure 3 With or without such repair, double-stranded nucleic acid molecules can be denatured into separate single-stranded nucleic acid molecules. Each individual single-stranded nucleic acid molecule can be circularized to form a single-stranded circular nucleic acid molecule. In some instances, separate single-stranded nucleic acid molecules (e.g., sense and antisense strands that are at least partially or completely complementary to each other) can be coupled (e.g., linked) and then circularized into a single-stranded circular nucleic acid molecule. Subsequently, the circular nucleic acid molecule can be subjected to conditions sufficient to hybridize one or more random hexamer primers to the complementary domains of the circular nucleic acid molecule. The resulting circular nucleic acid molecule with one or more hybridized hexamers can be analyzed by synthetic sequencing, such as detection of electrical signals or visualization via optical imaging.

[0264] In one instance, nanopore sequencing can be used. A polymerase binds to one of the hybridized hexamers and initiates an extension reaction. During the extension reaction, other hybridized hexamers can be displaced from the circular nucleic acid molecule by the direct or indirect activity of the polymerase. The polymerase can be located in proximity to or coupled to the nanopore (e.g., a protein nanopore or a solid nanopore).

[0265] Whole genome and targeted sequencing

[0266] For whole-genome sequencing, the binding site of the cyclic polynucleotide polymerase can be left uncontrolled. Non-specific nickases can be used to generate nicks at random locations on the strands of the cyclic polynucleotide, and the polymerase can bind to the non-specific nick site, for example, to perform an extension reaction. Alternatively, the binding site of the cyclic polynucleotide polymerase can be left uncontrolled. Nickases that recognize specific sequences can be used to bind and generate nicks, such as the Cas system (e.g., Cas nickases, like Cas9n), N.AlwI, Nb.BbvCl, Nt.BbvCl, Nb.BsmI, Nt.BsmAI, Nt.BspQl, Nb.BsrDI, Nt.BstNBI, Nb.BstsCI, Nt.CviPII, Nb.BpulOI, Nt.BpulOI, and Nt,Bst9I, variants thereof, and combinations thereof.

[0267] Targeted polynucleotide sequencing can be used to detect sequence variations (e.g., one or more mutations) at specific locations within a polynucleotide. Sequence variations can be, for example, single nucleotide polymorphisms (SNPs). For targeted sequencing, it may be necessary to bind a polymerase near the polynucleotide sequence of interest (e.g., the target site), such as a segment that may contain sequence variants (e.g., mutations). In some instances, a Cas system comprising a Cas nickase and sgRNA can be used. A polynucleotide segment with base pairs complementary to the adjacent polynucleotide sequence of interest can be generated as part of the sgRNA. The sgRNA can bind to the Cas9n nickase to form an sgRNA / Cas9n complex, which can bind to the polynucleotide at a segment identified by at least a portion of the sgRNA and generate a nick.

[0268] Figure 4A An example method for targeted sequencing using nanopore sequencing is illustrated schematically. Samples comprising genomic DNA / cDNA or cell-free DNA / cfDNA can be amplified. To enrich the DNA / cDNA mixture with the targeted polynucleotide sequence, the genomic DNA / cDNA can be reacted with a biotinylated sgRNA / CRISPR / Cas9 complex to cleave the DNA / cDNA in the region of interest. The DNA mixture can be enriched with the target DNA segment by purification using streptavidin beads. The enriched target DNA sample or cell-free DNA (such as ctDNA) can then be circularized. The circular DNA can be bound to an sgRNA / CRISPR / Cas9n cleavage enzyme to provide a cleavage site in the DNA strand. The sgRNA contains a nucleotide sequence complementary to the nucleotide sequence of the DNA adjacent to the region of interest (such as a region potentially having a sequence variant). The polymerase is then bound to the cleavage site. The polymerase / DNA complex is then associated with a nanopore, and the DNA can be sequenced using rolling circle amplification and transcription.

[0269] Figure 4BAn example method for genome sequencing using nanopore sequencing is illustrated schematically. This method may include the following steps: providing a sample containing genomic DNA; amplifying the genomic DNA; circularizing the genomic DNA to provide circular DNA; nicking the circular DNA with a nicking enzyme to provide nick sites on the strand of the circular DNA; binding a DNA polymerase to the nick sites; and amplifying and sequencing the circular DNA using a nanopore. In some instances, restriction nicking enzymes may be used to generate nicks at their respective recognition sequences, thus allowing no control over the distance between the nick and the target site (e.g., a mutation site). Alternatively, nicking enzymes (e.g., Cas9n complexes) may be used to target a specific sequence of interest to generate a nick in or near that specific sequence, thereby controlling the distance between the nick and the target site (e.g., a mutation site). The nick site may be a portion of single-stranded DNA that has been removed to expose the 3′ and 5′ ends. The 3′ end may be used as a template that the polymerase can bind to and amplify from.

[0270] sample

[0271] The sample used for analysis may contain multiple polynucleotides. Polynucleotides may be single-stranded DNA, double-stranded DNA, or combinations thereof. Polynucleotides may contain genomic DNA, genomic cDNA, cell-free DNA, cell-free cDNA, or any combination thereof.

[0272] Polynucleotides can include cell-free DNA, circulating tumor DNA, genomic DNA, and DNA from formalin-fixed and paraffin-embedded (FFPE) samples. In some instances, DNA extracted from FFPE samples may be damaged, and such damaged DNA can be repaired using available FFPE DNA repair kits. Samples can contain any suitable DNA and / or cDNA samples, such as urine, feces, blood, saliva, tissue, biopsy material, body fluids, or tumor cells.

[0273] Multiple polynucleotides can be single-stranded or double-stranded.

[0274] Polynucleotide samples can come from any suitable source. For example, samples can be obtained from patients, animals, plants, or the environment, such as naturally occurring or artificial sources like the atmosphere, water systems, soil, atmospheric pathogen collection systems, underground sediments, groundwater, or wastewater treatment plants.

[0275] The polynucleotides from the sample may include one or more different polynucleotides, such as DNA, RNA, ribosomal RNA (rRNA), transfer RNA (tRNA), microRNA (miRNA), messenger RNA (mRNA), fragments of any of the foregoing, or any combination thereof. The sample may contain DNA. The sample may contain genomic DNA. The sample may contain mitochondrial DNA, chloroplast DNA, plasmid DNA, bacterial artificial chromosomes, yeast artificial chromosomes, oligonucleotide tags, or any combination thereof.

[0276] Polynucleotides can be single-stranded, double-stranded, or a combination thereof. A polynucleotide can be a single-stranded polynucleotide, in which double-stranded polynucleotides may or may not be present.

[0277] The initial amount of polynucleotides in the sample can be, for example, less than 50 ng, such as less than 45 ng, 40 ng, 35 ng, 30 ng, 25 ng, 20 ng, 15 ng, 10 ng, 5 ng, 4 ng, 3 ng, 2 ng, 1 ng, 0.5 ng, 0.1 ng or less. The initial amount of polynucleotides in the sample can also be, for example, greater than 0.1 ng, such as greater than 0.5 ng, 1 ng, 2 ng, 3 ng, 4 ng, 5 ng, 10 ng, 15 ng, 20 ng, 25 ng, 30 ng, 35 ng, 40 ng, 45 ng, 50 ng or more. The initial amount of polynucleotides can be, for example, from 0.1 ng to 100 ng, from 1 ng to 75 ng, from 5 ng to 50 ng, or from 10 ng to 20 ng.

[0278] Polynucleotides in a sample may be single-stranded upon acquisition or may become single-stranded through treatment (e.g., denaturation). Further examples of suitable polynucleotides are described herein with respect to various aspects of this disclosure. Polynucleotides can be subjected to subsequent steps (e.g., cyclization and amplification) without an extraction step and / or a purification step. For example, a fluid sample may be treated to remove cells without an extraction step to produce a purified liquid sample and a cell sample, from which polynucleotides can then be isolated. Various methods can be used to isolate polynucleotides, such as by precipitation or non-specific binding to a substrate, followed by washing the substrate to release the bound polynucleotides. If polynucleotides are isolated from a sample without a cell extraction step, the polynucleotides will be predominantly extracellular or “cell-free” polynucleotides, which may correspond to dead or damaged cells. The identity of such cells can be used to characterize, for example, the cells or cell populations from which they originate in a microbial community.

[0279] The sample can be from the subject. The subject can be any suitable organism, including, for example, plants, animals, fungi, protozoa, anucleate protozoa, viruses, mitochondria, and chloroplasts. Sample polynucleotides can be isolated from the subject, such as cell samples, tissue samples, body fluid samples, or organ samples, or cell cultures derived from any of these, including, for example, cultured cell lines, biopsies, blood samples, buccal swabs, or fluid samples containing cells such as saliva. The subject can be an animal, such as a cow, pig, mouse, rat, chicken, cat, dog, or mammal, such as a human. The sample can contain tumor cells, as in a sample from tumor tissue from the subject.

[0280] The sample may not contain intact cells, may be processed to remove cells, or may be used to isolate polynucleotides without a cell extraction step, such as to isolate cell-free polynucleotides, like cell-free DNA.

[0281] Other examples of sample sources include blood, urine, feces, nasal cavity, lungs, intestines, other bodily fluids or excretions, their derivatives, or combinations thereof.

[0282] Samples from a single individual can be divided into multiple separate samples, such as 2, 3, 4, 5, 6, 7, 8, 9, 10 or more separate samples, which are independently subjected to the methods of this disclosure, for example, by performing analyses in duplicate, triplicate, quadruplicate or more copies. When the sample comes from a subject, the reference sequence can also be derived from the subject, such as from a common sequence in the analyzed samples or a polynucleotide sequence from another sample or tissue from the same subject. For example, ctDNA mutations in a blood sample can be analyzed, and cellular DNA from another sample (such as a cheek or skin sample) from the subject can be analyzed to determine the reference sequence.

[0283] Depending on any suitable method, polynucleotides may or may not be extracted from the cells in the sample.

[0284] Multiple polynucleotides can include cell-free polynucleotides, such as cell-free DNA (cfDNA) or circulating tumor DNA (ctDNA). Cell-free DNA circulates in both healthy and diseased individuals. cfDNA (ctDNA) from tumors is not limited to any specific cancer type but appears to be a common finding across various malignant tumor conditions. The concentration of free circulating DNA in the plasma of control subjects may be lower than that of patients with or suspected of having the disease. In one instance, the concentration of free circulating DNA in plasma could be, for example, 14 ng / mL to 18 ng / mL in control subjects, while it could be 18 ng / mL to 318 ng / mL in tumor-forming patients.

[0285] Apoptosis and necrotizing cell death may contribute to the production of cell-free circulating DNA in bodily fluids. For example, significantly elevated levels of circulating DNA can be observed in the plasma of patients with prostate cancer and other prostate diseases such as benign prostatic hyperplasia and prostatitis. Furthermore, circulating tumor DNA may be present in fluids originating from the primary tumor organ. In one instance, breast cancer detection can be achieved in catheter irrigation fluid; colorectal cancer detection in feces; lung cancer detection in sputum; and prostate cancer detection in urine or ejaculate. Cell-free DNA can be obtained from a variety of sources. An exemplary source could be a blood sample from a subject. However, cfDNA or other fragmented DNA can be derived from a variety of other sources, including, for example, urine and fecal samples, which can be sources of cfDNA, including ctDNA.

[0286] The methods for sequencing multiple nucleotides provided in this disclosure may include retrieving a biological sample having a set of multiple nucleotides to be sequenced, extracting or otherwise isolating the multiple nucleotide sample from the biological sample, and optionally preparing the multiple nucleotide sample for sequencing.

[0287] Methods for sequencing polynucleotide samples may include isolating polynucleotides from biological samples (e.g., tissue samples, fluid samples) and preparing polynucleotide samples for sequencing. In some cases, polynucleotide samples are extracted from cells. Examples of techniques for extracting polynucleotides include using lysozyme, sonication, extraction, high pressure, or any combination thereof. In some cases, the polynucleotides are cell-free polynucleotides and do not require extraction from cells.

[0288] In some cases, polynucleotide samples can be prepared for sequencing through processes involving the removal of proteins, cell wall debris, and other components from the polynucleotide sample. Many commercial products are available to perform this operation, such as rotating columns. Alternatively or additionally, ethanol precipitation and centrifugation can be used.

[0289] Nucleic acid fragmentation

[0290] Polynucleotides from a sample can be fragmented prior to further processing. Fragmentation can be accomplished by any suitable method, including chemical, enzymatic, and mechanical fragmentation. The average or median length of the fragment is at least 10 nucleotides. The average or median length of the fragment can be at least 10, 50, 100, 500, 1,000, 5,000, 10,000, 50,000, or more nucleotides. The average or median length of the fragment is at most 50,000, 10,000, 5,000, 1,000, 500, 100, 50, 10, or fewer nucleotides. The fragment can be 90 to 200 nucleotides, and / or have an average length of 150 nucleotides or any other suitable average length. Fragmentation of polynucleotides can be performed mechanically, including subjecting the sample polynucleotides to acoustic treatment. Fragmentation can include treating the sample polynucleotides with one or more enzymes under conditions suitable for generating double-stranded polynucleotide breaks with one or more enzymes. Examples of fragmenting enzymes include sequence-specific and non-sequence-specific nucleases. Suitable examples of nucleases include DNase I, fragmentases, restriction endonucleases, their variants, and combinations thereof. Fragmentation may include treating sample polynucleotides with one or more restriction endonucleases. Fragmentation may produce fragments with 5' overhangs, 3' overhangs, blunt ends, or combinations thereof. When fragmentation involves the use of one or more restriction endonucleases, the cleavage of the sample polynucleotides leaves overhangs with predictable sequences. Fragmented polynucleotides may undergo size-selective fragmentation steps using standard methods, such as column purification or separation from agarose gels.

[0291] Linear nucleic acid amplification

[0292] Polynucleotides in a sample can be amplified. Polynucleotides can be amplified using a variety of methods, such as primer extension reactions using any suitable combination of primers and DNA polymerase, including but not limited to polymerase chain reaction (PCR), reverse transcription, and combinations thereof. When the template for primer extension is RNA, the reverse transcription product is called complementary DNA (cDNA). Primers that can be used for primer extension reactions can contain sequences specific to one or more targets, random sequences, partially random sequences, and combinations thereof.

[0293] The amplified polynucleotide can be sequenced with or without enrichment, such as by enriching one or more target polynucleotides in the amplified polynucleotide through an enrichment step prior to sequencing. The enrichment step may include hybridizing the amplified polynucleotide with multiple probes attached to a substrate. The enrichment step may include amplifying a target sequence comprising sequences A and B oriented in a 5' to 3' direction in an amplification reaction mixture comprising: (a) the amplified polynucleotide; (b) a first primer comprising sequence A', wherein the first primer specifically hybridizes to sequence A of the target sequence via sequence complementarity between sequences A and A'; (c) a second primer comprising sequence B, wherein the second primer specifically hybridizes to sequence B' present in a complementary polynucleotide comprising the complementary sequence of the target sequence via sequence complementarity between B and B'; and (d) a polymerase extending the first and second primers to produce the amplified polynucleotide; wherein the distance between the 5' end of sequence A and the 3' end of sequence B of the target sequence is 75 nt or less.

[0294] enrichment of samples

[0295] Regions of interest in genomic DNA and cDNA can be selectively targeted and amplified. Targeted regions may include, for example, regions containing sequence variants of interest for diagnostic purposes.

[0296] The polynucleotide region of interest (PRNA) can be cleaved by binding a biotinylated sgRNA / CRISPR / Cas9 complex to a PRNA region adjacent to the PRNA. The sgRNA may contain a sequence complementary to the PRNA sequence at the adjacent PRNA region. The sgRNA / CRISPR / Cas9 complex cleaves the double-stranded polynucleotide into PRNA segments. One or more purification methods, such as binding the target segment to streptavidin beads, can be used to enrich the composition containing both the target and non-target PRNA segments with the target PRNA.

[0297] Then, a mixture of polynucleotides rich in the polynucleotide segments of interest can be cyclized.

[0298] Linear nucleic acid circularization

[0299] Polynucleotide samples can contain single-stranded polynucleotides. Cyclic polynucleotides can be achieved by linking multiple polynucleotides together. A cyclic polynucleotide can have a unique linker within the cyclic polynucleotide.

[0300] Cycloning can include attaching the 5' end of a polynucleotide to the 3' end of the same polynucleotide, the 3' end of another polynucleotide in the sample, or the 3' end of a polynucleotide from a different source (e.g., an artificial polynucleotide, such as an oligonucleotide adaptor). For example, the 5' end of a polynucleotide can be attached to the 3' end of the same polynucleotide (also known as "self-ligation"). The conditions of the cyclization reaction can be selected to favor the self-ligation of polynucleotides within a specific length range, thereby producing multiple cyclized polynucleotides characterized by a specific average length. For example, the cyclization reaction conditions can be selected to favor the self-ligation of polynucleotides with lengths less than 5,000, 2,500, 1,000, 750, 500, 400, 300, 200, 150, 100, 50, or fewer nucleotides. Polynucleotide fragments with lengths of 50 to 5000 nucleotides, 100 to 2500 nucleotides, or 150 to 500 nucleotides may be advantageous, such that the average length of the cyclized polynucleotides is within the desired range. For example, 80% or more of the cyclized polynucleotide fragments can be between 50 and 500 nucleotides in length, such as between 50 and 200 nucleotides. Optimizable reaction conditions include the duration of the ligation reaction, the concentrations of various reagents, and / or the concentration of the polynucleotide to be ligated. The cyclization reaction can preserve the fragment length distribution in the sample before cyclization. For example, in both the uncyclized and cyclized polynucleotides, one or more of the mean, median, modality, and standard deviation of fragment lengths in the sample may be between 75%, 80%, 85%, 90%, 95%, or more of each other.

[0301] One or more adaptor oligonucleotides can be used such that the 5' and 3' ends of a polynucleotide in a sample are linked by one or more adaptor oligonucleotides interposed therebetween to form a cyclic polynucleotide, rather than a self-linked cyclic polynucleotide. For example, the 5' end of a polynucleotide can be linked to the 3' end of an adaptor, and the 5' end of the same adaptor can be linked to the 3' end of the same polynucleotide. The adaptor oligonucleotide can include any suitable oligonucleotide having a sequence, at least a portion of which is known, that can be linked to the sample polynucleotide. The adaptor oligonucleotide can include, for example, DNA, RNA, nucleotide analogs, non-canonical nucleotides, labeled nucleotides, modified nucleotides, or any combination thereof. The adaptor oligonucleotide can be single-stranded, double-stranded, or partially double-stranded. A partially double-stranded adaptor can contain one or more single-stranded regions and one or more double-stranded regions. A double-stranded adaptor can contain two separate oligonucleotides hybridizing to each other, such as an oligonucleotide duplex, and the hybridization can leave one or more blunt ends, one or more 3' overhangs, one or more 5' overhangs, one or more protrusions caused by mismatched and / or unpaired nucleotides, or any combination thereof. A "bubble" structure is formed when the two hybridization regions of an adaptor are separated from each other by non-hybridization regions. Adaptors with different nucleotide sequences can be used. Different adaptors can be ligated to the sample polynucleotide in sequential reactions or simultaneously. The same adaptor can be added to both ends of the target polynucleotide. For example, the first and second adaptors can be added to the same reaction. The adaptor can be manipulated before binding to the sample polynucleotide. For example, the terminal phosphate ester can be added or removed.

[0302] Any suitable method can be used for the cyclization of polynucleotides. For example, cyclization can include enzymatic reactions, such as the use of ligases, like RNA or DNA ligases. Examples of suitable ligases include Circligase. TM(Epicentre; Madison, Wis.), RNA ligases, T4 RNA ligase 1 (ssRNA ligase), NAD-dependent ligases such as Taq DNA ligase, Thermus filiformis DNA ligase, Escherichia coli DNA ligase, Tth DNA ligase, Thermus scotoductus DNA ligase (I and II), thermostable ligases, Ampligase thermostable DNA ligase, VanC-type ligase, 9°N DNA ligase, Tsp DNA ligase, ATP-dependent ligases such as T4 RNA ligase, T4 DNA ligase, T3 DNA ligase, T7 DNA ligase, Pfu DNA ligase, DNA ligase 1, DNA ligase III, and DNA ligase IV.

[0303] For self-ligation cyclization, the concentrations of polynucleotides and enzymes can be adjusted to promote the formation of intramolecular rings rather than intermolecular structures. Reaction temperature and time can also be adjusted. An exonuclease step can be included after the cyclization reaction to digest any unligated polynucleotides. For example, cyclic polynucleotides do not contain a free 5' or 3' terminus, so introducing a 5' or 3' exonuclease will not digest the cyclization, but will digest the unligated polynucleotide.

[0304] Linkages between the two ends of a polynucleotide, either directly or via one or more intermediate adaptor oligonucleotides, to form a cyclic polynucleotide can produce a linker with a linker sequence. When the 5' and 3' ends of a polynucleotide are linked by an adaptor polynucleotide, the linker can refer to the junction between the polynucleotide and the adaptor (e.g., one of the 5' or 3' end links), or to the linker formed by and including the adaptor polynucleotide between the 5' and 3' ends of the polynucleotide. If the 5' and 3' ends of the polynucleotide are linked without an intermediate adaptor, such as the 5' and 3' ends of a single-stranded DNA, the linker can refer to the point where these ends are joined. The linker can be identified by the nucleotide sequence containing it, i.e., the linker sequence. The sample contains polynucleotides with a mixture of ends formed through natural degradation processes (such as cell lysis, cell death, and others), which release DNA from the cell into its surrounding environment where it is further degraded, such as into cell-free polynucleotides. Fragmentation is a byproduct of sample processing (such as fixation, staining, and / or storage processes), and can be performed by methods that cut DNA without being limited to a specific target sequence, such as mechanical fragmentation or treatment with non-sequence-specific nucleases, such as using DNase I or fragmentases. In cases where the sample may contain polynucleotides with a mixture of ends, the probability that two polynucleotides have the same 5' or 3' ends is low, and the probability that two polynucleotides independently have the same 5' and 3' ends is extremely low. In such mixtures, even if two polynucleotides contain portions with the same target sequence, the linker can be used to distinguish the different polynucleotides. If the polynucleotide ends are joined in the absence of an intervening adaptor, the linker sequence can be identified by comparison with a reference sequence. For example, in cases where the order of the two component sequences appears to be reversed relative to the reference sequence, the point where the inversion occurs can indicate the linker. In cases where polynucleotide ends are joined via one or more adaptor sequences, the joiner can be identified by proximity to a known adaptor sequence, or by alignment as described above if the sequencing read is long enough to obtain the sequence from the 5' and 3' ends of the cyclic polynucleotide. The formation of a particular joiner can be a very rare event, making it unique among the cyclic polynucleotides in a sample.

[0305] Circular nucleic acid amplification

[0306] The methods provided in this disclosure may include amplifying cyclic polynucleotides. For example, multiple different cyclic polynucleotides comprising a target sequence, wherein the target sequence comprises sequence A and sequence B oriented in a 5' to 3' orientation, may be amplified. A method for amplifying circular DNA may include subjecting the circular DNA to a polynucleotide amplification reaction, wherein the reaction amplification reaction mixture comprises (a) a plurality of circular polynucleotides, wherein each of the plurality of circular polynucleotides comprises a different linker formed by circularizing an individual polynucleotide having a 5' end and a 3' end; (b) a first primer comprising sequence A', wherein the first primer specifically hybridizes to sequence A of a target sequence via sequence complementarity between sequence A and sequence A'; (c) a second primer comprising sequence B, wherein the second primer specifically hybridizes to sequence B' present in a complementary polynucleotide comprising a complementary sequence of the target sequence via sequence complementarity between sequence B and B'; and (d) a polymerase extending the first primer and the second primer to produce the amplified polynucleotide; wherein sequences A and B are endogenous sequences, and the distance between the 5' end of sequence A and the 3' end of sequence B of the target sequence is 75 nucleotides or less.

[0307] Following circularization, circular double-stranded polynucleotides can be amplified. Various methods exist for amplifying circular polynucleotides (e.g., DNA and / or RNA). Amplification can be linear, exponential, or include both linear and exponential phases simultaneously in a polyphase amplification process. Amplification methods can involve temperature changes, such as a thermal denaturation step, or can be isothermal processes that do not require thermal denaturation. Examples of suitable amplification processes include rolling circle amplification (RCA). In RCA, the reaction mixture can contain one or more primers, a polymerase, and dNTPs, producing a tandem. The polymerase in the RCA reaction can be a polymerase with strand displacement activity. A variety of suitable polymerases are available, including, for example, DNA polymerase I large (Klenow) fragment lacking exonuclease activity, Phi29 DNA polymerase, and Taq DNA polymerase. As a result of RCA, the resulting tandem polynucleotide amplification product has two or more copies of the target sequence from the template polynucleotide, such as 2, 3, 4, 5, 6, 7, 8, 9, 10, or more copies of the target sequence. Amplification primers can have any suitable length, for example, at least 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 90, 100 or more nucleotides, any part or all of which can be complementary to the corresponding target sequence to which the primers hybridize. The RCA process can use, for example, random primers, target-specific primers, adaptor-targeting primers, or no primers.

[0308] purification

[0309] After circularizing polynucleotides such as double-stranded DNA to provide circular dsDNA, the circular dsDNA can be purified prior to amplification or sequencing, for example, by isolating the circular polynucleotides or removing one or more other molecules from the reaction, to increase the relative concentration or purity of the circular polynucleotides for subsequent steps. For example, a mixture containing double-stranded polynucleotides can be treated with an exonuclease to remove the non-circularized polynucleotides (e.g., single-stranded or double-stranded), or size exclusion chromatography can be performed to trap and discard small reagents, or to trap and release the cyclization product in a separate volume. Purification may include treatment to remove or degrade the ligase used in the cyclization reaction, and / or to purify the cyclized polynucleotides from such ligase. Treatment to degrade the ligase may include treatment with a protease.

[0310] incision

[0311] Mixtures of cyclic polynucleotides, such as circular DNA, circular cDNA, circular ctDNA, or combinations thereof, can react with cleaving enzymes (e.g., the CRISPR / Cas9n complex) to cleave single-stranded segments of the cyclic polynucleotides, thereby creating nicks in the polynucleotide chain. For cyclic double-stranded polynucleotides, the CRISPR / Cas9n complex can bind to and cleave segments of either the inner or outer strand.

[0312] The CRISPR / Cas9n nickase complex can be used to remove segments, i.e., to form nicks on polynucleotide regions targeted by at least a portion of the gRNA of the CRISPR / Cas9n complex. This approach can be applied to whole-genome sequencing and / or targeted sequencing. In some instances, the CRISPR-Cas9n system can be adapted to target specific nucleotide sequences by complexing with a short RNA molecule that recognizes a specific DNA target (called a short guide RNA (sgRNA)). To target a specific polynucleotide region of interest, the sgRNA / CRISPR / Cas9n complex can be used to bind the nickase to a polynucleotide region adjacent to the region of interest. The nickase can expose the 3′ and 5′ ends of circular DNA, and the 3′ end can be used as a binding site for amplification and / or reverse transcription sequencing (e.g., using polymerases such as DNA polymerase).

[0313] Therapeutic applications

[0314] The methods, systems, and compositions provided herein can be used for one or more therapeutic applications, such as characterizing patient samples and optionally diagnosing a subject's condition. Therapeutic applications may include informing patients of treatment options that may be most responsive to them and / or informing subjects requiring therapeutic intervention based on the results of the methods provided herein.

[0315] For example, the methods provided in this disclosure can be used to diagnose the presence, progression, and / or metastasis of tumors, such as when the polynucleotides analyzed contain or are composed of cfDNA, ctDNA, or fragmented tumor DNA. For instance, tumor treatment efficacy in a subject can be monitored by monitoring ctDNA over time; a decrease in ctDNA can serve as an indicator of treatment efficacy, and an increase in ctDNA can inform the selection of different treatments and / or different dosages. Other uses include assessing organ rejection in transplant recipients, such as using an increase in the amount of circulating DNA corresponding to the transplant donor's genome as an early indicator of transplant rejection, and genotyping / isotyping of pathogen infections (such as viral or bacterial infections). Detection of sequence variants in circulating fetal DNA can be used to diagnose fetal conditions.

[0316] The methods provided in this disclosure may include diagnosing a subject based on sequencing results, such as diagnosing the subject with a disease associated with a detected causal genetic variant, or reporting the likelihood that a patient has or will develop such a disease.

[0317] Causal genetic variants can include sequence variants associated with a specific type or stage of cancer or cancer with specific characteristics such as metastatic potential, drug resistance, and / or drug responsiveness. The methods provided in this disclosure can be used to inform, guide, and monitor cancer treatment decisions. For example, treatment efficacy can be monitored by comparing ctDNA samples from patients before, during, and after treatment, including specific molecularly targeted therapies such as monoclonal drugs, chemotherapy drugs, radiation regimens, and combinations of any of the foregoing methods. For example, ctDNA can be monitored to see if certain mutations increase or decrease after treatment, or if new mutations appear, allowing physicians to modify treatment regimens in a shorter time than monitoring methods that track patient symptoms. Methods can include diagnosing subjects based on polynucleotide sequencing results, such as diagnosing the subject with a specific stage or type of cancer associated with a detected sequence variant, or reporting the likelihood that a patient has or will develop such cancer.

[0318] For example, in patient-specific therapies based on molecular markers, patients can be tested to detect the presence of certain mutations in their tumors, and these mutations can be used to predict response to or resistance to the therapy, guiding the decision to use that therapy. Detecting and monitoring ctDNA during treatment helps guide treatment selection.

[0319] Sequence variants associated with one or more cancers can be used for diagnostic, prognostic, or treatment decisions. For example, suitable oncology-significant target sequences include alterations in the TP53, ALK, KRAS, PIK3CA, BRAF, EGFR, and KIT genes. Target sequences can be specifically amplified, and / or sequence variants of the target sequence can be specifically analyzed to determine whether they are likely to be all or part of a cancer-related gene.

[0320] The methods provided in this disclosure can be used to discover novel rare mutations associated with one or more cancer types, stages, or cancer characteristics. For example, in a population of individuals sharing the analyzed characteristics such as a specific disease, cancer type, and / or cancer stage, the methods provided in this disclosure can be used to identify sequence variants reflecting mutations in a specific gene or part of a gene. Sequence variants identified that occur at a statistically significantly higher frequency in the group of individuals sharing the characteristic may have an association with that characteristic compared to individuals without the characteristic. The sequence variants or types of sequence variants thus identified can then be used to diagnose or treat individuals found to carry them.

[0321] Additional therapeutic applications may include use in noninvasive fetal diagnostics. Fetal DNA can be found in the blood of pregnant women. The methods provided in this disclosure can be used to identify sequence variants in circulating fetal DNA and therefore can be used to diagnose one or more genetic diseases in the fetus, such as diseases associated with one or more causal genetic variants. Examples of causal genetic variants include trisomy, cystic fibrosis, sickle cell anemia, and Tay-Saks disease. The mother can provide a control sample and a blood sample for comparison. The control sample can be any suitable tissue, which can then be sequenced to provide a reference sequence. The cfDNA sequence corresponding to the fetal genomic DNA can then be identified as a sequence variant relative to the maternal reference. The father can also provide a reference sample to aid in the identification of the fetal sequence and sequence variants.

[0322] Different therapeutic applications may include the detection of exogenous polynucleotides, including those from pathogens such as bacteria, viruses, fungi, and microorganisms, which can provide information to guide treatment.

[0323] sequencing equipment

[0324] Impedance-based sequencing

[0325] In one aspect, this disclosure provides methods for processing or analyzing nucleic acid molecules. The method may include providing a nucleic acid molecule adjacent to a nanopore. The method may include contacting the nucleic acid molecule with a tagged nucleotide under conditions sufficient to incorporate a nucleotide into a nucleic acid chain complementary to at least a portion of the nucleic acid molecule. When the nucleotide is incorporated into the nucleic acid chain, at least a portion of the tag may be placed within the nanopore. The method may include detecting one or more signals indicating impedance or impedance changes in the nanopore while at least a portion of the tag is within the nanopore. The method may include using one or more signals to identify the nucleotide incorporated into the nucleic acid chain. The method may also include measuring a current or a change thereof while at least a portion of the tag is placed within the nanopore.

[0326] Nanopores can be disposed in close proximity to or adjacent to electrodes of a sensing circuit or coupled to a circuit (e.g., a CMOS or FET circuit). The circuit can be coupled to a voltage source. Alternatively, the nanopore can be part of a circuit. A constant voltage can be applied to the circuit, and changes in current can be measured. Alternatively, the voltage change required to maintain a steady-state current can be measured. The nanopore can be part of a circuit containing a tunnel junction. The nanopore can be a tunnel junction of the nanopore. Alternatively, the nanopore may not be part of a circuit containing a tunnel junction. The nanopore can be in an electrolyte solution (e.g., 0.5 M potassium acetate and 10 mM KCl). Alternatively, the nanopore may not be in an electrolyte solution.

[0327] One or more signals may be current or voltage measured from the sensing circuit. One or more signals may be both current and voltage measured from the sensing circuit. The signal may be a tunneling current. Alternatively, the signal may not be a tunneling current. The current may be a Faraday current. Alternatively, the current may not be a Faraday current. The current may be at least 1 picoampere (pA), 10 pA, 100 pA, 1 nanoampere (nA), 10 nA, 100 nA, 1 microampere (mA), 10 mA, 100 mA, or higher. The current may be at most 100 mA, 10 mA, 1 mA, 100 nA, 10 nA, 1 nA, 100 pA, 10 pA, 1 pA, or less. The current may be at least in the picoampere (pA) range, in the tens of pA range, in the hundreds of pA range, in the nanoampere (nA) range, in the tens of nA range, in the hundreds of nA range, in the microampere (mA) range, in the tens of mA range, or higher. The current can be in the range of tens of mA, mA, hundreds of nA, tens of nA, nA, hundreds of pA, tens of pA, pA, or lower. The voltage can be at least 0.1 mV, 0.5 mV, 1 mV, 5 mV, 10 mV, 50 mV, 100 mV, 500 mV, or higher. The voltage can be at most 500 mV, 100 mV, 50 mV, 10 mV, 5 mV, 1 mV, 0.5 mV, 0.1 mV, or lower. The voltage can be at least in the range of millivolts (mV), tens of mV, hundreds of mV, or higher. The voltage can be at most in the range of hundreds of mV, tens of mV, mV, or lower.

[0328] The circuit may include multiple electrodes (e.g., metal electrodes). The circuit may include at least 2, 3, 4, 5, 6, 7, 8, 9, 10, or more electrodes. The circuit may include up to 10, 9, 8, 7, 6, 5, 4, 3, or 2 electrodes. Multiple electrodes may not be in direct contact with the nanopore. Alternatively, multiple electrodes may be in direct contact with the nanopore. In another alternative, some electrodes may be in direct contact with the nanopore, while others may not. In some cases, the nanopore may contain multiple electrodes. Alternatively, the nanopore may not contain multiple electrodes. Identifying nucleotides incorporated into nucleic acid chains may include using multiple electrodes to detect one or more signals.

[0329] Nanopores can comprise protein nanopores or solid nanopores. Nanopores can have, for example, a characteristic width or diameter from about 0.1 nm to 1,000 nm. The width or diameter of the nanopore can be at least 0.1 nm, 0.5 nm, 1 nm, 5 nm, 10 nm, 50 nm, 100 nm, 500 nm, 1,000 nm or larger. The width or diameter of the nanopore can be at most 1,000 nm, 500 nm, 100 nm, 50 nm, 10 nm, 5 nm, 1 nm, 0.5 nm, 0.1 nm or smaller.

[0330] The method may further include releasing a tag from the nucleotide during incorporation into the nucleic acid chain before detecting one or more signals indicative of impedance or impedance change in the nanopore. At least one enzyme may incorporate the nucleotide into a nucleic acid chain complementary to at least a portion of the nucleic acid molecule. At least one enzyme may further release the tag from the nucleotide before, during, or after incorporation. Alternatively or additionally, an additional enzyme (which may be operatively coupled to at least one enzyme) may release the tag from the nucleotide before, during, or after incorporation. At least a portion of the released tag may enter the nanopore, and the method may include detecting one or more signals indicative of impedance or impedance change in the nanopore while at least a portion of the released tag is within the nanopore.

[0331] Nucleic acid molecules include circular nucleic acid molecules. Circular nucleic acid molecules can be single-stranded or double-stranded. Alternatively, the first part of a circular nucleic acid molecule can be single-stranded, and the second part can be double-stranded. In some instances, nucleic acid molecules can include single-stranded or double-stranded linear nucleic acid molecules.

[0332] Incorporation can be performed using at least one oligonucleotide primer. Alternatively, incorporation can be performed without the use of an oligonucleotide primer. The method may also include subjecting the nucleic acid molecule to RCA to generate a nucleic acid chain before detecting one or more signals indicating impedance or impedance changes in the nanopore. RCA can generate a nucleic acid chain containing at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more copies of at least a portion of the nucleic acid molecule (e.g., at least a portion of at least one strand of the nucleic acid molecule). RCA can generate a nucleic acid chain containing at most 10, 9, 8, 7, 6, 5, 4, 3, 2 or 1 copies of at least a portion of the nucleic acid molecule.

[0333] The provision may include coupling at least one enzyme that performs incorporation into (i) at least a portion of a nanopore or (ii) a membrane having nanopores. Coupling may include conjugating at least one enzyme to a nanopore or membrane. Coupling may include conjugating at least one enzyme to both a nanopore and a membrane. Coupling may be covalent (e.g., cross-linked or conjugated). Coupling may be carried out by another enzyme, such as transglutaminase, sorting enzyme, subtilisin, tyrosinase, laccase, etc., or by a chemical cross-linking agent, such as 1-ethyl-3-(3-dimethylaminopropyl)carbodiimide (EDC), N,N′-dicyclohexylcarbodiimide (DCC), N,N′-diisopropylcarbodiimide (DIC), etc. Alternatively or additionally, coupling may be non-covalent, for example, through hydrogen bonds, magnetic interactions, etc.

[0334] The membrane can be a lipid bilayer. The membrane can be a solid membrane (e.g., a thin film). The membrane can be a combination of a lipid bilayer and a solid membrane. At least one enzyme can be a polymerase, a nuclease, a functional variant thereof, or a combination thereof.

[0335] In another aspect, this disclosure provides a system for processing or analyzing nucleic acid molecules. The system may include a nanopore configured to receive at least a portion of a tag when a tagged nucleotide is incorporated into a nucleic acid chain. The nucleic acid chain may be complementary to at least a portion of the nucleic acid molecule. When at least a portion of the tag is within the nanopore, the nanopore may be configured to detect one or more signals indicating impedance or impedance changes within the nanopore. The one or more signals may be used to identify the nucleotide incorporated into the nucleic acid chain. The system for processing or analyzing nucleic acid molecules may be configured to perform one or more of the subject methods provided in this disclosure for processing or analyzing nucleic acid molecules or derivatives thereof.

[0336] One or more signals can be current or voltage. One or more signals can be both current and voltage. The current can be a Faraday current. Alternatively, the current may not be a Faraday current. One or more signals may not be tunneling current. Alternatively, one or more signals can be tunneling current. The nanopore can be part of a circuit containing a tunnel junction. Alternatively, the nanopore may not be part of a circuit containing a tunnel junction.

[0337] A nanopore can be configured to measure current or changes thereof when at least a portion of a tag is placed within the nanopore. Alternatively, the nanopore can be configured to measure current or changes thereof when at least a portion of the tag is released from the nanopore. The nanopore may include multiple electrodes configured to detect one or more signals. Alternatively, the nanopore may not include multiple electrodes configured to detect one or more signals, and the multiple electrodes may be operatively coupled to the nanopore to detect one or more signals.

[0338] Nanopores can contain protein nanopores or solid nanopores.

[0339] The system may also include at least one enzyme configured to perform incorporation. The at least one enzyme may incorporate a nucleotide into a nucleic acid chain complementary to at least a portion of a nucleic acid molecule. When the nucleotide is incorporated into the nucleic acid chain, at least a portion of the tag may be released from the nucleotide. The at least one enzyme may release the tag from the nucleotide before, during, or after incorporation. Alternatively or additionally, an additional enzyme (which may be operatively coupled to the at least one enzyme) may release the tag from the nucleotide before, during, or after incorporation. At least a portion of the released tag may enter a nanopore, and the method may include detecting one or more signals indicating impedance or impedance changes in the nanopore while at least a portion of the released tag is within the nanopore. The at least one enzyme may be a polymerase, a nuclease, a functional variant thereof, or a combination thereof.

[0340] Incorporation can be performed using at least one oligonucleotide primer. Therefore, the system may also contain at least one oligonucleotide primer. Alternatively, incorporation can be performed without using an oligonucleotide primer. Therefore, the system may not contain an oligonucleotide primer.

[0341] At least one enzyme (and / or additional enzymes) may be coupled to (i) at least a portion of a nanopore or (ii) a membrane having nanopores. The membrane may be a lipid bilayer or a solid membrane.

[0342] Coupling can be achieved by coupling enzymes and / or chemical cross-linking agents. The system may also contain coupling enzymes (e.g., transglutaminase, sorting enzymes, subtilisin, tyrosinase, laccase, etc.) or chemical cross-linking agents (e.g., EDC, DCC, DIC, etc.). Alternatively, the nanopore or membrane can be configured to bind to at least a portion of at least one enzyme (and / or additional enzymes). The nanopore or membrane may contain a binding portion (e.g., small molecule, nucleotide, peptide, polymer, combination thereof, etc.) capable of binding to at least a portion of at least one enzyme. In various alternatives, at least one enzyme may be configured to bind to at least a portion of the nanopore or at least a portion of the membrane. At least one enzyme may contain a binding portion (e.g., small molecule, nucleotide, peptide, polymer, combination thereof, etc.) capable of binding to at least a portion of the membrane.

[0343] Figures 5A to 5D An exemplary nanopore sequencing system for obtaining sequence information from one or more nucleic acid samples is schematically illustrated. (Reference) Figure 5A The nanopore sequencing system 510 may include a membrane 512 comprising at least one nanopore 514 (a cross-section of the nanopore 514 is shown). The membrane 512 may be a lipid bilayer and / or a solid membrane. The nanopore 514 may include a plurality of electrodes 516 configured to detect one or more signals from a circuit comprising the nanopore 514. The plurality of electrodes 516 may be disposed on one side of the membrane 512. The plurality of electrodes 516 may be coupled to the nanopore 514. The circuit may also include an ammeter and a voltage source. A nucleic acid molecule 520 may be provided near the nanopore 514. By using an enzyme 530 (e.g., a polymerase), the nucleic acid molecule 520 may be contacted with a nucleotide 541 having a tag 542, provided that conditions are sufficient to incorporate a nucleotide 541 into a nucleic acid chain 540 complementary to at least a portion of the nucleic acid molecule 520. For example, different types of nucleotides may have different tags 542, 544, 546, and 548, respectively. When nucleotide 541 is incorporated into nucleic acid chain 540, at least a portion of tag 542 may be placed within nanopore 514. When at least a portion of tag 542 is within the nanopore, one or more signals indicative of impedance or impedance changes within nanopore 514 can be detected. The one or more signals may include current or changes thereof. In this example, when tag 542 is disposed within nanopore 514, tag 542 may be attached to nucleic acid chain 540. One or more signals may be used to identify nucleotide 541 incorporated into nucleic acid chain 540.

[0344] refer to Figure 5BThe nanopore sequencing system 510 may include a membrane 512 comprising at least one nanopore 514 (a cross-section of the nanopore 514 is shown). The membrane 512 may be a lipid bilayer and / or a solid membrane. The nanopore 514 may include a plurality of electrodes 516 configured to detect one or more signals from a circuit comprising the nanopore 514. The plurality of electrodes 516 may be coupled to the nanopore 514. The circuit may also include an ammeter and a voltage source. A nucleic acid molecule 520 may be provided in proximity to the nanopore 514. By using an enzyme 530 (e.g., a polymerase), the nucleic acid molecule 520 may be contacted with a nucleotide 541 having a tag 542, provided that conditions are sufficient to incorporate a nucleotide 541 into a nucleic acid chain 540 complementary to at least a portion of the nucleic acid molecule 520. For example, different types of nucleotides may have different tags 542, 544, 546, and 548, respectively. When nucleotide 541 is incorporated into nucleic acid chain 540, tag 542 can be released from nucleotide 541, and at least a portion of the released tag 542 can be placed within nanopore 514. When at least a portion of the released tag 542 is within the nanopore, one or more signals indicating impedance or impedance change can be detected in nanopore 514. One or more signals can be used to identify nucleotide 541 incorporated into nucleic acid chain 540.

[0345] refer to Figure 5C The nanopore sequencing system 510 may include a membrane 512 comprising at least one nanopore 514 (a cross-section of the nanopore 514 is shown). The membrane 512 may be a lipid bilayer and / or a solid membrane. The nanopore 514 may be operatively coupled to a plurality of electrodes 516 configured to detect one or more signals transmembrane 512 from a circuit. The nanopore sequencing system 512 may be in an electrolytic solution. The circuit may also include an ammeter and a voltage source. A nucleic acid molecule 520 may be provided in proximity to the nanopore 514. By using an enzyme 530 (e.g., a polymerase), the nucleic acid molecule 520 may be contacted with a nucleotide 541 having a tag 542, provided that conditions are sufficient to incorporate a nucleotide 541 into a nucleic acid chain 540 complementary to at least a portion of the nucleic acid molecule 520. For example, different types of nucleotides may have different tags 542, 544, 546, and 548, respectively. When incorporating the nucleotide 541 into the nucleic acid chain 540, at least a portion of the tag 542 may be placed within the nanopore 514. When at least a portion of tag 542 is within the nanopore, one or more signals indicative of impedance or impedance changes within the nanopore 514 can be detected. The one or more signals may include current or changes thereof. In this example, when tag 542 is placed within the nanopore 514, tag 542 can be attached to the nucleic acid chain 540. One or more signals can be used to identify nucleotides 541 incorporated into the nucleic acid chain 540.

[0346] refer to Figure 5D The nanopore sequencing system 510 may include a membrane 512 comprising at least one nanopore 514 (a cross-section of the nanopore 514 is shown). The membrane may be a lipid bilayer and / or a solid membrane. The nanopore 514 may include a plurality of electrodes 516 configured to detect one or more signals from a circuit comprising the nanopore 514. The plurality of electrodes 516 may be disposed on opposite sides of the membrane 512. The plurality of electrodes 516 may be coupled to the nanopore 514. The circuit may also include an ammeter and a voltage source. A nucleic acid molecule 520 may be provided in proximity to the nanopore 514. By using an enzyme 530 (e.g., a polymerase), the nucleic acid molecule 520 may be contacted with a nucleotide 541 having a tag 542, provided that conditions are sufficient to incorporate a nucleotide 541 into a nucleic acid chain 540 complementary to at least a portion of the nucleic acid molecule 520. For example, different types of nucleotides may have different tags 542, 544, 546, and 548, respectively. When nucleotide 541 is incorporated into nucleic acid chain 540, at least a portion of tag 542 may be placed within nanopore 514. When at least a portion of tag 542 is within the nanopore, one or more signals indicative of impedance or impedance changes within nanopore 514 can be detected. The one or more signals may include current or changes thereof. In this example, when tag 542 is placed within nanopore 514, tag 542 may be attached to nucleic acid chain 540. One or more signals may be used to identify nucleotide 541 incorporated into nucleic acid chain 540.

[0347] exist Figure 5A , Figure 5B and 5D In this scenario, the nanopore 514 can be a solid nanopore, for example, a pore or channel guided through a solid substrate. Figure 5C In this context, nanopores can be porins, such as α-hemolysin molecules embedded in a lipid bilayer.

[0348] Equipment Overview

[0349] The methods and systems provided in this disclosure can be executed, and sequencing data can be acquired using any suitable sequencing equipment, such as equipment capable of performing massively parallel sequencing reactions. For example, a high-throughput sequencing system can be used.

[0350] Polynucleotide sequences can be analyzed, for example, to identify repeat unit lengths (e.g., monomer length), linkers formed by circularization, and any true variations relative to a reference sequence. Identifying repeat unit lengths can include calculating the regions of the repeat unit, finding reference sites for the sequence (e.g., when targeting one or more sequences for amplification, enrichment, and / or sequencing), the boundaries of each repeat region, and / or the number of repeats per sequencing run. Sequence analysis can include analyzing sequence data from both strands of a duplex. For example, identical variants of read sequences from different polynucleotides in a sample (e.g., circularized polynucleotides with different linkers) can be considered confirmed variants. If a sequence variant appears in more than one repeat unit of the same polynucleotide, it can also be considered a true variant, as the same sequence variation is also unlikely to appear at the same position in a target sequence repeated in the same tandem. When identifying variants and confirmed variants, sequence quality scores can be considered; for example, sequences and bases with quality scores below a threshold can be filtered out. Other bioinformatics methods can be used to further improve the sensitivity and specificity of variant identification.

[0351] For example, the system for detecting sequence variants provided in this disclosure may include: (a) a computer configured to receive a user's request to perform a detection reaction on a sample; (b) a polynucleotide preparation system that, in response to a user request, performs a polynucleotide amplification reaction on a sample or a portion thereof, wherein the amplification reaction includes the steps of: (i) cyclizing individual polynucleotides to form a plurality of cyclic polynucleotides, wherein each cyclic polynucleotide has a linker between a 5' end and a 3' end; and (ii) amplifying the cyclic polynucleotides, or amplifying double-stranded polynucleotides and cyclizing the amplified double-stranded polynucleotides; (iii) Creating a nick on the single strand of the cyclic polynucleotide to provide a nick site; (iv) coupling a polymerase to the nick site to provide a polymerase / cyclic polynucleotide complex; (v) associating the polymerase / cyclic polynucleotide complex with a nanopore; and (c) a sequencing system that generates sequencing reads of the polynucleotide, identifies sequence differences between the sequencing reads and a reference sequence, and determines sequence differences occurring in at least two cyclic polynucleotides with different linkages as sequence variants; and (d) a report generator that sends a report to a recipient containing the results of the sequence variant detection. In some embodiments, the recipient is a user. In some cases, the sequencing device may include a sensor array for sequencing nucleic acids, such as a nanopore array.

[0352] Each nanopore sequencing complex can be intercalated into a membrane, for example, a lipid bilayer, and positioned at a sensing electrode immediately adjacent to or near a sensing circuit, such as an integrated circuit of a nanopore-based sensor. Multiple nanopore sensors can be provided as an array, such as an array present on a chip or biochip. The nanopore array can have any suitable number of nanopores. The array can include approximately 200, approximately 400, approximately 600, approximately 800, approximately 1000, approximately 1500, approximately 2000, approximately 3000, approximately 4000, approximately 5000, approximately 10000, approximately 15000, approximately 20000, approximately 40000, approximately 60000, approximately 80000, approximately 100000, approximately 200000, approximately 400000, approximately 600000, approximately 800000, approximately 1000000 or more nanopores (or nanopore sequencing complexes).

[0353] During sequencing using one or more labeled nucleotides (or one or more labeled polynucleotides), the labeled nucleotides can be incorporated into the nanopore sequencing complex using an enzyme (e.g., polymerase). During polymerization, the tag can be detected through the nanopore, either by releasing and delivering the tag into or through the nanopore, or by presenting it to the nanopore. A single tag can be released and / or presented upon incorporation of a single nucleotide and detected through the nanopore. Multiple tags can be released and / or presented upon incorporation of multiple nucleotides. A nanopore sensor adjacent to (or coupled to) the nanopore can detect a single tag or multiple tags. One or more signals associated with multiple tags can be detected and processed to generate an average signal. The tag can be detected by the sensor as a function of time. Tags detected over time can be used to determine the nucleic acid sequence of a polynucleotide sample, such as by means of a computer system programmed to record sensor data and generate sequence information from that data.

[0354] Any device and system suitable for sequencing via RCA and transcription can be used. In some cases, the sequencing system can generate sequencing reads for polynucleotides amplified by the amplification system, identify sequence differences between the sequencing reads and a reference sequence, and determine sequence differences appearing in at least two circular polynucleotides with different linkers as sequence variants. The sequencing system and amplification system can be the same or include one or more overlapping devices. In one instance, the amplification system and sequencing system can utilize the same thermal cycler. A variety of sequencing platforms can be used, and selection can be based on the chosen sequencing method. Amplification and sequencing may involve the use of liquid processors. Several commercially available liquid handling systems can be used to automate these processes.

[0355] Sequencing systems may include, for example, computers, computer-readable media including computer-executable code, storage devices, communication devices, control algorithms, analysis algorithms, and / or reporting algorithms.

[0356] Sequencing equipment can be used to detect sequence variants. Detection of sequence variants can include detecting mutations, such as rare somatic mutations relative to a reference sequence or in a mutation-free background, where the sequence variant is associated with a disease. Sequence variants with statistical, biological, and / or functional evidence of association with a disease or trait are called “causal genetic variants.” A single causal genetic variant may be associated with more than one disease or trait. Causal genetic variants can be associated with Mendelian traits, non-Mendelian traits, or both. Causal genetic variants can manifest as variations in polynucleotides, such as 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 50, or more sequence differences (e.g., sequence differences at the same relative genomic position between a polynucleotide containing a causal genetic variant and a polynucleotide lacking a causal genetic variant). Examples of causal genetic variant types include single nucleotide polymorphisms (SNPs), deletion / insertion polymorphisms (DIPs), copy number variants (CNVs), short tandem repeats (STRs), restriction fragment length polymorphisms (RFLPs), simple sequence repeats (SSRs), variable number tandem repeats (VNTRs), random amplified polymorphic DNA (RAPDs), amplified fragment length polymorphisms (AFLPs), intertransposon amplification polymorphisms (IRAPs), long and short scattered elements (LINEs / SINEs), long tandem repeats (LTRs), mobile elements, retrotransposon microsatellite amplification polymorphisms, retrotranscribed insertion polymorphisms, sequence-specific amplification polymorphisms, and heritable epigenetic modifications, such as DNA methylation. Causal genetic variants can also be a group of closely related causal genetic variants. Some causal genetic variants can exert their influence as sequence variations in RNA polynucleotides. At this level, some causal genetic variants are also indicated by the presence or absence of a certain RNA polynucleotide. Some causal genetic variants result in sequence variations in protein polypeptides. Many causal genetic variants have been reported. Examples of SNP causal variants include the HbS variant of hemoglobin, which causes sickle cell anemia. Examples of DIP causal variants include the δ508 mutation in the CFTR gene, which causes cystic fibrosis. Examples of CNV causal variants include trisomy 21, which causes Down syndrome. Examples of STR causal variants include tandem repeat sequences, which cause Huntington's disease.

[0357] Nanoporous devices

[0358] The sequencing system may include a reaction chamber comprising one or more nanopore devices. The nanopore device may be an individually addressable nanopore device. An individually addressable nanopore may be individually readable. An individually addressable nanopore may be individually writable. An individually addressable nanopore may be both individually readable and individually writable. The system may include one or more computer processors for facilitating sample preparation and various operations of this disclosure, such as polynucleotide sequencing. The processor may be coupled to the nanopore device.

[0359] Nanopore devices can include multiple individually addressable sensing electrodes. Each sensing electrode can include a membrane adjacent to the electrode, and one or more nanopores within the membrane. The nanopores can be in a membrane such as a lipid bilayer, adjacent to or sensorily close to an electrode that is part of or coupled to an integrated circuit. The nanopores can be associated with a single electrode and a sensing integrated circuit or with multiple electrodes and sensing integrated circuits. The nanopores can comprise solid-state nanopores.

[0360] The apparatus and systems used in the methods provided by this disclosure can accurately detect individual nucleotide incorporation events, such as when a nucleotide is incorporated into a growing chain complementary to a template. Enzymes, such as DNA polymerases, RNA polymerases, or ligases, can incorporate nucleotides into a growing polynucleotide chain. Enzymes such as polymerases can generate polynucleotide chains.

[0361] The added nucleotide can be complementary to the corresponding template polynucleotide chain, which hybridizes with the growing chain. The nucleotide can include a tag or tagging material coupled to any position of the nucleotide, including but not limited to phosphates of the nucleotide such as gamma-phosphates, sugars, or nitrogenous base moieties. In some cases, the tag is detected during nucleotide tag incorporation when it associates with polymerase. The tag is detectable until it translocates through the nanopore after nucleotide incorporation and subsequent tag cleavage and / or release. Nucleotide incorporation events can release the tag from the nucleotide, which then passes through the nanopore and is detected. The tag can be released by polymerase or cleaved / released in any suitable manner, including but not limited to cleavage by an enzyme located near the polymerase. In this way, the incorporated base (i.e., A, C, G, T, or U) can be identified because a unique tag is released from each type of nucleotide (i.e., adenine, cytosine, guanine, thymine, or uracil). In non-released nucleotide incorporation events, the tag coupled to the incorporated nucleotide is detected by means of the nanopore. In some instances, the tag can move through or near the nanopore and be detected using the nanopore.

[0362] The methods and systems of this disclosure can enable the detection of polynucleotide incorporation events, for example, at a resolution of at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 100, 500, 1000, 5000, 10000, 50000, or 100000 polynucleotide bases within a given time period. For example, nanopore devices can be used to detect individual polynucleotide incorporation events, each associated with an individual nucleic acid base. In other instances, nanopore devices can be used to detect events associated with multiple bases. For example, the signal sensed by the nanopore device can be a combination signal from at least 2, 3, 4, or 5 bases.

[0363] In some sequencing methods, the tag does not pass through the nanopore. The tag can be detected through the nanopore and exit without passing through it, such as exiting in the opposite direction to where it entered the nanopore. The sequencing device can be configured to actively expel the tag from the nanopore.

[0364] In some sequencing methods, the tag is not released after a nucleotide incorporation event. A nucleotide incorporation event can present the tag to the nanopore without releasing it. Alternatively, the tag can be detected by the nanopore without being released. Another method involves attaching the tag to a nucleotide via a sufficiently long adapter to present the tag to the nanopore for detection.

[0365] Nanopores can detect nucleotide incorporation events in real time. Enzymes (such as DNA polymerases) attached to or near nanopores can facilitate the passage of polynucleotides through or adjacent to the nanopore. Nucleotide incorporation events, or the incorporation of multiple nucleotides, can release or present one or more tags that can be detected by nanopores. Detection can be performed when the tag passes through or is adjacent to the nanopore, when the tag resides in the nanopore, and / or when the tag is presented to the nanopore. In some cases, enzymes attached to or adjacent to the nanopore can aid in the detection of tags during the incorporation of one or more nucleotides.

[0366] Tags can be atoms, molecules, collections of atoms, or collections of molecules. Tags can provide optical, electrochemical, magnetic, or electrostatic (such as inductive or capacitive) signatures that can be detected using nanopores.

[0367] Nanopores can be formed or otherwise embedded in a film adjacent to a sensing electrode arrangement of a sensing circuit (such as an integrated circuit). The integrated circuit can be an application-specific integrated circuit (ASIC). The integrated circuit can be a field-effect transistor or a complementary metal-oxide-semiconductor (CMOS). The sensing circuit can be located within a chip or other device with nanopores, or outside the chip or device, such as in an off-chip configuration.

[0368] When a nucleic acid or tag passes through or is near a nanopore, a sensing circuit detects an electrical signal associated with the nucleic acid or tag. The nucleic acid can be a subunit of a larger chain. The tag can be a byproduct of a nucleotide incorporation event or other interactions between the tagged nucleic acid and material in or near the nanopore (such as an enzyme that cleaves the tag from the nucleic acid). The tag can remain attached to the nucleotide. The detected signal can be collected and stored in a storage location and then used to construct the nucleic acid sequence. The collected signal can be processed to interpret any anomalies, such as errors, in the detected signal.

[0369] Nanopores can be used for indirect sequencing of polynucleotides, in some cases via electrical detection. Indirect sequencing can be any method in which nucleotides incorporated into the growing chain do not pass through the nanopore. Polynucleotides can pass through at any suitable distance from and / or near the nanopore, in some cases allowing for the detection of tags released from nucleotide incorporation events within the nanopore.

[0370] Byproducts of nucleotide incorporation events can be detected by nanopores. A nucleotide incorporation event refers to the incorporation of a nucleotide into a growing polynucleotide chain. Byproducts may be associated with the incorporation of a given type of nucleotide. Nucleotide incorporation events can be catalyzed by enzymes such as DNA polymerases, which use base pair interactions with the template molecule to select available nucleotides for incorporation at each position.

[0371] Nucleic acid samples can be sequenced using labeled nucleotides or nucleotide analogs. In some instances, methods for sequencing nucleic acid molecules include: (a) incorporating (e.g., polymerizing) labeled nucleotides, wherein a tag associated with an individual nucleotide is released upon incorporation, and (b) detecting the released tag using a nanopore. In some cases, the method also includes guiding a tag attached to or released from an individual nucleotide through a nanopore. The released or attached tag can be guided using any suitable technique, in some cases by means of an enzyme (or molecular motor) and / or a voltage difference across the nanopore. Alternatively, the released or attached tag can be guided through a nanopore without the use of an enzyme. For example, as described herein, the tag can be guided by a voltage difference across the nanopore.

[0372] Tags can be detected using a nanopore device having at least one nanopore in a membrane. The tag can associate with the individually labeled nucleotide during incorporation. The nanopore device can detect the tag associated with the individually labeled nucleotide during incorporation. The labeled nucleotide, whether incorporated into the growing nucleic acid chain or not, can be detected, identified, or distinguished by the nanopore device within a given time period, in some cases by means of the electrodes and / or nanopores of the nanopore device. The detection time of the tag by the nanopore device can be shorter (in some cases significantly shorter) than the time the tag and / or the nucleotide coupled to the tag is held by an enzyme (such as an enzyme that promotes nucleotide incorporation into the nucleic acid chain (e.g., polymerase)). The electrode can detect the tag multiple times during the time the incorporated labeled nucleotide associates with the enzyme. For example, during the time period in which the incorporated labeled nucleotide associates with the enzyme, the tag can be detected by the electrode at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 1000, 10,000, 100,000, or 1,000,000 times.

[0373] Sequencing can be performed using preloaded tags. Preloaded tags may include guiding at least a portion of the tag through at least a portion of a nanopore, with the tag attached to a nucleotide that may already be incorporated into a nucleic acid chain (e.g., a growing nucleic acid chain), being incorporated into a nucleic acid chain, or not yet incorporated but potentially to be incorporated into a nucleic acid chain. Preloaded tags may be guided through at least a portion of the tag through at least a portion of the nanopore before or while the nucleotide is being incorporated into the nucleic acid chain. Preloaded tags may also be guided through at least a portion of the tag through at least a portion of the nanopore after the nucleotide has already been incorporated into the nucleic acid chain.

[0374] Tags associated with individual nucleotides can be detected by nanopores without being released from the nucleotides during incorporation. Tags can be detected without being released from the incorporated nucleotides during the synthesis of a nucleic acid strand complementary to the target strand. The tag can be attached to the nucleotide via a linker, such that the tag is presented to the nanopore (e.g., the tag hangs in at least a portion of the nanopore or otherwise extends through at least a portion of the nanopore). The linker can be long enough to allow the tag to extend into or through at least a portion of the nanopore. In some cases, the tag is presented to (i.e., moved into) the nanopore by a voltage difference. Other methods of presenting the tag into the pore can also be suitable (e.g., using enzymes, magnets, electric fields, voltage differences). In some cases, no active force is applied to the tag (i.e., the tag diffuses into the nanopore).

[0375] A chip for sequencing nucleic acid samples can contain multiple individually addressable nanopores. Each individually addressable nanopore may contain at least one nanopore formed in a membrane disposed adjacent to an integrated circuit. Each individually addressable nanopore is capable of detecting a tag associated with an individual nucleotide. Nucleotides can be incorporated (e.g., polymerized), and the tag may not be released from the nucleotide during incorporation.

[0376] The tag can be presented to the nanopore and released from the nucleotide after a nucleotide incorporation event. The released tag can pass through the nanopore. In some cases, the tag does not pass through the nanopore. Tags released during a nucleotide incorporation event are distinguished from tags that can pass through the nanopore but are not fully released within the nanopore's residence time after the nucleotide incorporation event. In some cases, tags that remain in the nanopore for at least 100 milliseconds (ms) are released during a nucleotide incorporation event, while tags that remain in the nanopore for less than 100 ms are not released during the nucleotide incorporation event. The tag can be captured and / or guided through the nanopore by a second enzyme or protein (e.g., a nucleic acid-binding protein). The second enzyme can cleave the tag during (e.g., during or after) nucleotide incorporation. The linker between the tag and the nucleotide can be cleaved.

[0377] Based on the tag's residence time in the nanopore or based on the signal detected by means of unincorporated nucleotides in the nanopore, tags associated with incorporated nucleotides and tags associated with unincorporated, growing complementary nucleotides can be distinguished. Signals generated by unincorporated nucleotides (e.g., voltage difference, current) can be detected within timeframes of 1 nanosecond (ns) to 100 milliseconds or 1 ns to 50 ms, while signals generated by incorporated nucleotides can have lifetimes of 50 ms to 500 ms or 100 ms to 200 ms. Signals generated by unincorporated nucleotides can be detected within timeframes of 1 ns to 10 ms or 1 ns to 1 ms. The average time for which nanopores can detect unincorporated tags is longer than the average time for which nanopores can detect incorporated tags.

[0378] Compared to unincorporated nucleotides, incorporated nucleic acids can be detected by nanopores for a shorter period of time. Alternatively, incorporated nucleic acids can be detected by nanopores for a longer period of time compared to unincorporated nucleotides. As described herein, the differences and / or proportions between these times can be used to determine whether nanopore-detectable nucleotides have been incorporated.

[0379] Detection time can be based on the free flow of nucleotides through the nanopore; unincorporated nucleotides may remain in or near the nanopore for a period of 1 nanosecond (ns) to 100 ms or 1 ns to 50 ms, while incorporated nucleotides may remain in or near the nanopore for a period of 50 ms to 500 ms or 100 ms to 200 ms. The time periods can vary depending on the processing conditions; however, the residence time of incorporated nucleotides can be longer than that of unincorporated nucleotides.

[0380] Tags or tagged substances may include detectable atoms or molecules, or multiple detectable atoms or molecules. Tags may include one or more adenine, guanine, cytosine, thymine, uracil, or derivatives thereof, attached to any position, including a phosphate group, sugar, or nitrogenous base of a nucleic acid molecule. Tags may include one or more adenine, guanine, cytosine, thymine, uracil, or derivatives thereof covalently linked to a phosphate group of a nucleic acid base.

[0381] The tag can have a length of at least 0.1 nanometers (nm), 1nm, 2nm, 3nm, 4nm, 5nm, 6nm, 7nm, 8nm, 9nm, 10nm, 20nm, 30nm, 40nm, 50nm, 60nm, 70nm, 80nm, 90nm, 100nm, 200nm, 300nm, 400nm, 500nm, or 1000nm.

[0382] The tag may include the tail of a repeating subunit, such as multiple adenine, guanine, cytosine, thymine, uracil, or derivatives thereof. For example, the tag may include the tail of at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 1000, 10,000, or 100,000 subunits of adenine, guanine, cytosine, thymine, uracil, or derivatives thereof. These subunits may be linked to each other and are terminally linked to phosphate groups of nucleic acids. Other examples of the tag portion include any polymeric material, such as polyethylene glycol (PEG), polysulfonates, amino acids, or any polymer that is wholly or partially positively charged, negatively charged, or uncharged.

[0383] polymerase

[0384] DNA polymerases bind to the 3' end of the nicked strand of a polynucleotide at the nick site. DNA sequencing can be accomplished by amplifying and transcribing the polynucleotide near the nanopore and the labeled nucleotide using an enzyme such as a DNA polymerase. Sequencing methods may involve incorporating or polymerizing labeled nucleotides using polymerases such as DNA polymerases or transcriptases. Polymerases can be mutated to accept labeled nucleotides. Polymerases can also be mutated to increase the time it takes for the nanopore to detect the tag.

[0385] For example, sequencing enzymes can be any suitable enzyme that produces polynucleotide chains through the phosphate bonds of nucleotides. For example, DNA polymerase can be 9°Nm TM Polymerase or its variants, Escherichia coli DNA polymerase I, bacteriophage T4 DNA polymerase, sequencer, Taq DNA polymerase, 9°Nm TM Polymerase (exo-) A485L / Y409V, Φ29 DNA polymerase, Bst DNA polymerase, or a variant, mutant, or homolog of any of the foregoing. Homologs may have any suitable percentage of homology, for example, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, or at least 95% sequence identity.

[0386] In some instances, for nanopore sequencing, polymerases can be attached to or located near nanopores. Suitable methods for attaching polymerases to nanopores include cross-linking the enzyme to or near the nanopore, such as by forming intramolecular disulfide bonds. Nanopores and enzymes can also be fusions, such as fusions encoded by a single polypeptide chain. Methods for generating fusion proteins can include fusing the enzyme's coding sequence within a frame and adjacent to the coding sequence of the nanopore, and expressing the fusion sequence from a single promoter. Staples or protein fingers can be used to attach or couple polymerases to nanopores. Polymerases can be attached to nanopores via intermediate molecules, such as biotin conjugated to both the enzyme and the nanopore, where a streptavidin tetramer is linked to two biotin molecules. The intermediate molecule can be referred to as a linker.

[0387] Sequencing enzymes can also be attached to nanopores using antibodies. Proteins that form covalent bonds with each other can be used to attach polymerases to nanopores. Phosphatases or enzymes that cleave tags from nucleotides can also be attached to nanopores.

[0388] Polymerases can be mutated relative to non-mutated polymerases to promote and / or improve the efficiency with which the mutated polymerase incorporates labeled nucleotides into growing polynucleotides. Polymerases can be mutated to allow nucleotide analogs (such as labeled nucleotides) to better enter the active site region of the polymerase and / or mutated to match nucleotide analogs in the active region.

[0389] Other mutations, such as amino acid substitutions, insertions, deletions, and / or exogenous polymerase characteristics, can lead to enhanced metal ion coordination, reduced exonuclease activity, decreased reaction rate in one or more steps of the polymerase kinetic cycle, reduced branching ratio, altered cofactor selectivity, increased yield, increased thermal stability, increased accuracy, increased speed, increased read length, and increased salt tolerance relative to non-mutant polymerases.

[0390] Suitable polymerases can possess kinetic rate characteristics suitable for detecting tags through nanopores. Rate characteristics typically refer to the overall rate of nucleotide incorporation and / or the rate of any step in nucleotide incorporation, such as nucleotide addition, enzyme isomerization (e.g., becoming closed or isomerizing from a closed state), cofactor binding or release, product release, incorporation of polynucleotides into growing polynucleotides, or the rate of translocation.

[0391] Polymerases can be adapted to allow for the detection of sequencing events. The rate characteristics of polymerases allow for the average duration of tag loading into nanopores (and / or detection by nanopores) of 0.1 ms, 1 ms, 5 ms, 10 ms, 20 ms, 30 ms, 40 ms, 50 ms, 60 ms, 80 ms, 100 ms, 120 ms, 140 ms, 160 ms, 180 ms, 200 ms, 220 ms, 240 ms, 260 ms, 280 ms, 300 ms, 400 ms, 500 ms, 600 ms, 800 ms, or 1000 ms. For example, the rate characteristics of polymerases can allow tags to be loaded into and / or detected by nanopores for at least 5 ms, at least 10 ms, at least 20 ms, at least 30 ms, at least 40 ms, at least 50 ms, at least 60 ms, at least 80 ms, at least 100 ms, at least 120 ms, at least 140 ms, at least 160 ms, at least 180 ms, at least 200 ms, at least 220 ms, at least 240 ms, at least 260 ms, at least 280 ms, at least 300 ms, at least 400 ms, at least 500 ms, at least 600 ms, at least 800 ms, or at least 1000 ms. Nanopores can detect tags on average between 80 ms and 260 ms, between 100 ms and 200 ms, or between 100 ms and 150 ms.

[0392] Nanopore / polymerase complexes can be configured to allow the detection of one or more events associated with the amplification and transcription of cyclic polynucleotides. These events can be kinetically observable and / or non-kinetically observable, such as the migration of nucleotides through the nanopore without contact with the polymerase.

[0393] In some cases, polymerase reactions exhibit two kinetic steps initiating at an intermediate where the nucleotide or polyphosphate product binds to the polymerase, and two kinetic steps initiating at an intermediate where neither the nucleotide nor the polyphosphate product binds to the polymerase. These two kinetic steps can include enzyme isomerization, nucleotide incorporation, and product release. In some cases, the two kinetic steps are template translocation and nucleotide binding.

[0394] A suitable polymerase can exhibit strong or enhanced chain substitution.

[0395] connector

[0396] Polymerases can contain linkers. Linkers can be used to couple polymerases to nanopores. Polymerases containing linkers can be bound to protein nanopores to form polymerase / nanopore complexes that can bind to cleavage sites of cyclic polynucleotides.

[0397] Polymerases containing linkers can react with nicked cyclic polynucleotides to form polymerase / cyclic polynucleotide complexes. These polymerase / cyclic polynucleotide complexes can then bind to nanopores such as protein nanopores (e.g., α-hemolysin) or solid nanopores via the linker.

[0398] Polymerase / cyclic polynucleotide / nanopore complexes can be used for multinucleotide sequencing. The linkage properties between DNA polymerase and nanopores can increase the concentration of effectively labeled nucleotides, thereby reducing the entropy barrier. Examples of optimizable linker aspects include linker length, which can increase the concentration of effectively labeled nucleotides, affecting capture kinetics and / or altering the entropy barrier; linker flexibility, which can affect the kinetics of linker conformational changes; and the number and location of links between the polymerase and nanopores, which can reduce the number of available conformational states, thereby increasing the likelihood of appropriate pore polymerase orientation, increasing the concentration of effectively labeled nucleotides, and reducing the entropy barrier.

[0399] The connector can be a polymer, such as a peptide, polynucleotide, or polyethylene glycol. The connector can be of any suitable length. For example, the connector length can be 5 nm, 10 nm, 15 nm, 20 nm, 40 nm, 50 nm, or 100 nm. The connector length can be at least 5 nm, at least 10 nm, at least 15 nm, at least 20 nm, at least 40 nm, at least 50 nm, or at least 100 nm. The connector length can be less than 5 nm, less than 10 nm, less than 15 nm, less than 20 nm, less than 40 nm, less than 50 nm, or less than 100 nm. The connector can be rigid, flexible, or a combination thereof.

[0400] In some implementations, no adapter is used, and the polymerase is directly attached to the nanopore.

[0401] Polymerases can attach to nanopores via two or more linkers. The number and location of these links between the polymerase and the nanopore can vary. Examples include links from the C-terminus of the αHL polymerase to its N-terminus, from the N-terminus of the αHL polymerase to its C-terminus, and links between amino acids not at the ends.

[0402] Connectors can be used to orient polymerase relative to nanopores, allowing tags to be detected via nanopores.

[0403] For example, in a method for sequencing a polynucleotide sample using a nanopore in a membrane adjacent to a sensing electrode, a labeled nucleotide is provided into a reaction chamber containing a nanopore, wherein the labeled individual nucleotides contain a tag coupled to the nucleotide, which can be detected by means of the nanopore. The method may include a polymerization reaction using a polymerase attached to the nanopore via a linker, thereby incorporating the labeled individual nucleotides of the labeled nucleotides into a growth chain complementary to a single-stranded polynucleotide from the polynucleotide sample. The method may include detecting the tag associated with the labeled individual nucleotide using the nanopore during the incorporation of the labeled individual nucleotide, wherein the tag is detected by means of the nanopore when the nucleotide associates with the polymerase.

[0404] Amplification and sequencing

[0405] Amplification and transcription can include rolling circle amplification (RCA).

[0406] In RCA, the reaction mixture may contain one or more primers, a polymerase, and dNTPs, producing a tandem polymerase. The polymerase in the RCA reaction may include a polymerase with strand displacement activity. Examples of polymerases with strand displacement activity include DNA polymerase I large (Klenow) fragment lacking exonuclease activity, Phi29 DNA polymerase, and Taq DNA polymerase.

[0407] During the sequencing process while amplifying, in order to prevent DNA polymerase from binding from the original template to the replaced single-stranded DNA, single-stranded cutting enzymes (e.g., truncated exonuclease VIII, T5 exonuclease, T7 exonuclease) can be used to cut the replaced single-stranded DNA into dNMP, dinucleotides, etc.

[0408] In some cases, amplified polynucleotides can be visualized as nanospheres under a fluorescence microscope or through particle size analysis.

[0409] Identification of sequence variants

[0410] The method provided in this disclosure can be used to identify sequence variants in polynucleotide samples. If the sequence difference occurs in at least two different polynucleotides, for example, two different cyclic polynucleotides, the sequence difference between the sequencing read and the reference sequence is called a true sequence variant, which can be distinguished by its different linkers. Because the location and type of sequence variants caused by amplification or sequencing errors are unlikely to be precisely replicated on two different polynucleotides containing the same target sequence, including this validation parameter can reduce the background of erroneous sequence variants while improving the sensitivity and accuracy of detecting actual sequence variations in a sample. The frequency of sequence variants can be less than 5%, 4%, 3%, 2%, 1.5%, 1%, 0.75%, 0.5%, 0.25%, 0.1%, 0.075%, 0.05%, 0.04%, 0.03%, 0.02%, 0.01%, 0.005%, 0.001%, or a lower frequency sufficiently higher than the background value to allow for accurate identification. The occurrence frequency of sequence variants may be less than 0.1%. When the frequency of a sequence variant is statistically significantly higher than the background error rate, for example, with a p-value less than 0.05, 0.01, 0.001, or 0.0001, then the frequency of the sequence variant can be considered sufficiently higher than the background. The frequency of a sequence variant can be considered sufficiently higher than the background when it is at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 25, 50, 100, or more times higher than the background error rate. The background error rate for accurately determining a sequence at a given location can be less than 1%, 0.5%, 0.1%, 0.05%, 0.01%, 0.005%, 0.001%, or 0.0005%.

[0411] Identifying sequence variants can include optimally aligning one or more sequencing reads to a reference sequence to identify differences between them, as well as identifying junctions. Alignment can include placing a sequence along another sequence, iteratively introducing gaps along each sequence, scoring the degree of match between the two sequences, and repeating at different positions along the reference sequence. The highest-scoring match is considered the alignment and represents an inference about the degree of correlation between the sequences.

[0412] The reference sequence compared to the sequencing read is a reference genome, such as the genome of a member of the same species as the subject. The reference genome can be complete or incomplete. The reference genome can only consist of regions containing target polynucleotides, such as regions derived from a reference genome or shared sequences generated from sequencing reads in the analysis. The reference sequence can contain or be composed of polynucleotide sequences from one or more organisms, such as sequences from one or more bacteria, archaea, viruses, protozoa, fungi, or other organisms. The reference sequence can consist of only a portion of the reference genome, such as regions corresponding to one or more target sequences in the analysis. For example, to detect pathogens, the reference genome can be the entire genome of the pathogen, or a portion thereof used for identification, such as a specific strain or serotype. Sequencing reads can be compared to multiple different reference sequences for screening multiple different organisms or strains.

[0413] Computer System

[0414] This disclosure provides computer systems programmed to implement one or more methods of this disclosure. The computer systems of this disclosure can be used to regulate various operations of nanopore sequencing, such as detecting one or more signals indicating impedance or impedance changes in the nanopore when a sample (e.g., at least a portion of a tagged nucleotide label) is within a nanopore (e.g., a protein nanopore or a solid nanopore).

[0415] Figure 6 A computer system 601 is illustrated, which is programmed or otherwise configured to communicate and regulate various aspects of sequencing according to this disclosure. For example, the computer system 601 may communicate with one or more circuits coupled to or including nanopores (or membranes containing nanopores), and with one or more means (e.g., machines) for preparing, processing, or maintaining one or more reaction mixtures for sequencing. The computer system 601 may also communicate with one or more controllers or processors according to this disclosure. The computer system 601 may be a user's electronic device or a computer system located remotely relative to the electronic device. The electronic device may be a mobile electronic device.

[0416] Computer system 601 includes a central processing unit (CPU, also referred to herein as a “processor” and “computer processor”) 605, which may be a single-core or multi-core processor, or multiple processors for parallel processing. Computer system 601 also includes memory or memory locations 610 (e.g., random access memory, read-only memory, flash memory), electronic storage units 615 (e.g., hard disks), a communication interface 620 (e.g., a network adapter) for communicating with one or more other systems, and peripheral devices 625, such as caches, other memories, data storage, and / or electronic display adapters. Memory 610, storage units 615, interface 620, and peripheral devices 625 communicate with CPU 605 via a communication bus (solid line) such as a motherboard. Storage unit 615 may be a data storage unit (or data repository) for storing data. Computer system 601 may be operatively coupled to computer network (“network”) 630 by means of communication interface 620. Network 630 may be the Internet, the Internet of Things, and / or an extranet, or an intranet and / or an extranet communicating with the Internet. In some cases, network 630 is a telecommunications and / or data network. Network 630 may include one or more computer servers that can enable distributed computing, such as cloud computing. In some cases, network 630 may enable a peer-to-peer network with the assistance of computer system 601, which allows devices coupled to computer system 601 to act as clients or servers.

[0417] CPU 605 can execute a series of machine-readable instructions, which can be embodied in a program or software. The instructions can be stored in a memory location, such as memory 610. The instructions can be directed to CPU 605, which can then be programmed or otherwise configured to implement the methods of this disclosure. Examples of operations performed by CPU 605 can include fetching, decoding, executing, and writing back.

[0418] CPU 605 may be part of a circuit such as an integrated circuit. The circuit may include one or more other components of system 601. In some cases, the circuit is an application-specific integrated circuit (ASIC).

[0419] Storage unit 615 may store files, such as drivers, libraries, and saved programs. Storage unit 615 may store user data, such as user preferences and user programs. In some cases, computer system 601 may include one or more additional data storage units outside of computer system 601, such as on a remote server that communicates with computer system 601 via an intranet or the Internet.

[0420] Computer system 601 can communicate with one or more remote computer systems via network 630. For example, computer system 601 can communicate with a user's remote computer system. Examples of remote computer systems include personal computers (e.g., portable PCs), tablets, or tablet computers (e.g., tablet PCs). iPad Galaxy Tab), telephone, smartphone (e.g., iPhone, Android-enabled devices (or personal digital assistant). Users can access computer system 601 via network 630.

[0421] The methods described herein can be implemented by means of machine-executable code (e.g., a computer processor) stored in an electronic storage location (e.g., memory 610 or electronic storage unit 615) of computer system 601. The machine-executable or machine-readable code may be provided in software form. During use, the code can be executed by processor 605. In some cases, the code can be retrieved from storage unit 615 and stored in memory 610 for access by processor 605 at any time. In some cases, electronic storage unit 615 may not be included, and the machine-executable instructions are stored in memory 610.

[0422] The code can be pre-compiled and configured for use with machines that have processors suitable for executing the code, or it can be compiled at runtime. The code can be provided in a programming language, and the programming language can be selected to enable the code to be executed in a pr...

Claims

1. A method for processing or analyzing double-stranded nucleic acid molecules, comprising: (a) Provide (i) the double-stranded nucleic acid molecule and (ii) a double-stranded linker having a cleavage site in its sense or antisense strand; (b) coupling the double-stranded linker to the double-stranded nucleic acid molecule; and (c) Circularize the double-stranded nucleic acid molecule coupled to the double-stranded linker to produce a circular double-stranded nucleic acid molecule.

2. The method according to claim 1, wherein the double-stranded nucleic acid molecule and the double-stranded adapter are heterologous to each other.

3. The method of claim 1, wherein the double-stranded nucleic acid molecule and the double-stranded adaptor are provided in the form of a cell-free composition.

4. The method of claim 1, wherein (b) or (c) is performed under cell-free conditions.

5. The method of claim 1, wherein the coupling comprises (i) coupling the sense strand of the double-stranded adapter to the sense strand of the double-stranded nucleic acid molecule, or (ii) coupling the antisense strand of the double-stranded adapter to the antisense strand of the double-stranded nucleic acid molecule.

6. A reaction mixture for processing or analyzing double-stranded nucleic acid molecules, comprising: A composition comprising (i) the double-stranded nucleic acid molecule and (ii) a double-stranded linker having a cleavage site within its sense or antisense strand; and At least one enzyme, said enzyme (i) couples the double-stranded adapter to the double-stranded nucleic acid molecule, and (ii) cyclizes the double-stranded nucleic acid molecule coupled with the double-stranded adapter to produce a cyclized double-stranded nucleic acid molecule.

7. A library of circularized double-stranded nucleic acid molecules comprising (i) a double-stranded nucleic acid domain coupled to (ii) a double-stranded linker domain, the double-stranded linker domain containing a nick site within the sense or antisense strand of the double-stranded linker domain, wherein each circularized double-stranded nucleic acid molecule in at least 5% of the library contains a recognition sequence.

8. A method for processing or analyzing circular nucleic acid molecules, comprising: (a) Providing a cell-free composition comprising the circular nucleic acid molecule, the circular nucleic acid molecule comprising (i) a target region and (ii) a nick site at a known distance from the target region; and (b) A nick is made at the nick site of the circular nucleic acid molecule.

9. A reaction mixture for processing or analyzing circular nucleic acid molecules, comprising: A cell-free composition comprising the circular nucleic acid molecule, the circular nucleic acid molecule comprising (i) a target site and (ii) a nick site at a known distance from the target site; and At least one enzyme that produces a nick at the nick site of the circular nucleic acid molecule.

10. A cell-free library of circular nucleic acid molecules, wherein each individual of at least 5% of the library comprises (i) a target site and (ii) a cut at a known distance from the target site.