Methods, systems, and compositions for nucleic acid sequencing
Patent Information
- Application Number
- HK42026127106
- Authority / Receiving Office
- HK · HK
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2018-05-22
- Filing Date
- 2026-08-05
- Publication Date
- 2026-09-18
- Estimated Expiration
- 2039-05-20
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
(19) State Intellectual Property Office (12) Invention Patent Application (10) Application Publication Number (43) Application Publication Date (21) Application Number 202511210432.X (22) Application Date 2019.05.21 (30) Priority Data 62 / 674,706 2018.05.22 US (62) Divisional Application Data 201980049362.0 2019.05.21 (71) Applicant Anxuyuan Co., Ltd. Address California, USA (72) Inventors Deng Suhua Vladimir Ivanovich Bashkilov Tian Hui Igor Konstantin Ivanov (74) Patent Agency Beijing Anxin Fangda Intellectual Property Agency Co., Ltd. 11262 Patent Attorneys Wei Changjin Wujing (51) Int.Cl. C12Q 1 / 6806 (2018.01) C12Q 1 / 6869 (2018.01) C40B 40 / 06 (2006.01) (54) Invention Title: Methods, Systems, and Compositions for Nucleic Acid Sequencing (57) Abstract: This disclosure describes methods, systems, and compositions for nucleic acid sequencing. This disclosure provides methods and systems for processing or analyzing nucleic acid molecules. Methods for processing or analyzing double-stranded nucleic acid molecules may include providing a double-stranded nucleic acid molecule and a double-stranded adaptor. The double-stranded adaptor may contain a nick site within its sense or antisense strand. The double-stranded adaptor may then be coupled to the double-stranded nucleic acid molecule, and the double-stranded nucleic acid molecule coupled to the double-stranded adaptor may be circularized to produce a circularized double-stranded nucleic acid molecule. Claims 1 page, Description 69 pages, Drawings 18 pages, CN 121472373 A 2026.02.06 CN 1 21 47 23 73 A 1. A method for processing or analyzing a double-stranded nucleic acid molecule, comprising: (a) providing (i) the double-stranded nucleic acid molecule and (ii) a double-stranded adaptor having a nick site within its sense or antisense strand; (b) coupling the double-stranded adaptor to the double-stranded nucleic acid molecule; and (c) cyclizing the double-stranded nucleic acid molecule coupled with the double-stranded adaptor to produce a cyclized double-stranded nucleic acid molecule. 2. The method of claim 1, wherein the double-stranded nucleic acid molecule and the double-stranded adaptor are heterologous to each other. 3. The method of claim 1, wherein the double-stranded nucleic acid molecule and the double-stranded adaptor are provided in the form of a cell-free composition. 4. The method of claim 1, wherein (b) or (c) is performed under cell-free conditions. 5. The method of claim 1, wherein the coupling comprises (i) coupling the sense strand of the double-stranded adapter to the sense strand of the double-stranded nucleic acid molecule, or (ii) coupling the antisense strand of the double-stranded adapter to the antisense strand of the double-stranded nucleic acid molecule.6. A reaction mixture for processing or analyzing double-stranded nucleic acid molecules, comprising: a composition comprising (i) the double-stranded nucleic acid molecule and (ii) a double-stranded adaptor having a nick site within its sense or antisense strand; and at least one enzyme, said enzyme (i) coupling the double-stranded adaptor to the double-stranded nucleic acid molecule, and (ii) cyclizing the double-stranded nucleic acid molecule coupled with the double-stranded adaptor to produce a cyclized double-stranded nucleic acid molecule. 7. A library of cyclized double-stranded nucleic acid molecules comprising (i) a double-stranded nucleic acid domain coupled to (ii) a double-stranded adaptor domain, said double-stranded adaptor domain containing a nick site within the sense or antisense strand of said double-stranded adaptor domain, wherein each cyclized double-stranded nucleic acid molecule in at least 5% of said library contains a recognition sequence. 8. A method for processing or analyzing a circular nucleic acid molecule, comprising: (a) providing a cell-free composition comprising the circular nucleic acid molecule, the circular nucleic acid molecule comprising (i) a target region and (ii) a nick site at a known distance from the target region; and (b) generating a nick at the nick site of the circular nucleic acid molecule. 9. A reaction mixture for processing or analyzing a circular nucleic acid molecule, comprising: a cell-free composition comprising the circular nucleic acid molecule, the circular nucleic acid molecule comprising (i) a target site and (ii) a nick site at a known distance from the target site; and at least one enzyme for generating a nick at the nick site of the circular nucleic acid molecule. 10. A cell-free library of circular nucleic acid molecules, wherein the circular nucleic acid molecule of each individual in at least 5% of the library comprises (i) a target site and (ii) a nick at a known distance from the target site. Claims 1 / 1 Page 2 CN 121472373 A Method, System, and Composition for Nucleic Acid Sequencing
[0001] This application is a divisional application of Chinese Patent Application No. 201980049362.0, filed on May 21, 2019, entitled "Method, System, and Composition for Nucleic Acid Sequencing" (the corresponding PCT application was filed on May 21, 2019, and is numbered PCT / US2019 / 033376).
[0002] Cross-Reference
[0003] This application claims the benefit of U.S. Provisional Application No. 62 / 674,706, filed on May 22, 2018, which is incorporated herein by reference in its entirety. Background of the Invention
[0005] Nucleic acid sequencing can be used to provide sequence information of nucleic acid samples. Such sequence information may help in the diagnosis or treatment of a subject's (e.g., an individual, a patient, etc.) condition (e.g., a disease). For example, a subject's nucleic acid sequence information can be used to identify, diagnose, or develop treatments for one or more genetic diseases. In another instance, the nucleic acid sequence information of one or more pathogens...Nucleic acid sequence information can guide the treatment of one or more infectious diseases.
[0006] In some cases, methods for nucleic acid sequencing may include generating a nick of a circular double-stranded nucleic acid chain and binding a polymerase to the nick. The resulting complex, comprising the polymerase and the circular double-stranded nucleic acid complex, may associate (e.g., conjugate or be adjacent to) a sequencing portion (e.g., nanopore) and may generate a growth chain complementary to at least a portion of the double-stranded nucleic acid (e.g., by rolling circle amplification (RCA)) for sequencing by the sequencing portion. Such methods can be used for whole-genome sequencing or for the detection of one or more sequence variants (e.g., mutations) in a nucleic acid library.
[0007] The detection of one or more rare sequence variants (e.g., mutations) may be valuable for healthcare. The detection of rare sequence variants may be important for the early detection of one or more pathological mutations. Detection of one or more cancer-related mutations (e.g., point mutations) in clinical samples may improve the identification of one or more minimal residual diseases in chemotherapy for relapsed patients or in tumor cell detection. Furthermore, such mutation detection may be important for assessing exposure to environmental mutagens, monitoring endogenous DNA repair, or studying the accumulation of one or more somatic mutations in aging individuals. Alternatively or additionally, detection of rare sequence variants can enhance prenatal diagnosis and characterize fetal cells present in maternal blood. Summary of the Invention
[0008] In one aspect, this disclosure provides a method for processing or analyzing double-stranded nucleic acid molecules, comprising: (a) providing (i) the double-stranded nucleic acid molecule and (ii) a double-stranded adaptor having a nick site within its sense or antisense strand; (b) coupling the double-stranded adaptor to the double-stranded nucleic acid molecule; and (c) cyclizing the double-stranded nucleic acid molecule coupled with the double-stranded adaptor to produce a cyclized double-stranded nucleic acid molecule.
[0009] In some embodiments, the double-stranded nucleic acid molecule and the double-stranded adaptor are heterologous to each other. In some embodiments, the double-stranded nucleic acid molecule and the double-stranded adaptor are provided in the form of a cell-free composition. In some embodiments, (b) or (c) is performed under cell-free conditions. In some embodiments, (b) and (c) are performed under cell-free conditions.
[0010] In some embodiments, the coupling includes (i) coupling the sense strand of the double-stranded adaptor to the sense strand of the double-stranded nucleic acid molecule, or (ii) coupling the antisense strand of the double-stranded adaptor to the antisense strand of the double-stranded nucleic acid molecule. In some embodiments, the coupling includes (i) coupling the sense strand of the double-stranded adaptor to the sense strand of the double-stranded nucleic acid molecule, and (ii) coupling the antisense strand of the double-stranded adaptor to the antisense strand of the double-stranded nucleic acid molecule.
[0011] In some embodiments, the nick site is part of the sense strand of the circularized double-stranded nucleic acid molecule. In some embodiments, the nick site is part of the antisense strand of the circularized double-stranded nucleic acid molecule.
[0012] In some embodiments, the method further includes sequencing the double-stranded nucleic acid molecule from the nick site of the double-stranded adaptor. In some embodiments, the sequencing includes (i) extending the double-stranded nucleic acid molecule from the nick site of the double-stranded adaptor to produce a growth strand that is sequence complementary to at least a portion of the strand of the double-stranded nucleic acid molecule, and (ii) obtaining sequence information of at least a portion of the growth strand. In some embodiments, obtaining the sequence information includes detecting the at least a portion of the growth strand. In some embodiments, the extension reaction includes contacting the double-stranded nucleic acid molecule with the nucleotide coupled to the tag under conditions sufficient to incorporate the nucleotide into the growth strand, and wherein obtaining the sequence information includes detecting the tag. In some embodiments, the method further includes releasing the tag from the nucleotide when incorporating the nucleotide into the growth strand. In some embodiments, the extension reaction is performed without the use of oligonucleotide primers. In some embodiments, the extension reaction includes rolling circle amplification.
[0013] In some embodiments, the sequencing includes (i) cleaving the double-stranded nucleic acid molecule from the cleavage site of the double-stranded connective to cleave at least a portion of the strand of the double-stranded nucleic acid molecule, and (ii) obtaining sequence information of the at least a portion of the strand. In some embodiments, obtaining the sequence information includes detecting the at least a portion of the strand. In some embodiments, the sequencing includes nanopore-based sequencing. In some embodiments, at least a portion of the double-stranded nucleic acid molecule has or is suspected of having one or more sequencing variants compared to at least one reference sequence, and wherein the sequencing is for identifying the presence of the at least a portion of the double-stranded nucleic acid molecule. In some embodiments, the one or more sequencing variants indicate mutations in a gene. In some embodiments, the at least one reference sequence includes a common sequence of at least a portion of the gene.
[0014] In some embodiments, the method further includes, prior to (b), amplifying the double-stranded nucleic acid molecule to generate multiple copies of the double-stranded nucleic acid molecule.
[0015] In some embodiments, the double-stranded nucleic acid molecule contains a recognition sequence, and the method further includes enriching the double-stranded nucleic acid molecule from a random nucleic acid molecule library at least in part based on the recognition sequence. In some embodiments, the enrichment includes generating a selected library of double-stranded nucleic acid molecules, wherein each double-stranded nucleic acid molecule in at least 5% of the selected library contains the recognition sequence. In some embodiments, at least 10%, 20%, 30%, 40%, ...Each double-stranded nucleic acid molecule contains the recognition sequence in 50%, 60%, 70%, 80%, 90%, or 95% of its components. In some embodiments, the probability of the recognition sequence appearing without any mismatch is at most once per 1x10⁴ base pairs. In some embodiments, the probability of the recognition sequence appearing without any mismatch is at most once per 5x10⁴, 7x10⁴, 1x10⁵, 1x10⁶, 1x10⁷, 1x10⁸, 1x10⁹, 1x10¹⁰, 1x10¹¹, or 1x10¹² base pairs. In some embodiments, the recognition sequence contains at least 5 bases. In some embodiments, the recognition sequence comprises at least 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 35, 40, 45, or 50 bases. In some embodiments, the enrichment comprises (i) binding a recognition portion complementary to the recognition sequence to the double-stranded nucleic acid molecule to form a recognition complex, and (ii) extracting the recognition complex. In some embodiments, the enrichment is performed before (a) or after (b). In some embodiments, the enrichment is performed before (a) and after (b).
[0016] In some embodiments, the double-stranded nucleic acid molecule is derived from or derived from a biological sample of the subject. In some embodiments, the biological sample comprises a cell-free biological sample of the subject. In some embodiments, the double-stranded nucleic acid molecule is derived from or derived from a cell-free nucleic acid molecule from the cell-free biological sample. In some embodiments, the cell-free nucleic acid molecule includes circulating tumor nucleic acid molecules or amniotic fluid nucleic acid molecules. In some embodiments, the biological sample includes a tissue sample of the subject. In some embodiments, the double-stranded nucleic acid molecule is derived from or derived from a genomic nucleic acid molecule from the tissue sample. In some embodiments, the tissue sample is derived from: infected tissue, diseased tissue, malignant tissue, calcified tissue, healthy tissue, and combinations thereof. In some embodiments, the tissue sample is from malignant tissue comprising a tumor, sarcoma, leukemia, or a derivative thereof.
[0017] In some embodiments, the double-stranded nucleic acid molecule includes DNA, complementary DNA, derivatives thereof, or combinations thereof. In some embodiments, the double-stranded nucleic acid molecule includes RNA.
[0018] In another aspect, this disclosure provides a reaction mixture for processing or analyzing double-stranded nucleic acid molecules, comprising: a composition comprising (i) the double-stranded nucleic acid molecule and (ii) having a cleavage within its sense or antisense strand.A double-stranded adaptor at an oral site; and at least one enzyme, said enzyme (i) coupling the double-stranded adaptor to the double-stranded nucleic acid molecule, and (ii) cyclizing the double-stranded nucleic acid molecule coupled with the double-stranded adaptor to produce a cyclized double-stranded nucleic acid molecule.
[0019] In some embodiments, the double-stranded nucleic acid molecule and the double-stranded adaptor are heterologous to each other. In some embodiments, the reaction mixture is a cell-free reaction mixture.
[0020] In some embodiments, said at least one enzyme (i) couples the sense strand of the double-stranded adaptor to the sense strand of the double-stranded nucleic acid molecule, or (ii) couples the antisense strand of the double-stranded adaptor to the antisense strand of the double-stranded nucleic acid molecule. In some embodiments, said at least one enzyme (i) couples the sense strand of the double-stranded adaptor to the sense strand of the double-stranded nucleic acid molecule, and (ii) couples the antisense strand of the double-stranded adaptor to the antisense strand of the double-stranded nucleic acid molecule. In some embodiments, said at least one enzyme links the double-stranded adaptor to the double-stranded nucleic acid molecule. In some embodiments, the at least one enzyme includes a ligase, a recombinase, a polymerase, a functional variant thereof, or a combination thereof.
[0021] In some embodiments, the cleavage site is part of the sense strand of the cyclic double-stranded nucleic acid molecule. In some embodiments, the cleavage site is part of the antisense strand of the cyclic double-stranded nucleic acid molecule.
[0022] In some embodiments, the reaction mixture further comprises at least a second enzyme that performs an extension reaction to produce a growth chain that is sequence complementary to at least a portion of the strand of the double-stranded nucleic acid molecule. In some embodiments, the at least second enzyme produces the growth chain before the double-stranded adaptor is coupled to the double-stranded nucleic acid molecule.
[0023] In some embodiments, after the double-stranded adaptor is coupled to the double-stranded nucleic acid molecule, the at least second enzyme performs the extension reaction from the cleavage site of the double-stranded adaptor to produce the growth chain. In some embodiments, the reaction mixture further comprises at least one nucleotide coupled to a tag, wherein the at least second enzyme incorporates the nucleotide into the growth chain. In some embodiments, when the nucleotide is incorporated into the growth chain, the at least second enzyme releases the tag from the nucleotide. In some embodiments, the at least second enzyme performs the extension reaction without using an oligonucleotide primer. In some embodiments, the extension reaction comprises rolling circle amplification. In some embodiments, the at least second enzyme comprises a polymerase.
[0024] In some embodiments, the reaction mixture further comprises at least a third enzyme that performs a cleavage reaction from the cleavage site of the double-stranded connective to cleave at least a portion of the strand of the double-stranded nucleic acid molecule.
[0025] In some embodiments, at least a portion of the double-stranded nucleic acid molecule has or is suspected of having one or more variants compared to at least one reference sequence. In some embodiments, the reaction mixture is used to prepare at least one composition, which is used for sequencing to identify the presence of the at least a portion of the double-stranded nucleic acid molecule. In some embodiments, the one or more sequencing variants indicate mutations in a gene. In some embodiments, the at least one reference sequence comprises a common sequence of at least a portion of the gene.
[0026] In some embodiments, the double-stranded nucleic acid molecule contains a recognition sequence. In some embodiments, the reaction mixture further comprises a recognition portion associated with the recognition sequence to enrich the double-stranded nucleic acid molecule from a random nucleic acid molecule library in the composition, at least partially based on the recognition sequence. In some embodiments, the recognition portion comprises at least one oligonucleotide complementary to at least the recognition sequence. In some embodiments, the composition comprises a selected library of double-stranded nucleic acid molecules, wherein each double-stranded nucleic acid molecule in at least 5% of the library contains the recognition sequence. In some embodiments, each double-stranded nucleic acid molecule in at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 95% of the selected library contains the recognition sequence. In some embodiments, the probability of the recognition sequence occurring without any mismatch is at most once per 1x10⁴ base pairs. In some embodiments, the probability of the recognition sequence occurring without any mismatch is at most once per 5x10⁴, 7x10⁴, 1x10⁵, 1x10⁶, 1x10⁷, 1x10⁸, 1x10⁹, 1x10¹⁰, 1x10¹¹, or 1x10¹² base pairs. In some embodiments, the recognition sequence contains at least 5 bases. In some embodiments, the identification sequence comprises at least 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 35, 40, 45, or 50 bases.
[0027] In some embodiments, the double-stranded nucleic acid molecule is derived from or derived from a biological sample of the subject. In some embodiments, the biological sample comprises a cell-free biological sample of the subject. In some embodiments, the double-stranded nucleic acid molecule is derived from or derived from a cell-free nucleic acid molecule derived from the cell-free biological sample. In some embodiments, the cell-free nucleic acid molecule comprises circulating tumor nucleic acid molecules or amniotic fluid nucleic acid molecules. In some embodiments, the biological sample comprises a tissue sample of the subject. In some embodiments, the double-stranded nucleic acid molecule is derived from or derived from...Genomic nucleic acid molecules derived from the tissue sample. In some embodiments, the tissue sample is derived from: infected tissue, diseased tissue, malignant tissue, calcified tissue, healthy tissue, and combinations thereof. In some embodiments, the tissue sample is derived from malignant tissue comprising tumors, sarcomas, leukemia, or derivatives thereof.
[0028] In some embodiments, the double-stranded nucleic acid molecule comprises DNA, complementary DNA, derivatives thereof, or combinations thereof. In some embodiments, the double-stranded nucleic acid molecule comprises RNA.
[0029] In various aspects, this disclosure provides a library of circularized double-stranded nucleic acid molecules comprising (i) a double-stranded nucleic acid domain coupled to (ii) a double-stranded adaptor domain, the double-stranded adaptor domain containing a cleavage site within the sense or antisense strand of the double-stranded adaptor domain, wherein each circularized double-stranded nucleic acid molecule in at least 5% of the library contains a recognition sequence.
[0030] In some embodiments, the cleavage site is present within the double-stranded adaptor domain before the double-stranded adaptor domain is coupled to the double-stranded nucleic acid domain. In some embodiments, the double-stranded nucleic acid domain and the double-stranded connective domain are heterologous to each other. In some embodiments, the library is in a cell-free composition.
[0031] In some embodiments, at least a portion of the circularized double-stranded nucleic acid domain has or is suspected of having one or more sequencing variants compared to at least one reference sequence. In some embodiments, the one or more sequencing variants indicate mutations in a gene. In some embodiments, the at least one reference sequence comprises a common sequence of at least a portion of the gene.
[0032] In some embodiments, each circularized double-stranded nucleic acid molecule in at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 95% of the selected library contains the recognition sequence. In some embodiments, as described on page 4 / 69 of CN 121472373 A, the probability of the recognition sequence appearing without any mismatch is at most once per 1 x 10⁴ base pairs. In some embodiments, the probability of the identification sequence occurring without any mismatch is at most once in every 5x10⁴, 7x10⁴, 1x10⁵, 1x10⁶, 1x10⁷, 1x10⁸, 1x10⁹, 1x10¹⁰, 1x10¹¹, or 1x10¹² base pairs. In some embodiments, the identification sequence contains at least 5 specific bases. In some embodiments, the identification sequence contains at least 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 35, 40, 45, or 50 bases.
[0033] In some embodiments, the nick site is part of the sense strand of the circularized double-stranded nucleic acid molecule. In some embodiments, the nick site is part of the antisense strand of the circularized double-stranded nucleic acid molecule.
[0034] In some embodiments, the circularized double-stranded nucleic acid molecule is derived from or derived from a biological sample of a subject. In some embodiments, the biological sample includes a cell-free biological sample of the subject. In some embodiments, the double-stranded nucleic acid molecule is derived from or derived from a cell-free nucleic acid molecule derived from the cell-free biological sample. In some embodiments, the cell-free nucleic acid molecule includes circulating tumor nucleic acid molecules or amniotic fluid nucleic acid molecules. In some embodiments, the biological sample includes a tissue sample of the subject. In some embodiments, the double-stranded nucleic acid molecule is derived from or derived from a genomic nucleic acid molecule derived from the tissue sample. In some embodiments, the tissue sample is derived from: infected tissue, diseased tissue, malignant tissue, calcified tissue, healthy tissue, and combinations thereof. In some embodiments, the tissue sample is derived from malignant tissue comprising a tumor, sarcoma, leukemia, or a derivative thereof.
[0035] In some embodiments, the circularized double-stranded nucleic acid molecule includes DNA, complementary DNA, derivatives thereof, or combinations thereof. In some embodiments, the circularized double-stranded nucleic acid molecule comprises RNA.
[0036] In various aspects, this disclosure provides methods for processing or analyzing circular nucleic acid molecules, comprising: (a) providing a cell-free composition comprising the circular nucleic acid molecule, the circular nucleic acid molecule comprising (i) a target region and (ii) a nick site at a known distance from the target region; and (b) creating a nick at the nick site of the circular nucleic acid molecule. In some embodiments, (b) is performed under cell-free conditions.
[0037] In some embodiments, the nick site is no more than 100,000 nucleotides away from the target site. In some embodiments, the nick site is no more than 50,000, 10,000, 5,000, 1,000, 500, 400, 300, 200, 100, 90, 80, 70, 60, 50, 40, 30, 25, 20, 15, 10, or 5 nucleotides away from the target site. In some embodiments, the target site comprises up to about 500,000 nucleotides. In some embodiments, the target site comprises up to about 100,000, 50,000, 10,000, 5,000, 1,000, 500, 400, 300, 200, 100, 50, 40, 30, 25, 20, 15, 10, 9, 8, 7, 6, 5, 4, 3, 2, or 1 nucleotide.
[0038] In some embodiments, the nucleic acid molecule comprises a double-stranded nucleic acid molecule. In some embodiments, the cleavage...The cleavage site is part of the sense strand of the circular nucleic acid molecule. In some embodiments, the cleavage site is part of the antisense strand of the circular nucleic acid molecule. In some embodiments, the method further includes determining the cleavage site based at least in part on the position of the target site relative to at least one reference sequence. In some embodiments, the at least one reference sequence comprises a common sequence of at least a portion of a gene. In some embodiments, the cleavage site is endogenous to the circular nucleic acid molecule. In some embodiments, the cleavage site is exogenous to the circular nucleic acid molecule, and the determination includes inserting the exogenous cleavage site into the circular nucleic acid molecule.
[0039] In some embodiments, the nucleic acid molecule further includes a cleavage enzyme binding site specific to the cleavage enzyme, and in (b) further includes providing the nucleic acid molecule with a cleavage enzyme under conditions sufficient to cause the cleavage enzyme to associate with the cleavage enzyme binding site and generate a cleavage. In some embodiments, the probability of the cleavage enzyme binding site occurring without any mismatch is at most once per 1 x 10⁴ base pairs. In some embodiments, the probability of the nicking enzyme binding site appearing without any mismatches (see page 5 / 69 of CN 121472373 A) is at most once per 5x10⁴, 7x10⁴, 1x10⁵, 1x10⁶, 1x10⁷, 1x10⁸, 1x10⁹, 1x10¹⁰, 1x10¹¹, or 1x10¹² base pairs. In some embodiments, the nicking enzyme binding site comprises at least 5 bases. In some embodiments, the nicking enzyme binding site comprises at least 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 35, 40, 45, or 50 bases. In some embodiments, the nicking enzyme binding site is endogenous to the nucleic acid molecule. In some embodiments, the nicking enzyme binding site is exogenous to the nucleic acid molecule. In some embodiments, the method further includes inserting the exogenous nicking enzyme binding site into the nucleic acid molecule prior to (b). In some embodiments, the method further includes circularizing the nucleic acid molecule prior to the insertion. In some embodiments, the method further includes circularizing the nucleic acid molecule after the insertion. In some embodiments, the nicking enzyme binding site is no more than 30 nucleotides away from the nicking site. In some embodiments, the nicking enzyme binding site is no more than 25, 20, 19, 18, 17, 16, 15, 14, 13, 12, 11, 10, 9, 8, 7, 6, 5, 4, 3, 2, or 1 nucleotide away from the nicking site. In some embodiments, the nicking enzyme binding site comprises the nicking site.
[0040] In some embodiments, the method further includes sequencing the circular nucleic acid molecule. In some embodiments, the sequencing includes (i) extending the circular nucleic acid molecule from the nick to produce a growth chain that is sequence complementary to at least a portion of the chain of the circular nucleic acid molecule, and (ii) obtaining sequence information of at least a portion of the growth chain. In some embodiments, obtaining the sequence information includes detecting the at least a portion of the growth chain. In some embodiments, the extension reaction includes contacting the circular nucleic acid molecule with the nucleotide coupled to the tag under conditions sufficient to incorporate the nucleotide into the growth chain, and wherein obtaining the sequence information includes detecting the tag. In some embodiments, the method further includes releasing the tag from the nucleotide when incorporating the nucleotide into the growth chain. In some embodiments, the extension reaction is performed without the use of oligonucleotide primers. In some embodiments, the extension reaction includes rolling circle amplification.
[0041] In some embodiments, the sequencing includes (i) subjecting the circular nucleic acid molecule to a cleavage reaction to cleave at least a portion of the strand of the double-stranded nucleic acid molecule, and (ii) obtaining sequence information of the at least a portion of the strand. In some embodiments, obtaining the sequence information includes detecting the at least a portion of the strand. In some embodiments, the sequencing includes nanopore-based sequencing. In some embodiments, at least a portion of the target site has or is suspected of having one or more sequencing variants compared to at least one reference sequence, and wherein the sequencing is for identifying the presence of the at least a portion of the target site. In some embodiments, the one or more sequencing variants indicate mutations in a gene. In some embodiments, the at least one reference sequence includes a common sequence of at least a portion of the gene.
[0042] In some embodiments, the method further includes, prior to (a), circularizing at least a linear nucleic acid molecule into a circular nucleic acid molecule. In some embodiments, the at least linear nucleic acid molecule is part of an amplification product. In some embodiments, the method further includes, prior to (b), amplifying the circular nucleic acid molecule to generate multiple copies of the circular nucleic acid molecule.
[0043] In some embodiments, the circular nucleic acid molecule contains a recognition sequence, and the method further includes enriching the circular nucleic acid molecule from a random nucleic acid molecule library at least in part based on the recognition sequence. In some embodiments, the enrichment includes generating a selected circular nucleic acid molecule library, wherein each circular nucleic acid molecule in at least 5% of the selected library contains the recognition sequence. In some embodiments, at least 10%, 20%, 30%, 40%, ...Each circular nucleic acid molecule contains the recognition sequence in 50%, 60%, 70%, 80%, 90%, or 95% of its components. In some embodiments, the probability of the recognition sequence occurring without any mismatch is at most once per 1x10⁴ base pairs. In some embodiments, the probability of the recognition sequence occurring without any mismatch is at most once per 5x10⁴, 7x10⁴, 1x10⁵, 1x10⁶, 1x10⁷, 1x10⁸, 1x10⁹, 1x10¹⁰, 1x10¹¹, or 1x10¹² base pairs. In some embodiments, the recognition sequence contains at least 5 bases. In some embodiments, the recognition sequence comprises at least 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 35, 40, 45, or 50 bases. In some embodiments, the enrichment comprises (i) binding a recognition portion complementary to the recognition sequence to the circular nucleic acid molecule to form a recognition complex, and (ii) extracting the recognition complex. In some embodiments, the enrichment is performed before (a) or after (b). In some embodiments, the enrichment is performed before (a) and after (b).
[0044] In some embodiments, the circular nucleic acid molecule is derived from or derived from a biological sample of the subject. In some embodiments, the biological sample comprises a cell-free biological sample of the subject. In some embodiments, the circular nucleic acid molecule is derived from or derived from a cell-free nucleic acid molecule derived from the cell-free biological sample. In some embodiments, the cell-free nucleic acid molecule comprises circulating tumor nucleic acid molecules or amniotic fluid nucleic acid molecules. In some embodiments, the biological sample comprises a tissue sample of the subject. In some embodiments, the circular nucleic acid molecule is derived from a genomic nucleic acid molecule derived from the tissue sample. In some embodiments, the tissue sample is derived from: infected tissue, diseased tissue, malignant tissue, calcified tissue, healthy tissue, and combinations thereof. In some embodiments, the tissue sample is derived from malignant tissue comprising a tumor, sarcoma, leukemia, or a derivative thereof.
[0045] In some embodiments, the circular nucleic acid molecule comprises DNA, complementary DNA, derivatives thereof, or combinations thereof. In some embodiments, the circular nucleic acid molecule comprises RNA.
[0046] In various aspects, this disclosure provides a reaction mixture for processing or analyzing circular nucleic acid molecules, comprising: a cell-free composition comprising the circular nucleic acid molecule, the circular nucleic acid molecule comprising (i) a target site and (ii) a nick site at a known distance from the target site; and at least one of the circular nucleic acid moleculesThe nick site produces the nicking enzyme. In some embodiments, the reaction mixture is a cell-free reaction mixture.
[0047] In some embodiments, the nick site is no more than 100,000 nucleotides away from the target site. In some embodiments, the nick site is no more than 50,000, 10,000, 5,000, 1,000, 500, 400, 300, 200, 100, 90, 80, 70, 60, 50, 40, 30, 25, 20, 15, 10, or 5 nucleotides away from the target site.
[0048] In some embodiments, the target site comprises up to about 500,000 nucleotides. In some embodiments, the target site comprises up to about 100,000, 50,000, 10,000, 5,000, 1,000, 500, 400, 300, 200, 100, 50, 40, 30, 25, 20, 15, 10, 9, 8, 7, 6, 5, 4, 3, 2, or 1 nucleotide.
[0049] In some embodiments, the circular nucleic acid molecule comprises a circular double-stranded nucleic acid molecule. In some embodiments, the cleavage is on the sense strand of the circular double-stranded nucleic acid molecule. In some embodiments, the cleavage is on the antisense strand of the circular double-stranded nucleic acid molecule. In some embodiments, the cleavage site is endogenous to the circular nucleic acid molecule. In some embodiments, the cleavage site is exogenous to the circular nucleic acid molecule.
[0050] In some embodiments, the nucleic acid molecule further comprises an enzyme binding site specific to the at least one enzyme. In some embodiments, the probability of the enzyme binding site occurring without any mismatch is at most once per 1x10⁴ base pairs. In some embodiments, the probability of the enzyme binding site occurring without any mismatch is at most once per 5x10⁴, 7x10⁴, 1x10⁵, 1x10⁶, 1x10⁷, 1x10⁸, 1x10⁹, 1x10¹⁰, 1x10¹¹, or 1x10¹² base pairs. In some embodiments, the enzyme binding site comprises at least 5 bases. In some embodiments, the enzyme binding site comprises at least 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 35, 40, 45, or 50 bases. In some embodiments, the enzyme binding site is endogenous to the circular nucleic acid molecule described on page 7 / 69 of the specification (CN 121472373 A). In some embodiments, the enzyme binding site is exogenous to the circular nucleic acid molecule. In some embodiments, the enzyme binding site is no more than 30 nucleotides away from the cleavage site. In some...In some embodiments, the enzyme binding site is no more than 25, 20, 19, 18, 17, 16, 15, 14, 13, 12, 11, 10, 9, 8, 7, 6, 5, 4, 3, 2, or 1 nucleotide from the nick site. In some embodiments, the enzyme binding site comprises the nick site.
[0051] In some embodiments, the reaction mixture further comprises at least a second enzyme that extends from the nick to produce a growth chain that is sequence complementary to at least a portion of the circular nucleic acid molecule. In some embodiments, the circular nucleic acid molecule is a circular double-stranded nucleic acid molecule, and the growth chain is sequence complementary to at least a portion of the chain of the circular double-stranded nucleic acid molecule. In some embodiments, the reaction mixture further comprises at least one nucleotide coupled to a tag, wherein the at least second enzyme incorporates the nucleotide into the growth chain. In some embodiments, when the nucleotide is incorporated into the growth chain, the at least second enzyme releases the tag from the nucleotide. In some embodiments, the at least second enzyme performs the extension reaction without using oligonucleotide primers. In some embodiments, the extension reaction comprises rolling circle amplification. In some embodiments, the at least second enzyme comprises a polymerase.
[0052] In some embodiments, the reaction mixture further comprises at least a third enzyme, which performs a cleavage reaction from the nick to cleave at least a portion of the circular nucleic acid molecule. In some embodiments, the circular nucleic acid molecule is a circular double-stranded nucleic acid molecule, and wherein the at least third enzyme cleaves at least a portion of the strand of the circular double-stranded nucleic acid molecule.
[0053] In some embodiments, at least a portion of the circular nucleic acid molecule has or is suspected of having one or more variants compared to at least one reference sequence. In some embodiments, the reaction mixture is used to prepare at least one composition for sequencing to identify the presence of at least a portion of the circular nucleic acid molecule. In some embodiments, the one or more sequencing variants indicate mutations in a gene. In some embodiments, the at least one reference sequence comprises a common sequence of at least a portion of the gene.
[0054] In some embodiments, the circular nucleic acid molecule contains a recognition sequence. In some embodiments, the reaction mixture further comprises a recognition moiety associated with the recognition sequence to enrich the circular nucleic acid molecules from a random nucleic acid molecule library in the composition, at least partially based on the recognition sequence. In some embodiments, the recognition moiety comprises at least one oligonucleotide complementary to at least the recognition sequence. In some embodiments, the composition comprises a selected library of circular nucleic acid molecules, wherein each circular nucleic acid molecule in at least 5% of the library contains the recognition sequence.The recognition sequence is specified. In some embodiments, each circular nucleic acid molecule in at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 95% of the selected library contains the recognition sequence. In some embodiments, the probability of the recognition sequence occurring without any mismatch is at most once per 1x10⁴ base pairs. In some embodiments, the probability of the recognition sequence occurring without any mismatch is at most once per 5x10⁴, 7x10⁴, 1x10⁵, 1x10⁶, 1x10⁷, 1x10⁸, 1x10⁹, 1x10¹⁰, 1x10¹¹, or 1x10¹² base pairs. In some embodiments, the recognition sequence contains at least 5 bases. In some embodiments, the identification sequence comprises at least 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 35, 40, 45, or 50 bases.
[0055] In some embodiments, the circular nucleic acid molecule is derived from or derived from a biological sample of the subject. In some embodiments, the biological sample comprises a cell-free biological sample of the subject. In some embodiments, the circular nucleic acid molecule is derived from or derived from a cell-free nucleic acid molecule derived from the cell-free biological sample. In some embodiments, as described on page 8 / 69 of CN 121472373 A, the cell-free nucleic acid molecule comprises circulating tumor nucleic acid molecules or amniotic fluid nucleic acid molecules. In some embodiments, the biological sample comprises a tissue sample of the subject. In some embodiments, the circular nucleic acid molecule is derived from a genomic nucleic acid molecule derived from the tissue sample. In some embodiments, the tissue sample is derived from: infected tissue, diseased tissue, malignant tissue, calcified tissue, healthy tissue, and combinations thereof. In some embodiments, the tissue sample is derived from malignant tissue comprising tumors, sarcomas, leukemia, or derivatives thereof.
[0056] In some embodiments, the circular nucleic acid molecule comprises DNA, complementary DNA, derivatives thereof, or combinations thereof. In some embodiments, the circular nucleic acid molecule comprises RNA.
[0057] In various aspects, this disclosure provides cell-free libraries of circular nucleic acid molecules, wherein the circular nucleic acid molecules of each individual in at least 5% of the library comprise (i) a target site and (ii) a cut at a known distance from the target site.
[0058] In some embodiments, the circular nucleic acid molecules of each individual in at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 95% of the library comprise (i) the target site and (ii) a cut at a known distance from the target site.The incision is made at the known distance. In some embodiments, the circularized nucleic acid molecule of the individual further includes a recognition sequence. In some embodiments, the probability of the recognition sequence occurring without any mismatch is at most once per 1x10⁴ base pairs. In some embodiments, the probability of the recognition sequence occurring without any mismatch is at most once per 5x10⁴, 7x10⁴, 1x10⁵, 1x10⁶, 1x10⁷, 1x10⁸, 1x10⁹, 1x10¹⁰, 1x10¹¹, or 1x10¹² base pairs. In some embodiments, the recognition sequence contains at least 5 bases. In some embodiments, the identification sequence comprises at least 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 35, 40, 45, or 50 bases.
[0059] In some embodiments, (i) the first target site of the first individual nucleic acid molecule and (ii) the second target site of the second individual nucleic acid molecule are the same. In some embodiments, (i) the first known distance between the first nick and the first target site of the first individual nucleic acid molecule and (ii) the second known distance between the second nick and the second target site of the second individual nucleic acid molecule are the same. In some embodiments, (i) the first known distance between the first nick and the first target site of the first individual nucleic acid molecule and (ii) the second known distance between the second nick and the second target site of the second individual nucleic acid molecule are different.
[0060] In some embodiments, (i) the first target site of the first individual nucleic acid molecule and (ii) the second target site of the second individual nucleic acid molecule are different. In some embodiments, (i) the first known distance between the first nick and the first target site of the first individual nucleic acid molecule and (ii) the second known distance between the second nick and the second target site of the second individual nucleic acid molecule are the same. In some embodiments, (i) the first known distance between the first nick and the first target site of the first individual nucleic acid molecule and (ii) the second known distance between the second nick and the second target site of the second individual nucleic acid molecule are different.
[0061] In some embodiments, the nick site is no more than 100,000 nucleotides away from the target site. In some embodiments, the nick site is no more than 50,000, 10,000, 5,000, 1,000, 500, 400, 300, 200, 100, 90, 80, 70, 60, 50, 40, 30, 25, 20, 15, 10, or 5 nucleotides away from the target site.
[0062] In some embodiments, the target site comprises up to about 500 nucleotides.,000 nucleotides. In some embodiments, the target site comprises up to about 100,000, 50,000, 10,000, 5,000, 1,000, 500, 400, 300, 200, 100, 50, 40, 30, 25, 20, 15, 10, 9, 8, 7, 6, 5, 4, 3, 2, or 1 nucleotide.
[0063] In some embodiments, at least a portion of the circular nucleic acid molecule has or is suspected of having one or more variants compared to at least one reference sequence. In some embodiments, the one or more sequencing variants indicate mutations in the gene. In some embodiments, the at least one reference sequence comprises a common sequence of at least a portion of the gene.
[0064] In some embodiments, the circular nucleic acid molecule comprises a circular double-stranded nucleic acid molecule. In some embodiments, the nick is on the sense strand of the circular double-stranded nucleic acid molecule. In some embodiments, the incision is made on the antisense strand of the circular double-stranded nucleic acid molecule.
[0065] In some embodiments, the circular nucleic acid molecule is derived from or derived from a biological sample of the subject. In some embodiments, the biological sample includes a cell-free biological sample of the subject. In some embodiments, the circular nucleic acid molecule is derived from or derived from a cell-free nucleic acid molecule from the cell-free biological sample. In some embodiments, the cell-free nucleic acid molecule includes circulating tumor nucleic acid molecules or amniotic fluid nucleic acid molecules. In some embodiments, the biological sample includes a tissue sample of the subject. In some embodiments, the circular nucleic acid molecule is derived from or derived from a genomic nucleic acid molecule from the tissue sample. In some embodiments, the tissue sample is derived from: infected tissue, diseased tissue, malignant tissue, calcified tissue, healthy tissue, and combinations thereof. In some embodiments, the tissue sample is from malignant tissue comprising a tumor, sarcoma, leukemia, or a derivative thereof.
[0066] In some embodiments, the circular nucleic acid molecule includes DNA, complementary DNA, derivatives thereof, or combinations thereof. In some embodiments, the circular nucleic acid molecule includes RNA.
[0067] In various aspects, this disclosure provides methods for processing or analyzing nucleic acid molecules, comprising: (a) providing the nucleic acid molecule adjacent to a nanopore, and contacting the nucleic acid molecule with a tagged nucleotide under conditions sufficient to incorporate a nucleotide into a nucleic acid chain complementary to at least a portion of the nucleic acid molecule, wherein at least a portion of the tag is located within the nanopore when the nucleotide is incorporated into the nucleic acid chain; and (b) detecting one or more signals indicating impedance or impedance changes in the nanopore when at least a portion of the tag is within the nanopore.(a) and (c) using the one or more signals to identify the nucleotide incorporated into the nucleic acid chain.
[0068] In some embodiments, the one or more signals are current or voltage. In some embodiments, the method further includes measuring a current or a change therein when at least a portion of the tag is located within the nanopore. In some embodiments, the one or more signals are not tunneling currents. In some embodiments, the current is not a Faraday current. In some embodiments, the nanopore is part of a circuit containing a tunnel junction. In some embodiments, the nanopore includes a plurality of electrodes, and wherein (c) includes using the plurality of electrodes to detect the one or more signals.
[0069] In some embodiments, the nanopore includes a protein nanopore or a solid nanopore.
[0070] In some embodiments, the method further includes, in (a), releasing the tag from the nucleotide when incorporating the nucleotide into the nucleic acid chain. In some embodiments, the nucleic acid molecule includes a circular nucleic acid molecule. In some embodiments, the method further includes, prior to (b), performing rolling circle amplification (RCA) on the nucleic acid molecule to generate the nucleic acid chain. In some embodiments, the circular nucleic acid molecule is a circular double-stranded nucleic acid molecule. In some embodiments, the incorporation is performed without the use of oligonucleotide primers. In some embodiments, the provision includes coupling at least one enzyme that performs incorporation into (i) at least a portion of the nanopore or (ii) a membrane having the nanopore. In some embodiments, the membrane is a lipid bilayer. In some embodiments, the membrane is a solid membrane. In some embodiments, the coupling includes conjugating the at least one enzyme to the nanopore or the membrane.
[0071] In various aspects, this disclosure provides a system for processing or analyzing nucleic acid molecules, comprising: a nanopore configured to (i) receive at least a portion of a tag when a tagged nucleotide is incorporated into a nucleic acid chain, wherein the nucleic acid chain is complementary to at least a portion of the nucleic acid molecule; and (ii) detect one or more signals indicative of impedance or impedance change in the nanopore when the at least a portion of the tag is within the nanopore, wherein the one or more signals can be used to identify the nucleotide incorporated into the nucleic acid chain.
[0072] In some embodiments, the one or more signals are current or voltage. In some embodiments, the nanopore is configured to measure current or a change thereof when at least a portion of the tag is located within the nanopore. In some embodiments, the one or more signals are not tunneling currents. In some embodiments, the current is not a Faraday current.Current. In some embodiments, the nanopore is part of a circuit comprising a tunnel junction. In some embodiments, the nanopore includes a plurality of electrodes configured to detect the one or more signals.
[0073] In some embodiments, the nanopore comprises a protein nanopore or a solid nanopore.
[0074] In some embodiments, when the nucleotide is incorporated into the nucleic acid chain, at least a portion of the tag is released from the nucleotide. In some embodiments, the system further includes at least one enzyme configured to perform the incorporation. In some embodiments, the at least one enzyme is coupled to (i) at least a portion of the nanopore or (ii) a membrane having nanopores. In some embodiments, the at least one enzyme is conjugated to (i) at least a portion of the nanopore or (ii) a membrane having nanopores.
[0075] In some embodiments, the membrane is a lipid bilayer. In some embodiments, the membrane is a solid membrane. In some embodiments, the nanopore or membrane is configured to bind at least a portion of the at least one enzyme. In some embodiments, the at least one enzyme is configured to bind to at least a portion of the nanopore or at least a portion of the membrane.
[0076] In various aspects, this disclosure provides a method for sequencing a plurality of polynucleotides, comprising: circularizing an individual polynucleotide to provide a plurality of cyclic polynucleotides; nicking one strand of the cyclic polynucleotides using a nicking enzyme to provide a nick site on each of the cyclic polynucleotides; binding a polymerase to the nick site; and sequencing the cyclic polynucleotides.
[0077] In some embodiments, the polynucleotides comprise double-stranded polynucleotides. In some embodiments, the polynucleotides comprise single-stranded polynucleotides, and the method comprises adding primer sequences to the adaptor region of the single-stranded polynucleotide.
[0078] In some embodiments of any of the subject methods, the method further comprises amplifying the plurality of polynucleotides prior to circularizing the individual polynucleotides. In some embodiments of any of the subject methods, the polynucleotides comprise DNA, cDNA, ctDNA, or any combination thereof. In some embodiments of any of the subject methods, circularizing the plurality of polynucleotides comprises reacting the plurality of polynucleotides with a ligase. In some embodiments of any of the subject methods, the cyclic polynucleotide comprises a cyclic double-stranded polynucleotide, and the cleavage comprises cleaving the inner strand of the cyclic double-stranded polynucleotide. In some embodiments of any of the subject methods, the cyclic polynucleotide comprises a cyclic double-stranded polynucleotide, and the cleavage comprises cleaving the outer strand of the cyclic double-stranded polynucleotide.
[0079] In some embodiments of any of the subject methods, the cleavage enzyme comprises sgRNA-CRISPR Cas9n (Cas9(D10A) nicking enzyme complex. In some embodiments, the sgRNA comprises a nucleotide sequence complementary to the target nucleotide sequence. In some embodiments of any subject method, the nicking enzyme comprises a Cas9n (Cas9 D10A) nicking enzyme.
[0080] In some embodiments of any subject method, the polymerase comprises a linker. In some embodiments, the method further comprises binding the linker to a nanopore after binding the polymerase comprising the linker to the nicking site. In some embodiments of any subject method, the polymerase comprises a linker and a protein nanopore bound to the linker.
[0081] In some embodiments of any subject method, the method further comprises cleaving the polynucleotide to provide a targeted polynucleotide fragment prior to cyclizing the plurality of polynucleotides. In some embodiments of any subject method, page 11 / 69 of CN 121472373 A, the cleavage comprises binding a biotinylated sgRNA-CRISPER Cas9n complex to the polynucleotide. In some embodiments, the biotinylated sgRNA comprises a nucleotide sequence complementary to the target polynucleotide sequence. In some embodiments of any of the subject methods, the method further includes enriching the targeted polynucleotide fragment.
[0082] In some embodiments of any of the subject methods, the polymerase exhibits strong strand substitution activity.
[0083] In some embodiments of any of the subject methods, the sequencing includes sequencing the circular polynucleotide more than once while binding to the nanopore. In some embodiments of any of the subject methods, the sequencing includes forward sequencing and reverse sequencing. In some embodiments of any of the subject methods, the polynucleotide comprises a double-stranded polynucleotide, and the uncut polynucleotide chain contains a template for amplification and sequencing. In some embodiments of any of the subject methods, the polynucleotide comprises genomic DNA, cDNA, cell-free DNA, ctDNA, or a combination thereof. In some embodiments of any of the subject methods, the sequencing includes rolling circle amplification and transcription. In some embodiments of any of the subject methods, the sequencing includes whole-genome sequencing. In some embodiments of any of the subject methods, the sequencing includes targeted sequencing. In some embodiments of any of the subject methods, the targeted sequencing includes identifying sequence variants.
[0084] Another aspect of this disclosure provides a non-transitory computer-readable medium comprising machine-executable code that, when executed by one or more computer processors, implements any of the methods described above or elsewhere herein.
[0085] Another aspect of this disclosure provides a computing medium comprising one or more computer processors and coupled thereto...A system of computer memory. The computer memory contains machine-executable code that, when executed by the one or more computer processors, implements any of the methods described above or elsewhere herein.
[0086] Other aspects and advantages of this disclosure will become apparent to those skilled in the art from the following detailed description of illustrative embodiments in which only the present disclosure is shown and described. As will be appreciated, this disclosure is capable of having other different implementations, and certain details thereof can be modified in various obvious ways without departing from this disclosure. Therefore, the drawings and detailed descriptions should be regarded as illustrative in nature and not restrictive.
[0087] Incorporation of References
[0088] All publications, patents and patent applications mentioned in this specification are incorporated herein by reference to the extent that each publication, patent or patent application is specifically and individually indicated to be incorporated by reference. Where a publication and patent or patent application incorporated by reference conflicts with the disclosure contained herein, this specification is intended to supersede and / or give preference to any such contradictory material. Brief Description of the Drawings
[0089] The novel features of the invention are set forth in the appended claims. A better understanding of the features and advantages of the invention will be obtained by referring to the following detailed description of illustrative embodiments in which the principles of the invention are utilized, and the accompanying drawings (also referred to herein as “Figures”), in which:
[0090] Figure 1A schematically illustrates an example method for providing a circular nucleic acid with a nick;
[0091] Figures 1B and 1C schematically illustrate example methods for isolating or enriching linear or circular nucleic acids containing recognition sites;
[0092] Figure 1D schematically illustrates an example method for providing a circular nucleic acid containing a nick at a specific location within the circular nucleic acid by using one or more uracil-specific enzymes; Specification 12 / 69 pages 14 CN 121472373 A
[0093] Figures 2A and 2B schematically illustrate an example method for generating a nick at a known distance from the target site within the circular nucleic acid;
[0094] Figures 2C and 2D schematically illustrate an example method for isolating or enriching a circular nucleic acid containing both a recognition site and a target site;
[0095] Figure 3 schematically illustrates an example method for double-stranded nucleic acid sequencing;
[0096] Figure 4A schematically illustrates an example method for targeted sequencing using nanopore sequencing;
[0097] Figure 4B schematically illustrates an example method for genome sequencing using nanopore sequencing;
[0098] Figures 5A to 5D schematically illustrate exemplary nanopore sequencing systems for obtaining sequence information from one or more nucleic acid samples;
[0099] Figure 6 illustrates a computer system programmed or otherwise configured to implement the methods provided herein.
[0100] Figure 7A shows an example of a gel electrophoresis image of a sample containing multiple circularized single-stranded nucleic acids, and Figure 7B shows an example of a fluorescence image of the RCA product from the circularized single-stranded nucleic acid;
[0101] Figure 7C shows an example of a gel electrophoresis image of a sample containing multiple circularized double-stranded nucleic acids, and Figure 7D shows an example of a fluorescence image of the RCA product from the circularized double-stranded nucleic acid; and
[0102] Figure 8 shows an example of a gel electrophoresis image of a circular double-stranded nucleic acid complexed with (i) a wild-type polymerase and (ii) a mutant polymerase. Detailed Description
[0103] Although various embodiments of the invention have been shown and described herein, it will be readily understood by those skilled in the art that such embodiments are provided by way of example only. Many modifications, alterations and substitutions may be conceived by those skilled in the art without departing from the invention. It should be understood that various alternatives to the embodiments of the invention described herein may be employed.
[0104] As used in the specification and claims, unless the context clearly indicates otherwise, the singular forms “a,” “an,” and “the” may include plural referents. For example, the term “transmembrane receptor” may include multiple transmembrane receptors.
[0105] As used herein, the term “about” or “approximately” may refer to an acceptable range of error for a particular value as determined by those skilled in the art, which will depend in part on how the value is measured or determined, i.e., the limitations of the measurement system. For example, according to practice in the art, “about” may mean within or greater than one standard deviation. Alternatively, “about” may mean a range of up to 20%, up to 10%, up to 5%, or up to 1% of a given value. Alternatively, particularly for biological systems or processes, the term may refer to an order of magnitude of a value, preferably within 5 times, more preferably within 2 times. Where a particular value is described in this application and claims, unless otherwise stated, the term “about” should be assumed to mean within an acceptable range of error for that particular value.
[0106] As used herein, the term “cell” generally refers to a biological cell or cell derivative. A cell can be the basic structural, functional, and / or biological unit of a living organism. A cell can be derived from any organism having one or more cells. Some non-limiting examples include: prokaryotic cells, eukaryotic cells, bacterial cells, archaea cells, cells of single-celled eukaryotes, protozoan cells, cells from plants (e.g., cells from plant crops, fruits, vegetables, grains, soybeans, corn, maize, wheat, seeds, tomatoes, rice, cassava, sugarcane, squash, hay, potatoes, cotton, hemp, tobacco, flowering plants, conifers, gymnosperms, ferns, lycophytes, hornworts, bryophytes, mosses), and algal cells (e.g., *Botryococcus*).Cells include *Chlamydomonas braunii*, *Chlamydomonas reinhardtii*, *Nannochloropsis gaditana*, *Chlorella pyrenoidosa*, *Sargassum patens C. Agardh*, etc., seaweed (e.g., kelp), fungal cells (e.g., yeast cells, mushroom cells), animal cells, cells from invertebrates (e.g., fruit flies, cnidarians, echinoderms, nematodes, etc.), cells from vertebrates (e.g., fish, amphibians, reptiles, birds, mammals), and cells from mammals (e.g., pigs, cattle, goats, sheep, rodents, rats, mice, non-human primates, humans, etc.). Sometimes cells do not originate from natural organisms (e.g., cells can be synthetic, sometimes called artificial cells).
[0107] As may be used interchangeably herein, the terms “nucleotide,” “nucleobase,” and “base” generally refer to a base-sugar-phosphate combination. A nucleotide may include synthetic nucleotides. A nucleotide may include synthetic nucleotide analogs. A nucleotide may be a monomeric unit of a nucleic acid sequence (e.g., deoxyribonucleic acid (DNA) and ribonucleic acid (RNA)). The term nucleotide may include ribonucleoside triphosphates (adenosine triphosphate (ATP), uridine triphosphate (UTP), cytosine triphosphate (CTP), guanosine triphosphate (GTP), uridine triphosphate (UTP)) and deoxyribonucleoside triphosphates (such as dATP, dCTP, dITP, dUTP, dGTP, dTTP) or derivatives thereof. Such derivatives may include, for example, [αS]dATP, 7-denitro-dGTP, and 7-denitro-dATP, as well as nucleotide derivatives that confer nuclease resistance on nucleic acid molecules containing them. As used herein, the term nucleotide generally refers to dideoxynucleotide triphosphates (ddNTPs) and their derivatives. Illustrative examples of dideoxynucleotide triphosphates may include, but are not limited to, ddATP, ddCTP, ddGTP, ddITP, and ddTTP. Nucleotides may be unlabeled or detectably labeled. Labeling may also be performed using quantum dots. Detectable labeling may include, for example, radioactive isotopes, fluorescent labels, chemiluminescent labels, bioluminescent labels, and enzyme labels. Fluorescent labeling of nucleotides may include, but is not limited to, fluorescein, 5-carboxyfluorescein (FAM), 2′7′-dimethoxy-4′5-dichloro-6-carboxyfluorescein (JOE), rhodamine, 6-carboxyrhodamine (R6G), N,N,N′,N′-tetramethyl-6-carboxyrhodamine (TAMRA), 6-carboxy-X-rhodamine (ROX), and 4-(4′-dimethylaminophenylazo)benzoic acid.(DABCYL), Cascade Blue, Oregon Green, Texas Red, Cyanide and 5-(2′-aminoethyl)aminonaphthalene-1-sulfonic acid (EDANS). Specific examples of fluorescently labeled nucleotides may include [R6G]dUTP, [TAMRA]dUTP, [R110]dCTP, [R6G]dCTP, [TAMRA]dCTP, [JOE]ddATP, [R6G]ddATP, [FAM]ddCTP, [R110]ddCTP, [TAMRA]ddGTP, [ROX]ddTTP, [dR6G]ddATP, [dR110]ddCTP, [dTAMRA]ddGTP, and [dROX]ddTTP, all available from Perkin Elmer, Foster City, Calif.; and FluoroLink deoxynucleotides, FluoroLink Cy3-dCTP, FluoroLink Cy5-dCTP, FluoroLink Fluor X-dCTP, FluoroLink Cy3-dUTP, and FluoroLink, all available from Amersham, Arlington Heights, Illinois. Cy5-dUTP; luciferin-15-dATP, luciferin-12-dUTP, tetramethyl-rhodamine-6-dUTP, IR770-9-dATP, luciferin-12-ddUTP, luciferin-12-UTP, and luciferin-15-2'-dATP, available from Boehringer Mannheim, Indianapolis, Ind.; and chromosome-marked nucleotides, BODIPY-FL-14-UTP, BODIPY-FL-4-UTP, BODIPY-TMR-14-UTP, BODIPY-TMR-14-dUTP, BODIPY-TR-14-UTP, and BODIPY-TR-14- , available from Molecular Probes, Eugene, Oreg. dUTP, Cascade Blue-7-UTP, Cascade Blue-7-dUTP, Fluorescein-12-UTP, Fluorescein-12-dUTP, Oregon Green 488-5-dUTP, Rhodamine Green-5-UTP, Rhodamine Green-5-dUTP, Tetramethylrhodamine-6-UTP, Tetramethylrhodamine-6-dUTP, Texas Red-5-UTP, Texas Red-5-dUTP, and Texas Red-12-dUTP. Nucleotides can also be labeled or marked by chemical modifications. A single chemically modified nucleotide can be a biotinylated dNTP. Some non-restricted biotinylated dNTPs...Examples of biotin-dATP may include bio-N6-ddATP (e.g., bio-N6-ddATP, bio-14-dATP), biotin-dCTP (e.g., biotin-11-dCTP, biotin-14-dCTP), and biotin-dUTP (e.g., biotin-11-dUTP, biotin-16-dUTP, biotin-20-dUTP).
[0108] Naturally occurring nucleotides guanine, cytosine, adenine, thymine, and uracil may be abbreviated as G, C, A, T, and U, respectively. Nucleotides may include any subunit that can be incorporated into the growing nucleic acid chain. Such subunits may be A, C, G, T, A, or U, or any other subunit that is specific to one or more complementary A, C, G, T, or U, or complementary to a purine (i.e., A or G or variants thereof) or a pyrimidine (i.e., C, T, or U or variants thereof). Subunits can break down individual nucleic acid bases or base sets (e.g., AA, TA, AT, GC, CG, CT, TC, GT, GT, TG, AC, CA, or their uracil counterparts).
[0109] The terms “polynucleotide,” “oligonucleotide,” “oligomer,” and “nucleic acid” are used interchangeably herein and generally refer to polymeric forms of nucleotides of any length, whether deoxyribonucleotides, ribonucleotides, or analogs thereof, whether single-stranded, double-stranded, or multi-stranded. Polynucleotides can be exogenous or endogenous to the cell. Polynucleotides can be present in a cell-free environment. Polynucleotides can be genes or fragments thereof. Polynucleotides can be DNA. Polynucleotides can be RNA. Polynucleotides can have any three-dimensional structure and can perform any known or unknown function. Polynucleotides can include one or more analogs (e.g., modified backbones, sugars, or nucleobases). In the presence of modifications, the nucleotide structure can be modified before or after polymer assembly. Some non-limiting examples of analogues include: 5-bromouracil, peptide nucleic acids, heteronucleic acids, morpholino derivatives, locked nucleic acids, glycol nucleic acids, threonine nucleic acids, dideoxynucleotides, cordycepin, 7-denitro-GTP, fluorophores (e.g., rhodamine or fluorescein linked to sugars), thiol-containing nucleotides, biotin-linked nucleotides, fluorescent base analogues, CpG islands, methyl-7-guanosine, methylated nucleotides, inosine, thiouridine, pseudouridine, dihydrouridine nucleoside, brassinoside, and woyoside. Non-limiting examples of polynucleotides include coding or non-coding regions of genes or gene segments, loci defined by linkage analysis, exons, introns, messenger RNA (mRNA), transfer RNA (tRNA), ribosomal RNA (rRNA), short interfering RNA (siRNA), short hairpin RNA (shRNA), microRNA (miRNA), ribozymes, and complementary DNA.(cDNA, such as double-stranded cDNA (dd-cDNA) or single-stranded cDNA (ss-cDNA)), circulating tumor DNA (ctDNA), damaged DNA, recombinant polynucleotides, branched polynucleotides, plasmids, vectors, isolated DNA of any sequence, isolated RNA of any sequence, cell-free polynucleotides (including cell-free DNA (cfDNA) and cell-free RNA (cfRNA)), nucleic acid probes (e.g., fluorescence in situ hybridization (FISH) probes), and primers. The sequence of nucleotides may be interrupted by non-nucleotide components. Polynucleotides may contain one or more modified nucleotides, such as methylated nucleotides and nucleotide analogs. In the presence of modifications, the nucleotide structure may be modified before or after polymer assembly. The sequence of nucleotides may be interrupted by non-nucleotide components. Polynucleotides may be further modified after polymerization, such as by conjugation with labeled components.
[0110] As used herein, the term “gene” generally refers to nucleic acids (e.g., DNA, such as genomic DNA and cDNA) that encode RNA transcripts and their corresponding nucleotide sequences. As used herein with respect to genomic DNA, the term "gene" can include intercalated non-coding regions as well as regulatory regions, and may include 5' and 3' ends. In some uses, the term includes transcribed sequences, including 5' and 3' untranslated regions (5'-UTR and 3'-UTR), exons, and introns. In some genes, the transcribed region will contain an "open reading frame" encoding a polypeptide. In some uses of the term, "gene" includes only the coding sequence necessary to encode a polypeptide (e.g., "open reading frame" or "coding region"). Genes may not encode polypeptides, such as ribosomal RNA (rRNA) genes and transfer RNA (tRNA) genes. The term "gene" can include not only transcribed sequences but also non-transcribed regions, including upstream and downstream regulatory regions, enhancers, and promoters. A gene can be an "endogenous gene" or a natural gene located naturally in the genome of an organism. A gene can be an "exogenous gene" or a non-natural gene. A non-natural gene can be a gene that is not normally found in a host organism but is introduced into the host organism through gene transfer (e.g., transgenes). Non-natural genes can be naturally occurring nucleic acid or polypeptide sequences (e.g., non-natural sequences) containing mutations, insertions, and / or deletions.
[0111] As used herein, the term “mutation” generally refers to an alteration in the nucleotide sequence of a normally conserved nucleic acid sequence, resulting in a mutant sequence that differs from the normal (unaltered) or wild-type sequence. Prior to sequencing, the location (e.g., relative to a gene or sample polynucleotide) and sequence of the mutation may be unknown. Alternatively, the location (e.g., relative to a gene or sample polynucleotide) and sequence of the mutation may be known prior to sequencing, in which case sequencing can be performed to detect the sample polynucleotide. (Instructions for use 15 / 69, page 17 CN)121472373 A Whether a mutation exists in the acid. Mutations may include base pair substitutions (e.g., single nucleotide substitutions) and frameshift mutations. Frameshift mutations may require the insertion or deletion of one or more nucleotide pairs.
[0112] As used herein, the term “probe” generally refers to a nucleotide or polynucleotide labeled with a marker (e.g., a fluorescent marker) that can be used to detect or identify its corresponding target nucleotide or polynucleotide in a hybridization reaction by hybridization with the corresponding target sequence. As used interchangeably herein, the terms “nucleotide probe,” “nucleotide tag,” and “labeled nucleotide” generally refer to a probe having a single nucleotide. As used interchangeably herein, the terms “polynucleotide probe,” “polynucleotide tag,” and “labeled polynucleotide” generally refer to a probe having a polynucleotide. A polynucleotide probe may be labeled with at least one marker (e.g., one marker per nucleotide of the polynucleotide probe). The probe may hybridize with one or more target nucleotides or polynucleotides. A polynucleotide probe may be fully complementary to one or more target polynucleotides in a sample, or may contain one or more nucleotides that are not complementary (i.e., mismatched) to one or more target polynucleotides in a sample.
[0113] As may be used interchangeably herein, the terms “complementary,” “complementary sequence,” “complementary,” and “complementarity” generally refer to a sequence that is completely complementary to and hybridizes with a given sequence. A sequence that hybridizes with a given nucleic acid is called a “complementary sequence” or “reverse complementary sequence” of a given molecule, provided that its base sequence in a given region can bind complementary to the base sequence of its binding partner, such that, for example, A-T, A-U, G-C, and G-U base pairs are formed. Typically, a first sequence that can hybridize with a second sequence can hybridize specifically or selectively with the second sequence, such that, during hybridization, hybridization with the second sequence or a group of second sequences is preferred (e.g., more thermodynamically stable under given conditions, such as stringent conditions commonly used in the art) compared to hybridization with a non-target sequence. Typically, hybridizable sequences share a degree of sequence complementarity, such as 25%–100%, across their respective lengths, including at least 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, and 100% sequence complementarity. Their respective lengths may include regions with at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50, or more nucleotides. For purposes such as assessing the percentage of complementarity, sequence identity can be measured using any suitable alignment algorithm, includingHowever, it is not limited to the Needleman-Wunsch algorithm (see, for example, the EMBOSS Needle alignment tool available at www.ebi.ac.uk / Tools / psa / emboss_needle / nucleotide.html, optionally using the default settings), the BLAST algorithm (see, for example, the BLAST alignment tool available at blast.ncbi.nlm.nih.gov / Blast.cgi, optionally using the default settings), or the Smith-Waterman algorithm (see, for example, the EMBOSS Water alignment tool available at www.ebi.ac.uk / Tools / psa / emboss_water / nucleotide.html, optionally using the default settings). Optimal alignment can be evaluated using any suitable parameters of the selected algorithm, including the default parameters.
[0114] Complementarity can be complete or substantial / sufficient. Complete complementarity between two nucleic acids can mean that the two nucleic acids can form a double helix, wherein each base in the double helix is bound to a complementary base via Watson-Crick pairing. Substantial or sufficient complementarity may mean that the sequence in one strand is not completely and / or imperfectly complementary to the sequence in the opposing strand, but that sufficient binding occurs between the bases on the two strands under a set of hybridization conditions (e.g., salt concentration and temperature) to form a stable hybrid complex. Such conditions can be predicted by using the sequence and standard mathematical calculations or by empirically determining the Tm using conventional methods.
[0115] As used herein, the term “hybridization” generally refers to a reaction in which one or more polynucleotides react to form a complex that is stabilized by hydrogen bonds between the bases of nucleotide residues. Hydrogen bonds can occur by Watson-Crick base pairing, Hoogstein binding, or in any other sequence-specific manner depending on base complementarity. The complex may comprise two strands forming a double-stranded structure, three or more strands forming a multi-stranded complex, a self-hybridized strand, or any combination thereof. Hybridization reactions can constitute a step in a broader process, such as initiating PCR or enzymatic cleavage of polynucleotides by endonucleases. A second sequence complementary to the first sequence may be referred to as the “complementary sequence” of the first sequence. The term “hybridizable” applied to polynucleotides generally refers to the ability of a polynucleotide to form a complex that is stabilized by hydrogen bonds between the bases of the nucleotide residues during hybridization.
[0116] As used herein, the term “target polynucleotide” generally refers to a polynucleotide in a nucleic acid molecule or population of nucleic acid molecules having a target sequence. This requires determining the presence, number, and / or variation of the nucleotide sequence, or one or more of them.The term “target sequence” generally refers to a nucleic acid sequence on a single strand of a nucleic acid. A target sequence can be a portion of a gene, a regulatory sequence, genomic DNA, cDNA, ctDNA, RNA (including mRNA, miRNA, rRNA), or others. A target sequence can be a target sequence derived from a sample or a secondary target, such as a product of an amplification reaction. A target polynucleotide can be a portion of a gene (or a fragment thereof) containing one or more mutations.
[0117] As used herein, the term “target site” generally refers to a polynucleotide sequence containing a target polynucleotide (or target nucleotide). The target polynucleotide (or target nucleotide) of a target site can be one or more sequence variants. Instances of one or more sequence variants may include single nucleotide variations, insertions or deletions of one or more nucleotides (e.g., continuous or discontinuous nucleotides), copy number variations (CNVs) consisting of one or more repeats of one or more nucleotides (e.g., CNVs with an average size of at least 1, 5, 10, 50, 100, 150, 200 or more kilobases (kb); CNVs with an average size of at most 200, 150, 100, 50, 10, 5, 1 or less kb), and microsatellite instability (MSI). Target sites may include at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 40, 50, 100, 200, 300, 400, 500, 1,000, 5,000, 10,000, 50,000, 100,000, 500,000 or more nucleotides. Target sites may include up to 500,000, 100,000, 50,000, 10,000, 5,000, 1,000, 500, 400, 300, 200, 100, 50, 40, 30, 25, 20, 15, 10, 9, 8, 7, 6, 5, 4, 3, 2, or 1 nucleotide. Target sites may contain at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 40, 50, 100, 200, 300, 400, 500, 1,000, 5,000, 10,000, 50,000, 100,000, 500,000, or more nucleotides than the target polynucleotide. The target site may contain up to 500,000, 100,000, 50,000, 10,000, 5,000, 1,000, 500, 400, 300, 200, 100, 50, 40, 30, 25, 20, 15, 10, 9, 8, 7, 6, 5, 4, 3, 2, or 1 nucleotide. In some instances, the target site may be the target polynucleotide.
[0118] As used herein, the term “strict condition” generally refers to one or more hybridization conditions under which the target polynucleotide is subjected to a specific cross-linking reaction.Nucleic acids with complementary sequences primarily hybridize with the target sequence and substantially do not hybridize with non-target sequences. Strict conditions can be sequence-dependent and can vary depending on many factors. In some cases, the longer the sequence, the higher the temperature at which the sequence can specifically hybridize with its target sequence.
[0119] As used herein, the term “recognition moiety” generally refers to a molecule (e.g., a small molecule, polynucleotide, protein, its variants, or combinations thereof) capable of interacting with a nucleic acid sequence, i.e., a “recognition sequence” or “recognition site,” such as a desired (or target) nucleic acid sequence. A recognition moiety may contain a domain (e.g., a component containing the domain) capable of binding (e.g., hybridizing) to the recognition sequence. Such a domain may contain one or more amino acids, one or more nucleotides, their variants, or combinations thereof. Alternatively or additionally, the recognition moiety may associate (e.g., bind) with a secondary molecule containing such a domain. In some instances, the recognition moiety may contain a nucleic acid molecule capable of hybridizing with the recognition sequence. In some instances, the recognition moiety may contain components exhibiting specific biological activities, including but not limited to one or more activities of nucleases (e.g., double-stranded nucleases), nickases, transcription activators, transcription repressors, nucleic acid methyltransferases, nucleic acid demethyltransferases, and recombinases. The recognition sequence may include at least 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50 or more nucleotides. The recognition sequence may include up to 50, 45, 40, 35, 30, 25, 24, 23, 22, 21, 20, 19, 18, 17, 16, 15, 14, 13, 12, 11, 10, 9, 8, 7, 6, 5, 4, 3 or fewer nucleotides. Specification 17 / 69 pages 19 CN 121472373 A
[0120] The recognition portion can be used to isolate a desired molecule containing a desired nucleic acid sequence from a plurality of molecules (e.g., a plurality of nucleic acid molecules). The recognition portion can be used to enrich a desired molecule containing a desired nucleic acid in a composition or reaction mixture. In some instances, the recognition portion can be captured by a capture system (e.g., magnetic beads) via one or more interactions (e.g., avidin-biotin binding, magnetic binding, etc.). In one instance, the recognition portion can contain biotin, which can be complexed with streptavidin magnetic beads for separation or enrichment.
[0121] Examples of the recognition portion may include a CRISPR-associated (Cas) system (e.g., Cas protein, including catalytically active or inactive Cas polypeptides); zinc finger nucleases (ZFNs); transcription activator-like effector nucleases (TALENs); broad-spectrum nucleases; RNA-binding proteins (RBPs); CasRNA-binding proteins; recombinases; flipases; transposases; Argonaute (Ago) proteins (e.g., prokaryotic Argonaute (pAgo), archaea Argonaute (aAgo), and eukaryotic Argonaute (eAgo)); variants thereof; and combinations thereof. The recognition portion may include a polynucleotide (e.g., a sequence at least 2, 3, 4, 5, 6, 7, 8, 9, 10, or more nucleotides long) that can be captured by a capture system through one or more interactions, for example, a biotin-tagged polynucleotide sequence that can be captured by one or more avidin-functionalized magnetic beads. At least a portion of the polynucleotide may share complementarity with the recognition sequence of the target nucleic acid molecule.
[0122] As used herein, the term “nicking enzyme” generally refers to a molecule (e.g., an enzyme) that cleaves one strand of a double-stranded nucleic acid molecule (i.e., a “nick” to the double-stranded molecule). A nicking enzyme may be a nuclease that cleaves only a single strand of DNA, either due to its natural function or because it has been engineered (e.g., modified by mutation and / or deletion of one or more nucleotides) to cleave only a single strand of DNA. A nicking enzyme can be an enzyme that produces a nick (e.g., restriction endonuclease, nicking endonuclease, etc.). A nicking enzyme binds to a nick site on a double-stranded nucleic acid molecule to create a nick (or gap) in one strand of the molecule. The nick can be created within the nick site. Alternatively, the nick can be created near the nick site. In some cases, the nicking enzyme can bind to a nicking enzyme binding site adjacent to the nick site. The length of the nick can be at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more nucleotides. The length of the nick can be at most 10, 9, 8, 7, 6, 5, 4, 3, 2 or 1 nucleotide. Examples of nicking enzymes may include Cas systems (e.g., Cas nicking enzymes, such as Cas9n), N.AlwI, Nb.BbvCl, Nt.BbvCl, Nb.BsmI, Nt.BsmAI, Nt.BspQl, Nb.BsrDI, Nt.BstNBI, Nb.BstsCI, Nt.CviPII, Nb.BpulOI, Nt.Bpu lOI, and Nt,Bst9I, variants thereof, and combinations thereof. In some instances, nucleic acid molecules (e.g., self-complementary double-stranded nucleic acid molecules or single-stranded nucleic acid molecules) may contain at least one nicking site that already includes at least one nick.
[0123] As may be used interchangeably herein, the terms “CRISPR-related system,” “Cas system,” and “Cas complex” generally refer to a two-component nucleoprotein complex having a guide RNA (gRNA) and a Cas polypeptide or protein (e.g., a Cas endonuclease, its catalytic or non-catalytic derivative, etc.) or other protein with endonuclease activity. The term "CRISPR" refers toClustered, regularly spaced short palindromic repeat sequences and their associated systems. At least a portion of the gRNA may be complementary to at least a portion of the target region. The target region may contain a “pre-intercalation sequence” and a “pre-intercalation sequence adjacent motif” (PAM), and both domains may be essential for the nuclease activity (e.g., cleavage) of the Cas polypeptide. The pre-intercalation sequence may be referred to as the target site (or genomic target site). The gRNA may pair (or hybridize) with the opposite strand of the pre-intercalation sequence (binding site) to guide the Cas polypeptide to the target region. The PAM site generally refers to a short sequence recognized by the Cas polypeptide, and in some cases may be essential for nuclease (or nickase) activity. The nucleotide sequence and number of PAM sites may vary depending on the type of Cas enzyme.
[0124] The Cas polypeptide may contain nuclease (or nickase) activity, and the gRNA may interact with the Cas polypeptide to guide the nuclease (or nickase) activity of the Cas polypeptide to the desired target region. Alternatively, the Cas polypeptide may be non-catalytic and may not contain nuclease activity. Non-catalytic Cas polypeptides may be referred to as dead or inactivated Cas (dCas).
[0125] Cas proteins may include proteins of the CRISPR-associated type I, II or III system or proteins derived from the CRISPR-associated type I, II or III system as described on pages 18 / 69 of the specification, which may have RNA-directed polynucleotide-binding or nuclease activity. Examples of suitable Cas proteins include Cas3, Cas4, Cas5, Cas5e (or CasD), Cas6, Cas6e, Cas6f, Cas7, Cas8a1, Cas8a2, Cas8b, Cas8c, Cas9 (also known as Csn1 and Csxl2), Cas10, Cas10d, CasF, CasG, CasH, Csy1, Csy2, Csy3, Cse1 (or CasA), Cse2 (or CasB), Cse3 (or CasE), Cse4 (or CasC), Csc1, Csc2, Csa5, Csn2, Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csz1, Csx15, Csf1, Csf2, Csf3, Csf4, Cu1966, their homologs, and their modified forms (e.g., catalytic or non-catalytic). In some cases, Cas proteins may include proteins of the CRISPR-associated type V or VI system or derived from CRISPR-associated type V or VI systems.Proteins of the type VI system, such as Cpf1 (or Cas12a), C2c1 (or Cas12b), C2c2, their homologs, and their modified forms (e.g., catalytic or non-catalytic).
[0126] Although some examples herein refer to Cas proteins, other proteins with endonuclease activity may be used. For example, such other proteins may not be Cas proteins but may be configured for use with gRNA.
[0127] Cas peptides or proteins may be engineered to modify nuclease activity to nicking enzyme activity. For example, an aspartic-to-alanine substitution (D10A) in the RuvC I catalytic domain of Cas9 from *S. pyogenes* can convert Cas9 from a two-strand nuclease to a single-strand nicking enzyme, Cas9n. The Cas9n nicking enzyme mutant can introduce gRNA-targeted single-strand breaks into the DNA, rather than the double-strand breaks produced by wild-type Cas peptides. Other examples of mutations that make Cas9 a nickase may include H840A, N854A, and N863A.
[0128] As used herein, the term “guide RNA (gRNA)” generally refers to an RNA molecule (e.g., DNA or gene) that can bind to the Cas polypeptide and help target the Cas polypeptide to a specific location within a target nucleic acid region. The complementarity between the gRNA and the specific location within the target nucleic acid region may be at least 50%, 60%, 75%, 80%, 85%, 90%, 95%, 97.5%, 99%, or higher. The guide RNA may comprise a CRISPR RNA (crRNA) segment and a trans-activated crRNA (tracrRNA) segment. As may be used interchangeably herein, the terms “crRNA” and “crRNA segment” generally refer to an RNA molecule or a portion thereof that includes a polynucleotide-targeting guide sequence, a stem sequence, and optionally a 5'-overhang sequence. As used interchangeably herein, the terms “tracrRNA” and “tracrRNA segment” generally refer to an RNA molecule or a portion thereof that includes a protein-binding segment (e.g., a protein-binding segment capable of interacting with CRISPR-related proteins such as Cas9). In some cases, the guide RNA may be a single guide RNA (sgRNA) in which the crRNA segment and the tracrRNA segment are located in the same RNA molecule. The gRNA may contain one or more peptide nucleic acids.
[0129] The crRNA may contain at least 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40 or more RNA bases. The crRNA may contain up to 40, 35, 30, 25, 24, 23, 22, 21, 20, 19, 18, 17, 16, 15 or fewer RNA bases.The target nucleic acid sequence of the gRNA in the Cas system may contain at least 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40 or more DNA bases. The target nucleic acid sequence of the gRNA in the Cas system may contain up to 40, 35, 30, 25, 24, 23, 22, 21, 20, 19, 18, 17, 16, 15 or fewer DNA bases. The crRNA sequence can be selected to target any target sequence. The target sequence can be a sequence within the cell genome. The target sequence can include a sequence unique within the target genome.
[0130] As used herein, the term “polymerase” generally refers to an enzyme (e.g., natural or synthetic) capable of catalyzing polymerization reactions. Examples of polymerases may include nucleic acid polymerases (e.g., DNA polymerase or RNA polymerase), transcriptases, and ligases. A polymerase may be a polymerization enzyme. The term “DNA polymerase” generally refers to an enzyme capable of catalyzing the polymerization reaction of DNA.
[0131] As used herein, the term “linked polymerase” generally refers to a polymerase that couples to (e.g., fuses to) a linker, such as a DNA polymerase. The linker may be capable of coupling to (e.g., binding to or conjugating to) another entity (e.g., a nanopore, such as a protein nanopore or solid nanopore as described on page 19 / 69 of this specification, CN 121472373 A).
[0132] As may be used interchangeably herein, the terms “sequence variant” and “sequencing variant” generally refer to any sequence variation relative to one or more reference sequences. Generally, for a given population of individuals for which a reference sequence is provided, sequence variants occur less frequently than the reference sequence. For example, a particular genus of bacteria may have a common reference sequence for a 16S rRNA gene, but individual species within that genus may have one or more sequence variants within the gene or a portion of the gene that can be used to identify that species within a bacterial population. As a further example, when optimal alignment occurs, sequences from multiple individuals of the same species or multiple sequencing reads from the same individual can generate a shared sequence, and sequence variants of this shared sequence can be used to identify mutants in the population that indicate dangerous contamination. Typically, a "shared sequence" refers to a nucleotide sequence reflecting the most common base selections at each position in a sequence in which the relevant nucleic acid series has undergone in-depth mathematical and / or sequence analysis, such as optimal alignment according to any of a variety of sequence alignment algorithms. A reference sequence is a single known reference sequence, such as the genome sequence of a single individual. A reference sequence can be a shared sequence formed by aligning multiple known sequences, such as the genome sequences of multiple individuals used as a reference population or multiple sequencing reads of polynucleotides from the same individual.The reference sequence can be a common sequence formed by optimally aligning sequences from the analyzed sample, such that sequence variants represent variations relative to the corresponding sequences in the same sample. Sequence variants may occur at a low frequency in the population (also known as “rare” sequence variants). For example, sequence variants may occur at a frequency of less than or equal to 5%, 4%, 3%, 2%, 1.5%, 1%, 0.75%, 0.5%, 0.25%, 0.1%, 0.075%, 0.05%, 0.04%, 0.03%, 0.02%, 0.01%, 0.005%, 0.001%, or lower. Sequence variants may occur at a frequency of less than or equal to 0.1%.
[0133] A sequence variant can be any variation relative to the reference sequence. Sequence variations may include alterations, insertions, or deletions of a single nucleotide or multiple nucleotides (such as 2, 3, 4, 5, 6, 7, 8, 9, 10, or more nucleotides). When a sequence variant contains two or more nucleotide differences, the different nucleotides may be continuous or discontinuous with each other. Examples of sequence variant types include single nucleotide polymorphisms (SNPs), deletion / insertion polymorphisms (DIPs), copy number variants (CNVs), short tandem repeats (STRs), simple sequence repeats (SSRs), variable number tandem repeats (VNTRs), amplified fragment length polymorphisms (AFLPs), reverse transcriptase-based insertion polymorphisms, sequence-specific amplification polymorphisms, and epigenetic marker differences that can be detected as sequence variants (e.g., methylation differences).
[0134] As used herein, the term “sequencing” generally refers to a procedure for determining the order in which nucleotides appear in a target nucleotide sequence. Sequencing methods may include high-throughput sequencing, such as next-generation sequencing (NGS). Sequencing may be whole-genome sequencing or targeted sequencing. Sequencing may be single-molecule sequencing or massively parallel sequencing. Next-generation sequencing methods can yield millions of sequences in a single run. In one instance, sequencing may be performed using one or more nanopore sequencing methods, such as synthetic sequencing, ligation sequencing, or cleavage sequencing.
[0135] As used herein, the term "nanopore" generally refers to a pore, channel, or pathway formed or otherwise provided in a membrane. The membrane may be an organic membrane, such as a lipid bilayer, or a synthetic membrane, such as a membrane formed from polymeric materials such as protein nanopores. The membrane may be a solid membrane (e.g., a silicon substrate). The nanopore may be located adjacent to or close to electrodes of a sensing circuit (e.g., a complementary metal-oxide-semiconductor (CMOS) or field-effect transistor (FET) circuit). The nanopore may be part of the sensing circuit. The nanopore may have a characteristic width or diameter, for example, from about 0.1 nanometers (nm) to 1000 nm. The nanopore may be a biological nanopore, a solid nanopore, a hybrid biological-solid nanopore, a variant thereof, or a combination thereof.Examples of bio-nanopores include, but are not limited to, OmpG from *Escherichia coli*, *Salmonella*, *Shigella*, and *Pseudomonas*, as well as α-hemolysin from *Staphylococcus aureus*, MspA from *Mycobacterium smegmatis*, their functional variants, or combinations thereof. Sequencing may include forward sequencing and / or reverse sequencing. Examples of solid-state nanopores include, but are not limited to, silicon nitride, silicon oxide, graphene, molybdenum sulfide, their functional variants, or combinations thereof. Solid-state nanopores can be fabricated by high-energy beam fabrication, imprinting (e.g., nanoimprinting), laser ablation, chemical etching, plasma etching (e.g., oxygen plasma etching), etc.
[0136] As used herein, the term “nanopore sequencing complex” generally refers to a nanopore linked or coupled to an enzyme, such as a polymerase, which in turn associates with a polymer, such as a polynucleotide template. The nanopore sequencing complex may be located in a membrane, such as a lipid bilayer, where its function is to identify polymer components, such as nucleotides or amino acids.
[0137] As used interchangeably herein, the terms “nanopore sequencing” and “nanopore-based sequencing” generally refer to methods for determining the sequence of a polynucleotide using nanopores. In some cases, the sequence of a polynucleotide can be determined in a template-dependent manner. In some cases, the methods, systems, or compositions disclosed herein may not be limited to any particular nanopore sequencing method, system, or apparatus.
[0138] As used herein, the term “barcode” generally refers to a known nucleic acid sequence that allows the identification of certain characteristics of a polynucleotide associated with a barcode (e.g., a polynucleotide containing at least a portion of the barcode or a polynucleotide complementary to at least a portion of the barcode). In some instances, the characteristics of the polynucleotide to be identified may be those of the sample from which the polynucleotide is derived. The length of a barcode can be at least about 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 or more nucleotides. The length of a barcode can be at most 20, 19, 18, 17, 16, 15, 14, 13, 12, 11, 10, 9, 8, 7, 6, 5, 4, 3 or 2 nucleotides. A barcode associated with a polynucleotide from a first sample can be different (e.g., a different sequence and / or a different length) from a barcode associated with a polynucleotide from a second sample different from the first sample. In this case, identifying the barcode in the corresponding polynucleotide can help identify the sample source of one or more polynucleotides. Therefore, different samples with different barcodes can be analyzed together (e.g., in batches) (e.g., sequencing).Furthermore, separation is at least partially based on barcodes during analysis. In some instances, barcodes can be accurately identified even after one or more nucleotides in the barcode sequence have been mutated, inserted, or deleted (e.g., at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more nucleotides). Multiple polynucleotides from the same sample can have the same barcode. Alternatively, multiple polynucleotides from the same sample can have different barcodes. The first barcode can differ from the second barcode by at least three nucleotide positions, such as at least 3, 4, 5, 6, 7, 8, 9, 10, or more nucleotide positions. Multiple barcodes can be represented in a sample library, where each sample contains a polynucleotide that contains one or more barcodes that are different from the barcodes contained in polynucleotides derived from other samples in the library. Polynucleotide samples containing one or more barcodes can be grouped according to the barcode sequences to which they are linked, such that all four nucleotide bases A, G, C, and T are approximately uniformly represented in the library along one or more positions of each barcode (such as at positions 1, 2, 3, 4, 5, 6, 7, 8, or more, or all, of the barcode). In some instances, the method of the present invention may include identifying the sample from which the target polynucleotide is derived based on the barcode sequence linked to the target polynucleotide. The barcode may contain a nucleic acid sequence that, when linked to the target polynucleotide, can act as an identifier of the sample from which the target polynucleotide is derived. In one instance, an oligonucleotide primer (e.g., an amplification primer) may contain one or more barcodes. In another instance, a nucleic acid molecule may be coupled (e.g., linked) to an adaptor nucleic acid (e.g., for circularization), and the adaptor nucleic acid may contain one or more barcodes.
[0139] As used herein, the term “sample” generally refers to any sample that may include one or more components (e.g., nucleic acid molecules) for processing or analysis. A sample may be a biological sample. A sample may be a cell or tissue sample. The sample may be a cell-free sample, such as blood (e.g., whole blood), plasma, serum, sweat, saliva, or urine. The sample may be obtained in vivo or cultured in vitro.
[0140] As used herein, the term “subject” generally refers to the individual or entity from which the sample is derived, such as a vertebrate (e.g., a mammal, such as a human) or an invertebrate. A mammal may be a mouse, monkey, human, farm animal (e.g., a cow, sheep, pig, or chicken) or pet (e.g., a cat or dog). The subject may be a plant. The subject may be a patient. The subject may not have symptoms of a disease (e.g., cancer). Alternatively, the subject may have symptoms of a disease.
[0141] Whenever the terms “at least,” “greater than,” or “greater than or equal to” are the first of a series of two or more numerical values.When a value precedes a number, the terms “at least,” “greater than,” or “greater than or equal to” apply to each value in the series. For example, greater than or equal to 1, 2, or 3 is equivalent to greater than or equal to 1, greater than or equal to 2, or greater than or equal to 3.
[0142] Whenever the terms “no more than,” “less than,” or “less than or equal to” precede the first value in a series of two or more values, the terms “no more than,” “less than,” or “less than or equal to” apply to each value in the series. For example, less than or equal to 3, 2, or 1 is equivalent to less than or equal to 3, less than or equal to 2, or less than or equal to 1.
[0143] Overview
[0144] Currently available sequencing methods can be used to sequence one or more nucleic acids. However, such methods can be expensive and may not provide sequence information within the time frame necessary for diagnosing or treating a subject (e.g., an individual, a patient, etc.) and within the range (or level) of accuracy.
[0145] Massive parallel nucleic acid sequencing can be used to identify one or more sequence variations within a population (e.g., a complex population). However, massively parallel sequencing using currently available sequencing technologies may be limited by its error frequency, which may be greater than the frequency of actual sequence variants in the population. In one instance, currently available high-throughput sequencing methods may exhibit an error rate of approximately 0.1–1 percent (%). In some cases, the detection of a sequence variant may have a high false positive rate when the frequency of one or more sequence variants (e.g., one or more rare sequence variants) is low, for example, when the frequency is equal to or lower than the error rate.
[0146] Methods and Compositions for Sequencing
[0147] Adaptor Coupling for Sequencing
[0148] In one aspect, this disclosure provides methods for processing or analyzing double-stranded nucleic acid molecules. The methods may include providing (i) a double-stranded nucleic acid molecule and (ii) a double-stranded adaptor having a cleavage site within its sense or antisense strand. The methods may include coupling the double-stranded adaptor to the double-stranded nucleic acid molecule. The methods may include circularizing the double-stranded nucleic acid molecule coupled with the double-stranded adaptor to produce a circularized double-stranded nucleic acid molecule. Before coupling between the double-stranded adaptor and the double-stranded nucleic acid molecule, the cleavage site may contain a cleavage. The cleavage may be a break in the sense strand of the double-stranded adaptor. Alternatively, the cleavage may be a break in the antisense strand of the double-stranded adaptor.
[0149] The length of the double-stranded nucleic acid molecule may be at least 2, 3, 4, 5, 10, 50, 100, 500, 1,000, 5,000, 10,000, 50,000, 100,000 or more nucleotides. The length of the double-stranded nucleic acid molecule may be at most 100,000, 50,000, 10,000, 5,000, 1,000, 500, 100, 50, 10 or fewer nucleotides. The length of the double-stranded adaptor may be at least 3, 4, 5, ...6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 45, 50 or more nucleotides. The length of the double-stranded adaptor can be up to 50, 45, 40, 35, 30, 25, 20, 19, 18, 17, 16, 15, 14, 13, 12, 11, 10, 9, 8, 7, 6, 5, 4, 3 or fewer nucleotides.
[0150] The double-stranded nucleic acid molecule and the double-stranded adaptor can be heterologous to each other. The double-stranded nucleic acid molecule and the double-stranded adaptor can be different parts of the same gene or different genes. One of the double-stranded nucleic acid molecule and the double-stranded adaptor can be derived from one species, while the other of the double-stranded nucleic acid molecule and the double-stranded adaptor can be from a different species. One of the double-stranded nucleic acid molecule and the double-stranded adaptor may be natural, while the other may be synthetic (e.g., the double-stranded adaptor may be a synthetic molecule). Alternatively, both the double-stranded nucleic acid molecule and the double-stranded adaptor may be synthetic.
[0151] In one example, the double-stranded nucleic acid molecule may be a fragmented double-stranded nucleic acid molecule from a genomic sample of a subject (e.g., cells from subject specification page 22 / 69, 24 CN 121472373 A), and the double-stranded adaptor may be a synthetic adaptor including a nick site. In another example, the double-stranded nucleic acid molecule may be a cell-free double-stranded nucleic acid molecule from a cell-free biological sample of a subject (e.g., blood, plasma, urine, etc.), and the double-stranded adaptor may be a synthetic adaptor including a nick site.
[0152] The double-stranded nucleic acid molecule and the double-stranded adaptor may be provided in the form of a cell-free composition. The cell-free composition may be substantially free of intact cells. The cell-free composition may contain cell lysates or extracts. Cell lysates may contain a fluid containing one or more contents of lysed cells. Cell lysates can be crude (i.e., unpurified) or at least partially purified (e.g., to remove cell debris or particles, such as damaged outer cell membranes). Methods for forming cell lysates may include sonication, homogenization, enzymatic lysis using lysozyme, freezing, grinding, and high-pressure lysis. Alternatively, cell-free compositions may be derived from cell-free biological samples.
[0153] The coupling of a double-stranded inverse to a double-stranded nucleic acid molecule can be performed under cell-free conditions, or the cyclization of a double-stranded nucleic acid molecule coupled with a double-stranded inverse to a cyclized double-stranded nucleic acid molecule can be performed under cell-free conditions. The coupling of a double-stranded inverse to a double-stranded nucleic acid molecule can be performed under cell-free conditions, and the cyclization of a double-stranded nucleic acid molecule coupled with a double-stranded inverse to a cyclized double-stranded nucleic acid molecule can be performed under cell-free conditions. Alternatively, in the presence of one or more cells, (i) the coupling of a double-stranded inverse to a double-stranded nucleic acid molecule can be performed, or (ii) the coupling of a double-stranded inverse to a double-stranded nucleic acid molecule can be performed.Double-stranded nucleic acid molecules are circularized into circularized double-stranded nucleic acid molecules. In one instance, one or more cells may be configured to express one or more enzymes capable of performing the processes in (i) and / or (ii).
[0154] Coupling may include coupling the sense strand of a double-stranded adaptor to the sense strand of a double-stranded nucleic acid molecule. Alternatively, coupling may include coupling the antisense strand of a double-stranded adaptor to the antisense strand of a double-stranded nucleic acid molecule. In various alternatives, coupling may include (i) coupling the sense strand of a double-stranded adaptor to the sense strand of a double-stranded nucleic acid molecule, and (ii) coupling the antisense strand of a double-stranded adaptor to the antisense strand of a double-stranded nucleic acid molecule. Coupling of two nucleic acid molecules may include ligation (e.g., by an enzyme, such as a ligase), hybridization (e.g., in the absence of an enzyme), or both.
[0155] The cleavage site may be part of the sense strand of the circularized double-stranded nucleic acid molecule. The cleavage site may be neither the 5′ end nor the 3′ end of the sense strand. The cleavage site may be at least 1, 2, 3, 4, 5, 10, 15, 20, 25, 30 or more nucleotides away from the 5' or 3' end of the sense strand. The cleavage site may be at most 30, 25, 20, 15, 10, 5, 4, 3, 2 or 1 nucleotide away from the 5' or 3' end of the sense strand. Alternatively or additionally, the cleavage site may be part of the antisense strand of a circularized double-stranded nucleic acid molecule. The cleavage site may not be the 5' or 3' end of the antisense strand. The cleavage site may be at least 1, 2, 3, 4, 5, 10, 15, 20, 25, 30 or more nucleotides away from the 5' or 3' end of the antisense strand. The cleavage site may be at most 30, 25, 20, 15, 10, 5, 4, 3, 2 or 1 nucleotide away from the 5' or 3' end of the antisense strand.
[0156] The method may further include sequencing double-stranded nucleic acid molecules from the nick site of a double-stranded connective. Sequencing may be whole-genome sequencing (or whole-genome sequencing) or targeted sequencing. Sequencing may include one or more NGS methods. Sequencing may include nanopore-based sequencing. The nanopore may be a protein nanopore (e.g., α-hemolysin) or a solid nanopore. Alternatively, the nanopore may be a hybrid nanopore comprising at least a portion of a protein nanopore (e.g., α-hemolysin) and at least a portion of a solid nanopore. Nanopore-based sequencing may utilize at least one enzyme (e.g., a polymerase or nuclease) to interact with at least one double-stranded nucleic acid molecule. At least one enzyme may be coupled to the nanopore. At least one enzyme may be fused, conjugated, or bound to a protein nanopore or a membrane containing a nanopore. At least one enzyme may be conjugated or bound to a solid nanopore or a membrane containing a solid nanopore. In some instances, at least one enzyme may have a binding portion capable of binding to a nanopore (or solid nanopore) or a membrane.
[0157] Sequencing may include (i) extending a double-stranded nucleic acid molecule from a nick site on a double-stranded adaptor to produce a product similar to...(ii) A growth chain having at least a portion of the strand of a double-stranded nucleic acid molecule that is sequence complementary, and obtaining at least one portion of the sequence information of the growth chain as described on page 23 / 69 of CN 121472373 A. The strand of the double-stranded nucleic acid molecule may be its sense strand or its antisense strand. Thus, the growth chain may exhibit complementarity with at least a portion of the sense strand or at least a portion of the antisense strand. Alternatively, the strand of the double-stranded nucleic acid molecule may be a sense strand and an antisense strand. Thus, a first growth chain may exhibit complementarity with at least a portion of the sense strand, while a second growth chain may exhibit complementarity with at least a portion of the antisense strand.
[0158] The extension reaction may include amplifying at least a portion of the double-stranded nucleic acid molecule (e.g., a double-stranded nucleic acid molecule and at least a portion of a double-stranded adaptor). Amplification can produce multiple copies (e.g., at least 1, 2, 3, 4, 5, 10, 15, 20 or more copies; up to 20, 15, 10, 5, 4, 3, 2 or 1 copy) of at least a portion of the sense and / or antisense strand of a double-stranded nucleic acid molecule. In some instances, the double-stranded nucleic acid molecule may be a portion of a circular (or circulated) double-stranded nucleic acid molecule, and the extension reaction may be RCA. RCA can generate a growth chain containing one or more copies (e.g., at least 1, 2, 3, 4, 5, 10, 15, 20 or more copies; up to 20, 15, 10, 5, 4, 3, 2 or 1 copy) of at least a portion of the sense or antisense strand of the circular double-stranded nucleic acid molecule.
[0159] Obtaining sequence information may include detecting at least a portion of the growth chain. The extension reaction may include contacting the double-stranded nucleic acid molecule with a nucleotide coupled to a tag under conditions sufficient to incorporate nucleotides into the growth chain. In this case, obtaining sequence information may include detecting the tag. In some cases, the method may also include releasing a tag from the nucleotide as it is incorporated into the growing chain, and detecting the released tag for sequencing.
[0160] The extension reaction may be performed using oligonucleotide primers. Alternatively, the extension reaction may be performed without the use of oligonucleotide primers. The cleavage within the double-stranded adapter may serve as a binding site for an enzyme (e.g., polymerase) capable of performing the extension reaction, and therefore may not require any oligonucleotide primers.
[0161] Alternatively, sequencing may include (i) cleaving the double-stranded nucleic acid molecule from the cleavage site of the double-stranded adapter to cleave at least a portion of the strand of the double-stranded nucleic acid molecule, and (ii) obtaining sequence information of at least a portion of the strand. The strand of the double-stranded nucleic acid molecule to be cleaved may be its sense strand or its antisense strand. Alternatively, the strand of the double-stranded nucleic acid molecule to be cleaved may be both a sense strand and an antisense strand. Subsequently, obtaining sequence information may include detecting at least a portion of the strand.
[0162] At least a portion of the double-stranded nucleic acid molecule may have or may be suspected of having a [missing information] compared to at least one reference sequence.One or more sequencing variants (e.g., one or more mutations). Therefore, sequencing can be performed to identify the presence of at least a portion of the double-stranded nucleic acid molecule. One or more sequencing variants can indicate a mutation in the gene. At least one reference sequence can contain a common sequence of at least a portion of the gene. The common sequence and the double-stranded nucleic acid molecule can be derived from the same species or different species. The common sequence can be a representative sequence from a collection of multiple sequences of a gene obtained from multiple samples of the same species (e.g., at least 2, 3, 4, 5, 10, 15, 20, 30, 40, 50 or more samples; up to 50, 40, 30, 20, 15, 10, 4, 3 or 2 samples). In one instance, both the double-stranded nucleic acid molecule and at least one reference sequence can be derived from a human sample, and at least one reference sequence can be a common sequence of at least a portion of the human gene of interest, such as a portion of a gene known to generally not have any mutations.
[0163] The method may include amplifying the double-stranded nucleic acid molecule to produce multiple copies of the double-stranded nucleic acid molecule before coupling a double-stranded adaptor to it. Amplification can produce at least 1, 2, 3, 4, 5, 10, 15, 20, 30, 40, 50, 100, 150, 200 or more copies of double-stranded nucleic acid molecules. Amplification can produce up to 200, 150, 100, 50, 40, 30, 20, 15, 10, 5, 4, 3, 2 or 1 copy of double-stranded nucleic acid molecules. Thus, the method can utilize multiple double-stranded linkers, for example, the same number or more as the copy number of the double-stranded nucleic acid molecules.
[0164] The double-stranded nucleic acid molecule may contain a recognition sequence. The recognition sequence for the double-stranded nucleic acid molecule may be endogenous or exogenous. The recognition sequence may contain at least one natural nucleotide, at least one synthetic nucleotide or both. The identification sequence may include at least 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 40, 50, 100 (see specification 24 / 69, page 26, CN 121472373 A). The identification sequence may also include up to 100, 50, 40, 30, 25, 24, 23, 22, 21, 20, 19, 18, 17, 16, 15, 14, 13, 12, 11, 10, 9, 8, 7, 6, 5, or fewer nucleotides. The method may also include enriching double-stranded nucleic acid molecules from a random nucleic acid library, at least partially based on the identification sequence. The library may contain at least 2, 3, 4, 5, 10, 50, 100, 500, 1,000, 5,000, 10,000, 50,000, 100,000, 500,000, 1,000,000 or more random nucleic acid molecules.It can contain up to 1,000,000, 500,000, 100,000, 50,000, 10,000, 5,000, 1,000, 500, 100, 50, 10, 5, 4, 3, or 2 random nucleic acid molecules. Enrichment may involve isolating double-stranded nucleic acid molecules (containing the recognition sequence) from at least one different nucleic acid molecule that does not contain the recognition sequence. Double-stranded nucleic acid molecules can be isolated from at least 1, 5, 10, 50, 100, 500, 1,000, 5,000, 10,000, 50,000, 100,000, 500,000, 1,000,000, or more different nucleic acid molecules that do not contain the recognition sequence. Double-stranded nucleic acid molecules can be isolated from up to 1,000,000, 500,000, 100,000, 50,000, 10,000, 5,000, 1,000, 500, 100, 50, 10, 5, or 1 different nucleic acid molecules that do not contain a recognition sequence.
[0165] Before amplification, double-stranded nucleic acid molecules in a random nucleic acid molecule library (one of which is a double-stranded nucleic acid molecule containing a recognition sequence) can be enriched. Alternatively or additionally, the random nucleic acid molecule library can be amplified before enriching the double-stranded nucleic acid molecules.
[0166] Double-stranded nucleic acid molecules in a random nucleic acid molecule library (one of which is a double-stranded nucleic acid molecule containing a recognition sequence) can be enriched before (i) at least coupling with a double-stranded adapter and (ii) circularization. Alternatively, the random nucleic acid molecule library containing double-stranded nucleic acid molecules can be coupled with a double-stranded adapter and then circularized. In various alternatives, a random nucleic acid library containing double-stranded nucleic acid molecules may be (i) coupled with at least a double-stranded linker and (ii) circularized before enriching with circularized double-stranded nucleic acid molecules (and any excess of linear double-stranded nucleic acid molecules) containing the recognition sequence.
[0167] Enrichment may include generating a library of selected double-stranded nucleic acid molecules (e.g., linear or circular). Each double-stranded nucleic acid molecule in at least 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 60%, 70%, 80%, 90%, 95% or more of the selected library may contain at least a portion of the recognition sequence. At most 100%, 95%, 90%, 80%, 70%, 60%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 15%, 10%, 5% or less of the selected library may contain at least a portion of the recognition sequence in each double-stranded nucleic acid molecule.
[0168] The probability of the recognition sequence appearing without any mismatch may be in the form of 1x10⁴, 5x10⁴, 7x10⁴, 1x10⁵, 5x10⁵, 1x10⁶, 5x10⁶, 1x10⁷, 5x10⁷, 1x10⁸, 5x10⁸, 1x10⁹, 5x10⁹, 1x10¹⁰, 5x10¹⁰, 1x10¹¹, ...At most once in every 5x10¹¹, 1x10¹², 5x10¹² or more nucleotides. Alternatively, the probability of the recognition sequence appearing without any mismatches may be at most two, three, four, five or more in every 1x10⁴, 5x10⁴, 7x10⁴, 1x10⁵, 5x10⁵, 1x10⁶, 5x10⁶, 1x10⁷, 5x10⁷, 1x10⁸, 5x10⁸, 1x10⁹, 5x10⁹, 1x10¹⁰, 5x10¹⁰, 1x10¹¹, 5x10¹¹, 1x10¹², 5x10¹² or more nucleotides.
[0169] The recognition sequence may contain at least 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 35, 40, 45, 50 or more nucleotides. The recognition sequence may contain up to 50, 45, 40, 35, 30, 29, 28, 27, 26, 25, 24, 23, 22, 21, 20, 19, 18, 17, 16, 15, 14, 13, 12, 11, 10, 9, 8, 7, 6, 5 or fewer nucleotides.
[0170] Enrichment may include (i) binding a recognition portion that is complementary to the recognition sequence to a double-stranded nucleic acid molecule to form a recognition complex, and (ii) extracting the recognition complex from a random nucleic acid molecule library. The recognition portion may have at least 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95% or more complementarity with the recognition sequence. The recognition portion may have at most 100%, 95%, 90%, 80%, 70%, 60%, 50%, 40%, 30% or less complementarity with the recognition sequence.
[0171] One or more double-stranded nucleic acid molecules containing the recognition sequence may be enriched in a random nucleic acid molecule library before (i) providing one or more double-stranded nucleic acid molecules and one or more double-stranded adapters having nick sites, or (ii) after coupling at least one double-stranded adapter to each of a plurality of double-stranded nucleic acid molecules from the library, as specified on page 25 / 69 of CN 121472373 A. Alternatively, one or more double-stranded nucleic acid molecules containing a recognition sequence may be enriched in a random nucleic acid molecule library before (i) providing one or more double-stranded nucleic acid molecules and one or more double-stranded adapters having nick sites, and (ii) after coupling at least one double-stranded adapter to each of a plurality of double-stranded nucleic acid molecules from the library.
[0172] Alternatively or additionally, the double-stranded adapter may contain at least one recognition sequence. After the double-stranded adapter is coupled (e.g., linked) to the double-stranded nucleic acid molecule, at least one recognition sequence (e.g., at least one provided herein) may be used.The recognition sequence of the double-stranded nucleic acid molecule is used to enrich the double-stranded nucleic acid molecule coupled with the double-stranded adapter. Enrichment can deplete one or more double-stranded nucleic acid molecules not coupled with the double-stranded adapter.
[0173] The recognition sequence of the double-stranded adapter may contain at least one natural nucleotide, at least one synthetic nucleotide, or both. The recognition sequence of the double-stranded adapter may contain at least 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 40, 50, 100 or more nucleotides. The recognition sequence of the double-stranded adapter may contain up to 100, 50, 40, 30, 25, 24, 23, 22, 21, 20, 19, 18, 17, 16, 15, 14, 13, 12, 11, 10, 9, 8, 7, 6, 5 or fewer nucleotides. The method may further include enriching double-stranded nucleic acid molecules coupled to double-stranded linkers from a random nucleic acid molecule library, at least in part, based on the recognition sequence of the double-stranded linker. The library may contain at least 2, 3, 4, 5, 10, 50, 100, 500, 1,000, 5,000, 10,000, 50,000, 100,000, 500,000, 1,000,000, or more random nucleic acid molecules. The library may also contain up to 1,000,000, 500,000, 100,000, 50,000, 10,000, 5,000, 1,000, 500, 100, 50, 10, 5, 4, 3, or 2 random nucleic acid molecules. Enrichment may involve isolating double-stranded nucleic acid molecules (coupled to double-stranded linkers containing the recognition sequence) from at least one different nucleic acid molecule that does not contain the recognition sequence. Double-stranded nucleic acid molecules can be isolated from at least 1, 5, 10, 50, 100, 500, 1,000, 5,000, 10,000, 50,000, 100,000, 500,000, 1,000,000 or more different nucleic acid molecules that do not contain a recognition sequence. Double-stranded nucleic acid molecules can be isolated from at most 1,000,000, 500,000, 100,000, 50,000, 10,000, 5,000, 1,000, 500, 100, 50, 10, 5 or 1 different nucleic acid molecules that do not contain a recognition sequence.
[0174] Double-stranded nucleic acid molecules can be derived from or derived from a biological sample of the subject. Biological samples can include cell-free biological samples of the subject. Cell-free biological samples may be selected from: blood, plasma, serum, urine, lymph, feces, saliva, semen, amniotic fluid, cerebrospinal fluid, bile, sweat, tears, sputum, synovial fluid, vomit, and combinations thereof. Double-stranded nucleic acid molecules may be derived from or derived from cell-free nucleic acid molecules from cell-free biological samples. Cell-free nucleic acid molecules may include circulating tumor nucleic acid molecules (e.g., ctDNA) or amniotic fluid nucleic acid molecules.
[0175] Biological samples may include tissue samples from a subject. Tissue samples may be derived from: bone, heart, thymus, arteries, blood vessels, lungs, muscles, stomach, intestines, liver, pancreas, spleen, kidneys, gallbladder, thyroid gland, adrenal glands, breast, ovaries, prostate, testes, skin, fat, eyes, brain, and combinations thereof. Double-stranded nucleic acid molecules may be derived from or derived from genomic nucleic acid molecules from tissue samples. Tissue samples may be derived from: infected tissue, diseased tissue, malignant tissue, calcified tissue, healthy tissue, and combinations thereof. Tissue samples may be derived from malignant tissue containing tumors, sarcomas, leukemia, or derivatives thereof.
[0176] Double-stranded nucleic acid molecules may include DNA, cDNA, ctDNA, derivatives thereof, or combinations thereof. Double-stranded nucleic acid molecules include RNA.
[0177] In another aspect, this disclosure provides reaction mixtures for processing or analyzing double-stranded nucleic acid molecules. The reaction mixture may comprise a composition comprising (i) a double-stranded nucleic acid molecule and (ii) a double-stranded linker having a cleavage site within its sense or antisense strand. The reaction mixture may contain at least one enzyme that (i) couples a double-stranded adaptor to a double-stranded nucleic acid molecule, or (ii) cyclizes the double-stranded nucleic acid molecule coupled to the double-stranded adaptor to produce a cyclized double-stranded nucleic acid molecule. A nick site may contain a nick before coupling between the double-stranded adaptor and the double-stranded nucleic acid molecule. The nick may be a break in the sense strand of the double-stranded adaptor. Alternatively, the nick may be a break in the antisense strand of the double-stranded adaptor. As provided in this disclosure, the reaction mixture may be used or identified in any of the subject methods for adaptor ligation. The reaction mixture may be used to prepare one or more libraries (e.g., libraries of nucleic acid molecules, enzymes, or combinations thereof) or compositions for one or more sequencing methods. One or more components of the reaction mixture may be used simultaneously in the same reaction. In one example, the reaction may be carried out in a single reaction vial (e.g., a reaction tube), thereby reducing purification steps and / or sample loss, and / or allowing sequencing with a small amount of nucleic acid sample input. One or more components of the reaction mixture may be used separately in different reactions.
[0178] The reaction mixture may contain at least 1, 2, 3, 4, 5, 10, 50, 100, 500, 1,000, 5,000, 10,000, 50,000, 100,000, 500,000, 1,000,000 or more double-stranded nucleic acid molecules. The reaction mixture may contain up to 1,000,000, 500,000, 100,000, 50,000, 10,000, 5,000, 1,000, 500, 100, 50, 10, 5, 4, 3 or 1 double-stranded nucleus.Acid molecules. The reaction mixture may contain at least 1, 2, 3, 4, 5, 10, 50, 100, 500, 1,000, 5,000, 10,000, 50,000, 100,000, 500,000, 1,000,000 or more double-stranded promoters. The reaction mixture may contain at most 1,000,000, 500,000, 100,000, 50,000, 10,000, 5,000, 1,000, 500, 100, 50, 10, 5, 4, 3 or 1 double-stranded promoter.
[0179] At least one enzyme may be able to (i) couple a double-stranded promoter to a double-stranded nucleic acid molecule, and (ii) cyclize the double-stranded nucleic acid molecule coupled with the double-stranded promoter to produce a cyclized double-stranded nucleic acid molecule. The reaction mixture may contain at least 0.01, 0.1, 1, 10, 100, 1,000, 10,000 or more units of at least one enzyme (e.g., Weiss units or Modrich-Lehman units). The reaction mixture may contain at most 10,000, 1,000, 100, 10, 1, 0.1, 0.01 or fewer units of at least one enzyme. Alternatively, the reaction mixture may contain (i) an enzyme that couples a double-stranded adaptor to a double-stranded nucleic acid molecule, and (ii) an additional enzyme that cyclizes the double-stranded nucleic acid molecule coupled with the double-stranded adaptor to produce a cyclized double-stranded nucleic acid molecule. The enzyme and the additional enzyme may be different. The reaction mixture may contain at least 0.01, 0.1, 1, 10, 100, 1,000, 10,000 or more units of an enzyme, and at least 0.01, 0.1, 1, 10, 100, 1,000, 10,000 or more units of an additional enzyme. The reaction mixture may contain up to 10,000, 1,000, 100, 10, 1, 0.1, 0.01 or fewer units of enzyme, and up to 10,000, 1,000, 100, 1, 0.1, 0.01 or fewer units of additional enzyme.
[0180] The reaction mixture may be a cell-free reaction mixture. The cell-free reaction mixture may be substantially free of intact cells. The cell-free reaction mixture contains cell lysates or extracts. Cell lysates may contain a fluid containing one or more contents of lysed cells. Cell lysates may be crude (i.e., unpurified) or at least partially purified (e.g., to remove cell debris or particles, such as damaged outer cell membranes). The cell-free reaction mixture may be the product of sonication, homogenization, enzymatic lysis using lysozyme, freezing, grinding, and autoclaving of one or more cells. Alternatively, the cell-free reaction mixture may be derived from a cell-free biological sample.
[0181] At least one enzyme may be able to (i) couple the sense strand of a double-stranded adaptor to the sense strand of a double-stranded nucleic acid molecule, or (ii) couple the antisense strand of a double-stranded adaptor to the antisense strand of a double-stranded nucleic acid molecule. At least one enzyme may be able to (i) couple the sense strand of a double-stranded adaptor to the antisense strand of a double-stranded nucleic acid molecule.(ii) The sense strand of the double-stranded adaptor is coupled to the sense strand of the double-stranded nucleic acid molecule, and (iii) the antisense strand of the double-stranded adaptor is coupled to the antisense strand of the double-stranded nucleic acid molecule. Alternatively, (i) a first enzyme may be able to couple the sense strand of the double-stranded adaptor to the sense strand of the double-stranded nucleic acid molecule, and (ii) a second enzyme, different from the first enzyme, may be able to couple the antisense strand of the double-stranded adaptor to the antisense strand of the double-stranded nucleic acid molecule. At least one enzyme may include a ligase, a recombinase, a polymerase, a functional variant thereof, or a combination thereof. In some instances, at least one enzyme may be able to link the double-stranded adaptor to the double-stranded nucleic acid molecule.
[0182] The reaction mixture also includes at least a second enzyme that performs an extension reaction to produce a growth chain that is sequence complementary to at least a portion of the strand of the double-stranded nucleic acid molecule. The at least second enzyme may produce the growth chain before the double-stranded adaptor is coupled to the double-stranded nucleic acid molecule. In some instances, at least the second enzyme can generate multiple growth chains to amplify double-stranded nucleic acid molecules. (Pages 27 / 69, CN 121472373 A) The amplified copies of the double-stranded nucleic acid molecule can be used in the same reaction mixture. Alternatively, the amplified copies of the double-stranded nucleic acid molecule can be split into multiple reaction samples for processing under the same or different reaction conditions. At least the second enzyme can include a polymerase. At least the second enzyme can include polymerases and recombinases, for example, recombinase polymerase amplification, which can be used for single-tube isothermal amplification instead of polymerase chain reaction (PCR) amplification.
[0183] After the double-stranded adapter is coupled to the double-stranded nucleic acid molecule, at least the second enzyme may be able to perform an extension reaction from the nick site (or nick of the nick site) of the double-stranded adapter to generate growth chains. The reaction mixture may also contain at least one nucleotide coupled to a tag, wherein at least the second enzyme incorporates the nucleotide into the growth chain. The tag can be a small molecule, nucleotide, polynucleotide, amino acid, polypeptide, polymer, metal and / or ceramic particle, etc. In some instances, each of the nucleotides G, C, A, T, and U may contain distinct tags that are distinguishable from each other. In some instances, the tags may not be released from the nucleotides when they are incorporated into the growth chain. Alternatively, at least a second enzyme may be able to release the tags from the nucleotides when they are incorporated into the growth chain. At least the second enzyme may perform the extension reaction with the aid of at least one oligonucleotide primer. Alternatively, at least the second enzyme may perform the extension reaction without the use of an oligonucleotide primer. In one instance, the extension reaction may include RCA, wherein at least the second enzyme includes a polymerase. The polymerase may bind to a nick at the nick site for use in RCA.
[0184] The reaction mixture may also include at least a third enzyme that performs a cleavage reaction from the nick site of the double-stranded linker to cleave at least a portion of the strand of the double-stranded nucleic acid molecule. Starting from the nick at the nick site, at least the third enzyme mayThe substitution and cleavage involves (i) at least a portion of the strand of a double-stranded adaptor containing the cleavage site and (ii) at least a portion of the strand of a double-stranded nucleic acid molecule coupled to the strand of the double-stranded adaptor. At least a third enzyme may include a nuclease (e.g., an endonuclease, such as a restriction endonuclease).
[0185] As provided in this disclosure, at least a portion of the double-stranded nucleic acid molecule may have or be suspected of having one or more variants compared to at least one reference sequence. The reaction mixture can be used to prepare at least one composition for sequencing to identify the presence of at least a portion of the double-stranded nucleic acid molecule.
[0186] As provided in this disclosure, the double-stranded nucleic acid molecule may contain a recognition sequence. Therefore, the reaction mixture may also contain a recognition portion associated with the recognition sequence to enrich the double-stranded nucleic acid molecule from a random library of nucleic acid molecules in the composition, at least partially based on the recognition sequence. The recognition portion may contain at least one oligonucleotide (e.g., at least a portion of a gRNA for Cas system variants) that is complementary to at least the recognition sequence. The oligonucleotide of the recognition portion may have at least 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95% or more complementarity with the recognition sequence. The oligonucleotide of the recognition portion may have at most 100%, 95%, 90%, 80%, 70%, 60%, 50%, 40%, 30% or less complementarity with the recognition sequence.
[0187] The composition of the reaction mixture may comprise a library of selected double-stranded nucleic acid molecules (e.g., linear or circular). Each double-stranded nucleic acid molecule in at least 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 60%, 70%, 80%, 90%, 95% or more of the selected library may contain at least a portion of the recognition sequence. Each double-stranded nucleic acid molecule in up to 100%, 95%, 90%, 80%, 70%, 60%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 15%, 10%, 5%, or less of the selected library may contain at least a portion of the recognition sequence.
[0188] The probability of the recognition sequence appearing without any mismatch may be at most once in every 1x10⁴, 5x10⁴, 7x10⁴, 1x10⁵, 5x10⁵, 1x10⁶, 5x10⁶, 1x10⁷, 5x10⁷, 1x10⁸, 5x10⁸, 1x10⁹, 5x10⁹, 1x10¹⁰, 5x10¹⁰, 1x10¹¹, 5x10¹¹, 1x10¹², 5x10¹², or more nucleotides (or base pairs). Alternatively, the probability of a recognized sequence occurring without any mismatches could be in the range of 1x10⁴, 5x10⁴, 7x10⁴, 1x10⁵, 5x10⁵, 1x10⁶, 5x10⁶, ...1x107, 5x107, 1x108, 5x108, 1x109, 5x109, 1x1010, 5x1010, 1x1011, 5x1011, 1x1012, 5x1012 or more of the specified nucleotides (or base pairs) at most two, three, four, five or more.
[0189] The identification sequence may contain at least 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 35, 40, 45, 50 or more nucleotides (or base pairs). The recognition sequence may contain up to 50, 45, 40, 35, 30, 29, 28, 27, 26, 25, 24, 23, 22, 21, 20, 19, 18, 17, 16, 15, 14, 13, 12, 11, 10, 9 or fewer nucleotides (or base pairs).
[0190] In various aspects, this disclosure provides libraries of circularized double-stranded nucleic acid molecules. The library may contain (i) a double-stranded nucleic acid domain coupled to (ii) a double-stranded connective domain containing a cleavage site within its sense or antisense strand. Each circularized double-stranded nucleic acid molecule in at least 5% of the library may contain a recognition sequence. The library of circularized double-stranded nucleic acid molecules may be the starting material or product of any of the subject methods or reaction mixtures provided in this disclosure.
[0191] A cleavage site may be present within the double-stranded adaptor domain before coupling the double-stranded adaptor domain to the double-stranded nucleic acid domain. The cleavage site may contain a cleavage before coupling between the double-stranded adaptor and the double-stranded nucleic acid molecule.
[0192] The double-stranded nucleic acid domain and the double-stranded adaptor domain may be heterologous to each other. Alternatively, the double-stranded nucleic acid domain and the double-stranded adaptor domain may not be heterologous to each other. The library may be in a cell-free composition. Alternatively, the library may not be in a cell-free composition.
[0193] At least a portion of the circularized double-stranded nucleic acid domain in the library may have or be suspected of having one or more sequencing variants compared to at least one reference sequence. One or more sequencing variants may indicate mutations in the gene. At least one reference sequence may contain a common sequence of at least a portion of the gene.
[0194] Each circularized double-stranded nucleic acid molecule in at least 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 60%, 70%, 80%, 90%, 95% or more of the selected library may contain at least a portion of the recognition sequence. At most 100%, 95%, 90%, 80%, 70%, 60%, 50%, 45%, 40%, 35% of the selected library may contain at least a portion of the recognition sequence.Each 30%, 25%, 20%, 15%, 10%, 5% or less of the circularized double-stranded nucleic acid molecule may contain at least a portion of the recognition sequence.
[0195] The probability of the recognition sequence appearing without any mismatch may be at most once in every 1x10⁴, 5x10⁴, 7x10⁴, 1x10⁵, 5x10⁵, 1x10⁶, 5x10⁶, 1x10⁷, 5x10⁷, 1x10⁸, 5x10⁸, 1x10⁹, 5x10⁹, 1x10¹⁰, 5x10¹⁰, 1x10¹¹, 5x10¹¹, 1x10¹², 5x10¹² or more nucleotides (or base pairs). Alternatively, the probability of a recognized sequence appearing without any mismatches may be up to two, three, four, five, or more times in every 1x10⁴, 5x10⁴, 7x10⁴, 1x10⁵, 5x10⁵, 1x10⁶, 5x10⁶, 1x10⁷, 5x10⁷, 1x10⁸, 5x10⁸, 1x10⁹, 5x10⁹, 1x10¹⁰, 5x10¹⁰, 1x10¹¹, 5x10¹¹, 1x10¹², 5x10¹², or more nucleotides (or base pairs).
[0196] The recognition sequence may contain at least 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 35, 40, 45, 50 or more nucleotides (or base pairs). The recognition sequence may contain up to 50, 45, 40, 35, 30, 29, 28, 27, 26, 25, 24, 23, 22, 21, 20, 19, 18, 17, 16, 15, 14, 13, 12, 11, 10, 9 or fewer nucleotides (or base pairs).
[0197] The cleavage site may be part of the sense strand of a circularized double-stranded nucleic acid molecule. Alternatively or additionally, the cleavage site may be part of the antisense strand of a circularized double-stranded nucleic acid molecule.
[0198] The circularized double-stranded nucleic acid molecule may be derived from or derived from a biological sample of the subject. The biological sample may include a cell-free biological sample of the subject. The circularized double-stranded nucleic acid molecule may be derived from or derived from a cell-free nucleic acid molecule derived from a cell-free biological sample. The cell-free nucleic acid molecule may include circulating tumor nucleic acid (e.g., ctDNA) molecules or amniotic fluid nucleic acid molecules. Alternatively, the biological sample may include a tissue sample of the subject. The circularized double-stranded nucleic acid molecule may be derived from or derived from genomic nucleic acid molecules derived from a tissue sample, as described on page 29 / 69 of the specification, CN 121472373 A. The tissue sample may be derived from: infected tissue, diseased tissue, malignant tissue, calcified tissue, healthy tissue, and combinations thereof. The tissue sample may be derived from malignant tissue comprising tumors, sarcomas, leukemia, or derivatives thereof.
[0199] Circulated double-stranded nucleic acid molecules may include DNA, cDNA, ctDNA, derivatives thereof, or combinations thereof. Circulated double-stranded nucleic acid molecules include RNA.
[0200] Figure 1A schematically illustrates an example method for providing a circular nucleic acid with a nick. A first double-stranded nucleic acid molecule 110 and a second double-stranded nucleic acid molecule 120 may be provided. The nucleic acid molecule may be at least a portion of a biological sample or derived from a biological sample. Nucleic acid molecule 120 may be an adaptor (e.g., an adaptor for circularization, polymerization, nuclease activity, etc.). The first double-stranded nucleic acid molecule 110 may contain a recognition sequence 112 and a target site 114. The target site 114 may have or be suspected of having one or more mutations compared to at least a reference sequence (e.g., a shared sequence of a portion of the genes of a species or multiple species). The second double-stranded nucleic acid molecule 112 may include a nick 122 within the nucleic acid molecule 112. The nick 122 may not be located directly at the 5′ or 3′ end of the (i) sense strand or (ii) antisense strand of the nucleic acid molecule 112. Alternatively, nucleic acid molecule 112 may be characterized by the removal of a phosphate group from either the 5' end of the forward strand or the 5' end of the reverse strand. In some instances, the phosphate-removed nucleic acid molecule 112 can be generated by DNA synthesis. In other instances, the phosphate-removed nucleic acid molecule 112 can be generated by PCR primers.
[0201] Referring to FIG1A, a first nucleic acid molecule 110 and a second nucleic acid molecule 120 may be coupled (e.g., by ligation and / or hybridization) to generate a coupled double-stranded nucleic acid molecule 130. Coupling may be performed at a constant temperature or at multiple temperatures (e.g., stepwise or gradient or multiple temperatures). Coupling may be performed by one or more enzymes, such as ligases and / or recombinases. Subsequently, nucleic acid molecule 130 may be circularized (e.g., by ligation and / or hybridization) to form a circular nucleic acid molecule 140. Circular nucleic acid molecule 140 may include at least a portion of nucleic acid molecule 110 (e.g., recognition sequence 112 and target site 114) and at least a portion of nucleic acid molecule 120 (e.g., nick 122). Circular nucleic acid molecule 140 can be a template for sequencing (e.g., nanopore sequencing). In one example, information about target site 114 of circular nucleic acid molecule 140 can be obtained by nanopore sequencing using at least one enzyme (e.g., polymerase or nuclease), and at least one enzyme can initiate its activity (e.g., extension reaction or cleavage) at nick 122 of circular nucleic acid molecule 140.
[0202] Figures 1B and 1C schematically illustrate example methods for isolating or enriching linear or circular nucleic acids containing recognition sites. Referring to Figure 1B, a random nucleic acid molecule library 150 may include a first double-stranded nucleic acid molecule 110 including a recognition site 112 and a target site 114. Library 150 may also include linear nucleic acid molecules containingRecognition site 112, but not target site 114. Library 150 may also contain linear nucleic acid molecules that do not contain recognition site 112. Library 150 may be treated with at least one recognition portion (e.g., dCas system) 113 to bind to recognition site 112 and form a recognition complex at recognition site 112. Any linear nucleic acid molecule containing such a recognition complex may be isolated from one or more nucleic acid molecules that do not contain the recognition complex to produce a purified or enriched library 155. In some instances, at least 5% of the nucleic acid molecules in library 155 may be characterized as having recognition sequence 112. In other instances, at least a majority of the nucleic acid molecules in library 155 may be characterized as having recognition sequence 112. Recognition portion 113 may be removed (isolated) from recognition site 112 during purification or enrichment.
[0203] Referring to FIG1C, random circular nucleic acid molecule library 160 may include double-stranded nucleic acid molecule 140, which contains recognition site 112 and target site 114. Library 160 may also contain other circular nucleic acid molecules that do not contain recognition site 112. Library 160 may be treated with at least one recognition portion (e.g., dCas system) 113 to bind to recognition site 112 and form a recognition complex at recognition site 112. Any circular nucleic acid molecule containing such a recognition complex may be isolated from one or more circular nucleic acid molecules that do not contain a recognition complex to produce a purified or enriched library 165. In some examples, at least 5% (or most) of the circular nucleic acid molecules in library 165 may be characterized as having recognition sequence 112. In some instances, at least a majority of the circular nucleic acid molecules in library 165 may be characterized as having recognition sequence 112. Recognition portion 113 may be removed (isolated) from recognition site 112 during purification or enrichment.
[0204] Figure 1D schematically illustrates an example method of providing a circular nucleic acid containing a nick at a specific location within the circular nucleic acid by using one or more uracil-specific enzymes. Double-stranded nucleic acid molecules can be provided. These molecules can be fragmented or complete nucleic acid molecules derived from biological samples. Double-stranded nucleic acid molecules may contain two blunt ends. They may also not contain two blunt ends (e.g., only one blunt end or no blunt ends). In such cases, end repair may be necessary so that the repaired double-stranded nucleic acid molecule (i) has no protruding ends and (ii) contains a 5′ phosphate group and a 3′ hydroxyl group in both the sense and antisense strands. The blunt ends can be obtained by end-filling with one or more enzymes (e.g., restriction endonucleases and / or exonucleases). 5′ phosphorylation can be achieved by one or more enzymes, such as kinases like T4 polynucleotide kinase. Alternatively or additionally, a non-template deoxyadenosine 5′-monophosphate (dAMP) can be incorporated into the double-stranded nucleus.The 3' end of the blunt end of the acid molecule (i.e., dA-tail). dA-tailing can prevent tandem formation (e.g., in one or more ligation steps). dA-tailing allows a double-stranded nucleic acid molecule to be linked to one or more adaptors containing a complementary deoxythymidine monophosphate (dTMP or "dT") overhang.
[0205] Referring to FIG1D, in process 170, a modified double-stranded nucleic acid molecule may be coupled to one or more adaptors. The adaptor may contain one or more uracil (U) nucleotides (e.g., within one or more strands of the adaptor). The adaptor may contain a dT overhang complementary to the dA tail of the modified double-stranded nucleic acid molecule. Coupling may include joining (e.g., by DNA ligase) the end of the modified double-stranded nucleic acid molecule containing the dA tail to the end of the adaptor containing the dT overhang. In this example, the modified double-stranded nucleic acid molecule may be conjugated to two adaptors, wherein the first adaptor contains one or more uracil residues, while the second adaptor does not contain any uracil residues. Subsequently, the free ends of the first and second adaptors can be coupled (e.g., connected) to each other to produce a circular nucleic acid molecule comprising at least a portion of the original double-stranded nucleic acid molecule, at least a portion of the first adaptor comprising one or more uracil residues, and at least a portion of the second adaptor. The circular nucleic acid molecule can then be treated with one or more uracil-specific enzymes (e.g., uracil-specific excision reagents or “USER”) to create a single nucleotide gap (i.e., a nick) at the location of the uracil residue. The resulting circular nucleic acid molecule with the nick at a specific site can be analyzed for sequencing.
[0206] Referring to FIG1D, in process 175, a modified double-stranded nucleic acid molecule can be coupled to one or more adaptors. The adaptor may comprise one or more uracil (U) nucleotides (e.g., within one or more strands of the adaptor). The adaptor may comprise a dT overhang complementary to the dA tail of the modified double-stranded nucleic acid molecule. The adaptor may also comprise a sticky end for hybridization. Coupling can include joining (e.g., by DNA ligase) the end of a modified double-stranded nucleic acid molecule containing a dA tail to the end of an adaptor containing a dT overhang. In this example, the modified double-stranded nucleic acid molecule can be conjugated to two adaptors, wherein the first adaptor contains one or more uracil residues, and the second adaptor does not contain any uracil residues. After coupling, both adaptors can include free sticky ends. Subsequently, the free ends of the first and second adaptors can be coupled to each other (e.g., by hybridization and ligation of sticky ends) to produce a circular nucleic acid molecule comprising at least a portion of the original double-stranded nucleic acid molecule, at least a portion of the first adaptor containing one or more uracil residues, and at least a portion of the second adaptor. Subsequently, one or more uracil-specific enzymes (e.g., uracil-specific enzymes) can be used.Circular nucleic acid molecules are treated with a cleavage reagent to create a single nucleotide gap (i.e., a nick) at the location of a uracil residue. The resulting circular nucleic acid molecules with a nick at a specific site can be analyzed for sequencing.
[0207] Other examples of uracil-specific enzymes may include, but are not limited to, uracil-DNA glycosylase (UDG) and / or Afu uracil-DNA glycosylase (Afu UDG), for example, to cleave the N-glycosidic bond of deoxyuridine and create a nick for DNA polymerase to perform DNA extension reactions. Alternatively, UDG and / or Afu UDG may be combined with one or more repair enzymes that are specific to depurine / depyrimidine sites (e.g., FPG, hOGG1, hNEIL1, etc.).
[0208] Controlling the distance from the nick to the target site for sequencing
[0209] In one aspect, this disclosure provides methods for processing or analyzing circular nucleic acid molecules. This method may include providing a cell-free composition comprising a circular nucleic acid molecule containing (i) a target region and (ii) a nick site at a known distance from the target region. The method may include creating a nick at the nick site of the circular nucleic acid molecule. The target region may be a gene of interest. The target region may be a suspected mutation site or another site adjacent to the suspected mutation site. In one instance, a suspected mutation site for a specific disease may be known, and the sequence of the gene including the suspected mutation site may also be known. Thus, by assigning a nick site to a specific region of a gene characterized by a low mutation probability, the circular nucleic acid molecule comprises (i) a target region and (ii) a nick site at a known distance from the target region.
[0210] During sequencing, it may be advantageous to position the nick at a known distance from the target site in one or more ways, including but not limited to: (i) increasing the probability of multiple amplifications of the target site during amplification (e.g., RCA), (ii) increasing the probability of at least one amplification of the target site before the activity of the enzyme responsible for amplification (e.g., polymerase) is depleted, and (iii) reducing sequencing errors (e.g., enzyme errors due to enzyme fatigue).
[0211] In some instances, multiple nicking enzymes (e.g., multiple different types of nicking enzymes or different variants of the same nicking enzyme, e.g., Cas nicking enzymes with different gRNAs) may be prepared to bind to multiple nicking sites of circular nucleic acid molecules. The binding of multiple nicking enzymes to the corresponding nicking sites, any off-target binding, or nicking activity can be evaluated. At least one of the multiple nicking sites may be selected to produce low off-target binding and / or high nicking activity. Thus, by selecting nicking sites from multiple nicking sites, nicking sites can be known before processing or analyzing one or more additional circular nucleic acid molecules.Distance to the target site. The nick site may be selected from at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 30, 40, 50, 100 or more nick sites of a circular nucleic acid molecule. The nick site may be selected from at most 100, 50, 40, 30, 20, 15, 10, 9, 8, 7, 6, 5, 4, 3 or 2 nick sites of a circular nucleic acid molecule.
[0212] The generation of nicks at the nick sites of the circular nucleic acid molecule can be performed under cell-free conditions. Cell-free conditions may be substantially free of intact cells. Cell-free conditions may include conditions with cell lysates or extracts. Cell lysates may contain a fluid containing one or more contents of lysed cells. Cell lysates may be crude (i.e., unpurified) or at least partially purified (e.g., to remove cell debris or particles, such as damaged outer cell membranes). Methods for forming cell lysates may include sonication, homogenization, enzymatic lysis using lysozyme, freezing, grinding, and high-pressure lysis. Alternatively, cell-free conditions may include conditions for cell-free biological samples. Alternatively, the generation of nicks at nick sites on circular nucleic acid molecules may be performed in the presence of one or more cells (e.g., live or dead).
[0213] The nick sites may be no more than 100,000, 50,000, 10,000, 5,000, 1,000, 500, 400, 300, 200, 100, 90, 80, 70, 60, 50, 40, 30, 25, 20, 15, 10, 5, or fewer nucleotides from the target site. The cleavage site may be at least 5, 10, 15, 20, 25, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 1,000, 5,000, 10,000, 50,000, 100,000 or more nucleotides away from the target site.
[0214] The circular nucleic acid molecule may include a circular double-stranded nucleic acid molecule. The cleavage site may be part of the sense strand of the circular nucleic acid molecule. Alternatively or additionally, the cleavage site may be part of the antisense strand of the circular nucleic acid molecule. The method may also include determining the cleavage site at least in part based on the position of the target site relative to at least one reference sequence. The at least one reference sequence may comprise a common sequence of at least a portion of a gene. The cleavage site may be endogenous to the circular nucleic acid molecule. Thus, determining the cleavage site may include selecting an endogenous sequence of the circular nucleic acid molecule. Alternatively, the cleavage site may be exogenous to the circular nucleic acid molecule. Therefore, the determination may include inserting an exogenous nick site into the circular nucleic acid molecule. In some examples, such as page 32 / 69 of CN 121472373 A, the endogenous sequence of the circular nucleic acid molecule can be selected, and then the exogenous nick site can be inserted into the circular nucleic acid molecule.The nucleotide is located within or near the endogenous sequence of the target nucleotide. This allows for control or knowledge of the distance between the exogenous nicking site and the target site.
[0215] The circular nucleic acid molecule may also contain a nicking enzyme binding site specific to the nicking enzyme. In this case, the method may further include providing the nicking enzyme to the circular nucleic acid molecule under conditions sufficient to cause the nicking enzyme to associate with the nicking enzyme binding site and generate a nick. The probability of the nicking enzyme binding site appearing without any mismatch may be at most once per 1x10⁴, 5x10⁴, 7x10⁴, 1x10⁵, 5x10⁵, 1x10⁶, 5x10⁶, 1x10⁷, 5x10⁷, 1x10⁸, 5x10⁸, 1x10⁹, 5x10⁹, 1x10¹⁰, 5x10¹⁰, 1x10¹¹, 5x10¹¹, 1x10¹², 5x10¹² or more nucleotides (or base pairs). Alternatively, the probability of a nickase binding site occurring in the absence of any mismatches may be up to two, three, four, five, or more times per 1x10⁴, 5x10⁴, 7x10⁴, 1x10⁵, 5x10⁵, 1x10⁶, 5x10⁶, 1x10⁷, 5x10⁷, 1x10⁸, 5x10⁸, 1x10⁹, 5x10⁹, 1x10¹⁰, 5x10¹⁰, 1x10¹¹, 5x10¹¹, 1x10¹², 5x10¹², or more nucleotides (or base pairs).
[0216] The nickase binding site may include at least 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 35, 40, 45, 50 or more nucleotides (or base pairs). The nickase binding site may include up to 50, 45, 40, 35, 30, 29, 28, 27, 26, 25, 24, 23, 22, 21, 20, 19, 18, 17, 16, 15, 14, 13, 12, 11, 10, 9 or fewer nucleotides (or base pairs).
[0217] The nickase binding site may be endogenous for circular nucleic acid molecules. In one example, a linear nucleic acid molecule containing an endogenous nickase binding site may be circularized to form a circular nucleic acid molecule. Alternatively, the nicking enzyme binding site may be exogenous for the circular nucleic acid molecule. In this case, the method may further include inserting the exogenous nicking enzyme binding site into the circular nucleic acid molecule or its starting material (e.g., a linear nucleic acid molecule) prior to nicking. In some instances, the exogenous nicking enzyme binding site may be inserted into the linear nucleic acid molecule prior to cyclization into a circular nucleic acid molecule. The linear nucleic acid molecule may contain at least one recognition site as described herein, and the recognition portion (e.g., a catalytically active recognition portion, such as Cas9) may be used to (i) bind to the recognition site, (ii) cleave the linear nucleic acid molecule at the recognition site, and (iii)This method facilitates the insertion of a foreign nickase binding site into at least one recognition site or adjacent to at least one recognition site on a linear nucleic acid molecule (e.g., via homology-directed repair). The method may also include circularizing the nucleic acid molecule after insertion of the foreign nickase binding site. In other examples, the method may further include circularizing the linear nucleic acid molecule into a circular nucleic acid molecule before inserting the foreign nickase binding site into the circular nucleic acid molecule. After circularization, the recognition portion (e.g., a catalytically active recognition portion, such as Cas9) can be used to (i) bind to the recognition site of the circular nucleic acid molecule, (ii) cleave the circular nucleic acid molecule at the recognition site, and (iii) facilitate the insertion of the foreign nickase binding site into at least one recognition site or adjacent to at least one recognition site on the circular nucleic acid molecule (e.g., via homology-directed repair).
[0218] The nicking enzyme binding site may be no more than 30, 25, 20, 19, 18, 17, 16, 15, 14, 13, 12, 11, 10, 9, 8, 7, 6, 5, 4, 3, 2, or 1 nucleotide from the nicking site. The nicking enzyme binding site may be at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, or more nucleotides from the nicking site. In some instances, the nicking enzyme binding site may contain the nicking site. In this case, the nicking site may be within the nicking enzyme binding site.
[0219] The method may also include sequencing a circular (or circularized) nucleic acid molecule from a nick in the circular nucleic acid molecule. Sequencing may be used for whole-genome sequencing (or whole-genome sequencing) or targeted sequencing. Sequencing may include one or more NGS methods. Sequencing may include nanopore-based sequencing. The nanopore can be a protein nanopore (e.g., α-hemolysin) or a solid nanopore. Alternatively, the nanopore can be a hybrid nanopore comprising at least a portion of a protein nanopore (e.g., α-hemolysin) and at least a portion of a solid nanopore. Nanopore-based sequencing can utilize at least one enzyme (e.g., a polymerase or nuclease) to interact with at least a circular nucleic acid molecule. At least one enzyme can be coupled to the nanopore. At least one enzyme can be fused, conjugated, or bound to a protein nanopore or a membrane containing a nanopore. At least one enzyme can be conjugated or bound to a solid nanopore or a membrane containing a solid nanopore. In some instances, at least one enzyme may have a binding portion capable of binding to a nanopore (or solid nanopore) or a membrane.
[0220] Sequencing can include sequencing a circular (or circularized) nucleic acid molecule. Sequencing can include causing the circular nucleic acid molecule to undergo an extension reaction starting from a nick to produce a growth chain that is sequence complementary to at least a portion of the chain of the circular nucleic acid molecule. The method may also include obtaining sequence information of at least a portion of the growth chain. Obtaining sequence information may include detection.At least a portion of the growth chain. The extension reaction may include contacting the circular nucleic acid molecule with the nucleotide coupled to the tag under conditions sufficient to incorporate the nucleotide into the growth chain. Obtaining sequence information may include detecting at least a portion of the tag. When performing sequencing analysis, at least a portion of the tag may be linked to the growth chain. Alternatively, the method may further include releasing the tag from the nucleotide as the nucleotide is incorporated into the growth chain, and detecting the released tag for sequencing.
[0221] The extension reaction may be performed using oligonucleotide primers. Alternatively, the extension reaction may be performed without the use of oligonucleotide primers. The cleavage within the double-stranded adapter may serve as a binding site for an enzyme (e.g., polymerase) capable of performing an extension reaction (e.g., rolling circle amplification), and therefore oligonucleotide primers may not be required.
[0222] Alternatively, sequencing may include (i) subjecting the circular nucleic acid molecule to a cleavage reaction to cleave at least a portion of the chain of the double-stranded nucleic acid molecule, and (ii) obtaining sequence information of at least a portion of the chain. The chain of the circular nucleic acid molecule may be its sense strand or its antisense strand. Therefore, the growth strand can exhibit complementarity with at least a portion of the sense strand or at least a portion of the antisense strand. Alternatively, the strands of the double-stranded nucleic acid molecule can be a sense strand and an antisense strand. Therefore, the first growth strand can exhibit complementarity with at least a portion of the sense strand, while the second growth strand can exhibit complementarity with at least a portion of the antisense strand.
[0223] The extension reaction can include amplifying at least a portion of the circular nucleic acid molecule (e.g., at least a portion of the cleavage site and at least a portion of the target site). Amplification (e.g., RCA) can produce multiple copies (e.g., at least 1, 2, 3, 4, 5, 10, 15, 20 or more copies; up to 20, 15, 10, 5, 4, 3, 2 or 1 copy) of at least a portion of the sense strand and / or antisense strand of the circular nucleic acid molecule. The extended product based on the circular nucleic acid molecule will have a first domain complementary to at least a portion of the nick site and a second domain complementary to at least a portion of the target site, and the distance between the first and second domains can be substantially the same as the known distance between the target site and the nick site in the circular nucleic acid molecule. Subsequently, obtaining sequence information may include detecting at least a portion of the strand.
[0224] At least a portion of the circular nucleic acid molecule may have or be suspected of having one or more sequencing variants (e.g., one or more mutations) compared to at least one reference sequence. Therefore, sequencing can be performed to identify the presence of at least a portion of the circular nucleic acid molecule. One or more sequencing variants may indicate mutations in a gene. At least one reference sequence may contain a common sequence of at least a portion of the gene. The common sequence and the double-stranded nucleic acid molecule may be derived from the same species or different species. The common sequence may be from multiple samples of the same species (e.g., at least 2, 3, 4, 5, 10, 15, 20, 30, 40, 50 or more).A representative sequence of a collection of multiple sequences of a gene obtained from multiple samples (up to 50, 40, 30, 20, 15, 10, 4, 3, or 2 samples). In one instance, the circular nucleic acid molecule and at least one reference sequence may both be derived from a human sample, and the at least one reference sequence may be a common sequence of at least a portion of the human gene of interest, such as a portion of a gene known to generally not have any mutations.
[0225] The method may include circularizing at least a linear nucleic acid molecule to produce a circular nucleic acid molecule. In some cases, the method may also include amplifying the linear nucleic acid molecule to produce multiple copies of the linear nucleic acid molecule. Amplification may produce at least 1, 2, 3, 4, 5, 10, 15, 20, 30, 40, 50, 100, 150, 200 or more copies of the linear nucleic acid molecule. Amplification may produce up to 200, 150, 100, 50, 40, 30, 20, 15, 10, 5, 4, 3, 2 or 1 copy of the linear nucleic acid molecule. The method may also include circularizing one or more copies of a linear nucleic acid molecule to produce a plurality of circular nucleic acid molecules. Circularization may include self-ligation (e.g., by one or more ligases), ligation by an adaptor, hybridization by an adaptor, or a combination thereof.
[0226] Alternatively or additionally, the method may include amplifying a circular nucleic acid molecule to produce a plurality of copies of the circular nucleic acid molecule. Amplification may produce at least 1, 2, 3, 4, 5, 10, 15, 20, 30, 40, 50, 100, 150, 200 or more copies of the circular nucleic acid molecule. Amplification may produce up to 200, 150, 100, 50, 40, 30, 20, 15, 10, 5, 4, 3, 2 or 1 copy of the circular nucleic acid molecule.
[0227] The circular nucleic acid molecule may contain a recognition sequence. The recognition sequence may be endogenous or exogenous for the circular nucleic acid molecule. The recognition sequence may contain at least one natural nucleotide, at least one synthetic nucleotide, or both. The recognition sequence may contain at least 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 40, 50, or 100 or more nucleotides. The recognition sequence may contain up to 100, 50, 40, 30, 25, 24, 23, 22, 21, 20, 19, 18, 17, 16, 15, 14, 13, 12, 11, 10, 9, 8, 7, 6, 5, or fewer nucleotides. The method may also include enriching circular nucleic acid molecules from a random nucleic acid molecule library, at least partially based on the recognition sequence. The library may contain at least 2, 3, 4, 5, 10, 50, 100, 500, or more nucleotides.1,000, 5,000, 10,000, 50,000, 100,000, 500,000, 1,000,000 or more random nucleic acid molecules. The library may contain up to 1,000,000, 500,000, 100,000, 50,000, 10,000, 500, 100, 50, 10, 5, 4, 3 or 2 random nucleic acid molecules. Enrichment may involve isolating circular nucleic acid molecules (containing the recognition sequence) from at least one different nucleic acid molecule that does not contain the recognition sequence. Circular nucleic acid molecules can be isolated from at least 1, 5, 10, 50, 100, 500, 1,000, 5,000, 10,000, 50,000, 100,000, 500,000, 1,000,000, or more different nucleic acid molecules that do not contain a recognition sequence. Circular nucleic acid molecules can also be isolated from at most 1,000,000, 500,000, 100,000, 50,000, 10,000, 5,000, 1,000, 500, 100, 50, 10, 5, or 1 different nucleic acid molecules that do not contain a recognition sequence.
[0228] Before amplification, linear nucleic acid molecules containing a recognition sequence can be enriched from a random nucleic acid molecule library (one of which is a linear nucleic acid molecule containing a recognition sequence). Alternatively or additionally, the random nucleic acid molecule library can be amplified before enriching linear nucleic acid molecules containing a recognition sequence. Alternatively, circular nucleic acid molecules in a random nucleic acid library (one of which is a circular nucleic acid molecule containing a recognition sequence) can be enriched prior to amplification. Alternatively or additionally, the random nucleic acid library can be amplified prior to enrichment of the circular nucleic acid molecules.
[0229] Enrichment may include generating a selected circular nucleic acid molecule library. Each circular nucleic acid molecule in at least 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 60%, 70%, 80%, 90%, 95% or more of the selected library may contain at least a portion of the recognition sequence. Each circular nucleic acid molecule in at most 100%, 95%, 90%, 80%, 70%, 60%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 15%, 10%, 5% or less of the selected library may contain at least a portion of the recognition sequence.
[0230] The probability of a recognized sequence occurring in the absence of any mismatches can be at most once in every 1x10⁴, 5x10⁴, 7x10⁴, 1x10⁵, 5x10⁵, 1x10⁶, 5x10⁶, 1x10⁷, 5x10⁷, 1x10⁸, 5x10⁸, 1x10⁹, 5x10⁹, 1x10¹⁰, 5x10¹⁰, 1x10¹¹, 5x10¹¹, 1x10¹², 5x10¹² or more nucleotides (or base pairs). Alternatively, in the absence of any mismatches...The probability of the recognition sequence appearing in any given situation can be up to two, three, four, five or more times in every 1x10⁴, 5x10⁴, 7x10⁴, 1x10⁵, 5x10⁵, 1x10⁶, 5x10⁶, 1x10⁷, 5x10⁷, 1x10⁸, 5x10⁸, 1x10⁹, 5x10⁹, 1x10¹⁰, 5x10¹⁰, 1x10¹¹, 5x10¹¹, 1x10¹², 5x10¹² or more nucleotides (or base pairs).
[0231] The identification sequence may contain at least 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 35, 40, 45, 50 or more nucleotides (or base pairs). The identification sequence may contain up to 50, 45, 40, 35, 30, 29, 28, 27, 26, 25, 24, 23, 22, 21, 20, 19, 18, 17, 16, 15, 14, 13, 12, 11, 10, 9 or fewer nucleotides (or base pairs).
[0232] Enrichment may include (i) binding a recognition moiety complementary to the recognition sequence to a circular nucleic acid molecule to form a recognition complex as described on pages 35 / 69 of the specification, CN 121472373 A, and (ii) extracting the recognition complex from a random nucleic acid molecule library. The recognition moiety may have at least 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95% or more complementarity to the recognition sequence. The recognition moiety may have at most 100%, 95%, 90%, 80%, 70%, 60%, 50%, 40%, 30% or less complementarity to the recognition sequence.
[0233] One or more circular nucleic acid molecules containing a recognition sequence from a random nucleic acid molecule library may be enriched, the enrichment being performed (i) prior to providing a cell-free composition containing a circular nucleic acid molecule comprising a target region and a nick site at a known distance from the target region, or (ii) after a nick has been created at the nick site of the circular nucleic acid molecule. Alternatively, one or more circular nucleic acid molecules containing a recognition sequence from a random nucleic acid molecule library may be enriched, the enrichment being performed (i) prior to providing a cell-free composition containing a circular nucleic acid molecule comprising a target region and a nick site at a known distance from the target region, and (ii) after a nick has been created at the nick site of the circular nucleic acid molecule.
[0234] The circular nucleic acid molecules may be derived from or derived from a biological sample of the subject. The biological sample may include a cell-free biological sample of the subject. Cell-free biological samples can be selected from: blood, plasma, serum, urine, lymph, feces, saliva, semen, amniotic fluid, cerebrospinal fluid, bile, sweat, tears, sputum, synovial fluid, vomit, and combinations thereof. Circular nucleic acid molecules can originate from or be derived from...Cell-free nucleic acid molecules derived from cell-free biological samples. Cell-free nucleic acid molecules may include circulating tumor nucleic acid molecules (e.g., ctDNA) or amniotic fluid nucleic acid molecules.
[0235] Biological samples may include tissue samples of a subject. Tissue samples may be derived from: bone, heart, thymus, arteries, blood vessels, lungs, muscles, stomach, intestines, liver, pancreas, spleen, kidneys, gallbladder, thyroid gland, adrenal glands, breast, ovaries, prostate, testes, skin, fat, eyes, brain, and combinations thereof. Circular nucleic acid molecules may be derived from or derived from genomic nucleic acid molecules from tissue samples. Tissue samples may be derived from: infected tissue, diseased tissue, malignant tissue, calcified tissue, healthy tissue, and combinations thereof. Tissue samples may be derived from malignant tissue containing tumors, sarcomas, leukemia, or derivatives thereof.
[0236] Circular nucleic acid molecules may include DNA, cDNA, ctDNA, derivatives thereof, or combinations thereof. Circular nucleic acid molecules include RNA.
[0237] In another aspect, this disclosure provides reaction mixtures for processing or analyzing circular nucleic acid molecules. The reaction mixture may contain a cell-free composition comprising a circular nucleic acid molecule. The circular nucleic acid molecule may contain (i) a target site and (ii) a nick site at a known distance from the target site. The reaction mixture may contain at least one enzyme that creates a nick at the nick site of the circular nucleic acid molecule. At least one enzyme may comprise a nuclease (e.g., a restriction endonuclease) or a nicking enzyme (e.g., a Cas9n nicking enzyme). At least one enzyme may comprise both a nuclease and a nicking enzyme. As provided in this disclosure, the reaction mixture may be used or identified in any subject method for sequencing at a known nick-to-target distance. The reaction mixture may be used to prepare one or more libraries (e.g., libraries of nucleic acid molecules, enzymes, or combinations thereof) or compositions for one or more sequencing methods. One or more components of the reaction mixture may be used simultaneously in the same reaction. In one example, the reaction may be carried out in a reaction vial (e.g., a reaction tube), thereby reducing purification steps and / or sample loss, and / or sequencing with a small amount of nucleic acid sample input. One or more components of the reaction mixture may be used separately in different reactions. The reaction mixture may be a cell-free reaction mixture. Alternatively, the reaction mixture may not be a cell-free reaction mixture.
[0238] The cleavage site may be no more than 100,000, 50,000, 10,000, 5,000, 1,000, 500, 400, 300, 200, 100, 90, 80, 70, 60, 50, 40, 30, 25, 20, 15, 10, 5 or fewer nucleotides away from the target site. The cleavage site may be at least 5, 10, 15, 20, 25, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 1,000, 5 nucleotides away from the target site.000, 10,000, 50,000, 100,000 or more nucleotides.
[0239] Circular nucleic acid molecules may include circular double-stranded nucleic acid molecules. The nick site may be a portion of the strand of the circular nucleic acid molecule. Alternatively or additionally, the nick site may be a portion of the antisense strand of the circular nucleic acid molecule. The method may also include determining the nick site based at least in part on the position of the target site relative to at least one reference sequence. The at least one reference sequence may comprise a common sequence of at least a portion of the gene. The nick site may be endogenous for the circular nucleic acid molecule. Therefore, determining the nick site may include selecting an endogenous sequence of the circular nucleic acid molecule. Alternatively, the nick site may be exogenous for the circular nucleic acid molecule. Therefore, the determination may include inserting an exogenous nick site into the circular nucleic acid molecule. In some instances, an endogenous sequence of the circular nucleic acid molecule may be selected, and then an exogenous nick site may be inserted within or near the endogenous sequence of the circular nucleic acid molecule. In this way, the distance between the exogenous nick site and the target site can be controlled or known.
[0240] The nucleic acid molecule may also contain an enzyme binding site that is specific to at least one enzyme. At least one enzyme may be included in some instances where the at least one enzyme can exhibit nick enzyme activity, and the enzyme binding site may be the same as the nick enzyme binding site, as provided in this disclosure.
[0241] The probability of an enzyme binding site occurring without any mismatch may be at most once per 1x10⁴, 5x10⁴, 7x10⁴, 1x10⁵, 5x10⁵, 1x10⁶, 5x10⁶, 1x10⁷, 5x10⁷, 1x10⁸, 5x10⁸, 1x10⁹, 5x10⁹, 1x10¹⁰, 5x10¹⁰, 1x10¹¹, 5x10¹¹, 1x10¹², 5x10¹² or more nucleotides (or base pairs). Alternatively, the probability of an enzyme binding site occurring in the absence of any mismatches may be up to two, three, four, five, or more times per 1x10⁴, 5x10⁴, 7x10⁴, 1x10⁵, 5x10⁵, 1x10⁶, 5x10⁶, 1x10⁷, 5x10⁷, 1x10⁸, 5x10⁸, 1x10⁹, 5x10⁹, 1x10¹⁰, 5x10¹⁰, 1x10¹¹, 5x10¹¹, 1x10¹², 5x10¹², or more nucleotides (or base pairs).
[0242] The enzyme binding site may contain at least 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 35, 40, 45, 50 or more nucleotides (or base pairs). The enzyme binding site may contain up to50, 45, 40, 35, 30, 29, 28, 27, 26, 25, 24, 23, 22, 21, 20, 19, 18, 17, 16, 15, 14, 13, 12, 11, 10, 9 or fewer nucleotides (or base pairs).
[0243] The enzyme binding site can be endogenous for the circular nucleic acid molecule. In one example, a linear nucleic acid molecule containing an endogenous enzyme binding site can be cyclized to form a circular nucleic acid molecule. Alternatively, the enzyme binding site can be exogenous for the circular nucleic acid molecule. In this case, the method may further include inserting the exogenous enzyme binding site into the circular nucleic acid molecule or its starting material (e.g., a linear nucleic acid molecule) before creating a nick. In some examples, the exogenous enzyme binding site can be inserted into the linear nucleic acid molecule before it is cyclized into a circular nucleic acid molecule. The linear nucleic acid molecule may contain at least one recognition site as provided herein, and the recognition portion (e.g., a catalytically active recognition portion, such as Cas9) may be used to (i) bind to the recognition site, (ii) cleave the linear nucleic acid molecule at the recognition site, and (iii) facilitate the insertion of a foreign enzyme binding site into at least one recognition site of the linear nucleic acid molecule or adjacent to at least one recognition site (e.g., via homology-directed repair). The method may also include circularizing the linear nucleic acid molecule after insertion of the foreign enzyme binding site. In other examples, the method may further include circularizing the linear nucleic acid molecule into a circular nucleic acid molecule before inserting the cleavage enzyme binding site into the circular nucleic acid molecule. After circularization, the recognition portion (e.g., a catalytically active recognition portion, such as Cas9) may be used to (i) bind to the recognition site of the circular nucleic acid molecule, (ii) cleave the circular nucleic acid molecule at the recognition site, and (iii) facilitate the insertion of a foreign enzyme binding site into at least one recognition site of the circular nucleic acid molecule or adjacent to at least one recognition site (e.g., via homology-directed repair).
[0244] The enzyme binding site may be no more than 30, 25, 20, 19, 18, 17, 16, 15, 14, 13, 12, 11, 10, 9, 8, 7, 6, 5, 4, 3, 2, or 1 nucleotide from the cleavage site. The enzyme binding site may be at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, or more nucleotides from the cleavage site. In some instances, the enzyme binding site may contain the cleavage site. In this case, the cleavage site may be within the enzyme binding site.
[0245] The reaction mixture also contains at least a second enzyme that performs an extension reaction from the cleavage to produce a growth chain that is sequence complementary to at least a portion of the circular nucleic acid molecule described in page 37 / 69 of the specification, CN 121472373 A. Circular nucleic acid molecules can be circular double-stranded nucleic acid molecules.The growth chain can exhibit sequence complementarity with at least a portion of the chain of the circular double-stranded nucleic acid molecule. The chain of the circular double-stranded nucleic acid molecule can be its sense strand, antisense strand, or both. In some instances, at least the second enzyme can be a polymerase. The reaction mixture can also contain at least one nucleotide coupled to a tag, wherein at least the second enzyme incorporates the nucleotide into the growth chain. The tag can be a small molecule, nucleotide, polynucleotide, amino acid, polypeptide, polymer, metal and / or ceramic particle, etc. In some instances, each of the nucleotides G, C, A, T, and U can contain a distinct tag that is distinguishable from each other. In some instances, the tag may not be released from the nucleotide when the nucleotide is incorporated into the growth chain. Alternatively, at least the second enzyme can release the tag from the nucleotide when the nucleotide is incorporated into the growth chain. At least the second enzyme can perform an extension reaction with the aid of at least one oligonucleotide primer. Alternatively, at least the second enzyme can perform an extension reaction without the use of an oligonucleotide primer. In one instance, the extension reaction can include RCA, wherein at least the second enzyme includes a polymerase. The polymerase can bind to a nick at a nick site for RCA.
[0246] The reaction mixture may further comprise at least a third enzyme that performs a cleavage reaction from a nick to cleave at least a portion of the circular nucleic acid molecule. Starting from the nick, the at least third enzyme may displace and cleave at least a portion of the chain of the circular nucleic acid molecule. The at least third enzyme may comprise a nuclease (e.g., an endonuclease, such as a restriction endonuclease).
[0247] As provided in this disclosure, at least a portion of the circular nucleic acid molecule may have or be suspected of having one or more variants compared to at least one reference sequence. The reaction mixture may be used to prepare at least one composition for sequencing to identify the presence of at least a portion of the circular nucleic acid molecule.
[0248] As provided in this disclosure, the circular nucleic acid molecule may comprise a recognition sequence. Therefore, the reaction mixture may further comprise a recognition portion associated with the recognition sequence to enrich the circular nucleic acid molecule from a random nucleic acid molecule library in the composition at least partially based on the recognition sequence. The recognition portion may comprise at least one oligonucleotide (e.g., at least a portion of a gRNA for Cas system variants) that is complementary to the at least recognition sequence. The oligonucleotide of the recognition portion may have at least 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95% or more complementarity with the recognition sequence. The oligonucleotide of the recognition portion may have at most 100%, 95%, 90%, 80%, 70%, 60%, 50%, 40%, 30% or less complementarity with the recognition sequence.
[0249] The composition of the reaction mixture may comprise a library of selected circular nucleic acid molecules (e.g., single-stranded or double-stranded). At least 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 60%, 70%, 80%, ...Each circular nucleic acid molecule in 90%, 95%, or more may contain at least a portion of the recognition sequence. Each circular nucleic acid molecule in the selected library in up to 100%, 95%, 90%, 80%, 70%, 60%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 15%, 10%, 5%, or less may contain at least a portion of the recognition sequence.
[0250] The probability of an identification sequence occurring in the absence of any mismatches may be at most once in every 1x10⁴, 5x10⁴, 7x10⁴, 1x10⁵, 5x10⁵, 1x10⁶, 5x10⁶, 1x10⁷, 5x10⁷, 1x10⁸, 5x10⁸, 1x10⁹, 5x10⁹, 1x10¹⁰, 5x10¹⁰, 1x10¹¹, 5x10¹¹, 1x10¹², 5x10¹² or more nucleotides (or base pairs). Alternatively, the probability of a recognized sequence appearing without any mismatches may be up to two, three, four, five, or more times in every 1x10⁴, 5x10⁴, 7x10⁴, 1x10⁵, 5x10⁵, 1x10⁶, 5x10⁶, 1x10⁷, 5x10⁷, 1x10⁸, 5x10⁸, 1x10⁹, 5x10⁹, 1x10¹⁰, 5x10¹⁰, 1x10¹¹, 5x10¹¹, 1x10¹², 5x10¹², or more nucleotides (or base pairs).
[0251] The recognition sequence may contain at least 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 35, 40, 45, 50 or more nucleotides (or base pairs). The recognition sequence may contain up to 50, 45, 40, 35, 30, 29, 28, 27, 26, 25, 24, 23, 22, 21, 20, 19, 18, 17, 16, 15, 14, 13, 12, 11, 10, 9 or fewer nucleotides (or base pairs). Specification 38 / 69 pages 40 CN 121472373 A
[0252] In different aspects, this disclosure provides cell-free libraries of circular nucleic acid molecules. Each individual circular nucleic acid molecule in at least 5% of the library may contain (i) a target site and (ii) a nick at a known distance from the target site. The circular nucleic acid molecule library may be a starting material or product of any of the subject methods or reaction mixtures provided in this disclosure.
[0253] A cell-free library of circular nucleic acid molecules may be a product or byproduct of enrichment of a random nucleic acid molecule library containing any nucleic acid molecule containing a recognition site. In one example, a random circular nucleic acid molecule library may be enriched with at least one circular nucleic acid molecule containing a recognition site and treated with at least one enzyme (e.g., a nicking enzyme, such as Cas9n nicking enzyme) to nick at a known distance from the target site.A nick is created at a known distance from the target site. In another example, a random circular nucleic acid library can be treated with at least one enzyme to create a nick at a known distance from the target site and enrich at least one circular nucleic acid molecule containing a recognition site. In a different example, at least one linear nucleic acid molecule containing a recognition site can be enriched from a random linear nucleic acid library, treated with at least one first enzyme (e.g., a ligase and / or a recombinase) to circularize one or more linear nucleic acid molecules, and treated with at least one second enzyme (e.g., a nicking enzyme, such as the Cas9n nicking enzyme) to create a nick at a known distance from the target site. In a different example, at least one circular nucleic acid molecule containing a recognition site can be enriched from a random circular nucleic acid library and treated with at least one enzyme (e.g., a recognition moiety) to insert a nick site (a nick may or may not exist before insertion). In a different example, a random circular nucleic acid library can be treated with at least one enzyme (e.g., a recognition moiety) to insert a nick site (a nick may or may not exist before insertion) and enrich at least one circular nucleic acid molecule containing a recognition site. In various instances, a random linear nucleic acid library may be treated with at least a first enzyme (e.g., a recognition moiety) to insert a nick site (which may or may not exist prior to insertion), treated with at least a second enzyme (e.g., a ligase and / or a recombinase) to circularize one or more linear nucleic acid molecules, and enriched with at least one circular nucleic acid molecule containing a recognition site. In another instance, a random linear nucleic acid library may be treated with at least a first enzyme (e.g., a recognition moiety) to insert a nick site (which may or may not exist prior to insertion), enriched with at least one linear nucleic acid molecule containing a recognition site, and treated with at least a second enzyme (e.g., a ligase and / or a recombinase) to circularize one or more linear nucleic acid molecules. In yet another instance, at least one linear nucleic acid molecule containing a recognition site may be enriched from a random linear nucleic acid library, treated with at least a first enzyme (e.g., a recognition moiety) to insert a nick site (which may or may not exist prior to insertion), and treated with at least a second enzyme (e.g., a ligase and / or a recombinase) to circularize one or more linear nucleic acid molecules.
[0254] The circular nucleic acid molecules of each individual in at least 5%, 10%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95% or more of the library may contain (i) a target site and (ii) a nick at a known distance from the target site. The circular nucleic acid molecules of each individual in at most 100%, 95%, 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20% or less of the library may contain (i) a target site and (ii) a nick at a known distance from the target site.The circularized nucleic acid molecule may contain (i) a target site and (ii) a nick at a known distance from the target site.
[0255] The circularized nucleic acid molecule of an individual may also contain a recognition sequence. The probability of the recognition sequence appearing without any mismatch may be at most once in every 1x10⁴, 5x10⁴, 7x10⁴, 1x10⁵, 5x10⁵, 1x10⁶, 5x10⁶, 1x10⁷, 5x10⁷, 1x10⁸, 5x10⁸, 1x10⁹, 5x10⁹, 1x10¹⁰, 5x10¹⁰, 1x10¹¹, 5x10¹¹, 1x10¹², 5x10¹² or more nucleotides (or base pairs). Alternatively, the probability of a recognized sequence appearing without any mismatches may be up to two, three, four, five, or more times in every 1x10⁴, 5x10⁴, 7x10⁴, 1x10⁵, 5x10⁵, 1x10⁶, 5x10⁶, 1x10⁷, 5x10⁷, 1x10⁸, 5x10⁸, 1x10⁹, 5x10⁹, 1x10¹⁰, 5x10¹⁰, 1x10¹¹, 5x10¹¹, 1x10¹², 5x10¹², or more nucleotides (or base pairs).
[0256] The identification sequence may contain at least 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 35, 40, 45, 50 or more nucleotides (or base pairs). The identification sequence may contain up to 50, 45, 40, 35, 30, 29, 28, 27, 26, 25, 24, 23, 22, 21, 20, 19, 18, 17, 16, 15, 14, 13, 12, 11, 10, 9 or fewer nucleotides (or base pairs).
[0257] The cell-free library may contain at least a first and a second somatic nucleic acid molecule. In some instances, (i) the first target site of the first somatic nucleic acid molecule and (ii) the second target site of the second somatic nucleic acid molecule may be the same. In this case, (1) the first known distance between the first nick and the first target site of the first somatic nucleic acid molecule and (2) the second known distance between the second nick and the second target site of the second somatic nucleic acid molecule may be the same. Alternatively, (1) the first known distance between the first nick and the first target site of the first somatic nucleic acid molecule and (2) the second known distance between the second nick and the second target site of the second somatic nucleic acid molecule may be different. In some instances, (i) the first target site of the first somatic nucleic acid molecule and (ii) the second target site of the second somatic nucleic acid molecule may be different. In this case, (1) the first known distance between the first nick and the first target site of the first somatic nucleic acid molecule and (2) the second known distance between the second nick and the second target site of the second somatic nucleic acid molecule may be different.The second known distance between the second cleavage site of the acid molecule and the second target site can be the same. Alternatively, (1) the first known distance between the first cleavage site of the first somatic nucleic acid molecule and the first target site and (2) the second known distance between the second cleavage site of the second somatic nucleic acid molecule and the second target site can be different.
[0258] The cleavage site can be no more than 100,000, 50,000, 10,000, 5,000, 1,000, 500, 400, 300, 200, 100, 90, 80, 70, 60, 50, 40, 30, 25, 20, 15, 10, 5 or fewer nucleotides from the target site. The nick site can be at least 5, 10, 15, 20, 25, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 1,000, 5,000, 10,000, 50,000, 100,000 or more nucleotides from the target site.
[0259] The circular nucleic acid molecules of the selected library may include circular double-stranded nucleic acid molecules. The nick can be on the sense strand of the circular double-stranded nucleic acid molecule or on the antisense strand of the circular double-stranded nucleic acid molecule. The nick can be on the sense strand of the circular double-stranded nucleic acid molecule and on the antisense strand of the circular double-stranded nucleic acid molecule.
[0260] Figures 2A and 2B schematically illustrate example methods for generating a nick at a known distance from the target site within a circular nucleic acid. Referring to Figure 2A, a circular nucleic acid 210a can be provided. The circular nucleic acid 210a may include a target site 214 and a nicking enzyme binding site 212 specific to a nicking enzyme 220 (e.g., Cas9n). The target site 214 may have or be suspected of having one or more mutations compared to at least a reference sequence (e.g., a shared sequence of a portion of a gene from a species or multiple species). The nicking enzyme binding site 212 may be located at a known distance 216 from the target site 214 (e.g., the number of nucleotides between the nicking enzyme binding site 212 and the target site 214 is known). The circular nucleic acid 210a may also include a recognition sequence 218 that can be specifically recognized by a recognition portion 230 (e.g., Cas or dCas). The circular nucleic acid 210a may be treated with the nicking enzyme 220. The nicking enzyme 220 may bind to the nicking enzyme binding site 212 and create a nick, thereby forming a circular nucleic acid 210b with a nick 222 located at, near, or within the nicking enzyme binding site 212. When nick 222 is formed, nicking enzyme 220 can be removed or separated (e.g., automatically separated). Nicking enzyme binding site 212 can be endogenous or exogenous for circular nucleic acid 210a.
[0261] Referring to FIG2B, circular nucleic acid 220a can be provided. Circular nucleic acid 220a may include target site 214 as provided herein and recognition sequence 218, which can be specifically recognized by recognition portion 230 (e.g., Cas). Circular nucleic acid 220aThe recognition portion 230 can be used to create a break at, near, or within the recognition site 218. After the break is created, a nicking enzyme binding site 212 can be inserted (e.g., by homology-directed repair) into the break, and the circular nucleic acid can be sealed to form a circular nucleic acid molecule 220b. Subsequently, the nicking enzyme 220 can bind to the nicking enzyme binding site 212 and create a nick, thereby forming a circular nucleic acid molecule 220c with a nick 222 at, near, or within the nicking enzyme binding site 212. When the nick 222 is created, the nicking enzyme 220 can be removed or isolated (e.g., automatically isolated).
[0262] Figures 2C and 2D schematically illustrate an example method for isolating or enriching circular nucleic acids containing recognition and target sites, as described on pages 40 / 69 of CN 121472373 A. Referring to Figure 2C, a random nucleic acid library may include a circular nucleic acid molecule 210a containing a target site 214, a nickase binding site 212 at a known distance 216 from the target site 214, and a recognition site 218. The random nucleic acid library may be treated with at least a recognition portion 150 (e.g., dCas) to form a recognition complex. In some instances, the recognition portion 150 may be conjugated to magnetic beads. Alternatively, the recognition portion 150 may contain one or more biotin molecules, which may subsequently be coupled to magnetic beads presenting avidin via an avidin-biotin interaction. The resulting recognition complex can be pulled out (separated) from other nucleic acid molecules lacking the recognition site 218 by magnetic bead separation. Similarly, the random nucleic acid library containing circular nucleic acid molecules 210b may be enriched (e.g., by using the recognition portion 150).
[0263] Referring to FIG2D, a random nucleic acid library may contain circular nucleic acid molecules 220a, which contain a target site 214 and a recognition site 218. Using a similar method provided in FIG2C (e.g., by using the recognition portion), the library may enrich the circular nucleic acid molecules 220a, or the circular nucleic acid molecules 220a may be isolated from the library. Alternatively, a random nucleic acid library may contain circular nucleic acid molecules 220b, which contain a target site 214, a nick site 212 at a known distance 216 from the target site 214, and a recognition site 218. Using a similar method provided in FIG2C (e.g., by using the recognition portion), the library may enrich the circular nucleic acid molecules 220b, or the circular nucleic acid molecules 220b may be isolated from the library. In different alternatives, the random nucleic acid library may contain circular nucleic acid molecules 220c, which include a target site 214, a nick 222 at a known distance 216 from the target site 214, and a recognition site 218. Using a similar method as shown in Figure 2C (e.g.,By using the identification portion, the library can be enriched with circular nucleic acid molecules 220c, or circular nucleic acid molecules 220c can be isolated from the library.
[0264] Figure 3 schematically illustrates an example method for sequencing double-stranded nucleic acids. Double-stranded nucleic acid molecules can be provided. Double-stranded nucleic acid molecules can be fragmented or whole nucleic acid molecules from biological samples. Double-stranded nucleic acid molecules can contain two blunt ends. Alternatively, double-stranded nucleic acid molecules may not contain two blunt ends (e.g., only one blunt end or no blunt ends). In this case, end repair may be required so that the repaired double-stranded nucleic acid molecules can (i) have no protruding ends, or (ii) contain a 5′ phosphate group and a 3′ hydroxyl group in the sense and antisense strands, as provided in Figure 1D. Referring to Figure 3, with or without such repair, double-stranded nucleic acid molecules can be denatured into isolated single-stranded nucleic acid molecules. The single-stranded nucleic acid molecules of each individual can be circularized to form single-stranded circular nucleic acid molecules. In some instances, isolated single-stranded nucleic acid molecules (e.g., sense and antisense strands that are at least partially or completely complementary to each other) can be coupled (e.g., linked) and then circularized into a single-stranded circular nucleic acid molecule. Subsequently, the circular nucleic acid molecule can be placed under conditions sufficient to allow one or more random hexamer primers to hybridize with complementary domains of the circular nucleic acid molecule. The resulting circular nucleic acid molecule having one or more hybridized hexamers can be analyzed by synthetic sequencing, such as detection of electrical signals or visualization via optical imaging.
[0265] In one instance, nanopore sequencing can be used. A polymerase can bind to one of the hybridized hexamers and initiate an extension reaction. During the extension reaction, other hybridized hexamers can be displaced from the circular nucleic acid molecule by the direct or indirect activity of the polymerase. The polymerase can be located or coupled to a nanopore (e.g., a protein nanopore or a solid nanopore).
[0266] Whole Genome and Targeted Sequencing
[0267] For whole genome sequencing, the binding site of the polymerase within the circular polynucleotide can be left uncontrolled. Nonspecific nicking enzymes can be used to generate nicks at random positions on the chain of a cyclic polynucleotide, and polymerases can bind to nonspecific nicking sites, for example, to perform an elongation reaction. Alternatively, the binding site of the polymerase within the cyclic polynucleotide may not be controlled. Nicking enzymes that recognize specific sequences can be used to bind and generate nicks, such as the Cas system (e.g., Cas nicking enzymes, like Cas9n), N.AlwI, Nb.BbvCl, Nt.BbvCl, Nb.BsmI, Nt.BsmAI, Nt.BspQl, Nb.BsrDI, Nt.BstNBI, Nb.BstsCI, Nt.CviPII, Nb.BpulOI, Nt.BpulOI, and Nt,Bst9I, variants thereof, and combinations thereof. Specification 41 / 69 pages 43 CN 121472373 A
[0268] Targeted polynucleotide sequencing can be used to detect sequence variations (e.g., one or more mutations) at specific locations of polynucleotides. Sequence variations can be, for example, single nucleotide polymorphisms (SNPs). For targeted sequencing, it may be necessary to bind a polymerase to the vicinity of the polynucleotide sequence of interest (e.g., the target site), such as a segment that may contain sequence variants (e.g., mutations). In some instances, a Cas system comprising a Cas nickase and sgRNA can be used. A polynucleotide segment with base pairs complementary to the polynucleotide sequence adjacent to the polynucleotide segment of interest can be generated as part of the sgRNA. The sgRNA can bind to the Cas9n nickase to form an sgRNA / Cas9n complex, which can bind to the polynucleotide at a segment identified by at least a portion of the sgRNA and generate a nick.
[0269] Figure 4A schematically illustrates an example method of targeted sequencing using nanopore sequencing. Samples comprising genomic DNA / cDNA or cell-free DNA / cfDNA can be amplified. To enrich the DNA / cDNA mixture with the targeted polynucleotide sequence, genomic DNA / cDNA can be reacted with a biotinylated sgRNA / CRISPER / Cas9 complex to cleave the DNA / cDNA in the region of interest. The DNA mixture can be enriched with the target DNA segment by purification using streptavidin beads. The enriched target DNA sample or cell-free DNA (such as ctDNA) can then be circularized. The circular DNA can be bound to an sgRNA / CRISPR / Cas9n cleavage enzyme to provide a cleavage site in the DNA strand. The sgRNA contains a nucleotide sequence complementary to the nucleotide sequence of DNA adjacent to the region of interest (such as a region potentially with sequence variants). The polymerase is then bound to the cleavage site. The polymerase / DNA complex is then associated with a nanopore, and the DNA can be sequenced using rolling circle amplification and transcription.
[0270] Figure 4B schematically illustrates an example method for genome sequencing using nanopore sequencing. This method may include the following steps: providing a sample containing genomic DNA; amplifying the genomic DNA; circularizing the genomic DNA to provide circular DNA; cleaving the circular DNA with a nicking enzyme to provide nick sites on the strand of the circular DNA; binding a DNA polymerase to the nick sites; and amplifying and sequencing the circular DNA using nanopores. In some instances, restriction nicking enzymes can be used to generate nicks at their respective recognition sequences, thus allowing control over the distance between the nick and the target site (e.g., a mutation site). Alternatively, nicking enzymes (e.g., Cas9n complexes) can be used to target a specific sequence of interest to generate a nick in or near that specific sequence, thereby controlling the distance between the nick and the target site (e.g., a mutation site). Nick siteThe dot can be a portion of single-stranded DNA that has been removed to expose the 3′ and 5′ ends. The 3′ end can be used as a template that polymerase can bind to and amplify from.
[0271] Sample
[0272] The sample used for analysis can contain multiple polynucleotides. The polynucleotide can be single-stranded DNA, double-stranded DNA, or a combination thereof. The polynucleotide can contain genomic DNA, genomic cDNA, cell-free DNA, cell-free cDNA, or any combination thereof.
[0273] The polynucleotide can include cell-free DNA, circulating tumor DNA, genomic DNA, and DNA from formalin-fixed and paraffin-embedded (FFPE) samples. In some instances, the DNA extracted from FFPE samples may be damaged, and such damaged DNA can be repaired using an available FFPE DNA repair kit. The sample can contain any suitable DNA and / or cDNA sample, such as urine, feces, blood, saliva, tissue, biopsy, body fluid, or tumor cells.
[0274] The multiple polynucleotides can be single-stranded or double-stranded.
[0275] The polynucleotide sample can be from any suitable source. For example, the sample can be obtained from a patient, animal, plant, or environment, such as naturally occurring or artificial atmosphere, water system, soil, atmospheric pathogen collection system, underground sediment, groundwater, or sewage treatment plant.
[0276] The polynucleotide from the sample may include one or more different polynucleotides, such as DNA, RNA, ribosomal RNA (rRNA), transfer RNA (tRNA), microRNA (miRNA), messenger RNA (mRNA), fragments of any of the foregoing, or any combination thereof. The sample may contain DNA. The sample may contain genomic DNA. The sample may contain mitochondrial DNA, chloroplast DNA, plasmid DNA, bacterial artificial chromosome, yeast artificial chromosome, oligonucleotide tag, or any combination thereof.
[0277] The polynucleotide may be single-stranded, double-stranded, or a combination thereof. The polynucleotide may be a single-stranded polynucleotide, wherein double-stranded polynucleotides may or may not be present.
[0278] The initial amount of polynucleotides in the sample can be, for example, less than 50 ng, such as less than 45 ng, 40 ng, 35 ng, 30 ng, 25 ng, 20 ng, 15 ng, 10 ng, 5 ng, 4 ng, 3 ng, 2 ng, 1 ng, 0.5 ng, 0.1 ng, or less. The initial amount of polynucleotides in the sample can be, for example, greater than 0.1 ng, such as greater than 0.5 ng, 1 ng, 2 ng, 3 ng, 4 ng, 5 ng, 10 ng, 15 ng, 20 ng, 25 ng,30 ng, 35 ng, 40 ng, 45 ng, 50 ng or more. The amount of starting polynucleotide can be, for example, 0.1 ng to 100 ng, 1 ng to 75 ng, 5 ng to 50 ng or 10 ng to 20 ng.
[0279] The polynucleotides in the sample may be single-stranded when obtained or may become single-stranded through treatment (e.g., denaturation). Further examples of suitable polynucleotides are described herein with respect to various aspects of this disclosure. Polynucleotides may be subjected to subsequent steps (e.g., cyclization and amplification) without an extraction step and / or a purification step. For example, a fluid sample may be treated to remove cells without an extraction step to produce a purified liquid sample and a cell sample, and then the polynucleotides may be isolated from the purified fluid sample. A variety of methods may be used to isolate polynucleotides, such as by precipitation or nonspecific binding to a substrate, followed by washing the substrate to release the bound polynucleotides. If polynucleotides are isolated from the sample without a cell extraction step, the polynucleotides will be primarily extracellular or “cell-free” polynucleotides, which may correspond to dead or damaged cells. The identity of such cells can be used to characterize, for example, the cells or cell populations from which they originate in a microbial community.
[0280] The sample may be from a subject. The subject may be any suitable organism, including, for example, plants, animals, fungi, protozoa, nucleolytic protozoa, viruses, mitochondria, and chloroplasts. The sample polynucleotides may be isolated from the subject, such as cell samples, tissue samples, body fluid samples, or organ samples, or cell cultures derived from any of these, including, for example, cultured cell lines, biopsies, blood samples, buccal swabs, or fluid samples containing cells such as saliva. The subject may be an animal, such as a cow, pig, mouse, rat, chicken, cat, dog, or mammal, such as a human. The sample may contain tumor cells, as in a sample from tumor tissue from the subject.
[0281] The sample may not contain intact cells and may be processed to remove cells or to isolate polynucleotides without a cell extraction step, for example to isolate cell-free polynucleotides, such as cell-free DNA.
[0282] Other examples of sample sources include blood, urine, feces, nasal cavity, lungs, intestines, other bodily fluids or excretions, their derivatives, or combinations thereof.
[0283] Samples from a single individual can be divided into multiple separate samples, such as 2, 3, 4, 5, 6, 7, 8, 9, 10, or more separate samples, which are independently subjected to the methods of this disclosure, for example, by performing analyses in duplicate, triplicate, quadruplicate, or more copies. When the sample is from a subject, the reference sequence can also be derived from the subject, such as a common sequence from the analyzed samples or a polynucleotide sequence from another sample or tissue from the same subject. For example, a blood sample can be analyzed...ctDNA mutations can be identified, and cellular DNA from another sample (such as a cheek or skin sample) from the subject can be analyzed to determine a reference sequence.
[0284] Depending on any suitable method, the polynucleotide may or may not be extracted from the cells in the sample.
[0285] The polynucleotides may comprise cell-free polynucleotides, such as cell-free DNA (cfDNA) or circulating tumor DNA (ctDNA). Cell-free DNA circulates in both healthy and diseased individuals. cfDNA (ctDNA) from tumors is not limited to any particular type of cancer but appears to be a common finding in various malignant tumor conditions. The concentration of free circulating DNA in the plasma of control subjects may be lower than that of patients with or suspected of having the disease. In one instance, the concentration of free circulating DNA in plasma may be, for example, 14 ng / mL to 18 ng / mL in control subjects, and 18 ng / mL to 318 ng / mL in tumor-forming patients.
[0286] Apoptosis and necrotic cell death may contribute to the production of cell-free circulating DNA in body fluids. For example, significantly elevated levels of circulating DNA can be observed in the plasma of patients with prostate cancer and other prostate diseases such as benign prostatic hyperplasia and prostatitis. Furthermore, circulating tumor DNA may be present in fluids from the primary organ of the tumor. In one instance, breast cancer detection can be achieved in catheter irrigation fluid; colorectal cancer detection in feces; lung cancer detection in sputum; and prostate cancer detection in urine or ejaculate. Cell-free DNA can be obtained from a variety of sources. An exemplary source could be a blood sample from a subject. However, cfDNA or other fragmented DNA can be derived from a variety of other sources, including, for example, urine and fecal samples, which can be sources of cfDNA including ctDNA.
[0287] The methods for sequencing polynucleotides provided in this disclosure may include retrieving a biological sample having a polynucleotide or set of polynucleotides to be sequenced, extracting or otherwise isolating the polynucleotide sample from the biological sample, and optionally preparing the polynucleotide sample for sequencing.
[0288] Methods for sequencing polynucleotide samples may include isolating polynucleotides from a biological sample (e.g., a tissue sample, a fluid sample) and preparing a polynucleotide sample for sequencing. In some cases, polynucleotide samples are extracted from cells. Examples of techniques for extracting polynucleotides include the use of lysozyme, sonication, extraction, high pressure, or any combination thereof. In some cases, the polynucleotides are cell-free polynucleotides and do not require extraction from cells.
[0289] In some cases, extraction can be achieved by removing proteins, cell wall debris, and other substances from the polynucleotide sample.The process of preparing polynucleotide samples for sequencing involves the composition of components. Many commercial products are available to perform this operation, such as spin columns. Alternatively or additionally, ethanol precipitation and centrifugation may be used.
[0290] Nucleic Acid Fragmentation
[0291] Polynucleotides from the sample may be fragmented prior to further processing. Fragmentation can be accomplished by any suitable method, including chemical, enzymatic, and mechanical fragmentation. The average or median length of the fragment is at least 10 nucleotides. The average or median length of the fragment may be at least 10, 50, 100, 500, 1,000, 5,000, 10,000, 50,000, or more nucleotides. The average or median length of the fragment is at most 50,000, 10,000, 5,000, 1,000, 500, 100, 50, 10, or fewer nucleotides. The fragment may be 90 to 200 nucleotides, and / or have an average length of 150 nucleotides or any other suitable average length. Fragmentation of polynucleotides can be performed mechanically, including subjecting sample polynucleotides to acoustic treatment. Fragmentation can include treating sample polynucleotides with one or more enzymes under conditions suitable for one or more enzymes to generate double-stranded polynucleotide breaks. Examples of fragmenting enzymes include sequence-specific nucleases and non-sequence-specific nucleases. Examples of suitable nucleases include DNase I, fragmentases, restriction endonucleases, variants thereof, and combinations thereof. Fragmentation can include treating sample polynucleotides with one or more restriction endonucleases. Fragmentation can produce fragments having 5' protrusions, 3' protrusions, blunt ends, or combinations thereof. When fragmentation involves the use of one or more restriction endonucleases, the cleavage of sample polynucleotides leaves protrusions with predictable sequences. Fragmented polynucleotides can undergo size-selective fragmentation steps by standard methods, such as column purification or separation from agarose gels.
[0292] Linear Nucleic Acid Amplification
[0293] Polynucleotides in the sample can be amplified. Polynucleotides can be amplified using a variety of methods, such as primer extension reactions using any suitable combination of primers and DNA polymerases, including but not limited to polymerase chain reaction (PCR), reverse transcription, and combinations thereof. When the template for primer extension is RNA, the reverse transcription product is called complementary DNA (cDNA). Primers that can be used for primer extension reactions may contain sequences, random sequences, partially random sequences, and combinations thereof that are specific to one or more targets.
[0294] The amplified polynucleotides can be sequenced with or without enrichment, such as by enriching one or more target polynucleotides in the amplified polynucleotides by performing an enrichment step prior to sequencing. The enrichment step may include hybridizing the amplified polynucleotides with multiple probes attached to a substrate. The enrichment step may include an amplification reaction mixture containing the followingThe amplification process involves a target sequence comprising sequences A and B oriented in a 5' to 3' direction: (a) an amplified polynucleotide; (b) a first primer containing sequence A', wherein the first primer specifically hybridizes to sequence A of the target sequence via sequence complementarity between sequence A and sequence A'; (c) a second primer containing sequence B, wherein the second primer specifically hybridizes to sequence B' present in a complementary polynucleotide containing the complementary sequence of the target sequence via sequence complementarity between B and B'; and (d) a polymerase extending the first primer and the second primer to produce the amplified polynucleotide; wherein the distance between the 5' end of sequence A and the 3' end of sequence B of the target sequence is 75 nt or less.
[0295] Sample Enrichment
[0296] Regions of genomic DNA and cDNA of interest can be selectively targeted and amplified. Targeted regions may include, for example, regions containing sequence variants of interest for diagnostic purposes.
[0297] Polynucleotide regions of interest can be cleaved by binding a biotinylated sgRNA / CRISPR / Cas9 complex to a polynucleotide region adjacent to the region of interest. sgRNA may contain a sequence complementary to a polynucleotide sequence in a region adjacent to the region of interest. The sgRNA / CRISPR / Cas9 complex can cleave double-stranded polynucleotides into polynucleotide segments. One or more purification methods, such as binding the target segment to streptavidin beads, can be used to enrich the composition containing the target polynucleotide segment and the non-target polynucleotide segment with the target polynucleotide segment.
[0298] The polynucleotide mixture enriched with the polynucleotide segment of interest can then be cyclized.
[0299] Linear nucleic acid cyclization
[0300] The polynucleotide sample may contain single-stranded polynucleotides. Cyclicizing such polynucleotides can be achieved by performing a linkage reaction of multiple polynucleotides. The cyclic polynucleotide may have a unique linker in the cyclized polynucleotide.
[0301] Cycling may include linking the 5' end of the polynucleotide to the 3' end of the same polynucleotide, the 3' end of another polynucleotide in the sample, or the 3' end of a polynucleotide from a different source (e.g., an artificial polynucleotide, such as an oligonucleotide adaptor). For example, the 5' end of a polynucleotide can be ligated to the 3' end of the same polynucleotide (also known as "self-ligation"). The conditions of the cyclization reaction can be selected to favor polynucleotide self-ligation within a specific length range, resulting in multiple cyclized polynucleotides characterized by a specific average length. For example, cyclization reaction conditions can be selected to favor polynucleotide self-ligation of nucleotides shorter than 5,000, 2,500, 1,000, 750, 500, 400, 300, 200, 150, 100, 50, or fewer nucleotides. Polynucleotide fragments with lengths of 50 to 5000 nucleotides, 100 to 2500 nucleotides, or 150 to 500 nucleotides can be...Advantageously, the average length of the cyclized polynucleotide is within the desired range. For example, 80% or more of the cyclized polynucleotide fragments may be between 50 and 500 nucleotides in length, such as between 50 and 200 nucleotides in length. Optimizable reaction conditions include the duration of the ligation reaction, the concentrations of various reagents, and / or the concentration of the polynucleotide to be ligated. The cyclization reaction can preserve the distribution of fragment lengths in the sample before cyclization. For example, in the pre-cyclized and cyclized polynucleotides, one or more of the mean, median, modality, and standard deviation of fragment lengths in the sample may be between 75%, 80%, 85%, 90%, 95%, or more of each other.
[0302] One or more adaptor oligonucleotides may be used such that the 5' and 3' ends of the polynucleotides in the sample are linked by one or more adaptor oligonucleotides in between to form a cyclic polynucleotide, rather than forming a self-ligated cyclic polynucleotide. For example, the 5' end of the polynucleotide may be linked to the 3' end of the adaptor, and the 5' end of the same adaptor may be linked to the 3' end of the same polynucleotide. The adaptor oligonucleotide may include any suitable oligonucleotide having a sequence as specified in the specification 45 / 69, page 47, CN 121472373 A, at least a portion of which is known and can be linked to the sample polynucleotide. The adaptor oligonucleotide may include, for example, DNA, RNA, nucleotide analogs, non-canonical nucleotides, labeled nucleotides, modified nucleotides, or any combination thereof. The adaptor oligonucleotide may be single-stranded, double-stranded, or partially double-stranded. A partially double-stranded adaptor may contain one or more single-stranded regions and one or more double-stranded regions. A double-stranded adaptor may contain two separate oligonucleotides hybridizing to each other, such as an oligonucleotide duplex, and hybridization may leave one or more blunt ends, one or more 3' overhangs, one or more 5' overhangs, one or more protrusions caused by mismatched and / or unpaired nucleotides, or any combination thereof. When the two hybridized regions of the adaptor are separated from each other by non-hybridized regions, a "bubble" structure is formed. Adaptors with different nucleotide sequences may be used. Different adaptors may be linked to the sample polynucleotide in a sequential reaction or simultaneously. The same adaptor may be added to both ends of the target polynucleotide. For example, a first and a second adaptor can be added to the same reaction. The adaptor can be manipulated before binding to the sample polynucleotide. For example, a terminal phosphate ester can be added or removed.
[0303] Any suitable method can be used for cyclization of polynucleotides. For example, cyclization can include enzymatic reactions, such as using ligases, such as RNA or DNA ligases. Examples of suitable ligases include Circligase™ (Epicentre; Madison, Wis.), RNA ligase, T4 ligase, etc.RNA ligase 1 (ssRNA ligase), NAD-dependent ligases such as Taq DNA ligase, *Thermus filiformis* DNA ligase, *Escherichia coli* DNA ligase, Tth DNA ligase, *Thermus scotoductus* DNA ligase (I and II), thermostable ligases, Ampligase thermostable DNA ligase, VanC-type ligase, 9°N DNA ligase, Tsp DNA ligase, ATP-dependent ligases such as T4 RNA ligase, T4 DNA ligase, T3 DNA ligase, T7 DNA ligase, Pfu DNA ligase, DNA ligase 1, DNA ligase III, and DNA ligase IV.
[0304] For self-ligation cyclization, the concentrations of polynucleotides and enzymes can be adjusted to promote the formation of intramolecular loops rather than intermolecular structures. The reaction temperature and time can also be adjusted. An exonuclease step can be included after the cyclization reaction to digest any unligated polynucleotides. For example, cyclic polynucleotides do not contain a free 5' or 3' end, so the introduction of a 5' or 3' exonuclease will not digest the closed loop, but will digest the unlinked polynucleotide.
[0305] The formation of a cyclic polynucleotide by linking the two ends of a polynucleotide directly or through one or more intermediate adaptor oligonucleotides can produce a linker with a linker sequence. In the case where the 5' and 3' ends of a polynucleotide are linked by an adaptor polynucleotide, the linker can refer to the linker between the polynucleotide and the adaptor (e.g., one of the 5' end linker or the 3' end linker), or to the linker formed by and including the adaptor polynucleotide between the 5' and 3' ends of the polynucleotide. If the 5' and 3' ends of a polynucleotide are linked without an intermediate adaptor, such as the 5' and 3' ends of a single-stranded DNA, the linker can refer to the point where these two ends are joined. The linker can be identified by the nucleotide sequence containing the linker, i.e., the linker sequence. The sample contains polynucleotides with a terminal mixture formed through natural degradation processes (such as cell lysis, cell death, and others), which release DNA from the cell into its surrounding environment where it is further degraded, such as into cell-free polynucleotides. Fragmentation is a byproduct of sample processing (such as fixation, staining, and / or storage processes), and fragmentation can be performed by methods that cut DNA without being restricted to specific target sequences, such as mechanical fragmentation or non-sequence-specific nuclease treatment, such as using DNase I or fragmentase. In cases where the sample may contain polynucleotides with a terminal mixture, the likelihood of two polynucleotides having the same 5' or 3' end is low, and the likelihood of two polynucleotides independently having the same 5' and 3' ends is also low.The probability is extremely low. In such mixtures, even if two polynucleotides contain portions with the same target sequence, the linker can be used to distinguish different polynucleotides. If the polynucleotide ends are joined in the absence of an intervening adaptor, the linker sequence can be identified by alignment with a reference sequence. For example, in cases where the order of the two component sequences appears to be reversed relative to the sequence in Reference Specification 46 / 69, page 48, CN 121472373 A, the point where the reversal occurs can indicate the linker. In cases where the polynucleotide ends are joined by one or more adaptor sequences, the linker can be identified by proximity to a known adaptor sequence, or by alignment as described above if the sequencing read is long enough to obtain the sequence from the 5' and 3' ends of the circular polynucleotide. The formation of a particular linker can be a very rare event, making it unique among the circular polynucleotides in the sample.
[0306] Circular Nucleic Acid Amplification
[0307] The methods provided in this disclosure may include amplifying circular polynucleotides. For example, multiple different circular polynucleotides containing a target sequence can be amplified, wherein the target sequence comprises sequence A and sequence B oriented in a 5' to 3' direction. A method for amplifying circular DNA may include subjecting the circular DNA to a polynucleotide amplification reaction, wherein the reaction amplification reaction mixture comprises (a) multiple circular polynucleotides, wherein individual circular polynucleotides among the multiple circular polynucleotides comprise different linkers formed by circularizing individual polynucleotides having 5' and 3' ends; (b) a first primer containing sequence A', wherein the first primer specifically hybridizes to sequence A of the target sequence via sequence complementarity between sequence A and sequence A'; (c) a second primer containing sequence B, wherein the second primer specifically hybridizes to sequence B' present in a complementary polynucleotide containing a complementary sequence of the target sequence via sequence complementarity between sequence B and B'; and (d) a polymerase extending the first and second primers to produce the amplified polynucleotide; wherein sequences A and B are endogenous sequences, and the distance between the 5' end of sequence A and the 3' end of sequence B of the target sequence is 75 nucleotides or less.
[0308] After cyclization, cyclic double-stranded polynucleotides can be amplified. There are various methods for amplifying cyclic polynucleotides (e.g., DNA and / or RNA). Amplification can be linear, exponential, or may include both linear and exponential phases in a polyphase amplification process. Amplification methods can involve temperature changes, such as a thermal denaturation step, or can be isothermal processes that do not require thermal denaturation. Examples of suitable amplification processes include rolling circle amplification (RCA). In RCA, the reaction mixture can contain one or more primers, a polymerase, and dNTPs, and produce tandem strands. The polymerase in the RCA reaction can be a polymerase with strand displacement activity.Various suitable polymerases are available, including, for example, DNA polymerase I (Klenow) fragment lacking exonuclease activity, Phi29 DNA polymerase, and Taq DNA polymerase. As a result of RCA, the formed tandem polynucleotide amplification product has two or more copies of the target sequence from the template polynucleotide, such as 2, 3, 4, 5, 6, 7, 8, 9, 10, or more copies of the target sequence. Amplification primers can have any suitable length, for example, at least 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 90, 100, or more nucleotides, any part or all of which can be complementary to the corresponding target sequence hybridized to by the primers. The RCA process can use, for example, random primers, target-specific primers, adaptor-targeting primers, or no primers.
[0309] Purification
[0310] After cyclizing polynucleotides such as double-stranded DNA to provide circular dsDNA, the circular dsDNA can be purified prior to amplification or sequencing, for example, by isolating the circular polynucleotides or removing one or more other molecules from the reaction, to increase the relative concentration or purity of the circular polynucleotides for subsequent steps. For example, the mixture containing double-stranded polynucleotides can be treated with an exonuclease to remove the non-circularized polynucleotides (e.g., single-stranded or double-stranded), or size exclusion chromatography can be performed to trap and discard small reagents, or to trap and release the cyclization product in a separate volume. Purification may include treatment to remove or degrade the ligase used in the cyclization reaction, and / or to purify the cyclized polynucleotides from such ligase. Treatment to degrade the ligase may include treatment with a protease.
[0311] Chapped Section
[0312] A mixture of circular polynucleotides, such as circular DNA, circular cDNA, circular ctDNA, or a combination thereof, can react with a cleaving enzyme (e.g., a CRISPR / Cas9n complex) to cleave segments of the single strand of the circular polynucleotide, thereby forming a cleavage in the chain of the polynucleotide. For circular double-stranded polynucleotides, the CRISPR / Cas9n complex can bind to and cleave segments of the inner or outer strand.
[0313] The CRISPR / Cas9n cleaving enzyme complex can be used to remove segments, i.e., to form cleavages on polynucleotide portions targeted by at least a portion of the gRNA of the CRISPR / Cas9n complex. This approach can be applied to whole-genome sequencing and / or targeted sequencing. In some instances, the CRISPR-Cas9n system can be adapted to target specific nucleotide sequences by complexing with a short RNA molecule (called a short guide RNA (sgRNA)) that recognizes a specific DNA target. To target specific multicores of interestThe nucleotide region can be used to bind a nicking enzyme to a polynucleotide region adjacent to the region of interest using an sgRNA / CRISPR / Cas9n complex. The nicking enzyme can expose the 3′ and 5′ ends of the circular DNA, and the 3′ end can be used as a binding site for amplification and / or reverse transcription sequencing (e.g., using a polymerase, such as DNA polymerase).
[0314] Therapeutic Applications
[0315] The methods, systems, and compositions provided herein can be targeted at one or more therapeutic applications, such as for characterizing patient samples and optionally diagnosing the condition of a subject. Therapeutic applications may include informing a patient of the treatment options that may be most responsive to them and / or informing a subject requiring treatment intervention based on the results of the methods provided herein.
[0316] For example, the methods provided herein can be used to diagnose the presence, progression, and / or metastasis of a tumor, such as when the polynucleotide analyzed contains or is composed of cfDNA, ctDNA, or fragmented tumor DNA. For example, the efficacy of tumor treatment in a subject can be monitored by monitoring ctDNA over time, a decrease in ctDNA can be used as an indicator of treatment efficacy, and an increase in ctDNA can indicate the selection of different treatments and / or different doses. Other uses include assessing organ rejection in transplant recipients, such as using an increase in the amount of circulating DNA corresponding to the transplant donor's genome as an early indicator of transplant rejection, and genotyping / isotyping of pathogen infections (such as viral or bacterial infections). Detection of sequence variants in circulating fetal DNA can be used to diagnose fetal conditions.
[0317] The methods provided in this disclosure may include diagnosing a subject based on sequencing results, such as diagnosing the subject with a disease associated with a detected causal genetic variant, or reporting the likelihood that a patient has or will develop such a disease.
[0318] Causal genetic variants may include sequence variants associated with a specific type or stage of cancer or cancer with specific characteristics (such as metastatic potential, drug resistance, and / or drug responsiveness). The methods provided in this disclosure may be used to inform, guide, and monitor treatment decisions for cancer treatment. For example, treatment efficacy can be monitored by comparing ctDNA samples from patients before, during, and after treatment, including specific molecularly targeted therapies such as monoclonal drugs, chemotherapy drugs, radiation regimens, and combinations of any of the foregoing methods. For example, ctDNA can be monitored to see if certain mutations increase or decrease after treatment, or if new mutations appear, which allows doctors to modify treatment plans in a shorter time than monitoring methods that track patient symptoms. Methods can include diagnosing subjects based on the results of multinucleotide sequencing, such as diagnosing a subject with a specific stage or type of cancer associated with a detected sequence variant, or reporting the likelihood that a patient has or will develop such cancer.
[0319] For example, for therapies based on molecular markers specifically targeting patients, patients can be tested to identify their...The presence of certain mutations in a tumor, and these mutations can be used to predict response to therapy or resistance, and guide the decision to use the therapy. Detection and monitoring of ctDNA during treatment helps guide treatment selection.
[0320] Sequence variants associated with one or more cancers can be used for diagnostic, prognostic, or treatment decisions. For example, suitable target sequences with oncological significance include alterations in the TP53 gene, ALK gene, KRAS gene, PIK3CA gene, BRAF gene, EGFR gene, and KIT gene. Target sequences can be specifically amplified, and / or sequence variants of target sequences can be specifically analyzed to determine whether they are likely to be all or part of a cancer-related gene.
[0321] The methods provided in this disclosure can be used to discover novel rare mutations associated with one or more cancer types, stages, or cancer characteristics. For example, in a population of individuals sharing the characteristics to be analyzed, such as a specific disease, cancer type, and / or cancer stage, the methods provided in this disclosure can be used to identify sequence variants reflecting mutations in a specific gene or part of a gene. Identified sequence variants that occur at a statistically significantly higher frequency in a group of individuals sharing the trait compared to individuals without the trait can have an association with the trait. The sequence variants or types of sequence variants thus identified can then be used to diagnose or treat individuals found to carry them.
[0322] Additional therapeutic applications may include use in noninvasive fetal diagnostics. Fetal DNA can be found in the blood of pregnant women. The methods provided in this disclosure can be used to identify sequence variants in circulating fetal DNA and therefore can be used to diagnose one or more genetic diseases in the fetus, such as diseases associated with one or more causal genetic variants. Examples of causal genetic variants include trisomy, cystic fibrosis, sickle cell anemia, and Tay-Saks disease. The mother can provide a control sample and a blood sample for comparison. The control sample can be any suitable tissue, which can then be sequenced to provide a reference sequence. The cfDNA sequence corresponding to the fetal genomic DNA can then be identified as a sequence variant relative to the maternal reference. The father can also provide a reference sample to aid in the identification of the fetal sequence and sequence variants.
[0323] Different therapeutic applications may include the detection of exogenous polynucleotides, including detection from pathogens such as bacteria, viruses, fungi, and microorganisms, which can indicate treatment.
[0324] Sequencing equipment
[0325] Impedance-based sequencing
[0326] In one aspect, this disclosure provides methods for processing or analyzing nucleic acid molecules. The method may include providing a nucleic acid molecule adjacent to a nanopore. The method may include contacting the nucleic acid molecule with a tagged nucleotide under conditions sufficient to incorporate a nucleotide into a nucleic acid chain complementary to at least a portion of the nucleic acid molecule. When the nucleotide is incorporated into the nucleus...When the tag is in the acid chain, at least a portion of the tag may be placed within the nanopore. The method may include detecting one or more signals indicating impedance or impedance changes in the nanopore when at least a portion of the tag is within the nanopore. The method may include using one or more signals to identify nucleotides incorporated into the nucleic acid chain. The method may also include measuring current or changes thereof when at least a portion of the tag is placed within the nanopore.
[0327] The nanopore may be disposed adjacent to or near an electrode of a sensing circuit or coupled to a circuit (e.g., a CMOS or FET circuit). The circuit may be coupled to a voltage source. Alternatively, the nanopore may be part of a circuit. A constant voltage may be applied to the circuit, and changes in current may be measured. Alternatively, changes in voltage required to maintain a steady-state current may be measured. The nanopore may be part of a circuit containing a tunnel junction. The nanopore may be a tunnel junction of the nanopore. Alternatively, the nanopore may not be part of a circuit containing a tunnel junction. The nanopore may be in an electrolyte solution (e.g., 0.5 M potassium acetate and 10 mM KCl). Alternatively, the nanopore may not be in an electrolyte solution.
[0328] One or more signals may be current or voltage measured from the sensing circuit. One or more signals can be current and voltage measured from sensing circuitry. The signal can be a tunneling current. Alternatively, the signal may not be a tunneling current. The current can be a Faraday current. Alternatively, the current may not be a Faraday current. The current can be at least 1 picoampere (pA), 10 pA, 100 pA, 1 nanoampere (nA), 10 nA, 100 nA, 1 microampere (mA), 10 mA, 100 mA, or higher. The current can be at most 100 mA, 10 mA, 1 mA, 100 nA, 10 nA, 1 nA, 100 pA, 10 pA, 1 pA, or smaller. The current can be at least in the picoampere (pA) range, in the tens of pA range, in the hundreds of pA range, in the nanoampere (nA) range, in the tens of nA range, in the hundreds of nA range, in the microampere (mA) range, in the tens of mA range, or higher. The current can be in the range of tens of mA, mA, hundreds of nA, tens of nA, nA, hundreds of pA, tens of pA, pA or lower. The voltage can be at least 0.1 mV, 0.5 mV, 1 mV, 5 mV, 10 mV, 50 mV, 100 mV, 500 mV or higher. The voltage can be at most 500 mV, 100 mV, 50 mV, 10 mV, 5 mV, 1 mV, 0.5 mV, 0.1 mV or lower. The voltage can be at least in the range of millivolts (mV), tens of mV, hundreds of mV or higher. The voltage can be at most hundreds of mV, tens of mV, mV or lower.
[0329] The circuit may include multiple electrodes (e.g., metal electrodes). The circuit may contain at least 2, 3, 4, 5, 6, 7, 8, 9,Ten or more electrodes. The circuit may contain up to 10, 9, 8, 7, 6, 5, 4, 3, or 2 electrodes. Multiple electrodes may not be in direct contact with the nanopore. Alternatively, multiple electrodes may be in direct contact with the nanopore. In another alternative, some electrodes may be in direct contact with the nanopore, while others may not. In some cases, the nanopore may contain multiple electrodes. Alternatively, the nanopore may not contain multiple electrodes. Identifying nucleotides incorporated into a nucleic acid chain may include using multiple electrodes to detect one or more signals.
[0330] The nanopore may contain protein nanopores or solid nanopores. The nanopore may contain protein nanopores and solid nanopores. The nanopore may have, for example, a characteristic width or diameter of about 0.1 nm to 1,000 nm. The width or diameter of the nanopore may be at least 0.1 nm, 0.5 nm, 1 nm, 5 nm, 10 nm, 50 nm, 100 nm, 500 nm, 1,000 nm, or greater. The width or diameter of the nanopore can be up to 1,000 nm, 500 nm, 100 nm, 50 nm, 10 nm, 5 nm, 1 nm, 0.5 nm, 0.1 nm or smaller.
[0331] The method may further include releasing a tag from the nucleotide during incorporation into the nucleic acid chain before detecting one or more signals indicative of impedance or impedance change in the nanopore. At least one enzyme may incorporate the nucleotide into a nucleic acid chain complementary to at least a portion of the nucleic acid molecule. At least one enzyme may further release the tag from the nucleotide before, during or after incorporation. Alternatively or additionally, an additional enzyme (which may be operatively coupled to at least one enzyme) may release the tag from the nucleotide before, during or after incorporation. At least a portion of the released tag may enter the nanopore, and the method may include detecting one or more signals indicative of impedance or impedance change in the nanopore while at least a portion of the released tag is within the nanopore.
[0332] The nucleic acid molecule includes a circular nucleic acid molecule. The circular nucleic acid molecule may be single-stranded or double-stranded. Alternatively, the first part of the circular nucleic acid molecule may be single-stranded, and the second part of the circular nucleic acid molecule may be double-stranded. In some instances, the nucleic acid molecule may include single-stranded or double-stranded linear nucleic acid molecules.
[0333] Incorporation may be performed using at least one oligonucleotide primer. Alternatively, incorporation may be performed without the use of an oligonucleotide primer. The method may also include subjecting the nucleic acid molecule to RCA to generate a nucleic acid chain before detecting one or more signals indicating impedance or impedance changes in the nanopore. RCA can generate a nucleic acid chain containing at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more copies of at least a portion of the nucleic acid molecule (e.g., at least a portion of at least one strand of the nucleic acid molecule). RCA can generate a nucleic acid chain containing at most 10,9, 8, 7, 6, 5, 4, 3, 2, or 1 copy.
[0334] The provision may include coupling at least one enzyme that performs incorporation into (i) at least a portion of a nanopore or (ii) a membrane having nanopores. The provision may include coupling at least one enzyme that performs incorporation into (i) at least a portion of a nanopore and (ii) a membrane having nanopores. Coupling may include conjugating at least one enzyme to a nanopore or a membrane. Coupling may include conjugating at least one enzyme to a nanopore and a membrane. Coupling may be covalent (e.g., cross-linked or conjugated). Coupling may be performed by another enzyme, such as transglutaminase, sorting enzyme, subtilisin, tyrosinase, laccase, etc., or by a chemical cross-linking agent, such as 1-ethyl-3-(3-dimethylaminopropyl)carbodiimide (EDC), N,N′-dicyclohexylcarbodiimide (DCC), N,N′-diisopropylcarbodiimide (DIC), etc. Alternatively or additionally, the coupling may be non-covalent, for example, through hydrogen bonds, magnetic interactions, etc.
[0335] The membrane may be a lipid bilayer. The membrane may be a solid membrane (e.g., a thin film). The membrane may be a combination of a lipid bilayer and a solid membrane. At least one enzyme may be a polymerase, a nuclease, a functional variant thereof, or a combination thereof.
[0336] In another aspect, this disclosure provides a system for processing or analyzing nucleic acid molecules. The system may include a nanopore configured to receive at least a portion of a tag when a tagged nucleotide is incorporated into a nucleic acid chain. The nucleic acid chain may be complementary to at least a portion of the nucleic acid molecule. When at least a portion of the tag is within the nanopore, the nanopore may be configured to detect one or more signals indicating impedance or impedance changes in the nanopore. One or more signals may be used to identify the nucleotide incorporated into the nucleic acid chain. The system for processing or analyzing nucleic acid molecules may be configured to perform one or more of the subject methods provided in this disclosure for processing or analyzing nucleic acid molecules or derivatives thereof.
[0337] One or more signals may be current or voltage. One or more signals may be current and voltage. The current may be a Faraday current. Alternatively, the current may not be a Faraday current. One or more signals may not be a tunneling current. Alternatively, one or more signals may be a tunneling current. The nanopore may be part of a circuit containing a tunnel junction. Alternatively, the nanopore may not be part of a circuit containing a tunnel junction.
[0338] The nanopore may be configured to measure current or changes thereof when at least a portion of the tag is placed within the nanopore. Alternatively, the nanopore may be configured to measure current or changes thereof when at least a portion of the tag is released from within the nanopore. The nanopore may include a plurality of electrodes configured to detect one or more signals. Alternatively, the nanopore may not include a plurality of electrodes configured to detect one or more signals, and the plurality of electrodes may be operatively coupled to the nanopore to detect one or more signals.Multiple signals.
[0339] The nanopore may comprise a protein nanopore or a solid nanopore. The nanopore may comprise both protein nanopores and solid nanopores.
[0340] The system may also comprise at least one enzyme configured to perform incorporation. At least one enzyme may incorporate a nucleotide into a nucleic acid chain complementary to at least a portion of a nucleic acid molecule. When the nucleotide is incorporated into the nucleic acid chain, at least a portion of the tag may be released from the nucleotide. At least one enzyme may release the tag from the nucleotide before, during, or after incorporation. Alternatively or additionally, an additional enzyme (which may be operatively coupled to at least one enzyme) may release the tag from the nucleotide before, during, or after incorporation. At least a portion of the released tag may enter the nanopore, and the method may include detecting one or more signals indicating impedance or impedance changes in the nanopore while at least a portion of the released tag is within the nanopore. At least one enzyme may be a polymerase, a nuclease, a functional variant thereof, or a combination thereof.
[0341] Incorporation may be performed using at least one oligonucleotide primer. Therefore, the system may also comprise at least one oligonucleotide primer. Alternatively, incorporation may be performed without the use of an oligonucleotide primer. Therefore, the system may not contain oligonucleotide primers.
[0342] At least one enzyme (and / or additional enzymes) may be coupled to at least a portion of (i) a nanopore or (ii) a membrane having nanopores. At least one enzyme (and / or additional enzymes) may be coupled to at least a portion of (i) a nanopore and (ii) a membrane having nanopores. At least one enzyme (and / or additional enzymes) may be conjugated to at least a portion of (i) a nanopore or (ii) a membrane having nanopores. At least one enzyme (and / or additional enzymes) may be conjugated to at least a portion of (i) a nanopore and (ii) a membrane having nanopores. The membrane may be a lipid bilayer or a solid membrane. The membrane may be a lipid bilayer and a solid membrane.
[0343] Coupling may be performed by a coupling enzyme and / or a chemical cross-linking agent. The system may also contain a coupling enzyme (e.g., transglutaminase, sorting enzyme, subtilisin, tyrosinase, laccase, etc.) or a chemical cross-linking agent (e.g., EDC, DCC, DIC, etc.). The system may comprise a coupling enzyme and a chemical cross-linking agent. Alternatively, the nanopore or membrane may be configured to bind to at least a portion of at least one enzyme (and / or additional enzymes). The nanopore or membrane may comprise a binding portion (e.g., a small molecule, nucleotide, peptide, polymer, combination thereof, etc.) capable of binding to at least a portion of at least one enzyme. In various alternatives, at least one enzyme may be configured to bind to at least a portion of the nanopore or at least a portion of the membrane. At least one enzyme may comprise a binding portion (e.g., a small molecule, nucleotide, peptide, polymer, combination thereof, etc.) capable of binding to at least a portion of the membrane. Specification 51 / 69 pages 53 CN 121472373 A
[0344] Figures 5A through 5D schematically illustrate exemplary nanopore sequencing systems for obtaining sequence information from one or more nucleic acid samples. Referring to Figure 5A, the nanopore sequencing system 510 may include a membrane 512 comprising at least one nanopore 514 (a cross-section of the nanopore 514 is shown). The membrane 512 may be a lipid bilayer and / or a solid membrane. The nanopore 514 may include a plurality of electrodes 516 configured to detect one or more signals from a circuit comprising the nanopore 514. The plurality of electrodes 516 may be disposed on one side of the membrane 512. The plurality of electrodes 516 may be coupled to the nanopore 514. The circuit may also include an ammeter and a voltage source. A nucleic acid molecule 520 may be provided in proximity to the nanopore 514. By using an enzyme 530 (e.g., a polymerase), the nucleic acid molecule 520 may be contacted with a nucleotide 541 having a tag 542, provided that conditions are sufficient to incorporate a nucleotide 541 into a nucleic acid chain 540 complementary to at least a portion of the nucleic acid molecule 520. For example, different types of nucleotides may have different tags 542, 544, 546, and 548, respectively. When nucleotide 541 is incorporated into nucleic acid chain 540, at least a portion of tag 542 may be placed within nanopore 514. When at least a portion of tag 542 is within nanopore, one or more signals indicative of impedance or impedance changes in nanopore 514 may be detected. One or more signals may include current or changes thereof. In this example, when tag 542 is placed within nanopore 514, tag 542 may be attached to nucleic acid chain 540. One or more signals may be used to identify nucleotide 541 incorporated into nucleic acid chain 540.
[0345] Referring to FIG5B, nanopore sequencing system 510 may include membrane 512 comprising at least one nanopore 514 (a cross-section of nanopore 514 is shown). Membrane 512 may be a lipid bilayer and / or a solid membrane. Nanopore 514 may include a plurality of electrodes 516 configured to detect one or more signals from circuitry comprising nanopore 514. Multiple electrodes 516 can be coupled to the nanopore 514. The circuit may also include an ammeter and a voltage source. A nucleic acid molecule 520 can be provided near the nanopore 514. By using an enzyme 530 (e.g., a polymerase), the nucleic acid molecule 520 can be contacted with a nucleotide 541 having a tag 542, under conditions sufficient to incorporate a nucleotide 541 into a nucleic acid chain 540 complementary to at least a portion of the nucleic acid molecule 520. For example, different types of nucleotides can have different tags 542, 544, 546, and 548, respectively. When the nucleotide 541 is incorporated into the nucleic acid chain 540, the tag 542 can be released from the nucleotide 541, and at least a portion of the released tag 542 can be placed within the nanopore 514. When at least a portion of the released tag 542 is within the nanopore, the tag 542 in the nanopore 514 can be detected.One or more signals indicating impedance or impedance change. One or more signals can be used to identify nucleotides 541 incorporated into nucleic acid strand 540.
[0346] Referring to FIG5C, nanopore sequencing system 510 may include a membrane 512 comprising at least one nanopore 514 (a cross-section of nanopore 514 is shown). Membrane 512 may be a lipid bilayer and / or a solid membrane. Nanopore 514 may be operatively coupled to a plurality of electrodes 516 configured to detect one or more signals across membrane 512 from a circuit. Nanopore sequencing system 512 may be in an electrolytic solution. The circuit may also include an ammeter and a voltage source. Nucleic acid molecule 520 may be provided in the vicinity of nanopore 514. By using an enzyme 530 (e.g., polymerase), the nucleic acid molecule 520 may be contacted with nucleotide 541 having tag 542 under conditions sufficient to incorporate nucleotide 541 into nucleic acid strand 540 complementary to at least a portion of nucleic acid molecule 520. For example, different types of nucleotides may have different tags 542, 544, 546, and 548, respectively. When nucleotide 541 is incorporated into nucleic acid chain 540, at least a portion of tag 542 may be placed within nanopore 514. When at least a portion of tag 542 is within nanopore, one or more signals indicative of impedance or impedance change in nanopore 514 may be detected. One or more signals may include current or a change thereof. In this example, when tag 542 is placed within nanopore 514, tag 542 may be attached to nucleic acid chain 540. One or more signals may be used to identify nucleotide 541 incorporated into nucleic acid chain 540.
[0347] Referring to FIG5D, nanopore sequencing system 510 may include membrane 512 comprising at least one nanopore 514 (a cross-section of nanopore 514 is shown). The membrane may be a lipid bilayer and / or a solid membrane. Nanopore 514 may include a plurality of electrodes 516 configured to detect one or more signals from circuitry comprising nanopore 514. Multiple electrodes 516 may be disposed on opposite sides of membrane 512, as per specification page 52 / 69, CN 121472373 A. The multiple electrodes 516 may be coupled to nanopore 514. The circuit may also include an ammeter and a voltage source. Nucleic acid molecule 520 may be provided near nanopore 514. By using an enzyme 530 (e.g., polymerase), the nucleic acid molecule 520 may be contacted with a nucleotide 541 having a tag 542, under conditions sufficient to incorporate a nucleotide 541 into a nucleic acid chain 540 complementary to at least a portion of the nucleic acid molecule 520. For example, different types of nucleotides may have different tags 542, 544, 546, and 548, respectively. When incorporating the nucleotide 541 into the nucleic acid chain 540, at least a portion of the tag 542 may be placed in the nanopore.Within the nanopore 514. When at least a portion of the tag 542 is within the nanopore, one or more signals indicative of impedance or impedance changes within the nanopore 514 can be detected. The one or more signals may include current or changes thereof. In this example, when the tag 542 is placed within the nanopore 514, the tag 542 may be attached to the nucleic acid chain 540. One or more signals may be used to identify nucleotides 541 incorporated into the nucleic acid chain 540.
[0348] In the scenarios of Figures 5A, 5B, and 5D, the nanopore 514 may be a solid nanopore, for example, a pore or channel guided through a solid substrate. In the scenario of Figure 5C, the nanopore may be a porin, for example, an α-hemolysin molecule embedded in a lipid bilayer.
[0349] Device Overview
[0350] The methods and systems provided in this disclosure may be performed, and sequencing data may be acquired using any suitable sequencing device, such as a device capable of performing massively parallel sequencing reactions. For example, a high-throughput sequencing system may be used.
[0351] Polynucleotide sequences can be analyzed, for example, to identify repeat unit lengths (e.g., monomer lengths), linkers formed by circularization, and any true variations relative to a reference sequence. Identifying repeat unit lengths can include calculating the regions of repeat units, finding reference sites for the sequence (e.g., when targeting one or more sequences for amplification, enrichment, and / or sequencing), the boundaries of each repeat region, and / or the number of repeats in each sequencing run. Sequence analysis can include analyzing sequence data from both strands of a duplex. For example, identical variants of read sequences from different polynucleotides (e.g., circularized polynucleotides with different linkers) from a sample can be considered confirmed variants. If a sequence variant appears in more than one repeat unit of the same polynucleotide, it can also be considered a true variant, as the same sequence variation is also unlikely to appear at the same position in a target sequence repeated in the same tandem. When identifying variants and confirmed variants, sequence quality scores can be considered; for example, sequences and bases with quality scores below a threshold can be filtered out. Other bioinformatics methods can be used to further improve the sensitivity and specificity of variant determination.
[0352] For example, the system for detecting sequence variants provided in this disclosure may include: (a) a computer configured to receive a user's request to perform a detection reaction on a sample; (b) a polynucleotide preparation system that, in response to a user request, performs a polynucleotide amplification reaction on a sample or a portion thereof, wherein the amplification reaction includes the steps of: (i) cyclizing individual polynucleotides to form a plurality of cyclic polynucleotides, wherein each cyclic polynucleotide has a linker between a 5' end and a 3' end; and (ii) amplifying the cyclic polynucleotides, or amplifying a double-stranded polynucleotide and cyclizing the amplified double-stranded polynucleotide; (iii) creating a nick on a single strand of the cyclic polynucleotide to provide a nick site; and (iv) coupling a polymerase to the nick site.The device comprises: (v) providing a polymerase / cyclic polynucleotide complex; (c) associating the polymerase / cyclic polynucleotide complex with a nanopore; and (d) a sequencing system that generates sequencing reads of the polynucleotide, identifies sequence differences between the sequencing reads and a reference sequence, and determines sequence differences occurring in at least two cyclic polynucleotides with different linkages as sequence variants; and (e) a report generator that sends a report to a recipient containing the results of the detection of sequence variants. In some embodiments, the recipient is a user. In some cases, the sequencing device may include a sensor array for sequencing nucleic acids, such as a nanopore array.
[0353] Each nanopore sequencing complex may be inserted into a membrane, for example, a lipid bilayer, and disposed at a sensing electrode immediately adjacent to or near a sensing circuit (such as an integrated circuit of a nanopore-based sensor). Multiple nanopore sensors may be provided as an array, such as an array present on a chip or biochip, as described on pages 53 / 69 of the specification, CN 121472373 A. The nanopore array may have any suitable number of nanopores. The array may include about 200, about 400, about 600, about 800, about 1000, about 1500, about 2000, about 3000, about 4000, about 5000, about 10000, about 15000, about 20000, about 40000, about 60000, about 80000, about 100000, about 200000, about 400000, about 600000, about 800000, about 100000, about 200000, about 400000, about 600000, about 800000, about 1000000 or more nanopores (or nanopore sequencing complexes).
[0354] During sequencing using one or more labeled nucleotides (or one or more labeled polynucleotides), the labeled nucleotides may be incorporated into each nanopore sequencing complex using an enzyme (e.g., a polymerase). During polymerization, the tag may be detected through the nanopore, such as by releasing and transferring the tag into or through the nanopore, or by presenting it to the nanopore. A single tag can be released and / or presented upon incorporation of a single nucleotide and detected through a nanopore. Multiple tags can be released and / or presented upon incorporation of multiple nucleotides. A nanopore sensor adjacent to (or coupled to) a nanopore can detect a single tag or multiple tags. One or more signals associated with multiple tags can be detected and processed to generate an average signal. Tags can be detected by the sensor as a function of time. Tags detected over time can be used to determine the nucleic acid sequence of a polynucleotide sample, such as by means of a computer system programmed to record sensor data and generate sequence information from that data.
[0355] Any device and system suitable for sequencing by RCA and transcription can be used. In some cases, the sequencing system can generate sequencing reads for polynucleotides amplified by an amplification system, identifying the sequence between the sequencing reads and a reference sequence.Differences are identified, and sequence differences appearing in at least two circular polynucleotides with different linkages are identified as sequence variants. The sequencing system and amplification system may be the same or include one or more overlapping devices. In one instance, the amplification system and sequencing system may utilize the same thermal cycler. A variety of sequencing platforms may be used, and selection may be based on the chosen sequencing method. Amplification and sequencing may involve the use of a liquid processor. Several commercially available liquid handling systems can be used to automate these processes.
[0356] The sequencing system may include, for example, a computer, a computer-readable medium including computer-executable code, a storage device, a communication device, control algorithms, analysis algorithms, and / or reporting algorithms.
[0357] The sequencing device may be used to detect sequence variants. Detection of sequence variants may include the detection of mutations, such as rare somatic mutations relative to a reference sequence or in a mutation-free background, wherein the sequence variant is associated with a disease. Sequence variants with statistical, biological, and / or functional evidence of association with a disease or trait are called “causal genetic variants”. A single causal genetic variant may be associated with more than one disease or trait. Causal genetic variants may be associated with Mendelian traits, non-Mendelian traits, or both. Causal genetic variants can manifest as variations in polynucleotides, such as 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 50 or more sequence differences (e.g., sequence differences at the same relative genomic location between a polynucleotide containing the causal genetic variant and a polynucleotide lacking the causal genetic variant). Examples of causal genetic variant types include single nucleotide polymorphisms (SNPs), deletion / insertion polymorphisms (DIPs), copy number variants (CNVs), short tandem repeats (STRs), restriction fragment length polymorphisms (RFLPs), simple sequence repeats (SSRs), variable number tandem repeats (VNTRs), random amplified polymorphic DNA (RAPDs), amplified fragment length polymorphisms (AFLPs), intertransposon amplification polymorphisms (IRAPs), long and short scattered elements (LINEs / SINEs), long tandem repeats (LTRs), mobile elements, microsatellite amplification polymorphisms of retrotransposons, retrotran-based insertion polymorphisms, sequence-specific amplification polymorphisms, and heritable epigenetic modifications, such as DNA methylation. Causal genetic variants can also be a group of closely related causal genetic variants. Some causal genetic variants can exert their influence as sequence variations in RNA polynucleotides. At this level, some causal genetic variants are also indicated by the presence or absence of a certain RNA polynucleotide. Some causal genetic variants cause sequence variations in protein polypeptides. Many causal genetic variants have been reported. An example of an SNP causal genetic variant is the HbS variant of hemoglobin, which causes sickle cell anemia. DIPAn example of a causal genetic variant is the δ508 mutation in the CFTR gene, which leads to cystic fibrosis. An example of a CNV causal genetic variant is trisomy 21, as described on pages 54 / 69 of the specification, which leads to Down syndrome. An example of a STR causal genetic variant is a tandem repeat sequence, which leads to Huntington's disease.
[0358] Nanopore Devices
[0359] The sequencing system may include a reaction chamber comprising one or more nanopore devices. The nanopore device may be an individually addressable nanopore device. An individually addressable nanopore may be individually readable. An individually addressable nanopore may be individually writable. An individually addressable nanopore may be both individually readable and individually writable. The system may include one or more computer processors for facilitating sample preparation and various operations of this disclosure, such as polynucleotide sequencing. The processor may be coupled to the nanopore device.
[0360] The nanopore device may include a plurality of individually addressable sensing electrodes. Each sensing electrode may include a membrane adjacent to the electrode and one or more nanopores in the membrane. Nanopores can be in membranes such as lipid bilayers, adjacent to or sensorimetrically close to electrodes that are part of or coupled to an integrated circuit. Nanopores can be associated with a single electrode and a sensing integrated circuit or with multiple electrodes and sensing integrated circuits. Nanopores can comprise solid-state nanopores.
[0361] Apparatus and systems used in the methods provided by this disclosure can accurately detect individual nucleotide incorporation events, such as when a nucleotide is incorporated into a growth chain complementary to a template. Enzymes, such as DNA polymerases, RNA polymerases, or ligases, can incorporate nucleotides into a growing polynucleotide chain. Enzymes such as polymerases can generate polynucleotide chains.
[0362] The added nucleotide can be complementary to a corresponding template polynucleotide chain that hybridizes with the growth chain. Nucleotides can include tags or tagging substances coupled to any position of the nucleotide, including but not limited to phosphates of the nucleotide such as γ-phosphates, sugars, or nitrogenous base moieties. In some cases, the tag is detected during nucleotide tag incorporation when the tag associates with the polymerase. The tag is detectable until it translocates through the nanopore after nucleotide incorporation and subsequent tag cleavage and / or release. Nucleotide incorporation events can result in the release of the tag from the nucleotide, which then passes through the nanopore and is detected. The tag can be released by polymerase or by cleavage / release in any suitable manner, including but not limited to cleavage by enzymes located near the polymerase. In this way, the incorporated base (i.e., A, C, G, T, or U) can be identified because a unique tag is released from each type of nucleotide (i.e., adenine, cytosine, guanine, thymine, or uracil). In non-released nucleotide incorporation events, the nanopore is used to detect the bases coupled to the incorporated nucleotide.Tags. In some instances, tags can move through or near nanopores and be detected by means of nanopores.
[0363] The methods and systems of this disclosure can enable the detection of polynucleotide incorporation events, for example, at a resolution of at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 100, 500, 1000, 5000, 10000, 50000, or 100000 polynucleotide bases over a given time period. For example, nanopore devices can be used to detect individual polynucleotide incorporation events, each associated with an individual nucleic acid base. In other instances, nanopore devices can be used to detect events associated with multiple bases. For example, the signal sensed by the nanopore device can be a combination signal from at least 2, 3, 4, or 5 bases.
[0364] In some sequencing methods, the tag does not pass through the nanopore. The tag can be detected through the nanopore and exit the nanopore without passing through it, such as exiting from the opposite direction from where the tag entered the nanopore. The sequencing device can be configured to actively expel the tag from the nanopore.
[0365] In some sequencing methods, the tag is not released after a nucleotide incorporation event. A nucleotide incorporation event can present the tag to the nanopore without releasing it. The tag can be detected by the nanopore without being released. The tag can be attached to a nucleotide with a sufficiently long adapter to present the tag to the nanopore for detection.
[0366] When a nucleotide incorporation event occurs, the nanopore can detect it in real time. An enzyme (such as DNA polymerase) attached to or near the nanopore can facilitate the passage of polynucleotides through or near the nanopore. A nucleotide incorporation event, or the incorporation of multiple nucleotides, can release or present one or more tags that can be detected by the nanopore. Detection can be performed when the tag passes through or near the nanopore, when the tag resides in the nanopore, and / or when the tag is presented to the nanopore. In some cases, an enzyme attached to or near the nanopore can help detect the tag during the incorporation of one or more nucleotides.
[0367] The tag can be an atom, molecule, collection of atoms, or collection of molecules. The tag can provide an optical, electrochemical, magnetic, or electrostatic (e.g., inductive or capacitive) signature that can be detected by means of a nanopore.
[0368] The nanopore can be formed or otherwise embedded in a membrane adjacent to a sensing electrode arrangement of a sensing circuit (e.g., an integrated circuit). The integrated circuit can be an application-specific integrated circuit (ASIC). The integrated circuit can be a field-effect transistor or a complementary metal-oxide-semiconductor (CMOS). The sensing circuit can be located in a chip or other device having nanopores, or outside the chip or device, such as in an off-chip configuration.
[0369] When a nucleic acid or tag passes through or is adjacent to the nanopore, the sensing circuit detects an electrical signal associated with the nucleic acid or tag.Nucleic acids can be subunits of larger chains. Tags can be byproducts of nucleotide incorporation events or other interactions between the tagged nucleic acid and a nanopore or adjacent material (such as an enzyme that cleaves the tag from the nucleic acid). Tags can remain attached to the nucleotide. Detected signals can be collected and stored in a storage location and then used to construct nucleic acid sequences. The collected signals can be processed to interpret any anomalies, such as errors, in the detected signals.
[0370] Nanopores can be used for indirect sequencing of polynucleotides, in some cases by electrical detection. Indirect sequencing can be any method in which nucleotides incorporated into the growing chain do not pass through the nanopore. Polynucleotides can pass through at any suitable distance from and / or near the nanopore, in some cases such that a tag released from a nucleotide incorporation event can be detected in the nanopore.
[0371] Byproducts of nucleotide incorporation events can be detected by nanopores. A nucleotide incorporation event refers to the incorporation of a nucleotide into a growing polynucleotide chain. Byproducts may be associated with the incorporation of a given type of nucleotide. Nucleotide incorporation events can be catalyzed by enzymes (such as DNA polymerase) and the incorporation of available nucleotides at each position can be selected using base pair interactions with the template molecule.
[0372] Nucleic acid samples can be sequenced using labeled nucleotides or nucleotide analogs. In some instances, methods for sequencing nucleic acid molecules include: (a) incorporating (e.g., polymerizing) a labeled nucleotide, wherein a tag associated with an individual nucleotide is released upon incorporation, and (b) detecting the released tag using a nanopore. In some cases, the method further includes guiding a tag attached to or released from an individual nucleotide through a nanopore. The released or attached tag can be guided by any suitable technique, in some cases by means of an enzyme (or molecular motor) and / or a voltage difference across the pore. Alternatively, the released or attached tag can be guided through a nanopore without the use of an enzyme. For example, as described herein, the tag can be guided by a voltage difference across the nanopore.
[0373] The tag can be detected by means of a nanopore device having at least one nanopore in a membrane. The tag can associate with an individual labeled nucleotide during the incorporation of the individual labeled nucleotide. Nanopore devices can detect tags associated with individually labeled nucleotides during incorporation. The labeled nucleotide, whether incorporated into the growing nucleic acid chain or not, can be detected, identified, or distinguished by the nanopore device within a given time period, in some cases by means of the nanopore's electrodes and / or nanopores. The detection time of the nanopore device can be shorter than (in some cases significantly shorter than) the time the tag and / or the nucleotide coupled to the tag is held by an enzyme (such as an enzyme that promotes nucleotide incorporation into the nucleic acid chain, e.g., polymerase). The electrode can detect the tag multiple times during the time the incorporated labeled nucleotide associates with the enzyme. For example, during incorporation...During the time period during which the labeled nucleotide associates with the enzyme, the tag can be detected by the electrode at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 1000, 10,000, 100,000, or 1,000,000 times.
[0374] Sequencing can be performed using a preloaded tag. The preloaded tag may include guiding at least a portion of the tag through at least a portion of the nanopore, while the tag may be attached to a nucleotide that may have been incorporated into a nucleic acid chain (e.g., a growing nucleic acid chain), is being incorporated into a nucleic acid chain, or has not yet been incorporated into a nucleic acid chain but may be incorporated into a nucleic acid chain. Preloaded tags may include guiding at least a portion of the tag through at least a portion of a nanopore before or while the nucleotide is being incorporated into the nucleic acid chain. Preloaded tags may also include guiding at least a portion of the tag through at least a portion of a nanopore after the nucleotide has been incorporated into the nucleic acid chain.
[0375] Tags associated with individual nucleotides can be detected by the nanopore without being released from the nucleotide upon incorporation. Tags can be detected without being released from the incorporated nucleotide during the synthesis of a nucleic acid chain complementary to the target chain. Tags can be attached to nucleotides via adapters such that the tag is presented to the nanopore (e.g., the tag hangs in at least a portion of the nanopore or otherwise extends through at least a portion of the nanopore). The length of the adapter can be long enough to allow the tag to extend into or through at least a portion of the nanopore. In some cases, the tag is presented to (i.e., moved into) the nanopore by a voltage difference. Other methods of presenting the tag into the pore may also be suitable (e.g., using enzymes, magnets, electric fields, voltage differences). In some cases, no active force is applied to the tag (i.e., the tag diffuses into the nanopore).
[0376] A chip for sequencing nucleic acid samples may comprise a plurality of individually addressable nanopores. Each individually addressable nanopore may contain at least one nanopore formed in a membrane disposed adjacent to an integrated circuit. Each individually addressable nanopore is capable of detecting a tag associated with an individual nucleotide. Nucleotides may be incorporated (e.g., polymerized), and the tag may not be released from the nucleotide upon incorporation.
[0377] The tag may be presented to the nanopore and released from the nucleotide after a nucleotide incorporation event. The released tag may pass through the nanopore. In some cases, the tag does not pass through the nanopore. A tag released during a nucleotide incorporation event is distinguished from a tag that may pass through the nanopore but is not fully released during its residence time in the nanopore after the nucleotide incorporation event. In some cases, a tag that has resided in the nanopore for at least 100 milliseconds (ms) is released during a nucleotide incorporation event, and...Tags that remain in the nanopore for less than 100 ms are not released during the nucleotide incorporation event. The tag can be captured and / or guided through the nanopore by a second enzyme or protein (e.g., a nucleic acid-binding protein). The second enzyme can cleave the tag during (e.g., during or after) nucleotide incorporation. The linker between the tag and the nucleotide can be cleaved.
[0378] Based on the residence time of the tag in the nanopore or based on the signal detected by means of the unincorporated nucleotide, tags coupled to the incorporated nucleotide and tags associated with the unincorporated growing complementary strand can be distinguished. The signal generated by the unincorporated nucleotide (e.g., voltage difference, current) can be detectable for a time period between 1 nanosecond (ns) and 100 milliseconds or between 1 ns and 50 ms, while the lifetime of the signal generated by the incorporated nucleotide can be between 50 ms and 500 ms or between 100 ms and 200 ms. The signal generated by the unincorporated nucleotide can be detectable for a time period between 1 ns and 10 ms or between 1 ns and 1 ms. The time period (on average) during which the nanopore can detect unincorporated tags is longer than the time period during which the nanopore can detect incorporated tags.
[0379] The time period during which incorporated nucleic acids can be detected by the nanopore is shorter than that of unincorporated nucleotides. Alternatively, the time period during which incorporated nucleic acids can be detected by the nanopore is longer than that of unincorporated nucleotides. As described herein, the differences and / or proportions between these times can be used to determine whether nucleotides detected by the nanopore have been incorporated.
[0380] The detection time can be based on the free flow of nucleotides through the nanopore; unincorporated nucleotides may remain in or near the nanopore for a time period between 1 nanosecond (ns) and 100 ms or between 1 ns and 50 ms, while incorporated nucleotides may remain in or near the nanopore for a time period between 50 ms and 500 ms or between 100 ms and 200 ms. The time periods may vary depending on the processing conditions; however, the residence time of incorporated nucleotides may be longer than that of unincorporated nucleotides.
[0381] The tag or tag material may include detectable atoms or molecules, or multiple detectable atoms or molecules. The label may include one or more adenine, guanine, cytosine, thymine, uracil, or derivatives thereof, which are attached to any position, including the phosphate group, sugar, or nitrogenous base of the nucleic acid molecule. The label may include one or more adenine, guanine, cytosine, thymine, uracil, or derivatives thereof covalently linked to the phosphate group of a nucleic acid base. Specification 57 / 69 pages 59 CN 121472373 A
[0382] The label may have a size of at least 0.1 nanometers (nm), 1 nm, 2 nm, 3 nm, 4 nm, 5 nm, 6 nm, 7 nm, 8 nm, 9 nm,The length can be 10 nm, 20 nm, 30 nm, 40 nm, 50 nm, 60 nm, 70 nm, 80 nm, 90 nm, 100 nm, 200 nm, 300 nm, 400 nm, 500 nm, or 1000 nm.
[0383] The tag may include the tail of a repeating subunit, such as a plurality of adenine, guanine, cytosine, thymine, uracil, or derivatives thereof. For example, the tag may include the tail of at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 1000, 10,000, or 100,000 subunits of adenine, guanine, cytosine, thymine, uracil, or derivatives thereof. These subunits may be linked to each other and are connected at the ends to phosphate groups of nucleic acids. Other examples of the tag portion include any polymeric material, such as polyethylene glycol (PEG), polysulfonates, amino acids, or any polymer that is wholly or partially positively charged, negatively charged, or uncharged.
[0384] Polymerase
[0385] DNA polymerase can bind to the 3′ end of the cleavage chain of a polynucleotide at a cleavage site. DNA sequencing can be accomplished by using an enzyme (such as DNA polymerase) to amplify and transcribe the polynucleotide in the vicinity of the nanopore and the labeled nucleotide. Sequencing methods may involve incorporating or polymerizing the labeled nucleotide using a polymerase (such as DNA polymerase or transcriptase). Polymerases can be mutated to receive the labeled nucleotide. Polymerases can also be mutated to increase the time for nanopore detection of the tag.
[0386] For example, the sequencing enzyme can be any suitable enzyme that generates a polynucleotide chain through the phosphate bond of the nucleotide. For example, the DNA polymerase may be 9°NmTM polymerase or a variant thereof, E. coli DNA polymerase I, bacteriophage T4 DNA polymerase, sequencer, Taq DNA polymerase, 9°NmTM polymerase (exo-) A485L / Y409V, Φ29 DNA polymerase, Bst DNA polymerase, or a variant, mutant, or homolog of any of the foregoing. Homologs may have any suitable percentage of homology, for example, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, or at least 95% sequence identity.
[0387] In some instances, for nanopore sequencing, the polymerase may be attached to or located near the nanopore. Suitable methods for attaching the polymerase to the nanopore include crosslinking the enzyme to or near the nanopore, such as by forming an intramolecular disulfide bond. The nanopore and the enzyme may also be a fusion, such as a fusion encoded by a single polypeptide chain. Methods for generating fusion proteins may include fusing the enzyme's coding sequence within a frame and adjacent to the coding sequence of the nanopore, and expressing the fusion sequence from a single promoter. Polymerases can be attached to or coupled to nanopores using molecular staples or protein fingers.Pores. Polymerases can be attached to nanopores via intermediate molecules, such as biotin conjugated to the enzyme and nanopore, wherein a streptavidin tetramer is linked to two biotins. The intermediate molecule may be referred to as a linker.
[0388] Sequencing enzymes can also be attached to nanopores using antibodies. Proteins that form covalent bonds with each other can be used to attach polymerases to nanopores. Phosphatases or enzymes that cleave tags from nucleotides can also be attached to nanopores.
[0389] Polymerases can be mutated relative to non-mutated polymerases to promote and / or improve the efficiency with which the mutated polymerase incorporates labeled nucleotides into growing polynucleotides. Polymerases can be mutated to allow nucleotide analogs (such as labeled nucleotides) to better enter the active site region of the polymerase and / or mutated to match nucleotide analogs in the active region.
[0390] Other mutations, such as amino acid substitutions, insertions, deletions, and / or exogenous polymerase features, can result in enhanced metal ion coordination, reduced exonuclease activity, reduced reaction rates of one or more steps in the polymerase kinetic cycle, reduced branching ratios, altered cofactor selectivity, increased yield, increased thermal stability, increased accuracy, increased speed, increased read length, and increased salt tolerance relative to non-mutant polymerases.
[0391] Suitable polymerases may have kinetic rate characteristics suitable for detecting tags through nanopores. Rate characteristics generally refer to the total rate of nucleotide incorporation and / or the rate of any step of nucleotide incorporation, such as nucleotide addition, enzyme isomerization (e.g., becoming closed or isomerizing from a closed state), cofactor binding or release, product release, incorporation of polynucleotides into the growth specification (pages 58 / 69, CN 121472373 A), or the rate of translocation.
[0392] Polymerases may be adapted to allow the detection of sequencing events. The rate characteristics of the polymerase can allow the tag to be loaded into the nanopore (and / or detected by the nanopore) for an average duration of 0.1 ms, 1 ms, 5 ms, 10 ms, 20 ms, 30 ms, 40 ms, 50 ms, 60 ms, 80 ms, 100 ms, 120 ms, 140 ms, 160 ms, 180 ms, 200 ms, 220 ms, 240 ms, 260 ms, 280 ms, 300 ms, 400 ms, 500 ms, 600 ms, 800 ms, or 1000 ms. For example, the rate characteristics of the polymerase can allow the tag to be loaded into the nanopore and / or detected by the nanopore for at least 5 ms, at least 10 ms, at least 20 ms, at least 30 ms, at least 40 ms, at least 50 ms, at least 60 ms, at least 80 ms, at least 100 ms, at least 120 ms, at least 140 ms, at least 160 ms, at least 180 ms, or even longer.The reaction time is at least 200 ms, at least 220 ms, at least 240 ms, at least 260 ms, at least 280 ms, at least 300 ms, at least 400 ms, at least 500 ms, at least 600 ms, at least 800 ms, or at least 1000 ms. The nanopore can detect the tag in an average range of 80 ms to 260 ms, 100 ms to 200 ms, or 100 ms to 150 ms.
[0393] The nanopore / polymerase complex can be configured to allow the detection of one or more events associated with the amplification and transcription of cyclic polynucleotides. One or more events can be kinetically observable and / or non-kinetically observable, such as nucleotide migration through the nanopore without contact with the polymerase.
[0394] In some cases, the polymerase reaction exhibits two kinetic steps that begin in an intermediate in which the nucleotide or polyphosphate product binds to the polymerase, and two kinetic steps that begin in an intermediate in which neither the nucleotide nor the polyphosphate product binds to the polymerase. The two kinetic steps may include enzyme isomerization, nucleotide incorporation, and product release. In some cases, the two kinetic steps are template translocation and nucleotide binding.
[0395] Suitable polymerases can exhibit strong or enhanced strand substitution.
[0396] Connector
[0397] The polymerase may contain a connector. The connector can be used to couple the polymerase to a nanopore. A polymerase containing a connector can be bound to a protein nanopore to form a polymerase / nanopore complex that can bind to the cleavage site of a cyclic polynucleotide.
[0398] A polymerase containing a connector can react with a cleaved cyclic polynucleotide to form a polymerase / cyclic polynucleotide complex. The polymerase / cyclic polynucleotide complex can bind to nanopores such as protein nanopores (e.g., α-hemolysin) or solid nanopores via the connector.
[0399] The polymerase / cyclic polynucleotide / nanopore complex can be used for polynucleotide sequencing. The linkage properties between the DNA polymerase and the nanopore can increase the concentration of effectively labeled nucleotides, thereby reducing the entropy barrier. Examples of optimizable linker aspects include linker length, which can increase the concentration of effectively labeled nucleotides, affect capture kinetics, and / or alter the entropy barrier; linker flexibility, which can affect the kinetics of linker conformational changes; and the number and location of links between the polymerase and the nanopore, which can reduce the number of available conformational states, thereby increasing the likelihood of appropriate pore polymerase orientation, increasing the concentration of effectively labeled nucleotides, and reducing the entropy barrier.
[0400] Linkers can be polymers, such as peptides, polynucleotides, or polyethylene glycol. Linkers can be of any suitable length. For example, linker lengths can be 5 nm, 10 nm, 15 nm, 20 nm, 40 nm, 50 nm, or 100 nm. Linker lengths can be at least5 nm, at least 10 nm, at least 15 nm, at least 20 nm, at least 40 nm, at least 50 nm, or at least 100 nm. The length of the connector can be less than 5 nm, less than 10 nm, less than 15 nm, less than 20 nm, less than 40 nm, less than 50 nm, or less than 100 nm. The connector can be rigid, flexible, or a combination thereof.
[0401] In some embodiments, no connector is used, and the polymerase is directly attached to the nanopore.
[0402] The polymerase can be attached to the nanopore via two or more connectors. The number and location of the connections between the polymerase and the nanopore can be varied. Examples include connections between the αHL C-terminus and the N-terminus of the polymerase, the αHL N-terminus and the C-terminus of the polymerase, and connections between amino acids not at the ends as described on pages 59 / 69 of the specification, CN 121472373 A.
[0403] The connector can be used to orient the polymerase relative to the nanopore so that the tag can be detected by means of the nanopore.
[0404] For example, in a method for sequencing a polynucleotide sample using a nanopore in a membrane adjacent to a sensing electrode, a labeled nucleotide is provided into a reaction chamber containing a nanopore, wherein the labeled individual nucleotides of the labeled nucleotides contain a tag coupled to the nucleotide, which can be detected by means of the nanopore. The method may include a polymerization reaction using a polymerase attached to the nanopore via a linker, thereby incorporating the labeled individual nucleotides of the labeled nucleotides into a growth chain complementary to a single-stranded polynucleotide from the polynucleotide sample. The method may include detecting the tag associated with the labeled individual nucleotides using the nanopore during the incorporation of the labeled individual nucleotides, wherein the tag is detected by means of the nanopore when the nucleotides associate with the polymerase.
[0405] Amplification and Sequencing
[0406] Amplification and transcription may include rolling circle amplification (RCA).
[0407] In RCA, the reaction mixture may contain one or more primers, a polymerase, and dNTPs, and produce a tandem. The polymerase in the RCA reaction may include a polymerase having chain displacement activity. Examples of polymerases with strand substitution activity include DNA polymerase I large (Klenow) fragment, Phi29 DNA polymerase, and Taq DNA polymerase, which lack exonuclease activity.
[0408] During sequencing while amplifying, to prevent DNA polymerase from binding from the original template to the substituted single-stranded DNA, single-stranded cutting enzymes (e.g., truncated exonuclease VIII, T5 exonuclease, T7 exonuclease) can be used to cleave the substituted single-stranded DNA into dNMP, dinucleotides, etc.
[0409] In some cases, the amplified polynucleotides can be visualized as nanospheres under a fluorescence microscope or by particle size analysis.
[0410] Identification of Sequence Variations
[0411] The methods provided in this disclosure can be used to identify sequence variations in polynucleotide samples. If the sequence differences are significant...When a sequence difference occurs in at least two distinct polynucleotides, such as two distinct cyclic polynucleotides, the sequence difference between the sequencing read and the reference sequence is called a true sequence variant. This can be distinguished by the different linkers. Because the location and type of sequence variants caused by amplification or sequencing errors are unlikely to be precisely replicated on two distinct polynucleotides containing the same targe...
Claims
1. A method for processing or analyzing double-stranded nucleic acid molecules, comprising: (a) Provide (i) the double-stranded nucleic acid molecule and (ii) a double-stranded linker having a cleavage site in its sense or antisense strand; (b) coupling the double-stranded linker to the double-stranded nucleic acid molecule; and (c) Circularize the double-stranded nucleic acid molecule coupled to the double-stranded linker to produce a circular double-stranded nucleic acid molecule.
2. The method according to claim 1, wherein the double-stranded nucleic acid molecule and the double-stranded adapter are heterologous to each other.
3. The method of claim 1, wherein the double-stranded nucleic acid molecule and the double-stranded adaptor are provided in the form of a cell-free composition.
4. The method of claim 1, wherein (b) or (c) is performed under cell-free conditions.
5. The method of claim 1, wherein the coupling comprises (i) coupling the sense strand of the double-stranded adapter to the sense strand of the double-stranded nucleic acid molecule, or (ii) coupling the antisense strand of the double-stranded adapter to the antisense strand of the double-stranded nucleic acid molecule.
6. A reaction mixture for processing or analyzing double-stranded nucleic acid molecules, comprising: A composition comprising (i) the double-stranded nucleic acid molecule and (ii) a double-stranded linker having a cleavage site within its sense or antisense strand; and At least one enzyme, said enzyme (i) couples the double-stranded adapter to the double-stranded nucleic acid molecule, and (ii) cyclizes the double-stranded nucleic acid molecule coupled with the double-stranded adapter to produce a cyclized double-stranded nucleic acid molecule.
7. A library of circularized double-stranded nucleic acid molecules comprising (i) a double-stranded nucleic acid domain coupled to (ii) a double-stranded linker domain, the double-stranded linker domain containing a nick site within the sense or antisense strand of the double-stranded linker domain, wherein each circularized double-stranded nucleic acid molecule in at least 5% of the library contains a recognition sequence.
8. A method for processing or analyzing circular nucleic acid molecules, comprising: (a) Providing a cell-free composition comprising the circular nucleic acid molecule, the circular nucleic acid molecule comprising (i) a target region and (ii) a nick site at a known distance from the target region; and (b) A nick is made at the nick site of the circular nucleic acid molecule.
9. A reaction mixture for processing or analyzing circular nucleic acid molecules, comprising: A cell-free composition comprising the circular nucleic acid molecule, the circular nucleic acid molecule comprising (i) a target site and (ii) a nick site at a known distance from the target site; and At least one enzyme that produces a nick at the nick site of the circular nucleic acid molecule.
10. A cell-free library of circular nucleic acid molecules, wherein each individual of at least 5% of the library comprises (i) a target site and (ii) a cut at a known distance from the target site.