Nanopore sequencing DNA by replacement of nucleotides with analogs

By incorporating analog dNTPs with modified electrical properties into nanopore sequencing, the method enhances the accuracy of nanopore sequencing, addressing the high error rates in existing technologies.

WO2025128542A1PCT designated stage expired Publication Date: 2025-06-19UNIV OF WASHINGTON
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/US2024/059342
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-11
Filing Date
2024-12-10
Publication Date
2025-06-19

AI Technical Summary

Technical Problem

Nanopore sequencing methods suffer from higher error rates in raw sequence reads compared to other sequencing methods, necessitating improvements in accuracy.

Method used

A method involving the use of analog dNTPs with modifications relative to corresponding reference dNTPs is employed to improve the accuracy of nanopore sequencing. These analog dNTPs are incorporated into nascent DNA strands by DNA polymerase, altering the electrical resistance and charge, thereby enhancing sequencing accuracy.

Benefits of technology

The incorporation of analog dNTPs significantly improves the accuracy of nanopore sequencing by altering the electrical properties of the nucleotides, leading to higher confidence in nucleobase identification and reduced sequencing errors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2024059342_19062025_PF_FP_ABST
    Figure US2024059342_19062025_PF_FP_ABST
Patent Text Reader

Abstract

Devices, kits, compositions, and methods for high-accuracy nanopore sequencing of polynucleotides. Methods include copying target sequences with polymerase reactions having analog deoxynucleotide triphosphates (dNTPs; e.g., dNTPs that differ chemically or physically from canonical dNTPs), in place of canonical dNTPs, that are incorporated into copies of the target sequences in place of canonical dNTPs. Due to the presence of the analog dNTPs in the copy sequences, the copy sequences exhibit different electrical signatures compared with the target sequences and can be used for improving confidence in base-calling and reducing the complexity of the sequencing problem.
Need to check novelty before this filing date? Find Prior Art

Description

NANOPORE SEQUENCING DNA BY REPLACEMENT OF NUCLEOTIDES WITH ANALOGSCROSS-REFERENCE(S) TO RELATED APPLICATION S)

[0001] This international application claims priority to, and the benefit of, U.S. Provisional Application No.: 63 / 608,712 filed December 11, 2023; the contents of which are incorporated by reference herein in their entirety for all purposes.STATEMENT OF GOVERNMENT LICENSE RIGHTS

[0002] This invention was made with government support under Grant No. R01HG005115, awarded by the National Human Genome Research Institute. The government has certain rights in the invention.STATEMENT REGARDING SEQUENCE LISTING

[0003] The Sequence Listing XML associated with this application is provided in XML format and is hereby incorporated by reference into the specification. The name of the XML file containing the sequence listing is 3915- P1330WO.UW_Sequence_Listing.xml. The XML file is 29,721 bytes; was created on Friday, December 06, 2024; and is being submitted electronically via Patent Center with the filing of the specification.BACKGROUND

[0004] Nanopore sequencing (e.g., the Oxford Nanopore Technologies® platform) is based on measuring changes in electrical signal generated due to DNA or RNA molecules passing through nano-scaled pores (“nanopores”). The approach offers long read sequencing (reads in which the mean read length can exceed 10 kb, and the maximal read length can reach 880 kb or more) as well as real-time analysis and a low initial investment cost. However, while nanopore sequencing provides a generally effective approach for sequencing a variety of DNA and RNA fragments, these approaches typically suffer from a higher error rate on raw sequence reads compared with other sequencing methods, such as standard next-generation sequencing (NGS), (e.g., the Illumina® platform). Improvements in the accuracy of nanopore sequencing are greatly needed.

[0005] Accordingly, there is a need for methods for improving the accuracy of nanopore sequencing devices and methods. The present disclosure addresses this and other long-felt and unmet needs in the art.SUMMARY

[0006] This summary is provided to introduce a selection of concepts in a simplified form that are further described below in the Detailed Description. This summary is not intended to identify key features of the claimed subject matter, nor is it intended to be used as an aid in determining the scope of the claimed subject matter.

[0007] In an aspect, the disclosure provides a method for preparing a DNA library for high-accuracy nanopore sequencing, the method comprising: contacting a DNA strand comprising a target polynucleotide sequence with a DNA polymerase and a dNTP pool comprising an analog dNTP that comprises a modification relative to a corresponding reference dNTP that is absent from the dNTP pool; wherein the DNA polymerase incorporates the analog dNTP into a nascent DNA strand as an analog nucleotide in place of incorporation of the corresponding reference dNTP as a corresponding reference nucleotide.

[0008] In embodiments, an electrical resistance of the corresponding reference nucleotide differs from an electrical resistance of the analog nucleotide.

[0009] In embodiments, the electrical resistance of the corresponding reference nucleotide is greater than the electrical resistance of the analog nucleotide.

[0010] In embodiments, the electrical resistance of the corresponding reference nucleotide is less than the electrical resistance of the analog nucleotide.

[0011] In embodiments, the modification of the analog nucleotide increases an electrical charge of the analog nucleotide relative to an electrical charge of the corresponding reference nucleotide.

[0012] In embodiments, the modification of the analog nucleotide decreases an electrical charge of the analog nucleotide relative to an electrical charge of the corresponding reference nucleotide.

[0013] In embodiments, the method is isothermal.

[0014] In embodiments, the contacting the DNA strand is performed a plurality of times with a plurality of dNTP pools, wherein dNTP pools of the plurality of dNTP poolscomprise pool-specific analog dNTPs that comprise pool-specific modifications relative to corresponding reference dNTPs that are absent from the dNTP pools.

[0015] In embodiments, the contacting the DNA strand is performed three times.

[0016] In embodiments, the method further comprises removing previous dNTP pools from a reaction mixture between contacting steps.

[0017] In embodiments, the removing previous dNTP pools comprises immobilizing the DNA strand and the nascent DNA strand and washing the reaction mixture such that a previous analog dNTP is removed from the reaction mixture.

[0018] In embodiments, the immobilizing step comprises attachment of a DNA molecule comprising the DNA strand, the nascent DNA strand, or both, to a substrate.

[0019] In embodiments, a majority of dNTPs of the dNTP pool are not analog dNTPs.

[0020] In embodiments, no more than 25% of dNTPs of the dNTP pool are analog dNTPs.

[0021] In embodiments, a majority of dNTPs of the dNTP pool are analog dNTPs.

[0022] In embodiments, 100% of dNTPs of the dNTP pool are analog dNTPs.

[0023] In embodiments, the method further comprises ligating a sequencing adaptor to the DNA strand and the nascent DNA strand to configure the DNA library for nanopore sequencing.

[0024] In embodiments, the sequence adaptor comprises a barcode that corresponds with an identity of the analog dNTP.

[0025] In embodiments, the method further comprises reverse transcribing an RNA strand to produce the DNA strand comprising the target polynucleotide sequence, wherein the DNA library represents at least part of an RNA fraction of a sample.

[0026] In an aspect, the disclosure provides a DNA library prepared according to a method of the disclosure.

[0027] In an aspect, the disclosure provides a method for high-accuracy analysis of nanopore sequencing data obtained from nanopore sequencing of a DNA library of the disclosure, the method comprising: aligning the nanopore sequencing data such that sequence data of the DNA strand is aligned with sequence data of the nascent DNA strand; computing an electrical property of the nanopore sequencing data that corresponds to an analog nucleobase of the nascent DNA strand for an analog nucleobase measurement; comparing the analog nucleobase measurement with a reference analog nucleobasemeasurement associated with a known analog nucleobase identity and assigning a nucleobase identity of the analog nucleobase consistent with the known analog nucleobase identity; and assigning, based on the nucleobase identity of the analog nucleobase, a nucleobase identity to an unknown nucleobase of the DNA strand that positionally corresponds to the analog nucleobase of the nascent DNA strand.

[0028] In embodiments, a plurality of nucleobase identities of analog nucleobases of the nascent DNA strand correspond to a plurality of nucleobase identities of unknown nucleobases of the DNA strand.

[0029] In embodiments, the plurality of nucleobase identities of analog nucleobases of the nascent DNA strand comprises one distinct nucleobase identity that corresponds to one contacting step used for DNA library preparation.

[0030] In embodiments, the plurality of nucleobase identities of analog nucleobases of the nascent DNA strand comprises two distinct nucleobase identities that correspond to one or more contacting steps used for DNA library preparation.

[0031] In embodiments, the plurality of nucleobase identities of analog nucleobases of the nascent DNA strand comprises three distinct nucleobase identities that correspond to one or more contacting steps used for DNA library preparation.

[0032] In embodiments, the plurality of nucleobase identities of analog nucleobases of the nascent DNA strand comprises four distinct nucleobase identities that correspond to one or more contacting steps used for DNA library preparation.

[0033] In embodiments, the method is performed at least in part by a programmable processor, a processor circuitry, a computational device, a computational system, a computational network, or any combination thereof. These and other devices for methods of the disclosure can comprise circuitry configured for performance of all or part of a method of the disclosure. Accordingly, in various aspects, the disclosure also provides a processor, a processor circuitry, a computational device, a computational system, a computational network, or any combination thereof, comprising circuitry configured for performance of all or part of a method of the disclosure, in any order or combination of steps, by the processor, processor circuitry, computational device, computational system, computational network, or combination thereof, as the case may be.DESCRIPTION OF THE DRAWINGS

[0034] The foregoing aspects and many of the attendant advantages of this disclosure will become more readily appreciated as the same become better understood by reference to the following detailed description, when taken in conjunction with the accompanying drawings.

[0035] FIGs 1A-1D show examples of variable voltage nanopore sequencing, according to aspects of the disclosure. FIG. 1A) Schematic of nanopore setup. FIG. IB) Because the DNA is elastic, different applied voltages (forces) stretch the DNA to differing degrees. FIG. 1C) (left) Constant-voltage sequencing yields only information about the average current. Current-level degeneracies contribute significantly to sequencing errors, (right) Using a variable- voltage to floss the DNA back and forth, one can sample various locations along the DNA and extract smooth curve segments. These smooth curve segments are what one would see if DNA were smoothly and continuously translocated through the pore. FIG. ID) Evaluation of sequencing accuracy involves an alignment step in which the called bases are aligned to the known sequence. Alignment of the sequence to the reference boosts sequencing “accuracy” of random sequences from 25% to nearly 60%. Variablevoltage sequencing increases single-passage sequencing accuracy significantly above random and serves as a significant step towards 100% accuracy. This boost in sequencing accuracy comes from the added information supplied by a “curve-segment” as compared to an individual ion current level. One can understand this by noting that a curve segment provides additional sequence information in the slope and curvature in addition to the average ion current that can be measured using a constant voltage. The added information available with variable voltage sequencing makes it an attractive method for sequencing 8-, 10-, and 12-letter DNA.

[0036] FIG. 2A shows examples of histograms of constant-voltage conductance caused by translocation of DNA substrates containing 16 nucleotide homopolymers of each Hachimoji base (excluding G), according to aspects of the disclosure. Homopolymers of P and Z give the lowest and highest conductance signals, evidence that Hachimoji system has an expanded nanopore signal range relative to the standard alphabet. The number of bases which contribute to sequencing kmers and limited current space in which these kmer- induced currents exist contributes to the difficulty of high-accuracy nanopore sequencing. Using previous approaches, many kmers are difficult to distinguish from one another.

[0037] FIGs 2B and 2C show example (FIG. 2B) variable-voltage consensus patterns of Hachimoji single-base substitutions within a pseudorandom sequence, according to aspects of the disclosure. Variable-voltage nanopore sequencing contains more sequence information than obtained from current levels produced from constantvoltage analysis. The figure also shows an example (FIG. 2C) confusion matrix of the basecalling algorithm’s accuracy, which shows that Hachimoji single-base substitutions are distinguished with high confidence using variable-voltage nanopore sequencing. Reads are included in base calling only if they align to one of the consensus patterns with confidence > 90%.

[0038] FIG. 3 shows an example scheme for boosting nanopore sequencing accuracy of supernumerary DNA using additional AEGIS alphabets, according to aspects of the disclosure. Many AEGIS bases are interchangeable with one another; z.e., there exist two versions of the base Z, four versions of the base C, etc. which are compatible with polymerases. Using molecular cloning techniques, strands can be built which contain multiple copies of the original sequence using differing AEGIS alphabets. Briefly, by A- tailing (step 2), hairpin adapters can be ligated to the sequence of interest (step 3). Using a strand displacing polymerase, the strand can then be zipped up using an alternative AEGIS alphabet (step 4). This hairpin construct can be again A-tailed (step 5), and another hairpin adapter ligated (step 6). This hairpin can again be zipped up with a strand displacing polymerase using a third AEGIS alphabet. Then, as the strand is read in a nanopore sequencing read, comparisons of reads of the sense (S) and antisense (S') strands from each alphabet can be combined (FIGs 4A and 4B) to identify individual bases within the read and significantly simplify the sequencing problem.

[0039] FIGs 4A and 4B show example schematic comparisons of reads of the same letter sequence (8 letters ACGTPZKX or 4 letters, in FIG. 4A or FIG. 4B, respectively), according to aspects of the disclosure, using different heterocyclic variants of an 8-letter AEGIS alphabet or 4-letter alphabet (S, S' in FIGs 3, 4A, and 4B). If a strand constructed as shown in FIG. 3 is read by the nanopore, several reads of the same “letter sequence,” as identified by their base-pairing behavior, can be obtained using different AEGIS alphabets. In the first comparison of the sequence, “S,” including heterocyclic variants of X and P variants are used. These base “doppl egangers” are indicated by “ " “. The two reads can then be aligned based on shared ion currents (displayed here with a constant applied voltage for simplicity). Ion current increases indicate the location of X and current decreasesindicate the location of P in the primary strand. Reads of the complementary sequence, “S',” using a third alphabet, can be used to gain additional sequence information on additional bases, in this case C and A. Combining this information significantly simplifies the sequencing problem for supernumerary DNA. A similar application to 4-letter DNA in FIG 4B using a G base analog allows for identification of the location of all G’s in the primary sequence and C’s in the complementary sequence.

[0040] FIG. 5A shows examples of how rearranging hydrogen bond donor and acceptor groups on purine-pyrimidine pairs increases the number of independently replicable information units in a DNA-like evolvable system from 4 to 12, according to aspects of the disclosure. The bases A, T, C, G, P, and Z of this artificially expanded genetic information system (AEGIS) are Rokumoji DNA (from the Japanese, “roku” = “6”, moji = “letter”, as in “emoji”). Following the nomenclature, Hachimoji DNA has 8 letters (“hachi” = “8”), where the S:B pair is added.

[0041] FIG. 5B shows a range of example AEGIS Z variants provided herein, according to aspects of the disclosure. All have the same general hydrogen bonding pattern, but may be implemented with different heterocycles and sugar substituents. This allows the tuning of the acid-base properties of the heterocycle, influencing Z contribution to AEGISZyme catalysis and influencing the Z contribution to the measured ion current, enabling base identification, according to aspects of the disclosure.

[0042] FIG. 6A shows (left) a range of example AEGIS B provided herein, according to aspects of the disclosure, with the same general hydrogen bonding pattern implemented with different heterocycles. These allow the tuning of tautomerism properties of the heterocycle to prefer the keto or enol tautomer, (right) N- and C-glycoside variants of C and T. This perturbs the stability of the nucleotide to chemical degradation and repair enzymes, which can be desired or not depending on the application. Each nucleotide variant is configured to produce a different ion current when held in the pore constriction. These differences can be used to locate particular bases.

[0043] FIG. 6B shows some examples of the functionalized AEGIS variants provided herein, according to aspects of the disclosure. These increase the catalytic potential of AEGIS libraries. It is noted how the boronate side chain (right) allows an AEGIS library to be enriched in components that bind glycoproteins. Each nucleotide variant is configured to produce a different ion current when held in the pore constriction. These differences can be used to locate particular bases.

[0044] FIG. 6C shows an example generic synthesis of C-glycosides, illustrated for a variant of AEGIS Z, according to aspects of the disclosure.

[0045] FIG. 7A shows validation of a strand with agarose gel electrophoresis, according to aspects of the disclosure. Lane 1 : sequencing adapter. Result of elongation with Lane 2: normal Cytosine, Lane 3: 5mC, or Lane 4: 5-hmC. Lanes 5, 6, and 7 show the full sequence after sequencing adaptor ligation for C, 5mC, and 5-hmC strands, respectively. Identified bands are a) HP2 ~ 70 bases (~35 bp), b) original sequence “S” ~ 40 bp. c) sequencing adaptor ~ 90 bp. d) extended sequence - 180 bp. c) Successful full sequence - 270 bp. Bands a, b, and c are due to the partial success of hairpin ligation steps. 5-hydroxy methyl showed a low yield which is reflected as fainter bands in lanes 4 and 7.

[0046] FIG. 7B shows a kmer-map prediction for ACGT DNA (black) compared to nanopore read of DNA in which 5mC replaces all Cs, according to aspects of the disclosure. The first 45 bases of the read are from the sequencing adapter and contain C. The vertical dashed line in the lower figure indicates the boundary between the adapter and primary sequence containing 5mC. The effect of 5mC on the ion current depends on the surrounding sequence context. An element of this result is that substitution of all C with 5mC significantly alters the kmer map. Reads of the same sequence containing C or 5mC produce orthogonal information, which can be used to significantly boost sequencing accuracy.

[0047] FIG. 8A shows results from performing one round of replication using an alphabet of dNTPs which produces a high-current contrast signal in the nanopore which enables high accuracy sequencing of 6-letter DNA (ACGTPZ); selection of various modified dNTPs enables bases to be easily distinguished from one another.

[0048] FIG. 8B shows an example scheme for generating DNA strands which enable multiple reads of the same primary sequence in multiple different alphabets, according to aspects of the disclosure.

[0049] FIG. 8C shows an ion current histogram for ACGT 4-mers with M2- MspA, according to aspects of the disclosure.

[0050] FIG. 9A shows an example of library preparation to generate data for sequencing model training, according to aspects of the disclosure.

[0051] FIG. 9B shows the joint distribution of kmer and anti -kmer ion currents, according to aspects of the disclosure. A slice through the distribution at the most populous kmer current bin shows how the anti-kmer current can help break sequencing degeneracies.Adding kmer currents in additional alternative alphabets could be used to further-resolve kmers in particular by targeting bases responsible for kmers which are still poorly resolved from one another in duplex sequencing (cluster in light gray).

[0052] FIG. 9C shows sequencing with multiple data streams amounts to a multiple sequence alignment problem in which data stream A is a nanopore read of the sense strand and data stream B is a nanopore read of the antisense strand, according to aspects of the disclosure. The third axis of the “multiple sequence alignment” (MSA) is the kmer map. Steps in the horizontal plane amount to a Needleman-Wunch alignment of A and B while vertical steps in the cube are a Viterbi sequencing algorithm which takes match scores for both data streams. The faces of the cube reduce to alignment of A to B (top face), sequencing of A (right face) and sequencing of B (left face). Additional streams of data can be used by performing higher dimensional alignments, which are difficult to visualize but straightforward to implement on a computer.

[0053] FIG. 9D shows a model demonstrating that one need only consider a small subset of the elements of the MSA hyper-cube, according to aspects of the disclosure. These elements can be determined from pairwise alignments of the reads and individual read sequencing which form the faces of the hyper-cube.

[0054] FIG. 10 shows an example diagram of sequencing of + and - strands (left), sequencing + and - sense strands with a second high-contrast alphabet (middle), and sequencing + and - sense strands with two or more high-contrast alphabets (right), according to aspects of the disclosure. The methods of the disclosure can be performed iteratively, with iterative use of two or more high-contrast dNTP pools for improved nanopore sequencing.

[0055] FIG. 11A shows the set of Hachimoji DNA bases, which includes A, T, G, C, Z, P, S, and B, in a hydrogen bonded configuration, according to aspects of the disclosure.

[0056] FIG. 11B shows an illustration of an example nanopore, according to aspects of the disclosure. The figures shows single-stranded DNA (ssDNA) in the nanopore, and motion of the DNA due to a voltage potential across the bilayer, Hel308 helicase, and MspA. The letter x designates a span of nucleotides (nt) positioned within an interior space of the MspA protein.

[0057] FIG. 11C shows ion current (pA) as a function of time (s), according to aspects of the disclosure. The figure shows ion current changes as the non-canonical basesof a DNA polynucleotide strand transition from the pore (left portion of graph between two vertical dotted lines) to the helicase (right portion of graph between two vertical dotted lines)DETAILED DESCRIPTION

[0058] Nucleic acids have a wide array of natural and synthetic applications. Synthetic nucleic acids can be comprised of the canonical four nucleobases, as well as additional, non-canonical nucleobase analogs. Nanopore sequencing of natural, as well as synthetic nucleic acids, is in significant need of improvements in accuracy.

[0059] By way of background, synthetic biologists have shown that DNA can be assembled from nucleotides other than the canonical four, G, A, T, and C. In particular, the two base-complementarity rules (size and hydrogen bonding) that make canonical basepairing function can be exploited to accommodate 12 independently replicating nucleotides forming 6 orthogonal base pairs (FIG. 5A). As in canonical DNA, this “artificially expanded genetic information system” (AEGIS) can store, transmit, and evolve information. Thus, AEGIS-DNA can support a broad range of applications. For example, AEGIS can transform clinical diagnostics, and AEGIS aptamers (AEGISBodies) generatedby laboratory in vitro evolution (LIVE) can bind whole cancer cells, and AEGIS-LIVE has produced reagents that deliver drugs selectively to cancer cells. However, despite these successes, there are no methods to sequence AEGIS DNA containing more than 6-letters.

[0060] In various aspects, this disclosure provides innovative tools to de novo sequence full 12 letter AEGIS DNA that also make variable voltage sequencing more accessible.

[0061] As described herein, differentiation between canonical nucleobases of a target sequence can be facilitated by “doping” a copy of a target sequence with one or more nucleobase analogs (containing at least one nucleobase analog), such that the “doped” copy produces a distinct electric current signature relative to a corresponding “canonical-only” copy of the target sequence (containing only G, A, T, C).

[0062] Alignment of sequences enables identification of the location of the nucleobase analog in the doped sequence, which corresponds to the location of a corresponding canonical nucleobase in the target sequence and enables identification of that corresponding canonical nucleobase in the target sequence. This increases the accuracy of nanopore sequencing builds and reduces the size of the sequencing problem. In the context of various aspects of the disclosure, “doping”, “doped”, “substituted”, and similar terms, refer to the replacement of all or a subset of canonical dNTPs of a target sequence with analog dNTPs in a copy of the target sequence generated with a reaction of the disclosure; the number of analog dNTPs in a given reaction can be 1, 2, 3, 4, or more, for example.

[0063] Highly accurate nanopore sequencing as disclosed herein not only enhances nanopore sequencing of any target polynucleotide sequence, but also enables AEGIS-LIVE to produce a new class of research, diagnostics, and therapeutic tools at low cost with fast response (in just weeks) that are catered to specific needs. These include AEGIS Bodies and AEGIS-Zymes (30-40 nucleotide DNA strands that, depending on how they are evolved, bind targets or, after binding, attach themselves to the targets, modify the targets, or are transported with the targets). AEGIS-Zymes can activate, inactivate, or modulate targets as “manipulating evolvable drugs” (MEDs); cross bio-barriers to enter cells, brain, and other privileged tissues in “mirror image” stable and active forms; allow cargo molecules (drugs) to acquire the pharmacodynamics of bound proteins, e.g., albumin or IgG; and can be made in weeks to target cells from individual patients, potentially providing a low-cost “personalized pharmacopeia”.

[0064] There is also a need for reagents that deliver selective and potent binding “on demand” for research, diagnostics, and therapeutics, and improved nanopore sequencing methods that enable development and use of these reagents. In addition, adding functional groups to a DNA library increases its intrinsic value as a source of catalytic species, and is also an evolvable system that allows hydrophobic units to be added sparingly to evolved nucleic acid products. Such a system is able to deliver receptors and ligands that can provide utility in a range of contexts. The disclosure provides improved approaches for nanopore sequencing that enable development of these and other technologies to create these and other macromolecules that can be easily reproduced and can reliably bind targets.

[0065] As such, it is proposed to utilize more nucleotides (e.g., 12), more functional groups (FIG. 5A), better folding, more opportunities for compact folds, sub-picomolar affinity, and increased catalytic power; e.g., an AEGIS. Since evolving informational biopolymers have a repeating backbone charge, they are able to have generally constant overall physical properties with changing genetic information, including solubility. However, a polyelectrolyte backbone discourages backbone backbone interactions to form compact folds, which can be useful to achieve precise positioning for high affinity binding and catalyzed reactivity. This can be resolved by using base-base interactions to form folds, instead of the backbone backbone interactions.

[0066] By way of background, standard 4-letter DNA has one base:base interaction that forms a compact fold, the “G quadruplex.” It often emerges when evolving 4-letter DNA; libraries can be biased to be G-rich to create aptamers with this compact fold. AEGIS DNA has many additional base-base interactions that create core folds. Rokumoji DNA has one new fold that can be of interest, the Z:Z- “fZ-motif ’. NMR and biophysical studies of the fZ motif shows that it gives compact folds with parallel strands. Here, it is analogous to the i-motif formed by protonated C.

[0067] While chemistry has been developed for the complete set of 12 letter DNA, and polymerases can work with many different AEGIS pairs, existing sequencing technology is limited and exists only for the six-letter subset of AEGIS (GACTZP, Rokumoji DNA, from the Japanese roku = 6, and moji = letters, in the box in FIG. 5A). This limitation in sequencing technologies may not enable the full benefit of AEGISBodies and AEGISZymes. Further, sequence data can be used to modify AEGISConstructs, assess mechanisms, add nanotrains, understand folds, and re-design AEGIS building blocks to give libraries that are richer reservoirs for targeted functions. However, furtherdevelopment of 8-, 10-, and 12-letter AEGIS systems needs technology capable of sequencing 8-, 10-, and 12-letter DNA.

[0068] Analytical tools can determine what has emerged by evolution of AEGIS DNA. For example, with Rokumoji LIVE experiments, pools of Rokumoji DNA were sequenced using technology that used transliteration. Transliteration used PCR conditions that convert P to a mixture of A and G, and Z to a mixture of C and T. The mixtures were then subjected to deep sequencing. Bioinformatics identifies reads that descend from a single component in the pool, and deconvolutes the sequence. The sequence of the ancestral DNA molecule in the pool at positions where all of its descendants have G, A, C, or T are assigned to be G, A, C, or T. If a position in the descendant has both T and C, a Z is assigned at that position in the ancestral molecule in the pool. If a position in the descendant has both G and A, a P is assigned at that position in the ancestral molecule.

[0069] However, the transliteration strategy can only be applied to 6-nucleotides, and a broader sequencing approach is needed if additional orders of magnitude in performance are to be obtained from AEGIS-LIVE that exploits more of the expanded genetic alphabet.

[0070] Nanopore sequencing technology offers such a broader approach. It is “label free,” and thus is less challenging to develop and use than the “cyclic reversible terminator strategy” used, for example, on an Illumina® platform. However, a challenge in adapting nanopore DNA sequencing for AEGIS-systems is the sheer number of 4-base-long “kmers” for 8-, 10- and 12-letter DNA (4096, 10000, and 20736, respectively). Add to that a large number of heterocyclic variations for many of the AEGIS nucleotides and there is a vast parameter space.

[0071] In the examples disclosed herein, focus was given to the eight-letter “Hachimoji” subset (GACTZPSB) of the AEGIS 12-letter alphabet. Here, synthetic Hachimoji DNA was presented for sequencing using the MspA (Mycobacterium smegmatis porin A) nanopore. Success was also found using ONT to sequence 12-letter DNA containing sparse xenonucleotide insertions. The MspA -based variable-voltage sequencing (FIGs 1 A-1D) can be utilized over ONT in at least some embodiments due to the beneficial information content of the nanopore signal in the former, which can be helpful for full factorial sequencing of 8-, 10- and 12-letter DNA. In addition, the pore and enzyme used by ONT is specifically tuned for the sequencing of 4-letter, ATCG DNA. ONT may not be able to be readily tuned for sequencing with AEGIS bases.

[0072] As shown at FIG. 2A, initial results with MspA -variable-voltage sequencing were unexpectedly good. Hachimoji DNA exhibits a broader signal range in nanopore sequencing than standard DNA alone. Information about the additional bases is encoded across a broader current range, facilitating base-recognition. Hachimoji single-base substitutions are distinguishable with high confidence.

[0073] As nanopore sequencing can use a helicase molecular motor to control the motion of DNA, the compatibility of the Hel308 motor enzyme with Hachimoji nucleotides was assessed by tracking the translocation of single Hel308 molecules along Hachimoji DNA, monitoring the enzyme kinetics and premature enzyme dissociation from the DNA. As a result, it was found that Hel308 is compatible with Hachimoji DNA but can dissociate more frequently when walking over C-glycoside nucleosides, compared to N-glycosides. These results highlight the possibility to improve nanopore sequencing motors to handle different glycosidic bonds and can also inform designs of alternative DNA systems that can be sequenced with existing motors and pores. These results also show that sequencing of AEGIS DNA is possible via nanopore sequencing, and in at least some instances, can include only small adjustments.

[0074] High-accuracy nanopore sequencing as per approaches of the disclosure can enable: development of sub-picomolar receptors and ligands on demand, with turnaround times in weeks; enable delivery of macromolecular receptors and ligands into cells, not today possible, with AEGISBodies that bind cell-surface proteins like XRCC5 and are internalized, or bind transferrin, to move across the blood brain barrier; evolvable drugs that manipulate targets that they bind; and “personalized pharmacopeias,” where the low cost and rapid turn-around of AEGIS-LIVE allows, for example, anti-cancer drug delivery systems to evolve with the evolution of cancer in a specific patient.

[0075] Accordingly, the disclosure provides example approaches for developing sequencing of the 8-letter AEGIS-alphabet for which there are working polymerases (ATCGPZKX). AEGIS-LIVE is possible with 10- and 12-letter DNA, and as such, the disclosure provides approaches for 10- and 12-letter sequencing. Variable voltage sequencing and the sequence-dependent kinetics of the motor enzyme can be used to extract additional sequence information and gain high accuracy.

[0076] As an example, the enzyme motor can be modified by single amino acid replacements, which can involve mutagenesis of a gene encoding Hel308 from Thermococcus gammatolerans, codon-optimized for expression in E. coli. Sites for aminoacid replacement can be chosen by their proximity and interaction with the nucleic acid and ATP -binding pocket. Variant forms of the protein can be characterized by expression level and activity on the nanopore.

[0077] As another example, AEGIS components can be modified. Modification can be achieved by attaching acetylene linkers at the 5-position of the small “pyrimidine” analogs, or the 7-position of the large “purine” analogs, using palladium-catalyzed coupling and the corresponding iodo-heterocycle, obtained by iodinating the base heterocycle by N- iodosuccinimide.

[0078] Yet another aspect of improved nanopore sequencing involves building sequencing models of expanded DNA alphabets. This can include designing and building synthetic oligos containing the various kmers within the expanded DNA libraries. These oligos can be measured on the nanopore system and incorporated into sequencing models.

[0079] By way of background, the kmer size for MspA is ~4 nucleotides, where the central dimer contributes the majority of the sequence information. It is thus possible to generate a rough kmer map by measuring all possible dimers within the given alphabet. Even for 12-letter DNA, there exist only 144 possible dimers which can be measured with a few synthetic oligos. While this map may have poor sequencing accuracy, it can serve as an initial guide for automated alignment of subsequent strands and provide insight into which bases can be tweaked for improvements to sequencing accuracy. (Because of the large number of heterocyclic variations available for each base, there are a number of possible AEGIS-LIVE configurations that can be simultaneously optimized for both effectiveness and facilitation of sequencing.) Once a particular configuration is developed, one can gather additional data of each 4-nucleotide-long kmer in many different contexts using a library of synthetic oligos containing many instances of each kmer in various sequence contexts. The rough dimer map will be sufficient for strand identification within these pools and individual reads can be incorporated into the kmer map as they are acquired. As the size of the dataset grows, the fidelity of the map will increase and converge to the true map.

[0080] In addition, motor enzymes used for the purpose of nanopore sequencing have displayed kinetics that are dependent upon the DNA sequence passing through the motor enzyme, and it was found that AEGIS bases P, Z, S, and B also interact with the motor enzyme in a unique way, affecting both step dwell-time and enzyme backsteps. For high-accuracy nanopore sequencing, one can use information derived from enzyme kineticsas a secondary source of information about the DNA sequence. Advances in this domain can feed directly into work in sequencing expanded DNA alphabets. Helicase mutants developed to increase the amount of sequence information contained within the enzyme’s kinetics can also be applied to the sequencing of AEGIS DNA.

[0081] Furthermore, when testing the feasibility of sequencing 8-letter DNA, it was found that the Hel308 helicase motor (from Thermococcus gammatolerans EJ3; accession number WP 015858487.1) would dissociate prematurely from the DNA strands when encountering long sections of C-glycoside bases. Without wishing to be bound by any particular theory, this may be due to the difference in sugar pucker for C-glycosides vs. the more typical N-glycosides seen in canonical DNA bases. As it turns out, C-glycosides adopt a 3'-endo pucker rather than the 2'-endo pucker typically observed in B-form DNA. It is possible that other AEGIS nucleotides may also not work well with the standard motor enzyme. This problem can be addressed in either or both of two ways: 1) mutation of the helicase to accommodate C-glycosides; and 2) selection of AEGIS-nucleotides that are compatible with the motor enzyme.

[0082] In addition, by analyzing where in the strand Hel308 falls off, the position within the enzyme at which the glycosidic nature of the DNA is important has been identified. Further, it is noted that the 3'-endo sugar pucker characteristic of C-glycosides is typically seen in RNA which adopts an A-form helix. Without wishing to be bound by any particular theory, this may be, at least in part, how Hel308 distinguishes between DNA and RNA substrates. Based on this, since one of the prime differences between 3 '-endo and 2'-endo DNA is the inter-phosphate distance: ~6 A vs. ~7 A, one can add flexibility to the helicase core where it interacts with the DNA backbone to better accommodate C- gly cosides. That there exist helicases which walk on both DNA and RNA (e.g., nspl3143- 145, upfl 146) this is possible to achieve. In addition, even if modification of Hel308 proves difficult, the large number of heterocyclic variations available among the AEGIS nucleotides can be leveraged to find an AEGIS library that is enzyme-compatible and useful in AEGIS-LIVE. For this, one can adapt a paradigm in which one AEGIS alphabet is used to generate binders while a PCR-compatible dual alphabet is used in sequencing applications, similar to how Z-NO2 and Z-CONH2can be used interchangeably for binders or PCR fidelity, respectively. It should be noted that the current Hel308 processivity on strands containing C-glycosides is sufficient for sequencing of short oligos such as thoseproduced by AEGIS-LIVE. Enzyme dissociation is primarily a problem in sequencing long AEGIS strands or those with several consecutive C-glycoside bases.

[0083] Furthermore, as currently implemented, Hel308-MspA variable-voltage sequencing is performed on a single pore with a single amplifier. This research / development setup allows for flexibility and rapid prototyping but may lack a level of throughput. As sequencing AEGIS-DNA occurs, one can build parallelized setups capable of collecting significantly larger volumes of data. These setups can be for collecting large datasets to train kmer models and ultimately for sequencing large, diverse pools of AEGISBodies and AEGISZymes. One can develop in-house, low noise amplifiers that can be produced at minimal cost and integrated in parallel.

[0084] With respect to aspects of the disclosure, addressing the challenge of fullfactorial sequencing of 8-, 10- and 12-letter DNA, despite promising initial results, it may be the case that direct, full-factorial de novo sequencing of AEGIS DNA with sufficient accuracy is not possible given the range of ion currents available and sheer number of 4- base long kmers for expanded alphabets (84 = 4096, 104 = 10000, and 124 = 20736). Additional strategies can be employed to boost accuracy. As a first pass, one can employ the strategy of sequencing both the positive sense and negative sense strands. Here too, the artificial nature of the AEGIS system is an advantage as one can construct the negative sense strand using an alternative set of AEGIS nucleotides which have differing error modes in nanopore sequencing. For example, since AEGIS nucleotides P and Z have high nanopore signal contrast (FIG. 2A), one can engineer other AEGIS nucleotides to increase their signal contrast relative to the rest of the alphabet.

[0085] Using molecular cloning techniques, strands including several concatenated “dual alphabets” (alphabets with the same base-pairing scheme but using different heterocyclic variants) can be combined and sequenced serially (FIGs 8A-8B). In the shown embodiments, the base-pairing nature of each base is preserved, and reads are from singlemolecules. Various single-molecule re-reading strategies can also be employed to further improve single-molecule sequencing accuracy.

[0086] In addition to full-factorial sequencing, the ability to sequence subsets of the full factorial kmer map can be useful. As noted above, 8-, 10-, and 12-letter versions of AEGISBodies and AEGISZymes are expected to include 4-6 “structural” bases along with a handful of other “functional” bases which are incorporated (“doped”) far more sparsely. Thus, the vast kmer explosion implied by de novo sequencing of all combinations of 8-,10-, or 12-letter DNA can be developed, or alternatively, with use of a limited alphabet subset, there can be extensive utility for AEGIS-LIVE. This can be considered a far smaller subspace, and sequencing of strands composed of this subset of kmers can be achieved by modification of the allowed transitions between kmer states within the sequencing model excluding kmers that include multiple functional groups.

[0087] Accordingly, in an aspect, the disclosure provides a method for preparing a DNA library for high-accuracy nanopore sequencing, the method comprising: contacting a DNA strand comprising a target polynucleotide sequence with a DNA polymerase and a dNTP pool comprising an analog dNTP that comprises a modification relative to a corresponding reference dNTP that is absent from the dNTP pool. The DNA polymerase incorporates the analog dNTP into a nascent DNA strand in place of the corresponding reference dNTP.

[0088] In various implementations, a particular target polynucleotide sequence can be copied using one or more dNTP pools comprising one or more analog dNTPs that are “swapped out” for one or more canonical dNTPs. The modified sequence(s) having the analog dNTPs incorporated therein and resulting from copying the target polynucleotide sequence can be compared with each other for increasing statistical confidence for nucleotide base-calling for the original target polynucleotide sequence.

[0089] In embodiments, an electrical resistance of the corresponding reference dNTP differs from an electrical resistance of the analog dNTP. In embodiments, the electrical resistance of the corresponding reference dNTP is greater than the electrical resistance of the analog dNTP. In embodiments, the electrical resistance of the corresponding reference dNTP is less than the electrical resistance of the analog dNTP. In embodiments, the modification of the analog dNTP increases an electrical charge of the analog dNTP relative to an electrical charge of the corresponding reference dNTP. In embodiments, the modification of the analog dNTP decreases an electrical charge of the analog dNTP relative to an electrical charge of the corresponding reference dNTP.

[0090] In embodiments, the method is isothermal and does not require heat to be applied for various steps of the method to be performed. This can be advantageous for library preparations involving many different nucleic acid sequences. For example, if strands are melted and re-hybridized, they can form off-target duplexes; by using strand displacing polymerases, off-target hybridization can be avoided. This is also advantageous because instrumentation for thermocycling the reaction is not necessarily required.

[0091] In embodiments, the contacting step is performed a plurality of times with a plurality of dNTP pools, wherein dNTP pools of the plurality of dNTP pools comprise pool-specific analog dNTPs that comprise pool-specific modifications relative to corresponding reference dNTPs that are absent from the dNTP pools. The contacting step can be performed two times, three times, four times, five times, or more. In at least some embodiments, the contacting step is performed three times. In embodiments, the method further comprises removing previous dNTP pools from a reaction mixture between contacting steps.

[0092] In embodiments, the removing step comprises immobilizing the DNA strand and the nascent DNA strand and washing the reaction mixture such that a previous analog dNTP is removed from the reaction mixture. In embodiments, the immobilizing step comprises attachment of a DNA molecule comprising the DNA strand, the nascent DNA strand, or both, to a substrate. Any of various substrates can be used, at least relatively immobile, configured to hold the DNA molecule comprising the DNA strand, the nascent DNA strand, or both, thereto.

[0093] In embodiments, a majority of dNTPs of the dNTP pool are not analog dNTPs and are, instead, canonical dNTPs; this can be beneficial at least in implementations wherein a direct comparison is made between a sequence read from one dNTP pool and a sequence read from another (different) dNTP pool during sequence analysis. In embodiments, no more than 25% of dNTPs of the dNTP pool are analog dNTPs. Maintaining a maximum threshold for percent analog dNTPs in the dNTP pool at a given step can help ensure that the copy of the target sequence maintains enough nucleobases in common with the original target sequence, which can be helpful for alignment of the copy of the target sequence with the original target sequence during analysis. However, in other embodiments, a majority of dNTPs of the dNTP pool are analog dNTPs; in at least some of these embodiments, 100% of dNTPs of the dNTP pool are analog dNTPs, and there are no canonical dNTPs in the dNTP pool.

[0094] In at least some embodiments, methods of sequence analysis are contemplated that make it possible to analyze or combine sequence reads from copied sequences comprising entirely different analog dNTP substitutions relative to each other. For example, with reference to the examples herein, one can simultaneously use the sense sequence S and antisense sequence S’ to inform the sequencer (FIGs 3, 4A-4B). For illustration of this concept, and with reference to these examples, S’ can be treated as acopy of the sequence S wherein A, T, C, and G (in S) have been swapped out for the alphabet T, A, G, and C, respectively (in S’). While sequencing S and S’ can improve sequencing accuracy, as this is currently done, S is sequenced separately from S’, then the letter sequences are combined. In embodiments, replacement of the entire set of dNTPs enables tuning of ion currents to maximize sequencing accuracy by sequencing the modified strand alone. The modified strand sequence can be compared against the sequence from the original strand, for example, as done when sequencing both sense and antisense strands (e.g., Pac Bio or ONT). These and other methods can be implemented at least in part by algorithms configured as logic code capable of being carried out, at least in part, by a computational device or processor circuitry.

[0095] In embodiments, the method further comprises ligating a sequencing adaptor to the DNA strand and the nascent DNA strand to configure the DNA library for nanopore sequencing. While any sequencing adaptor can be used, in various embodiments, the sequencing adaptor comprises a barcode that corresponds with an identity of the analog dNTP.

[0096] In embodiments, the method further comprises reverse transcribing an RNA strand to produce the DNA strand comprising the target polynucleotide sequence, wherein the DNA library represents at least part of an RNA fraction of a sample.

[0097] In an aspect, the disclosure provides a DNA library prepared according to a method of the disclosure. In another aspect, the disclosure provides a method for high- accuracy analysis of nanopore sequencing data obtained from nanopore sequencing of a DNA library prepared according to a method of the disclosure. The method comprises: aligning the nanopore sequencing data such that sequence data of the DNA strand is aligned with sequence data of the nascent DNA strand; computing an electrical property of the nanopore sequencing data that corresponds to an analog nucleobase of the nascent DNA strand for an analog nucleobase measurement; comparing the analog nucleobase measurement with a reference analog nucleobase measurement associated with a known analog nucleobase identity and assigning a nucleobase identity of the analog nucleobase consistent with the known analog nucleobase identity; and assigning, based on the nucleobase identity of the analog nucleobase, a nucleobase identity to an unknown nucleobase of the DNA strand that positionally corresponds to the analog nucleobase of the nascent DNA strand. These and other methods of the disclosure can be performed, in whole or in part, by computational devices and systems.

[0098] In embodiments, a plurality of nucleobase identities of analog nucleobases of the nascent DNA strand correspond to a plurality of nucleobase identities of unknown nucleobases of the DNA strand. In addition, while swapping of analog dNTPs (for canonical dNTPs) can be performed in parallel, such as by splitting a sample containing multiple copies of a target polynucleotide, in at least some embodiments the swapping of analog dNTPs (for canonical dNTPs) can be advantageously performed using sequent! al / additi ve modifi cati on .

[0099] Accordingly, in embodiments, the plurality of nucleobase identities of analog nucleobases of the nascent DNA strand comprises one distinct nucleobase identity that corresponds to one contacting step used for DNA library preparation. In embodiments, the plurality of nucleobase identities of analog nucleobases of the nascent DNA strand comprises two distinct nucleobase identities that correspond to one or more contacting steps used for DNA library preparation. In embodiments, the plurality of nucleobase identities of analog nucleobases of the nascent DNA strand comprises three distinct nucleobase identities that correspond to one or more contacting steps used for DNA library preparation.

[0100] In embodiments, the plurality of nucleobase identities of analog nucleobases of the nascent DNA strand comprises four distinct nucleobase identities that correspond to one or more contacting steps used for DNA library preparation.

[0101] In various embodiments, a method is performed at least in part by a programmable processor, a processor circuitry, a computational device, a computational system, a computational network, or any combination thereof. These and other devices for methods of the disclosure can comprise circuitry configured for performance of all or part of a method of the disclosure.

[0102] Accordingly, in various aspects, the disclosure also provides a processor, a processor circuitry, a computational device, a computational system, a computational network, or any combination thereof, comprising circuitry configured for performance of all or part of a method of the disclosure, in any order or combination of steps, by the processor, processor circuitry, computational device, computational system, computational network, or combination thereof, as the case may be.

[0103] In a further aspect, the disclosure also provides a kit comprising an instructional material and at least one element for performing a method of the disclosure.Circuitry, Processor, and Computer Implementations

[0104] Embodiments of devices and any systems disclosed herein, including embodiments that include or utilize a processor and / or processor executable instructions can utilize circuitry to implement those technologies and methodologies. Such circuitry can operatively connect two or more components, generate information, determine operation conditions, control an appliance, device, or method, and / or the like. Circuitry of any type can be used. In embodiments, circuitry includes dedicated hardware having electronic circuitry configured to perform operations or computations on a dedicated basis, without any use of microprocessors, central processing units, or software or firmware or processorexecutable instructions. However, in embodiments, circuitry includes, among other things, one or more computing devices such as one or more processors (e.g., microprocessor(s)), one or more central processing units (CPU), one or more digital signal processors (DSP), one or more application-specific integrated circuits (ASIC), one or more field- programmable gate arrays (FPGA), or the like, or any variations or combinations thereof, and can include discrete digital and / or analog circuit elements or electronics, or combinations thereof.

[0105] In embodiments, circuitry includes one or more ASICs having a plurality of predefined logic components. In embodiments, circuitry includes one or more FPGA having a plurality of programmable logic components. In embodiments, circuitry includes hardware circuit implementations (e.g., implementations in analog circuitry, implementations in digital circuitry, and the like, and combinations thereof). In embodiments, circuitry includes combinations of circuits and computer program products having software or firmware processor-executable instructions stored on one or more computer readable memories, e.g., non-transitory computer-readable storage mediums, that work together to cause a device or system to perform one or more methodologies or technologies described herein.

[0106] In embodiments, circuitry includes circuits, such as, for example, microprocessors or portions of microprocessors, that require software, firmware, and the like for operation. In embodiments, circuitry includes an implementation comprising one or more processors or portions thereof and accompanying software, firmware, hardware, and the like. In embodiments, circuitry includes a baseband integrated circuit or applications processor integrated circuit or a similar integrated circuit in a server, a cellular network device, other network device, or other computing device. In embodiments,circuitry includes one or more remotely located components. In embodiments, remotely located components (e.g., server, server cluster, server farm, virtual private network, etc.) are operatively connected via wired and / or wireless communication to non-remotely located components (e.g., desktop computer, workstation, mobile device, controller, etc.). In embodiments, remotely located components are operatively connected via one or more receivers, transmitters, transceivers, or the like.

[0107] Embodiments include one or more data stores that, for example, store instructions and / or data. Non-limiting examples of one or more data stores include volatile memory (e.g., Random Access memory (RAM), Dynamic Random Access memory (DRAM), or the like), non-volatile memory (e.g., Read-Only memory (ROM), Electrically Erasable Programmable Read-Only memory (EEPROM), Compact Disc Read-Only memory (CD-ROM), or the like), persistent memory, or the like. Further non-limiting examples of one or more data stores include Erasable Programmable Read-Only memory (EPROM), flash memory, or the like. The one or more data stores can be connected to, for example, one or more computing devices by one or more instructions, data, or power buses.

[0108] In embodiments, circuitry includes one or more computer-readable media drives, interface sockets, Universal Serial Bus (USB) ports, memory card slots, or the like, and one or more input / output components such as, for example, a graphical user interface, a display, a keyboard, a keypad, a trackball, a joystick, a touch-screen, a mouse, a switch, a dial, or the like, and any other peripheral device. In embodiments, circuitry includes one or more user input / output components that are operatively connected to at least one computing device to control (electrical, electromechanical, software-implemented, firmware-implemented, or other control, or combinations thereof) one or more aspects of the embodiment.

[0109] In embodiments, circuitry includes a computer-readable media drive or memory slot configured to accept signal -bearing medium (e.g., computer-readable memory media, computer-readable recording media, or the like). In embodiments, a program for causing a system to execute any of the disclosed methods can be stored on, for example, a computer-readable recording medium (CRMM), a signal -bearing medium, or the like. Nonlimiting examples of signal-bearing media include a recordable type medium such as any form of flash memory, magnetic tape, floppy disk, a hard disk drive, a Compact Disc (CD), a Digital Video Disk (DVD), Blu-Ray Disc, a digital tape, a computer memory, or the like, as well as transmission type medium such as a digital and / or an analog communicationmedium (e.g., a fiber optic cable, a waveguide, a wired communications link, a wireless communication link e.g., transmitter, receiver, transceiver, transmission logic, reception logic, etc.). Further non-limiting examples of signal-bearing media include, but are not limited to, DVD-ROM, DVD-RAM, DVD+RW, DVD-RW, DVD-R, DVD+R, CD-ROM, Super Audio CD, CD-R, CD+R, CD+RW, CD-RW, Video Compact Discs, Super Video Discs, flash memory, magnetic tape, magneto-optic disk, MINIDISC, non-volatile memory card, EEPROM, optical disk, optical storage, RAM, ROM, system memory, web server, or the like.Terminology

[0110] The complete disclosure of all patents, patent applications, and publications, and electronically available material cited herein are incorporated by reference in their entirety. Supplementary materials referenced in publications (such as supplementary tables, supplementary figures, supplementary materials and methods, and / or supplementary experimental data) are likewise incorporated by reference in their entirety. In the event that any inconsistency exists between the disclosure of the present application and the disclosure(s) of any document incorporated herein by reference, the disclosure of the present application shall govern. The foregoing detailed description and examples have been given for clarity of understanding only. No unnecessary limitations are to be understood therefrom. The disclosure is not limited to the exact details shown and described, for variations obvious to one skilled in the art will be included within the disclosure defined by the claims.[OHl] The description of embodiments of the disclosure is not intended to be exhaustive or to limit the disclosure to the precise form disclosed. While the specific embodiments of, and examples for, the disclosure are described herein for illustrative purposes, various equivalent modifications are possible within the scope of the disclosure.

[0112] Specific elements of any foregoing embodiments can be combined or substituted for elements in other embodiments. Moreover, the inclusion of specific elements in at least some of these embodiments can be optional, wherein further embodiments can include one or more embodiments that specifically exclude one or more of these specific elements. Furthermore, while advantages associated with certain embodiments of the disclosure have been described in the context of these embodiments,other embodiments can also exhibit such advantages, and not all embodiments need necessarily exhibit such advantages to fall within the scope of the disclosure.

[0113] As used herein, the term “dNTP pool” refers to a plurality of dNTPs, for example, of a composition or solution, that includes at least one analog dNTP.

[0114] As used herein, the term “analog dNTP” refers to any dNTP that is not a reference, naturally-occurring, or canonical dNTP.

[0115] As used herein, an “instructional material” includes a publication, a recording, a diagram, or any other medium of expression which can be used to communicate the usefulness of one or more elements of a kit of the disclosure for carrying out a testing method, a screening method, a treatment method, or other method of the disclosure, including methods for alleviation of one or more diseases or disorders as described herein or as known in the art. Optionally, or alternately, the instructional material can describe one or more methods of alleviating the diseases or disorders in a cell or a tissue of a mammal. The instructional material of the kit can, for example, be affixed to a container which contains an identified compound or can be shipped together with a container which contains the identified compound. Alternatively, the instructional material can be shipped separately from the container with the intention that the instructional material and the compound be used cooperatively.

[0116] As used herein, the term “non-transitory machine-readable storage medium”, refers to a computer readable medium or processor readable medium that can store information as data, as well as instructions for performance of logic operations by one or more processors for manipulation of the data. The non-transitory machine-readable storage medium can include any type of memory, such as volatile memory like random access memory (RAM), dynamic random-access memory (DRAM), static random-access memory (SRAM), or non-volatile memory like read-only memory (ROM), flash memory, magnetic or optical disks, or compact-disc read-only memory (CD-ROM), among other devices used to store data or programs on a temporary or permanent basis. The non- transitory machine-readable storage medium can be configured to store instructions. The instructions are executable by the one or more processors to cause the computing device to perform any of the functions or methods described herein. The non-transitory machine- readable storage medium can also be configured to store a computational model. The computational model can take the form of a convolutional neural network or any other type of artificial neural network. The computational model can take other forms as well.

[0117] As used herein, the term “computational device” refers to a device, e.g., an electronic device, capable of performing logic operations. Example computational devices include computers, laptops, and tablets, as are known in the art. A computing device includes one or more processors, a non-transitory computer readable medium, a communication interface, a display, and a user interface. Components of the computing device are linked together by a system bus, network, or other connection mechanism. The one or more processors can be any type of processor(s), such as a microprocessor, a digital signal processor, a multicore processor, etc., coupled to the non-transitory computer readable medium. The communication interface can include hardware to enable communication within the computational device and / or between the computational device and one or more other devices. The hardware can include transmitters, receivers, and antennas, for example. The communication interface can be configured to facilitate communication with one or more other devices, in accordance with one or more wired or wireless communication protocols. For example, the communication interface can be configured to facilitate wireless data communication for the computational device according to one or more wireless communication standards, such as one or more Institute of Electrical and Electronics Engineers (IEEE) 801.11 standards, ZigBee standards, Bluetooth standards, etc. As another example, the communication interface can be configured to facilitate wired data communication with one or more other devices. The communication interface can also include anal og-to-digi tai converters (ADCs) or digital- to-analog converters (DACs) that the computational device can use to control various components. The display can be any type of display component configured to display data. As one example, the display can include a touchscreen display. As another example, the display can include a flat-panel display, such as a liquid-crystal display (LCD) or a lightemitting diode (LED) display. A user interface can be included as part of the computational device, and can include one or more pieces of hardware used to provide data and control signals to the computing device. For instance, the user interface can include a mouse or a pointing device, a keyboard or a keypad, a microphone, a touchpad, or a touchscreen, among other possible types of user input devices. Generally, the user interface can enable an operator to interact with a graphical user interface (GUI) provided by the computing device (e.g., displayed by the display)

[0118] As used herein, the term “system” refers to one or more computational devices, or one or more elements thereof, configured to perform one or more tasks,methods, or processes. A system of the disclosure can comprise a computing device. The computing device includes one or more processors, a non-transitory computer readable medium, a communication interface, a display, and a user interface. Components of the computing device are linked together by a system bus, network, or other connection mechanism.

[0119] As used herein and unless otherwise indicated, the terms “a” and “an” are taken to mean “one”, “at least one” or “one or more”. Unless otherwise required by context, singular terms used herein shall include pluralities and plural terms shall include the singular.

[0120] Unless the context clearly requires otherwise, throughout the description and the claims, the words “comprise”, “comprising”, and the like are to be construed in an inclusive sense as opposed to an exclusive or exhaustive sense; that is to say, in the sense of “including, but not limited to”.

[0121] Unless the context clearly requires otherwise, the phrase “consisting essentially of’ limits the scope of a claim to the specified materials or steps and those that do not materially affect the basic and novel characteristic(s) of the claim.

[0122] Unless the context clearly requires otherwise, the phrase “consisting of’ excludes any element, step, or ingredient not specified.

[0123] If an element is described or claimed herein such that it “comprises” a feature, that description or claim also includes embodiments wherein the element “consists essentially of’ and embodiments wherein the element “consists of’ the feature, unless something else is specifically stated to the contrary.

[0124] Words using the singular or plural number also include the plural and singular number, respectively. Additionally, the words “herein,” “above,” and “below” and words of similar import, when used in this application, shall refer to this application as a whole and not to any particular portions of the application.

[0125] Unless otherwise indicated, all numbers expressing quantities of components, molecular weights, and so forth used in the specification and claims are to be understood as being modified in all instances by the term “about.” Accordingly, unless otherwise indicated to the contrary, the numerical parameters set forth in the specification and claims are approximations that can vary depending upon the desired properties sought to be obtained by the present disclosure. At the very least, and not as an attempt to limit the doctrine of equivalents to the scope of the claims, each numerical parameter should at leastbe construed in light of the number of reported significant digits and by applying ordinary rounding techniques.

[0126] Notwithstanding that the numerical ranges and parameters setting forth the broad scope of the disclosure are approximations, the numerical values set forth in the specific examples are reported as precisely as possible. All numerical values, however, inherently contain a range necessarily resulting from the standard deviation found in their respective testing measurements.

[0127] All headings are for the convenience of the reader and should not be used to limit the meaning of the text that follows the heading, unless so specified.

[0128] All of the references cited herein are incorporated by reference. Aspects of the disclosure can be modified, if necessary, to employ the systems, functions, and concepts of the above references and application to provide yet further embodiments of the disclosure. These and other changes can be made to the disclosure in light of the detailed description.

[0129] It will be appreciated that, although specific embodiments of the disclosure have been described herein for purposes of illustration, various modifications can be made without deviating from the spirit and scope of the disclosure. Accordingly, the disclosure is not limited except as stated by the claims.ExamplesExample 1. Nanopore sequencing DNA by replacement of nucleotides with analogs.

[0130] The disclosure provides approaches for generating DNA strands which enable multiple reads of the same primary sequence in multiple alternative alphabets. As an example, using phosphoramidite-synthesized oligos assembled by ligation, a short, selfpriming hairpin structure was created. For simplicity, this strand included a 4-letter ACGT- DNA. The sequence was elongated using 5-methyl cytosine (5mC) and 5 -hydroxymethyl cytosine (5hmC) triphosphates in place of normal cytosine (C) and then was ligated to the nanopore sequencing adaptor as in FIGs 5A-6C. The product was validated using gel electrophoresis which showed promising results (FIG. 7A). Unsuccessful ligation products were attributed to incomplete A-tailing.

[0131] Using a motor enzyme for nanopore sequencing, such as PcrA-X, sequencing reads of the desired strand can be obtained. FIG. 7B shows one such readaligned to the predicted current values for unmodified C. Replacement of C with 5mC significantly alters nanopore ion currents, and reads of the target sequence in both ACGT and A(5mC)GT alphabets can yield unique sequence information and significantly increase sequencing accuracy. Other modifications can have larger effects, and this approach can be extended to use multiple modified bases, including with use of target sequences derived from plasmids and hairpin adapters synthesized enzymatically via PCR.Example 2. Binders on demand and alternative approaches to them.

[0132] Molecules that interact with biomolecules in complex environments are used throughout medical research, diagnostics, and therapy. Antibodies are frequently used for pull-down assays to interrogate chromatin organization and epigenetic markers, e.g. chromatin immuno-precipitation, CUT&RUN, etc. In medicine, antibodies are used for diagnostics, as for example in liquid biopsies that aim to detect small numbers of cancer cells circulating in complex biological fluids. Antibodies are central to immunotherapy, as drugs themselves (e.g., herceptin), or conjugated to drugs (e.g., Kadcyla), where the antibody delivers toxic drugs (e.g., emtansine) to tumor cells.

[0133] Generally, antibodies are generated by slow Darwinian evolution in animal immune systems. If the desired antibody is not in stock, a customer must pay to raise it (or an equivalent), a process of months and considerable expense. Furthermore, there are limitations to what antibodies can achieve. Even when they are in hand, antibodies as protein biologies have severe limitations. In particular, antibodies are either polyclonal or converted to monoclonal form. Here, “one gets what one gets”. Having a monoclonal antibody in hand allows researchers to do what that particular monoclonal can do, but only that. Researchers have little opportunity to modify an antibody to get a set of antibodies with a range of affinities, a reagent tool kit that would be valuable to analyze systems over a dynamic range. Further, they cannot be evolved to do things outside the evolvable repertoire of the immune system. Also, antibodies as biologies have been implicated in “irreproducibility crises” in medical research. The limits of antibody technology also highlight the unmet need for reagents that deliver selective and potent binding “on demand” to research, diagnostics, and therapeutics. The extent of the unmet need is illustrated by the many efforts to replace antibodies by other protein scaffolds, such as fibronectins, darpins, and ankyrins. These were hoped to be easier to create, diminish cost, and manage irreproducibility. However, these scaffolds give binders with only nM affinities. Clinicaloncologists also tolerate these problems, especially with antibody-drug conjugates, because they need specific binding, and have no better way to get it.

[0134] This background drives the significance of various aspects of this disclosure, which provides nanopore sequencing approaches that can enable for rapid and inexpensive creation and sequence verification of macromolecules that bind targets for a significant impact in medicine and other industries. Such evolvable biomolecules could be raised against whole cells and counter-selected to not bind other cells; be able to chemically transform targets or attach themselves covalently to those targets; and / or be able to act after entering living cells or crossing the blood-brain barrier.

[0135] Aptamers and catalysts (aptazymes) can be rapidly obtained by applying selective pressure to libraries of natural DNA and RNA (nucleic acids; NA). Variously called in vitro selection, “Systematic Evolution of Ligands by Exponential enrichment” (seLex), or “laboratory in vitro evolution” (Live), these techniques expose diverse libraries of short NA sequences to iterative rounds of selective pressure for desired traits (e.g., binding) followed by replication with low-level mutation. A few rounds of such a process can produce new binders and catalysts comprised of short NA strands within a relatively short time-span.

[0136] After this random evolution, the successful aptamer can be sequenced to “trim” it, optimize its performance, and attach other things. Further, an advantage of the sequenced aptamer over antibodies is that aptamers can later be synthesized as defined chemical entities and are no longer irreproducible biologies. The binding of aptamers made from regular 4-letter NA was initially hoped to “rival antibodies”. This potential has not yet been realized. NA binders for targets such as carbohydrates, small molecules, and peptides are now known, and some have entered the clinic after modification of the species originally selected. But the affinities of NAs built from standard 4-letter DNA with unfunctionalized building blocks does not match the pM affinities of antibodies.

[0137] Compared to proteins, functional RNAs can have a far more complex alphabet containing well over 100 known modified bases. For this reason, functional groups can be added to the four “letters” of the standard DNA / RNA “alphabet”. SOMAamers add a hydrophobic side chain to uracil in their evolvable system and have a hydrophobic group attached on average to every 4th nucleotide.

[0138] There is some support for the premise that adding functional groups to a DNA library increases its intrinsic value as a source of catalytic species. However, addingfunctional groups to standard nucleotides does not solve all of the problems with standard 4-letter (ACGT) aptamers and aptazymes. One problem arises because of folding ambiguity. With just four building blocks (and just two pairs, A:T and G:C), standard nucleic acid molecules have difficulty avoiding alternative folds that have similar energy. This gives ambiguous folding. Further, over-functionalizing NA creates problems of its own. If every copy of a particular base has functionality, DNA no longer behaves like DNA, especially if the group is hydrophobic. However, adding hydrophobic groups sparingly, one or two per molecule, gives DNA molecules antibody-like affinity without nonspecificity. Thus, sub-picomolar aptamers for targets such as the protein VEGF can be obtained by this approach. This establishes another premise: An evolvable system that allows hydrophobic units to be added sparingly to evolved nucleic acid products results in receptors and ligands that can replace antibodies.

[0139] The disclosure provides approaches for nanopore sequencing including results from re-engineering DNA to get a better evolvable platform, with more nucleotides (now 12), more functional groups, better folding, more opportunities for compact folds, sub-picomolar affinity, and increased catalytic power, complete with organic synthesis pipelines, replicating polymerases and analytical chemistry for this “artificially expanded genetic information system” (Aegis).

[0140] By adding additional complexity to the DNA alphabet, new DNA folds become possible, further increasing the structural variability accessible to the evolvable polymer. Standard 4-letter DNA has one base:base interaction that forms a compact fold, the “G quadruplex”. It often emerges when evolving 4-letter DNA; some bias libraries to be G-rich get its evolution. Aegis DNA has many additional base-base interactions that create core folds. Rokumoji DNA (ACGTPZ) has one new fold of particular interest, the Z:Z- “fZ-motif ’. NMR and biophysical studies of the fZ motif shows that it gives compact folds with parallel strands. Here, it is analogous to the i-motif formed by protonated C123.

[0141] However, this added complexity makes prior efforts at sequencing a challenge. Prior sequencing technologies exist for only the six-letter subset of Aegis (ACGTPZ, FIG. 5 A), however, these approaches utilize a cumbersome conversion of P and Z probabilistically into A or G and C or T, respectively. This strategy does not scale well to larger alphabets or complex libraries. Sequence data are used to rationally modify AegisConstructs enabling drug delivery, assessment of catalytic mechanisms,understanding novel folds, and re-designing of Aegis building blocks to give libraries that are richer reservoirs for targeted functions.

[0142] Even with just half of the 12 Aegis building blocks able to be sequenced, Aegis-Live is delivering reagents, ligands, and catalysts with five orders of magnitude better performance than those emerging from standard Live based on a 4-letter DNA alphabet. Live applied to Rokumoji DNA libraries gave RokuBodies that bind toxins, breast cancer cells, liver Hep G2 cancer cells, engineered cells, and proteins such as VEGF. RokuZymes were evolved that cleave targeted RNA sequences. Aegis nanotrains that hold ~50 doxorubicin drug molecules selfattached themselves to RokuBodies that bind liver cancer cells; the conjugate selectively delivered drugs to the liver cancer cells (not normal liver cells) and selectively kill them. Here, the RokuBody replaced an antibody. Thus, the significance of increasing the number of building blocks from 4 to 6 in an evolvable biopolymer have been established. Adding additional bases will give still better control over folding, with more compact folds and with more functional groups leading to still better outcomes. If Live could be extended to the full Aegis platform, with all 12 building blocks, more functional groups, folds, and performance could be transformative. Additional folds, additional functional groups, and additional control over secondary structure emerge with this expansion. DNA can be made more like protein, but better in this application, because it’s more intrinsically soluble and it is itself directly evolvable.

[0143] The lack of adequate sequencing approaches has prevented extension of Live to the full Aegis alphabet. The sequencing tools to sequence the products that 8-, 10-, or 12-letter Aegis-Live would deliver are provided by various aspects and embodiments of this disclosure.

[0144] Development of NGS technologies involved investment and development of polymerases, fluorescently-labeled bases, and reversible terminators for ACGT. Adapting NGS to enable sequencing of DNA with more than 6 nucleotides is prohibitively expensive. Nanopore sequencing technology offers an attractive alternative. It is “label free”, and can easily distinguish various non-canonical DNA and RNA bases. The principal challenge in adapting nanopore DNA sequencing for Aegis-systems is that nanopore sequencing reads NA in sets of overlapping 4-base-long “kmers.” Each k-mer is associated with a specific ion current. For 8-, 10- and 12-letter DNA there are 4096, 10000, and 20736, respectively, making it increasingly difficult to resolve individual kmers with confidence. Rather than tackle full de novo sequencing head on, preliminary work focused on thequestion of feasibility. Is it indeed possible to sequence 8-, 10- and 12-letter DNA with nanopores given the “kmer explosion?” It is proposed herein to use nanopores to sequence Aegis DNA. In this example, the eight-letter Hachimoji subset (GACTZPSB) of the full Aegis 12-letter alphabet was used as a non-limiting example. Synthetic Hachimoji DNA was presented for sequencing using the MspA (Mycobacterium smegmatis porin A) nanopore.

[0145] The results were surprisingly and unexpectedly good. In summary: Result 111. (FIG. 2A) Hachimoji DNA exhibits a broader nanopore ion current signal range than standard DNA alone. This means that adding letters need not force the nanopore sequencing approach to resolve hopelessly unresolvable nucleotides. Using variable voltage nanopore sequencing, the Hachimoji single base substitutions are distinguishable with high confidence (FIGs 2B-2C). The helicase Hel308 can be used to pull Hachimoji DNA through MspA. In addition, Hel308 tends to dissociate more frequently when walking over C-glycoside nucleosides, compared to N-glycosides. These results highlight the need to develop nanopore sequencing motors to handle different glycosidic bonds for sequencing of long DNA strands containing C-gly cosides. However, Hel308 already has sufficient processivity to sequence short strands which form AegisBodies and AegisZymes.

[0146] The ability of the Oxford Nanopore Technologies sequencer to identify sparse incorporations of PZSBJVXK within a background of ACGT was evaluated. The ability to identify such sparse insertions was surprisingly good and indicative of an ability for full de novo sequencing of 12-letter DNA. In sum, these results show that alternative DNA systems can be sequenced with existing motors and pores, and the next challenge involved extracting more information out of the nanopore signal.

[0147] This disclosure provides devices, systems and methods to sequence yet- more-expanded DNA molecules; increase the power of Aegis-Live, to create sub- picomolar receptors and ligands on demand, with turnaround times in weeks not today possible; create replacements for antibodies by chemicals that suffer none of their cost or challenges as biologies; get macromolecular receptors and ligands into cells, which is impossible to date; offer the possibility of evolvable drugs that manipulate targets that they bind; and offer the possibility of “personalized pharmacopeias”, where the low cost and rapid turn-around of Aegis-Live allows, for example, anti-cancer drug delivery systems to evolve with the evolution of cancer in a specific patient.

[0148] These approaches are novel, in a broad sense, to many fields. This novelty encompasses many areas including at least synthetic chemistry, physical organic chemistry, analytical chemistry, structural biology, molecular evolution, and medicine. The disclosure is to enable de novo sequencing of expanded alphabets with meaningful accuracy. Given the limitations of commercial nanopore sequencing, straightforward application of the technology or incremental improvement may be insufficient. Sequencing innovations in the course of this project will have impact beyond the sequencing of Aegis-DNA. Techniques and lessons learned will be directly applicable to ongoing NHGRI-supported efforts in sequencing DNA, RNA, and peptides, identifying DNA and RNA modifications.

[0149] Aegis alphabets, and nanopore sequencing, can be refined and dramatically improved to enable sequencing of expanded DNA alphabets, de novo sequencing of 8-, 10-, and 12-letter DNA is not possible using current techniques. Innovations described herein make it possible. This is a new and ambitious application of nanopore sequencing to de novo sequence the large kmer set implied by 6-, 8-, 10- and 12-letter DNA.

[0150] The choice of using MspA-based variable-voltage sequencing (FIGs 1A- 1D) can be beneficial as additional bases are added to the DNA alphabet, due to advantageous information content. In addition, the existing pore and enzyme used by Oxford Nanotechnologies (ONT) is tuned for the sequencing of 4-letter, ATCG DNA, and cannot be readily tuned to improve sequencing with 6-, 8-, 10- and 12-letter DNA. Further, the kmer size of the current “R10 chemistry” from ONT is estimated to be 9-10 nucleotides. This large kmer size means that fully trained sequencing models can involve massive datasets to fully train (an 8-letter alphabet would have 810 > 1 billion unique kmers).

[0151] There have been polymerases capable of replicating the 6-letter alphabet consisting of ACGTPZ, however, this expanded alphabet produces binders with pM affinities and catalytic abilities. As noted herein (FIG. 2A), P generally produces currents lower than the four standard bases, while Z produces ion currents that are substantially higher than the four standard bases. From these data, it is conceivable that direct nanopore sequencing of 6-letter DNA is possible without modification, especially if P and Z are incorporated sparsely within the strand (as they are for existing 6-letter AegisBodies.) Sequencing of this 6-letter alphabet would have immediate utility in the further development and refinement of a number of AegisBodies and AegisZymes currently at various stages of production. Typical Live experiments yield a library of AegisBodies containing just a handful of the most dominant binders in the form of -20-50 basesequences flanked by replication primers of known DNA sequence. The goal of this sub aim is to consensus sequence one such library with > 90% accuracy.

[0152] Given the apparent ion current separation of P and Z from natural bases (FIG. 2A), sequencing of 6-letter is possible as an extension of existing nanopore technology. Accuracy can be achieved by building empirical datasets which can be used to train sequencing models. This, in combination with practices such as consensus sequencing and obtaining duplex reads, can be used to achieve the needed accuracy.

[0153] Furthermore, heterocyclic variants of each base retain the base-pairing identity of their respective bases but have different chemical characteristics, which have various uses in AegisBodies and AegisZymes. Modified bases can be readily distinguished from one another in the nanopore, and slight chemical differences can lead to profound differences in nanopore ion current, and as such, sequencing of alternative versions of P and Z can be performed. Furthermore, analogs such as 5-methylcytosine, 5- hydroxymethylcytosine, 5-formylcytosine, and 5-carboxycytosine, all of which base pair in the same way as C does with G, exist for all other natural bases and can be synthesized as triphosphates which are compatible with replicative polymerases. Chemists have developed still more modified bases for various biochemical techniques. Similarly, the bases of Aegis-DNA have many heterocyclic variants (FIG. 3). As such, one could use an alternative alphabet comprised of several modified bases to significantly boost the accuracy of nanopore reads. In addition, one could use a polymerase with an alphabet of modified dNTPs to generate a strand with a high-contrast nanopore signal (FIGs 4A-4B). In this scheme, base mods can be selected to ensure maximum sequencing accuracy. Such a scheme can enable direct sequencing of 6-letter DNA by choosing an alphabet that reduces the similarities in ion current between different kmers. This can be achieved by targeting bases which are particularly hard to distinguish (such as S and G in FIGs 2B-2C) to try to make them more distinguishable.

[0154] For applications involving high single-molecule sequencing accuracy, such as sequencing diverse, 6-letter libraries in which duplicate reads will be rare or sequencing of still larger alphabets, such as 8-, 10-, and 12-letter DNA, various approaches can be implemented. The number of kmers within these expanded alphabets will make even variable voltage nanopore sequencing and consensus sequencing insufficient. A strategy involves obtaining “duplex reads,” in which both the sense strand S and antisense strandS’ are read by the nanopore. The errors inherent in the antisense read are often orthogonal to those of the sense read, facilitating construction of the sequence.

[0155] Furthermore, the complementary strand is itself an alternative alphabet in which A is swapped for T, C for G, G for C, and T for A. Building on this further using the ideas of FIGs 4A-4B, there is a way of generating still more orthogonal reads in additional alternative alphabets. This scheme is called multi-alphabet substitution sequencing. Multialphabet sequencing is a novel library preparation technique in which iterative rounds of strand-displacing replication using base-analog dNTP alphabets generate a concatemer strand containing multiple successive copies of a sequence S and its complement S’ (FIG. 5B). A nanopore read of the concatemer strand will yield significantly more sequence information because the different alphabets’ constituents can be selected to complement one another’s weaknesses. Such a concatemer strand would include at least 4 duplicate reads of the original sequence: i) the primary sequence in alphabet 1, ii) the complement in alphabet 1, iii) the primary sequence in alphabet 2 and iv) the complementary sequence in alphabet 2. Information from a single read of a multi-alphabet concatemer can yield high single-molecule sequencing accuracy, leveraging the orthogonal information available in each of the four encodings of the original sequence. If still more information is needed, additional rounds of replication with additional alphabets can be made repeating steps 5 and 6.

[0156] This scheme can involve implementing sets of alternative alphabets with orthogonal sequence information and sequencing algorithms capable of combining reads from several alphabets. Once implemented, their use can be put to practice to demonstrate their utility.

[0157] Within the dynamic range of the nanopore ion current measurement (~40pA to ~100pA for the standard conditions), characteristic ion currents are not uniformly distributed. Many kmers bunch up in the middle, while fewer lie at the extremes (FIG. 6C). C and G comprise much of the highly degenerate peak at ~80 pA. One can expand and flatten this distribution by selective modification of these bases, which produce mid-level currents. Various base modifications can either increase or decrease the measured ion current relative to the baseline ion current. For example, 5 -methyl cytosine has a higher current than standard C while 5-hydroxymethylcytosine has a lower ion current, therefore, it is possible to use non-standard nucleotides to fine-tune the distinguishability of bases not well-resolved in the original alphabet.

[0158] Results from selectively modified bases can be used to create alphabets which complement one another’s weaknesses. The scheme of FIG. 6B enables nanopore reads of the same primary sequence in four separate alphabets. Modified Aegis components can be authenticated. Sequencing accuracy with a particular sequencer can be tested using blind strands generated by FfAME. 6-letter ACGTPZ DNA will serve as the proving ground for this technique. It is anticipated that multialphabet sequencing is capable of achieving arbitrarily high single-molecule sequencing accuracies for short DNA strands such as AegisBodies. Because one can continue to tack on additional copies of the sequence S and S’ in still more different alphabets. Accuracy may be limited by the purity of modified dNTPs and polymerase fidelity.

[0159] The sheer number of possible modifications to each of the 6 bases ACTGPZ means that there will be a large parameter space to search for optimal sequencing alphabets, and one can seek to produce alphabets whose errors complement one another well. For example, in standard ACGT sequencing with M2-MspA, the natural bases A and T are easily identified. Thus a secondary alphabet would optimize for identification of C and G. Combining information from reads in both alphabets can enable sequencing with high accuracy from a single-molecule read. Modified bases themselves are also modular and can be mixed and matched across alphabets once their effect on the ion current is understood.

[0160] Collecting sufficient empirical data to fully train sequencing models involves significant amounts of data, however, one can evaluate the sequencing power of a given alphabet with significantly less data. Tools for evaluating the mutual information between a nanopore signal and the measured DNA sequence allows to evaluate the sequencing power of a given alphabet after sampling a small subspace of the full kmer- map. Comparing ion current signals from just a few oligos can enable rapid and robust comparison across different alphabets. One can develop standardized sequences which provide a sparse sampling of the kmer space that can be used to determine the sequencing power of a given alphabet, which will be produced via solid phase synthesis by FfAME and evaluated for sequencing power based on their ability to resolve various nucleotides.

[0161] As noted above, ultra-high fidelity PCR replicative polymerases exist for the ACGTPZ alphabet, and have been developed for PZSB alphabets, and transcriptive RNA polymerases exist for the GACUZPSB alphabet. Further, a feature of Aegis is that its components are themselves continuously evolving and base “cores” replaced with those that interact better with enzymes, such as polymerases (for better PCR amplification) orribosomes (to let expanded DNA / RNA alphabets encode extra amino acids in an expanded protein lexicon). This development is not slowing as more teams internationally seek to become involved in this new molecular biology. This development would benefit from a degree of flexibility in nanopore sequencing that mirrors the flexibility implemented when using nanopores to sequence, for example, modified RNA nucleotides. In parallel with this progress, the goal of this sub aim is continuing to develop high-contrast alphabets that can be used to sequence these increasingly complex alphabets. Development of polymerases capable of replicating these expanded alphabets go hand-in-hand with efforts to sequence them, as one may use such polymerases to do the work of multi-alphabet library preparation. Sequencing can also play a role in evaluation of polymerase fidelity.

[0162] One can also develop multiple sequencing alphabets that can be used in concert to determine a strand’s primary sequence. Rather than hope for individual sequencing alphabets which are capable of sequencing all 8, 10, or 12 bases simultaneously, one can adopt a modular approach. One can develop alphabets in which 2 or 3 individual non-complementary bases can be well distinguished (ie., A and C). By acquiring reads of both sense and antisense sequences of such an alphabet, it can be possible to determine the locations of all A, C, G, and T bases because the antisense read reveals A&C in the antisense sequence or T&G in the primary sequence. In a modular fashion, individual alphabets can then contribute information about a subset of the bases in 8-, 10- or 12-letter DNA. Combining this information can yield the sequence with good accuracy.

[0163] Even with this technique, one can anticipate that sequencing of 12-letter DNA can be exceptionally challenging. One may have to add additional rounds of replication with a third and perhaps 4th alphabet. In the scheme of FIGs 4A-4B, each additional iteration of steps 5 and 6 doubles the length of the underlying strand. Because the AegisBody sequences are quite short (20-50 nucleotides), this is not expected to be a problem for the first several rounds of replication. (Note there is no thermocycling in this process so aggregation of complementary concatemers is not an issue.) Still, if substantially more alternative alphabets could be beneficial to achieve sequencing accuracy, one can pursue other biochemical techniques to instead support linear strand growth during replication.

[0164] Machine Learning (ML) based sequencers serve as the primary sequencing models for nanopore sequencing and are fast and effective. Simultaneous use of data streams from multiple different alphabets to determine the primary sequence of a DNAstrand is a ML problem in which multiple features are used as inputs to produce the desired output.

[0165] With a set of sequencing alphabets, one can construct training libraries for ML via a combination of micro-array solid-phase synthesis and polymerase extension. Each sequence produced by the microarray can have a unique barcode that can be used to establish the ground truth sequence and also include a sense and antisense sequence region and a self-priming hairpin. These micro-array libraries can then be subjected to one round of polymerase extension in which the primer is extended with a secondary alphabet (plus a third or fourth as may be used by larger alphabets as described above). Libraries can then be nanopore sequenced ensuring a broad random sampling of the available sequence space. Depending on the read throughput and library complexity, there may not be duplicate nanopore reads of any given training oligo. This is by design to ensure that each read provides maximal information about the current-to-sequence map without redundancy. Strand barcodes can be long enough to enable easy identification of each read’s ground truth. As the training set grows, some training data can be set aside and used for benchmarking. The strand’s ground truth can be known from the barcode.

[0166] ML approaches benefit significantly from large training sets which fully capture the range of variation present within the data. Generation of sufficiently large, complex libraries containing expanded alphabets can be replaced or combined with generation hidden-Markov model (HMM-based) which need considerably less training data. One can begin sequencing with even just one representative measurement for each of the possible 2-mers, albeit with low accuracy.

[0167] Sequencing of nanopore data was first achieved using HMM-based models and HMM-based models are still used in various applications. While ML-based sequencers perform quite well and are exceptionally fast, which is useful in high-throughput nanopore experiments with limited computation, HMM- sequencers still find utility for applications in which large training sets are unavailable.

[0168] HMM sequencers have the added benefit of not being “black box” sequencers so one can understand how raw data is converted into sequence. Such sequencing models can be developed with multiple data streams. Sequencers can be developed which simultaneously take into account measurements from various applied voltages, or which use information from both ion current and the sequence dependent stepping of the motor enzyme. One can use these lessons in developing a HMM-basedsequencer which simultaneously uses reads from different alphabets. Duplex sequencing via HMM frequently makes independent measurements of the individual reads and merges them by alignment of the base sequences. Discrepancies between the two reads are presumably resolved by evaluation of respective base-calling confidence. This is a sub- optimal approach, because the joint distributions contain significantly more information than either the sense read or antisense read alone. One can sequence using the information from the sense and antisense strands simultaneously.

[0169] Connected to the work in developing the motor enzyme’s sequencedependent kinetics into a second nanopore “read head,” one can develop HMM-based frameworks which can take multiple different data streams into account. This method essentially breaks down into a multiple-sequence alignment (MSA) problem (FIG. 9C). Performing sequencing in this manner preserves the partial information from each independent data stream and determines the best possible sequence given all of the available data.

[0170] The MSA approach can be computationally intensive, especially as more and more separate data streams are incorporated. Thankfully, the MSA problem has been explored extensively in the literature in the context of alignment of biological sequences. It has been shown that nearly all entries of a MSA hyper-cube can be excluded from calculation using pair-wise alignments of all constituent reads (FIG. 9D). The accrued “cost” of stepping to the vast majority of hyper-cube elements is such that they need not be calculated. Validate sequencing on real libraries and begin to plug it into the development of new binders, catalysts and reagents.

[0171] AegisBodies and AEGISZymes have been created using the 6-letter alphabet ACGTPZ. The sequences of these AegisBodies and AEGISZymes are comprised within libraries that can enable validation of a nanopore sequencing work-flow. AegisBodies that bind to (i) Binding immunoglobulin protein (BiP) also known as 78kDa glucose-regulated protein (GRP-78), for which an AegisBody exists (“LH5b”), (ii) Ku90, encoded by the human XRCC5 gene, for which an AegisBody exists (“LZH8”), and (iii) Human serum albumin, a generic carrier protein for AegisBodies. These can serve as a proving ground for other practical sequencing applications for libraries of 6-letter DNA. One can follow sequencing protocols developed in the course of Aims 1 & 2 using specialized sequencing alphabets. Should Aims 1 and 2 prove successful, one can anticipatesequencing such 6-letter DNA with >90% accuracy. The existence of the ground truth sequencing data can enable error characterization and further validate the methods.

[0172] One can apply sequencing to aid in the further development of Aegis systems. Of particular interest is supporting the development of polymerases capable of using additional bases. While a strategy for sequencing DNA with high accuracy relies on such polymerases, the burden of proof for validation of a known sequence (essentially reference sequencing) is lower than that used for de novo sequencing. Thus, using a small subset of sequences which can be independently synthesized via solid phase synthesis, one can use the nanopore to validate strands made by evolved polymerases. In this manner, the nanopore can resolve polymerase error rates as small as a few percent and, given the right controls, could even help determine the mode of polymerase error.

[0173] Sequences can be validated by using them to generate a second generation of AegisBodies that can be tested for affinity vs. those derived from Aegis-Live. In each of these targets, sequencing “fidelity” can be borne out in the binding affinity achieved by AegisBodies derived from nanopore sequencing reads. This can also enable determination of the AegisBody structure enabling pruning of unnecessary loops and attachment of other cargomolecules for downstream applications. Sequencing of Aegis-DNA is envisioned in this proposal as an enabling technology, success can be gauged on this metric.

[0174] One can establish existing AEGIS components as authenticated and reproducible reagents through their manufacturing and sale through an affiliate, Firebird Biomolecular Sciences LLC. Nanopore sequencing, as enabled by the disclosure, can also be important for further improving the authentication / reproducibility pipeline for DNA molecules coming from these reagents to enable to ensure that AEGIS components, AEGISBodies, AEGISZymes, and AEGIS constructs are effective and high purity reagents. While multiple reads of the same primary sequence in various alternative alphabets as enabled by this disclosure can facilitate the kmer-explosion of Aegis-DNA, the disclosed library preparation schemes can also enable single molecule sequencing of natural ACGT DNA and ACGU RNA to almost arbitrary precision with significantly lower coverage than is used by standard nanopore sequencing workflows. For example, one can use a conspicuously modified A in one alphabet and a modified G in a second alphabet allowing straightforward identification of all bases. If more certainty is desired, a third and fourth alphabet can be used, as disclosed herein.

[0175] There are numerous applications for high-accuracy (>Q30), kilobase-long, single-molecule reads of DNA and RNA that are enabled by this disclosure. Alternative ACGT bases developed as part of the alternative alphabets also have direct utility in multialphabet sequencing of naturally occurring, genetically engineered, and biomedically relevant DNA and / or RNA. The library preparation technique also preserves the original strand as a part of the sequenced molecule. This means that genomic DNA or natural RNA containing modifications can be simultaneously sequenced to high precision and simultaneously screened for modifications or damage. This can be revolutionary as it can address the single most important limitation of nanopore sequencing: low single-molecule accuracy.Tables

[0176] Table 1. Examples of analog dNTPs useful by aspects and embodiments of the disclosure.

[0177] Table 2. Example analog base pairs, the sequencing of which is improved due to aspects and embodiments of the disclosure. Aspects and embodiments of this disclosure can aid in sequencing at least the following base pairs, including those in the art of synthetic biology.NON-LIMITING EMBODIMENTS

[0178] While general features of the disclosure are described and shown and particular features of the disclosure are set forth in the claims, the following non-limiting embodiments relate to features, and combinations of features, that are explicitly envisioned as being part of the disclosure. The following non-limiting Embodiments contain elements that are modular and can be combined with each other in any number, order, or combination to form a new non-limiting Embodiment, which can itself be further combined with other non-limiting Embodiments.

[0179] Embodiment 1. A method for preparing a DNA library for high- accuracy nanopore sequencing, the method comprising: contacting a DNA strand comprising a target polynucleotide sequence with a DNA polymerase and a dNTP poolcomprising an analog dNTP that comprises a modification relative to a corresponding reference dNTP that is absent from the dNTP pool; wherein the DNA polymerase incorporates the analog dNTP into a nascent DNA strand as an analog nucleotide in place of incorporation of the corresponding reference dNTP as a corresponding reference nucleotide.

[0180] Embodiment 2. The method of Embodiment 1 or any other Embodiment, wherein an electrical resistance of the corresponding reference nucleotide differs from an electrical resistance of the analog nucleotide.

[0181] Embodiment s. The method of Embodiment 2 or any other Embodiment, wherein the electrical resistance of the corresponding reference nucleotide is greater than the electrical resistance of the analog nucleotide.

[0182] Embodiment 4. The method of Embodiment 2 or any other Embodiment, wherein the electrical resistance of the corresponding reference nucleotide is less than the electrical resistance of the analog nucleotide.

[0183] Embodiment 5. The method of any of Embodiments 1-4 or any other Embodiment, wherein the modification of the analog nucleotide increases an electrical charge of the analog nucleotide relative to an electrical charge of the corresponding reference nucleotide.

[0184] Embodiment 6. The method of any of Embodiments 1-4 or any other Embodiment, wherein the modification of the analog nucleotide decreases an electrical charge of the analog nucleotide relative to an electrical charge of the corresponding reference nucleotide.

[0185] Embodiment 7. The method of any of Embodiments 1-6 or any other Embodiment, wherein the method is isothermal.

[0186] Embodiment 8. The method of any of Embodiments 1-7 or any other Embodiment, wherein the contacting the DNA strand is performed a plurality of times with a plurality of dNTP pools, wherein dNTP pools of the plurality of dNTP pools comprise pool-specific analog dNTPs that comprise pool-specific modifications relative to corresponding reference dNTPs that are absent from the dNTP pools.

[0187] Embodiment 9. The method of any of Embodiments 1-8 or any other Embodiment, wherein the contacting the DNA strand is performed three times.

[0188] Embodiment 10. The method of any of Embodiments 1-9 or any otherEmbodiment, further comprising removing previous dNTP pools from a reaction mixture between contacting steps.

[0189] Embodiment 11. The method of any of Embodiments 1-10 or any other Embodiment, wherein the removing previous dNTP pools comprises immobilizing the DNA strand and the nascent DNA strand and washing the reaction mixture such that a previous analog dNTP is removed from the reaction mixture.

[0190] Embodiment 12. The method of any of Embodiments 1-11 or any other Embodiment, wherein the immobilizing step comprises attachment of a DNA molecule comprising the DNA strand, the nascent DNA strand, or both, to a substrate.

[0191] Embodiment 13. The method of any of Embodiments 1-12 or any other Embodiment, wherein a majority of dNTPs of the dNTP pool are not analog dNTPs.

[0192] Embodiment 14. The method of any of Embodiments 1-13 or any other Embodiment, wherein no more than 25% of dNTPs of the dNTP pool are analog dNTPs.

[0193] Embodiment 15. The method of any of Embodiments 1-12 or any other Embodiment, wherein a majority of dNTPs of the dNTP pool are analog dNTPs.

[0194] Embodiment 16. The method of any of Embodiments 1-12 and 15 or any other Embodiment, wherein 100% of dNTPs of the dNTP pool are analog dNTPs.

[0195] Embodiment 17. The method of any of Embodiments 1-16 or any other Embodiment, further comprising ligating a sequencing adaptor to the DNA strand and the nascent DNA strand to configure the DNA library for nanopore sequencing.

[0196] Embodiment 18. The method of any of Embodiments 1-17 or any other Embodiment, wherein the sequence adaptor comprises a barcode that corresponds with an identity of the analog dNTP.

[0197] Embodiment 19. The method of any of Embodiments 1-18 or any other Embodiment, further comprising reverse transcribing an RNA strand to produce the DNA strand comprising the target polynucleotide sequence, wherein the DNA library represents at least part of an RNA fraction of a sample.

[0198] Embodiment 20. A DNA library prepared according to the method of any of Embodiments 1-19.

[0199] Embodiment 21. A method for high-accuracy analysis of nanopore sequencing data obtained from nanopore sequencing of the DNA library of Embodiment 1or any other Embodiment, the method comprising: aligning the nanopore sequencing data such that sequence data of the DNA strand is aligned with sequence data of the nascent DNA strand; computing an electrical property of the nanopore sequencing data that corresponds to an analog nucleobase of the nascent DNA strand for an analog nucleobase measurement; comparing the analog nucleobase measurement with a reference analog nucleobase measurement associated with a known analog nucleobase identity and assigning a nucleobase identity of the analog nucleobase consistent with the known analog nucleobase identity; and assigning, based on the nucleobase identity of the analog nucleobase, a nucleobase identity to an unknown nucleobase of the DNA strand that positionally corresponds to the analog nucleobase of the nascent DNA strand.

[0200] Embodiment 22. The method of Embodiment 21 or any other Embodiment, wherein a plurality of nucleobase identities of analog nucleobases of the nascent DNA strand correspond to a plurality of nucleobase identities of unknown nucleobases of the DNA strand.

[0201] Embodiment 23. The method of any of Embodiments 21-22 or any other Embodiment, wherein the plurality of nucleobase identities of analog nucleobases of the nascent DNA strand comprises one distinct nucleobase identity that corresponds to one contacting step used for DNA library preparation.

[0202] Embodiment 24. The method of any of Embodiments 21-23 or any other Embodiment, wherein the plurality of nucleobase identities of analog nucleobases of the nascent DNA strand comprises two distinct nucleobase identities that correspond to one or more contacting steps used for DNA library preparation.

[0203] Embodiment 25. The method of Embodiment 24 or any other Embodiment, wherein the plurality of nucleobase identities of analog nucleobases of the nascent DNA strand comprises three distinct nucleobase identities that correspond to one or more contacting steps used for DNA library preparation.

[0204] Embodiment 26. The method of Embodiment 25 or any other Embodiment, wherein the plurality of nucleobase identities of analog nucleobases of the nascent DNA strand comprises four distinct nucleobase identities that correspond to one or more contacting steps used for DNA library preparation.

[0205] Embodiment 27. The method of any of Embodiments 21-26 or any other Embodiment, wherein the method is performed at least in part by a programmableprocessor, a processor circuitry, a computational device, a computational system, a computational network, or any combination thereof.

[0206] While illustrative embodiments have been illustrated and described, it will be appreciated that various changes can be made therein without departing from the spirit and scope of the disclosure.

Claims

CLAIMSThe embodiments of the invention in which an exclusive property or privilege is claimed are defined as follows:

1. A method for preparing a DNA library for high-accuracy nanopore sequencing, the method comprising: contacting a DNA strand comprising a target polynucleotide sequence with a DNA polymerase and a dNTP pool comprising an analog dNTP that comprises a modification relative to a corresponding reference dNTP that is absent from the dNTP pool; wherein the DNA polymerase incorporates the analog dNTP into a nascent DNA strand as an analog nucleotide in place of incorporation of the corresponding reference dNTP as a corresponding reference nucleotide.

2. The method of claim 1 , wherein an electrical resistance of the corresponding reference nucleotide differs from an electrical resistance of the analog nucleotide.

3. The method of claim 2, wherein the electrical resistance of the corresponding reference nucleotide is greater than the electrical resistance of the analog nucleotide.

4. The method of claim 2, wherein the electrical resistance of the corresponding reference nucleotide is less than the electrical resistance of the analog nucleotide.

5. The method of claim 1, wherein the modification of the analog nucleotide increases an electrical charge of the analog nucleotide relative to an electrical charge of the corresponding reference nucleotide.

6. The method of claim 1, wherein the modification of the analog nucleotide decreases an electrical charge of the analog nucleotide relative to an electrical charge of the corresponding reference nucleotide.

7. The method of claim 1, wherein the method is isothermal.

8. The method of claim 1, wherein the contacting the DNA strand is performed a plurality of times with a plurality of dNTP pools, wherein dNTP pools of the plurality of dNTP pools comprise pool-specific analog dNTPs that comprise pool-specificmodifications relative to corresponding reference dNTPs that are absent from the dNTP pools.

9. The method of claim 1 , wherein the contacting the DNA strand is performed three times.

10. The method of claim 1, further comprising removing previous dNTP pools from a reaction mixture between contacting steps.

11. The method of claim 1, wherein the removing previous dNTP pools comprises immobilizing the DNA strand and the nascent DNA strand and washing the reaction mixture such that a previous analog dNTP is removed from the reaction mixture.

12. The method of claim 1, wherein the immobilizing step comprises attachment of a DNA molecule comprising the DNA strand, the nascent DNA strand, or both, to a substrate.

13. The method of claim 1, wherein a majority of dNTPs of the dNTP pool are not analog dNTPs.

14. The method of claim 1, wherein no more than 25% of dNTPs of the dNTP pool are analog dNTPs.

15. The method of claim 1, wherein a majority of dNTPs of the dNTP pool are analog dNTPs.

16. The method of claim 1, wherein 100% of dNTPs of the dNTP pool are analog dNTPs.

17. The method of claim 1, further comprising ligating a sequencing adaptor to the DNA strand and the nascent DNA strand to configure the DNA library for nanopore sequencing.

18. The method of claim 1, wherein the sequence adaptor comprises a barcode that corresponds with an identity of the analog dNTP.

19. The method of claim 1, further comprising reverse transcribing an RNA strand to produce the DNA strand comprising the target polynucleotide sequence, wherein the DNA library represents at least part of an RNA fraction of a sample.

20. A DNA library prepared according to the method of claim 1.

21. A method for high-accuracy analysis of nanopore sequencing data obtained from nanopore sequencing of the DNA library of claim 1, the method comprising: aligning the nanopore sequencing data such that sequence data of the DNA strand is aligned with sequence data of the nascent DNA strand; computing an electrical property of the nanopore sequencing data that corresponds to an analog nucleobase of the nascent DNA strand for an analog nucleobase measurement; comparing the analog nucleobase measurement with a reference analog nucleobase measurement associated with a known analog nucleobase identity and assigning a nucleobase identity of the analog nucleobase consistent with the known analog nucleobase identity; and assigning, based on the nucleobase identity of the analog nucleobase, a nucleobase identity to an unknown nucleobase of the DNA strand that positionally corresponds to the analog nucleobase of the nascent DNA strand.

22. The method of claim 21, wherein a plurality of nucleobase identities of analog nucleobases of the nascent DNA strand correspond to a plurality of nucleobase identities of unknown nucleobases of the DNA strand.

23. The method of claim 21, wherein the plurality of nucleobase identities of analog nucleobases of the nascent DNA strand comprises one distinct nucleobase identity that corresponds to one contacting step used for DNA library preparation.

24. The method of claim 21, wherein the plurality of nucleobase identities of analog nucleobases of the nascent DNA strand comprises two distinct nucleobase identities that correspond to one or more contacting steps used for DNA library preparation.

25. The method of claim 24, wherein the plurality of nucleobase identities of analog nucleobases of the nascent DNA strand comprises three distinct nucleobase identities that correspond to one or more contacting steps used for DNA library preparation.

26. The method of claim 25, wherein the plurality of nucleobase identities of analog nucleobases of the nascent DNA strand comprises four distinct nucleobase identities that correspond to one or more contacting steps used for DNA library preparation.

27. The method of claim 21, wherein the method is performed at least in part by a programmable processor, a processor circuitry, a computational device, a computational system, a computational network, or any combination thereof.

Citation Information

Patent Citations

  • Method of preparation of nanopore and uses thereof

    US20190309008A1

  • Compositions and methods for improving nanopore sequencing

    US20210172013A1