Long contiguous DNA sequence reads from longer amplified inserts by maintaining phase and improving detection
By preparing amplicons with longer inserts and using crosslinking oligonucleotides and enhanced detection, the technology achieves longer and more accurate sequence reads, addressing errors in existing sequencing methods.
Patent Information
- Application Number
- PCT/CN2025/078176
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-02-20
- Filing Date
- 2025-02-20
- Publication Date
- 2025-08-28
AI Technical Summary
Existing DNA sequencing technologies face limitations in obtaining accurate and long sequence reads due to the length of each read, which can introduce errors, particularly in areas with high sequence complexity or repetition, leading to decreased accuracy in variant detection.
The technology involves preparing amplicons with longer fragment inserts, arraying them closer together using crosslinking oligonucleotides and high magnesium concentration, and improving primer extension product phase maintenance through enhanced detection protocols, such as two color four image (2c4i) detection, to achieve reads of 600 to 1000 base pairs or more.
This approach enhances the length and accuracy of sequence reads, overcoming challenges with longer inserts by maintaining primer extension phase and increasing signal intensity, resulting in improved sequencing accuracy and cost-effectiveness.
Smart Images

Figure PCTCN2025078176-FTAPPB-I100001 
Figure PCTCN2025078176-FTAPPB-I100002 
Figure PCTCN2025078176-FTAPPB-I100003
Abstract
Description
Long contiguous DNA sequence reads from longer amplified inserts by maintaining phase and improving detectionRELATED APPLICATION
[0001] This disclosure claims the priority benefit of international patent application PCT / CN2024 / 077740, filed on February 20, 2024. The priority application is hereby incorporated herein by reference in its entirety for all purposes.TECHNICAL FIELD
[0002] This disclosure relates generally to the fields of oligonucleotide chemistry and DNA sequencing. It provides several technological enhancements that extend sequence reads and improve accuracy of the data obtained.BACKGROUND
[0003] Massively parallel sequencing (MPS) , also known as next-generation sequencing (NGS) , has become an important tool in disease research, personalized medicine, genetic testing, and disease tracking worldwide, to improve the standard of clinical diagnosis and care. MPS is used in whole genome sequencing (WGS) to study an organism’s complete DNA sequence, providing a detailed map of its entire genetic material. By sequencing both coding and non-coding regions of the genome, WGS provides an extensive view of genetic information, including genetic variations, mutations, and structural changes, capturing the entire genome in a single process. MPS is also used for whole transcriptome analysis (WTA) to study a cell or an organism’s entire mRNA production, providing a snapshot at a particular moment of which and to what extent individual genes are being expressed.
[0004] Sequencing a large target DNA entails making a fragment library, obtaining sequence reads from each of the fragments, then assembling the sequence reads to obtain and characterize the nucleotide sequence of the target. Potential limitations of MPS technology arise from the length of each sequence read -in previous technology, typically ranging from 100 to 300 base pairs. Longer reads can introduce errors in sequencing of samples that don’ t have a close reference sequence. Errors in base calls and read assembly adversely affect the accuracy of variant detection, particularly in areas with high sequence complexity or repetition.
[0005] The technology put forth in this disclosure improves the length and accuracy of sequence reads that can be obtained, providing important benefits for the sequencing of genome and expression libraries.SUMMARY OF THE INVENTION
[0006] This disclosure provides a technology for improving the length and accuracy of sequence reads obtained from a DNA fragment library. Amplicons (such as DNA nanoballs or PCR clusters) are prepared that have longer fragment inserts. This makes the amplicons larger, resulting in lower signal intensity per amplicon, and increasing the likelihood that primer extension products will go out of phase. To compensate, the amplicons are arrayed closer together, and compacted using crosslinking oligonucleotides and a high magnesium concentration. Primer extension products are kept in phase by improving incorporation of bases, ensuring that reactions go to completion, and purging unused reagents. Data collection is improved with brighter detection reagents using a two color four image (2c4i) detection protocol. When these features are combined, accurate reads of 600 to 1000 base pairs or more are routinely obtained.
[0007] To the extent applicable, these features may be incorporated for iterative sequencing or characterization of amplicons of various kinds. Amplicons of a fragment library can be produced by processes such as emulsion PCR (ePCR) , rolling circle amplification (RCA) , microwell array amplification, nanowire or nanoparticle-based amplification, solid-phase isothermal amplification, laser-induced cluster formation, surface tension PCR, microfluidics-based amplification.
[0008] By way of illustration and without implying any limitation, these features can be applied to amplicons referred to as DNA nanoballs (DNBs) or concatemers. This may be done by preparing circular DNA molecules comprising DNA inserts to be sequenced, linked end to end with an adaptor, and forming DNBs or concatemers from each of the circular DNA molecules (for example, by rolling circle replication, RCR) . The concatemers can be arrayed on a surface that has regions or “spots” that bind DNA, interspersed by regions that are inert from DNA binding. By using a spot diameter that is proportional (at least not larger) than the diameter of the concatemers, saturating the binding regions with concatemers will result in most of the binding sites presenting just one concatemer.
[0009] To the extent applicable, these features may be incorporated to the extent applicable for obtaining sequence reads by a variety of methods. Such methods include but are not limited to sequencing by synthesis (SBS) , sequencing by binding (SBB) , nanopore sequencing, sequencing by ligation (SBL) , and hybridization-based sequencing.
[0010] To practice SBS, oligonucleotide primers can be annealed to a sequencing primer binding site in adaptors that are ligated or positioned adjacent (nearby) the insert DNA in the amplicon that is being sequenced. The primers are extended one base at a time, using a DNA polymerase and nucleotide triphosphates or analogs thereof to form primer extension products. For each iteration, the user determines which base is added to the primer extension products on each concatemer. The nucleotide triphosphates used to elongate or extend the primers can be reversibly blocked terminators (RTs) that are chemical analogs of the four bases, which can be detected using a labeled detecting means specific for each of the four RTs. Suitable detecting means include monoclonal antibodies, fragments thereof that contain at least one antigen combining site, aptamers, nanobodies, lectins, and molecularly imprinted polymers (MIPs) . After detection, the RTs are converted to regular nucleotides for the next sequencing cycle.
[0011] One of the improvements provided in this disclosure is to use a longer DNA insert in each template of the amplicon (Feature 1) : rather than 100 or 300 bases, the template could have a median length of at least 400, 600, 800, 1000, or 1200 bases, or from 400 or 600 to 1000, 1200, or 1500. By keeping the number of replicates of the primer about the same, the mass of each amplicon will be increased. The number of replicates or copies may be at least 25, 35, 50, 80, 100, 150 or more, or from 30 to 50, 30 to 75 or 100, or 50 to 100 or 150. Starting with the circular DNA, the replication time under standard conditions may be at least 50, 75, 100, 150, 200, or 300 minutes.
[0012] Before development of the technology described here, most prior art methods do not operate reliably or accurately with longer inserts and longer sequence reads. Using amplicons with larger inserts will generally result in a lower adaptor to insert size ratio, which in turn will decrease the signal being detected for each extension product (as a proportion of total DNA on each binding site) . Long primer extensions can introduce more sequencing errors that accumulate with each sequencing cycle.
[0013] The adverse effects of using longer inserts can be compensated by implementing other features provided in this disclosure. Feature 2 comprises arraying the amplicons on the surface closer together and / or in a more compact form. Feature 3 comprises minimizing the occurrence of primer extension products getting out of phase with each other. Feature 4 comprises increasing signal intensity and / or sensitivity of detection of the primer extension products. Thus, Feature 1 can be implemented with any one of Features 2, 3, and 4. Feature 1 can be implemented with both Features 2 and 3, Features 2 and 4, or Features 3 and 4; or all four features can be implemented together.
[0014] Feature 2 can be implemented by arraying the amplicons on the surface closer together. If the binding sites on the surface are in a grid pattern, the pitch (the distance between the center of spots in successive rows) may be, for example, equal to or no more than 250, 300, 400, 500, 600, or 800 nm, or 1.0, 1.2, or 1.5 microns. The diameter of each amplicon binding site or spot may be equal to or no more than 150, 200, 250, 300, 400, 500, or 600 nm.
[0015] Alternatively or in addition, Feature 2 can be implemented by compacting and / or cross-linking each amplicon when they are arrayed on a surface. For example, amplicons may be compacted before or after plating on the surface using one or more crosslinking oligonucleotides that bridge between different parts of the amplicons: such as a crosslinker binding site on adapters in different copies of the template replicated in each amplicon. The crosslinker oligonucleotides may have two, three, or more than three binding sites for adaptors and / or for other crosslinker oligonucleotides, as illustrated in a following section of this disclosure (such as SEQ ID NOS. 1 to 14) . The adaptors may each comprise one, two, or a plurality of binding sites for crosslinking oligonucleotides. The oligonucleotide binding sites may be located upstream (towards the 5’ end) from a binding site for sequencing primers and / or a bar code, which precede the insert to be sequenced, so that the oligonucleotide does not complicate sequence determination.
[0016] Alternatively or in addition, Feature 2 can be implemented by compacting each amplicon using adaptors within adaptors in the amplicon that have palindromic or self-complementary sequences, whereby adapters in different copies of the template replicated in each amplicon may hybridize with each other (such as SEQ ID NOS. 15 to 17) .
[0017] Alternatively or in addition, Feature 2 can be implemented by including in the reaction mixture (either during amplicon formation, during plating on the surface, or both) a solute that has the effect of compacting and / or stabilizing the amplicon in a compact form. Such solutes may include one or more of the following: alcohol (such as ethanol) , polyethylene glycol (PEG) , polypropylene glycol (PPG) , dimethyl sulfoxide (DMSO) , a DNA minor groove binder moiety such as CDPI3, or a divalent cation such as calcium or magnesium. For example, the amplicons may be compacted during preparation or plating using magnesium cation in the reaction mixture at a concentration between 20 and 50 mM, 15 and 50 mM, 25 and 60 mM, or at a concentration that is at least 20, 25, 30, 35, or 40 mM. In some circumstances, the compacting effect of crosslinking oligomers, self-complementary adaptors, and / or solutes such as magnesium can be promoted after amplicon formation or plating by heating for a short time (for example, 20-60 sec) above room temperature but below melting temperature (for example, 40-50℃) .
[0018] Feature 3 constitutes taking action to minimize the occurrence of primer extension products getting out of phase with each other. Since multiple replicates of the insert fragment in each amplicon are sequenced together, if one of the extension products lags out of sequence (by skipping a sequencing cycle) or leads out of sequence (by incorporating more than one base in a sequencing cycle) , then the extension product that is out of phase will be in conflict with and dilute the signal of the other extension products in the amplicons. The idea here is ensuring that the step of adding each nucleotide or analog to primer extension products in each iteration of step (d) is at least 98%, 99%, 99.5%, or 99.8%complete. Alternatively or in addition, the extension products can be rephased every 50, 100, or 200 cycles or whenever needed, as described in EP 4121554 B1.
[0019] Feature 3 can be implemented by using more than one step of reversible terminator (RT) incorporation for each of the four bases for each iteration of step (d) . Thus, the amplicons are reacted with the DNA polymerase and one or more RTs, washed, and then reacted with the same RTs a second, third, or fourth time before detection.
[0020] Alternatively or in addition, Feature 3 can be implemented by more thorough washing between sequencing cycles to purge reagents from the previous cycle. Reagents involved in sequence determination (particularly RTs and / or RT detecting means such as antibodies) are removed after primer extension in each sequencing cycle. The completeness of washing can be enhanced by including in the wash solution unlabeled RTs and / or unlabeled RT detector means in the wash solution. In this context, an “unlabeled RT” is a terminator analog that does not bear a label and / or is not specifically recognized by a secondary reagent (such as an antibody) that bears a label, but it may be able to compete for binding to the amplicons or the DNA polymerase. If specific antibody is used as the RT detector means, then its unlabeled form may be an unlabeled antibody of different specificity, or a non-reactive immunoglobulin or other protein.
[0021] Alternatively or in addition, Feature 3 can be implemented by using a DNA polymerase (such as a Taq polymerase) for primer extension that is adapted (typically by mutation or directed evolution) to decrease preference for 3’-OH nucleotides. That is, it incorporates 3’-OH nucleotides at least 1.5, 2, 5, or 10-fold less well than nucleotides without a 3’-OH. Exemplary is Taq polymerase incorporating one or more amino acid substitutions at or around position 617 of its amino acid sequence.
[0022] Feature 4 can be implemented by increasing the intensity of the signal or label used in each sequencing cycle. This may be done, for example, by using a fluorophore that is at last 1.2 1.5, or 2 or 3-fold as bright as a standard fluorophore having a similar emission frequency, such as fluoresceine or Cy3, when coupled to the RT detecting means used for sequencing. Exemplary are iF647 and zF647, which are less subject to quenching when coupled to a protein such as an antibody. Alternatively or in addition, when an antibody (or similar RT detecting means) is used, the antibody can be labeled with a plurality of the same or different fluorescent labels: at least 3, 5, or 8 per molecule.
[0023] Alternatively or in addition, Feature 4 can be implemented by increasing the number of copies of the DNA inserts in the amplicon after arraying on the surface by in situ amplification: for example, as described in U.S. Patent No. 8,785,127, or by multiple displacement amplification (MDA) as set forth below. Alternatively or in addition, Feature 4 can be implemented by increasing the sensitivity of detection: for example, using a two label four image (2c4i) detection protocol.
[0024] By using amplicons with longer inserts (Feature 1) and optionally one or more of the other features described in the sections that follow, sequence reads can be obtained up to about the length of the insert: for example, a median length of at least 400, 600, 800, 1000, 1200, or 1500 bases.
[0025] The improvements set forth in this disclosure can constitute a particular method for nucleotide sequencing: for example, forming a plurality of concatemers, each of which comprises one of fragments to be sequenced that are replicated to form a linear DNA; and then obtaining sequence reads in the fragment in each concatemer by a process that comprises primer extension; wherein the median size of fragments in the concatemers is at least a certain number of bases indicated above such as 600, and / or the median size of sequence reads obtained therefrom is at least a certain number of bases indicated above such as 400.
[0026] The concatemers may be arrayed on a surface separated from each other center-to-center by no more than a certain distance or pitch, such as 300 or 500 nm. The concatemers may be compacted by a plurality of oligonucleotides that hybridize between and crosslink different adaptors between adjacent replicates of the fragment in the linear DNA. The concatemers may be compacted by adaptors within said linear DNA that cross-hybridize with each other. The concatemers may be compacted during preparation and / or while plating onto a surface using magnesium cation at a concentration indicated above, such as between 20 and 50 mM.
[0027] The sequence reads may be obtained by extending primers hybridized to adaptors within the linear DNA one base at a time (1) using more than one step of reversible terminator (RT) incorporation for each base in each cycle of sequencing; and / or (2) by removing labeled reagents in each sequencing cycle by including unlabeled RTs and / or non-specific antibody in a wash solution. To increase intensity, sequence reads may be obtained using a moiety that binds to nucleotide analogs that has a plurality of fluorophores, and / or using two label four image (2c4i) detection.
[0028] The improvements set forth in this disclosure can also constitute a composition of matter: for example, one or a plurality of linear DNA concatemers, each comprising replications of a DNA fragment to be sequenced, wherein the median size of such fragments in the concatemers is at least 600, 800, or 1000 bases configured to obtain sequence reads of the fragment of at least 400 bases or more. The concatemers may be arrayed on a surface separated from each other by a specified distance, such as no more than 300 or 500 nm. The concatemers may be compacted by a plurality of oligonucleotides that hybridize to and crosslink between different adaptors ligated between adjacent replicates of the template in the concatemer. The concatemers may also comprise adaptors having palindromic or self-complementary sequences that cross-hybridize with each other. The concatemers may be suspended or arrayed on a surface in an aqueous solution that contains magnesium cation at a particular concentration, such as between 20 and 50 mM.
[0029] Other features, aspects, and embodiments of the invention are presented in the sections that follow.BRIEF DESCRIPTION OF THE DRAWINGS
[0030] FIG. 1A provides data from a working illustration of the technology in this disclosure, as applied to sequencing-by-synthesis of DNA nanoballs. The graphs are a plot of RHO (signal intensity) for each of the four bases for a sequencing run of 600 bases. In each panel, the curves from top to bottom show the results for DNBs amplified for 200, 150, or 100 min. A longer amplification time increases the size of the DNB, which in turn increases signal strength.
[0031] FIG. 1B depicts some of the challenges in DNA sequencing that are resolved using the long sequence read technology of this disclosure.
[0032] FIG. 2 depicts the underlying technology of how fragments from a DNA library are prepared for sequencing by forming an array of DNA nanoballs (DNBs) .
[0033] FIGS. 3A and 3B depict an implementation of concatemer based sequencing that uses sequencing by synthesis (FIG. 3A) . Reversable terminators that are analogs of each of the four bases are detected using specific antibodies bearing fluorescent labels (FIG. 3B) . FIG. 3C depicts the making of complementary strands for pair-end (PE) sequencing on the platform.
[0034] FIGS. 4A and 4B depict other methods for making amplicons from DNA sequencing libraries: specifically, emulsion PCR and bridge PCR.
[0035] FIGS. 5A and 5B show the size distribution of library inserts for making DNBs for PE sequencing. X axis indicates length (bp) of inserts. Y axis indicates the corresponding percentage (%) of inserts.
[0036] FIGS. 6A and 6B provide site occupancy analysis of a human 600 bp library and a 1000 bp library, respectively.
[0037] FIGS. 7A to 7D show that larger DNBs provide a stronger signal. The signal intensity (rho) of each of the four bases is graphed across 50 cycles of sequencing. The data for moderately enlarged DNBs (50 min amplification time) is similar to regular DNBs (25 min) , while even larger DNBs. (75 min and 100 min) show a corresponding increase.
[0038] FIGS. 8A to 8D compare the signal to noise (SNR) ratio observed from DNBs of different sizes. Each of the four bases is graphed across 50 cycles of sequencing. Large DNBs show higher initial SNR (>10 or 12) .
[0039] FIGS. 9A to 9D show the Q30 score (ameasure of sequencing accuracy) obtained across 50 cycles of sequencing from DNBs of different sizes. Large DNBs scored better than regular DNBs under these conditions.
[0040] FIGS. 10A and 10B depict the structure of cross-linker oligonucleotides that can be used for compacting large DNBs so that they may be plated closer together and generate a brighter signal. FIG. 10A provides a pair of “Z-linker” type stapler oligonucleotide, each comprising a stapler sequence and a primer sequence, with the stapler sequence located 5’ to the primer sequence. FIG. 10B shows “X-linker” type oligonucleotides, which include sequences that can hybridize to another linker oligonucleotide and to an adaptor in a concatemer. FIG. 10C is a graph of data demonstrating that both X-linkers and Z-linkers reduce error rate over the course of 200 sequencing cycles.
[0041] FIGS. 11A and 11B depict the structure of a two unit and a three unit crosslinking oligonucleotide hybridized to two sites within an adapter of a DNB. FIG. 11A are the sequences (SEQ ID NOS. 5 to 14) . FIG. 11B depict the portions that anneal to a complementary sequence in the adaptor molecule. FIG. 11C shows the structure of two alternative designs for adaptors within each concatemer that increase sites for hybridizing crosslinkers next to the insert being sequenced.
[0042] FIG. 11D provides data that illustrate the benefit of using crosslinkers for condensing DNBs. The graph shows sequencing quality scores (%Q30) for DNBs condensed with X-linkers. Even at sequencing cycle 600, over 79%of DNBs had Q scores greater than 30.
[0043] FIGS. 12A to 12F provide data that illustrate the benefit of condensing DNBs using high concentrations of Mg++ during preparation and plating. Box = interquartile range IQR; whiskers = 1.5x the IQR of the data above or below the box; dots = any outlier datapoints beyond the fence. Yield of binding sites (spots) with single DNBs increase substantially (FIG. 12A) , while spots lost due to a mixed signal (FIG. 12B) or split DNBs (FIG. 12D) decrease. There is also a modest increase in signal intensity (FIG. 12F) .
[0044] FIGS. 13A and 13B show the benefit of using fluorescent labels with higher intensity. The graphed data are from a staged sequencing run using iF647 and zF647 respectively. There is no reduction after 4 days at 4℃.
[0045] FIGS. 13C and 13D demonstrate a method of keeping primer extension products in phase by using second and third incorporations of the reversible terminators (RTs) with reagent exchange between each incorporation for each sequencing cycle. This decreases the occurrence of primer extension products falling behind phase.
[0046] FIGS. 14A and 14B demonstrate another method of keeping primer extension products in phase: in this case, using improved washing conditions. The two-dimensional plots show cross-talk between nucleotide bases observed in subsequent sequencing cycles. Including unlabeled reversible terminators (RTs) in the wash solution (FIG. 14B) improves sequencing accuracy compared with a wash solution not containing unlabeled RTs.
[0047] FIGS. 15A to 15C demonstrate a third method of keeping primer extension products in phase. DNA polymerase PX617 was used for primer extension, which is mutated so that it has a lower preference for 3’-nucleotides. In FIG. 15A, the average run-on is flattened, reducing sequence lag. Comparing FIGS. 15B and 15C shows that using PX617 results in fewer incorporation mismatches.
[0048] FIGS. 16A and 16B demonstrate the benefit of using two color four image (2c4i) detection. In FIG. 16A, DNBs prepared with a 25 min RCE reaction were sequenced using 2c4i for the first 150 cycles, and 2c2i for subsequent cycles. FIG. 16B shows similar results obtained from larger DNBs prepared by a 100-min RCR reaction. Switching away from 2c4i detection in both cases resulted in a decrease in sequence quality (Q30) .
[0049] FIGS 17A to 17F demonstrate the efficacy of technological advancements provided in this disclosure. The data show the effect of spot diameter on spot diameter and pitch on DNB spot yield, split rate, and RHO. Load conditions used for the data marked ■ and □ have higher copy number and therefore give higher signals than control (marked Δ) .
[0050] FIGS. 17D and 17F show the effect of DNB size and magnesium concentration on yield (the number of spots bearing one and only one DNB) and the percentage of DNBs split between neighboring spots. FIG. 17F compares the effect of spot pitch on split DNBs. Increasing the insert size of DNBs causes more splitting, which decreases yield. Adjusting magnesium concentration and spot pitch compensates and lowers the split rate.DETAILED DESCRIPTIONOverview of the technology of this disclosure
[0051] For DNA nanoball (DNB) sequencing, each fragment insert to be sequenced is ligated to an oligonucleotide adaptor to form a template, which is then replicated linearly to form a concatemer. The concatemer is then contacted with a sequencing primer and a DNA polymerase. The primer hybridizes to the adaptor at a site adjacent to the insert. The DNA polymerase extends the primer one base at a time depending on the sequence of the insert. Using labeled and blocked primers, each nucleotide added can be determined iteratively by determining the label on each the nucleotide, then unblocking it and removing the label so that the next nucleotide can be added and determined.
[0052] The total length of the sequence reads obtained by this process (the number of sequencing cycles) is limited by the length of the DNA fragment being sequenced in each concatemer, and the accuracy of sequencing: typically, about 200 base pairs (bp) . The technology provided in this disclosure improves concatemer sequencing by increasing the length of the fragment library. If about the same number of copies are present in each amplicon, the longer fragment makes the amplicons larger (Feature 1) .
[0053] However, increasing read length isn’ t as simple as just making the amplicons bigger. The longer fragments in each concatemer dilute the adaptor: fragment ratio, which in turn decreases the relative signal intensity. When plated in an array, larger DNBs split between binding sites, which reduces the number of DNBs that can be read. Accuracy also decreases, because a proportion of primer extension products go out of phase in each sequencing cycle, accumulating with longer reads.
[0054] The technology put forth in this disclosure can overcome the challenges of working with large amplicons. Aspects of the technology include the following: ● arraying the concatemers on a surface more densely and closer together (Feature 2) ; ● decreasing the frequency that each primer extension product gets out of phase (Feature 3) ; and / or ● increasing the signal obtained in each primer extension cycle (Feature 4) .
[0055] These advancements can be achieved by one or a combination of the following structural and procedural changes: ● using smaller DNA spots on the patterned array to improve signal to noise ratio (Feature 2A) ; ● compacting and stabilizing each amplicon to improve plating accuracy, decrease the frequency of splitting, and maintain the amplicon during sequencing using crosslinking oligonucleotides (Feature 2B) and / or using special solutes such as magnesium (Feature 2C) ; ● ensuring that the step of adding each nucleotide to the primer extension products goes to completion (Feature 3A) ; ● improving wash conditions for removing the label each sequencing cycle (Feature 3B) ; ● using a DNA polymerase that has lower affinity for 3’-OH nucleotides (Feature 3C) ; ● increasing signal intensity for each primer extension product (Feature 4A) ; and ● using two label four image (2c4i) detection of each nucleotide added (Feature 4B) .
[0056] Depending on the implementation, a selected combination of these technical features increases the read length from each fragment; improves the accuracy of sequencing determination; and makes the system more cost-effective. Better results at lower cost
[0057] The technology put forth in this disclosure compensates for the dilution effect of longer inserts. Accurate sequence reads of 600 to 1000 or more bp can be obtained.
[0058] FIG. 1A is a plot of rho (signal intensity) for each of the four bases for a sequencing run of 600 bp. In each panel, the curves from top to bottom show the results for DNBs that were amplified for 200, 150, or 100 min. A longer amplification time increases the size of the DNB, which in turn increases signal strength.
[0059] FIG. 1B schematically depicts difficult challenges in DNA sequencing that are resolved using the long sequence read technology of this disclosure: (1) Regions that are difficult to sequence in a human genome, such as repeated regions; (2) structural variation of chromosomes; (3) determining T-cell and B-cell repertoires (TCR / BCR) and bacterial diversity from complex microbiomes or environments (16S) ; (4) sequencing DNA from a whole community of a microorganism (Meta WGS) ; microorganism DNA; and (5) precisely measuring the levels of transcripts in a transcriptome and their isoforms (RNA-seq) .
[0060] A current working model of this technology being developed by MGI Tech Co., Ltd. is referred to as DNBSEQ-G800RSTM. A 2 x 150 sequencing kit equates to $5 USD per gigabase or less. A cartridge has two independent flow cells; effectively two instruments in one, with 8 individually auto-loaded lanes. G800RS supports up to SE600 as an all-in-one box solution. Auto-wash is included, and there is incorporated library compatibility: no extra cost for primers, no extra manual steps for library preparation. Definitions and abbreviations
[0061] As used in this disclosure, the term “iterative sequencing” refers to a process for sequencing a polynucleotide (usually single-stranded DNA) which comprises a series of steps that are reiterated through a plurality of cycles. A polynucleotide to be sequenced is contacted with a reagent that binds in a sequence dependent manner to one or a small number of bases, which is then detected o identify the nucleotide base or bases it is attached to, and subsequently removed whereby the next cycle of sequencing may be performed to identify the next nucleotide base (s) . Exemplary are sequencing by synthesis (SBS) , sequencing by binding (SBB) , sequencing by ligation (SBL) , and cPAL (combinatorial probe anchor ligation) .
[0062] A full iteration or cycle is a procedure in which nucleotides or analogs for each of the four bases adenine (A) , cytosine (C) , guanine (G) , and thymine (T) is presented to each of the elongating primer extension product in each amplicon in the presence of a DNA polymerase, such that any one of the four will hybridize to the template and be ligated to the primer extension product as the next base in the sequence. The bases may be presented with the product simultaneously, separately, or in pairs (as in 2c4i) as part of the same iteration. If blocked, they may each be offered to the primer extension product more than once in any order as part of the same iteration. The primer extension product may be washed in between each full iteration and optionally in between any or all separate presentations of the nucleotides or analogs.
[0063] An “amplicon” is the product of a polynucleotide amplification reaction, namely, a population of polynucleotides that are replicated from one or more starting sequences. Each amplicon is a plurality of substantially identical DNA molecules grouped together in physical space: for example, contiguously in a linear DNA (such as a DNA nanoball) or constrained in proximity (within 50, 100, 200, or 300 nm center-to-center) , for example, by being attached near to each other on a flat surface or a bead. Amplicons may be produced by a variety of amplification reactions, including but not limited to polymerase chain reactions (PCRs) , bridge PCR, linear polymerase reactions, nucleic acid sequence-based amplification, rolling circle amplification (U.S. Patent Nos. 7,115,400, 4,683,195; 5,210,015; 6,174,670; 5,399,491; 6,287,824 and 5,854,033; and publication US 2006 / 0024711 A1) .
[0064] The four “standard” nucleotides in a DNA are adenine (A) , cytosine (C) , guanine (G) , and thymine (T) . A “nucleotide analog” is a chemical analog of one of the standard nucleotides with one or more structural differences in its covalent structure that behaves chemically in a similar manner to the corresponding standard analog in certain contexts, such as primer extension. This includes standard nucleotides that are conjugated covalently with a label, and blocked forms of standard nucleotides.
[0065] DNA molecules or strands are “substantially identical” to each other if they comprise a sequence of at least 50%of their length that is at least 90%or 95%identical with each other, as scored with the BLAST algorithm (Altschul, S. F. et al. (1990) J. Mol. Biol., 215 (3) , 403-410) .
[0066] Amplicons are “adjacent” to each other if they are within 500, 300, or 200 nm of each other on a surface or in three-dimensional space. Molecules are “adjacent” to each other if they are within 100, 50, 20, or 10 nm of each other in three-dimensional space. Features of a linear DNA (such as an insert and an adaptor) are adjacent if they are within 50 or 20 bases of each other, irrespective of bases between them. A feature of a DNA strand that is “upstream” from another feature is positioned more towards the direction that is opposite to the direction of primer extension.
[0067] A “concatemer” is a DNA macromolecule in which multiple replicates of a DNA template are present in the same DNA strand adjacent or nearby each other in the same orientation -such as may be achieved by rolling circle amplification (RCA) of the template presented in a single or double stranded DNA. The template comprises a portion of a target nucleic acid or DNA that is being sequenced or otherwise characterized (an “insert” ) , plus one or more artificial sequences or adaptors.
[0068] A “DNA nanoball” or DNB is a DNA macromolecule such as a concatemer that adopts a globular structure in a compatible buffer. “Concatemer sequencing” is the use of concatemers to determine the nucleotide sequence of a target DNA. Sequence reads are obtained from copies of the insert that is replicated in each concatemer, and assembled with sequence reads from other concatemer inserts to obtain at least part of the sequence of the target nucleic acid.
[0069] An “insert” is a DNA fragment of at least 100, 200, 400, or 800 bases in length (usually of unknown sequence) that derived from a fragment library to be sequenced, and is ultimately incorporated and replicated in an amplicon. Each insert may be an entire fragment taken directly from the library, or it may be a fragment thereof obtained, for example, by digestion with a restriction endonuclease or controlled primer extension (CPE) .
[0070] A “template” is generally a DNA strand that is replicated by a DNA polymerase to form a complementary strand, for a purpose such as iterative sequencing or replication of the entire strand. In the context of DNA sequencing according to this disclosure, the template is a DNA molecule that comprises an insert ligated at one or both ends by at least one adaptor, and may be replicated to form an amplicon.
[0071] In reference to DNA fragments, extension products and other polynucleotides, the terms “length” and “size” refer synonymously to the number of bases in the polynucleotide.
[0072] A “sized library” is a library of DNA fragments intended for sequencing that have a predetermined median length and length distribution. A sized library can be obtained, for example, by separating DNA fragments of different length, or by controlled primer extension of fragments according to this disclosure. The median length is “predetermined” in the sense that conditions of a reaction (such as primer extension) are empirically determined and / or chosen to generate polynucleotide products that have an intended median length. Lengths of fragments in a preparation are limits wherein less than 10%, 5%, or 2% (in terms of DNA mass) are below the stated “minimum” size or length, or above the stated “maximum” size or length.
[0073] “Adaptors” are oligonucleotides having an artificial (human designed) nucleotide sequence (single or double stranded, as appropriate) that are ligated on or into another polynucleotide, such as a DNA fragment intended for sequencing. Each adaptor potentially constitutes a tool kit for manipulation of an adjacent section of DNA on either or both sides. They may be referred to with an adjective (areplication or sequencing primer) that is non-limiting and used for purposes of labeling and discrimination from other adaptors in other locations. “Binding sites” on adaptors and other components are hybridization sites for other oligonucleotides, such as replication primers, sequencing primers, and crosslinking oligonucleotides. Adaptors may also contain one or more oligonucleotide barcodes, such as a unique molecule identifier (UMI) , which identify a larger DNA molecule or fragment from which the insert sequence was obtained, and / or other bar codes that indicate, for example, a source cell or biological material. Types of oligonucleotide barcodes and their use are described further in US 2024 / 0240174 A1.
[0074] “Primers” are oligonucleotides that are capable, upon forming a duplex with a polynucleotide template, of acting as a point of initiation of nucleic acid synthesis, whereby the primer is extended along a DNA template so that an extension product is formed, usually annealed to the template as the second strand of a double stranded DNA. A primer may be extended for any suitable purpose, such as PCR amplification, rolling circle amplification, and DNA sequencing by primer extension.
[0075] The term “controlled primer extension” (CPE) refers to a type of DNA strand replication in which an oligonucleotide primer is hybridized to a primer binding site on a DNA being replicated, and extended by DNA polymerase to make the complementary strand. The reaction can be controlled by reaction conditions, such as reagent concentration, buffer composition, and / or reaction time so that bases are added to make an extended primer having a predetermined median length and distribution. By adjusting one or more of these parameters, the user may correspondingly adjust the length of the DNA strand that is produced.
[0076] A “label” used with the technology of this disclosure is an atom or molecule attached to or contained within a chemical structure (such as a nucleotide) that provides a detectable signal differentiating the structure from structures bearing other labels or no labels. Exemplary are labels that emit a fluorescent signal at an emission frequency (wavelength between 200 and 700 nm) upon excitation with electromagnetic radiation at an excitation frequency. Other types of labels are radioisotopes, chromophores, mass labels, spin labels, and chemiluminescent labels, which, depending on context, can be used as an equivalent of fluorescent labels mutatis mutandis.
[0077] An “antibody” is an immunoglobulin, portion thereof, or fusion product thereof comprising at least one antigen recognition site that binds specifically to and distinguishes a particular antigen or epitope, such as a nucleotide analog, reversible terminator, or hapten. An “antigen binding moiety” may be an antibody, or a macromolecule with a different structure exemplified below that binds specifically to and distinguishes a particular antigen or epitope, such as a nucleotide analog, reversible terminator, or hapten. “target nucleic acid” (or polynucleotide) or “nucleic acid of interest” refers to any nucleic acid (or polynucleotide) suitable for processing and sequencing by the methods described herein. In some approaches, the target nucleic acid is a genomic fragment, generated by fragmenting genomic DNA extracted from a sample. It is noted that while genomic fragments are used for illustration of the methods and compositions disclosed herein, sequencing libraries can also be prepared using these methods and compositions to sequence any target nucleic acid or fragments thereof, including those that contain modifications of the nucleotides, e.g., nucleotide analogs.
[0078] A “target nucleic acid” or polynucleotide is an nucleic acid or library thereof from a source that is intended for sequencing according to this technology. Typically, its sequence will not previously be known in all details before sequence reads are obtained. Target nucleic acids may be genomic DNA or it could be an expression library (mRNA or cDNA) from one or more individuals or a selected or combined population of humans, other mammals, other vertebrates, other animals, plants, bacteria, and viruses, obtained as a sample of the target organism, or from one or more environmental samples.
[0079] Other terms are defined as they arise in this disclosure. Terms not explicitly defined have their ordinary meaning, adapted to the context in which they are used. Underlying technology: amplicon based nucleic acid sequencing
[0080] Determining the nucleotide sequence of inserts in an array of amplicons has emerged as a preferred method of massive parallel sequencing (MPS) . This section outlines aspects of sequencing technology that is already in commercial practice. Preparation of DNA nanoballs (DNBs)
[0081] FIG. 2 depicts how an array of DNA nanoballs is prepared. The user obtains fragments of a target DNA molecule for sequencing -genomic DNA from cells, reverse-transcribed RNA, or DNA from any other source. A circular DNA is produced with the ends of each fragment joined together via an adaptor sequence. The DNA being sequenced is referred to as an insert, which in combination with the adjacent adaptor constitutes a template.
[0082] Each circular DNA is amplified by rolling circle amplification (RCA) to produce the concatemer. The concatemer assumes a substantially spherical shape, known as a DNA nanoball (DNB) . Concatemers of a plurality of the templates are distributed on a surface, such as by random distribution on a patterned surface of DNA binding regions -thereby forming a patterned array. U.S. Patent No. 9,944,984. Sequencing can be done by hybridizing a primer to the adaptor, and extending the primer by synthesis or by ligation to form an extension product that is complementary to the insert being sequenced.
[0083] The adaptor is effectively a tool kit for manipulation and analysis of the insert during set-up and sequencing. There is typically a hybridization site for a sequencing primer at or near one or both ends of the adaptor, which anchors the primer extension product produced in the course of sequencing by synthesis or by ligation. There may be a hybridization site for an oligonucleotide that anchors the concatemer to a surface. There may be a binding or recognition site for a sequence specific endonuclease. There may be a barcode that is sequenced concurrently with the fragment insert. There may also be hybridization sites for structural oligonucleotides that bridge between adaptors in the same concatemer. Methods of iterative sequencing
[0084] Iterative or cyclical sequencing can be used to generate sequence reads because it builds a DNA strand complementary to the DNA template being sequined one nucleotide at a time. This can be done by hybridizing a primer to the template, and extending the template by cycles of single nucleotide addition and detection.
[0085] Sequencing by synthesis (SBS) begins with the binding of a short, single-stranded oligonucleotide primer to a complementary region of the DNA template strand: usually a primer binding site on an adaptor upstream from the insert being sequence. A DNA polymerase extends the primer from the 3’-hydroxyl group, incorporating a blocked nucleotide analog complementary to the template that bears a label are blocked from further extension. . After each incorporation, the added nucleotide is identified by way of a direct or secondary label. The nucleotide analog may be fluorescent itself or tethered with a fluorescent group, or it may be identified using a fluorescently labeled antibody or other antigen binding moiety that recognizes the terminator nucleotide just added, or a hapten (such as biotin) attached thereto. The label and the block are then removed to prepare for the next cycle of primer extension and detection.
[0086] Sequencing by binding (SBB) is also done by determining the sequence of a template, one nucleotide at a time. A fluorescent nucleotide analog is added immediately downstream to the primer that is complementary to the template being sequenced. Unlike SBS, the added nucleotide is not coupled to the 3’-OH of the primer. Instead, it is detected, removed, and then replaced with a regular nucleotide to prepare for the next cycle of sequencing. SBS using reversible terminators and SBB are both suitable for obtaining long reads, because they generate a primer extension product that (except for the terminal analog) contains regular nucleotides.
[0087] Depending on context and objectives, sequence reads may be obtained, for example, using SBS sequencing (such as SOLiDTM 5500, Life Technologies Corporation, Carlsbad, CA) , ion semiconductor sequencing (such as Ion PGMTM or Ion ProtonTM sequencers, Life Technologies Corporation, Carlsbad, CA) , zero-mode waveguides (such as PacBio RSTM sequencer, Pacific Biosciences, Menlo Park, CA) , nanopore sequencing (available from Oxford Nanopore Technologies Ltd., Oxford, United Kingdom) , pyrosequencing (available from 454 Life Sciences, Branford, CT) , or other sequencing technologies. Detection of reversible terminators using labeled antibody
[0088] FIG. 3A depicts an implementation of concatemer based sequencing that uses sequencing by synthesis and base detection using antibodies bearing fluorescent labels. U.S. Patent No. 10,851,410; US 2022 / 0162693 A1. Primer extension products are extended base-by-base with unlabeled reversibly terminated nucleotides. The identity of each base added is determined in each cycle of sequencing using antibodies (FIG. 3B) that are specific for each of the four 3’ blocked nucleotides (FIG. 3B) . Removal of the bound antibodies and 3’ blocking moiety on the sugar group of the nucleotide removes the label and regenerates natural nucleotides with no scar on the base.
[0089] The feature of reversion to a natural nucleotide allows further extension of the strand in a new cycle of sequencing without any interference from the prior cycle. Unlabeled RTs are easier and less costly to make, and they can be incorporated more efficiently. The antibodies can carry multiple labels, amplifying the sequencing signal compared with single dye molecule per base on standard labeled RTs. S. Drmanac et al., bioRxiv preprint, doi. org / 10.1101 / 2020.02.19.953307.
[0090] The technology can include controlled a multiple displacement amplification (MDA) process. U.S. Patent No. 10,227,647. After the first read is generated on DNBs, extended products (optionally using an additional primer) are further extended using natural unblocked nucleotides in a controlled and sufficiently synchronized way by a strand displacement polymerase such as Phi29. The process generates single-stranded (ss) DNA branches complementary to original DNBs and still bound to DNBs through regions that are not displaced.
[0091] FIG. 3C illustrates complementary strand making and pair-end sequencing on the MPS platform. (A) DNA nanoball (DNB) , as a concatemer, containing copies of adaptor sequence and inserted genomic DNA, is hybridized with a primer for the first-end sequencing. (B) After generating the first-end read, controlled, continued extension is performed by a strand displacing DNA polymerase to generate a plurality of complementary strands. (C) When the 3’ ends of the newly synthesized strands reach the 5’ ends of the downstream strands, the 5’ ends are displaced by the DNA polymerase generating ssDNA overhangs creating a “branched DNB” . (D) A second-end sequencing primer is hybridized to the adaptor copies in the newly created branches to generate a second-end read. (C) and (D) are idealized drawings; not all branches are necessarily generated, and they may be different in size. Preparing amplicons by emulsion PCR
[0092] FIG. 4A illustrates the formation of amplicons by emulsion PCR. DNA inserts ligated to adaptors are replicated within aqueous droplets suspended in an oil phase. Each droplet acts as a microreactor, containing a single DNA insert fragment, a bead, replication primers, and PCR reagents. The DNA fragments in each droplet become attached to the beads by way of primers on the bead surface that are complementary to part of the adaptor. During PCR, the DNA fragments bound to the bead in each droplet are amplified within the droplets, coating each bead with replicates of the insert. After PCR, the emulsion is broken; the beads are recovered, washed, and immobilized on a surface for sequencing. Preparing amplicons by bridge PCR
[0093] FIG. 4B illustrates how bridge PCR can be used as another method for generating clusters of an insert library on a surface for sequencing. Each insert is ligated with two different adaptor sequences, one at each end. The surface presents two types of anchor oligonucleotides having sequences that are complementary to each of the two adaptor sequences. Single stranded inserts hybridize at one end to a complementary anchor oligonucleotide on the surface. A DNA polymerase extends the anchored fragments, creating a complementary strand that remains tethered to the surface. The original template strand is removed, leaving the newly synthesized strand. The new strand bends over so that the adaptor at the other end hybridizes to a complementary anchor oligonucleotide on the surface forming a bridge. Subsequent rounds of amplification generate replicates of the insert DNA in clusters on the surface. Technical features that can be incorporated for sequencing long inserts
[0094] The sections that follow provide a range of different features that can be incorporated into the amplicon sequencing process. They combine to improve read length, sequence accuracy, and cost effectiveness, amongst other benefits. Feature 1: Amplicons having a larger insert size and larger mass
[0095] Increasing the length of a fragment insert to be sequenced in an amplicon decreases the adaptor: insert ratio. Everything else being the same, this reduces the number of primer extension products formed during sequencing, which decreases signal. Part of the technology in this disclosure is to compensate by making the amplicons larger. This restores the number of primer extension products made for each amplicon, which restores signal intensity.
[0096] In a working example, DNBs containing human genomic DNA were prepared with insert lengths of 600 bp and 1000 bp for read sequencing.
[0097] FIGS. 5A and 5B show the distribution of library insert size determined by PE sequencing. X axis indicates length (bp) of inserts and Y axis indicates corresponding percentage (%) of inserts. The insert size is distributed in roughly a bell-shaped curve, with an average size that increases with the time of the replication reaction.
[0098] FIGS. 6A and 6B show site occupancy analysis of the human 600 bp library and 1000 bp library, respectively. Upper curves in each graph: all DNBs; lower curves: seeds and singular DNBs (omits DNBs that are split between binding sites) . Feature 2A: Small DNA spots and high spot density
[0099] Another means to improve detectability of terminal nucleotdies and increase signal strength is to improve the signal to noise ratio (SNR) . This can be accomplished by making the amplicons more compact, and distributing them in an array closer together.
[0100] Potential sizes are <0.5 μm, <0.7 μm, or <1 μm spot pitch, with each DNA binding site or spot <300 nm or <500 nm in diameter. This achieves an initial SNR of >8, >10, or >12. Occupancy analysis is shown in TABLE 1. The data demonstrate much better loading, in terms of lower mixed sites and nanoball split when large DNB are distributed on a high density surface (500 nm pitch) .
[0101] FIGS. 7A to 7D compare signal intensity (rho) for each of the four bases across 50 cycles of sequencing. Large DNB (50 min) is similar to regular DNB (25 min) while 75 min and 100 min large DNB show proportional increase in signal intensity.
[0102] FIGS. 8A to 8D compare the signal to noise (SNR) ratio for each of the four bases across 50 cycles of sequencing. Large DNBs have better SNR (>10 or 12) .
[0103] FIGS. 9A and 9D show the Q30 quality score: a measure of sequencing accuracy, indicating the probability of incorrect base call of 1: 1000 (99.9%accuracy) . FIGS. 9B and 9C show DNBs out of phase because of run-on (at least +1 out of phase) and lag average (at least -1 out of phase) . FIG. 9D is a summary table of the data in the graphs. Large DNBs had better Q quality scores than regular DNBs under these conditions. Feature 2B: Compacting and stabilizing amplicons using crosslinking oligonucleotides
[0104] Making the amplicons smaller and more dense achieves several benefits. Smaller amplicons can be arrayed on a tighter grid and still be optically distinguishable from each other, improving the signal to noise ratio. They are more stable, preventing disassociation of the amplicon and the primer extension products each sequencing cycle. They will also generate a more intense signal, in terms of fluorescent emmsion per pixel in the imaging apparatus.
[0105] Increasing the number of copies of DNA templates yields larger DNA nanoballs. However, increasing the number of copies of DNA templates may degenerate the consolidated structure of DNA nanoballs. If the diameter of the DNA nanoballs surpasses the pore size of nanoarray, it may generate crosstalk between adjacent DNA nanoballs, further impeding sequencing quality.
[0106] Linker oligonucleotides are described generally in PCT publication WO 2021 / 083195. Use of oligonucleotide linkers reduces template copy loss by reducing enzymatic or chemical DNA cutting, and / or by keeping the cleaved copy attached to the rest of the amplicon by way of a primer or other oligonucleotide linker. This can help to reduce RHO loss, reduce worsening of signal-to-noise (SNR) , lessen a decrease in Q30 score (sequence quality) , and reduce sequencing error rate, producing a higher quality of long reads.
[0107] For the technology described here, the primers used for extension of the sequencing strand can be designed to incorporate longer tail sequences at the 5’ end that facilitate the linking process. In addition, other regions of the adapter can be designed to enhance linking between subunits. These secondary linking sites may again be linked via oligo tail sequences: however, the 3’ end of these oligos is blocked to prevent extension during sequencing.
[0108] In addition to providing stability during sequencing, oligonucleotide linkers can also be used during synthesis of the amplicon: for example, making DNBs from single stranded circles and strand displacing polymerase. By including the 3’ blocked linkers that hybridize to the secondary linker site of the adapter, condensation of the DNB will be promoted as the DNB is being made. When plated in an array, large DNBs can potentially split between two neighboring binding spots, which can be minimized by including the linker in the DNB making process.
[0109] FIG. 10A depicts an example of a pair of “Z-linker” stapler oligonucleotides, each comprising a stapler sequence and a primer sequence. The stapler sequence is located 5’ to the primer sequence. The stapler sequences of the pair of linker oligonucleotides are complementary, and when annealed will connect the linker oligonucleotides at the 5’ end to form a 2-arm Z-linker. A primer hybridizes to a DNB template (afirst strand) and each linker oligonucleotide is extended to form a second strand, which results in two second strands linked at 5’ (bottom right panel) . In this instance, the extension is carried out by a strand-displacement DNA polymerase, forming a branched structure in which each second strand is partially hybridized to the DNA template.
[0110] In reference to an “X-linker” of this disclosure, two linker oligonucleotides that are identical in sequence are linked through a palindromic sequence, each linker oligonucleotide possessing an extra sequence at the 5’ ends that can hybridize to another linker oligonucleotide having a non-palindromic stapler sequence. Each linker oligonucleotide also comprises a primer sequence that is complementary and can hybridize to an adaptor in the DNA template replicated in the amplicon. The X-linker structure allows for multiple 3’ extendable primer sequences, possibly four or more.
[0111] FIG. 10B shows an X-linker that comprises a pair of D-Bb-A linker oligonucleotides ( “Oligo Linker 1” ) and a pair of d-cC-A linker oligonucleotides ( “Oligo Linker 2” ) . The two D-Bb-A linker oligonucleotides hybridize to each other via the palindromic stapler sequences Bb, and the two d-cC-A linker oligonucleotides hybridize to each other via the palindromic stapler sequences cC. Each of the D-Bb-A linker oligonucleotide is also hybridized to one of the d-cC-A linker oligonucleotide via the complementary stapler sequences D and d. This results in “Structure 1” with four primer sequences that can hybridize to a DNA template. Each of the four primer sequences comprises an extendible 3’ end and can be extended to produce a second strand, which is a reverse complement of the DNA template.
[0112] The X-linker may also take a form shown in “Structure 2” . Two D-Bb-A linker oligonucleotides are annealed with four d-cC-A linker oligonucleotides, which results in a structure that has multiple primer sequences (A) , and excess single-stranded arms ( “D” or “d” ) . These excess strand arms allow for continued structure growth as a random network. For example, “D” is readily complementary to any “d” region, and is capable of annealing to “d” to expand the structure in any form.
[0113] In the context of the technology put forth in this disclosure, compact and / or stabilized DNBs have been constructed using an X-linker, a barcode Z-linker, and compact oligos. In the following example, X-linker 1 and X-linker 2 were mixed in equal proportions. BC X-linker 1 sequence: (tot 3’) : (SEQ. ID NO: 1) BC X-linker 2 sequence (5’ to 3’) : (SEQ. ID NO: 2) BC Z-linker sequence (5’ to 3’) : (SEQ. ID NO: 3) Compact Oligo sequence (5’ to 3’) : (SEQ ID NO: 4) . Example 1: The underlined TTTTT can be replaced with other linkers with different lengths) .
[0114] Use of oligonucleotide linkers also reduces template copy loss by reducing enzymatic or chemical DNA cutting and / or keeping the cleaved copy still attached by primer / other linker to uncleaned copies. This can help to reduce RHO loss, SNR loss, slow down Q30 drop, and reduce error rate, producing higher quality long reads.
[0115] FIG. 10C shows that a combination of linker oligonucleotides reduces error rate over the course of 200 sequencing cycles. Using multiple crosslinking oligonucleotides to compact amplicons
[0116] During the DNB forming process, linker oligonucleotides were used to hybridize to multiple regions of the same DNB. Linkers were typically designed as two or three, tandem repeats of an identical hybridizing sequence. The hybridizing sequence is typically 15-45 or more bases in length, preferably 20-30 bases in length. Each hybridizing unit can hybridize to a region of the adapter, and since multiple hybridizing units are linked within the oligonucleotide, a single oligonucleotide can link to multiple adapters of the DNB.
[0117] Between each hybridizing sequence would typically be a short spacer sequence such as 3’ to 5’poly-T or poly-A nucleotides. At the 3’ end of the replicated hybridizing sequence is a terminal set of 3 nucleotides that are resistant to degradation by exonuclease activity of Phi29 polymerase and resistant to extension by incorporation of nucleotides. This is achieved by using a 2’-O-methyl group on the three terminal nucleotides (available from IDT DNA technologies) .
[0118] FIGS. 11A and 11B depict a two unit (1, 1) linker and a three unit (2, 2, 2) linker hybridized to two sites within an adapter of a DNB. FIG. 11A shows the linker sequences (SEQ ID NOS. 5 to 14) . FIG. 11B depicts the portions of the linkers that anneal to a complementary sequence in the adaptor molecule in the template. Each linker has the potential to hybridize to other adapters due to the tandem arrangement of hybridizing sites. Each linker is blocked from exonuclease digestion or polymerase extension by virtue of one or more blocking nucleotides (4) . A region within the adapter may be available for later hybridization of a sequencing primer (3) . Adaptor design optimized for binding of crosslinker oligonucleotides
[0119] FIG. 11C shows two alternative adapter designs that increase sites for hybridizing linkers next to the insert being sequenced. Version 1 is an example of the traditional design in which a barcode sequencing primer (1) is located upstream of a barcode region (3) of the adapter. Downstream to the barcode region is a site for primer (2) binding which allows sequencing into the insert (8) . In Version 2, the barcode region (3) is moved closer to the 3’ end of the adapter and is sequenced by primer (4) during the initial cycles of sequencing. The extended strand from primer (4) continues on to sequence the insert after passing through the barcode region. This arrangement allows for additional adapter sequences upstream that could accommodate two (for example, (6) and (7) ) or more blocked linker oligonucleotides: for example, (6) and (7) .
[0120] As an example of the use of these linking oligonucleotides, 2 unit and 3 unit versions were tested. Also tested was using 1 or 2 oligonucleotides that bind to different regions of the same adapter. The data in Table 2 demonstrate that in lane 1, a single 2-unit linker had sequence splits ( “vSeqSplits” ) , or the percent of DNBs that can be scored to have split across two binding sites on the array, of over 5%. However, in lane 4 with two oligonucleotides, each of 3 units, that percentage of DNBs had dropped to 2%. Reverse complement adapter sequences with self-annealing palindromic sequences
[0121] Another feature that can compact DNBs and potentially other amplicons is reverse complement hairpin sequences in the adapter. In this arrangement, sequences that are in tandem with each other contain the reverse complement sequence of the other, which self-hybridize in newly generated DNB strands. For example, the 10-base sequence AAAGGTCCTT (SEQ ID NO. 15) can be used in tandem with AAGGACCTTT (SEQ ID NO. 16) to create the 20 base sequence AAAGGTCCTTAAGGACCTTT (SEQ ID NO. 17) between surrounding sequences in the adaptor. Compaction occurs by hairpin formation of the 10-mers or hybridization to more distal 20-mer sequences. In these arrangements, the reverse complement sequences can be separated by intervening sequences within the adapter so that hairpin formation can occur within an adapter. Hybridization can also occur to other adapters in the DNB. Reverse complement sequences can be from 6 x 6 bases, up to 20 x 20 bases in length. Preferably, reverse complement sequences are 10 x 10 bases, or from 8 x 8 to 12 x 12 bases in length. Feature 2C: Compacting and stabilizing amplicons using high magnesium concentration and other additives Solutes that promote DNB compacting
[0122] Compacting and / or reading DNBs can be assisted by including in the reaction mixture metal cations such as Ca++ or Mg++, aprotic polar solvents such as dimethyl sulfoxide or cyrene, and DNA minor groove binders such as CDPI3, duocarmycin A, Hoechst 33258 or Hoechst 33342. Once distributed in an array, DNBs can be condensed and stabilized further using alcohol and / or polyethylene glycol, and coated with protein. U.S. Patents 10,837,879 and 11,835,437. Amplicon compacting effect of high magnesium concentration
[0123] The makers of this technology have discovered that using unusually high concentrations of magnesium helps compact DNBs: concentrations of 15 to 100 mM, or 20 to 50 mM, or 15 to 60 mM, or at least 15, 20, 25, or 35 mM. This is considerably higher than is typically used in either the preparation or the plating of DNA amplicons. This is considerably higher than magnesium concentrations used in prior art methods. See, for example, L. Chen et al., “Locally denatured DNA compaction by divalent cations, J Phys Chem B. 2023 Jun 1; 127 (21) : 4783-4789, who studied this question and recommends a Mg++concentration of 3 to 10 mM.
[0124] The effect of the magnesium is to prevent DNBs from agglomerating, and also to prevent DNBs from splitting between neighboring binding sites, which would render the spot in which it lands void in terms of data capture or interpretation. Loss can be due to DNB splitting between two spots, and / or two different signals coming from one spot. DMSO can be include to increase melting temperature, compensating the effect of the magnesium. A high magnesium concentration may also have a minor effect of increasing the proportion of DNBs that have low signal intensity (often because they are empty) . However, the disadvantage of this effect is outweighed by the benefits in improving yield of spots occupied by a single DNB.
[0125] FIGS. 12A to 12F demonstrate the benefit of high magnesium concentrations in action. The data shown are from a working example in which concatemers were formed from a human genomic DNA library and plated on an array at four different concentrations of Mg++: 10 mM, 15 mM, 25 mM, and 35 mM. Interpretation of the symbols is as follows: Box = interquartile range IQR (i.e. 25 percentile to 75 percentile) ; line in the box = median; whiskers = lines extending from the box edge to the highest and lowest value inside the “fence” , which is 1.5x the IQR of the data above or below the box. Dots = any outliers beyond the fence.
[0126] FIG. 12A shows the yield of binding sites (spots) on the array that have a single productive DNB attached thereto. The numbers shown are each a count of such spots as a total of concatemer binding positions within the field observed. With increasing magnesium concentration, yield increases from 75%to almost 82% (an improvement of about 7%) . FIG. 12B shows spots that have a mixed signal, which is a position in which portions of two different concatemers are bound. These spots decrease about 3%with higher magnesium concentration. FIG. 12C shows spots that have a low signal, which is a position which lacks a complete concatemer. These spots increase slightly with higher magnesium concentration, but by only about 1%. FIG. 12D shows adjacent spots with a concatemer split between them. These spots decease by about 5%at higher magnesium concentration.
[0127] FIG. 12E shows together the spots that have mixed signals, low signals, or split spots. Signal from all these spots are eliminated from the dataset during sequence assembly, thereby reducing the total yield of sequence data. The decrease in spots with mixed signal and split concatemers correlate with the increase in single spot yield shown in FIG. 12A. Relative signal intensity (the brightness of each productive spot) shown in FIG. 12F increases from about 14000 to about 16000 (an increase of between 10 and 15%) . Making 100 kb+ DNBs using 2 or 3 compact-oligo crosslinkers at high magnesium concentration
[0128] Adapters with 2-3 compact-oligo binding sites 15-25 bases in length are used to make a long insert library (for example, one kilobase in size, or an average size between 500-1500 bases) . A low but sufficient concentration of each compact oligo is used in a reaction mixture configured to make large DNBs, containing 20-50 mM Mg++ and other additives, such as DMSO. Rolling circle time is typically 100+. 150+, or 200+ minutes. After being made, the DNBs are preferably heated 20-60 sec at about 45 ℃(40-50 ℃) to condense them further and improve stability. Preferable arraying conditions are in neutral pH. Crosslinking oligos act as staplers, keeping 100 kb+ DNA concatemers or inserts packed tightly together to prevent them from splitting in the pre-loading, loading or post-loading process. Two to three compact oligos with 2 copies provide multiple binding of each copy to other copies compare to one compact oligo with 3-4 copies. Unused two-copy oligos are shorter, reducing impact on DNB loading compared with 3-4 copy oligos. Long read sequencing demonstration
[0129] As a demonstration of large DNB technology for long read sequencing, 150 minute RCR time DNBs were prepared with an additional 35 mM MgCl2 added to the make the DNB mixture. Two linkers, each with two DNB binding subunits were included in the RCR reaction. The DNBs were loaded for 60 minutes on flow cells before hybridization of X-linker sequencing primers for sequencing extension. To provide further structural support for the DNBs, two X-linkers were hybridized to the DNBs that were blocked at the 3’ end from further incorporation extension. Blocking was achieved using inverted-T base incorporation (IDT Technologies) at the 3’ end. Each of these X-linkers was prepared by pre-hybridization of two oligonucleotides to form a single X-linker, with complementary sequences in the oligonucleotide pair. One thousand and fifty sequencing cycles were performed, and the DNBs from randomly selected field of views were analyzed for mapping rates to the human genome.
[0130] Table 3 shows that for one field, mapping of 600 cycles had rates of around 72%. Mapping of the same field for 800 cycles showed mapping rates of around 62%.
[0131] Rates of DNBs with Q scores (sequence quality scores) greater than 30 were determined for DNBs in the lanes. FIG. 11D shows that by cycle 600, over 79%of DNBs had Q scores greater than 30. By cycle 1000, the number of DNBs having Q scores greater than 30 dropped to 35%. Feature 3A: Keeping primer extension products in phase by performing multiple steps of reversible terminator (RT) incorporation in each sequencing cycle
[0132] To ensure accurate base calling, the multiple reads within a amplicon need to be in-phase with each other in relation to the insert position being read. Asynchrony of primer extension can lead to run-on (+1) extension products with an extra base near the sequencing end, and lag products (-1) that are missing a base at the sequencing end.
[0133] Factors that may cause individual extension products to be one base ahead of the correct position (plus-one out of phase) may arise from extraneous unblocked nucleotides becoming incorporated before or during incorporation of blocked nucleotides . Unblocked nucleotides could arise from contamination of the polymerase, contamination of the blocked nucleotide set, or residual nucleotides from amplicon synthesis. Another potential cause of plus-one out-of-phase extension products is the premature unblocking of incorporated blocked nucleotides before all incorporation steps are complete.
[0134] To help primer extension products from getting out of phase, reaction conditions may be adjusted such that RT incorporation in each sequencing cycle goes as far as possible to completion. To achieve this, multiple RT incorporations are performed with reagent exchange between each incorporation for each sequencing cycle. This can help to remove and replace dysfunctional incorporations from the earlier steps, thus minimizing the extension products that go out of phase by minus one (missing one of the bases in the sequence) .
[0135] FIGS. 13C and 13D provide a demonstration. After each imaging step, second and third RT incorporation were conducted. Over 80 cycles, ~3.5%combined out-of-phase was accumulated (lag is 0.044%per cycle) . If polymerase was removed in the third incorporation, lag would increase from 0.037%to 0.046%per cycle. If polymerase was removed both in the second and third RT incorporation, lag would be doubled. Therefore, two additional RT incorporations is very effective for maintaining the low (-1) out-of-phase.
[0136] In the event that sequencing of large amplicons get out of phase before reaching the end of the insert, the primer extension products can be brought back into phase using the technology described in granted European patent EP 4121554 B1.
[0137] For example, sequencing cycles can be conducted wherein one of the nucleotides is a reversible terminator blocked with a first blocking group and the other three nucleotides are not blocked, and then unblocking the extended primers. Particularly effective is dinucleotide-frequency rephasing (DFR) , in which each different mixtures of blocked and unblocked nucleotides are used until each primer is extended to a selected dinucleotide -whereupon all nucleotides are unblocked, and sequencing of the rephased primers is resumed. Feature 3B: Keeping primer extension products in phase by improving washing conditions
[0138] It is beneficial to purge the DNA polymerase, antibodies, and other reagents from the previous sequencing cycle that may interfere with the subsequent cycle. This can improve signal-to-noise ratio (SNR) , and potentially reduce signal loss (or dye quenching by unwashed molecules) though successive cycles of primer extension.
[0139] Following incorporation of each RT (reversible terminator) , a wash can be performed to purge incorporation reaction reagents and potentially bound polymerase on the amplicon, before the addition of labeled antibodies for the detection of the incorporated azidomethyl modified and base specific nucleotide. In the first antibody binding step of a sequencing cycle, only two of the four labeled antibodies are combined and allowed to bind (2c4i detection) . Each of the two antibodies is labeled with a fluorescent dye distinct from each other to minimize spectral cross talk of the dyes. To minimize non-specific cross-binding of the labeled antibodies, the antibodies for the other two bases can be included in unlabeled form.
[0140] The first imaging step typically collects the images for the two antibody bindings at two different wavelengths. After imaging, the antibodies are removed using incorporation mix to provide a pool of competitive binding antigens (free azidomethyl nucleotides) that promote dissociation of the antibodies from the incorporated azidomethyl nucleoside. A second function of the incorporation step is to further complete incorporation of the nucleotide to ensure all subunits of the amplicon are in-phase for the correct reading position of the insert. With sufficient removal of the first antibody, a second binding of antibodies can occur for the second pair of bases, followed by imaging with the same pair of dyes as used for the first imaging. Again, after imaging the antibodies can be removed with a third incorporation step using the same nucleotide / polymerase mix. A total of three incorporation steps within one sequencing cycle, with intervening washing, may ensure highly efficient in-phase alignment of all subunits to the same insert position. The 3’a zidomethyl group of the incorporated nucleoside can be unblocked with a phosphine such as THPP to return the 3’-azidomethyl group to a hydroxyl group.
[0141] FIGS. 14A and 14B show the cross-talk between nucleotide bases observed in consecutive sequencing cycles. A flat line in the lower scatter indicates accurate base determination, whereas an upwards sloping line indicates accumulating error. Including unlabeled RTs in the wash solution (FIG. 14B) improves sequencing accuracy compared with a wash solution not containing unlabeled RTs. Feature 3C: Keeping primer extension products in phase using a mutated DNA polymerase that has a lower preference for 3’-OH nucleotides
[0142] To help primer extension products from getting out of phase with each other in multiple cycles of sequencing, it is helpful to use a DNA polymerase that can add blocked nucleotides to the 3’ end of the growing chain with similar efficiency as regular nucleotides.
[0143] PCT publication WO 2022 / 194244 describes a family of DNA polymerases (such as Taq polymerase) that have amino acid mutations that permit template-directed incorporation of 3’-phosphate nucleotides and 3’-O-blocked nucleotides, such as 3’-O-nitrobenzyl (NB) -modified nucleotides, into primer extension products. Effective mutations are at any of positions 611 to 617 in motif A, or positions 655, 657, 681, 742, or 747 of the Taq polymerase amino acid sequence.
[0144] WO 2022 / 247055 describes a family of thermostable DNA polymerase enzymes based on the Pyrococcus abyssi DNA polymerase exo-mutant. Exemplary is a DNA polymerase comprising at least three amino acid mutations in the following four sites or functionally equivalent sites: position 409, position 410, position 411, and position 486. The modified enzymes have greatly improved incorporation efficiency of specific non-natural dNTPs, which thereby improves the sequencing speed and sequencing quality of the sequencing by synthesis (SBS) . WO 2022 / 247055 provides a thermostable B-family DNA polymerase variant with mutations at amino acid residues in the following: A-group sites (408 / 409 / 410) , B-group sites (451 / 485) , C-group sites (389 / 383 / 384) , D-group sites (589 / 67 / 680) , E-group sites (491 / 493 / 494 / 497) , and F-group sites (474 / 478 / 486 / 480 / 484) . Compared with non-mutant DNA polymerases, the mutant DNA polymerases exhibit a strong polymerization capability in catalyzing 3'blocking-modified nucleotides for high-throughput sequencing.
[0145] To place the improved DNA polymerase enzymes in the context of the technology described here, mutated DNA polymerase PX617 was used for SE300 sequencing. 40 μg / mL (2x) PX617 and 6 μM (3x) cold dNTP were applied for prepare incorporation mixture, compare to Pz1095 (3x) , PX617 reduce RHO loss from 0.21%-0.24%to 0.14%-0.18%and 0.25%-0.27%to 0.19%-0.22%per cycle under sequencing primer hybridization with and without BC-X-linker, respectively.
[0146] FIG. 15A shows that using PX617 DNA polymerase flattens the average run-on, reducing sequence lag from 0.04%to 0.03%per cycle. This confirms that PX617 has less preference for 3’-OH nucleotides. Comparing FIGS. 15B and 15C shows that using PX617 results in fewer incorporation mismatches (< 0.04%after 300 cycles) than Pz1095 (< 0.10%after 300 cycle) . There was very low out-of-phase template copies either -1 or +1. There was minimal incorporation of mismatches that block further extension, despite mismatches to reduce signal loss. Feature 4A: Increasing signal intensity
[0147] Efficient labeling with bright or multiple dye molecules generates better images, which produces more productive reads in later cycles of sequencing that have a smaller number of usable template copies.
[0148] For example, more than one (2, 3, 5, 10, or 15) fluorophores can be attached to each RT-specific antibody. To help increase the fluorophore to antibody ratio, whole IgG or IgM monoclonal antibody molecules can be used, rather than Fab or other smaller fragments. Another option is to leave the primary antibody unlabeled, and detect it using a multiply labeled secondary antibody or binding agent (such as rabbit anti-mouse IgG, or avidin) that specifically binds to the primary antibody. In tests of this technology, “A” specific antibody and the “G” specific antibody were typically labeled with the shorter wavelength AF532 dye and the “T” and “C” specific antibodies were labeled with the longer wavelength AF647 dye, although other dye / antibody combinations can be used instead.
[0149] In the following demonstration, different dyes were tested for increasing labeling quality: iF647, zF647, XFD647 and Cy5.
[0150] FIGS. 13A and 13B show results of a staged sequencing run in which C and T were labeled with iF647 and zF647 respectively. Signals from iF647 and zF647 did not show reduction after 4 days incubation at 4℃. iF647 labeling produced higher signal than zF647 signal, shown by a drop of the intensity at the 20th cycle.
[0151] In other experiments (not shown) , C-iF647 and T-iF647 signals were higher than C-EF660 and T-EF660 with DMTU, while A and G signals were also higher. C-iF647 and T-iF647 signal loss was similar to C-EF660 and T-EF660 with DMTU. The error rate and mismatch from iF664 were comparable with EF660. C-XFD647 and T-XFD647 signals were higher than C-EF660 and T-EF660. C-XFD647 and T-XFD647 signal loss was less than C-EF660 and T-EF660 signal loss. Mismatch of T-A, T-C, T-G from XFD647 was lower than EF660 mismatch.
[0152] Another way to increase signal intensity is to increase the number of primer extension products within each amplicon. The number of template copies can be 150-300, 200-500, 300-500, or greater than 1000, 500, 300, 200, or 150. With 150-300 copies, there is only 20%signal loss per 100 bases, with 75%signal loss after 600 cycles and 87%loss at 900 cycles. With long inserts, the number of copies in the amplicon will be about 30 to 50. One way to achieve a higher copy number is by amplification of the amplicon in situ (after arraying on a surface) . U.S. Patent No. 7,960,104. Another way is multiple displacement amplification, creating single strand DNA branches complementary to the insert sequence on the amplicon. U.S. Patent No. 10,227,647. US 2023 / 0295696 A1 describes a method for loading nucleic acid molecules such as DNBs on a solid support by performing amplification reactions both before and after loading. Feature 4B: Increasing detection using two color four image (2c4i) data capture
[0153] Four different reversible terminator (RT) bases on primer extension products can be detected using four differently colored labels. However, the signal from primer extension products diminishes with successive cycles of sequencing, because some of the primer extension products on each amplicon become structurally deficient or go out of phase. It gets harder to differentiate between labels from each of the four bases, even with each label imaged separately.
[0154] The technology in this disclosure provides the insight that cross-talk between the labels inhibits signal detection deeper into the sequence read. This is because even though each label has an average frequency distribution and color, there is some spread to higher and lower frequencies. To remedy this, two labels are chosen that are sufficiently different in color so that there is very little cross-talk, and used to determine successively added bases in a pair-wise fashion.
[0155] To measure all four bases, the amplicons are contacted with antibodies specific for the first pair of RTs (say, A and G) labeled with two distinguishable labels, and detected by imaging at each of the two peak emission frequencies. The RTs that have been imaged are converted to regular nucleotides, and the corresponding antibodies are washed away. The amplicons are then contacted with antibody specific for the second pair of RTs (C and T) , also labeled with two distinguishable labels, and again detected by each of the two peak emission frequencies. The second pair of RTs are then converted for the next round of sequencing. The two labels used on the first pair of antibodies can be the same labels used on the second pair of antibodies: hence the term two color four image (2c4i) data capture, with two separate labeling reactions each cycle of sequencing. The same problem can be addressed using one color and four images (1c4i) , but this would require four separate labeling reactions, which would generally be less cost effective.
[0156] 2c4i processing and other modes of labeling and imaging primer extension products with labeled antibody are described in PCT publication WO 2020 / 097607. By way of illustration, using antibody-labeling of RTs the following protocol may be used.
[0157] After hybridization of the primer to the amplicon the 3’ azidomethyl blocked nucleotide is incorporated. The incorporating enzyme can belong to the class of type B DNA polymerases including those of the Thermococcus sp. 9° N parental strain and Kod parental strain with targeted amino acid modifications to enhance recognition and incorporation of the modified nucleotides. Type A DNA polymerases such as those of the Taq family have also been demonstrated to incorporate 3’ modified nucleotides with optimal amino acid mutation. Typically, the incorporation reaction with a type B polymerase of occurs in a pH buffered solution of pH 8.5 to pH 9.2, pH 8.8 to pH 9.0, pH 9.0 to pH 9.2. Additional additives in the incorporation buffer can include DMSO glycerol, ammonium sulfate, and EDTA. High concentrations of MgCl2 are beneficial. It is also beneficial to provide one sufficiently incorporated nucleotide to produce as close to maximal signal as possible with the binding of antibodies. However, to achieve fast cycle times the incorporation efficiency could be less than 100%, but ideally greater than 90%or 95%. Reaction times of 10 to 20 seconds are usually sufficient.
[0158] Following incorporation, a wash can be performed to remove incorporation reaction reagents and potentially bound polymerase on the amplicon, before the addition of labeled antibodies for the detection of the incorporated azidomethyl modified and base specific nucleotide. In the first antibody binding step of a sequencing cycle only two of the four labeled antibodies are combined and allowed to bind. To minimize chances of cross-binding of antibodies, unlabeled antibodies specific to the other pair of RTs could be included. Each of the two antibodies is labeled with a fluorescent dye distinct from each other to enable low levels of spectral cross talk of the dyes.
[0159] The first imaging step typically collects the images for the two antibody bindings at two different wavelengths. After imaging, the antibodies are removed by using incorporation mix to provide a pool of competitive binding antigens (free azidomethyl nucleotides) to promote dissociation of the antibodies from the incorporated azidomethyl nucleoside. In addition, the dissociation is promoted using higher temperature at high pH. A second function of the incorporation step is to further complete incorporation of the nucleotide to ensure all subunits of the amplicon are in-phase for the correct reading position of the insert. With sufficient removal of the first antibody binding, a second binding of antibodies can occur for the second pair of bases, followed by imaging with the same pair of dyes as used for the first imaging. Again, after imaging, the antibodies can be removed with a third incorporation step using the same nucleotide / polymerase mix. A total of three incorporation steps within one sequencing cycle, with intervening washing, may act to ensure highly efficient in-phase alignment of all subunits to the same insert position. Finally, the 3’-azidomethyl group of the incorporated nucleoside can be unblocked with a phosphine such as tris (hydroxypropyl) phosphine (THPP) to return the 3’-azidomethyl group to a hydroxyl group.
[0160] FIGS. 16A and 16B show that compared with 2c2i labeling and imaging, 2c4i exhibits higher accuracy (Higher Q30 and lower AvgErrorRate) . In FIG. 16A, DNBs from an E. coli genomic library with less than 125 copies of 320 base average template size (prepared by a 25-min RCR reaction) were sequenced using 2c4i with a high Q30 of 94.5%for the first 150 cycles, following which detection was switched to 2c2i, which had a lower accuracy. FIG. 16B shows similar results obtained from larger DNBs prepared by a 100-min RCR reaction. Again, 2c4i image detection was more accurate. Various implementations of 2c4i data capture
[0161] The two-color four-detection methodology can fortify long sequence reads when using other kinds of labeling. 2c4i is cleaner than using four different labels for each of the bases, because there is less overlap in the emission curves with two labels. On the other hand, it is faster and more cost effective than using a single label for each base, which requires twice as many imaging steps for each iteration of sequencing.
[0162] 2c4i detection can be done using terminators that are directly labeled, or tethered to a label at the time of incorporation into the primer extension product. Each iteration of sequencing can comprise the following steps: (a) contacting the primer extension product with analogs of two of the four nucleotide bases (A) , (T) , (C) , and (G) , each bearing a different colored label, under conditions whereby an analog that is complementary to the template will hybridize to the template at a position next to the primer extension product; (b) imaging the different colored labels referred to in step (a) separately; (c) removing colored labels used in step (a) from the primer extension product; (d) contacting the primer extension product with analogs of the other two nucleotide bases, each bearing a different colored label, under conditions whereby an analog that is complementary to the template will hybridize to the template at a position next to the primer extension product; (e) imaging the different colored labels referred to in step (d) separately; (f) removing colored labels used in step (d) from the primer extension product; and (g) preparing the primer extension product for the next iteration of sequencing.
[0163] The nucleotide analogs can themselves constitute a colored label, in which case they are chemically converted to a non-colored form in step (c) or step (f) . Alternatively, the analogs can be conjugated or linked to a colored label, which can be unconjugated or neutralized in step (c) or step (f) . Optimally, when the label is removed, the analog becomes a standard nucleotide. The two colored labels used in step (d) can be the same or different from the colored labels used in step (a) .
[0164] The direct labeling approach works for both sequencing by synthesis (SBS) and sequencing by binding (SBB) . In SBS, one of the nucleotide analogs extends the primer extension product in step (a) or step (d) (depending on which analog is complementary to the template) . In SBB, the nucleotide analogs hybridizes to the template next to the primer extension product in step (a) or step (d) , and washed way or otherwise removed entirely before step (g) . The final step includes extending the primer extension product by one base to replace the analog that has been removed. U.S. Patent Nos. 9,951,385 and 10,161,003 (Pacific Biosciences) .
[0165] 2c4i detection can also be done using terminators that are not directly labeled, but can be distinctively labeled using a secondary reagent after incorporation. Each iteration can comprise the following steps: (a) contacting the primer extension product blocked analogs of all four nucleotide bases (A) , (T) , (C) , and (G) , each bearing a different epitope, under conditions whereby an analog that is complementary to the template will hybridize to the template at a position next to the primer extension product; (b) contacting the primer extension product with detection reagents for two of the epitopes referred to in step (a) , each detection agent bearing a different colored label; (c) imaging the different colored labels referred to in step (b) separately; (d) removing the detection reagents used in step (b) from the primer extension product; (e) contacting the primer extension product with detection reagents for the other two epitopes referred to in step (a) , each detection reagent bearing a different colored label; (f) imaging the different colored labels referred to in step (e) separately; (g) removing the detection reagents used in step (e) from the primer extension product; and (h) preparing the primer extension product for the next iteration of sequencing.
[0166] The antigenically distinguishable epitopes incorporated in step (a) can be the nucleotide analogs themselves. is an embodiment wherein the epitopes are reversable terminators (RT) , which, when unblocked, revert to a standard nucleotide leaving no scar on the extension product. Alternatively, the epitopes can be different molecular haptens, which are antigenically distinguishable small molecules (less than 3000 or 1000 Daltons) conjugated or tethered to the nucleotides. Suitable haptens include biotin, digoxigenin, fluorescein, rhodamine, dinitrophenyl (DNP) , and trinitrophenol.
[0167] The epitope-distinguishing detection reagents used in steps (b) and (e) can each be, for example, an antibody, as defined above. Other antigen binding moieties suitable for this purpose include aptamers, nanobodies, lectins, molecularly imprinted polymers (MIPs) , and specialized binding conjugates. For example, if one of the haptens is biotin, then the antigen binding moiety could be avidin or streptavidin. The two colored labels used in step (e) can be the same or different from the colored labels used in step (b) .
[0168] The secondary labeling approach works for both sequencing by synthesis (SBS) and sequencing by binding (SBB) . For SBS, one of the nucleotide analogs extend the primer extension product in step (a) and is unblocked before or during step (g) . For SBB, one of the nucleotide analogs hybridizes to the template next to the primer extension product in step (a) , and is removed entirely before step (h) . TO prepare for the next iteration, the primer is extended using a standard nucleotide by one base to re place the analog that has been removed after imaging.
[0169] The 2c4i detection protocol is suitable for long sequence reads of any DNA template. As described earlier in this disclosure, it can be used to concurrently sequence template copies in an amplicon in phase. Effectiveness of the technology put forth in this disclosure for obtaining long sequence reads
[0170] A previous limitation to the length of sequence reads is that large DNA amplicons containing longer fragment inserts may split between adjacent binding sites when distributed on a DNA array. The rate that DNBs split between two binding spots on a surface depends on multiple factors, including the spot density, or in other terms the linear spot pitch, and the DNB size, or template copy number. Increasing DNB size can improve long-run spot yield and accuracy by increasing the number of reaction sites per DNB spot, which is one site per template per copy, but since this extra mass also increases the probability of splitting, there is a yield trade-off. Any gains in yield from increased copy number per DNB, which produces higher signal and faster and more accurate reads, were largely offset by losses from more frequent splitting.
[0171] FIGS. 17A to 17F demonstrate that the technological advancements provided in this disclosure greatly reduce the split rate during load for larger DNBs. The improvements in DNB structure and load chemistry lowers split rates and increases data yield, even at three or more times the copy number of previous methods.
[0172] FIG. 17A shows the effect of spot diameter on spot yield vs pitch under various load conditions. Control, standard load, is found in Run 2 L02, triangle markers above. FIG. 17B shows the effect of spot diameter on DNB split rate at various pitches and load conditions. Split rate is much higher for run 2 lane_L02 control conditions, (triangle markers) which is the main cause of yield loss. FIG. 17C: DNB average RHO is the DNB signal above image background. Load conditions, such as those used in run_3 lane_L01 and run_4 lane_L01 have much higher copy number, and therefore give higher signals than control (run_2 lane_L02 triangle markers) .
[0173] FIG. 17D shows the effect on yield (the number of spots bearing one and only one DNB) of spot diameter, pitch (the distance between rows, DNB size (number of nucleotides) , and magnesium concentration. FIG. 17E shows the effect of the same parameters on the percentage of DNBs split between neighboring spots. FIG. 17F shows the same data, graphed to compare the effect of spot pitch on split DNBs. Increasing the insert size of DNBs increases the rate of splitting between DNB binding sites. The presence of high concentration of magnesium, and other factors, compensates and lowers the split rate.
[0174] The advancements put forth in this disclosure combine to allow higher throughput and lower cost DNA sequencing than was previously possible. Other technologies to increase sequence read length and accuracy Fragment library sizing by controlled primer extension (CPE)
[0175] Making amplicons with large inserts requires a library of fragments of the target polynucleotide of minimum size. Conventional modes of library preparation (enzymatic cleavage, shearing, and other means) typically generate libraries with a wide range of fragment sizes -many of which are too small for long fragment reads.
[0176] The scientists at Complete Genomics Inc. have developed a technology of library preparation wherein fragment size is between a preselected minimum and a preselected maximum. A starting preparation of random fragments is generated from a target DNA to be sequenced. A first primer is hybridized to an adaptor at one end of the fragments, and extended by CPE in such a way that only fragments above the minimum are retained, and other fragments are discarded. A second primer is hybridized to an adaptor at the other end of the fragments, and extended by CPE in such a way that only fragments below the maximum are retained, and other fragments are discarded. The chosen size range can be matched to the desired insert size of the DNA nanoballs.
[0177] Patent disclosures for CPE library preparations are being filed separately in the name of Complete Genomics Inc. Said disclosures are hereby incorporated herein by reference in their entireties as enhancements of and embodied combinations with the technology described herein. Sequencing both strands of a circular DNA
[0178] The scientists at Complete Genomics Inc. have also developed a technology for two strand single end (2xSE) DNA sequencing. A target DNA is characterized by simultaneously or consecutively sequencing both strands of each fragment in a fragment library. Sense and antisense concatemers are formed, either by replicating a double stranded circular DNA in opposite directions, or by replicating separate sense and antisense single stranded circular DNAs. Each replicate contains an adaptor with binding sites for sequencing primers, plus a unique a molecule identifier (UMI) barcode sequence or a complementary sequence (cUMI) that can be used for matching and sequence assembly.
[0179] Benefits of having complementary strands of each DNA fragment prepared before beginning sequence reads include the following: ● Longer combined sequence reads: reads begun on opposite strands can be assembled to obtain a combined read that is twice as long; ● Fewer sequencing cycles: Even if insert fragments are short enough for a single read, the same sequence information can be obtained in half the number of sequencing cycles; ● Error correction: Sequence reads in opposite directions that overlap can be used to verify sequencing and identifying missed calls. Sequencing errors can be distinguished from genuine mutations in comparison with a reference sequence; ● No need to make a second strand in situ during sequencing, which may interrupt or interfere with reading of the first strand.
[0180] Patent disclosures for two-strand sequencing technology have been filed separately: for example, U.S. provisional patent application 63 / 733,958, filed December 13, 2024. That application is hereby incorporated herein by reference in its entirety as an enhancement of and embodied combinations with the technology described herein. Other publications
[0181] Other recent advancements in polynucleotide sequencing include the following:
[0182] U.S. Patent No. 10,351,909 describes single molecule arrays for genetic analysis. U.S. Patent No. Nos. 10,125,392 and 11,389,779 describe long read fragment (LFR) nucleic acid analysis by barcoded random mixtures of non-overlapping fragments. U.S. Patent No. 10,557,166 describes multiple tagging of individual long DNA fragments. Granted European patent EP 3790967 B1 describes single tube bead-based DNA co-barcoding for accurate and cost-effective sequencing, haplotyping, and assembly.
[0183] Publication US 2021 / 0189483 A1 describes controlled strand-displacement for paired end or single end sequencing. U.S. Patent No. 7,767,400 describes paired-end reads in sequencing by synthesis. European patent EP 4121554 B1 describes restoring phase in massively parallel sequencing. U.S. Patent No. 10,954,559 describes bubble-shaped adaptor elements for constructing a sequencing library. US 2024 / 0240174 A1 describes nick-ligate single tube long fragment read (stLFR) sequencing. US 2024 / 00423924 A1 describes the determination of long DNA sequences using short MPS reads. US 2024 / 0279644 A1 describes template mutagenesis for improved assembly of sequence reads. Application WO 2024 / 022207 describes methods of in-solution positional co-barcoding for sequencing long DNA molecules.
[0184] U.S. Patent No. 10,301,346 and WO 2008 / 037568 describe nucleotide analogs with a cleavable protective group: reversible terminators. U.S. Patent No. 11,254,961 puts forth polymerase-tethered nucleoside triphosphates for use in nucleic acid synthesis.
[0185] Selective DNA amplification from complex genomes using universal double-sided adapters is described by Callow, M, et al., Nucleic Acids Research, 2004, Vol. 32, No. 2, e21. Drmanac, R. et al. describes human genome sequencing using unchained base reads on self-assembling DNA nanoarrays. Science 327 (5961) : 78-81 (2010) . Additionally, Peters, B., et al. describe accurate whole-genome sequencing and haplotyping from 10 to 20 human cells. Nature, 487: 190-195 (2012) .
[0186] Levy, S. et al. provide a general review of Advancements in Next-Generation Sequencing. Annu. Rev. Genom. Hum. Genet. 2016.17: 95–115. Drmanac, S. et al. describe amethod of advanced massively parallel sequencing using antibodies specific to each natural nucleobase. bioRxiv February 2020, DOI: 10.1101 / 2020.02.19.953307. Wang, Q. Drmanac, R. et al. describe efficient and unique co-barcoding of second-generation sequencing reads from long DNA molecules enabling cost-effective and accurate sequencing, haplotyping, and de novo assembly. Genome Research 2019, 29 (5) : 798–808. Hahn, O., et al. describe for robust sequencing of single-nuclear RNAs captured by droplet-based method. Nucleic Acids Research, 2021, Vol. 49, No. 2, e11.
[0187] Wang, L. et al. describe 3’ branch ligation as a novel method to ligate non-complementary DNA to recessed or internal 3COM ends in DNA or RNA. DNA Research, 2019, 26 (1) , 45–53. Mao, Q. et al. describe whole genome sequences and experimentally phased haplotypes of over 100 personal genomes. GigaScience (2016) 5: 42. Siotlos, S., Drmanac, R. et al. describe whole genome sequence analysis of BT-474 using complete Genomics’ standard and long fragment read technologies. Ciotlos et al. GigaScience (2016) 5: 8.
[0188] Peters, B. et al. describe co-barcoded sequence reads from long DNA fragments: a cost-effective solution for “perfect genome” sequencing. Front. Genet. January 2015, Vol. 5, Article 486; doi: 10.3389 / fgene. 2014.00466. Z. Dong, et al. describe the development of coupling controlled polymerizations by adapter-ligation inmate-pair sequencing for detection of various genomic variants in one single assay. DNA Research, 2019, 26 (4) , 313–325.
[0189] Fang, CB. et al. describe high-resolution single-molecule long-fragment rRNA gene amplicon sequencing of bacterial and eukaryotic microbial communities. 2023, Cell Reports Methods 3, 100437. Murigneux, V. et al. compare long-read methods for sequencing and assembly of a plant genome. GigaScience (2020) , 9, 1–11. Cai, Y. et al. describe assembly and analysis of the genome of Notholithocarpus densiflorus. G3 Genes|Genomes|Genetics, 2024, 14 (5) , Article jkae043.
[0190] Controlled primer extension has been used, for example, in US 2022 / 0195624 A1: High coverage single tube Long Fragment Read (stLFR) technology; and US 2024 / 0043924 A1: Determining long DNA sequence using short MPS reads. Trademarks
[0191] The wordmark CoolMPS, the wordmark DNBSEQ, the wordmark MGIEasy, and the MGI logo are all registered trademarks of MGI Tech Co., Ltd. Incorporation by reference
[0192] For all purposes in the United States of America and other jurisdictions where allowed, each and every publication and patent document referred to in this disclosure is incorporated herein by reference in its entirety for all purposes to the same extent as if each such publication or document was specifically and individually indicated to be incorporated herein by reference. Practice of the claimed invention
[0193] The technology provided in this disclosure and its use are described within an academic understanding of general principles of protein and oligonucleotide chemistry and DNA sequencing technology. These discussions are provided for the edification and interest of the reader, and are not intended to limit the practice of the claimed invention. All of the products and methods claimed in this application may be used for any suitable purpose without restriction, unless explicitly indicated or otherwise required.
[0194] While this disclosure has been described with reference to the specific embodiments, changes can be made and equivalents can be substituted to adapt this disclosure to a particular context or intended use as a matter of routine experimentation, thereby achieving benefits of this disclosure without departing from the scope of what is claimed. Descriptions and illustrations of the technology using DNA nanoballs or concatemers may be adapted wherever applicable to other types of amplicons.
Claims
1.An improvement in a method for nucleotide sequencing, wherein the method comprises:(a) forming amplicons that each contain a plurality of copies of a DNA insert obtained from a fragment library;(b) iteratively sequencing copies of the inserts in the amplicons by extending primers hybridized to an adaptor upstream of each copy of the inserts one base at a time using nucleotides or analogs thereof, thereby forming primer extension products; and(c) determining for each iteration of step (b) whether a nucleotide or analog added to the primer extension products in each amplicon is adenine (A) , cytosine (C) , guanine (G) , or thymine (T) ;the improvement comprising:Feature 1: using a longer DNA insert in each amplicon; and incorporating one or more of the following features in any combination:Feature 2: arraying the amplicons on a surface closer together and / or in a more compact form;Feature 3: minimizing the occurrence of primer extension products in each amplicon getting out of phase with each other; and / orFeature 4 increasing signal intensity and / or sensitivity of detection of which nucleotide or analog has been added to primer extension products in each amplicon in step (c) .2.The method of claim 1, wherein the amplicons formed in step (a) are DNA concatemers, DNA nanoballs (DNBs) , PCR clusters, or amplicons on beads generated by emulsion PCR.3.The method of claim 1, wherein the sequencing in step (b) is sequencing-by-synthesis or sequencing-by-binding.4.The method of claims 1 to 3, wherein the nucleotide analogs used in step (b) contain or are tethered to a fluorescent label, a hapten such as biotin, and / or a reversible terminator.5.The method of claims 1 to 4, wherein the determining in step (c) comprises contacting the primer extension products with detection reagents which are antibodies or other antigen binding moieties bearing fluorescent labels.6.The method of any of claims 1 to 5, comprising Features 1, 2, and 3.7.The method of any of claims 1 to 5, comprising Features 1, 2, and 4.8.The method of any of claims 1 to 5, comprising Features 1, 3, and 4.9.The method of any of claims 1 to 5, comprising all of Features 1, 2, 3, and 4.10.The method of any preceding claim, wherein the median length of the DNA inserts in the amplicons is at least 600, between 600 and 1000, or at least 1000 bases (Feature 1) .11.The method of any preceding claim, wherein there are 30 to 50 copies of each DNA insert in each amplicon.12.The method of any preceding claim, wherein the amplicons are DNBs, and the improvement includes arraying the DNBs on a surface closer together (Feature 2) .13.The method of claim 12, wherein the DNBs are arrayed on a surface at a median center-to-center distance of no more than 300 nm or 500 nm.14.The method of any preceding claim, wherein the amplicons are DNBs, and the improvement includes compacting and / or cross-linking each DNB before sequencing (Feature 2) .15.The method of claim 14, wherein each DNB comprises copies of an adaptor between copies of an insert, wherein the DNBs are compacted using oligonucleotides that crosslink between copies of the adaptor.16.The method of claim 15, wherein the crosslinking oligonucleotides hybridize to the copies of the adaptor at a position that is adjacent to and upstream from a sequencing primer binding site.17.The method of claims 14 to 16, wherein each DNB comprises copies of an adaptor between copies of an insert, and the adaptor contains self-complementary sequences that cross-hybridize between different copies of the adaptor, thereby compacting the DNB.18.The method of claims 14 to 17, wherein the DNBs are compacted during or after plating on a surface using alcohol, polyethylene glycol (PEG) , polypropylene glycol (PPG) , dimethyl sulfoxide, a DNA minor groove binder moiety such as CDPI3, or a combination thereof.19.The method of claims 14 to 18, wherein the DNBs are compacted in a reaction mixture during preparation or plating on a surface using magnesium cation in the reaction mixture at a concentration between 20 and 50 mM.20.The method of any preceding claim, wherein the improvement includes minimizing occurrence of the primer extension products getting out of phase with each other in each amplicon (Feature 3) .21.The method of claim 20, wherein the improvement includes ensuring that the step of adding each nucleotide or analog to primer extension products in each iteration of step (d) is at least 99%complete.22.The method of claim 20 or 21, wherein the primer extension products are contacted with each of the nucleotides or nucleotide analogs more than once during each iteration of the sequencing.23.The method of claims 20 to 22, wherein the primer extension products are washed between each iteration of the sequencing with a solution that contains unlabeled nucleotide analogs and / or unlabeled antibody.24.The method of claims 20 to 23, wherein a DNA polymerase used to catalyze primer extension in step (b) is a DNA polymerase that has been mutated or otherwise adapted to decrease preference for 3’-OH nucleotides.25.The method of any preceding claim, wherein determining which nucleotide or analog has been added to the primer extension products in each amplicon step (c) is improved by using labeled detection reagents with brighter signal intensity (Feature 4) .26.The method of claim 25, wherein the signal intensity is increased by using specific detection reagents that bear fluorophores at least 1.5-fold brighter than fluorescein, such as iF647 or zF647.27.The method of claim 25, wherein the signal intensity is increased by using specific detection agents such as antibodies that bear at least five fluorescent labels each.28.The method of claim 25, wherein the signal intensity is increased by increasing the number of copies of the DNA inserts in the amplicons after arraying on a surface by in situ amplification.29.The method of any preceding claim, wherein sensitivity of detection of the primer extension products in each amplicon in step (d) is increased by using two label four image (2c4i) detection (Feature 4) .30.The method of any preceding claim, whereby sequence reads having a median length of at least 400, 600, 800 or 1000 bases are obtained.31.A method for nucleotide sequencing of a target polynucleotide, comprisingforming a plurality of amplicons, each of which comprises a DNA insert from a fragment library replicated to form a plurality of substantially identical DNA molecules linked or grouped together; andobtaining sequence reads of the insert in each amplicon by an iterative process that comprises primer extension;wherein the median size of inserts in the amplicons is at least 600 bases, and the median size of sequence reads obtained therefrom is at least 400 bases.32.The method of claim 31, wherein the amplicons are DNA nanoballs.33.The method of claim 31, wherein the amplicons are PCR clusters on a surface or emulsion PCR products on beads.34.The method of claim 31, wherein the amplicons are arrayed on a surface separated from each other by a center-to-center distance of no more than 300 or 500 nm.35.The method of claims 31 to 34, wherein the amplicons are DNBs, compacted by a plurality of oligonucleotides that hybridize between and crosslink different adaptors in each DNB.36.The method of claims 31 to 35, wherein the amplicons are DNBs, compacted by a plurality of adaptors within each amplicon that cross-hybridize with each other.37.The method of claims 31 to 36, wherein the amplicons are DNBs, compacted during preparation and / or while plating onto a surface using magnesium cation at a concentration between 20 and 50 mM.38.The method of claims 31 to 37, wherein the products of primer extension in each amplicon are kept in phase by using more than one incorporation step of analogs of each base (A) , (T) , (C) , and (G) using DNA polymerase for each iteration of sequencing.39.The method of claims 31 to 38, wherein the products of prier extension in each amplicon are kept in phase by washing the amplicons between each base incorporation or between each iteration of the sequencing with a solution that contains unlabeled nucleotides, unlabeled analogs, and / or unlabeled antibody.40.The method of claims 38 and 39, wherein said analogs are reversible terminators (RTs) .41.The method of claims 31 to 40, wherein the sequence reads are obtained using antibodies or other antigen binding moieties that specifically bind to terminal nucleotides, wherein each of the antibodies or antigen binding moieties bear a plurality of fluorophores.42.The method of claims 31 to 41, wherein the sequence reads are obtained using two label four image (2c4i) detection.43.A plurality of DNA nanoballs (DNBs) , each comprising replicate of a DNA insert to be sequenced, wherein the median size of such inserts in the DNBs is at least 600, 800, or 1000 bases;wherein the DNA nanoballs are configured to obtain sequence reads of the insert of at least 400 bases.44.The DNBs of claim 43, arrayed on a surface separated from each other by a center-to-center distance of no more than 300 or 500 nm.45.The DNBs of claim 43 or 44, each compacted by a plurality of crosslinking oligonucleotides that hybridize between and crosslink different adaptors situated between adjacent replicates of the insert in the DNBs.46.The DNBs of claims 45, wherein the adaptors comprise a plurality of binding sites for crosslinking oligonucleotides that are upstream of and adjacent to a binding site for a sequencing primer and the insert to be sequenced.47.The DNBs of claims 43 to 46, each compacted by adaptors situated between adjacent replicates in the DNBs that cross-hybridize with each other.48.A method for keeping primer extension products in phase during iterative sequencing of an amplicon;wherein the amplicon comprises a DNA insert from a fragment library replicated to form a plurality of substantially identical DNA molecules linked or grouped together;wherein the iterative sequencing comprises iteratively adding and identifying one or more nucleotides to a primer extension product hybridized to each copy of the insert;wherein the primer extension products are kept in phase:(i) by using more than one incorporation step for each of the four nucleotide bases (A) , (T) , (C) , and (G) for each iteration of sequencing;(ii) by washing each amplicon with a solution that contains unlabeled nucleotides, unlabeled nucleotide analogs and / or unlabeled antibody between each elongation reaction and / or each iteration of the sequencing; and / or(iii) by forming the primer extension products using a DNA polymerase that is modified to decrease preference for 3’-OH nucleotides by incorporating one or more amino acid substitutions.49.The method of claim 48, wherein the method incorporates all three of (i) , (ii) , and (iii) .50.The method of claim 48 or 49, wherein the amplicon is a DNA nanoball.51.The method of claims 48 or 49, wherein the amplicon is a PCR cluster on a surface or a product of emulsion PCR on a bead.52.A direct two-color four-image (2c4i) method for obtaining sequence reads of at least 400 bases of a DNA template by iteratively extending a primer extension product in a sequence dependent manner, wherein each iteration comprises:(a) contacting the primer extension product with analogs of two of the four nucleotide bases (A) , (T) , (C) , and (G) , each bearing a different colored label, under conditions whereby an analog that is complementary to the template will hybridize to the template at a position next to the primer extension product;(b) imaging the different colored labels referred to in step (a) separately;(c) removing colored labels used in step (a) from the primer extension product;(d) contacting the primer extension product with analogs of the other two nucleotide bases, each bearing a different colored label, under conditions whereby an analog that is complementary to the template will hybridize to the template at a position next to the primer extension product;(e) imaging the different colored labels referred to in step (d) separately;(f) removing colored labels used in step (d) from the primer extension product; and(g) preparing the primer extension product for the next iteration of sequencing.53.The method of claim 52, which is a method of sequencing by synthesis (SBS) ,wherein the nucleotide analogs are blocked,wherein one of the nucleotide analogs extends the primer extension product in each iteration of sequencing in step (a) or step (d) and is unblocked before or during step (g) .54.The method of claim 52, which is a method of sequencing by binding (SBB) ,wherein one of the nucleotide analogs hybridizes to the template next to the primer extension product in each iteration of sequencing in step (a) or step (d) , and is removed entirely before step (g) , andwherein step (g) includes extending the primer extension product by one base to replace said analog.55.An indirect two-color four-image (2c4i) method for obtaining sequence reads of at least 400 bases of a DNA template by iteratively extending a primer extension product in a sequence dependent manner, wherein each iteration comprises:(a) contacting the primer extension product with analogs of all four nucleotide bases (A) , (T) , (C) , and (G) , each bearing a different epitope, under conditions whereby an analog that is complementary to the template will hybridize to the template at a position next to the primer extension product;(b) contacting the primer extension product with detection reagents for two of the epitopes referred to in step (a) , each detection agent bearing a different colored label;(c) imaging the different colored labels referred to in step (b) separately;(d) removing the detection reagents used in step (b) from the primer extension product;(e) contacting the primer extension product with detection reagents for the other two epitopes referred to in step (a) , each detection reagent bearing a different colored label;(f) imaging the different colored labels referred to in step (e) separately;(g) removing the detection reagents used in step (e) from the primer extension product; and(h) preparing the primer extension product for the next iteration of sequencing.56.The method of claim 55, wherein in each iteration of sequencing, each of the epitopes referred to in step (a) is a blocked reversable terminator (RT) that reverts to a standard nucleotide in step (d) or step (g) .57.The method of claim 55, wherein in each iteration of sequencing, each of the epitopes is a different hapten conjugated each of the four analogs.58.The method of claim 57, wherein each hapten is selected from biotin, digoxigenin, fluorescein, rhodamine, dinitrophenyl (DNP) , and trinitrophenol.59.The method of claims 55 to 58, wherein each of the detection reagents is an epitope-specific antibody or antigen binding moiety.60.The method of claims 55 to 59, which is a method of sequencing by synthesis (SBS) ,wherein the nucleotide analogs are blocked,wherein in each iteration of sequencing, one of the nucleotide analogs extends the primer extension product in step (a) and is unblocked before or during step (g) .61.The method of claims 55 to 59, which is a method of sequencing by binding (SBB) ,wherein in each iteration of sequencing, one of the nucleotide analogs hybridizes to the template next to the primer extension product in step (a) , and is removed entirely before step (h) , andwherein step (h) includes extending the primer extension product by one base to replace said analog.62.The method of any of claims 50 to 61, wherein a plurality of copies of said template is present in a DNB or other amplicon and are sequenced simultaneously.
Citation Information
Patent Citations
Method for constructing long fragment DNA library
US20180195060A1
Adding nucleotides during sequence detection
US20200190576A1
Method for sequencing long-fragment nucleic acid
US20210324466A1
Compositions and methods for pairwise sequencing
US20230203564A1
Determining long DNA sequence using short MPS reads
US20240043924A1