Methods for detecting a target polynucleotide
The enzymatic extension of immobilized polynucleotide sequences with thermal cycling enhances detection sensitivity by increasing binding sites, addressing the limitations of traditional methods in polynucleotide sequence detection.
Patent Information
- Application Number
- JP2020544825
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2018-02-26
- Filing Date
- 2019-02-25
- Publication Date
- 2025-09-03
- Estimated Expiration
- 2039-02-25
AI Technical Summary
Existing methods for detecting target polynucleotide sequences, such as DNA or RNA, suffer from limited sensitivity due to only one hybridization event per probe, leading to reduced signal-to-noise ratios and the need for specialized instrumentation and meticulous sample preparation.
A method involving the enzymatic extension of immobilized polynucleotide sequences containing tandem repeats, using thermal cycling to increase the number of binding sites for target sequences, thereby enhancing detection sensitivity.
The method improves detection accuracy and allows for the use of less sensitive detectors by increasing the number of binding sites for target sequences, overcoming the limitations of signal-to-noise ratios in traditional hybridization-based detection systems.
Smart Images

Figure 0007733382000006 
Figure 0007733382000007 
Figure 0007733382000008
Abstract
Description
[Technical Field]
[0001] The present invention provides a thermal cycling method for increasing the number of tandem repeats of a unit sequence having a length of 1 to 60 nucleotides in a linear polynucleotide. The present invention also provides a solid substrate having a surface on which at least one linear probe polynucleotide is immobilized, wherein the at least one linear probe polynucleotide comprises at least two tandem repeats of a unit sequence having a length of 1 to 60 nucleotides. Also provided is a method for determining the presence of a linear target polynucleotide sequence in a test sample using the solid substrate. [Background technology]
[0002] Methods for determining the presence of target polynucleotide sequences in test samples are routinely used in basic research and are highly useful in diagnostics and therapeutics. Libraries of polynucleotides with defined sequences can also be used as identifiers (DNA barcodes) for security applications.
[0003] The presence or absence of target polynucleotide sequence (such as DNA or RNA) in test sample can be easily determined by several methods known in the art.Non-limiting examples of such methods include in situ hybridization, microarray analysis, PCR and next-generation sequencing.Such methods have been found to be useful for, for example, identifying or monitoring the biomarker or therapeutic target of disease.
[0004] In situ hybridization (ISH) is based on the complementary pairing of labeled DNA or RNA probes with normal or abnormal polynucleotide sequences in intact chromosomes, cells, or tissue sections. Because ISH is similar to immunohistochemistry, it is more compatible with histopathologists than other molecular biology techniques that can be applied to anatomical pathology. It offers the unique advantage of enabling the localization and visualization of target polynucleotide sequences within morphologically identifiable cells or cellular structures, as opposed to other molecular biology techniques that are primarily based on probe hybridization with polynucleotides extracted from homogenized tissue samples.
[0005] Polynucleotide arrays, or simply DNA arrays, are a group of technologies in which specific DNA sequences are deposited or synthesized on a surface in a two-dimensional (and sometimes three-dimensional) array, with the DNA covalently or noncovalently attached to the surface. In typical use, DNA arrays are used to probe a solution of a mixture of labeled polynucleotides, and the binding (by hybridization) of these "targets" to the "probes" on the array is used to measure the relative concentrations of polynucleotide species in the solution. By generalizing to a very large number of DNA spots, arrays can be used to quantify an arbitrarily large number of different polynucleotide sequences in solution. An alternative labeling strategy is to label the hybridized species after they are formed. Two advantages of this technique are that the target probes do not require labeling, and no signal is generated in the absence of target. A disadvantage is that there is only one hybridization event, and therefore one signaling event, per probe bound to the surface, which limits the sensitivity of the technique overall.
[0006] The most common method for analyzing disease-state DNA or disease-state RNA is the use of polymer chain reaction (PCR)-based assays. Biological samples will contain DNA, but usually at very low concentrations, making it difficult to detect. PCR is used to amplify DNA, or more specific regions of DNA containing sequences of interest. PCR products are analyzed by a variety of methods, including electrophoresis, where fragments are separated by size on an agarose gel and the fragment bands are visualized with ethidium bromide staining and UV light. Alternatively, real-time polymerase chain reaction (real-time PCR), also known as quantitative polymerase chain reaction (qPCR), monitors the amplification of target DNA molecules during PCR, i.e., in real time, rather than at the end point as in conventional PCR.
[0007] Sanger sequencing, also known as chain termination, is a DNA sequencing technique based on the selective incorporation of non-natural chain-terminating dideoxynucleotides (ddNTPs) by DNA polymerases during in vitro DNA replication. NGS is Sanger sequencing, but performed in parallel on several samples simultaneously.
[0008] ISH, PCR, and NGS are powerful techniques for elucidating the functions of DNA and RNA and are therefore excellent diagnostic tools. However, these techniques require meticulous sample preparation, multi-step processes, and specialized instrumentation. In contrast, microarrays offer a more simplified approach to identifying DNA / RNA indicative of disease states and are therefore excellent "point-of-care" diagnostic tools for providing rapid results and informing patient treatment decisions. However, microarrays are typically limited in their level of sensitivity because typically only one hybridization event occurs per surface-bound probe, resulting in one signaling event per hybridization reaction.
[0009] There is a need for improved methods for determining the presence of a target polynucleotide sequence in a test sample. Summary of the Invention
[0010] The present inventors have developed a novel method for the enzymatic extension of immobilized polynucleotide sequences containing tandem repeats. This method can be used to generate solid substrates, such as microarrays, bearing one or more immobilized polynucleotide sequences, each of which contains multiple tandem repeats. Advantageously, these solid substrates can be used to detect the presence of target polynucleotide sequences complementary to the repeat sequences with a higher level of sensitivity than other substrates known in the art, because each immobilized polynucleotide (also referred to herein as a "probe polynucleotide" bound to the surface) provides multiple binding sites for the target sequence.
[0011] The present invention overcomes the fundamental problem of signal-to-noise ratio that arises in DNA hybridization-based detection systems because the signal depends on surface coverage, which is no longer a limiting factor in the present invention. In other words, the present invention improves detection of the signal relative to background noise, thereby overcoming the sensitivity problems of the prior art. Advantageously, this can increase detection accuracy and / or allow for the use of less sensitive (e.g., less expensive) detectors in the detection method, as data collection is enhanced.
[0012] The present invention has a wide range of applications and can be useful in several different target polynucleotide detection technologies. For example, the present invention can be used in connection with single base mismatch detection, SNP detection, gene sequencing technology, and medical diagnostics (e.g., rapid diagnosis of colorectal cancer such as Lynch syndrome using BAT25 repeat sequence DNA probes). The present invention can also be used in connection with molecular diagnostics, such as biomarker detection, treatment response, disease stratification, and point-of-care applications.
[0013] Provided is a thermal cycling method for increasing the number of tandem repeats of a unit sequence that is 1 to 60 nucleotides in length in a linear polynucleotide, the method comprising the steps of: i) providing a solid substrate having a surface on which a single-stranded primer polynucleotide containing at least two tandem repeats of a unit sequence having a length of 1 to 60 nucleotides is immobilized; ii) contacting the immobilized primer polynucleotide with a single-stranded template polynucleotide comprising at least two tandem repeats complementary to the unit sequence of the primer polynucleotide under hybridization conditions that permit mismatch duplex formation between the unit sequence and its complementary strand, such that a 5' overhang of the template polynucleotide is generated, the 5' overhang comprising at least one tandem repeat complementary to the unit sequence of the primer polynucleotide; and iii) contacting the mismatched duplex with a thermostable 5' to 3' polymerase and nucleotides under extension conditions that permit polynucleotide extension in the 5' to 3' direction.
[0014] Preferably, the solid substrate having a surface onto which single-stranded linear primer polynucleotides comprising at least two tandem repeats of a unit sequence having a length of 1 to 60 nucleotides are immobilized is prepared by immobilizing a double-stranded linear primer polynucleotide comprising at least two tandem repeats of a unit sequence having a length of 1 to 60 nucleotides onto the surface of the solid substrate, and denaturing the double-stranded linear primer polynucleotide to obtain the single-stranded linear primer polynucleotide.
[0015] Preferably, the method further comprises the steps of: iv) denaturing the duplex of iii) under denaturing conditions to produce immobilized single-stranded polynucleotides; and v) repeating steps ii) and iii) at least once to increase the number of tandem repeats in the immobilized polynucleotide.
[0016] Preferably, the immobilized primer polynucleotide comprises at least 2, at least 5, at least 10, or at least 15 tandem repeats of the unit sequence.
[0017] Provided is a solid substrate having a surface on which at least one linear probe polynucleotide is immobilized, wherein the at least one linear probe polynucleotide comprises at least two tandem repeats of a unit sequence having a length of 1 to 60 nucleotides.
[0018] Preferably, the surface comprises the following elements: i) a plurality of spaced, discontinuous regions onto which linear probe polynucleotides are immobilized; and ii) inter-region areas between said spaced apart, discontinuous regions, said inter-region areas being substantially free of linear probe polynucleotides.
[0019] Preferably, the spaced, discontinuous regions to which the linear probe polynucleotides are immobilized form an array.
[0020] Preferably, a plurality of identical linear probe polynucleotides are immobilized within a single, spaced, discontinuous region.
[0021] Preferably, each of the plurality of spaced, discontinuous regions contains a different linear probe polynucleotide.
[0022] Preferably, the linear probe polynucleotide comprises at least three tandem repeats of the unit sequence.
[0023] Preferably, the unit sequence is a microsatellite sequence having 2 to 9 nucleotides.
[0024] Preferably, the unit sequence is a minisatellite sequence having 10 to 60 nucleotides.
[0025] Preferably, the linear polynucleotide is single-stranded or double-stranded DNA.
[0026] Preferably, the surface comprises glass, silica, gold, graphene or graphene oxide, epoxy, plastic, metal, a gel matrix, metal prepared by template stripping, or a composite thereof.
[0027] Preferably, the linear polynucleotide is immobilized on the surface by covalent or non-covalent attachment. Optionally, the linear polynucleotide is non-covalently immobilized on a chemically modified region of the surface.
[0028] Preferably, the polynucleotides are immobilized on the surface by a linker.
[0029] Preferably, the linker comprises a silane linker molecule, a biotin-streptavidin complex, a thiol-Au linker, a Si-C covalent bond to silicon, a Si-O covalent bond to silicon, a Si-N covalent bond to silicon, a nanoparticle linker, or a dynamic covalent bond.
[0030] Provided is a method for determining the presence of a linear target polynucleotide sequence in a test sample, comprising the steps of: i) providing a solid substrate as described herein, wherein the unit sequences of immobilized linear probe polynucleotides comprise nucleic acid sequences that are complementary to the sequence of a linear target polynucleotide sequence of interest; ii) contacting a test sample with the immobilized linear probe polynucleotide under conditions that allow duplex formation between the linear target polynucleotide sequence and complementary portions of the unit sequences of the immobilized linear probe polynucleotide; and iii) detecting duplex formation, which indicates the presence of the target polynucleotide sequence in the test sample.
[0031] Preferably, the test sample is a blood, saliva, cerebrospinal fluid, pleural effusion, milk, lymph, sputum, semen or needle aspirate sample.
[0032] Preferably, duplex formation is detected using fluorescent intercalators, fluorescently tagged DNA, fluorescein, redox tagged DNA, ferrocene, nanoparticles or magnetically tagged DNA.
[0033] Throughout the specification and claims of this application, the words "comprise" and "contain" and variations thereof mean "including but not limited to" and are not intended to exclude (and do not exclude) other moieties, adjuncts, components, integers or steps.
[0034] Throughout the specification and claims of this application, the singular encompasses the plural unless the context otherwise requires. In particular, where the indefinite article is used, it is to be understood that the specification contemplates plural as well as singular, unless the context otherwise requires.
[0035] It is to be understood that any feature, integer, property, compound, chemical moiety, or chemical group described in connection with a particular aspect, embodiment, or example of the invention can also be applied to any other aspect, embodiment, or example described herein, except to the extent that it is incompatible.
[0036] The patent, scientific, and technical literature referred to in this specification demonstrates the knowledge available to those skilled in the art at the time of filing. The entire disclosures of issued patents, published and pending patent applications and other publications referred to in this specification are hereby incorporated by reference as if each were specifically and individually indicated to be incorporated by reference. In the case of conflict, the present disclosure will control.
[0037] Unless otherwise defined herein, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the art to which this invention pertains. For example, Singleton and Sainsbury, Dictionary of Microbiology and Molecular Biology, 2nd Ed., John Wiley and Sons, NY (1994), and Hale and Marham, The Harper Collins Dictionary of Biology, Harper Perennial, NY (1991), provide those skilled in the art with a general dictionary of many of the terms used in this invention. Although any methods and materials similar or equivalent to those described herein are useful in practicing the present invention, preferred methods and materials are described herein. Therefore, the terms defined immediately below are more fully described by reference to the entire specification. Also, as used herein, the singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise. Unless otherwise indicated, polynucleotides are written from left to right in a 5' to 3' orientation, and amino acid sequences are written from left to right in an amino to carboxy orientation, respectively. It is to be understood that this invention is not limited to the particular methodology, protocols, and reagents described, as these may vary depending on the circumstances in which one of skill in the art would use them.
[0038] Various aspects of the invention are described in further detail below.
[0039] Hereinafter, embodiments of the present invention will be further described with reference to the accompanying drawings. [Brief explanation of the drawings]
[0040] [Figure 1]Schematic illustrating the surface functionalization steps required to produce multiple repeat sequence DNA by surface-confined PCR-based reactions: i) immobilization of bifunctional linkers, ii) attachment of 5'-amino-modified oligoseeds, iii) hybridization of the oligoseed complementary strand, and iv) PCR-based extension of the immobilized dsDNA to obtain long DNA brushes. For sensing applications, the dsDNA surface must be modified to obtain long, repetitive single-stranded DNAs carrying multiple target sequences. [Figure 2] A) Silicon wafer functionalized with oligo seeds cut to fit into a standard PCR Eppendorf tube and B) schematic of the enzymatic extension mechanism of surface-immobilized oligo seeds. [Figure 3] Fluorescence images of surfaces with Picogreen applied a) dsDNA and b) extDNA. c) Bar graph highlighting the difference in fluorescence intensity between the two surfaces. d) Fluorescence image of a patterned DNA surface where extended dsDNA is confined to circular islands. Above are data for the formation of [GATC]n sequences on a glass / siloxane / amide linker surface in a), b), and c), and for the formation of photolithographically patterned [G:C]n sequences on a glass / siloxane / streptavidin-biotin linker. [Figure 4] A) AFM image of stretched [GATC] ssDNA removed from a glass surface by dehybridization in water at 90°C. The DNA sample was combed onto a freshly cleaved mica substrate and imaged in tapping mode. B) Height profile of a single DNA strand confirming the expected dimensions of ssDNA. [Figure 5] This represents a change in the gene shape that occurs when CFTR is mutated, resulting in the deletion of the amino acid phenylalanine14. [Figure 6]Figure 1 shows 20 cycles of heat-cool cycle extension with Tgo-Pol Z3 exo- and [GCATCTTTCG(SEQ ID NO:1)]2 / [CGTAGAAAGC(SEQ ID NO:2)]2. A) Agarose gel: Lane 1 is the extension product after 20 cycles; B) UV-Vis plot; and C) ImageJ analysis of the extension products, showing band intensities as percentages of the highest intensity compared to the ladder. L = DNA ladder. [Figure 7] A) Fluorescence intensity bar graph highlighting the differences between surfaces decorated with short oligomers compared to long extDNA with respect to the CFTR sequence, and B) respective fluorescence images of the different surfaces. [Figure 8] Twenty cycles of heat-cool cycle extension with Tgo-Pol Z3 exo- and [GAAAAAAAAAAC(SEQ ID NO:3)]2 / [CTTTTTTTTTTG(SEQ ID NO:4)]2; A) Agarose gel: lane 1 is the extension product after 20 cycles; B) UV-Vis plot; C) ImageJ analysis of the extension products, where band intensity is expressed as a percentage of the highest intensity compared to the ladder. L = DNA ladder; and D) Sanger sequencing results. [Figure 9] AFM images of solution-stretched DNA, as well as the average height and length of the DNA strands, confirming the expected dimensions for dsDNA. [Figure 10] A) Fluorescence intensity bar graph highlighting the differences between surfaces modified with short oligoseeds compared to long extDNA for the BAT25 sequence, and B) respective fluorescence images of the different surfaces. [Figure 11] AFM image of long single-stranded extDNA dehybridized from the surface for the BAT25 sequence, as well as the average height and length of the single DNA strands, confirming the dimensions expected for ssDNA. [Figure 12] A) Agarose gel length comparison for 1) GATC, 2) CFTR, and 3) BAT25 and B) comparison of fluorescence intensity for surfaces modified with each sequence, highlighting the differences between short oligo seeds and extDNA. [Figure 13] Comparison of fluorescence intensity between surfaces stained with DAPI and PG. [Figure 14] A) Fluorescence intensity of extDNA and PG in TE buffer and H2O; B) Fluorescence image of exDNA in PG / TE solution; and C) Fluorescence image of extDNA in PG / H2O solution. [Figure 15] Fluorescence intensity for different PG binding times. [Figure 16] Fluorescence intensity of extDNA grown on a silicon dioxide surface. [Figure 17] Change in contact angle when BAT25 sequences are stretched on a glass surface. See Table 3 for BAT25 sequences. [Figure 18] Changes in fluorescence of the extended BAT25 surface with targets of different numbers of mismatches. This demonstrates the sensitivity of the method for detecting small numbers of mismatches in VNTR sequences. See Table 5 for mismatch sequences rehybridized with the BAT25 probe. DETAILED DESCRIPTION OF THE INVENTION
[0041] The present inventors have identified a method for enzymatically extending surface-immobilized report oligonucleotide sequences (oligoseeds). Surprisingly, the inventors have shown that immobilized oligoseeds can be extended using PCR with typical heating-cooling cycles when the oligoseeds are immobilized on a solid support. This method results in long DNA brushes immobilized directly to the surface via linker molecules (see Figure 1 for an overview). Subsequent denaturation to ssDNA and rehybridization with target complementary DNA results in increased fluorescence intensity compared to short dsDNA on the surface. Extension increases the number of target binding sites per probe molecule, thus increasing target detection.
[0042] The inventors have demonstrated that this method is highly versatile, as the surface, linker, and DNA sequence can be individually tailored to suit the desired application. As demonstrated in the Examples section herein, three different oligo seeds were successfully extended using this method, using two different linkers and two different solid support substrates. The data presented herein clearly demonstrate an increase in the fluorescent signal in extended dsDNA samples compared to shorter dsDNA strands. These data also clearly demonstrate that the described method is suitable for integration into both optical (glass) and electronic (silicon) devices.
[0043] Heat Cycle Method The present inventors have now identified a method by which immobilized linear polynucleotides containing tandem repeats of unit sequences up to 60 nucleotides in length can be generated using thermal cycling. This discovery provides a new method for generating immobilized polynucleotide probes with several binding sites for target polynucleotide sequences, and thus has great potential for improving the sensitivity of such methods, for example, in array formats.
[0044] "Thermocycling" refers to a method that involves multiple repeated cycles, each cycle including a change in temperature from a first temperature to a second (or further) temperature. A well-known example of a thermocycling method is the polymerase chain reaction (PCR).
[0045] As used herein, the terms "nucleic acid sequence," "oligonucleotide," "polynucleotide," "nucleic acid molecule," and variations thereof, are used interchangeably to refer to a plurality of nucleotides in an ordered or irregular arrangement. Polynucleotides are typically single-stranded or double-stranded (duplex), but may also adopt higher-order structures containing triple-stranded (triplexed) or quadruple-stranded (quadruplexed / i-motif) strands, or may contain combinations of these configurations at different loci under appropriate conditions. Polynucleotides can be short or long. A polynucleotide has at least two contiguous nucleotides.
[0046] Nucleotide sequences can be of genomic, synthetic, or recombinant origin and can be double-stranded or single-stranded (corresponding to the sense or antisense strand). The term "nucleotide sequence" encompasses genomic DNA, cDNA, synthetic DNA, and RNA (e.g., mRNA), as well as analogs of DNA or RNA generated, for example, using nucleotide analogs. In other words, modified DNA or RNA bases are also encompassed. Thus, a polynucleotide can contain one or more modified DNA or RNA bases. Polynucleotides carrying multiple modifications at specific sites have applications in synthetic biology, nanomaterial fabrication, bioanalytical applications, and sequencing applications. For example, DNA can be chemically modified at any or all of its three components—phosphate linkage, sugar ring, and nucleobase. A variety of modified nucleotides are commercially available as deoxynucleotide triphosphates (dNTPs) or phosphoramidite derivatives. These and other modified nucleotides can be synthesized as dNTPs and inserted enzymatically into DNA or RNA, or as phosphoramidites and inserted into DNA or RNA by automated DNA synthesis.
[0047] Nucleotide residues are typically derived from the naturally occurring purine bases adenine (A), guanine (G), hypoxanthine (I), and xanthine (X), and the pyrimidine bases cytosine (C), thymine (T), and uracil (U). Nucleotide analogs can be used at one or more positions within a polynucleotide sequence, and such nucleotide analogs are modified, for example, at the base moiety and / or sugar moiety and / or phosphate linkage. Any nucleotide analog can be used as long as it does not interfere with polynucleotide hybridization and is accepted by the polymerase as both a template and a substrate.
[0048] Nucleic acid sequences presented herein are conventionally written 5' to 3' (left to right). A "linear polynucleotide" refers to a polynucleotide that is neither branched nor circularized (i.e., the 3' end is not circularized with the 5' end).
[0049] In one example, polynucleotide is DNA. In another example, polynucleotide is RNA. DNA or DNA can contain natural bases or modified bases, including combinations thereof. Some modified bases are known in the art, so those skilled in the art can easily identify suitable modified bases.
[0050] A thermal cycling method is provided for increasing the number of tandem repeats of a unit sequence that is 1 to 60 nucleotides in length in a linear polynucleotide.
[0051] At the start of the thermal cycling method, the linear primer polynucleotide contains at least two copies of a unit sequence. The unit sequence is 1 to 60 nucleotides in length. The linear polypeptide may be 2 or more, 3 or more, 4 or more, 5 or more, 6 or more, 7 or more, 8 or more, 9 or more, 10 or more, 11 or more, 12 or more, 13 or more, 14 or more, 15 or more, etc. sameThe repeating unit sequences may comprise a unit sequence (i.e., a linear polypeptide may comprise multiple repeats of the same unit sequence). This may include at least 20 copies, at least 25 copies, at least 30 copies, at least 35 copies, at least 40 copies, at least 45 copies, at least 50 copies, at least 55 copies, at least 60 copies, at least 65 copies, at least 70 copies, at least 75 copies, at least 80 copies, at least 85 copies, at least 90 copies, at least 95 copies, or at least 100 copies of the unit sequence. The repeated unit sequences may be in tandem (i.e., they may be referred to as "tandem repeats"). Tandem repeats occur in a polynucleotide sequence when a pattern of nucleotides (in this case, the unit sequence) is repeated and the repeats are immediately adjacent to each other. As an example, if the nucleotide sequence is ATTCG, a polynucleotide containing two tandem repeats of the nucleotide sequence would contain the sequence ATTCGATTCG (SEQ ID NO: 5), a polynucleotide containing three tandem repeats of the nucleotide sequence would contain the sequence ATTCGATTCGATTCG (SEQ ID NO: 6), a polynucleotide containing four tandem repeats of the nucleotide sequence would contain the sequence ATTCGATTCGATTCGATTCG (SEQ ID NO: 7), etc. The number of tandem repeats can also be referred to as the "copy number" of the nucleotide sequence.
[0052] The nucleotide sequence may have any permutation of bases. Non-limiting examples of common nucleotide sequences that may be tandemly repeated in a linear polynucleotide include the following: (AT)n, (GC)n, (GGC)n, (CAG)n, (GCAT)n, (GATC)n, (AAAG)n, (AAAAAAAAG)n (SEQ ID NO: 8), (ACTGATCAGC)n (SEQ ID NO: 9), where (xxxx) represents the nucleotide sequence and n represents the number of tandem repeats (i.e., n = nucleotide sequence copy number).
[0053] The unit sequence is between 1 and 60 nucleotides in length. Thus, the unit sequence may contain at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, or 58 nucleotides, or at least 59 nucleotides (in each case, the upper limit is 60 nucleotides).
[0054] In one example, the unit sequence is a microsatellite sequence having 2 to 9 nucleotides. Alternatively, the unit sequence can be a minisatellite sequence having 10 to 60 nucleotides.
[0055] The method increases the number of tandem repeats of a unit sequence in a polynucleotide. In other words, if the starting polynucleotide (i.e., the initial immobilized primer polynucleotide) has two tandem repeats of a unit sequence (called ATTCG) (i.e., ATTCGATTCG (SEQ ID NO: 5)), the method will increase this to at least three tandem repeats (i.e., ATTCGATTCGATTCG (SEQ ID NO: 6)). Similarly, if the starting polynucleotide has three tandem repeats of the unit sequence (i.e., ATTCGATTCGATTCG (SEQ ID NO: 6)), the method will increase this to at least four tandem repeats (i.e., ATTCGATTCGATTCGATTCG (SEQ ID NO: 7)), and so on.
[0056] Tandem repeats occur naturally in genomic DNA. They are termed "minisatellites" for repeats of a single sequence between 10 and 60 nucleotides in length, and "microsatellites" for repeats of a single sequence between 2 and 9 nucleotides in length. The number of tandem repeats is referred to as the copy number. Variable-number tandem repeats (VNTRs) are tandem repeats whose copy number varies between individuals. VNTR analysis in DNA fingerprinting is invaluable in modern forensic science and identity verification, as well as species typing of pathogens, fungi, and plants. Abnormalities in trinucleotide repeats are associated with genetic disorders, including Huntington's disease (CAG), Friedreich's ataxia (GAA), myotonic dystrophy (CTG), and fragile X syndrome (CGG).
[0057] The method comprises the steps of: i) preparing a solid substrate having a surface on which a single-stranded primer polynucleotide containing at least two tandem repeats of a unit sequence having a length of 1 to 60 nucleotides is immobilized; "Solid substrate" refers to a material or group of materials having one or more rigid or semi-rigid surfaces. In many embodiments, at least one surface of the solid substrate will be substantially flat (or planar). In other embodiments, it may be desirable to physically separate the regions onto which different polynucleotides are immobilized, for example, using wells, raised regions, etched grooves, or combinations thereof. In other embodiments, the solid substrate can take the form of beads, resins, gels, microspheres, or other geometric configurations. Thus, the substrate can include a semi-solid substrate (e.g., a gel or other matrix) and / or a porous substrate (e.g., a nylon or other membrane).
[0058] The surface of the solid substrate on which the single-stranded primer polynucleotides are immobilized can be composed of any suitable material, including, but not limited to, glass, polyacrylamide-coated glass, epoxy, ceramics, fused silica, silicon, quartz, various plastics, metals such as gold or silver (e.g., metals such as gold or silver prepared by template stripping), nylon, a gel matrix, graphene, or graphene oxide. Combinations or composites of these materials are also contemplated.
[0059] A single-stranded primer polynucleotide containing at least two tandem repeats of a unit sequence, 1 to 60 nucleotides in length, is immobilized on the surface of a solid substrate. The term "primer polynucleotide" refers to the polynucleotide's function in the thermal cycling method as a primer for the polymerase chain reaction. The "primer polynucleotide" is single-stranded and can thus hybridize with at least a partially complementary polynucleotide sequence, thereby initiating elongation of the immobilized primer polynucleotide sequence and increasing the number of tandem repeats in the immobilized sequence.
[0060] The single-stranded primer polynucleotide is immobilized on the surface of the solid substrate using any suitable immobilization means. A linker (or any other means) may be used to immobilize the polynucleotide on the surface. The polynucleotide may be immobilized on the surface using an appropriate surface chemistry (e.g., a surface chemistry compatible with either inkjet printing or inkjet spotting techniques).
[0061] In one example, linear polynucleotides are immobilized to a surface by covalent or non-covalent bonds. Linear polynucleotides can be non-covalently immobilized to chemically modified regions of a surface.
[0062] Suitable linker / surface (or surface chemistry) combinations are well known in the art. For example, a linear polynucleotide may contain one or more nucleotide analogs modified with a functional group that can be used to immobilize the polynucleotide to the surface of a solid substrate, and the functional group may be attached to the nucleotide by a flexible or rigid linker. Exemplary functional groups include, but are not limited to, amines, which covalently react with succinimidyl ester-modified labels; azides, which covalently react with alkyne-modified labels; alkynes, which covalently react with azide-modified labels; digoxigenin, which forms strong noncovalent interactions with anti-digoxigenin antibodies; or biotin, which forms strong noncovalent interactions with avidin or streptavidin labeled with reporter groups such as fluorescent dyes, either noncovalently with biotin-conjugated labels or covalently attached to proteins. Specific examples include 5-(3-aminoallyl)-uracil, 5-aminoallylcytosine, 5-aminoallyluracil, 7-deaza-7-propargylaminoadenine, 7-deaza-7-propargylaminoguanine, 5-propargylaminocytosine, 5-propargylaminouracil, 8-[(6-amino)hexyl-biotin]-aminoadenosine, γ-[N-(biotin-6-amino-hexanoyl)]-7-propargylamino-7-deazaadenine, γ-[N-(biotin-6-aminohexanoyl)]-5-aminoallyl-uracil, γ-[N-(biotin-6-amino-hexanoyl-6-amino hexanoyl)]-5-(3-aminoallyl)-uracil, digoxigenin-X-5-aminoallyl-uracil, 5-(3-azidopropyl)-uracil, 5-azido-PEG4-uracil, 5-azido-PEG4-cytosine, 5-(octa-1,7-diynyl)-uracil, 5-(octa-1,7-diynyl)-cytosine (5-C8-alkyne-C), 5-dibenzylcyclooctyl-PEG4-uracil, 5-dibenzylcyclooctyl-PEG4-cytosine, 5-trans-cyclooctene-PEG4-uracil, or any combination thereof.Suitable linkers include silane linker molecules, biotin-streptavidin complexes, thiol-Au linkers, Si-C covalent bonds to silicon, Si-O covalent bonds to silicon, Si-N covalent bonds to silicon, nanoparticle linkers, or dynamic covalent bonds.
[0063] The functional group / linker is typically attached to the 5' end of the single-stranded linear polynucleotide (or at least attached to the linear polynucleotide in a manner that allows for immobilization of the 5' end of the linear polynucleotide to the surface of a solid substrate).
[0064] The tandem repeat is typically located at the 3' end of the immobilized single-stranded primer polynucleotide. In some instances, the tandem repeat may constitute the last nucleotide (in the 5' to 3' direction) of the linear polynucleotide.
[0065] The method further comprises the steps of: i) contacting the immobilized primer polynucleotide with a single-stranded template polynucleotide comprising at least two tandem repeats complementary to the unit sequence of the primer polynucleotide under hybridization conditions that permit mismatch duplex formation between the unit sequence and its complementary strand, such that a 5' overhang of the template polynucleotide is generated, the 5' overhang comprising at least one tandem repeat complementary to the unit sequence of the primer polynucleotide.
[0066] This step is also referred to herein as the "annealing step," the "hybridization step," or variations thereof.
[0067] In this context, a "single-stranded template polynucleotide comprising at least two tandem repeats complementary to the unit sequence of a primer polynucleotide" can also be referred to as a "template polynucleotide" (or "template").
[0068] The template polynucleotide contains at least two tandem repeats that are complementary to the unit sequence of the primer polynucleotide. The term "complementary" is used in its usual context in the art. As an example, if the unit sequence of the primer polynucleotide is 5' ATCG 3', the tandem repeat of the template polynucleotide will be 5' CGAT 3' (i.e., the template polynucleotide will contain at least two tandem repeats, and therefore will contain the sequence 5' CGATCGAT 3'). In other words, in this example, the tandem repeats in the template and primer are 100% complementary.
[0069] Where the template polynucleotide has three or more tandem repeats, the invention encompasses template polynucleotide sequences in which at least two of the 3'-end tandem repeats are 100% complementary to the corresponding tandem repeats in the primer polynucleotide (and additional tandem repeats are 100% or less complementary to the corresponding tandem repeats in the primer polynucleotide). In other words, as long as the 3'-end of the immobilized primer strand has at least two repeats that are complementary to the template strand, the remainder of the template strand need not be 100% complementary.
[0070] The immobilized primer polynucleotide is contacted with the template polynucleotide under hybridization conditions that allow mismatch duplex formation between the unit sequence and its complementary strand. In this context, "contacting" refers to direct contact between the primer polynucleotide and the template polynucleotide, for example, in a suitable buffer and container for the subsequent thermal cycling step of the method. Suitable buffers and containers are well known, for example, PCR buffer and PCR Eppendorf tubes.
[0071] When a primer polynucleotide is contacted with a template polynucleotide under appropriate hybridization conditions, the complementary sequences (i.e., the unit sequence(s) of the primer polynucleotide and the complementary tandem repeat(s) of the template polynucleotide) hybridize to form a duplex. Because the primer polynucleotide and the template polynucleotide each have at least two tandem repeats, and the primer repeats and template repeats are complementary to each other, hybridization between the template polynucleotide and the primer polynucleotide results in a misaligned alignment of the resulting polynucleotide duplex in a certain percentage of reactions, where not all of the complementary sequences are aligned (see Figure 2). In other words, mismatch duplex formation occurs (where "mismatch" refers to less than perfect alignment between all of the complementary sequences of the unit sequences and tandem repeats in the primer polynucleotide and the template polynucleotide, respectively). As shown in Figure 2, a 5' overhang of the template polynucleotide can be generated that contains at least one tandem repeat that is complementary to the unit sequence of the primer polynucleotide, which can then serve as a template in a thermal cycling method to extend the immobilized primer polynucleotide, thereby increasing the number of tandem repeats in the immobilized linear polynucleotide.
[0072] As used herein, "hybridization conditions" refers to the reagents and reaction conditions (e.g., temperature, time, etc.) used. This describes the conditions for hybridization and washing. Typically, hybridization conditions can be stringent or moderate. Hybridization conditions used in connection with the methods described herein allow for mismatch duplex formation and therefore can be moderate or stringent. Preferably, hybridization between the unit sequence of the immobilized polynucleotide and the complementary sequence of the template oligonucleotide will form a stable duplex at 65°C or below. Preferably, mismatch duplexes can be formed at temperatures up to 65°C, e.g., 55°C to 65°C, optionally for a period of 1 to 30 seconds.
[0073] Moderate and stringent conditions are known to those skilled in the art and can be found in available references (e.g., Current Protocols in Molecular Biology, John Wiley & Sons, NY, 1989, 6.3.1-6.3.6). Aqueous and non-aqueous methods are described in the above references, and either method can be used. One preferred example of stringent hybridization conditions is hybridization in 6× sodium chloride / sodium citrate (SSC) at about 45°C, followed by one or more washes in 0.2× SSC, 0.1% (w / v) SDS at 50°C. Another preferred example of stringent hybridization conditions is hybridization in 6× SSC at about 45°C, followed by one or more washes in 0.2× SSC, 0.1% (w / v) SDS at 55°C. Yet another preferred example of stringent hybridization conditions is hybridization in 6×SSC at about 45° C., followed by one or more washes in 0.2×SSC, 0.1% (w / v) SDS at 60° C. Preferably, stringent hybridization conditions are hybridization in 6×SSC at about 45° C., followed by one or more washes in 0.2×SSC, 0.1% (w / v) SDS at 65° C. Particularly preferred stringency conditions (and conditions to be used if the practitioner is unsure what conditions to apply to determine whether a molecule is within the hybridization limits of the invention) are 0.5 molar sodium phosphate, 7% (w / v) SDS at 65° C., followed by one or more washes in 0.2×SSC, 1% (w / v) SDS at 65° C.
[0074] The method further comprises the steps of: iii) contacting the mismatched duplex with a thermostable 5' to 3' polymerase and nucleotides under extension conditions that permit polynucleotide extension in the 5' to 3' direction.
[0075] This step is also referred to herein as the "extension step" or variations thereof.
[0076] The term "contacting" has been defined above and applies equally in this context. It therefore refers to direct contact between the mismatched duplex and the thermostable 5' to 3' polymerase (and nucleotides), e.g., in a suitable buffer and container for the subsequent thermocycling step of the method. Suitable buffers and containers are well known and include, for example, PCR buffers and PCR Eppendorf tubes.
[0077] There are several well-known thermostable 5' to 3' polymerases available that may be used in the methods described herein. The polymerase is preferably thermostable and highly stable, so that its activity is substantially maintained during the long incubation times required for the extension reaction. The polymerase preferably has high processivity. The polymerase preferably does not exhibit nonspecific nuclease activity. The polymerase preferably has good fidelity but can also tolerate a range of nucleotide analogs as both templates and substrates. Therefore, high-fidelity polymerases with efficient proofreading activity are inappropriate. Preferably, the polymerase lacks 3'→5' exonuclease activity [3'→5' exo(-)], and therefore has low fidelity due to the lack of proofreading function. Those skilled in the art can determine whether a particular polymerase has the required properties defined above. Exemplary polymerases include, but are not limited to, Tgo-Pol Z3 exo(-) [Jozwiakowski et al., 2011 Chembiochem 12:35-37], Deep Vent exo(-) (New England Biolabs), Vent exo(-) (New England Biolabs), Pfu exo(-) (Agilent Technologies), and Taq polymerase (many suppliers). In one example, the polymerase of choice is the Thermococcus gorgonarius family B polymerase (Tgo-Pol) enzyme variant Z3.
[0078] Suitable nucleotides for thermocycling reactions are well known in the art.
[0079] Contact between the mismatched duplex, polymerase, and nucleotides occurs under extension conditions that allow polynucleotide extension in the 5' to 3' direction. As used herein, "extension conditions" refers to the reagents and reaction conditions (e.g., temperature, time, etc.) used. This describes the conditions for extending a primer polynucleotide. Suitable extension conditions are well known in the art. Preferably, extension is carried out at a temperature of about 65°C to 75°C, optionally for a period of 30 to 120 seconds. Suitable conditions can be found, for example, in Whitfield CJ, Turley AT, Tuite EM, Connolly BA, Pike AR. "Enzymatic Method for the Synthesis of Long DNA Sequences with Multiple Repeat Units," Angewandte Chemie International Edition 2015, 54(31), 8971-8974.
[0080] Steps i) through iii) of this method can be repeated at least once. To repeat these steps, the extended duplex produced by step iii) is denatured to produce a single, extended, immobilized polynucleotide. The single-stranded polynucleotide can then serve as a primer polynucleotide in repeated cycles of steps i) through iii). Polynucleotides of different lengths can be obtained by varying the number of repeated cycles. The average polynucleotide length is controlled by the temperature, the length of time per step, and the number of cycles.
[0081] Thus, the extended duplex produced by step iii) may be subjected to denaturing conditions that allow for dissociation of the duplex into single-stranded polynucleotides. This is also referred to herein as the "melting step." Suitable denaturing conditions are well known in the art and include subjecting the extended duplex to a temperature of about 75-100°C, preferably about 90-98°C, optionally for about 15-30 seconds.
[0082] Therefore, the method may further comprise the steps of: iv) denaturing the duplex of iii) under denaturing conditions to produce immobilized single-stranded polynucleotides; and v) repeating steps ii) and iii) at least once to increase the number of tandem repeats in the immobilized polynucleotide.
[0083] The method described above provides a starting material, a solid substrate having a surface on which single-stranded linear primer polynucleotides containing at least two tandem repeats of a unit sequence having a length of 1 to 60 nucleotides are immobilized. The starting material contains at least two tandem repeats of a unit sequence having a length of 1 to 60 nucleotides. Two The double-stranded linear primer polynucleotide can be derived from a solid substrate having a surface onto which the stranded linear primer polynucleotide is immobilized, and the double-stranded linear primer polynucleotide is denatured to produce the immobilized single-stranded linear primer polynucleotide starting material for the method.
[0084] Suitable denaturing conditions are described elsewhere herein.
[0085] Therefore, a solid substrate having a surface on which single-stranded linear primer polynucleotides containing at least two tandem repeats of a unit sequence having a length of 1 to 60 nucleotides are immobilized can be provided by the following steps: a) immobilizing a double-stranded linear primer polynucleotide comprising at least two tandem repeats of a unit sequence having a length of 1 to 60 nucleotides on the surface of a solid substrate; and b) denaturing the double-stranded linear primer polynucleotide to obtain the single-stranded linear primer polynucleotide.
[0086] A double-stranded linear primer polynucleotide containing at least two tandem repeats of a unit sequence having a length of 1 to 60 nucleotides is also referred to herein as an "oligoseed."
[0087] The immobilized primer polynucleotide of step i) comprises at least two tandem repeats of the unit sequence. As noted above, the polynucleotide can have from at least two to at least 100 tandem repeats of the unit sequence. In one example, the immobilized primer polynucleotide comprises at least two, at least five, at least 10, or at least 15 tandem repeats of the unit sequence.
[0088] solid substrate Also provided is a solid substrate having a surface having at least one linear probe polynucleotide immobilized thereon, wherein the at least one linear probe polynucleotide comprises at least two tandem repeats of a unit sequence that is 1 to 60 nucleotides in length.
[0089] The solid substrate can be generated by the thermal cycling method described above, in which an initial immobilized linear primer polynucleotide is extended to increase the number of tandem repeats of its unit sequence, and the extended linear polynucleotide corresponds to a linear probe polynucleotide immobilized on the solid substrate (the extended linear polynucleotide is referred to herein as a linear probe polynucleotide immobilized on the solid substrate). Alternatively, the solid substrate can be obtained by generating a linear probe polynucleotide (e.g., in solution) and then immobilizing the linear probe polynucleotide on the surface of the solid substrate.
[0090] The definitions provided for linear primer polynucleotides apply equally to linear probe polynucleotides unless the context specifically states otherwise. Similarly, the definitions provided for solid substrates (or their surfaces) also apply to the methods described herein (which require the presence of a solid substrate) unless the context specifically states otherwise.
[0091] The linear probe polynucleotides may be 2 or more, 3 or more, 4 or more, 5 or more, 6 or more, 7 or more, 8 or more, 9 or more, 10 or more, 11 or more, 12 or more, 13 or more, 14 or more, 15 or more, etc. same It may comprise an identity sequence (i.e., a linear polynucleotide may comprise several repeats of the same identity sequence), which may include at least 20 copies, at least 25 copies, at least 30 copies, at least 35 copies, at least 40 copies, at least 45 copies, at least 50 copies, at least 55 copies, at least 60 copies, at least 65 copies, at least 70 copies, at least 75 copies, at least 80 copies, at least 85 copies, at least 90 copies, at least 95 copies, at least 100 copies, at least 110 copies, at least 120 copies, at least 140 copies, at least 150 copies, at least 160 copies, at least 170 copies, at least 180 copies, at least 190 copies, at least 200 copies, or at least 250 copies of the identity sequence.
[0092] The surface of the solid substrate of the present invention may comprise the following elements: i) a plurality of spaced, discontinuous regions onto which linear probe (or primer) polynucleotides are immobilized; and ii) an inter-region area between said spaced apart, discontinuous regions, said inter-region area being substantially free of linear probe (or primer) polynucleotides;
[0093] Linear (probe or primer) polynucleotides can be immobilized on a substrate surface in regions referred to herein as "spaced, discrete regions." This term is used to refer to regions of a surface that are distinct from other, spatially separated regions of the surface to which different polynucleotides (or copies of the same polynucleotide) may be immobilized. Thus, a surface can include multiple spaced, discrete regions, each of which is spatially separated such that each spaced region is optically separable (i.e., optically resolvable) from adjacent "spaced, discrete regions" (such that any optical signal generated from a region is optically distinct or distinguishable from its adjacent regions). Spaced, discrete regions are separated by providing inter-region zones between the spaced, discrete regions that are substantially free of linear probe polynucleotides. Such "inter-region zones" are typically inert, in the sense that linear polynucleotides (or other polymeric structures) described herein do not bind to such regions. In some instances, such inter-region areas may be treated with blocking agents, such as other polymers, oxides, "chemically unreactive / incompatible sites," and the like.
[0094] The distinction between "separate, discrete regions" and "inter-region regions" can be determined by the surface chemistry of these regions (where the surface chemistry of the separate, discrete regions allows for suitable linear polynucleotide immobilization, while the surface chemistry of the inter-region regions does not). Alternatively, the surface chemistry may be the same for both the separate, discrete regions and the inter-region regions, and the distinction between them may be determined by the placement (immobilization) of linear polynucleotides in certain regions of the surface (thereby generating separate, discrete regions), in which case the regions where no linear polynucleotides are placed (immobilized) become "inter-region regions."
[0095] Each of the spaced, discrete regions may have a predetermined location on the surface of the solid substrate. The required spacing between each of the spaced, discrete regions will depend on the method and device used to optically resolve or measure any (direct or indirect) signals generated from the immobilized polynucleotides. It may have a size that allows for the immobilization of only one linear polynucleotide described herein. Alternatively, multiple (e.g., at least 2, at least 5, at least 10, at least 20) identical linear probe (or primer) polynucleotides may be immobilized within a single spaced, discrete region.
[0096] Methods for determining the appropriate spatial separation of such regions and for determining the size of such regions are well known in the art.
[0097] The spaced apart, discontinuous regions can be arranged on the surface in virtually any pattern, i.e., any regular array, in which the regions have predetermined locations and improve the efficiency of signal collection and data analysis functions. Such patterns include, but are not limited to, concentric circular regions, spiral patterns, rectilinear patterns, hexagonal patterns, etc. Preferably, the regions are arranged in a rectilinear or hexagonal pattern.
[0098] Thus, the spaced, discrete areas to which the linear probe polynucleotides are immobilized can form an array, eg, a microarray.
[0099] The surface of the solid substrate may include a plurality of spaced, discrete regions, each of the spaced, discrete regions containing a different linear probe polynucleotide, or in other words, there may be at least two different linear probe polynucleotides (immobilized in different spaced, discrete regions) on the surface of the solid substrate.
[0100] As used herein, "plurality" refers to more than one, i.e., 2 or more, 3 or more, 4 or more, 5 or more, 6 or more, 7 or more, 8 or more, 9 or more, 10 or more, 11 or more, 12 or more, 13 or more, 14 or more, 15 or more, etc. This includes at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, at least 70, at least 75, at least 80, at least 85, at least 90, at least 95, or at least 100, at least 200, at least 300, at least 400, at least 500, at least 1000, etc.
[0101] The number of tandem repeats (and the size and sequence of the "unit sequence") in a linear polynucleotide (e.g., a linear probe polynucleotide or a linear primer polynucleotide) has been discussed in detail elsewhere herein. As noted elsewhere herein, the linear polynucleotides described herein can be single-stranded or double-stranded DNA. Alternatively, they can be RNA or cDNA.
[0102] The unit sequences of the linear polynucleotides described herein (whether primer polynucleotides or probe polynucleotides) can comprise any desired sequence, for example, any nucleotide sequence relevant to diagnosis.
[0103] Relevant non-limiting examples include sequences that are diagnostic of disease, e.g., microsatellite instability (MSI) such as BAT25, and single or multiple base mutations that occur in the CFTR gene, e.g., in cystic fibrosis. Thus, relevant unit sequences include 5' GCATCTTTCG 3' (SEQ ID NO: 1) (derived from the CFTR gene in cystic fibrosis; resulting from a three-base frameshift mutation in the CFTR gene), 5' AGA TAC ATT GAC CTT' 3 (SEQ ID NO: 10) (derived from the CYP450 liver enzyme CYP29C, which affects the metabolism of warfarin), and 5' GCATCTTTCG 3' (SEQ ID NO: 1) (derived from the CFTR gene in cystic fibrosis; resulting from a three-base frameshift mutation in the CFTR gene). * 2 / *3; corresponding to a single base mutation), or 5' GAG GAC CGT GTT CAA'3 (SEQ ID NO: 11), or many other sequences (including sequences that are the complements of the sequences set forth above). One of skill in the art can readily identify a sequence of interest, and the gene (mutation) of interest (or its complement) is contained within that sequence. Examples of suitable sequences for warfarin genetic analysis are Johnson J, Caudle K, Gong L, Whirl-Carrillo M, Stein C, Scott S, et al., "Clinical Pharmacogenetics Implementation Consortium (CPIC) Guideline for Pharmacogenetics-Guided Warfarin Dosing: 2017 Update," Clin Pharmacol Ther. 2017 Sep 1;102(3):397-404; Stubbins MJ, Harries LW, Smith G, Tarbit MH, Wolf CR. "Genetic analysis of the human cytochrome P450 CYP2C9 locus," Pharmacogenetics. 1996;6(5):429-39; Rettie AE, Wienkers LC, Gonzalez FJ, Trager WF, Korzekwa KR. "Impaired (S)-warfarin metabolism catalyzed by the R144C allelic variant of CYP2C9” Pharmacogenetics.1994 Feb;4(1):39-42;Steward DJ, Haining RL, Henne KR, Davis G, Rushmore TH, Trager WF, et al. “Genetic association between sensitivity to warfarin and expression of CYP2C9 *3," Vol. 7, Pharmacogenetics. 1997. p. 361-7; and Lee CR, Goldstein JA, Pieper JA. "Cytochrome P450 2C9 polymorphisms: a comprehensive review of the in-vitro and human data," Pharmacogenetics. 2002 Apr; 12(3): 251-63.
[0104] Method for determining the presence of a linear target sequence in a test sample The solid substrates described herein can be used to determine whether a linear target sequence of interest is present in a test sample. The solid substrate has a surface on which one or more linear probe polynucleotides are immobilized. The linear probe polynucleotides comprise at least two tandem repeats of a unit sequence. Each unit sequence can act as a probe (i.e., a binding site) for a complementary linear target sequence of interest if at least a portion of each unit sequence comprises a nucleic acid sequence complementary to the linear target sequence of interest. Therefore, the solid substrates described herein provide a detection technique with improved sensitivity for detecting target polynucleotide sequences in a test sample, since each immobilized probe polynucleotide comprises several target binding sites.
[0105] Therefore, there is provided a method for determining the presence of a linear target polynucleotide sequence in a test sample, the method comprising the steps of: i) providing a solid substrate as described elsewhere herein, wherein the unit sequences of the immobilized linear probe polynucleotides comprise nucleic acid sequences that are complementary to the sequence of the linear target polynucleotide sequence of interest; ii) contacting a test sample with the immobilized linear probe polynucleotide under conditions that allow duplex formation between the linear target polynucleotide sequence and complementary portions of the unit sequences of the immobilized linear probe polynucleotide; and iii) detecting duplex formation, which indicates the presence of the target polynucleotide sequence in the test sample.
[0106] Linear target polynucleotide can be any polynucleotide sequence of interest. Linear target polynucleotide sequence must be able to hybridize with the complementary part of the unit sequence of linear probe polynucleotide, therefore, suitable linear probe polynucleotide must be immobilized on a solid substrate. For example, if linear target polynucleotide has the sequence 5'ATCGAA 3', the unit sequence of linear probe polynucleotide should contain the sequence 5'TTCGAT 3'. Those skilled in the art can easily identify suitable unit sequences for linear target polynucleotide.
[0107] The unit sequence of the immobilized linear probe polynucleotide therefore comprises a nucleic acid sequence that is complementary to the sequence of the linear target polynucleotide sequence of interest. In this example, the unit sequence may also comprise other (additional) nucleic acids that are not complementary to the sequence of the linear target polynucleotide of interest. In this example, the additional nucleic acids in the unit sequence may act as "spacers" between target binding sites in the immobilized linear probe polynucleotide.
[0108] In another example, the unit sequence consists of a nucleic acid sequence that is complementary to the sequence of the linear target polynucleotide sequence of interest (i.e., the unit sequence does not include "additional nucleic acids," in contrast to the example above).
[0109] Linear target polynucleotide can be a part of a longer polynucleotide molecule in test sample.Therefore, " linear target polynucleotide " as used herein does not limit the total length (or sequence) of the polynucleotide that forms a duplex with immobilized linear probe polynucleotide, but only refers to the sequence that has the ability to hybridize with the corresponding sequence in the unit sequence of immobilized linear probe polynucleotide (therefore, the sequence in test sample that is of interest and / or provides information (for example, diagnostic / prognostic)).Therefore, it can be a part of a longer sequence in test sample.
[0110] A test sample can be any suitable sample that can contain a linear target polynucleotide of interest. The term "test sample" generally refers to a quantity of material from a biological, environmental, medical, or patient source, in which detection or measurement of a linear target polynucleotide of interest is desired. On the one hand, it encompasses specimens or cultures (e.g., microbial cultures). On the other hand, it encompasses both biological and environmental samples. A sample can also include specimens of synthetic origin. Biological samples can be bodily fluids, solids (e.g., feces), or tissues of animals, including humans, as well as liquid and solid food and feed products and ingredients, such as dairy products, vegetables, meat and meat by-products, and waste products. Biological samples can include materials collected from patients, including, but not limited to, cultures, blood, saliva, cerebrospinal fluid, pleural effusion, milk, lymph, sputum, semen, needle aspirates, and the like. Biological samples can be obtained from various families of domestic animals, as well as feral or wild animals, including, but not limited to, ungulates, bears, fish, rodents, etc. Environmental samples include environmental materials such as surface matter, soil, water, and industrial samples, as well as samples obtained from food and dairy processing equipment, devices, facilities, tools, disposable and non-disposable items. These examples should not be considered limiting of the sample types applicable to the present invention.
[0111] Several standard methods for detecting duplex formation are known. These include the use of fluorescent intercalators such as Picogreen, DAPI, or Sybergreen, fluorescently tagged DNA, fluorescein, redox-tagged DNA, ferrocene, nanoparticles, or magnetically tagged DNA. Such standard methods are discussed, for example, in "Comparison of DNA detection methods using nanoparticles and silver enhancement" by B. Foultier; L. Moreno-Hagelsieb; D. Flandre; and J. Remacle, Volume 152, Issue 1, IEE Proceedings-Nanobiotechnology; February 2005, pp. 3-12; and in the review article "DNA Biosensors" by Kavita V, J Bioengineer & Biomedical Sci 2017, 7:2; or Mikkelsen, SR (1996) "Electrochemical biosensors for DNA sequence detection" Electroanalysis, 8:15-19.
[0112] Detection of duplex formation indicates the presence of the target polynucleotide sequence in the test sample.
[0113] Method for determining the presence of a target nucleotide-binding molecule in a test sample The solid substrates described herein can also be used to determine whether a target nucleotide-binding molecule of interest is present in a test sample. The solid substrate has a surface on which one or more linear probe polynucleotides are immobilized. The linear probe polynucleotides comprise at least two tandem repeats of a unit sequence. Each unit sequence can act as a probe (i.e., a binding site) for a target nucleotide-binding molecule of interest if at least a portion of each unit sequence comprises a nucleic acid sequence that acts as a binding sequence for the target nucleotide-binding molecule of interest. The target nucleotide-binding molecule of interest can be any molecule (e.g., a protein) that binds to a specific nucleotide sequence (which can be represented by a unit sequence described herein). Therefore, the solid substrates described herein provide a detection technique with improved sensitivity for detecting a nucleotide-binding molecule of interest in a test sample, since each immobilized probe polynucleotide contains several target binding sites.
[0114] Therefore, there is provided a method for determining the presence of a nucleotide binding molecule in a test sample, the method comprising the steps of: i) providing a solid substrate as described elsewhere herein, wherein the unit sequences of the immobilized linear probe polynucleotides comprise nucleic acid sequences that are binding sites for a nucleotide-binding molecule of interest; ii) contacting a test sample with the immobilized linear probe polynucleotide under conditions that allow binding between the nucleotide-binding molecule and a portion of the unit sequence of the immobilized linear probe polynucleotide that acts as a binding site for the nucleotide-binding molecule; and iii) detecting binding between the immobilized linear probe polynucleotide and the nucleotide-binding molecule, wherein binding indicates the presence of said nucleotide-binding molecule in the test sample.
[0115] The nucleotide-binding molecule can be any molecule of interest that can bind to the appropriate portion of the unit sequence of the linear probe polynucleotide. Non-limiting examples include transcription factors, DNA repair proteins, or histones. The appropriate linear probe polynucleotide must be immobilized on a solid substrate. For example, if the nucleotide-binding molecule binds to the sequence 5' ATCGAA 3', the unit sequence of the linear probe polynucleotide should contain 5' ATCGAA 3'. Suitable unit sequences can be easily identified by those skilled in the art.
[0116] A test sample can be any suitable sample that can contain a nucleotide-binding molecule of interest. The term "test sample" generally refers to a quantity of material from a biological, environmental, medical, or patient source, in which detection or measurement of a nucleotide-binding molecule of interest is desired. On the one hand, it encompasses specimens or cultures (e.g., microbial cultures). On the other hand, it encompasses both biological and environmental samples. A sample can also include specimens of synthetic origin. Biological samples can be bodily fluids, solids (e.g., feces), or tissues of animals, including humans, as well as liquid and solid food and feed products and ingredients, such as dairy products, vegetables, meat and meat by-products, and waste products. Biological samples can include materials collected from patients, including, but not limited to, cultures, blood, saliva, cerebrospinal fluid, pleural effusion, milk, lymph, sputum, semen, needle aspirates, and the like. Biological samples can be obtained from various families of domestic animals, as well as feral or wild animals, including, but not limited to, ungulates, bears, fish, rodents, etc. Environmental samples include environmental materials such as surfaces, soil, water, and industrial samples, as well as samples obtained from food and dairy processing equipment, devices, facilities, tools, disposable and non-disposable items. These examples should not be considered limiting of the sample types applicable to the present invention.
[0117] Any standard method known in the art for detecting binding between a nucleotide-binding molecule and a linear probe polynucleotide may be used in the context of the present invention [e.g., "Crystal structure of Δ-[Ru(bpy)2dppz] 2+ "bound to mismatched DNA reveals side-by-side metalloinsertion and intercalation" Nature Chemistry,2012 Volume 4,No 8,615-620;Hang Song,Jens T.Kaiser&Jacqueline K.Barton;or "Label-free detection of DNA-binding proteins based on microfluidic solid-state molecular beacon sensor" Anal Chem.2011,83(9),3528-32.Wang J,Onoshima D,Aki M,Okamoto Y,Kaji N,Tokeshi M,Baba Y.; or Annu Rev Anal Chem,2011,4(1),105-128."Metal Ion Sensors Based on DNAzymes and Related DNA Molecules"Xiao-Bing Zhang,Rong-Mei Kong, and Yi Lu].
[0118] Comparison of sequences and determination of percent identity or similarity between two sequences can be accomplished using a mathematical algorithm. In a preferred embodiment, percent identity between two nucleotide sequences is determined using the GAP program in the GCG software package (available at http: / / www.gcg.com), using the NWSgapdna.CMP matrix and a gap weight of 40, 50, 60, 70, or 80 and a length weight of 1, 2, 3, 4, 5, or 6. A particularly preferred set of parameters (and one that should be used if the practitioner is unsure what parameters to apply to determine whether a molecule falls within the sequence identity or homology limits of the invention) is the BLOSUM62 scoring matrix with a gap penalty of 12, a gap extend penalty of 4, and a frameshift gap penalty of 5.
[0119] Alternatively, percent identity between two nucleotide sequences can be determined using the algorithm of Meyers et al. (1989) CABIOS 4:11-17), which has been incorporated into the ALIGN program (version 2.0), using a PAM120 weight residue table, a gap length penalty of 12, and a gap penalty of 4.
[0120] The nucleic acid and protein sequences described herein can be used as "query sequences" to perform searches against public databases, e.g., to identify other family members or related sequences. Such searches can be performed using the NBLAST and XBLAST programs (version 2.0) of Altschul, et al. (1990) J. Mol. Biol. 215:403-410. To obtain nucleotide sequences homologous to the nucleic acid molecules of the present invention, BLAST nucleotide searches can be performed with the NBLAST program, score = 100, word length = 12. To obtain amino acid sequences homologous to the protein molecules of the present invention, BLAST protein searches can be performed with the XBLAST program, score = 50, word length = 3. To obtain gapped alignments for comparison, gapped BLAST, as described in Altschul et al. (1997, Nucl. Acids Res. 25:3389-3402), can be used. When utilizing BLAST and gapped BLAST programs, the default parameters of the respective programs (eg, XBLAST and NBLAST) can be used.<htttp: / / www.ncbi.nlm.nih.gov> Please refer to.
[0121] Aspects of the present invention are illustrated by the following non-limiting examples. [Example]
[0122] 1. Covalent immobilization of ssDNA onto solid glass surfaces via APEGDMES linkers Glass microscope slides were first cleaned with acetone, IPA, and nanopure water, and then treated with O2 plasma to remove any residual organic contaminants and activate the surface with OH groups. 8 。The surface was then modified with the acetal-protected aldehyde-terminated siloxane linker APEGDMES (SEQ ID NO: 12) (acetal polyethylene glycol dimethylethoxysilane) by heating overnight at 80 °C in toluene. APEGDMES modification yields a terminal acetal-protected aldehyde surface, which is easily removed in 10% aqueous acetic acid to display an aldehyde functionality on top of the linker. Amino-modified oligoseeds 5'-NH2-[GATC]5-3' were covalently coupled to the aldehyde surface using sodium cyanoborohydride to facilitate a reductive amination reaction between the aldehyde linker and the amino functionality on the DNA strand. The oligoseed-functionalized surface was then washed with nanopure water for 30 minutes and then with 0.5x PBS buffer to remove any physically adsorbed DNA strands. This resulted in short ssDNAs being covalently tethered to the surface.
[0123] 2. Generation of immobilized oligo seeds and their PCR extension The covalently tethered short ssDNA generated in Example 1 was then hybridized with its complementary strand, 5'-[CTAG]5-3', to form the 20-base starting oligoseed duplex required for PCR-based enzymatic extension reactions. The dsDNA surface was further rinsed with nanopure water and 0.5x PBS buffer. The DNA strands were then subjected to PCR-based heating-cooling extension cycles using Thermococcus gorgonarius family B polymerase (Tgo-Pol) enzyme variant Z3 as previously described (Pike et al., Angew 2015). The procedure was similar to the reported solution-based heating-cooling method, except that the oligoseed-modified silicon chip was scaled to a 0.5 cm diameter to fit into a standard PCR Eppendorf tube for thermal cycling using a heating-cooling block. 2The total volume of the reagents, Z3 enzyme, dNTPs, and buffer was 180 μL, sufficient to ensure constant immersion of the silicon surface. First, short dsDNA was dehybridized by heating to 95°C and then cooled to 55°C for rehybridization. However, repeat CATG / GTAC sequences do not always result in perfectly matched duplexes, and this mismatch is exploited to extend DNA from the surface. Upon rehybridization, the complementary strands can be shifted by one, two, or three units (assuming a minimum of eight dimer repeat duplexes is required to form a stable duplex), creating a 5'-overhang suitable for DNA polymerase extension. Next, DNA extension was performed at 72°C. Here, the enzyme adds the corresponding NTPs to the surface-bound sequences, thereby increasing the length of the DNA duplex by the number of misaligned repeat units. The increase in the length of the immobilized DNA occurs in the z-direction away from the surface, maintaining the surface packing density of the short oligoseeds. This heating-cooling method is repeated for up to 20 cycles to obtain DNA brushes approximately 700 bases long (Figure 2).
[0124] 3. Visualization of Extended Polynucleotide Sequences Visualization of the increased number of DNA bases packed into the same surface area was demonstrated by the addition of the fluorescent dye Picogreen (PG), which intercalates with dsDNA. PG exhibits a >1000-fold increase in fluorescence upon binding to dsDNA. After 30 minutes of incubation in a 200-fold dilution of the stock solution, the enzyme-treated surface exhibited an increase in fluorescence intensity compared to the untreated starting oligonucleotide strands on the surface after 20 heating-cooling cycles, as shown in Figures 3a and 3b.
[0125] While the change in fluorescence is visible to the naked eye (Figures 3a and 3b), it is more evident when integrated density analysis is performed using ImageJ software on the two Picogreen-labeled surfaces (Figure 3c). The extension reaction dramatically increases the number of Picogreen binding sites per surface attachment point, resulting in an increase in fluorescence intensity. There was a clear difference between the short oligo seed surface and the extDNA surface, attributed to the increased number of PG binding sites per probe molecule.
[0126] The increase in fluorescence intensity is a good indicator that the DNA has been extended from the surface. To further confirm the increase in DNA length, the surface extDNA was dehybridized, removing and collecting the long complementary strands that were not covalently attached to the surface. For dehybridization, the extDNA surface was immersed twice in nanopure water heated to 95°C. A 5 μL sample was taken and placed on freshly cleaved mica by molecular combing, which is known to stretch flexible DNA strands, to visualize the long DNA using AFM. 10 An AFM image of ssDNA can be seen in Figure 4a. Many of the ssDNA strands are aggregated on the surface, which is typical for ssDNA. 11 Some single strands exhibit an average height of 0.5 nm (see Figure 4b), which is comparable to the reported dimensions of ssDNA. 12 Analysis of the average length of the ssDNA strands gave a range of 160–300 bp, confirming successful extension of oligo seeds from the surface. The strand lengths obtained from the surface were shorter than those observed in agarose gels from solution-based extension. While not wishing to be bound by any particular theory, this may be because steric hindrance of the surface limits sufficient movement of the enzyme, leading to a reduced reaction rate.
[0127] 4.CTFR gene sequence detection Long DNA brushes with repeat sequences can distinguish single-base mismatches in diseases, making them ideal for DNA biosensing. Many diseases are the result of mutations in genes that cause differences in DNA sequences. 14A platform for discriminating between specific sequences would allow for early detection of disease or determination of defect type.
[0128] One non-limiting example in which this type of device would be particularly beneficial is the cystic fibrosis transmembrane conductance regulator (CFTR) gene. The CFTR gene encodes a protein that acts as a channel to control the transport of chloride ions in and out of cells, thereby regulating water content for mucus production. This protein is also required for sodium regulation in the lungs and pancreas. The most common frameshift mutation found in CFTR is a triple deletion of the bases CTT, known as delta F508, which alters the gene coding by removing the amino acid phenylalanine (see Figure 5). The amino acid deletion results in a complete distortion of the gene's shape, which causes the ion channel to malfunction. Without a functioning ion channel, cells lining the pancreas, lungs, and other organs produce thick mucus, which leads to blockage of airways and glands. Early detection of this alteration would allow for immediate treatment, reducing the impact of the disease and improving quality of life. Therefore, the DNA sequence for the CFTR delta F508 mutation was chosen as the next sequence to be tested using the surface-extension method.
[0129] To ensure that this particular sequence could be extended using this method, the CFTR sequence was first extended in solution as described herein. Duplexes were formed from the amino-modified probes and target DNA sequences listed in Table 1.
[0130] [Table 1]
[0131] The DNA product was analyzed by agarose gel electrophoresis (Figure 6a), and a modal length of 300 bp was determined using ImageJ analysis software. This DNA sequence was shorter and less concentrated than the GATC sequence. However, it was long enough to provide 30 repeats per probe strand, so further extension from the surface was attempted.
[0132] Having established that surface-immobilized oligoseeds can be extended and that extended sequences exhibit increased fluorescent signals, glass slides were prepared with oligoseeds of the sequence 5' GCATCTTTCG 3' (SEQ ID NO: 1), a CTFR gene that, upon mutation, results in a three-base mismatch frameshift (bold) 5' GCATTCGAGC 3' (SEQ ID NO: 15). Clearly, the ability to rapidly analyze this shift in base sequence to aid in the early detection of cystic fibrosis would be of diagnostic importance.
[0133] The CFTR oligoseed sequence was immobilized on a glass surface and subjected to reaction conditions for 20 cycles of enzymatic extension. After adding PG to the surface, fluorescence images were obtained, and an increase in fluorescence intensity was observed for the extDNA sample (see Figure 7). The extDNA exhibited increased fluorescence intensity compared to the short oligoseed, but the difference was not as great as for the GATC sequence. Without wishing to be bound by any particular theory, this may be a result of the fact that the CFTR sequence in solution did not extend as well as the GATC sequence, and therefore, there were fewer DNA bp for PG to bind, which resulted in a decrease in fluorescence intensity compared to the GATC surface.
[0134] To investigate the hydrophobicity of different surface modifications, the contact angles of ssDNA, dsDNA, and extDNA surfaces were determined (see Table 2). For dsDNA, the contact angle was 32.33°. After extension from the surface, the contact angle dramatically increased to 73.11°. Without wishing to be bound by any particular theory, the large increase in hydrophobicity is attributed to the hydrophobicity of the DNA bases. In the duplex conformation of short oligomers, the bases are shielded from any water molecules, but for longer DNA, the structural rigidity is reduced, allowing more bases to come into contact with the water droplet, resulting in a hydrophobic terminal monolayer and an increased contact angle. 15 .
[0135] [Table 2]
[0136] The approach described here can enhance the fluorescence response by providing elongated, densely packed arrays of target genes, which holds promise for early detection of single- and few-base mismatches.
[0137] 5. Covalent immobilization and extension of ssDNA on silicon surfaces Silicon surface, 25mm 2n-type Si <111> Wafers were cleaned and modified as described above for glass slides. Another diagnostically important oligo seed sequence, 5'-[GTTTTTTTTTTC]2-3' (SEQ ID NO: 16), consisting of a long stretch of 10 T bases, was immobilized using similar siloxane chemistry and then enzymatically extended to yield multiple repeats. In this case, the target sequence is important due to the single-base misincorporation that occurs when consecutive Ts are extended or shortened by mistranscription. This is an example of a potential MSI target involved in the development of colorectal cancer [see, e.g., Arq Bras Cir Dig. 2012 Oct-Dec;25(4):240-4, "Microsatellite instability - MSI markers (BAT26, BAT25, D2S123, D5S346, D17S250) in rectal cancer," Losso GM1, Moraes Rda S, Gentili AC, Messias-Reason IT].
[0138] Fluorescence data from this stretched surface showed the same amplified response as the glass surface.
[0139] 6. Array [G 10 :C 10 Covalent immobilization of ] and extension from glass surfaces Using a biotin / streptavidin substrate interface, we patterned target ssDNA on glass substrates so that it was confined to small islands on the substrate surface [see, e.g., Nakamura S, Mitomo H, Aizawa M, Tani T, Matsuo Y, Niikura K, Pike AR, Naya N, Shishido A, Ijiro K. "DNA Brush-Directed Vertical Alignment of Extensive Gold Nanorod Arrays with Controlled Density," ACS Omega, 2017, 2(5), 2208-2213]. This demonstrates that the processivity of the Z3 enzyme is not hindered by the type of substrate or by more complex functionalization approaches involving the protein streptavidin within the linker layer. After hybridization and extension as previously described, SYBR green dye was applied to the surface. The fluorescence image in Figure 3d shows that the DNA regions exhibit fluorescence, while the bare tracks between them are unresponsive. This provides direct evidence for the feasibility of multiplexing this approach, where every DNA spot has a different sequence. To confirm that the DNA remained tethered to the surface while increasing in length, the final dsDNA on the surface was dehybridized by heating the surface to 95°C in water (1 mL x 2). This complementary strand, therefore, not bound to the surface, was removed from the surface during two heated washes. AFM was used to confirm the presence of long ssDNA by spotting a 5 μL sample from the combined washes onto freshly cleaved mica using molecular combing in an attempt to stretch the flexible molecule. The ssDNA strands appeared to aggregate with each other. While this is typical for ssDNA on mica surfaces, some single strands exhibited an average height of 0.5 nm, similar to previous dimensions reported for ssDNA. Length analysis of the single strands yielded a range of 160–300 bp, confirming successful extension from the surface. Elongation of [GATC]5 / [CTAG]5 in solution resulted in an average length of 750 bp as observed by agarose gel electrophoresis and AFM.
[0140] Interestingly, contact angle measurements revealed an increase from 25.6° on the initial oligo-seeded dsDNA surface to 74.8° on the enzymatically extended dsDNA surface.
[0141] 7. Extension of Bat25 sequence Microsatellite instability (MSI) is a marker of genetic instability found in the majority of tumors in patients with hereditary colorectal cancer and in some sporadic colorectal cancers. 16 MSI is a non-coding mononucleotide repeat sequence that exhibits differences in allele length due to deletions or insertions in tumor cells compared to normal DNA alleles from the same patient. One of the most commonly used mononucleotide repeat markers used to identify MSI is the BAT25 sequence, a poly-T repeat unit. The BAT25 sequence can be used without comparison to normal DNA, and is associated with significant base deletions in virtually all tumors exhibiting MSI. 16 Understanding the types of MSI allows for tumor type identification and can predict a patient's chemotherapy response. Rapid, sensitive, and reproducible methods for MSI identification are needed to enable rapid diagnosis and treatment.
[0142] Current methods for MSI recognition use a specially designed panel, the Bethesda panel, that screens for five microsatellite markers: the mononucleotide repeat markers BAT25 and BAT26 and the dinucleotide repeat markers D2S123, D5S346, and D17S250. If two to five of these markers are mutated, the patient is considered to have high microsatellite instability (MSI-H). However, there are discrepancies in the consistent specificity and sensitivity of this panel, which limits the reliability of this test. 17 .
[0143] Therefore, identification of MSI using the extension method described here will reduce diagnostic time and increase sensitivity, and when used in arrays, it will be possible to simultaneously screen for multiple MSI markers from a single DNA sample. The BAT25 mononucleotide repeat sequence was used to test the efficiency of this microsatellite instability detection method. The probe and target strands used to form the BAT25 oligoseeds are listed in Table 3. Extension of the amino-modified duplex was attempted in solution before extension from a surface was attempted.
[0144] [Table 3]
[0145] The BAT25 extension product was analyzed by gel electrophoresis (see Figure 8a) and had a modal length of 2000 bp. DNA Sanger sequencing was performed on the 5-cycle extension product of the BAT25 sequence by GATC-biotech (see Figure 8d). The sequencing results confirmed the accurate incorporation of each base by the DNA polymerase. DNA sequencing could not be achieved for the GATC and CFTR sequences. Without wishing to be bound by any particular theory, this is likely due to challenges associated with GC-rich DNA sequences. Both the GATC and CFTR sequences contain 50% GC content.
[0146] The DNA length of the BAT25 sequence was longer than that of the GATC and CFTR sequences, and showed improved elongation at a high concentration of 618 ng / µL. The bp length ranged from approximately 750 bp to 3000 bp, which was confirmed by AFM analysis of the dsDNA elongation products (Figure 9). The average dsDNA height was calculated to be 0.87 nm, and the DNA length ranged from 130 to 1500 bp, consistent with the gel analysis. After attachment of the amino-modified oligoseeds, the BAT25 sequence was extended from the surface. Upon addition of PG, the extDNA revealed increased fluorescence compared to the short immobilized oligoseeds (see Figure 10).
[0147] The same increase in contact angle measurements as seen with the CFTR sequence was also observed with the BAT25 extDNA sequence.
[0148] [Table 4]
[0149] Double-stranded extended DNA (dsextDNA) strands were dehybridized as described in 8.5 below and combed onto freshly cleaved mica for AFM analysis (see Figure 11). Single-stranded AT-rich DNA is prone to folding and stacking. 18,19 Although many of the strands appeared to be aggregated, some strands were elongated and available for height and length analysis. The average single strand height was 0.5 nm and the average length was 70 nm, corresponding to 200 bp.
[0150] The enhanced fluorescence intensity for the BAT25 sequence was greater than that observed for either the GATC or CFTR sequences. Comparison of gel electrophoresis and fluorescence intensity for the GATC, CFTR, and BAT25 sequences (see Figure 12) confirmed the longer extensions both in solution and from the surface.
[0151] Because the BAT25 sequence extension exhibited the strongest increase in fluorescence intensity compared to the other sequences studied and is medically important, the BAT25 sequence was used for further investigation of the sensing applications of this device.
[0152] [Table 5]
[0153] 8. Parameter Optimization Although differences in fluorescence intensity were distinguishable for extDNA compared to short oligo seeds, several parameters were implemented to improve sensitivity and reliability.
[0154] PG has been established as a base-pair independent dsDNA intercalator. Other sequence-specific fluorescent dye candidates may further enhance the fluorescence intensity change of extDNA samples. Because the BAT25 sequence is AT-rich, the fluorescent stain DAPI was analyzed for its suitability. DAPI is a commonly used stain due to its 20-fold increase in fluorescence upon binding to AT regions in dsDNA. 20 The DAPI solution was applied to the surface for 20 minutes, and fluorescence images were acquired at an excitation wavelength of 360 nm to observe differences in fluorescence intensity (Figure 13). The fluorescence intensity (FI) of the extDNA sample was lower than that of the short dsDNA. The difference between dsDNA and extDNA was consistently greater when PG was used, so we continued to use PG.
[0155] PG is traditionally used in Tris-EDTA buffer TE. 21 When the PG-TE solution was applied to the surface, several salt spots were observed, but when PG was deposited in nanopure H2O, these salt spots were not observed (see Figure 14b). Also, because the FI of PG in H2O was significantly higher than when PG was in TE, in subsequent studies, PG was dissolved in H2O. Many papers state that the binding time of PG is almost instantaneous upon interaction with DNA. 21 In the case of surface-localized DNA, it may take longer for PG to intercalate into the DNA. PG was applied to the extDNA surface for various times and the fluorescence intensity was compared (see Figure 15). FI increased continuously with each 5-minute increment, followed by a decrease over 20 minutes. Therefore, in subsequent studies, PG was applied to the surface for 20 minutes and then washed. All stretching from the surface was first performed on cleaved glass microscope slides. To see if the same phenomenon could be observed on different surfaces, the stretching protocol was performed on silicon wafers. The stretching procedure was performed on 0.5 cm 2This was performed on a cleaved p-type (100) silicon wafer and PG was applied to the surface for fluorescence imaging (see Figure 16). An increase in fluorescence was observed from short dsDNA to extDNA, but the difference was smaller than that seen on glass surfaces. Nevertheless, there was still a clear enhancement for extDNA, highlighting the versatility of this extension method. Not only the sequence but also the surface can be tailored to the desired application.
[0156] 9. Discussion of Results The method described herein enables the synthesis of multiple probe DNA sequences on a surface to enhance target detection. The invention utilizes an enzymatic DNA synthesis method (patent number GB17000531.5) that allows for the rapid synthesis of DNA with repeat units in solution. The repeat units can be tailored to match the binding regions of molecules of interest, such as DNA fragments indicative of disease states. After denaturation (unwinding of the double-stranded DNA), each single-stranded DNA has multiple binding sites along the perpendicular axis, maximizing the binding opportunities per probe strand bound to the surface. Binding to complementary target fragments results in increased fluorescence intensity compared to short dsDNA on the surface. Elongation increases the number of target binding sites per probe molecule and, therefore, target detection and diagnostic sensitivity.
[0157] The inventors have demonstrated that the present invention has superior sensitivity, thereby addressing one of the key challenges in molecular diagnostics: detecting a signal against background noise. Because signal generation is enhanced, the sensitivity of the present invention can be leveraged to detect low-abundance target DNA molecules, thereby eliminating the need for extended use of expensive and time-consuming sample preparation techniques such as PCR. Alternatively, the true cost-benefit of the present invention may lie in enhanced data collection. This, in turn, means that less detector sensitivity and signal processing are required, facilitating improved portability of diagnostic devices. Point-of-care devices incorporating this technology could be economically manufactured using less expensive parts, software, and design elements, ultimately reducing overall economic expenditures. Point-of-care devices could benefit from enabling technologies such as those presented herein, making them inexpensive to manufacture and suitable for portable implementation.
[0158] Experimental materials and methods Chemical Reagents All chemical reagents were purchased from Sigma-Aldrich and used as received without further purification. Glass slides were purchased from Henso Labware Manufacturing Co., Ltd. (Hangzhou, China). APEGDMES was purchased from NewChem Technologies Limited. DNA was purchased from Eurofins Genomics (Ebersberg, Germany). Tgo-Pol Z3 exo- was prepared and purified in-house. (Jozwiakowski, SK & Connolly, BA "A modified family-B archaeal DNA polymerase with reverse transcriptase activity" ChemBioChem 12, 35-37 (2011); Evans, SJ et al. "Improving dideoxynucleotide-triphosphate utilization by the hyperthermophilic DNA polymerase from the archaeon Pyrococcus furiosis" Nucleic Acids Res 28, 1059-1066 (2000)).
[0159] Primer-template annealing Primer duplexes were prepared for extension: DNA annealing buffer (10 mM Hepes pH 7.5, 100 mM NaCl, and 1 mM EDTA) was added to the oligomers and heated to 95°C for 10 minutes. The duplex solution was slowly cooled to room temperature and stored at -20°C.
[0160] DNA polymerase For extension, we used Tgo-Pol Z3 exo-DNA polymerase, a low-fidelity variant of archaeal family B polymerases in which the 3'→5' exonuclease activity has been ablated and the fingers domain has been modified.
[0161] DNA elongation in solution 0.5 μM DNA duplex, 200 nM Tgo-Pol Z3 exo-DNA polymerase, DNA polymerase reaction buffer (200 mM Tris-HCl (pH 8.8, 25 °C), 100 mM (NH4)2SO4, 100 mM KCl, 1% Triton X-100, 1 mg / mL bovine serum albumin (BSA), and 20 mM MgSO4), and 0.5 mM deoxynucleotide triphosphates (dNTPs) (dCTP, dATP, dTTP, and dGTTP) were mixed. Heat-cool thermal cycling was performed in an Applied Bioscience Veritt 96-well thermal cycler for the following cycles: Cycle number (20) × 30 s at 95°C, 30 s at 55°C, and 2 min at 72°C. After the reaction, the solution was cooled to 4°C. The DNA extension product was purified using a QIAquick PCR purification kit (25) (QIAGEN, Manchester, UK) according to the manufacturer's protocol.
[0162] Agarose gel electrophoresis DNA extension products were analyzed by gel electrophoresis in TBE buffer (Tris, boric acid, and Na2EDTA.2H2O). 1% agarose (Melford, Ipswich, UK) was added to 1% TBE buffer and heated until completely dissolved. The gel mixture was cooled to 50°C and poured into a gel set to solidify. DNA ladders 1 kb and 1 kb+ (Thermo Scientific) were supplemented with loading dye (2.5% Ficoll-400, 11 mM EDTA, 3.3 mM Tris-HCl (pH 8.0, 25°C), 0.017% SDS, and 0.015% bromophenol blue). 2 μL of gel loading dye was added to the DNA sample (20 ng / μL) and loaded into the gel wells. The gel was run at 100 V, 100 mA, and 10 W for approximately 1 hour. The gel was post-stained with a 5 μg / mL solution of ethidium bromide and visualized using a UV transilluminator.
[0163] UV UV-Vis spectroscopy was performed using a Nanodrop. The spectrometer was blanked with nanopure H2O.
[0164] Surface Pretreatment Glass slide or n-type Si <111> Dicing the wafer into 0.5cm 2 The chips were wiped with acetone, IPA, and NP-H2O, sonicated sequentially in acetone, IPA, and NP-H2O for 15 min, and dried with N2. The chips were then subjected to O2 plasma treatment for 15 min.
[0165] Attachment of oligo seed DNA via APEGDMES linker The cleaned chips were immersed in an APEGDMES / toluene solution (233 μM, 3 mL) preheated to 65 °C for 16 h. The chips were washed three times sequentially with toluene, ethanol, and NP-HO, and then placed in a vacuum oven at 120 °C for 40 min. An amino-tagged DNA probe solution (40 μL, 100 μM) in 10% acetic acid solution was drop-cast onto the chip for 1 h in a humid environment. NaCNBH3 (40 μL, 16 μM) in 50% MeOH solution was placed on top of the probe solution for an additional 2 h in a humid environment. The chips were washed with phosphate-buffered saline (0.5x) and excess water to remove any excess DNA molecules. The chips were dried with a stream of nitrogen.
[0166] DNA hybridization A complementary DNA target solution (40 μL, 200 nM) in 0.5x PBS buffer was drop-cast onto the silicon chip for 15 min in a humid environment. The chip was washed with 0.5x PBS buffer for 30 min, then with NP-HO for 30 min, and then dried under a stream of nitrogen.
[0167] DNA extension from a surface For heating-cooling cycling, the chip was placed in an Eppendorf with the following required solutions: 200 nM DNA polymerase, DNA polymerase reaction buffer (200 mM Tris-HCl (pH 8.8, 25 °C), 100 mM (NH4)2SO4, 100 mM KCl, 1% Triton X-100, 1 mg / mL bovine serum albumin (BSA), and 20 mM MgSO4), and 0.5 mM deoxynucleotide triphosphates (dNTPs) (dCTP, dATP, dTTP, and dGTTP). Heating-cooling thermal cycling was performed in an Applied Bioscience Veritt 96-well thermal cycler over the following cycles: Cycle number (20) × 30 s at 95°C, 30 s at 55°C, and 2 min at 72°C. After the reaction, the solution was cooled to 4° C. The chip was removed from the solution, washed in NP-H2O for 30 minutes, and dried under a stream of nitrogen.
[0168] Attachment of oligo-seed DNA via a biotin-streptavidin-biotin linker The glass surface was cleaned with piranha solution, then rinsed with H2O and dried. At <25% humidity, a solution of 2-carbomethoxyethyltrichlorosilane in dry toluene was added. After 1 hour, the glass was washed with acetone, ethanol, and then H2O, then flooded with HCl, and left overnight. The glass was rinsed extensively with H2O and covered with 50 mM 1-ethyl-3-(3-dimethylaminopropyl)carbodiimide and 1 mM amine-PEG2-biotin in 10 mM Hepes. After 1 hour, the glass was washed, and then 50 μL of 0.1 mg / mL streptavidin in 10 mM Tris, pH 7.9, was pipetted onto the surface. After 1 hour, the glass was washed and then patterned using a UV photomask and UV exposure. The surface was then rinsed with 10 mM Tris, pH 7.9, and covered with 1 μM C in 10 mM Tris, pH 7.9, and 200 mM NaCl. 15 After 1 hour, the surface was washed with 10 mM Tris-HCl and 50 μL of 1 μM G 1510 mM Tris-HCl, pH 7.9, and 200 mM NaCl were added. After 1 hour, the glass surface was washed and filled with polymerization solution: 0.5 mM dCTP and dGTP, 200 nM Tgo-Pol Z3 exo-, 200 mM NaCl, and 0.5 mM MgCl2 in 15 mM Tris-HCl, pH 7.6. The reaction was stopped by removing the reaction solution and washing with 10 mM Tris, pH 7.9, and 200 mM NaCl. The surface was stained with SYBR Green.
[0169] Fluorescence microscopy imaging The samples were mounted on an Axioshop 2 plus (Zeiss, Germany) imaging platform equipped with a Plan-NEOFLUAR 10× / 0.3 objective (Zeiss) set at filter 44. Samples were excited at 490 nm from an ebq100 mercury lamp (LEJ, Germany) and imaged using an AxioCam HRm (Zeiss).
[0170] AFM The top layer of the mica surface was cleaved using adhesive tape. 5 μL of DNA sample (2 ng / μL or 4 ng / μL) was placed on the mica surface, which was held at a 25° angle to allow the DNA to flow across the mica surface. After 5 min, 5 μL of nanopure H2O was dropped onto the DNA sample, again at a 25° angle. A gentle stream of N2 was passed over the surface, which was then further dried under laminar airflow for 1 h. AFM images were acquired using a Dimension V and a nanoscope controller (Vecco Instruments Inc., Metrology Group, Santa Barbara, CA) on an isolation table (Veeco Inc., Metrology Group) to reduce interference. Data were acquired using NanoScope Analysis 1.8 software.
[0171] contact angle Contact angle measurements were performed on a KSV Cam 101 (KSV Instruments Ltd., Finland) using the built-in CAM 2008 software. A 1 μL droplet of NP-H2O was placed on the surface. The software was used to estimate the angle of the water droplet on the surface. Ten measurements were taken for each sample. If the difference between the angle on the left and right sides of the droplet was more than 2°, the collected measurement was discarded. Measurements that were more than two standard deviations away from the mean were also discarded.
[0172] The reader is directed to all papers and documents related to this application that have been filed contemporaneously or previously hereto and are published herewith, the entire contents of which are incorporated herein by reference.
[0173] All features disclosed in this specification (including any accompanying claims, abstract and drawings), and / or all steps of any method or process so disclosed, may be combined in any combination, except combinations in which at least some of such features and / or steps are mutually exclusive.
[0174] Each feature disclosed in this specification (including any accompanying claims, abstract, and drawings), unless expressly stated otherwise, may be replaced by alternative features serving the same, equivalent, or similar purpose. Thus, unless expressly stated otherwise, each feature disclosed is only an example of a generic series of equivalent or similar features.
[0175] The invention is not limited to the details of any of the embodiments described above, and extends to any novel one or any novel combination of features disclosed in this specification (including any accompanying claims, abstract and drawings), or to any novel one or any novel combination of any method or process so disclosed.
[0176] References 1 Gong, P. & Levicky, R. DNA surface hybridization regimes. Proceedings of the National Academy of Sciences 105, 5301-5306 (2008). 2 Metzker, M. L. Sequencing technologies - the next generation. Nature Reviews Genetics 11, 31 (2009). 3 Valignat, M.-P., Theodoly, O., Crocker, J. C., Russel, W. B. & Chaikin, P. M. Reversible self-assembly and directed assembly of DNA-linked micrometer-sized colloids. Proceedings of the National Academy of Sciences of the United States of America 102, 4225-4229 (2005). 4 Bracha, D., Karzbrun, E., Shemer, G., Pincus, P. A. & Bar-Ziv, R. H. Entropy-driven collective interactions in DNA brushes on a biochip. Proceedings of the National Academy of Sciences 110, 4534-4538 (2013). 5 Wang, C. et al. DNA microarray fabricated on poly(acrylic acid) brushes-coated porous silicon by in situ rolling circle amplification. Analyst 137, 4539-4545 (2012). 6 Jozwiakowski, S. K. & Connolly, B. A. A Modified Family-B Archaeal DNA Polymerase with Reverse Transcriptase Activity. ChemBioChem 12, 35-37 (2011). 7 Whitfield, C. J., Turley, A. T., Tuite, E. M., Connolly, B. A. & Pike, A. R. Enzymatic Method for the Synthesis of Long DNA Sequences with Multiple Repeat Units. Angewandte Chemie International Edition 54, 8971-8974, doi:10.1002 / anie.201502971 (2015). 8 Terpilowski, K. & Rymuszka, D. Surface properties of glass plates activated by air, oxygen, nitrogen and argon plasma. Glass Physics and Chemistry 42, 535-541 (2016). 9 Dragan, A. I. et al. Characterization of PicoGreen Interaction with dsDNA and the Origin of Its Fluorescence Enhancement upon Binding. Biophysical Journal 99, 3010-3019 (2010). 10 Li, J. et al. A convenient method of aligning large DNA molecules on bare mica surfaces for atomic force microscopy. Nucleic Acids Research 26, 4785-4786 (1998). 11 Hansma, H. G., Sinsheimer, R. L., Li, M.-Q. & Hansma, P. K. Atomic force microscopy of single-and double-stranded DNA. Nucleic Acids Research 20, 3585-3590 (1992). 12 Hansma, H. G., Revenko, I., Kim, K. & Laney, D. E. Atomic Force Microscopy of Long and Short Double-Stranded, Single-Stranded and Triple-Stranded Nucleic Acids. Nucleic Acids Research 24, 713-720 (1996). 13 Mattheyses, A. L., Simon, S. M. & Rappoport, J. Z. Imaging with total internal reflection fluorescence microscopy for the cell biologist. Journal of Cell Science 123, 3621-3628 (2010). 14 R. A. Bartoszewski et al. A Synonymous Single Nucleotide Polymorphism in ΔF508 CFTR alters the Secondary Structure of the mRNA and the Expression of the Mutant Protein. J. Bio. Chem 285, 28741-28748 (2010). 15 Costa, D., Miguel, M. G. & Lindman, B. Responsive Polymer Gels:Double-Stranded versus Single-Stranded DNA. The Journal of Physical Chemistry B 111, 10886-10896 (2007). 16 Zhou, X.-P. et al. Determination of the replication error phenotype in human tumors without the requirement for matching normal DNA by analysis of mononucleotide repeat microsatellites. Genes, Chromosomes and Cancer 21, 101-107 (1998). 17 Umar, A. et al. Revised Bethesda Guidelines for Hereditary Nonpolyposis Colorectal Cancer (Lynch Syndrome) and Microsatellite Instability. Journal of the National Cancer Institute 96, 261-268 (2004). 18 Luzzati, V., Mathis, A., Masson, F. & Witz, J. Structure transitions observed in DNA and poly A in solution as a function of temperature and pH. Journal of Molecular Biology 10, 28-41 (1964). 19 Mills, J. B., Vacano, E. & Hagerman, P. J. Flexibility of single-stranded DNA: use of gapped duplex helices to determine the persistence lengths of Poly(dT) and Poly(dA)11Edited by B. Honig. Journal of Molecular Biology 285, 245-257 (1999). 20 Kapuscinski, J. DAPI: a DNA-Specific Fluorescent Probe. Biotechnic & Histochemistry 70, 220-233 (1995). 21 Singer, V. L., Jones, L. J., Yue, S. T. & Haugland, R. P. Characterization of PicoGreen Reagent and Development of a Fluorescence-Based Solution Assay for Double-Stranded DNA Quantitation. Analytical Biochemistry 249, 228-238 (1997). 22 Dodge, A., Turcatti, G., Lawrence, I., de Rooij, N. F. & Verpoorte, E. A Microfluidic Platform Using Molecular Beacon-Based Temperature Calibration for Thermal Dehybridization of Surface-Bound DNA. Analytical Chemistry 76, 1778-1787 (2004). 23 Lockett, M. R. & Smith, L. M. Fabrication and Characterization of DNA Arrays Prepared on Carbon-on-Metal Substrates. Analytical Chemistry 81, 6429-6437 (2009). 24 Eda, G., Fanchini, G. & Chhowalla, M. Large-area ultrathin films of reduced graphene oxide as a transparent and flexible electronic material. Nature Nanotechnology 3, 270 (2008). 25 Le, L. T., Ervin, M. H., Qiu, H., Fuchs, B. E. & Lee, W. Y. Graphene supercapacitor electrodes fabricated by inkjet printing and thermal reduction of graphene oxide. Electrochemistry Communications 13, 355-358 (2011). 26 He, Q. et al. Centimeter-Long and Large-Scale Micropatterns of Reduced Graphene Oxide Films: Fabrication and Sensing Applications. ACS Nano 4, 3201-3208 (2010). 27 Mohanty, N. & Berry, V. Graphene-Based Single-Bacterium Resolution Biodevice and DNA Transistor: Interfacing Graphene Derivatives with Nanoscale and Microscale Biocomponents. Nano Letters 8, 4469-4476 (2008).
Claims
1. 1. A thermal cycling method for increasing the number of tandem repeats of a unit sequence that is 1 to 60 nucleotides in length in a linear polynucleotide, the method comprising the steps of: i) providing a solid substrate having a surface on which single-stranded primer polynucleotides consisting of two or more tandem repeats of a unit sequence having a length of 1 to 60 nucleotides are immobilized; ii) contacting the immobilized primer polynucleotide with a single-stranded template polynucleotide consisting of two or more tandem repeats complementary to the unit sequence of the primer polynucleotide under hybridization conditions that allow mismatch duplex formation between the unit sequence and its complementary strand, such that a 5' overhang of the template polynucleotide is generated, the 5' overhang comprising at least one tandem repeat complementary to the unit sequence of the primer polynucleotide; and iii) contacting the mismatched duplex with a thermostable 5' to 3' polymerase and nucleotides under extension conditions that permit polynucleotide extension in the 5' to 3' direction.
2. The method of claim 1, wherein the solid substrate having a surface onto which single-stranded linear primer polynucleotides consisting of two or more tandem repeats of a unit sequence having a length of 1 to 60 nucleotides are immobilized is prepared by immobilizing double-stranded linear primer polynucleotides consisting of two or more tandem repeats of a unit sequence having a length of 1 to 60 nucleotides onto the surface of the solid substrate, and denaturing the double-stranded linear primer polynucleotides to obtain the single-stranded linear primer polynucleotides.
3. 3. The method of claim 1 or 2, further comprising the steps of: iv) denaturing the duplex of iii) under denaturing conditions to produce immobilized single-stranded polynucleotides; and v) repeating steps ii) to iii) at least once to increase the number of tandem repeats in the immobilized polynucleotide.
4. The method of any one of claims 1 to 3, wherein the immobilized primer polynucleotide comprises 2 or more, 5 or more, 10 or more, or 15 or more tandem repeats of the unit sequence.
5. 5. The method of any one of claims 1 to 4, wherein the linear polynucleotide is covalently or non-covalently immobilized on the surface, optionally wherein the linear polynucleotide is non-covalently immobilized on a chemically modified region of the surface.
6. the polynucleotide is immobilized on the surface by a linker; 6. The method of any one of claims 1 to 5, wherein the linker comprises a silane linker molecule, a biotin-streptavidin complex, a thiol-Au linker, a Si-C covalent bond to silicon, a Si-O covalent bond to silicon, a Si-N covalent bond to silicon, a nanoparticle linker, or a dynamic covalent bond.
7. The method according to any one of claims 1 to 6, wherein the unit sequence is a microsatellite sequence having 2 to 9 nucleotides.
8. The method according to any one of claims 1 to 7, wherein the unit sequence is a minisatellite sequence having 10 to 60 nucleotides.
9. The method of any one of claims 1 to 8, wherein the linear polynucleotide is single-stranded DNA or double-stranded DNA.
10. 10. The method of any one of claims 1 to 9, wherein the surface comprises glass, silica, gold, graphene or graphene oxide, epoxy, plastic, metal, a gel matrix, metal created by template stripping, or a composite thereof.
Citation Information
Patent Citations
Genetic analysis and authentication method
JP2005537799A
Oligonucleotide repeat arrays
WO1995030774A1