Sequencing method
Luminescent perovskite markers with narrow emission spectra address the challenge of nucleotide identification errors in sequencing, enhancing accuracy and precision by providing distinct luminescence signatures for each nucleotide.
Patent Information
- Application Number
- GB2024003949
- Authority / Receiving Office
- GB · GB
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-03-19
- Publication Date
- 2025-10-01
AI Technical Summary
Current sequencing technologies face challenges in accurately determining the incorporation of nucleotides complementary to the next base of a template strand, particularly due to overlapping emission spectra of organic fluorophores leading to errors in nucleotide identification.
The use of luminescent perovskite markers with narrow emission spectra, such as particulate perovskites, to label nucleotides, allowing for precise identification of incorporated nucleotides through distinct luminescence detection, and optionally using cleavable linkers to facilitate sequential sequencing cycles.
Enhances the accuracy of nucleotide incorporation detection by reducing spectral overlap, thereby improving the reliability and precision of sequencing results.
Abstract
Description
The present disclosure provides a method of determining whether a first test nucleotide comprises a base complementary to the next base of a template strand immediately downstream of a primer in a primed template nucleic acid molecule, the method comprising: providing a primed template nucleic acid molecule; providing a labelled nucleotide comprising the first test nucleotide labelled with a first luminescent marker comprising a first luminescent material; contacting the primed template nucleic acid molecule with a reaction mixture that comprises a polymerase and the first test nucleotide, to thereby incorporate the first test nucleotide into the primed strand only if the first test nucleotide comprises a base complementary to the next base of the template strand; exciting the first luminescent marker with an excitation light; and detecting light emitted by the first luminescent marker, wherein the detection of emitted light identifies the incorporation of the first test nucleotide into the primed strand, and thereby indicates that the first test nucleotide comprises a base complementary to the next base of the template strand, and wherein the first luminescent marker comprises a particulate perovskite. Optionally, the particulate perovskite is a compound of formula ABX3 wherein A is a cation; B is a metal; and X in each occurrence is independently an anion. Optionally, A is an alkali metal or ammonium cation. Optionally, B is Sn or Pb. Optionally, each X is a halogen. Optionally, each X is the same halogen. Optionally, ABX3 comprises two different halogen groups X. Optionally, the luminescent marker comprises a silica shell. Optionally: the reaction mixture comprises a plurality of test nucleotides including the first test nucleotide comprising the first luminescent marker comprising a particulate perovskite and two or three further different species of test nucleotides; one or more of the two or three further different species of test nucleotides comprise a further luminescent marker which is different from the first luminescent marker; each of the plurality of test nucleotides has an emission characteristic selected from a characteristic light-emission in the case of a test nucleotide comprising a luminescent marker, or no light emission in the case of test nucleotide which does not comprise a luminescent marker; detection of the emission characteristic of one of the plurality of different species of test nucleotides identifies the incorporation of that particular test nucleotide into the primed strand, and thereby indicates that the particular test nucleotide comprises a base complementary to the next base of the template strand. Optionally, the one or more further luminescent markers include a luminescent marker comprising a particulate perovskite as described herein which is different from 5 the particulate perovskite of the first luminescent marker. Optionally, each of the luminescent markers comprise a particulate perovskite. The present disclosure provides a marked nucleotide comprising a nucleotide attached to a luminescent marker by a cleavable linker, wherein the luminescent marker comprises a particulate luminescent perovskite. 10 Optionally, the particulate luminescent perovskite comprises a shell. Optionally, the shell comprises silica. DETAILED DESCRIPTION Unless the context clearly requires otherwise, throughout the description and the claims, the words "comprise," "comprising," and the like are to be construed in an 15 inclusive sense, as opposed to an exclusive or exhaustive sense; that is to say, in the sense of "including, but not limited to." As used herein, the terms "connected," "coupled," or any variant thereof means any connection or coupling, either direct or indirect, between two or more elements. Additionally, the words "herein," "above," "below," and words of similar import, when used in this application, refer to this 20 application as a whole and not to any particular portions of this application. Reference to an element of the Periodic Table includes all isotopes of that element unless stated otherwise. Where the context permits, words in the Detailed Description using the singular or plural number may also include the plural or singular number respectively. The word "or," in reference to a list of two or more items, covers all of the following 25 interpretations of the word: any of the items in the list, all of the items in the list, and any combination of the items in the list. The teachings of the technology provided herein can be applied to other methods, not necessarily the methods described below. The elements and acts of the various examples described below can be combined to provide further implementations of 30 the technology. Some alternative implementations of the technology may include not only additional elements to those implementations noted below, but also may include fewer elements. These and other changes can be made to the technology in light of the following detailed description. While the description describes certain examples of the technology, and describes the best mode contemplated, no matter how detailed the description appears, the technology can be practiced in many ways. Details of the disclosed methods and systems may vary considerably in their specific implementation, while still being encompassed by the technology disclosed herein. As noted above, particular terminology used when describing certain features or aspects of the technology should not be taken to imply that the terminology is being redefined herein to be restricted to any specific characteristics, features, or aspects of the technology with which that terminology is associated. In general, the terms used in the following claims should not be construed to limit the technology to the specific examples disclosed in the specification, unless the Detailed Description section explicitly defines such terms. Accordingly, the actual scope of the technology encompasses not only the disclosed examples, but also all equivalent ways of practicing or implementing the technology under the claims. To reduce the number of claims, certain aspects of the technology are presented below in certain claim forms, but the applicant contemplates the various aspects of the technology in any number of claim forms. In the following description, for the purposes of explanation, numerous specific details are set forth in order to provide a thorough understanding of implementations of the disclosed technology. It will be apparent, however, to one skilled in the art that embodiments of the disclosed technology may be practiced without some of these specific details. The disclosed methods may comprise sequencing template fragments derived from a target nucleic acid. Unless defined otherwise, all technical and scientific terms used herein have the same meaning as is commonly understood by one of ordinary skill in the art. For clarity, the following specific terms have the specified meanings. The term "nucleic acid" can refer to at least two nucleotide monomers linked together. Examples include, but are not limited to DNA, such as genomic or cDNA; RNA, such as mRNA, sRNA or rRNA; or a hybrid of DNA and RNA. Thus, a "nucleic acid" is a polynucleotide, such as DNA, RNA, or any combination thereof, that can be acted upon by a polymerizing enzyme during nucleic acid synthesis. The term "nucleic acid" includes single-, double-, or multiple-stranded DNA, RNA and analogs (derivatives) thereof. As apparent from the disclosure below and elsewhere herein, a nucleic acid can have a naturally occurring nucleic acid structure or a non-naturally occurring nucleic acid analog structure. A nucleic acid can contain phosphodiester bonds; however, in some embodiments, nucleic acids may have other types of backbones, comprising, for example, phosphoramide, phosphorothioate, phosphorodithioate, O-methylphosphoroamidite and peptide nucleic acid backbones and linkages. Nucleic acids can have positive backbones; non-ionic backbones, and non-ribose based backbones. Nucleic acids may also contain one or more carbocyclic sugars. The nucleic acids used in methods or compositions herein may be single stranded or, alternatively double stranded, as specified. In some embodiments a nucleic acid can contain portions of both double stranded and single stranded sequence, for example, as demonstrated by forked adapters. A nucleic acid can contain any combination of deoxyribo- and ribo-nucleotides, and any combination of bases, including uracil, adenine, thymine, cytosine, guanine, inosine, xanthanine, hypoxanthanine, isocytosine, isoguanine, and base analogs such as nitropyrrole (including 3-nitropyrrole) and nitroindole (including 5-nitroindole), etc. A "template nucleic acid" is a nucleic acid to be detected or sequenced using any sequencing method disclosed herein. As used herein, a "primed template nucleic acid" (or alternatively, "primed template nucleic acid molecule") is a template nucleic acid primed with (i.e., hybridized to) a primer, wherein the primer is an oligonucleotide having a 3'-end with a sequence complementary to a portion of the template nucleic acid. The primer can optionally have a free 5'-end (e.g., a portion of the primer being non-hybridized with the template), be fully hybridized to the template or can be continuous with the template (e.g., via a hairpin structure). The primed template nucleic acid includes the complementary primer and the template nucleic acid to which it is bound. Unless explicitly stated, a primed template nucleic acid can have either a 3'-end that is extendible by a polymerase, or a 3'-end that is blocked from extension. In preferred embodiments, genomic DNA fragments, or amplified copies thereof, are used as the target nucleic acid. In other preferred embodiments, mitochondrial or chloroplast DNA is used. Other embodiments are targeted to RNA or derivatives thereof such as mRNA or cDNA. The term "nucleotide sequence" is intended to refer to the order and type of nucleotide monomers in a nucleic acid polymer. A nucleotide sequence is a characteristic of a nucleic acid molecule and can be represented in any of a variety of formats including, for example, a depiction, image, electronic medium, series of symbols, series of numbers, series of letters, series of colors, etc. A series of "A," "T," "G," and "C" letters is a well-known sequence representation for DNA that can be correlated, at single nucleotide resolution, with the actual sequence of a DNA molecule. A similar representation is used for RNA except that "T" is replaced with "U" in the series. A "nucleotide" is a molecule that includes a nitrogenous base, a five-carbon sugar (ribose or deoxyribose), and at least one phosphate group. The term embraces, but is not limited to, ribonucleotides, deoxyribonucleotides, nucleotides modified to include exogenous labels or reversible terminators, and nucleotide analogs. The test nucleotide is preferably a native nucleotide. A "native" nucleotide refers to a naturally occurring nucleotide that does not include an exogenous label (e.g., a luminescent dye, or other label) or chemical modification such as may characterize a nucleotide analog. The term "dNTP" refers to any deoxyribonucleotide triphosphate, and a dNTP for use in the disclosed method may comprise a native nucleotide. Examples of native nucleotides that may be used as a test nucleotide in the disclosed methods include: dATP (2'-deoxyadenosine-5'-triphosphate); dGTP (2'-deoxyguanosine-5'-triphosphate); dCTP (2'-deoxycytidine-5'-triphosphate); dTTP (2'-deoxythymidine-5'-triphosphate); and dUTP (2'-deoxyuridine-5'-triphosphate). The test nucleotide may be nucleotide analog. A "nucleotide analog" has one or more modifications, such as chemical moieties, which replace, remove and / or modify any of the components (e.g., nitrogenous base, five-carbon sugar, or phosphate group(s)) of a native nucleotide. Nucleotide analogs may be either incorporable or non-incorporable by a polymerase in a nucleic acid polymerization reaction. Optionally, the 3'-OH group of a nucleotide analog is modified with a moiety. The moiety may be a 3' reversible or irreversible terminator of polymerase extension. The base of a nucleotide may be any of adenine, cytosine, guanine, thymine, or uracil, or analogs thereof. Optionally, a nucleotide has an inosine, xanthine, hypoxanthine, isocytosine, isoguanine, nitropyrrole (including 3-nitropyrrole) or nitroindole (including 5-nitroindole) base. Nucleotides may include, but are not limited to, ATP, UTP, CTP, GTP, ADP, UDP, CDP, GDP, AMP, UMP, CMP, GMP, dATP, dTTP, dUTP, dCTP, dGTP, dADP, dTDP, dCDP, dGDP, dAMP, dTMP, dCMP, and dGMP. Nucleotides may also contain terminating -7- inhibitors of DNA polymerase, dideoxynucleotides or 2',3' dideoxynucleotides, which are abbreviated as ddNTPs (ddGTP, ddATP, ddTTP, ddUTP and ddCTP). The "next correct nucleotide" (also referred to as the "cognate" nucleotide) refers to the nucleotide type that will bind and / or incorporate at the 3' end of a primer to complement a base in a template strand to which the primer is hybridized. The base in the template strand is referred to as the "next template nucleotide" and is immediately 5' of the base in the template that is hybridized to the 3' end of the primer. The next correct nucleotide can be, but need not necessarily be, capable of being incorporated at the 3'end of the primer or 3'end of the nascent growing strand. A nucleotide having a base that is not complementary to the next template base is referred to as an "incorrect" (or "non-cognate") nucleotide. A "marked nucleotide" refers to a nucleotide conjugated to any marker (e.g., a fluorophore) by a linker, wherein the nucleotide may or may not be incorporated into the primed strand. The terms "label" and "marker" may be used interchangeably to refer to any group or moiety that may be used to identify, detect, and / or distinguish between nucleotides. A label may be a luminescent label, in which case, it may be referred to as an "emitter". A "polymerase" refers to any nucleic acid synthesizing enzyme, including but not limited to, DNA polymerase, RNA polymerase, reverse transcriptase, primase and transferase. Typically, the polymerase includes one or more active sites at which nucleotide binding and / or catalysis of nucleotide polymerization may occur. The polymerase may catalyze the polymerization of nucleotides to the 3'-end of a primer bound to its complementary nucleic acid strand. For example, a polymerase can catalyze the addition of a next correct nucleotide to the 3' oxygen of the primer via a phosphodiester bond, thereby chemically incorporating the nucleotide into the primer. The term "providing", as used, for example, in relation to a test nucleotide, a marked nucleotide, a template, a primer, or a primed template nucleic acid, refers to the preparation and delivery of one or many of the relevant reagents, for example to a reaction mixture or reaction chamber. The term "contacting" refers to the mixing together of reagents (e.g., mixing a primed template nucleic acid molecule with a reaction mixture that comprises a polymerase and the test nucleotide) so that a physical binding reaction or a chemical reaction may take place. The term, "incorporating" or "chemically incorporating" refers to the inclusion of the cognate nucleotide, for example, by correct base pairing with the corresponding base in the template strand, or by attachment to the primer by formation of a phosphodiester bond. Accordingly, the term "incorporating" refers to the process of joining a nucleotide to the 3'-end of a primer or nascent strand by formation of a phosphodiester bond. Thus, the incorporation of a nucleotide at the 3' end of the primer or nascent strand leads to extension of the primer or nascent strand. The incorporated nucleotide thereby provides the 3' end of the primer or nascent strand in the subsequent sequencing cycle. The 3' end of the primer or nascent strand thereby advances by one position along the template strand in each sequencing cycle. The terms "primer" and "nascent strand" may be interchangeably to refer to an oligonucleotide having a 3'-end with a sequence complementary to a portion of the template nucleic acid. As used herein, "extension" refers to the process in a polymerase enzyme catalyzes addition of one or more nucleotides at the 3'-end of the primer or nascent strand, thereby leading to extension of the primer or nascent strand. In some embodiments, the sequencing method may comprise sequencing-by-synthesis (SBS) method. In some embodiments, a SBS method may comprise four steps: 1. library preparation; 2. cluster generation; 3. sequencing; and 4. data analysis. 1. Library Preparation Library preparation is a molecular biology protocol that converts a nucleic acid template, such as a genomic DNA sample, or cDNA sample, into a sequencing library, which can then be sequenced, for example, using a Next Generation Sequencing (NGS) instrument. A target nucleic acid sample can, in some embodiments, be processed prior to performing other modifications. For example, a target nucleic acid sample can be amplified prior to attaching to a bead, or prior to attaching to the surface of a solid support. Amplification is particularly useful when samples are in low abundance or when small amounts of a target nucleic acid are provided. Methods that amplify the vast majority of sequences in a genome are referred to as "whole genome amplification" methods. Examples of such methods include multiple displacement amplification (MDA), strand displacement amplification (SDA), or hyperbranched strand displacement amplification, each of which can be carried out using degenerate primers. Particularly useful methods are those that are used during sample preparation methods recommended by commercial providers of whole genome sequencing platforms (e.g., Illumina Inc., San Diego and Life Technologies Inc., Carlsbad). The sequencing library may be prepared by random fragmentation of the nucleic acid sample. The term "fragment," when used in reference to a first nucleic acid, is intended to mean a second nucleic add consisting of a part or portion of the sequence of the first nucleic acid. In some embodiments, fragmentation inherently results from amplification, for example, in cases where the portion of the template that occurs between sites where flanking primers hybridize is selectively copied. In other embodiments, fragmentation may be achieved using chemical, enzymatic or physical techniques known in the art. Fragments in a desired size range can be obtained using separation methods known in the art such as gel electrophoresis. Fragmentation can be carried out to obtain template nucleic acid fragments that have a minimum size of at least about 0.1 kb, 0.5 kb, 1 kb, 2 kb, 3, kb, 4 kb, 5 kb, 10 kb or longer in length. Adapters, which may be referred to as "library adapters" may be ligated to the template fragments, such as, for example, ligation of 5' and 3' adapters to each DNA fragment. "Tagmentation" may be used to combine the fragmentation and ligation reactions into a single step that may increase the efficiency of the library preparation process. Adapter-ligated fragments may be amplified and purified by any suitable method currently used in the art. For example, adapter-ligated fragments may be PCR amplified and gel purified. The fragments that are produced from one or more nucleic acid templates can be captured randomly at locations on a solid support surface. Solid supports can be two-or three-dimensional and can be a planar surface (e.g., a glass slide) or can be shaped. Useful materials include glass (e.g., controlled pore glass (CPG)), quartz, plastic (such as polystyrene (low cross-linked and high crosslinked polystyrene), polycarbonate, polypropylene and poly(methylmethacrylate)), acrylic copolymer, polyamide, silicon, metal (e.g., alkanethiolate-derivatized gold), cellulose, nylon, latex, dextran, gel matrix (e.g., silica gel), polyacrolein, or composites. Suitable three-dimensional solid supports include, for example, spheres, microparticles, beads, membranes, slides, plates, micromachined chips, tubes (e.g., capillary tubes), microwells, microfluidic devices, channels, filters, or any other structure suitable for anchoring a nucleic acid. Solid supports can include planar microarrays or matrices capable of having regions that include populations of nucleic acids or primers. Examples include nucleoside-derivatized CPG and polystyrene slides; derivatized magnetic slides; polystyrene grafted with polyethylene glycol, and the like. A solid support to which nucleic acids may be attached in the sequencing method have a continuous or monolithic surface. Thus, fragments can attach at spatially random locations wherein the distance between nearest neighbor fragments (or nearest neighbor clusters derived from the fragments) may be variable. The resulting arrays may have a variable or random spatial pattern of features. Different template fragments that are at different sites of an array can be differentiated from each other according to the locations of the sites in the array. An individual site of an array can include one or more molecules of a particular type. For example, a site can include a single target nucleic acid molecule having a particular sequence or a site can include several nucleic acid molecules having the same sequence (and / or complementary sequence, thereof). The sites of an array can be different features or locations on the same substrate. Exemplary sites include, for example, wells in a substrate, beads (or other particles) in or on a substrate, projections from a substrate, ridges on a substrate or channels in a substrate. The sites of an array can be separate substrates each bearing a different molecule. Exemplary arrays in which separate substrates are located on a surface include, for example, those having beads in wells. The disclosed methods can advantageously use arrays having a high density of features such as, for example, at least about 10 features / cm2, 100 features / cm2, 500 features / cm2, 1,000 features / cm2, 5,000 features / cm2, 10,000 features / cm2, 50,000 features / cm2, 100,000 features / cm2, 1,000,000 features / cm2, 5,000,000 features / cm2, 107 features / cm2, 5xl07 features / cm2, 108 features / cm2, 5xl08 features / cm2, 109 features / cm2, 5xl09 features / cm2, or higher. Flow cells provide a convenient format for housing an array of nucleic acid fragments for use in the disclosed methods. As used herein, the term "flow cell" is intended to mean a chamber having a surface across which one or more fluid reagents can be flowed. Generally, a flow cell will have an ingress opening and an egress opening to facilitate flow of fluid. Flow cells provide a convenient format for use in the disclosed method that involves repeated delivery of reagents in cycles. For example, to initiate a first SBS cycle, one or more dNTPs, DNA polymerase, etc., can be flowed into / through a flow cell that houses an array of nucleic acid fragments. Washes can easily be carried out in the flow cell between the various delivery steps. The cycle can be repeated n times to extend the primer by n nucleotides, thereby detecting a sequence of length n. For cluster generation, the library of adapter-ligated template fragments may be loaded into a flow cell where fragments are captured on a lawn of surface-bound binding molecules, such as oligonucleotides complementary to the library adapters. DNA nanoballs can also be used in the disclosed methods and methods for preparing and using DNA nanoballs for genomic sequencing are known in the art. Briefly, following genomic DNA fragmentation consecutive rounds of adaptor ligation, amplification and digestion results in head to tail concatamers of multiple copies of the circular genomic DNA template / adaptor sequences which are circularized into single stranded DNA by ligation with a circle ligase and rolling circle amplified. The adaptor structure of the concatamers promotes coiling of the single stranded DNA thereby creating compact DNA nanoballs. The DNA nanoballs can be captured on substrates, preferably to create an ordered or patterned array such that distance between each nanoball is maintained thereby allowing sequencing of the separate DNA nanoballs. Sequencing utilizing the methods and compositions described herein can also be performed in a microtiter plate, for example in high density reaction plates or slides. For example, genomic targets can be prepared by emPCR technologies. Reaction plates or slides can be created from fiber optic material capable of capturing and recording light generated from a reaction, for example from a luminescent reaction. The core material can be etched to provide discrete reaction wells capable of holding at least one emPCR reaction bead. Such slides / plates can contain over a 1.6 million wells. The created slides / plates can be loaded with the target sequencing reaction emPCR beads and mounted to an instrument where the sequencing reagents are provided and sequencing occurs. 2. Cluster Generation Cluster generation is a process of clonal amplification of target nucleic acid templates which may be used or required, for example, for imaging systems which cannot detect single luminescence events. Various suitable methods for clonally amplifying nucleic acid template molecules to produce clusters of cloned templates will be known to the skilled person. Any suitable method may be used, and typically, these methods comprise polymerase chain reaction (PCR)-based techniques. The cluster generation procedure is relatively complicated and time-consuming, and may introduce errors into the cloned template nucleic acids. In the cluster generation step, each template fragment is clonally amplified into distinct clusters. The result is a clonal grouping of identical template fragments bound to the surface of the flow cell. Each cluster on the flow cell produces a single sequencing read. For example, 10,000 clusters on the flow cell would produce 10,000 sequence reads. Where paired-end reads are implemented, a second read of each sequence would also be performed. Each cluster is seeded by a single template nucleic acid fragment and is clonally amplified, for example, using a PCR-based approach, such as involving the use of forward and reverse primers that are attached to the support within the flow cell. In current sequencing methods, a typical cluster has in the order of 1000 copies. To enable the generation of defined clusters, the template fragments may be captured onto surfaces that are patterned for example, with embedded beads (typically 1-2 pm in diameter) or wells (typically 200-600 nm in diameter). Each bead or well only captures a single template fragment and the size of the bead or well defines the maximum size of the cluster. The structured organization provided by the patterned surface of the flow cell provides improved, regular spacing of template clusters, and increased cluster density, which provides advantages over non-pattered clusters, such as in relation to signal detection. Bridge amplification may be used to generate clusters. Bridge amplification may occur on the surface of the flow cell. For example, in currently used methods, the surface of the flow cell is coated with a "lawn" of oligonucleotides. In the first step of bridge amplification, a single-stranded sequencing library (with complementary adapter ends) is loaded into the flow cell. Individual molecules in the library bind to complementary oligos as they "flow" across the oligo lawn. Priming occurs as the opposite end of a ligated fragment bends over and "bridges" to another complementary oligo on the surface. Repeated denaturation and extension cycles (similar to PCR) result in localized amplification of single molecules into millions of unique, clonal clusters across the flow cell. Other suitable amplification methods known in the art can also be used to produce immobilized amplicons from immobilized nucleic acid fragments. For example, one or more clusters can be formed via solid-phase PCR, solid-phase MDA, solid-phase RCA etc., whether one or both primers of each pair of amplification primers are immobilized. 3. Sequencing In some embodiments, the disclosed method comprises the detection of single nucleotides as they are incorporated at the 3' end of the primer or 3' end of the nascent growing strand. In some embodiments, nucleotides are added to a nucleic acid primer thereby extending the primer in a template-dependent manner. Detection of the order and type of nucleotides added to the primer can be used to determine the sequence of the template. At least some of the nucleotides are labelled with a luminescent marker which may be used to detect and identify the nucleotide. During each sequencing cycle, a nucleotide which is the next complementary nucleotide (i.e., comprises a base complementary to the next base of the template strand), is incorporated into the nucleic acid chain in a template-dependent manner due to complementary hydrogen bonding with the corresponding nucleotide in the template fragment. A polymerase enzyme may subsequently catalyze the chemical addition of the next complementary nucleotide into the nucleic acid chain. In each sequencing cycle, the next complementary nucleotide is identified by irradiation of the nucleotide and detection of any luminescence attributable to a luminescent marker that the nucleotide is labelled with. A plurality of different species of test nucleotide may be included in the reaction mixture in a single sequencing cycle. Thus, labels associated with each different nucleotide species may be distinguishable and identifiable, for example on the basis of different luminescence emission spectra. In these embodiments, the luminescence emission that is detected is used to determine which of the different species of test nucleotides is incorporated into the primed strand and, thus represents the next complementary nucleotide. In some embodiments, one or more of the nucleotides may be pre-labelled with the relevant marker prior to inclusion in the reaction mixture and / or prior to incorporation into the primer or nascent strand of the nucleic acid that is being sequenced. In other embodiments, one or more of the nucleotides may be labelled after correctly base pairing with the relevant nucleotide in the template strand, and / or after incorporation into the growing strand. For example, one of more of the nucleotides may comprise a moiety, such as a linker, to which a detectable label may be directly or indirectly attached to thereby detect and identify the nucleotide. In some sequencing methods, each nucleotide is associated with a different luminescent label that may be used to detect and identify the nucleotide. Thus, the number of possible different nucleotides to be detected is equal to the number of different labels used, and the number of detection channels required. At least one of the luminescent labels comprise a luminescent perovskite, preferably a luminescent perovskite having a FWHM (full width at half maximum) of no more than 40 nm, preferably no more than 30 nm. Suitably, 2, 3 or 4 different luminescent labels are used. At least one of the luminescent labels comprises a particulate luminescent perovskite. In some embodiments, the luminescent labels may further include one or more nonperovskite luminescent labels. Known non-perovskite luminescent labels include organic luminescent labels such as 4\6-diamid;no-2-phenylindole (DAPI), fluorescein isothiocyanate (FITC) and AlexaFluor 647. However, these organic fluorophores commonly have broad emission spectra - for example a FWHM of 50 nm or more - and overlap in emission spectra of these organic fluorophores may lead to errors in determining which nucleotide has been incorporated into a strand. Accordingly, it is preferred that each luminescent label comprises a luminescent perovskite, more preferably a luminescent perovskite having a FWHM of no more than 40 nm, preferably no more than 30 nm. In some sequencing methods, the number of possible nucleotides is the same as the number of different labels. In some sequencing methods, the number of possible different nucleotides is greater than the number of different labels used and / or the number of detection channels required. In some embodiments, one of the nucleotides is not labelled and all other nucleotides are labelled. The presence of the unlabeled nucleotide may be determined if no luminescence arising from a luminescent marker is detected. In some embodiments, a first nucleotide may be labelled with an optionally cleavable luminescent marker. In some embodiments, an additional step may be employed to cleave the optionally cleavable luminescent marker from the first nucleotide and attach the optionally cleavable luminescent marker to a second nucleotide. In some sequencing methods, emissions from at least two different luminescent markers are detectable by the same detector. In some embodiments, a single nucleotide may be labelled with a first and second luminescent marker, wherein the first and second luminescent markers can be excited by different laser frequencies to produce emissions having first and second peak wavelengths wherein the first and second peak wavelengths are detectable by the same or different detectors. In these embodiments, a single nucleotide may be labelled with a first and second luminescent marker, excitable by different frequencies by detectable by the same or different detectors. In any event, at the point of detection, such as for example, when electromagnetic radiation at an excitation wavelength is applied to an emitter, the nucleotide that is being detected comprises the label, in the sense that the label (or absence of a label) is associated with the nucleotide in such a way that it may be used to specifically detect and identify the nucleotide. The nucleotide may comprise a "terminator" which may be a "reversible terminator". A nucleotide having a terminator or reversible terminator moiety can be used such that subsequent extension cannot occur until a deblocking agent is delivered to remove the terminator moiety. Thus, after nucleotide incorporation and identification, the terminator may be removed, such as by enzymatic cleavage, to allow the next sequencing cycle to commence. The linker attaching the label to the nucleotide in the disclosed method may be cleavable or otherwise arranged to allow dissociation of the label from the nucleotide. Thus, after incorporation and identification of the test nucleotide, the emitter or other label may be removed to allow the next sequencing cycle to commence and the next complementary nucleotide to be identified. The result is base-by-base sequencing of the template fragment nucleic acids. The term "linker" is intended to mean a chemical bond or moiety that bridges two moieties, for example by covalent linkage or the formation of a stable complex. A linker can be, for example, the sugar-phosphate backbone that connects nucleotides in a nucleic acid moiety. In the disclosed method, a linker may be used to conjugate a test nucleotide to a luminescent material. The linker can include, for example, one or more of a nucleotide moiety, a nucleic acid moiety, a non-nucleotide chemical moiety, a nucleotide analogue moiety, amino acid moiety, polypeptide moiety, or protein moiety. The terms "cleave", "cleavage site", and similar terms, refer to a moiety in a molecule, such as a linker, that can be modified or removed to physically separate two other moieties of the molecule. In some embodiments of the disclosed method, a linker may be used to associate the test nucleotide to a terminator. The linker may be a cleavable linker, such that at the appropriate point in the sequencing cycle the linker is cleaved, thereby removing the terminator from the nucleotide. In some embodiments of the disclosed method, the linker that is used to associate the test nucleotide and terminator comprises the same type of cleavable linkage that is present in the linker conjugating the test nucleotide to the fluorophore or other detectable label. Thus, at the appropriate point in the sequencing cycle, a single agent, such as a single type of cleaving enzyme, may be used to remove both the terminator and label from the test nucleotide. In some embodiments of the disclosed method, a single linker may be arranged to conjugate both the terminator and label to the test nucleotide. In some embodiments of the disclosed method, the linker that is used to associate the test nucleotide and terminator may comprise a different type of cleavable linkage to that present in the linker conjugating the test nucleotide to the label. In some embodiments, a plurality of different nucleic acid fragments can be 5 sequenced simultaneously under conditions where events occurring for different templates can be distinguished, for example due to being present at different locations in an array. 4. Data Analysis During data analysis, the newly identified sequence reads of the template fragment io are aligned, and the target nucleic acid sequence may thus be determined. Following alignment, many variations of analysis are possible, including single nucleotide polymorphism (SNP), insertion-deletion (indel) identification, and read counting for RNA methods, phylogenetic or metagenomic analysis. The skilled person will appreciate that the disclosed methods are not limited to 15 sequencing-by-synthesis methods, and may be applied to various other methods involving the use of luminescence to identify a nucleotide. Such methods include, for example, methods for the analysis of short tandem repeat (STR.) markers, single nucleotide polymorphisms (SNPs), methylation patterns, ChIP analysis, and RNA transcription. Also included are methods of analysis of any nucleic acid template, 20 including, for example, DNA from any organism or mixed population such as a microbiome, whole genomes, RNA transcripts for expression analysis, cancer samples (such as methods of analysing somatic variants and / or tumour subclones), and mitochondrial DNA. Signal States 25 The disclosed methods may include, for each imaging event, correlating one or more nucleotide species to a dark state, correlating one or more nucleotide species to a signal state, correlating one or more nucleotide species to a grey state, and / or correlating one or more nucleotide species to a change in state between two imaging events (such as first and second imaging events), between a dark state, a grey state 30 or a signal state. A "signal state," when used in reference to a detection event, means a condition in which a specific signal is produced in the detection event. For example, a nucleotide can be in a signal state and detectable when attached to a luminescent label that is excited at a specific excitation wavelength and detected by emission in an emission detection step in a sequencing method, using a specific detection filter. The term "dark state," when used in reference to a detection event, means a condition in which a specific signal is not produced in the detection event. For example, a nucleotide can be in a dark state when the nucleotide does not emit above a threshold level in an emission detection step in a sequencing method, using a single detection filter. For example, a nucleotide may lack a luminescent label, and / or may be attached to a luminescent label that is excited at a specific excitation wavelength that is different to the excitation wavelength used in relation to that particular detection event. Dark state detection may also include any background emission which may be present in the absence of a luminescent label that is excited at the specific excitation wavelength being used in relation to that particular detection event. For example, some reaction components, which may include luminescent labels that are excited at a different excitation wavelength to the specific excitation wavelength being used, may demonstrate minimal emission in response to the excitation wavelength being used. As such, there may be background emission from such components. Further, background emission may be due to light scatter, for example from adjacent sequencing reactions, which may be detected by a detector. In addition, "dark state" can include background emission produced when an emissive moiety is not specifically included, such as when a nucleotide lacking a luminescent label is used. However, such background emission is distinguishable from a signal state or a grey state and, as such, nucleotide incorporation of an unlabeled nucleotide (or "dark" nucleotide) is still discernible. Likewise, nucleotide incorporation of a nucleotide that is attached to a luminescent label that is excited at a specific excitation wavelength that is different to the excitation wavelength used in relation to that particular detection event can also be distinguished from a signal state or a grey state. The term "grey state," when used in reference to a detection event, means a condition in which an attenuated signal is produced in the detection event. For example, a population of nucleotides of a particular type can be in a grey state when a first subpopulation of the nucleotides attached to a first luminescent label that is detected in a luminescence detection step of a sequencing method, while a second subpopulation of the nucleotides lacks the first luminescent label and does not give emission that is specifically detected in the luminescence detection step when excited at the excitation wavelength of the first luminescent label. The second subpopulation of the nucleotides, may, for example, comprise a second luminescent label that may be detected when excited at an excitation wavelength that is different to the excitation wavelength of the first luminescent label, such that the first and second luminescent labels may be detected in the same detector but are distinguishable on the basis that they are excited using different excitation wavelengths. Typically, a reaction cycle will be carried out by delivering all four nucleotide types to a nucleic acid sample in the presence of a polymerase, for example a DNA or RNA polymerase, during a primer extension reaction. The presence of four nucleotide types provides an advantage of increasing polymerase fidelity compared to the use of fewer than four nucleotide types. In some embodiments, fewer than four different types of nucleotides can be present during a polymerase extension reaction. The disclosed methods are not limited to nucleic acid sequencing and may also be used in other applications where detection of more than one analyte (i.e., nucleotide, protein, or fragments thereof) in a sample is desired. It will be understood that reference to the use of a luminescent label includes the use of a plurality of different luminescent labels having the same or similar excitation and emission properties. In preferred embodiments, the disclosed method is performed on a substrate (reaction surface), such as a glass, plastic, semiconductor chip or composite derived substrate, for example, within a flow cell. Sequencing may be in a multiplex format, wherein multiple nucleic acid targets are detected and sequenced in parallel, for example in a flow cell or array type of format. The disclosed method is particularly advantageous when practicing parallel sequencing or massive parallel sequencing. For purposes of illustration and not intended to limit embodiments, a general strategy sequencing cycle can be described by a sequence of steps. The following example is based on a sequence by synthesis sequencing reaction, however the disclosed methods are not limited to any particular sequencing reaction methodology. The four nucleotide types A, C, T and G, typically modified nucleotides designed for sequencing reactions such as reversibly blocked (rb) nucleotides (e.g., rbA, rbT, rbC, rbG) wherein three of the four types of nucleotide are luminescently labelled, are simultaneously added, along with other reaction components, to the reaction surface (e.g., flow cell, chip, slide, etc.). Following incorporation of a nucleotide into a growing sequence nucleic acid chain based on the target sequence, a first specific excitation wavelength, which is capable of selectively exciting a first emitter that is associated with a first of the four possible nucleotides is provided to the reaction surface. Any resulting emission is recorded using a specific detection filter; this constitutes a first imaging event and a first luminescence detection pattern. Following the first imaging event, a second specific excitation wavelength, which is capable of selectively exciting a second emitter that is associated with a second of the four possible nucleotides is provided to the reaction surface. Any resulting emission is recorded using a specific detection filter, and this constitutes a second imaging event and a second luminescence detection pattern. In the same way, third and fourth imaging events and luminescence detection patterns are subsequently produced in respect of the third and fourth of the possible nucleotides. In some embodiments, the same excitation wavelength may be capable of exciting two or more different luminescent labels associated with different nucleotides. In particular, two different luminescent perovskites may be selected which have overlapping absorption spectra at an excitation wavelength but little or no overlap in emission spectra, preferably no emission overlap at half maximum, more preferably no emission overlap at 25% or 10% of emission maximum. Accordingly, the required number of different excitation wavelengths may be fewer than the number of different luminescent labels. The sequence of the target nucleic acid, for that particular cycle (i.e., the identity of the next cognate residue) is determined by identifying the nucleotide that has been incorporated from the four luminescence detection patterns. Specifically, one of the luminescence detection patterns will comprise a signal state and three of the luminescence detection patterns will comprise a dark state. Various reagents present after the fourth imaging event are washed away in preparation for the next sequencing cycle. Exemplary chemical reagents that may be removed include, but are not limited to, blocking agents, luminescent labels, quenchers, cleavage reagents, or any other reagents that may directly or indirectly 5 cause an identifiable and measurable change in luminescence, or which may inhibit the incorporation of the next correct nucleotide in the sequence. Luminescent markers The luminescent markers used in the process described herein include at least one particulate marker comprising a luminescent perovskite. 10 Luminescent perovskites as described herein include, without limitation: - compounds of formula ABX3 wherein A is a cation; B is a metal; and X in each occurrence is independently an anion; and - compounds of formula A'2B'B"O6 wherein A' is a divalent cation; and B' and B" are each a metal. 15 Preferably, A is an alkali metal or ammonium cation, more preferably an alkali metal selected from Li+, Na+, K+ and Cs+, most preferably Cs+. Ammonium cations may have formula N(R1)4+ wherein R1 in each occurrence is independently selected from H and a C1-12 hydrocarbyl group. A C1-12 hydrocarbyl group as described anywhere herein may be selected from Ci-12 alkyl and phenyl which is optionally substituted 20 with one or more Ci 6 alkyl groups. Preferably, B is Pb or Sn. Preferably, X in each occurrence is a halide anion, preferably F-, Cl~, Br or T, more preferably Cl-, Br or T. In some embodiments, each X is the same. In some embodiments, the perovskite contains two different anions X. 25 A' is preferably an alkali earth metal cation, more preferably Be2+, Mg2+, Ca2+, Sr2+ or Ba2+. Exemplary perovskites are disclosed in S. Ghosh et al, "Recent developments of lead-free halide double perovskites: a new superstar in the optoelectronic field", Material Advances, Issue 9, 2022 and S Vasala, "A2B'B"O6 perovskites: A review", Progress in Solid State Chemistry, Volume 43, Issues 1-2, May 2015, Pages 1-36 The luminescent perovskite may emit light having a peak wavelength in the range of 350-1000 nm. Emission of a perovskite may be tuned across a wide wavelength range by any method known to the skilled person, for example by selection of anions X and mixtures of anions X such as is described in L. Protescu et al, "Nanocrystals of Cesium Lead Halide Perovskites (CsPbXs, X = Cl, Br, and I): Novel Optoelectronic Materials Showing Bright Emission with Wide Color Gamut", Nano Lett. 2015, 15, 3692-3696, the contents of which are incorporated herein by reference. A blue luminescent marker as described herein may have a photoluminescence spectrum with a peak of no more than 500 nm, preferably in the range of 400-500 nm, optionally 400-490 nm. A green luminescent marker as described herein may have a photoluminescence spectrum with a peak of more than 500 nm up to 580 nm, optionally more than 500 nm up to 540 nm. A red luminescent marker as described herein may have a photoluminescence spectrum with a peak of no more than more than 580 nm up to 950 nm, optionally up to 630 nm, optionally 585 nm up to 625 nm. The luminescent marker may have a shift between excitation and emission maxima in the range of 20-400 nm. Photoluminescence spectra of luminescent markers as described herein may be measured in methanol suspension using a Jobin Yvon Horiba Fluoromax-3. UV / vis absorption spectra of luminescent markers as described herein may be as measured in methanol suspension using a Cary 5000 UV-vis-IR spectrometer. In some embodiments, the particulate marker comprising a luminescent perovskite consists of the luminescent perovskite. In some embodiments, the surface of the luminescent perovskite particle is bound directly to a linking group for binding the particle to a nucleotide. In some embodiments, the particulate marker comprises a core comprising or consisting of the luminescent perovskite and a shell. The shell at least partially, and optionally completely, encloses the core. The shell may comprise or consist of an organic or inorganic material. An exemplary 5 organic shell material is an organic polymer. An exemplary inorganic shell material is an inorganic oxide. Preferably, the shell comprises or consists of silica, alumina or a combination thereof. A silica shell may be formed by reacting a solution of a compound that forms silica in the presence of particulate perovskite. 10 Optionally, the silica precursor is an alkoxysilane, preferably a trialkoxy or tetraalkoxysilane, optionally a C1-12 trialkoxy or tetra-alkoxysilane, for example tetraethyl orthosilicate. The silica precursor may be substituted only with alkoxy groups or may be substituted with one or more groups. The reaction may take place in the presence of an acid or a base. 15 Optionally, the solvent of the solution is selected from water, one or more C1-8 alcohols or a combination thereof. Optionally, at least 50 % of total weight of the core-shell particles core consists of the luminescent perovskite particles. The shell may comprise one or more surface groups bound thereto, preferably 20 covalently bound thereto. Surface groups may be selected from linking groups for binding the luminescent marker to a nucleotide; and stabilising groups. The shell may at least partially isolate the luminescent perovskite from the surrounding environment. This may limit any undesirable interactions between the luminescent perovskite and the external environment. 25 Linking groups may be formed at the surface of the luminescent marker by reaction of a compound of formula (X) at the surface of the luminescent marker: RG1-L-RG2 wherein RG1 is a reactive group capable of reacting with the surface of the luminescent marker particle; L is a linker group; and RG2 is a group capable of binding to a nucleotide or a functionalised nucleotide. In the case where the luminescent marker comprises a silica shell, RG1 is preferably a group of formula -SiR23 wherein R2 in each occurrence is a leaving group, preferably a halogen, more preferably Cl, Br or I. L may be any cleavable linker comprising a cleavable group known to the skilled person, for example as disclosed in G Leriche et al, "Cleavable linkers in chemical biology", Bioorganic &Medicinal Chemistry, Volume 20, Issue 2, 15 January 2012, Pages 571-582, the contents of which are incorporated herein by reference. Cleavable groups of a cleavable linker include, without limitation: - an enzymatically cleavable group, for example an ester, a peptide or a glycoside; - a disulfide; - an ethylene glycol group, preferably an ethylene glycol group substituted with an azide; and a photocleavable group.RG2 may be any group known to the skilled person for reaction at the 3'-OH group or modified 3'-OH group of a nucleotide. Preferably, the luminescent particles as described herein have a number average diameter of no more than 50 nm in methanol as measured by dynamic light scattering (DLS) using a Malvern Zetasizer Nano ZS, optionally no more than 25 nm or no more than 15 nm. In some embodiments, the luminescent markers include at least one particulate marker comprising a luminescent perovskite and one or more further markers. Further markers may be selected from organic luminescent markers and inorganic non-perovskite markers. Exemplary inorganic non-perovskite markers include non-perovskite light-emitting quantum dots, for example metal chalcogenide quantum dots. Exemplary organic fluorescent materials include, without limitation, non-polymeric organic fluorescent materials and light-emitting polymers. Non-polymeric organic fluorescent materials include, without limitation, fluorescein, fluorescein isothiocyanate (FITC); fluorescein NHS; Alexa Fluor 488; Dylight 488; Oregon green; DAF-FM; 6-FAM; 2,7-dichlorofluorescein; 3'-(p-aminophenyl)fluorescein; 3'-(hydroxyphenyl)fluorescein; rhodamines, for example 5 Rhodamine 6G and Rhodamine 110 chloride; coumarins; boron-dipyrromethenes (BODIPYs); naphthalimides; perylenes; benzanthrones; benzoxanthrones; benzothiooxanthrones; 2-(4-pyridyl)-5-phenyl-oxazole; 2-quinolinyl-5-phenyl-oxazole; 2-(4-pyridyl)-5-naphthyl-oxazole; 2-(4-pyridyl)-5-phenyl-thiazole; 2- quinolinyl-5-phenyl-thiazole; 2-(4-pyridyl)-5-naphthyl-thiazole; 2-(4-pyridyl)-5-io phenyl-thiophene; 2-quinolinyl-5-phenyl- thiophene; 2-(4-pyridyl)- 5-naphthyl-thiophene and salts thereof, each of which may be unsubstituted or substituted with one or more substituents. Exemplary substituents are ionic or non-ionic substituents as described herein, optionally chlorine, alkyl amino; phenylamino; and hydroxyphenyl. 15 Exemplary polymeric fluorescent materials include conjugated fluorescent polymers, for example fluorescent polyfluorenes.
Claims
1. A method of determining whether a first test nucleotide comprises a base complementary to the next base of a template strand immediately downstream of a primer in a primed template nucleic acid molecule, the method comprising: providing a primed template nucleic acid molecule;providing a labelled nucleotide comprising the first test nucleotide labelled with a first luminescent marker comprising a first luminescent material;contacting the primed template nucleic acid molecule with a reaction mixture that comprises a polymerase and the first test nucleotide, to thereby incorporate the first test nucleotide into the primed strand only if the first test nucleotide comprises a base complementary to the next base of the template strand;exciting the first luminescent marker with an excitation light; and detecting light emitted by the first luminescent marker,wherein the detection of emitted light identifies the incorporation of the first test nucleotide into the primed strand, and thereby indicates that the first test nucleotide comprises a base complementary to the next base of the template strand, andwherein the first luminescent marker comprises a particulate perovskite.
2. The method according to claim 1 wherein the particulate perovskite is a compound of formula ABX3 wherein A is a cation; B is a metal; and X in each occurrence is independently an anion.
3. The method according to claim 2 wherein A is an alkali metal or ammonium cation.
4. The method according to claim 2 or 3 wherein B is Sn or Pb.
5. The method according to any one of claims 2-4 wherein each X is a halogen.
6. The method according to claim 5 wherein each X is the same halogen.
7. The method according to claim 5 wherein ABX3 comprises two different halogengroups X.
8. The method according to any one of the preceding claims wherein the luminescent marker comprises a silica shell.
9. The method according to any one of claims 1-8 wherein:the reaction mixture comprises a plurality of test nucleotides including the first test nucleotide comprising the first luminescent marker comprising a particulate perovskite and two or three further different species of test nucleotides;one or more of the two or three further different species of test nucleotides comprise a further luminescent marker which is different from the first luminescent marker;each of the plurality of test nucleotides has an emission characteristic selected from a characteristic light-emission in the case of a test nucleotide comprising a luminescent marker, or no light emission in the case of test nucleotide which does not comprise a luminescent marker;detection of the emission characteristic of one of the plurality of different species of test nucleotides identifies the incorporation of that particular test nucleotide into the primed strand, and thereby indicates that the particular test nucleotide comprises a base complementary to the next base of the template strand.
10. The method according to claim 9 wherein the one or more further luminescent markers include a luminescent marker comprising a particulate perovskite according to any one of claims 2-8 which is different from the particulate perovskite of the first luminescent marker.
11. The method according to claim 10 wherein each of the luminescent markers comprise a particulate perovskite.
12. A marked nucleotide comprising a nucleotide attached to a luminescent marker by a cleavable linker, wherein the luminescent marker comprises a particulate luminescent perovskite.
13. The marked nucleotide according to claim 12 wherein the particulate luminescent perovskite comprises a shell.
14. The marked nucleotide according to claim 13 wherein the shell comprises silica.
Citation Information
Patent Citations
Sequencing method
GB2602063A
Systems and methods for multicolor imaging
US20220214278A1
Primary analysis in next generation sequencing
US20230326065A1
Quality measurement of base calling in next generation sequencing
WO2023230279A1