Systems, methods and compositions for detecting epigenetic modifications of nucleic acids
Patent Information
- Application Number
- JP2023561697
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2022-06-02
- Publication Date
- 2025-06-10
AI Technical Summary
Current methods for detecting epigenetic modifications in nucleic acids, such as methylation, are cumbersome, require significant sample preparation time, and can cause DNA degradation, limiting high-resolution genome-wide analysis.
A method involving transient binding of short oligonucleotides to nucleic acids, monitoring binding kinetics, and using a signature of modification detected on a single molecule to differentiate between modified and unmodified bases, allowing for super-resolution imaging and repeated binding events to confirm the presence of epigenetic modifications.
Enables high-resolution, efficient detection of epigenetic modifications without DNA degradation, reducing sample preparation time and improving the accuracy of genome-wide methylation analysis.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Provisional Patent Application No. 63,171,566, entitled "Detecting Epigenetic Modifications in Nucleic Acids," filed April 6, 2021, which is hereby incorporated by reference in its entirety for all purposes.
[0002] The present disclosure is directed to systems, methods, and compounds for detecting epigenetic modifications of nucleic acids. [Background technology]
[0003] The information for generating functional biological systems is written in the sequence of bases along the length of nucleic acids. The Watson-Crick double helix view of DNA's structure does not take into account epigenetic modifications to the nucleic acid polymer found in living organisms. Many modifications, such as 5-methylcytosine (5mC), 5-hydroxymethylcytosine (5hmC), 5-formylcytosine (5fC), 5-carboxylcytosine (5caC), and N6 methadenine (6mA), have biological roles. Of these modifications, 5-methylcytosine (5mC) is abundant in the human genome and is often referred to as the fifth base. In contrast to the standard bases, it is not replicated by the cellular replication machinery and is therefore not inherited in the normal way.
[0004] In mammals, methylation is most prevalent in the 5'-CpG-3' context. In the double helix, both Cs of a complementary CpG dyad are usually methylated, but a hemimethylated state, where only one C is methylated, is also observed and is thought to represent a transition state between methylation and demethylation. CpG sites (CpG islands) are frequently found upstream of the coding regions of genes and are involved in the regulation of gene activity and tissue-specific transcription.
[0005] In contrast to sequencing the human genome, mapping the human methylome is a more complex task: determining comprehensive, high-resolution genome-wide methylation patterns from a given sample has been challenging due to sample preparation requirements and the need to amplify or modify nucleic acids before sequencing.
[0006] A variety of methods exist for isolating or detecting methylated portions of the genome. The most widely used method involves the treatment of DNA with bisulfite, which converts unmethylated cytosines to uracil but not 5-methylcytosines. The DNA is then amplified (converting all uracils to thymine) and subsequently analyzed by a variety of methods, including microarray-based techniques and second-generation (e.g., Illumina) sequencing. Although bisulfite-based techniques have greatly advanced the analysis of methylated DNA, they also have some drawbacks. First, bisulfite sequencing requires significant sample preparation time. Second, the stringent reaction conditions required for the complete conversion of unmethylated cytosines to uracil cause DNA degradation, thus requiring large starting sample volumes, which may be problematic for some applications. Furthermore, bisulfite sequencing relies on microarray or second-generation DNA sequencing techniques to read out the methylation status and therefore suffers from the same limitations as these methodologies.
[0007] In addition to functional modifications, damage-induced modifications of DNA by various agents lead to genetic mutations. In the case of RNA, it has long been known that tRNAs have a myriad of modifications, but it is now increasingly recognized that RNA in general also has modifications associated with it. Furthermore, nucleic acids can be epigenetically modified non-covalently by the binding of ligands, ranging from metals to DNA-binding proteins.
[0008] In view of the above background, what is needed in the art are improved systems, methods, and compounds for detecting epigenetic modifications of nucleic acids. Summary of the Invention
[0009] The present disclosure addresses the above-disclosed shortcomings by providing systems, methods, and compounds for detecting epigenetic modifications of nucleic acids.
[0010] Thus, one aspect of the present disclosure aims to provide a method for directly determining the presence of epigenetic modifications on nucleic acid molecules, i.e., detection of methylation, hydroxymethylation and other modifications on DNA or RNA.
[0011] In certain aspects of the present invention, a method is provided for detecting modification in nucleic acid molecules.Generally, a sample is provided that contains a nucleic acid sequence that may be modified, and at least one probe that can bind to the nucleic acid sequence is provided.In some embodiments, the nucleic acid is repeatedly bound by the probe, and the kinetics of binding is monitored.
[0012] In some embodiments, detection of modifications in nucleic acid molecules is performed by directly detecting the modifications on a single molecule. In some embodiments, detection is performed via distinct signatures of modifications detected on a single molecule. In some embodiments, the signature is a binding profile of one or more oligonucleotides (oligos) targeted to their respective complementary sequences on the target molecule. In some embodiments, the binding profile includes the degree (e.g., amount, rate, lifetime) of hybridization to the target sequence. In some embodiments, the target molecule is immobilized on a flat surface. In some embodiments, the oligonucleotide binding is transient and repeatable. Transient binding allows for super-resolution imaging (1), and repeated binding provides confidence that a true signal is observed. Here, the binding profile and degree of hybridization include the number of binding events on an individual molecule (number of repeated bindings), the "on" or "dwell" time of these binding events, and the "off" time (time between binding events). In some embodiments, the oligos are short, seven bases or less, typically three to five bases in length. In some embodiments, the oligos are optically labeled (eg, with fluorescent dyes, nanoparticles, or light scattering particles) and the dwell time is the "light" time and the off time is the "dark" time.
[0013] A surprising feature discovered by the present disclosure and forming the basis of an important embodiment of the present invention is that when hybridization is performed under conditions where the oligos bind transiently, the light time, the dark time, and / or the number of repeat bindings can be used to classify the binding sites as unmodified or modified, and to confirm that different modifications fall into different classes. For example, in some embodiments, the dwell time of the oligo binding is longer when the target sequence has a methylation site than when it does not. In some embodiments, the binding profile of multiple oligonucleotides complementary to the sequence around the base (e.g., a single base has five 5mers that include the base in the binding base footprint) is taken into account to determine whether the base is modified. Having multiple oligonucleotides target the same base increases redundancy and the robustness of the measurement. [Brief description of the drawings]
[0014] [Figure 1] FIG. 1 provides a diagram of a method for detecting transient oligonucleotide binding to a target nucleic acid. The nucleic acid is immobilized on a surface, for example, via streptavidin / biotin interactions. An oligonucleotide probe labeled with a fluorescent dye, for example, Cy3, binds to the nucleic acid in a sequence-specific manner. The oligonucleotide binds to the target transiently, but long enough (e.g., 200 ms) for the fluorescent dye to be excited by a laser and the emission detected. Hybridization kinetics are measured over multiple binding and unbinding events. [Diagram 2] Provide an illustrative example of the difference in binding profile kinetics between 5-MeC and unmethylated target DNA. ton is the time at which the oligonucleotide fluorescent signal is detected on the nucleic acid molecule. Toff is the time at which no fluorescent signal is detected on the nucleic acid molecule. [Figure 3A]Collectively, examples of experimentally generated oligo binding profiles for 5-MeC and unmethylated DNA targets in different contexts are shown. Two separate flow cells are used: one containing a target nucleic acid with 5-MeC at a designated site on the sequence, and the other flow cell containing a target nucleic acid of the same sequence but without 5-MeC at the designated site. The same Cy3-labeled oligonucleotide is added to each flow cell and the average ton is determined. The graph displays the average ton for 10 different oligonucleotides, each binding to methylated and unmethylated versions of the target DNA molecule. [Figure 3B] Collectively, examples of experimentally generated oligo binding profiles for 5-MeC and unmethylated DNA targets in different contexts are shown. Two separate flow cells are used: one containing a target nucleic acid with 5-MeC at a designated site on the sequence, and the other flow cell containing a target nucleic acid of the same sequence but without 5-MeC at the designated site. The same Cy3-labeled oligonucleotide is added to each flow cell and the average ton is determined. The graph displays the average ton for 10 different oligonucleotides, each binding to methylated and unmethylated versions of the target DNA molecule. [Figure 3C] Collectively, examples of experimentally generated oligo binding profiles for 5-MeC and unmethylated DNA targets in different contexts are shown. Two separate flow cells are used: one containing a target nucleic acid with 5-MeC at a designated site on the sequence, and the other flow cell containing a target nucleic acid of the same sequence but without 5-MeC at the designated site. The same Cy3-labeled oligonucleotide is added to each flow cell and the average ton is determined. The graph displays the average ton for 10 different oligonucleotides, each binding to methylated and unmethylated versions of the target DNA molecule. [Figure 3D]Collectively, examples of experimentally generated oligo binding profiles for 5-MeC and unmethylated DNA targets in different contexts are shown. Two separate flow cells are used: one containing a target nucleic acid with 5-MeC at a designated site on the sequence, and the other flow cell containing a target nucleic acid of the same sequence but without 5-MeC at the designated site. The same Cy3-labeled oligonucleotide is added to each flow cell and the average ton is determined. The graph displays the average ton for 10 different oligonucleotides, each binding to methylated and unmethylated versions of the target DNA molecule. [Figure 3E] Collectively, examples of experimentally generated oligo binding profiles for 5-MeC and unmethylated DNA targets in different contexts are shown. Two separate flow cells are used: one containing a target nucleic acid with 5-MeC at a designated site on the sequence, and the other flow cell containing a target nucleic acid of the same sequence but without 5-MeC at the designated site. The same Cy3-labeled oligonucleotide is added to each flow cell and the average ton is determined. The graph displays the average ton for 10 different oligonucleotides, each binding to methylated and unmethylated versions of the target DNA molecule. [Figure 3F] Collectively, examples of experimentally generated oligo binding profiles for 5-MeC and unmethylated DNA targets in different contexts are shown. Two separate flow cells are used: one containing a target nucleic acid with 5-MeC at a designated site on the sequence, and the other flow cell containing a target nucleic acid of the same sequence but without 5-MeC at the designated site. The same Cy3-labeled oligonucleotide is added to each flow cell and the average ton is determined. The graph displays the average ton for 10 different oligonucleotides, each binding to methylated and unmethylated versions of the target DNA molecule. [Figure 3G]Collectively, examples of experimentally generated oligo binding profiles for 5-MeC and unmethylated DNA targets in different contexts are shown. Two separate flow cells are used: one containing a target nucleic acid with 5-MeC at a designated site on the sequence, and the other flow cell containing a target nucleic acid of the same sequence but without 5-MeC at the designated site. The same Cy3-labeled oligonucleotide is added to each flow cell and the average ton is determined. The graph displays the average ton for 10 different oligonucleotides, each binding to methylated and unmethylated versions of the target DNA molecule. [Figure 3H] Collectively, examples of experimentally generated oligo binding profiles for 5-MeC and unmethylated DNA targets in different contexts are shown. Two separate flow cells are used: one containing a target nucleic acid with 5-MeC at a designated site on the sequence, and the other flow cell containing a target nucleic acid of the same sequence but without 5-MeC at the designated site. The same Cy3-labeled oligonucleotide is added to each flow cell and the average ton is determined. The graph displays the average ton for 10 different oligonucleotides, each binding to methylated and unmethylated versions of the target DNA molecule. [Figure 3I] Collectively, examples of experimentally generated oligo binding profiles for 5-MeC and unmethylated DNA targets in different contexts are shown. Two separate flow cells are used: one containing a target nucleic acid with 5-MeC at a designated site on the sequence, and the other flow cell containing a target nucleic acid of the same sequence but without 5-MeC at the designated site. The same Cy3-labeled oligonucleotide is added to each flow cell and the average ton is determined. The graph displays the average ton for 10 different oligonucleotides, each binding to methylated and unmethylated versions of the target DNA molecule. [Figure 4A]Collectively, we demonstrate that the addition of a cap such as 3'-Uaq or pyrene allows for the use of the 3' wobble base of an oligo to distinguish between methylated and unmethylated sites. In this example experiment, it is the terminal 3' base of the oligo that base pairs with a methylated cytosine residue in the target molecule. The kinetic profile of a single oligo binding to methylated and unmethylated versions of the same DNA target molecule is shown. [Figure 4B] Collectively, we demonstrate that the addition of a cap such as 3'-Uaq or pyrene allows for the use of the 3' wobble base of an oligo to distinguish between methylated and unmethylated sites. In this example experiment, it is the terminal 3' base of the oligo that base pairs with a methylated cytosine residue in the target molecule. The kinetic profile of the same oligo and DNA target molecule is shown; however, in this example, the oligo is 3'-Uaq capped. [Figure 5A] Examples of experimentally generated kinetic profiles for multiple oligos overlapping a single 5-MeC site are shown together, demonstrating the ability to perform multiple independent reads for each location on a target molecule. Two separate flow cells are used; one containing a nucleic acid with 5-MeC at a designated site on the sequence, and the other flow cell containing a nucleic acid of the same sequence but without 5-MeC at the designated site. Five Cy3-labeled oligonucleotides are added sequentially to each flow cell. The graph displays the average ton for each of the five overlapping oligonucleotides binding to methylated and unmethylated versions of the same target DNA molecule. [Figure 5B]Examples of experimentally generated kinetic profiles for multiple oligos overlapping a single 5-MeC site are shown together, demonstrating the ability to perform multiple independent reads for each location on a target molecule. Two separate flow cells are used; one containing a nucleic acid with 5-MeC at a designated site on the sequence, and the other flow cell containing a nucleic acid of the same sequence but without 5-MeC at the designated site. Five Cy3-labeled oligonucleotides are added sequentially to each flow cell. The graph displays the average ton for each of the five overlapping oligonucleotides binding to methylated and unmethylated versions of the same target DNA molecule. [Figure 6] Collectively, they indicate that in some cases, an oligo may hybridize to a region of a target molecule that contains multiple CpG sites. These CpG sites may be either all unmethylated, all methylated, or a combination of methylated and unmethylated. Two example experiments are shown in which these three contexts are distinguished based on differences in oligo hybridization kinetics. A shows that unmethylated, singly methylated, and doubly methylated DNA target molecules containing the sequence TTCCCG are immobilized on a surface, and the same Cy3-labeled oligonucleotide with the sequence GGGAA is added to each flow cell. B shows that unmethylated, singly methylated, and doubly methylated DNA target molecules containing the sequence CCCGCG are immobilized on a surface, and the same Cy3-labeled oligonucleotide with the sequence GCGGG is added to each flow cell. The graph displays the average ton for unmethylated, singly methylated, and doubly methylated target molecules with each probe. [Figure 7]Collectively, we show the process of methylation haplotype discrimination. When an oligo hybridizes at a position on a target molecule that contains multiple CpG sites, its kinetic profile is determined by the methylation status of both sites (A). In the case of mixed signals (e.g., a kinetic signal generated by hybridization to one methylated and one unmethylated cytosine residue), additional probes overlapping each of the CpG sites can be used to discern which are methylated and which are unmethylated (B). [Figure 8A] Collectively, we demonstrate detection of spike-ins and discrimination of methylation in a mixed background of cell-free DNA. An equal mixture of unmethylated and methylated synthetic single-stranded DNA targets was spiked into synthetic plasma at a high percentage (approximately 50%). DNA was extracted from the mixture, biotinylated using terminal transferase, and loaded onto a flow cell. Ten oligo probes were added sequentially to identify spike-ins (eight that bind, two that do not), followed by one oligo probe that hybridizes to a differentially methylated site. The kinetic profile of the final probe was used to determine the methylation state of each target molecule identified as a spike-in (Figure 8A). A map of all molecules detected on a portion of the flow cell surface is shown in (Figure 8B). [Figure 8B] Collectively, we demonstrate detection of spike-ins and discrimination of methylation in a mixed background of cell-free DNA. An equal mixture of unmethylated and methylated synthetic single-stranded DNA targets was spiked into synthetic plasma at a high percentage (approximately 50%). DNA was extracted from the mixture, biotinylated using terminal transferase, and loaded onto a flow cell. Ten oligo probes were added sequentially to identify spike-ins (eight that bind, two that do not), followed by one oligo probe that hybridizes to a differentially methylated site. The kinetic profile of the final probe was used to determine the methylation state of each target molecule identified as a spike-in (Figure 8A). A map of all molecules detected on a portion of the flow cell surface is shown in (Figure 8B). [Figure 8C]Collectively, we demonstrate detection of spike-ins and discrimination of methylation in a mixed background of cell-free DNA. An equal mixture of unmethylated and methylated synthetic single-stranded DNA targets was spiked into synthetic plasma at a high percentage (approximately 50%). DNA was extracted from the mixture, biotinylated using terminal transferase, and loaded onto a flow cell. Ten oligo probes were added sequentially to identify spike-ins (eight that bind, two that do not), followed by one oligo probe that hybridizes to a differentially methylated site. The kinetic profile of the final probe was used to determine the methylation state of each target molecule identified as a spike-in (Figure 8A). A map of all molecules detected on a portion of the flow cell surface is shown in (Figure 8B). [Figure 9A] Collectively, we show the discrimination between cytosine and hydroxymethylcytosine in nucleic acids using a 5mer oligo, where the dwell time of the oligo binding to hydroxymethylcytosine is longer than the dwell time of the oligo binding to cytosine. [Figure 9B] Collectively, we show the discrimination between cytosine and hydroxymethylcytosine in nucleic acids using a 5mer oligo, where the dwell time of the oligo binding to hydroxymethylcytosine is longer than the dwell time of the oligo binding to cytosine. [Figure 10A] Collectively show the discrimination of cytosine (blue / dark grey), hydroxymethylcytosine (pink / red / yellow / light grey), and methylcytosine (green / purple / grey) in nucleic acids using 4mer oligos. Experimentally generated examples of oligo binding profiles for 5-hydroxymethylated (5-hmC), 5-methylated, and unmethylated DNA. Shown are three separate flow cells used, one containing a target nucleic acid with two 5-hmC residues at the indicated sites, one containing two 5-MeC at the indicated sites, and one containing a DNA target of identical sequence but without any epigenetic modifications. The same Cy3-labeled oligonucleotide is added to each flow cell and the average ton is determined. [Figure 10B]Collectively show the differentiation of cytosine (blue / dark grey), hydroxymethylcytosine (pink / red / yellow / light grey), and methylcytosine (green / purple / grey) in nucleic acids using 4mer oligos. Examples of experimentally generated oligo binding profiles for 5-hydroxymethylated (5-hmC), 5-methylated, and unmethylated DNA. Average ton for individual molecules are displayed, color-coded by type or absence of epigenetic modification. Residence time of oligo bound to cytosine is shorter than hydroxymethylcytosine, which is shorter than methylcytosine. [Figure 11] 1 illustrates an exemplary system topology including a computer system, according to an exemplary embodiment of the present disclosure.
[0015] It should be understood that the accompanying drawings are not necessarily to scale and that they illustrate various features illustrating the underlying principles of the invention in somewhat simplified form. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0016] Reference will now be made in detail to the embodiments, examples of which are illustrated in the accompanying drawings. In the following detailed description, numerous specific details are set forth in order to provide a thorough understanding of the present disclosure. However, it will be apparent to those skilled in the art that the present disclosure may be practiced without these specific details. In other instances, well-known methods, procedures, and components have not been described in detail so as not to unnecessarily obscure aspects of the embodiments.
[0017] It is also understood that terms such as first, second, etc. may be used herein to describe various elements, but these elements should not be limited by these terms. These terms are used only to distinguish one element from another. For example, a first subject matter may be referred to as a second subject matter, and similarly, a second subject matter may be referred to as a first subject matter, without departing from the scope of the present disclosure. Although both the first subject matter and the second subject matter are subjects, the first subject matter and the second subject matter are not the same subject matter.
[0018] The terms used in this disclosure are for the purpose of describing particular embodiments only and are not intended to limit the invention. When used in the description of the invention and the appended claims, the singular forms "a", "an" and "the" are intended to include the plural forms unless the context clearly indicates otherwise. It is also understood that the term "and / or" as used herein refers to and includes any and all possible combinations of one or more of the associated listed items. It is also understood that the terms "comprise" and / or "comprising" as used herein specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof.
[0019] The foregoing description has included exemplary systems, methods, techniques, instruction sequences, and computing machine program products embodying exemplary implementations. For purposes of explanation, numerous specific details are set forth to provide an understanding of various implementations of the inventive subject matter. However, it will be apparent to those skilled in the art that the inventive subject matter may be practiced without these specific details. In general, well-known instruction instances, protocols, structures, and techniques have not been shown in detail.
[0020] The above description has been set forth with reference to specific implementations for purposes of explanation. However, the exemplary description below is not intended to be exhaustive or to limit the embodiments to the precise forms disclosed. Many modifications and variations are possible in light of the above teachings. The embodiments have been chosen and described in order to best explain the principles and their practical application, and thereby enable others skilled in the art to best utilize the embodiments, and various modifications thereof, as suited to the particular use intended.
[0021] For clarity, not all of the routine features of the implementations described herein are shown and described. It will be understood that in developing any such actual implementation, numerous implementation-specific decisions will be made to achieve the particular goals of the designer, such as compliance with use cases and business-related constraints, and that these particular goals will vary from implementation to implementation and from designer to designer. Moreover, it will be understood that such design activities may be complex and time-consuming, but are nevertheless routine engineering activities for those of ordinary skill in the art having the benefit of this disclosure.
[0022] As used herein, the term "if" may be interpreted to mean "when" or "upon" or "in response to determining" or "in response to detecting," depending on the context. Similarly, the phrase "if it is determined" or "if [a stated condition or event] is detected" may be interpreted to mean "upon determining" or "in response to determining" or "upon detecting [the stated condition or event]" or "in response to detecting [the stated condition or event]," depending on the context.
[0023] As used herein, the term "about" or "approximately" may mean within an acceptable error range for a particular value as determined by one of ordinary skill in the art, which may depend in part on how the value is measured or determined, e.g., the limitations of the measurement system. For example, "about" may mean within 1 or more than 1 standard deviation per practice in the art. "About" may mean within a range of ±20%, ±10%, ±5%, or ±1% of a given value. Unless otherwise stated, when a particular value is described in this application and claims, the term "about" means within an acceptable error range for the particular value. The term "about" has the meaning as commonly understood by one of ordinary skill in the art. The term "about" may refer to ±10%. The term "about" may refer to ±5%.
[0024] In this disclosure, unless expressly stated otherwise, descriptions of devices and systems include one or more computer implementations. For example, for purposes of illustration in FIG. 11, computer system 1900 is depicted as a single device that includes all of the functionality of computer system 1900. However, the invention is not so limited. For example, in some embodiments, the functionality of computer system 1900 is distributed across any number of networked computers, and / or by residing on each of several networked computers, and / or by being hosted on one or more virtual machines and / or containers at remote locations accessible via a communication network (e.g., communication network 1906 in FIG. 11). Those skilled in the art will appreciate that a wealth of different computer topologies are contemplated for analysis computer system 1900 as well as other devices and systems of the present disclosure, and all such topologies are within the scope of the present disclosure. Moreover, rather than relying on a physical communication network 1906, the illustrated devices and systems can wirelessly transmit information to each other. Thus, the exemplary topology depicted in FIG. 11 is merely useful for illustrating features of embodiments of the present disclosure in a manner readily understood by those skilled in the art.
[0025] 11 illustrates a block diagram of a distributed computer system according to some embodiments of the present disclosure, such as computer system 1900. Computer system 1900 facilitates at least the transmission of one or more instructions for detecting epigenetic modifications of nucleic acids.
[0026] In some embodiments, communications network 1906 optionally includes the Internet, one or more local area networks (LANs), one or more wide area networks (WANs), other types of networks, or a combination of such networks.
[0027] Examples of communication networks 1906 include the World Wide Web (WWW), intranets, and / or cellular telephone networks, wireless networks such as wireless local area networks (LANs) and / or metropolitan area networks (MANs), and other devices communicating via wireless communication. Wireless communications may include, but are not limited to, Global System for Mobile Communications (GSM), Enhanced Data GSM Environment (EDGE), High Speed Downlink Packet Access (HSDPA), High Speed Uplink Packet Access (HSUPA), Evolution, Data Only (EV-DO), HSPA, HSPA+, Dual Cell HSPA (DC-HSPADA), Long Term Evolution (LTE), Near Field Communication (NFC), Wideband Code Division Multiple Access (W-CDMA), Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Bluetooth, Wireless Fidelity (Wi-Fi) (e.g., IEEE 802.11a, IEEE 802.11ac, IEEE 802.11ax, IEEE 802.11b, IEEE 802.11g and / or IEEE Optionally, the communication may use any of a number of communications standards, protocols, and technologies, including 802.11n), Voice over Internet Protocol (VoIP), Wi-MAX, protocols for email (e.g., Internet Message Access Protocol (IMAP) and / or Post Office Protocol (POP)), instant messaging (e.g., Extensible Messaging and Presence Protocol (XMPP), Session Initiation Protocol for Instant Messaging and Presence Enhancements (SIMPLE), Instant Messaging and Presence Service (IMPS)), and / or Short Message Service (SMS), or other suitable communications protocols (including communications protocols not yet developed as of the filing date of this document).
[0028] In various embodiments, computer system 1900 includes one or more processing units (CPUs) 1902 , a network or other communication interface 1904 , and memory 1912 .
[0029] In some embodiments, computer system 1900 includes a user interface 1906. User interface 1906 typically includes a display 1908 for presenting media. In some embodiments, display 1908 is integrated within the computer system (e.g., housed in the same chassis as CPU 1902 and memory 1912). In some embodiments, computer system 1900 includes one or more input devices 1910 that allow a subject to interact with computer system 1900. In some embodiments, input device 1910 includes a keyboard, a mouse, and / or other input mechanisms. Alternatively, or additionally, in some embodiments, display 1908 includes a touch-sensitive surface (e.g., if display 1908 is a touch-sensitive display or computer system 1900 includes a touchpad).
[0030] In some embodiments, the computer system 1900 presents media to the user through the display 1908. Examples of media presented by the display 1908 include one or more images (e.g., a user interface on the display 1908 presenting a chart of the 3Cs), video, audio (e.g., a waveform of an audio sample), or a combination thereof. In a typical embodiment, one or more images, videos, audio, or a combination thereof are presented by the display 1908 through a client application. In some embodiments, the audio is presented through an external device (e.g., a speaker, headphones, an input / output (I / O) subsystem, etc.) that receives audio information from the computer system 1900 and presents audio data based on the audio information. In some embodiments, the user interface 1906 also includes an audio output device, such as a speaker, or an audio output for connecting to a speaker, earphones, or headphones.
[0031] The memory 1912 includes high-speed random access memory, such as DRAM, SRAM, DDR RAM, or other random access solid-state memory devices, and optionally also includes non-volatile memory, such as one or more magnetic disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile solid-state storage devices. The memory 1912 may optionally include one or more storage devices located remotely from the CPU(s) 1902. The memory 1912, or alternatively the non-volatile memory device(s) within the memory 1912, includes a non-transitory computer-readable storage medium. Access to the memory 1912 by other components of the computer system 1900, such as the CPU(s) 1902, is optionally controlled by a controller. In some embodiments, the memory 1912 may include mass storage devices located remotely relative to the CPU(s) 1902. In other words, some data stored in memory 1912 may be hosted on a device that is in fact external to the computer system 1900, but that is electronically accessible by the computer system 1900 via the Internet, an intranet, or other form of network 106 using the communications interface 1904 or an electronic cable.
[0032] In some embodiments, the memory 1912 of the computer system 1900 stores: an operating system 1920 including procedures for handling various basic system services (e.g., an embedded operating system such as ANDROID, iOS, DARWIN, RTXC, LINUX, UNIX, OS X, WINDOWS, or VxWorks); · an electronic address associated with the computer system 1900 that identifies the computer system 1900 (e.g., within the communications network 1906); · a control module 1922 including one or more modules 1924 for controlling one or more processes (e.g., methods) associated with the computer system 1900; Optionally, a client application for presenting information (eg, media) using the display 1908 of the computer system 1900.
[0033] In some embodiments, the control module 1922 includes one or more models 1924 configured to perform one or more steps of the methods of the disclosure (e.g., a first method for determining the methylation state of at least a portion of a nucleic acid molecule, a second method for determining the modification state of a plurality of nucleic acid molecules comprising different sequences, a third method for determining the modification (e.g., methylation) state of one or more nucleic acid molecules, a fourth method for determining the sequence and epi-sequence of at least a portion of a nucleic acid molecule, etc.).
[0034] Additionally, in some embodiments, the computer system includes one or more reference libraries (e.g., one or more reference databases, such as a cancer reference database, a nucleic acid reference database, etc.).
[0035] Modifications detectable by the methods provided herein include chemically modified bases, enzyme modified bases, DNA damage, abasic sites, non-natural bases, secondary structures, and drugs bound to nucleic acids.Exemplary modifications that can be detected by the methods of the present invention include, but are not limited to, methylated bases (e.g., 5-methylcytosine, N6-methyladenosine, etc.), pseudouridine bases, 7,8-dihydro-8-oxoguanine bases, 2'-O-methyl derivative bases, nicks, non-purine sites, non-pyrimidinic sites, pyrimidine dimers, cisplatin crosslink products, oxidative damage, hydrolytic damage, bulky base adducts, thymine dimers, photochemical reaction products, interstrand crosslink products, mismatched bases, secondary structures, and drug binding.
[0036] In some embodiments, the binding properties of the oligos are tuned by using different modifications in the probe, for example, there are several options to enhance the binding stability of short oligos (LNA, PNA, locked nucleic acid (LNA), peptide nucleic acid (PNA), analog bases with altered stability, minor groove binding, stacking, intercalation, cationic complexes, etc.), and certain modifications can accentuate the binding differences between modified and unmodified bases. In some embodiments, the binding properties are tuned by the buffer composition, especially the concentration and type of salt, the presence of denaturants, binding enhancers, pH, temperature, and oligo-probe concentration.
[0037] In some embodiments, in order to measure oligo binding, the target must be immobilized on a surface so that measurements to determine discrimination and methylation state (and repeats thereof) can be performed on the same molecule.
[0038] In some embodiments, the invention involves determining only the modification state of a target molecule where the sequence or identity of the target molecule is already known or identified.
[0039] In some embodiments, the methods include attaching one or more oligos to a nucleic acid molecule that exhibit different binding to the sequence when modified compared to when unmodified.
[0040] In some embodiments, when the identity of the target molecule is not known in advance, for example, when examining random or shotgun fragments of genomic DNA, it is necessary to determine the identity of the nucleic acid molecule in order to detect epigenetic modifications in such nucleic acid samples. Thus, this embodiment includes two main aspects. In some embodiments, the first main aspect includes obtaining sequence information from the nucleic acid molecule to determine its identity. In some embodiments, the second main aspect includes binding one or more oligos to the nucleic acid molecule that exhibit different binding to the sequence compared to the modified and unmodified versions.
[0041] In some embodiments, the sequence information is obtained by sequencing, e.g., PacBio sequencing, Helicos sequencing, Oxford nanopore sequencing, XGenomes sequencing (2, 3). In some embodiments, the sequence information is obtained by molecular probing (2, 3).
[0042] In some embodiments, the method for detecting epigenetic modifications in a nucleic acid sample comprises two main aspects. In some such embodiments, the first aspect comprises binding one or more oligos to a nucleic acid molecule and determining its discrimination. In some such embodiments, the second aspect comprises binding one or more oligos to a nucleic acid molecule that exhibit different binding to the sequence compared to modified and unmodified.
[0043] In some embodiments, the assigned discrimination is based on determining whether one or multiple oligos bind to the target and what their binding profile (number of repeat bindings, time in light, time in dark) is.
[0044] In some embodiments, the identity of a nucleic acid includes its genomic origin.
[0045] In some embodiments, the modification state is determined by matching the binding profile of one or more oligos with an expected binding profile when one or more modifications are present. In some embodiments, when the discrimination has been determined, the data of molecules whose modification state is of interest can be selectively processed (e.g., processed by model 1924 of control module 1922 in Figure 11).
[0046] In some embodiments, historical or training data (e.g., historical or training data of the model 1924 of the control module 1922 in FIG. 11) is used to determine whether an obtained binding profile corresponds to the presence of a modification.
[0047] In some embodiments, the spike-in control is used as a reference to determine whether the obtained binding profile corresponds to the presence of a modification.
[0048] In some embodiments, the extent of binding of the probe to the modified and unmodified complementary sequences is predetermined or is determined in situ by observing binding to a reference spike-in target.
[0049] In some embodiments, spike-in controls allow for the establishment of normal signal levels for modified and unmodified bases. For example, such controls include one or more synthetic oligonucleotides with modified bases in specific sequence contexts and their sequence-matching oligonucleotides without the modified bases.
[0050] In some embodiments, the binding comprises duplex formation.Thus, in some embodiments, the disclosure provides a method for determining the methylation state of at least a portion of a nucleic acid molecule.
[0051] In some embodiments, the methods include measuring the degree of hybridization of a complementary oligonucleotide probe to a test target sequence.
[0052] In some embodiments, the methods include optionally measuring the degree of hybridization of different complementary oligonucleotide probes to the test target sequence.
[0053] In some embodiments, the method comprises determining that methylation is present if the degree of hybridization is greater than the degree of hybridization to a reference unmethylated target sequence and / or is equivalent to the degree of hybridization to a reference methylated target sequence.
[0054] In some embodiments, a pattern of hybridization of two or more oligonucleotides to a target sequence is obtained.
[0055] In some embodiments, the present disclosure relates to providing methods for determining the modification state of a plurality of nucleic acid molecules that comprise distinct sequences.
[0056] In some such embodiments, the methods include obtaining a sample of the nucleic acid molecule.
[0057] In some such embodiments, the method comprises dispersing and immobilizing / fixing nucleic acid molecules onto a surface, thereby obtaining an array of nucleic acid molecules, each molecule in the array being fixed to a distinct location on the surface.
[0058] In some such embodiments, the method comprises exposing one or more oligos (typically a repertoire or panel of oligos) of known sequence to a nucleic acid, wherein the one or more of said oligos can determine the identity of each of the individual nucleic acid molecules, detecting binding of the one or more of said oligos to each of the individual nucleic acids, and determining the identity of said nucleic acids.
[0059] In some such embodiments, the method comprises exposing one or more oligos of known sequence to a nucleic acid, where one or more of said oligos may have a different binding profile when the sequence is modified compared to when the sequence is not modified, detecting the binding profile of one or more of said oligos to each of the individual nucleic acids, and determining whether the binding profile more closely matches the binding profile when the sequence is modified or when the sequence is not modified.
[0060] In some such embodiments, the methods include recording the modification state of the identified molecules.
[0061] In some embodiments, the present disclosure relates to providing methods for determining the modification state of a plurality of nucleic acid molecules that comprise distinct sequences.
[0062] In some embodiments, the method includes obtaining a sample of the nucleic acid molecule.
[0063] In some embodiments, the method comprises dispersing and immobilizing / fixing nucleic acid molecules onto a surface, thereby obtaining an array of nucleic acid molecules, each molecule in the array being immobilized on the surface.
[0064] In some embodiments, the methods involve exposing a nucleic acid to one or more oligos (typically a repertoire or panel of oligos) of known sequence, wherein the one or more oligos are capable of determining the respective discriminatory properties of individual nucleic acid molecules, and the one or more oligos have a different binding profile when the sequence of the nucleic acid molecule is modified compared to when it is not modified.
[0065] In some embodiments, the method comprises detecting binding of one or more of said oligos to each of the distinct nucleic acids, determining the discrimination of said nucleic acids, detecting a binding profile of one or more of said oligos to each of the distinct nucleic acids, and determining whether the binding profile more closely matches the binding profile when the sequence is modified or the binding profile when the sequence is unmodified.
[0066] In some embodiments, the method includes recording the modification state of the identified molecules.
[0067] In some embodiments, the present disclosure relates to methods for determining the modification (eg, methylation) status of one or more nucleic acid molecules.
[0068] In some embodiments, the method involves exposing the nucleic acid to one or more oligos of known sequence.
[0069] In some embodiments, the methods involve detecting whether one or more oligos hybridize to any of the nucleic acid molecules.
[0070] In some embodiments, the methods involve exposing a nucleic acid molecule to one or more oligos with known sequence and known hybridization behavior with respect to nucleic acid modifications, such as methylation.
[0071] In some embodiments, the methods involve detecting whether one or more oligos hybridize to any of the nucleic acid molecules.
[0072] In some embodiments, the method comprises constructing a binding profile for each nucleic acid molecule. The binding profile comprises a set of one or more binding calls for each oligo exposed to the nucleic acid molecule. Further, the binding profile comprises a confidence metric that obtains an estimate of the probability of error for each binding call.
[0073] In some embodiments, the methods include a set of one or more reference nucleic acid sequences.
[0074] In some embodiments, the methods utilize a computer database (eg, computer system 1900 of FIG. 11) that stores the positions of perfect matches between one or more oligonucleotide sequences and one or more reference sequences.
[0075] In some embodiments, the method utilizes a computer program (e.g., control module 1922 of computer system 1900 of FIG. 11 ) that uses a database of exact matches between one or more reference sequences to convert each binding profile into one or more intervals within any subset of reference sequences that are most likely to encompass the nucleic acid sequence corresponding to the nucleic acid molecule.
[0076] In some embodiments, the method utilizes a computer program (e.g., control module 1922 of computer system 1900 of FIG. 11 ) that uses a subset of binding profiles corresponding to one or more matching intervals between one or more reference sequences and modification-sensitive oligonucleotides to construct a modification profile for each of the nucleic acid molecules.
[0077] In some embodiments, the molecules are dispersed on the surface such that they are located, on average, less than 250 nanometers (nm) apart on the surface and can be resolved by super-resolution imaging.
[0078] In some embodiments, binding events on individual molecules are localized by single molecule localization.
[0079] The same approach as above can be used for a range of modifications, although depending on the modification the extent of binding may be greater or less than the unmodified nucleic acid.
[0080] Although methylation is the most common modification in genomes, other modifications coexist in genomes in real-world samples, and a means to distinguish and classify different modifications is useful. Due to this complexity, in some embodiments, if information about a particular modification is only desired, that modification can be tagged and the binding profile to the tag measured. For example, β-glucosyltransferase (βGT) tags hydroxymethyl to provide a significant detectable signal.
[0081] In some embodiments, multiple modifications may be present within the footprint of a single oligo along the target sequence, which will affect the binding profile. Historical and training data (e.g., historical and training data of model 1924 of control module 1922 in FIG. 11) can be used to distinguish whether there are one, two, or three CpG modifications in the footprint of a 5-mer. Aggregated data from multiple oligos targeting a locality can also aid in this decision. In other cases, different types of modifications may be present within the footprint of an oligonucleotide. For example, methyl and hydroxymethyl sites may be found in the same location. Similarly, historical and training data can be used to determine what type of modification is present. In both cases, appropriately selected spike-in controls are used to assist in the decision.
[0082] In some embodiments, machine learning can be used to determine binding profiles with different numbers of modifications or types of modifications.
[0083] In some embodiments, the measurements provide estimates that are relevant to determining the state of the modification and how it may affect a biological process or medical condition. In many cases, the modification state is used as a biomarker and may be one of multiple biomarkers that are used in combination to provide a probability and therefore a clinical decision or to serve as the basis for a hypothesis regarding a molecular event.
[0084] In some embodiments, the ends of the sample DNA molecules are modified to facilitate immobilization and immobilization to a surface. In some embodiments, the ends are modified by adding one or more nucleotides using a terminal transferase. In some embodiments, a single modified nucleotide is added using a terminal transferase. In some embodiments, homopolymers are added using deoxynucleotides. Some of the nucleotides are modified to facilitate capture to a surface, for example, modified with biotin for immobilization to a streptavidin / neutravidin surface, or modified with aminoallyl for immobilization to a COOH surface. In some embodiments, ligation is used to add short oligos to the ends, which may carry modifications that facilitate capture to a surface. In some embodiments, the oligonucleotides or homopolymers hybridize to complementary sequences attached to the surface, thus immobilizing the target sequence.
[0085] In some embodiments, multiple oligos are bound in multiple cycles. In some embodiments, there is one or more wash steps between binding of one oligo to another. In some embodiments, multiple oligos are bound and exposed in one cycle (multiplexed). In some embodiments, the oligos are labeled with the same label (e.g., a fluorescent label, a light scattering label, or a plasmon resonance label). In some embodiments, the oligos are labeled with different labels. In some embodiments, the different labels are represented by different emission and / or excitation wavelengths. In some embodiments, the different labels are represented by different physical properties including fluorescence lifetime, anisotropy, optical permittivity.
[0086] In some embodiments, the label is at one end of the oligonucleotide. In some embodiments, the oligonucleotide has labels at both ends. In some embodiments, the oligonucleotide has a label internally.
[0087] In some embodiments, a methylation-sensitive reagent (e.g., an antibody or modified binding protein or other ligand) occupies the site where the modified nucleotide is located and regulates binding of the oligonucleotide, which may be complementary to the site occupied by the methylation-sensitive reagent.
[0088] In some embodiments, the identity and modification state of each molecule are aggregated to provide insight into a biological process or medical condition.
[0089] In some embodiments, the extent of modification for each target molecule is estimated.
[0090] In some embodiments, a modified haplotype is determined.
[0091] In some embodiments, the double-stranded DNA is denatured before or after immobilization so that the molecules being interrogated are single-stranded.
[0092] The method does not require modification to modify (e.g., β-glucosyltransferase (βGT) labeling of hmC is not required), does not require separate sequencing of treated and untreated portions of the sample, e.g., bisulfite sequencing for methylation detection, and does not require an amplification step such as polymerase chain reaction.
[0093] Nonetheless, in some embodiments, the differential oligonucleotide binding methods of the invention are able to distinguish between base-modified intermediates in common methylation / hydroxymethylation kits, including Tet-assisted pyridine-borane sequencing (TAPS, Base Genomics / Exact Sciences) and the enzymatic methyl-seq (New England Biolabs).
[0094] In some embodiments, individual nucleic acid molecules are attached to a surface as part of an array of nucleic acid molecules. In some embodiments, the array contains a range of molecules of different species or sequences (e.g., a fragment containing the entire transcriptome or the entire human genome). Many molecules in the array may share sequences completely or partially. In some embodiments, the array is a single molecule array. Yhe sample molecules are randomly arranged at different locations on the surface. In some embodiments, the molecules remain fixed at distinct locations throughout the molecular identity and modification detection process. In some embodiments, the exact location to which a particular molecule is attached is not known until the molecular identity is determined as part of the method of the invention.
[0095] In some embodiments, the identity of the target molecule is already known and only the modification state of the molecule is determined. The modification state may include a pattern of modifications along the target sequence. In some embodiments, the target molecule forms part of a spatially addressable array or microarray. In some such embodiments, the identity of the molecule in each element / spot of the microarray is known. In some embodiments, the microarray element / spot comprises multiple molecules and the modification state is determined as a bulk measurement. In some embodiments, the microarray spot comprises multiple molecules and the modification state is determined for individual molecules within the element / spot.
[0096] In some embodiments, the target molecule is not attached to a substrate, but is free in solution. In some such embodiments, the target molecule is single-stranded. In some such embodiments, the epigenetic state is determined by adding an oligonucleotide probe for detecting epigenetic modification to the solution, and then measuring a melting curve. The modification is detected by the fact that a different temperature (higher in the case of hydroxymethyl C and methyl C) is required to melt the heteroduplex formed between the target molecule and the probe when the modification is present compared to when the modification is absent.
[0097] In some embodiments, the present disclosure relates to providing methods for determining the sequence and epi-sequence of at least a portion of a nucleic acid molecule.
[0098] In some embodiments, the method comprises immobilizing the nucleic acid molecule on a test substrate if the nucleic acid molecule is a single-stranded molecule, or denaturing the nucleic acid molecule into a single-stranded molecule and immobilizing the single-stranded nucleic acid molecule on the test substrate if the nucleic acid molecule is a double-stranded molecule, or immobilizing the nucleic acid molecule on the test substrate if the nucleic acid molecule is a double-stranded molecule and denaturing the nucleic acid molecule on the test substrate into a single-stranded molecule, thereby forming a single-stranded nucleic acid immobilized on the test substrate.
[0099] In some embodiments, the method includes exposing the immobilized single-stranded nucleic acid to each oligonucleotide probe species in a set of oligonucleotide probe species, each of which is capable of hybridizing to its complementary portion located at one or more locations on the immobilized single-stranded nucleic acid and has (i) a unique respective predefined sequence, (ii) a predefined length, and (iii) a respective label selected from the group consisting of dyes, fluorescent nanoparticles, plasmon resonant particles, light scattering particles, nanoparticles, and fluorescence resonance energy transfer (FRET) partners (capable of producing a fluorescent signal). The exposing step is carried out under conditions such that: i) the oligonucleotide probes of each oligonucleotide probe species of the set of oligonucleotide probe species repetitively, transiently and reversibly bind to one or more locations on the fixed single-stranded nucleic acid on the test substrate, thereby forming a respective transient heteroduplex at each of the one or more locations on the fixed single-stranded nucleic acid on the test substrate; ii) respective emissions of optical activity from the respective labels are generated and detected by the repetitively, transiently and reversibly binding of the oligonucleotide probes of each oligonucleotide probe species of the set of oligonucleotide probe species to one or more locations on the fixed single-stranded nucleic acid on the test substrate, which are detected at each of the one or more locations on the fixed single-stranded nucleic acid on the test substrate.
[0100] In some embodiments, the method includes determining whether one or more portions of the fixed single-stranded nucleic acid are complementary to each oligonucleotide probe species of the set of oligonucleotide probe species by counting and measuring the duration of each occurrence of optical activity at each of the one or more locations on the fixed single-stranded nucleic acid on the test substrate during the exposure step using a two-dimensional imager capable of detecting each occurrence of optical activity generated from each label, thereby obtaining a first set of one or more locations on the fixed single-stranded nucleic acid that are complementary to each oligonucleotide probe species of the set of oligonucleotide probe species.
[0101] In some embodiments, the method includes washing the test substrate to remove each oligonucleotide probe species of the set of oligonucleotide probe species from the test substrate.
[0102] In some embodiments, the method includes repeating the exposing, measuring, and washing steps by exposing the immobilized single-stranded nucleic acid on the test substrate to each different oligonucleotide probe species in the set of oligonucleotide probe species, thereby obtaining a second set of one or more locations on the immobilized single-stranded nucleic acid that are complementary to each different oligonucleotide probe species in the set of oligonucleotide probe species.
[0103] In some embodiments, the methods include determining a sequence of at least a portion of a nucleic acid based at least in part on a first set of one or more positions on the fixed, single-stranded nucleic acid that are complementary to each oligonucleotide probe species of the set of oligonucleotide probe species and a second set of one or more positions on the fixed, single-stranded nucleic acid that are complementary to another respective oligonucleotide probe species of the set of oligonucleotide probe species.
[0104] The method also includes determining whether a portion of the nucleic acid molecule has one or more epigenetic modifications based on the observed differential binding behavior of the oligonucleotide probes of each oligonucleotide probe species in the set of oligonucleotide probe species to their complementary portions located at one or more locations on the fixed single-stranded nucleic acid when the one or more locations have the epigenetic modification compared to when the one or more locations do not have the epigenetic modification.
[0105] In some embodiments, the differential binding behavior includes differences in the number of repeat bindings obtained by counting the onset of optical activity, on-rate, residence time, In some embodiments, the differential binding behavior includes differences in the number of repeat bindings, dark time (when there is no detectable optical activity) and / or light time (when there is optical activity).
[0106] In some embodiments, the binding kinetics of the probes (e.g., time in light, time in dark, number of repeat binding events) are used to determine the methylation state of cytosines in each fragment by either or both of the following: (i) comparing the probe binding data to a database of previously collected data from probes binding to unmodified and modified cytosines (e.g., computer system 1900 of FIG. 11 ), and / or (ii) comparing the probe binding data for each DNA fragment to the probe binding dynamics of that probe for all DNA fragments in the sample (and / or the dynamics of the probe binding to control DNA spiked into the sample).
[0107] In some embodiments, the binding profile of each oligonucleotide capable of binding to a nucleic acid sequence having a potential modification is pre-characterized by testing against synthetic modified and unmodified versions of the nucleic acid sequence having a potential modification, thus serving as a reference for comparing the binding profile obtained for the sample molecule.
[0108] In some embodiments, information regarding the modification state of sample molecules obtained by the methods of the invention is used as the basis for determining a biological or medical condition.
[0109] composition
[0110] In some embodiments, compositions for oligonucleotides of known sequence are used in the present invention. Some embodiments include oligonucleotides less than 8, less than 7, less than 6, less than 5, less than 4 nucleotides in length. In some embodiments, the compositions include a repertoire or panel of oligonucleotides all of the same length. In some embodiments, the oligos are different lengths. In some embodiments, the oligos include LNA nucleotides. In some embodiments, the oligos include LNA / DNA oligos. In some embodiments, the oligos include DNA, LNA, LNA / DNA oligos of one or more lengths. In some embodiments, one or more positions on the oligonucleotide are methylated. In some embodiments, the modification is at the 5 position of the base. In some embodiments, some of the oligos include an undefined N or universal base position. In some embodiments, some of the oligos include a conjugate. In some embodiments, the conjugate is a ZNA, spermine residue, or other positively charged residue. In some embodiments, the conjugate is an intercalating structure or a stacking / capping structure. In some embodiments, the capping structure comprises UAQ (e.g., attached via a reagent: 5'-dimethoxytrityl-uridine, 2'-(anthraquinone-2-ylcarboxamido)-3'-succinoyl-long chain alkylamino-CPG), pyrene, thiazole orange. In some embodiments, some of the probes comprise multiple copies of an oligo linked together. In some embodiments, copies of the probe sequence are connected in series with or without a spacer (e.g., hexatheylene glycol), and in one such embodiment, a label is attached to one of the nucleosides. In some embodiments, the probes are connected to a dendrimer via branching amidites, each probe being an arm or branch of a dendritic structure. In some embodiments, the label is on one branch of the dendrimer. In some embodiments, the label is on multiple branches of the dendrimer. In some embodiments, some of the probes are PNA or other non-natural scaffolds.In some embodiments, some of the bases are modified to increase duplex stability. In some embodiments, some of the bases are modified to increase nucleation ability. In some embodiments, some of the bases are modified to decrease duplex stability. In some embodiments, the conjugate is at one end. In some embodiments, the conjugate is at both ends. In some embodiments, the conjugate is internal. Some embodiments include buffer compositions useful in the present invention: TMACl, SSC, ethylene carbonate, dextran sulfate, formamide, PEG, urea, betaine, etc.
[0111] RNA modifications including N6-methyladenosine (m6A), N6,2'-O-dimethyladenosine (m6Am), 8-oxo-7,8-dihydroguanosine (8-oxoG), pseudouridine (Ψ), 5-methylcytidine (m5C), and N4-acetylcytidine (ac4C) are suitable for the methods of the invention.
[0112] DNA modifications including 5-methylcytosine (5mC), 5-hydroxymethylcytosine (5hmC), 5-formylcytosine (5fC), and 5-carboxylcytosine (5caC) are suitable for the methods of the present invention.
[0113] Further details and information regarding the systems, methods, and compounds of the present disclosure can be found in U.S. Patent Application Publication No. 2019 / 0149681 A1, published May 16, 2019, entitled "Image Processing System, Image Processing Apparatus, Control Method of Imaging Processing Apparatus, and Program"; Jungmann et al., 2010, "Single-molecule kinetics and super-resolution microscopy by fluorescence imaging of transient binding on DNA origami." Nano letters, 10(11), pg. 4756-4761; and U.S. Patent Application Publication No. 2022 / 0064712 A1, published March 3, 2022, entitled "Sequencing by emergence"; and U.S. Patent Application Publication No. 2020 / 0082913, published March 12, 2020, entitled "Systems and Methods for Determining Sequencing." A1; U.S. Patent Application Publication No. 2020 / 0056229 A1, published on February 20, 2020, entitled "Sequencing by emergence"; and U.S. Patent No. 10,982,260 B2, published on April 20, 2021, entitled "Sequencing by emergence," each of which is incorporated by reference herein in its entirety. EXAMPLES
[0114] Example 1: Detection of methylated single-stranded nucleic acid molecules (Figure 4)
[0115] Nucleic Acid A single stranded DNA molecule was synthesized with the following sequence: Target 1: Biotin-TTTTTTTTTTTTTTTTTTTTTTTTTTTTTCCATTCCCGCCACCATCGCCTCAATCCCTGTGCGCTAATTTTTTTTTTTTTTTTTTTTTTTT and a methylated version were synthesized. Target 2: Biotin-TTTTTTTTTTTTTTTTTTTTTTTTTTTCCATTCccGCCACCATcGCCTCAATCcCTGTGCGCTAATTTTTTTTTTTTTTTTTTTTTT.
[0116] A custom flow cell was prepared using a streptavidin-coated cover glass attached to a plastic flow cell chamber. The flow cell was washed with BX (5 mM Tris, 10 mM MgCL2, 0.05% Tween20 and 1 mM EDTA). 10 pM of target 1 and target 2 were added to separate chambers of the flow cell. After 5 minutes, the flow cell was washed with 150 ul of BX. 5 nM of oligonucleotides (Cy3-GGTAG, Cy3-GTAGC, Cy3-TAGCG, Cy3-AGGTA, Cy3-GGTGG, Cy3-CGGAG) in BX+F (5 mM Tris, 10 mM MgCL2, 0.05% Tween20 and 1 mM EDTA, 30% formamide) were added sequentially to each flow cell. Washing between oligonucleotide imaging was performed three times with 150 ul of BX+F. Imaging was performed on an ONI Nanoimager S in TIRF mode with focus lock enabled and 10% laser (532 nm) power. For each oligo, a total of 2000 (200, or as few as 200, or as many as 8,000) frames were captured at 200 ms (or as low as 25 ms) per frame.
[0117] The data was processed using drift correction algorithms, single molecule localization algorithms, and algorithms (e.g., model 1924 of control module 1922 in Figure 11) that determine the number of repeated binding events, the dark time and the light time of binding for each individual molecule. Statistics regarding the residence times of oligos targeting target 1 (unmethylated) and target 2 (methylated) were compared and the data are summarized in Figures 3A-5B.
[0118] Example 2: Detection of methylation status from FFPE tumor samples
[0119] Eight FFPE sections with a thickness of 5um and a service area of 250mm2 were obtained from tumor samples from lung cancer patients. DNA was isolated using the QIAamp DNA FFPE Tissue Kit (following the manufacturer's instructions). DNA was then fragmented to approximately 150bp using a ME220 Focused-ultrasonicator (Covaris). Biotin labeling of DNA was then performed using the following terminal transferase (TdT) reaction: 200ng cfDNA, 1x TdT reaction buffer, 4uM ddATP-biotin, 250nM CoCl2, 40U TdT, incubated for 90 minutes. Samples were then purified using a GeneJet PCR purification kit following the manufacturer's instructions. A unique pool of control oligonucleotides was spiked into the samples. The pool of oligonucleotides includes methylated and unmethylated cytosines.
[0120] A custom flow cell was prepared using a streptavidin-coated coverslip (Schott AG, Mainz Germany) mounted in a plastic flow cell chamber (Sticky-Slide VIV 0.4, Ibidi, Martinstried, Germany). The flow cell was washed with BX (5 mM Tris, 10 mM MgCL2, 0.05% Tween20 and 1 mM EDTA). Biotin-labeled tumor DNA prepared by the TdT protocol described above was added to a single channel of the flow cell. To make the DNA single-stranded and to wash the flow cell, the channel is washed five times with freshly prepared 0.5 M NaOH (including one 2 min incubation in NaOH at room temperature) followed by four washes with BX buffer. 5 nM of discriminating oligonucleotide (for genome-wide analysis, up to 250 probes randomly selected from the complete repertoire of 5-mers are tested) in BX+F (5 mM Tris, 10 mM MgCL2, 0.05% Tween20 and 1 mM EDTA, 30% formamide) was added to each flow cell. Sets of oligos from the complete 1024 repertoire of 5-mers were added sequentially. Washing between oligonucleotide imaging was performed three times with 150 ul of BX+F. Imaging was performed in TIRF mode on an ONI Nanoimager S. A total of 500 frames were captured for each oligo at 200 ms per frame. After molecular identification, methylation probes (see Table 1) in BX+F (5 mM Tris, 10 mM MgCL2, 0.05% Tween20 and 1 mM EDTA, 30% formamide) were added sequentially to each flow cell. Imaging was performed on an ONI Nanoimager S in TIRF mode with focus lock enabled and 10% laser (532 nm) power. A total of 2000 frames were captured for each oligo at 200 ms per frame.
[0121] The data was processed using drift correction algorithms, single molecule localization algorithms, and algorithms (e.g., one or more models 1924 of control module 1922 in FIG. 11) that determine the number of repeated binding events, the dark and light times of the fluorescent signal due to the binding of each individual molecule. The resulting processed data is further processed using statistical algorithms that estimate the identity of the molecules by reference to a cancer database, providing the methylation probability of all sites containing cytosine, C residues, based on the measured residence times in the sample DNA and the control DNA. Alternatively, machine learning algorithms are used to classify molecules as containing methylated C or not.
[0122] Example 3: Whole genome methylation assay of plasma DNA.
[0123] To detect the methylation status of cell-free DNA in plasma, the following steps are performed.
[0124] (i) The blood sample is centrifuged to obtain plasma.
[0125] (ii) Purify cfDNA from plasma using a commercially available kit (e.g., ThermoFisher MagMax cell-free DNA isolation kit) (for subsequent steps, cfDNA needs to be in EDTA-free buffer; elute the DNA using the kit's EDTA-free buffer or exchange the buffer (e.g., use a commercially available PCR purification kit such as the GeneJet PCR purification kit).
[0126] (ii) Biotin labeling is performed by incubating DNA with terminal deoxyribonuclease (TdT) and ddATP-biotin in TdT buffer.
[0127] (iv) Purify the samples using a commercially available DNA purification kit, such as the Genejet PCR purification kit (biotinylated control DNA can be spiked in at this point, or non-biotinylated DNA can be spiked in earlier in the process).
[0128] (v) A custom flow cell is prepared using a streptavidin-coated coverslip (Schott AG, Mainz Germany) mounted in a plastic flow cell chamber (Sticky-Slide VIV 0.4, Ibidi, Martinstried, Germany).
[0129] (vi) Wash the flow cell with PBS and BX buffer.
[0130] (vii) Add the biotin-labeled DNA prepared above to a single channel of the flow cell. After 4 minutes, wash the flow cell 4 times with 150 ul of Bx.
[0131] (viii) To single-strand the DNA and wash the flow cell, the channel is washed five times with freshly prepared 0.5 M NaOH (including one 2 min incubation in NaOH at room temperature), followed by four washes with BX buffer.
[0132] (ix) The flow cell is mounted onto the ONI nanoimager.
[0133] (x) The channel is primed with an imaging buffer appropriate for the first round of oligos, after which fluorescently labeled oligos are added to the channel.
[0134] (xi) Imaging is performed in TIRF mode with 200 ms frames. Multiple fields of view can be collected in each round of imaging.
[0135] (xii) Further rounds of oligos are passed through the flow cell. For example, up to 1024 5-mers can be passed. Between subsequent rounds of imaging of oligos, the channel is washed twice with wash buffer.
[0136] (xiii) After all rounds of fluorescently labeled oligos, the patterns of probes that bind and do not bind to each DNA fragment on the flow cell are compared to a reference genome to identify the location of the DNA within the genome and the methylation patterns along the identified fragments.
[0137] (xiv) The binding kinetics of the probes (e.g., time in light, time in dark, number of repeated binding events) are used to determine the methylation status of cytosines in each fragment.
[0138] Below is a list of methyl-detection probe sequences (all combinations of probes that can bind to CpG motifs): CCCCG;CCCGGG;CCGGGG;CGCCC;CCGCC;CCCGC;GCCCG;CGGGGG;CCCGT;CGGCC;CCGGC;GGCCG;GCCGG;CCCGA;CGCCCG;CCGCCG;CCGGT;CCGGA;CGGGC;GGGCG,GGCGG;GCGGG;GCGCC;GCCGC;CGGCCG;CGCGG;CGGGGT;CGCCCA;CCGCA;GCCGT;TCCCCG;CGGGA;GGCGC ;GCGGC;GCCGA;ACCCG;TGCCG;AGCCG;CGCCT;CCGCT;CGCGC;GCGCG;CGGCA;GGCGT;GCGGT;GGCGA;GCGGA;TCCGG;ACCGG;TGGCG;TGCGG;CGCGT;CCGTG;CGGC T;AGGCG;AGCGG;CGCGA;GCGCA;CCGAG;TCGCC;TCGGC;TCGGG;ACGCC;ACCGC;GCGCT;TGCGC;AGCGC;CGGTG;CGTGG;ACGGG;CGGAG;CGAGG;TCCGT;ACCGT;AGC GT;TGCGT;CCGTC;CGTCC;GTCCG;TCGGC;TCCGA;CGTGC;GCGTG;GTGCG;ACGGC;ACCGA;AGCGA;CCACG;CACCG;TGCGA;CGCAG;CAGCG;CCGTA;CCGAC;CGACC;CG AGC;GCGAG;GACCG;GAGCG;TCGCG;CCTCG;CTCCG;ACGCG;TCGGT;CGTGT;CGCTG;CTGCG;ACGGT;CCGAT;CGGTC;GGTCG;GTCGG;CGAGT;CGTGA;TCGGA;ACGGA;C ACGG;CGGTA;CCGTT;CGTCG;CGAGA;CGGAC;GGACG;GACGG;ACGCA;TCGCA;CTCGG;GCGTC;GTCGC;CGGAT;CGCAC;CACGC;CGACG;GCACG;GCGTA;ACGCT;TCGCT; GTCGT;GCGAC;GACGC;CGCAT;CGCTC;CTCGC;GCTCG;CGGTT;CACGT;CGTCA;GCGAT;CGCTA;CCGAA;GTCGA;GACGT;CGTCT;CTCGT;TAGCG;CGACA;CACGA;ATCCG;GCGTT;TACCG;ATGCG;AGTCG;GACGA;ACGTG;TCGTG;TGTCG;CGACT;CTCGA;ACGAG;AGACG;CGCTT;CGTAG;TCGAG;CGGAA;TGACG;ATCGG;CGCAA;TACGG;TTCCG;TTG CG;GCGAA;CGATG;ATCGC;TACGC;AAGCG;ACGTC;AACCG;CGTTG;TCGTC;ATCGT;ACGTA;TACGT;ACACG;TCGTA;TCACG;CGTAC;GTACG;TTCGG;ACGAC;CGTAT;ACGAT;T CGAC;ACTCG;ATCGA;TCGAT;CATCG;TCTCG;TACGA;TTCGC;CGATC;GATCG;ACGTT;AACGG;CGAAG;CTACG;CGATA;TCGTT;TTCGT;AACGC;CGTTC;GTTCG;CGTTA;AACG T;TTCGA;CGATT;CTTCG;CGTAA;CAACG;ACGAA;AACGA;CGTTT;CGAAT;TCGAA;CGAAC;GAACG;ATACG;TATCG;ATTCG;TTACG;CGAAA;AATCG;TAACG;TTTCG; and AAACG. ;
[0139] Cited References and Alternative Embodiments All references cited in this specification are incorporated by reference in their entirety for all purposes to the same extent as if each individual publication or patent or patent application was specifically and individually indicated to be incorporated by reference in its entirety for all purposes.
[0140] The present invention can be implemented as a computer program product that includes a computer program mechanism embedded in a non-transitory computer readable storage medium. For example, the computer program product can include instructions for operating the user interface disclosed herein and described with respect to the figures. These program modules can be stored on a CD-ROM, DVD, magnetic disk storage product, USB key, or other non-transitory computer readable data or program storage product.
[0141] It will be apparent to those skilled in the art that many modifications and variations of this invention can be made without departing from the spirit and scope of the invention. The specific embodiments described herein are provided by way of example only. The embodiments have been chosen and described in order to best explain the principles of the invention and its practical application, and thereby enable those skilled in the art to best utilize the invention and its various modifications as suited to the particular use intended. The present invention is limited only by the terms of the appended claims, along with the full scope of equivalents to which such claims are entitled.
Claims
1. 1. A method for determining the identity and modification state of a nucleic acid molecule, the method comprising: a. immobilizing the nucleic acid on a surface to obtain a nucleic acid attached to the surface; b. Exposing one or more oligos of known sequence to the nucleic acid and detecting binding of the oligos to the nucleic acid to determine the identity of the nucleic acid, wherein one or more or a combination of the oligos can determine the identity of the nucleic acid; c. exposing one or more oligos of known sequence to the nucleic acid molecule, detecting binding of the oligos to the nucleic acid, and measuring the binding characteristics of the oligos, wherein one or more of the oligos can bind differently to a sequence when the sequence is modified compared to when the sequence is not modified; d. assigning a modification status to the molecule of determined identity by evaluating the measured property signature; The method comprising:
2. The method of claim 1, wherein the oligonucleotide comprises a labeled oligonucleotide.
3. The method of claim 2 , wherein the label comprises one or more fluorophores, nanoparticles, proteins, or nanostructures.
4. 2. The method of claim 1, wherein the one or more oligos are 8 or less, 7 or less, 6 or less, 5 or less, 4 or less, or 3 or less nucleotides in length.
5. The method of claim 1, wherein the oligo comprises one or more modifications including an LNA residue, 3'-Uaq, or pyrene.
6. 10. The method of claim 1, wherein said binding of one or more oligos is transient, and each site on each target molecule is capable of transient binding multiple times.
7. The method of claim 1, wherein the difference in binding of the oligonucleotide to the nucleic acid is measured as a function of on-time and off-time and / or fluorescence intensity of the signal.
8. The method of claim 1 , wherein monitoring of binding is performed in real time during the process.
9. The method of claim 1 , wherein multiple oligos are added in a single cycle.
10. The method of claim 1, wherein multiple oligos are added over multiple cycles.
11. The method of claim 1, wherein the same oligo can determine discrimination and determine the modification state.
12. The method of claim 1 , wherein the nucleic acid comprises DNA.
13. 2. The method of claim 1, wherein a plurality of nucleic acids are immobilized at optically resolvable reaction sites on a substrate, and a single nucleic acid at one of the reaction sites is optically resolvable from any other nucleic acid molecule immobilized at any other site.
14. The method of claim 1 , wherein the nucleic acid comprises RNA.
15. The method of claim 1 , wherein the identity of a nucleic acid comprises its genomic origin.
16. 2. The method of claim 1, wherein said determining of distinctiveness is performed by comparing the obtained binding pattern to a database comprising matching in silico binding patterns of segments of a genome.
17. 2. The method of claim 1, wherein the modifications are chemical modifications including 5-methylcytosine (5mC), 5-hydroxymethylcytosine (5hmC), 5-formylcytosine (5fC), 5-carboxylcytosine (5caC), and N6-methadenine (6mA), or other modifications common to nucleic acids found in living organisms.
18. The method of claim 1 , wherein the modification is due to DNA damage.
19. The method of claim 1 , wherein the modification is the binding of a ligand, including a protein, a nucleic acid, a small molecule, or a metal.
20. The method of claim 1 , wherein a spike-in control is used as a reference.
21. The method of claim 1, wherein the repeating bonds differ in their properties, but an average is taken to determine whether the properties assign a modified or unmodified state to the nucleic acid.
22. The method of claim 1, wherein the nucleic acid is subjected to a treatment to alter the modification prior to binding of the oligonucleotide.
23. The method of claim 1, wherein the target is elongated and the position of the modification along its length is identified.
24. The method of claim 1 , wherein the molecules are densely arranged and super-resolution is used to resolve individual molecules.
25. The method of claim 1 , wherein the number of modifications on the molecule is counted or estimated.
26. The method of claim 1, using oligos that can distinguish between different types of modifications.
27. The method of claim 1 , wherein different types of modifications are tested under different conditions.
28. A composition comprising a mixture of short fluorescently labeled oligonucleotide probes designed to detect nucleic acid modifications.
29. 1. A method for determining the modification state of a nucleic acid, comprising: a. immobilizing the nucleic acid on a surface to obtain the nucleic acid at an immobilized location on the surface; b. exposing one or more oligos of known sequence to the nucleic acid molecule, where one or more of the oligos can have a different binding profile when the sequence is modified compared to when the sequence is not modified; c. detecting binding of the oligos to the nucleic acid and determining whether the binding profile more closely matches the binding profile when the sequence is modified or the binding profile when the sequence is unmodified; d. assigning a modification state to said nucleic acid molecule or to one or more positions on said nucleic acid molecule; A method comprising:
30. 1. A method for determining the identity and modification state of a nucleic acid molecule, comprising: a. immobilizing the nucleic acid molecule on a surface to obtain a nucleic acid molecule attached to the surface; b. exposing one or more oligos of known sequence to the nucleic acid, where one or more or a combination of the oligos can determine the identity of each individual nucleic acid molecule, and where one or more of the oligos have a different binding profile when the sequence of the nucleic acid molecule is modified compared to when the sequence is not modified; c. Detecting binding of the oligos to each of one or more distinct nucleic acids to determine the discrimination of the nucleic acids and determining whether the binding profile more closely matches the binding profile when the sequence is modified or the binding profile when the sequence is unmodified; d. recording the modification state of the identified molecules; The method comprising:
31. 31. The method of claim 30, wherein the nucleic acid molecule is a cell-free nucleic acid molecule.
32. 1. A method for determining the sequence and epi-sequence of at least a portion of a nucleic acid molecule, comprising: immobilizing the nucleic acid molecule on a test substrate if the nucleic acid molecule is a single stranded molecule, or if the nucleic acid molecule is a double stranded molecule, denaturing the nucleic acid molecule to a single stranded molecule and immobilizing the single stranded nucleic acid molecule on the test substrate, or if the nucleic acid molecule is a double stranded molecule, immobilizing the nucleic acid molecule on the test substrate and denaturing the nucleic acid molecule on the test substrate to a single stranded molecule, thereby forming a single stranded nucleic acid immobilized on the test substrate; exposing the fixed single-stranded nucleic acid to each oligonucleotide probe species in a set of oligonucleotide probe species, each oligonucleotide probe species of the set of oligonucleotide probe species capable of hybridizing to its complementary portion located at one or more locations on the fixed single-stranded nucleic acid and having (i) a unique respective pre-defined sequence, (ii) a pre-defined length, and (iii) a respective label selected from the group consisting of dyes, fluorescent nanoparticles, plasmon resonant particles, light scattering particles, nanoparticles, and fluorescence resonance energy transfer (FRET) partners (capable of producing a fluorescent signal); said exposing step comprising: i) the oligonucleotide probes of each oligonucleotide probe species of the set of oligonucleotide probe species repeatedly bind transiently and reversibly to the one or more locations on the immobilized single-stranded nucleic acid on the test substrate, thereby forming a respective transient heteroduplex at each of the one or more locations on the immobilized single-stranded nucleic acid on the test substrate; and ii) each generation of optical activity from each of the labels is generated and detected by repeatedly, transiently and reversibly binding the oligonucleotide probes of each of the set of oligonucleotide probe species to the one or more locations on the fixed single-stranded nucleic acid on the test substrate, which are detected at each of the one or more locations on the fixed single-stranded nucleic acid on the test substrate; said exposing being carried out under conditions; determining whether one or more portions of the fixed single-stranded nucleic acid are complementary to each of the oligonucleotide probe species of the set of oligonucleotide probe species by measuring each occurrence of optical activity at each of the one or more locations on the fixed single-stranded nucleic acid on the test substrate that occurs during the exposing step using a two-dimensional imager capable of detecting each occurrence of optical activity generated from the respective labels, thereby obtaining a first set of one or more locations on the fixed single-stranded nucleic acid that are complementary to each of the oligonucleotide probe species of the set of oligonucleotide probe species; washing the test substrate to remove each oligonucleotide probe species of the set of oligonucleotide probe species from the test substrate; repeating the exposing, measuring, and washing steps by exposing the fixed, single-stranded nucleic acid on the test substrate to another respective oligonucleotide probe species of the set of oligonucleotide probe species, thereby obtaining a second set of one or more locations on the fixed, single-stranded nucleic acid that are complementary to another respective oligonucleotide probe species in the set of oligonucleotide probe species; determining a sequence of at least the portion of the nucleic acid based at least in part on the first set of one or more locations on the fixed, single-stranded nucleic acid that are complementary to each oligonucleotide probe species of the set of oligonucleotide probe species and the second set of one or more locations on the fixed, single-stranded nucleic acid that are complementary to another each oligonucleotide probe species of the set of oligonucleotide probe species; determining whether the portion of the nucleic acid molecule has an epigenetic modification based on an observed differential binding behavior of the oligonucleotide probes of the respective oligonucleotide probe species of the set of oligonucleotide probe species to their complementary portions located at one or more locations on the fixed single-stranded nucleic acid when the one or more locations have an epigenetic modification compared to when the one or more locations do not have an epigenetic modification; The method comprising:
33. The method of claim 32, wherein the binding profile of each oligonucleotide capable of binding to a nucleic acid sequence having a possible modification is pre-characterized by testing against synthetic modified and unmodified versions of the nucleic acid sequence having a possible modification, thus serving as a reference for comparing the binding profile obtained for a sample molecule.