Method for simultaneously analyzing genetic information and epigenetic information
By using the same target nucleic acid to simultaneously collect genetic and epigenetic information, the problem of difficulty in detecting gene mutations and gene methylation in the prior art is solved, and high-accurate genome and methylation sequencing is achieved.
Patent Information
- Application Number
- CN202311842967.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-28
- Publication Date
- 2025-07-01
AI Technical Summary
The prior art is difficult to detect gene mutations and gene methylation simultaneously, resulting in complex detection processes, high cost and difficult to apply to clinical practice.
By one method, the same target nucleic acid is used to simultaneously collect genetic and epigenetic information, the specific steps include providing contact with a single-stranded target nucleic acid with a nucleotide mixture containing a modified cytosine, synthesize complementary strands and sequence.
It significantly reduces the probability of false positive events in genome sequencing results, improves the accuracy of sequencing, especially reduces the error rate of genome C>T mutation events, and improves the accuracy of UMI splitting.
Smart Images

Figure BDA0004639687400000261 
Figure BDA0004639687400000271 
Figure BDA0004639687400000281
Abstract
Description
Technical Field
[0001] The present invention relates to a method for preparing double-stranded nucleic acid molecules, a method for preparing a DNA library, and a method for analyzing epigenetic information and / or genetic information. In addition, the present invention also relates to double-stranded nucleic acid molecules and DNA libraries respectively prepared by the said methods. Background Art
[0002] Gene mutation refers to changes in the DNA sequence, including point mutations, insertions, deletions, etc. Gene mutation is an important form of genetic variation and plays an important role in aspects such as the growth, development, and disease occurrence of organisms. Gene mutation detection is of great significance in studying the mechanism of disease occurrence, diagnosing diseases, and formulating individualized treatment plans. For example, in tumor research, gene mutation detection can be used to screen for mutations in tumor-related genes, evaluate the occurrence and development process of tumors, and formulate individualized treatment plans for different mutation types. In genetic disease research, gene mutation detection can be used to determine the disease-causing genes and mutation types of genetic diseases, providing an important basis for the prevention and treatment of familial genetic diseases. Modern gene mutation detection technologies include methods such as Sanger sequencing, PCR amplification sequencing, Next-Generation Sequencing (NGS), etc. These technologies have high sensitivity and specificity, can detect mutations in multiple genes simultaneously, and can diagnose diseases quickly and accurately.
[0003] Gene methylation is an epigenetic modification, which refers to the process in which a methyl group (CH3) on the DNA molecule covalently binds to the 5th carbon atom on the cytosine ring. Gene methylation plays an important role in aspects such as gene expression, cell differentiation, development, and diseases. In particular, it is of great significance for understanding gene regulation mechanisms, studying the mechanism of disease occurrence, evaluating disease risks, and formulating individualized treatment plans.
[0004] At present, the detection of gene mutations and gene methylation is widely used in aspects such as early diagnosis of tumors, tumor typing, prognosis assessment, and monitoring of treatment effects, bringing significant clinical benefits. Gene methylation has received increasing attention due to its good diagnostic performance and tissue tracing performance. However, the detection products on the current market are mainly divided into two categories. One category is based on technologies such as qPCR, Sanger, and NGS to detect gene mutation information, and the other category is based on the above technologies to detect gene methylation information. There are few products based on the fluorescence quantitative PCR methodology that can simultaneously detect gene mutations and gene methylation analysis. The main reason for the above problems is that the existing detection technologies cannot achieve the simultaneous detection of gene mutations and gene methylation in one sample. The current common practice is to detect gene mutations and methylation separately in different tubes, which will consume a large amount of template DNA, and at the same time, the operation complexity and detection cost increase significantly, making it often difficult to apply and popularize clinically.
[0005] Currently, to simultaneously detect gene methylation and gene mutations, it is necessary to construct methylation and mutation libraries separately, which requires more original input DNA samples and two wet experimental procedures for methylation library construction and mutation library construction. However, in special application scenarios such as plasma-free DNA, FFPE small sample DNA, and single-cell DNA, the DNA from the sample source is often very limited and cannot meet the requirements for separate library construction.
[0006] Therefore, providing a method that can achieve the simultaneous detection of genetic information and epigenetic information through one sample / one process will have broad application prospects. Summary of the Invention
[0007] The inventors of the present application have conducted extensive research and obtained a method that can collect genetic information and epigenetic information using the same target nucleic acid. Further, a method for simultaneously performing genome sequencing and methylome sequencing using the same target nucleic acid has been obtained. This method can significantly reduce the probability of false positive events in the genome sequencing results and greatly improve the accuracy of sequencing. Specifically, the error rate of the genomic C>T mutation event is reduced by at least two orders of magnitude, and the correct rate of UMI splitting is also significantly improved. Further, the present application has also made design improvements to the single linker / double linker and sequencing linker used in the method, thereby further improving the accuracy of the method.
[0008] Method for preparing double-stranded nucleic acid molecules
[0009] Therefore, in the first aspect, the present application provides a method for preparing a double-stranded nucleic acid molecule, the method comprising:
[0010] (a-1) Provide at least one single-stranded target nucleic acid, and a nucleotide mixture comprising adenine (A), cytosine (C), guanine (G), and thymine (T), wherein the cytosine comprises or consists of propargyl-modified cytosine;
[0011] (b-1) Contact the single-stranded target nucleic acid with the nucleotide mixture under conditions that permit synthesis of the complementary strand of the single-stranded target nucleic acid.
[0012] In the method of the present invention, under conditions that permit synthesis of the complementary strand of the single-stranded target nucleic acid, after the single-stranded target nucleic acid is contacted with the nucleotide mixture, the complementary strand of the single-stranded target nucleic acid (i.e., the second strand) will be synthesized using the single-stranded target nucleic acid (i.e., the first strand) as a template. Further, since the cytosine in the nucleotide mixture comprises or consists of propargyl-modified cytosine, the complementary strand of the single-stranded target nucleic acid (i.e., the second strand) contains propargyl-modified cytosine. And the cytosine contained in the single-stranded target nucleic acid (i.e., the first strand) is unmodified cytosine. Therefore, in certain embodiments, some or all of the cytosines in the complementary strand of the nucleic acid molecule (i.e., the second strand) are resistant to cytosine conversion. That is, during the cytosine conversion process, the propargyl-modified cytosine in the complementary strand of the nucleic acid molecule (i.e., the second strand) will not change.
[0013] As used herein, the expression "conditions that permit synthesis of the complementary strand of the single-stranded target nucleic acid" has the meaning commonly understood by those skilled in the art, which refers to conditions that permit a nucleic acid polymerase (such as a DNA polymerase) to synthesize another nucleic acid strand using one nucleic acid strand (i.e., the single-stranded target nucleic acid) as a template and form a duplex. Such conditions are well known to those skilled in the art and may involve factors such as temperature, pH value, composition, concentration, and ionic strength of the hybridization buffer. Suitable conditions can be determined by conventional methods (see, for example, Joseph Sambrook, et al., Molecular Cloning, A Laboratory Manual, Cold Spring Harbor Laboratory Press, Cold Spring Harbor, N.Y. (2001)). In the method of the present invention, the "conditions that permit synthesis of the complementary strand of the single-stranded target nucleic acid" are preferably the working conditions of a nucleic acid polymerase (such as a DNA polymerase).
[0014] In certain embodiments, the cytosine further comprises other (e.g., methyl, hydroxymethyl, carboxyl, halogenated) modified cytosines, and / or unmodified cytosine.
[0015] In certain embodiments, the propargyl-modified cytosine further has one or more other (e.g., methyl, hydroxymethyl, carboxyl, halogenated) modifications.
[0016] In certain embodiments, the other modified cytosines are selected from 5-methylcytosine, 5-hydroxymethylcytosine, 5-carboxylpyrimidine, 5-hydroxycytosine, 5-fluorocytosine, 5-chlorocytosine, 5-bromocytosine, 5-iodocytosine, or any combination thereof.
[0017] In certain embodiments, the cytosine comprises or consists of 5-propynylcytosine.
[0018] In certain embodiments, the cytosine comprises or consists of a modified cytosine, wherein the modified cytosine comprises at least 10% (e.g., at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 99%) of 5-propynylcytosine.
[0019] In certain embodiments, the modified cytosine comprises at least 50%, 60%, 70%, 80%, 90%, 95% or 99% of 5-propynylcytosine.
[0020] In certain embodiments, the cytosine comprises or consists of 5-propynylcytosine and 5-hydroxymethylcytosine.
[0021] In certain embodiments, the adenine comprises a modified adenine and / or an unmodified adenine.
[0022] In certain embodiments, the thymine comprises a modified thymine and / or an unmodified thymine.
[0023] In certain embodiments, the guanine comprises a modified guanine and / or an unmodified guanine.
[0024] Target nucleic acid
[0025] In certain embodiments, the target nucleic acid is DNA (e.g., genomic DNA, cfDNA) and / or RNA.
[0026] In certain embodiments, the single-stranded target nucleic acid is a naturally occurring single-stranded nucleic acid or is derived from one strand of a double-stranded nucleic acid.
[0027] In certain embodiments, the method further comprises: obtaining a single-stranded target nucleic acid from a sample, or obtaining a double-stranded target nucleic acid from a sample and preparing it into a single-stranded target nucleic acid.
[0028] Methods for extracting and / or purifying nucleic acids (e.g., genomic DNA) from cells or from biological tissues composed of cells are well known in the art, and those skilled in the art can select appropriate methods or commercial kits according to the biological tissues.
[0029] In some embodiments, extracting nucleic acids (e.g., genomic DNA) from tissues requires cell disruption or cell lysis. In such embodiments, tissue samples can be treated by chemical and physical methods, such as mixing, grinding, or sonication; or membrane lipids can be removed by adding detergents or surfactants for cell lysis, such as removing proteins by adding proteases; for example, removing RNA by adding RNase. Purification of nucleic acids (e.g., genomic DNA) can be carried out by precipitation with ethanol or isopropanol, or by phenol-chloroform extraction.
[0030] Sample
[0031] In certain embodiments, the sample or target nucleic acid is obtained from prokaryotes, eukaryotes (e.g., protozoa, parasites, fungi, yeast, plants, animals including mammals and humans), or viruses (e.g., Herpes virus, HIV, influenza virus, Epstein-Barr virus, hepatitis virus, poliovirus, etc.) or viroids.
[0032] In certain embodiments, the sample is a sample containing cells and / or tissues.
[0033] In certain embodiments, the sample is selected from whole blood, serum, plasma, cerebrospinal fluid, sputum, feces, urine, saliva, sweat, tears, ear discharge, lymph fluid, lavage fluid, bone marrow suspension, vaginal discharge, cervical lavage fluid, ascites, milk, secretions of the respiratory tract, intestine, and urogenital tract, amniotic fluid, or any combination thereof.
[0034] In certain embodiments, the sample is a sample that can be easily obtained by non-invasive methods, such as whole blood, plasma, serum, sweat, tears, sputum, urine, ear discharge, saliva, or excreta. In certain embodiments, the sample is a peripheral blood sample, or the plasma and / or serum fraction of a peripheral blood sample. In certain embodiments, the sample is a swab or smear, biopsy sample, or cell culture. In certain embodiments, the sample is a mixture of two or more samples, for example, the sample can contain two or more of a biological fluid sample, a tissue sample, and a cell culture sample. As used herein, the terms "whole blood", "plasma", and "serum" expressly cover their fractions or processed parts. Similarly, in cases where the sample is taken from a biopsy, swab, smear, etc., "sample" expressly covers the processed fractions or parts derived from the biopsy, swab, smear, etc.
[0035] Although samples are typically taken from a subject (e.g., a patient), they can also be obtained from any mammal (including but not limited to dogs, cats, horses, goats, sheep, cows, pigs, etc.), as well as samples from mixed populations (such as microbial populations from the wild or viral populations from patients).
[0036] The sample can be used directly after being obtained from a biological source or after being pretreated to change the characteristics of the sample. For example, such pretreatment can include preparing plasma from blood, diluting viscous fluids, etc. The methods of pretreatment can also include but are not limited to filtration, precipitation, dilution, distillation, mixing, centrifugation, freezing, lyophilization, concentration, amplification, nucleic acid fragmentation, inactivation of interfering components, addition of reagents, lysis, etc. If such pretreatment methods are used for the sample, then such pretreatment methods generally keep the target nucleic acid in the test sample.
[0037] In certain embodiments, the method further includes, before step (b-1): (a-2) providing one or more adapters and a ligase, and contacting the single-stranded target nucleic acid with the adapter and the ligase under conditions permitting nucleic acid ligation.
[0038] In the method of the present invention, under conditions permitting nucleic acid ligation, the single-stranded target nucleic acid is contacted with the adapter and the ligase, so that the adapter is ligated to the single-stranded target nucleic acid.
[0039] As used herein, "conditions permitting nucleic acid ligation" has the meaning commonly understood by those skilled in the art and can be determined by conventional methods. Such ligation conditions can involve factors such as temperature, pH value, composition, and ionic strength of the hybridization buffer, etc.
[0040] In certain embodiments, the ligase is selected from T4 DNA ligase, T3 DNA ligase, T7 DNA ligase, Escherichia coli DNA ligase, HIFI Taq DNA ligase, T4 RNA ligase, RTCB ligase, circular ligase, thermostable DNA ligase, or any combination thereof.
[0041] Adapter of single-stranded target nucleic acid
[0042] In certain embodiments, the adapter is selected from single-stranded adapters, double-stranded adapters, or any combination thereof.
[0043] In certain embodiments, the 3'-end of the adapter is blocked; for example, by adding a chemical moiety (such as biotin or an alkyl group) to the 3'-OH of the last nucleotide of the single-stranded adapter, by removing the 3'-OH of the last nucleotide of the probe, or by replacing the last nucleotide with a dideoxynucleotide, thereby blocking the 3'-end of the adapter.
[0044] Single-stranded adapter
[0045] In certain embodiments, the single-stranded linker comprises: a first complementary sequence, and a second complementary sequence that is partially or fully complementary to the first complementary sequence.
[0046] In certain embodiments, the linker further comprises: a linking sequence that connects the first complementary sequence and the second complementary sequence.
[0047] In certain embodiments, the linking sequence is located downstream of the first complementary sequence.
[0048] In certain embodiments, the second complementary sequence is located downstream of the linking sequence.
[0049] In certain embodiments, under conditions that permit nucleic acid hybridization or annealing, the linker is in a stem-loop conformation.
[0050] In certain embodiments, the linker contains one or more (e.g., 2, 3, 4, 5) Spacers.
[0051] In certain embodiments, the one or more Spacers are located in the linking sequence of the linker.
[0052] In certain embodiments, the Spacer is selected from Spacer C3, Spacer C6, Spacer C12, Spacer9, Spacer 18, abasic linker (dSpacer), PC linker, or any combination thereof.
[0053] In certain embodiments, the single-stranded linker further comprises one or more (e.g., 2, 3, 4, 5) cleavable moieties.
[0054] As used herein, the term "cleavable moiety" refers to any cleavable or excisable portion contained within a nucleic acid. For example, a cleavable moiety can include uracil, ribonucleotides, or other modified nucleotides that can be excised or cleaved using a nucleic acid cleavage agent (e.g., UDG, RNase, endonuclease, etc.).
[0055] In certain embodiments, the nucleic acid cleavage agent is capable of cleaving the cleavable moiety in the linker.
[0056] In certain embodiments, the nucleic acid cleavage agent is selected from uracil DNA glycosylase (UDG), apurinic / apyrimidinic endonuclease (APE), endonuclease (e.g., endonuclease VIII (EndoVIII) or V (EndoV)), uracil-specific excision reagent (USER) enzyme, formamidopyrimidine DNA glycosylase (Fpg), 8-oxoguanine glycosylase (OGG1), ribonuclease, or any combination thereof.
[0057] It will be appreciated that when the linker contains different cleavable moieties, multiple corresponding nucleic acid cleavage agents can be selected for corresponding cleavage.
[0058] In certain embodiments, when the cleavable moiety contains ribonucleotides (e.g., ribothymidine), the nucleic acid cleavage agent contains ribonuclease (e.g., ribonuclease H (RNase H)).
[0059] In certain embodiments, when the cleavable moiety contains uracil deoxyribonucleic acid, the nucleic acid cleavage agent contains UDG or USER enzyme.
[0060] In certain embodiments, the second complementary sequence contains at least one cleavable moiety, or the 3'-end of the second complementary sequence is linked to at least one cleavable moiety.
[0061] In certain embodiments, the 3'-end of the second complementary sequence is linked to uracil deoxyribonucleic acid or ribonucleotides.
[0062] In certain embodiments, the linker sequence further contains one or more (e.g., 2, 3, 4, 5) cleavable moieties.
[0063] In certain embodiments, multiple cleavable moieties in the linker sequence are located on both sides of the Spacer.
[0064] In certain embodiments, the linker sequence contains two cleavable moieties, and the two cleavable moieties are adjacent to both sides of the Spacer respectively.
[0065] In certain embodiments, the single-stranded linker further contains: unique molecular identifier (UMI).
[0066] As used herein, the term "UMI" or "unique molecular identifier" refers to one or more nucleotide sequences that can be used to identify one or more specific nucleic acids, which can be used to distinguish different nucleotide sequences from each other. For methods of using UMI to identify or distinguish nucleotide sequences, see, for example, Kivioja, Nature Methods 9, 72-74 (2012).
[0067] In certain embodiments, the UMI is selected from random UMI, non-random UMI, or any combination thereof.
[0068] In certain embodiments, the random UMI comprises N bases of 4-28 bp (e.g., 4-10 bp, 10-16 bp, 16-22 bp, 22-28 bp), where N is any one of A, T, C, G.
[0069] In certain embodiments, the non-random UMI comprises multiple (e.g., 4, 8, 16, 32, 64, 96, or more) specific sequences of 4-28 bp (e.g., 4-10 bp, 10-16 bp, 16-22 bp, 22-28 bp).
[0070] In certain embodiments, each non-random UMI differs from other non-random UMIs by at least 1 (e.g., 1, 2, 3, 4) nucleotide at its corresponding sequence position.
[0071] In certain embodiments, the single-linker header comprises a random UMI of 4-10 bp (e.g., 4 bp, 5 bp, 6 bp, 7 bp, 8 bp, 9 bp, 10 bp).
[0072] In certain embodiments, the random UMI is located downstream of the second complementary sequence.
[0073] In certain embodiments, a cleavable moiety is included between the random UMI and the second complementary sequence.
[0074] In certain embodiments, the single-linker header sequentially comprises, from the 5' to 3' direction: a first complementary sequence, a linker sequence, a second complementary sequence, a cleavable moiety, and a UMI.
[0075] Wherein, the linker sequence may or may not include a cleavable moiety.
[0076] In certain embodiments, the linker sequence does not include a cleavable moiety. In certain embodiments, the single-linker header has the sequence shown in SEQ ID NO:1 or SEQ ID NO:8.
[0077] In certain embodiments, the linker sequence includes a cleavable moiety. In certain embodiments, the single-linker header has the sequence shown in SEQ ID NO:7.
[0078] Double-stranded adapter
[0079] In certain embodiments, the dual linker comprises: a first oligonucleotide strand, and a second oligonucleotide strand; wherein, the first oligonucleotide strand comprises a first hybridization sequence, and a first template sequence; the second oligonucleotide strand comprises a second hybridization sequence, and a second template sequence; wherein, the second hybridization sequence is partially or fully complementary to the first hybridization sequence; the second template sequence is partially or fully non-complementary to the first template sequence.
[0080] In certain embodiments, in the first oligonucleotide strand, the first template sequence is located downstream of the first hybridization sequence; preferably, the first template sequence is a free 3'-single-stranded arm.
[0081] In certain embodiments, the end of the 3'-single-stranded arm is blocked; for example, by adding a chemical moiety (such as biotin or an alkyl group) to the 3'-OH of the last nucleotide of the single linker, by removing the 3'-OH of the last nucleotide of the probe, or by replacing the last nucleotide with a dideoxynucleotide, thereby blocking the 3'-end of the linker.
[0082] In certain embodiments, in the second oligonucleotide strand, the second hybridization sequence is located downstream of the second template sequence. In certain embodiments, the second template sequence is a free 5'-single-stranded arm.
[0083] In certain embodiments, the end of the 5'-single-stranded arm is blocked; for example, by adding a chemical moiety (such as biotin or an alkyl group) to the 3'-OH of the last nucleotide of the single linker, by removing the 3'-OH of the last nucleotide of the probe, or by replacing the last nucleotide with a dideoxynucleotide, thereby blocking the 3'-end of the linker.
[0084] In certain embodiments, the dual linker further comprises a cleavable moiety as described above.
[0085] In certain embodiments, the dual linker further comprises a UMI as described above.
[0086] In certain embodiments, the first oligonucleotide strand of the dual linker, in the 5' to 3' direction, sequentially comprises: a first hybridization sequence, a first template sequence, wherein the first template sequence is a free 3'-single-stranded arm.
[0087] In certain embodiments, the second oligonucleotide strand of the dual linker, in the 5' to 3' direction, sequentially comprises: a second template sequence, a second hybridization sequence, a cleavable moiety, a UMI, wherein the second template sequence is a free 5'-single-stranded arm.
[0088] In certain embodiments, the 3'-end of the second oligonucleotide strand is blocked; for example, by adding a chemical moiety (e.g., biotin or an alkyl group) to the 3'-OH of the last nucleotide of the single-stranded linker, by removing the 3'-OH of the last nucleotide of the probe, or by replacing the last nucleotide with a dideoxynucleotide, thereby blocking the 3'-end of the linker.
[0089] In certain embodiments, the double-stranded linker has the sequences shown in SEQ ID NO: 9 and 10.
[0090] Thus, in certain embodiments, the method further comprises: (a-3) providing a nucleic acid cleaving agent and contacting the nucleic acid cleaving agent with the product of (a-2).
[0091] In certain embodiments, after contacting the nucleic acid cleaving agent with the product of (a-2), the nucleic acid cleaving agent is capable of cleaving the cleavable portion in the linker.
[0092] In the method of the present invention, after the nucleic acid cleaving agent cleaves the cleavable portion in the linker, the linker will produce a broken nucleotide strand at the cleavable portion. Thus, the UMI downstream of the cleavable portion and the blocked 3'-end will fall off from the nucleotide strand of the linker accordingly.
[0093] In such embodiments, after cleavage, since there is no blocked 3'-end at the break of the linker, after contacting with the nucleotide mixture, a complementary strand of the single-stranded target nucleic acid (i.e., the first strand) will be synthesized downward at the break using the single-stranded target nucleic acid as a template.
[0094] In certain embodiments, the method is implemented by the following steps (1) to (4):
[0095] (1) Provide a single-stranded target nucleic acid; optionally, the 5'-end of the single-stranded target nucleic acid does not have a free phosphate group;
[0096] (2) Provide one or more of the linkers and a ligase, and under conditions permitting nucleic acid ligation, contact the single-stranded target nucleic acid with the linker and the ligase;
[0097] (3) Provide the nucleic acid cleaving agent, and under conditions permitting the nucleic acid cleaving agent to cleave nucleic acid, contact the product of step (2) with the nucleic acid cleaving agent;
[0098] (4) Provide the nucleotide mixture, and under conditions permitting synthesis of the complementary strand of the single-stranded target nucleic acid, contact the single-stranded target nucleic acid with the nucleotide mixture.
[0099] In certain embodiments, in step (1), the single-stranded target nucleic acid is contacted with alkaline phosphatase so that the 5' end of the single-stranded target nucleic acid does not have a free phosphate group.
[0100] Nucleic acid molecule or its amplification product
[0101] In a second aspect, the present application provides a nucleic acid molecule or an amplification product thereof, which is prepared by the method described in the first aspect.
[0102] In certain embodiments, the amplification product is an amplification product of the first strand of the nucleic acid molecule, and the first strand has the same sequence as the single-stranded target nucleic acid sequence.
[0103] In certain embodiments, the amplification product is an amplification product of the second strand of the nucleic acid molecule, and the second strand has a sequence complementary to the single-stranded target nucleic acid sequence.
[0104] In certain embodiments, the amplification product is an amplification product of the first strand and the second strand of the nucleic acid molecule, wherein the first strand has the same sequence as the single-stranded target nucleic acid sequence, and the second strand has a sequence complementary to the single-stranded target nucleic acid sequence.
[0105] Use of nucleic acid molecule or its amplification product
[0106] In a third aspect, the present application provides the use of the nucleic acid molecule or an amplification product thereof described in the second aspect for cytosine conversion.
[0107] In certain embodiments, the nucleic acid molecule or an amplification product thereof described in the second aspect is used for epigenetic information (e.g., DNA methylation, DNA mutation) analysis.
[0108] In certain embodiments, the nucleic acid molecule or an amplification product thereof described in the second aspect is used for genetic information (e.g., genome) analysis.
[0109] Method for preparing DNA library
[0110] In a fourth aspect, the present application provides a method for preparing a DNA library, the method comprising:
[0111] (i) providing the nucleic acid molecule or an amplification product thereof described in the third aspect;
[0112] (ii) subjecting the nucleic acid molecule or an amplification product thereof to cytosine conversion treatment under conditions that allow unmodified cytosine to be converted to uracil.
[0113] In certain embodiments, the DNA library is selected from a genomic library, a methylome library, or any combination thereof.
[0114] In certain embodiments, the method is used to simultaneously prepare a genomic library and a methylome library separately.
[0115] In some embodiments, the above method further includes: fragmenting the nucleic acid molecule or its amplification product. Those skilled in the art can select a suitable DNA fragmentation method to make the DNA library suitable for sequencing (including but not limited to next-generation sequencing). In certain embodiments, the DNA is fragmented by physical fragmentation, enzymatic fragmentation, and chemical shearing methods to generate fragments of a suitable and / or sufficient length for the DNA library.
[0116] In this article, cytosine conversion can be performed using standard experimental procedures or commercially available kits. In certain embodiments, NaOH is used to denature the nucleic acid molecule or its amplification product. After the nucleic acid molecule or its amplification product is denatured, it can be treated with sodium bisulfite or metabisulfite at a final concentration of, for example, 2 M (pH between about 5 and 6) at 55 °C for 4 - 16 hours. This step covalently modifies unmodified cytosine with sulfite. After conversion, the nucleic acid molecule or its amplification product is desalted and then desulfonated by incubating the nucleic acid molecule or its amplification product at alkaline pH and room temperature, which results in deamination - and conversion to uracil. In one example, commercially available kits such as the EZ DNA Methylation-Gold, EZ DNA Methylation-Direct, or EZ DNA Methylation-Lightning kits (purchased from Zymo Research Corp (Irvine, CA)) are used for bisulfite conversion.
[0117] In certain embodiments, cytosine can also be converted to uracil by methods that do not use bisulfite ions. For example, the irreversible hydrolytic deamination of cytidine and deoxycytidine to uridine and deoxyuridine, respectively, is catalyzed by cytidine deaminase. In some embodiments, the conversion of unmodified cytosine to uracil is accomplished by an enzymatic reaction. In some embodiments, the nucleic acid molecule is incubated with cytidine deaminase. In some embodiments, the cytidine deaminase includes activation-induced cytidine deaminase (AID) and apolipoprotein B mRNA editing enzyme, catalytic polypeptide-like protein (APOBEC). In some embodiments, the APOBEC enzyme is selected from the human APOBEC family, which consists of the following groups: APOBEC-1 (Apo1), APOBEC-2 (Apo2), AID, APOBEC-3A, -3B, -3C, -3DE, -3F, -3G, -3H, and APOBEC-4 (Apo4). In some embodiments, the enzyme is a variant of APOBEC (US20130244237). In some embodiments, the conversion uses a commercially available kit. In one example, a kit such as APOBEC-Seq (NEBiolabs; US20130244237) is used.
[0118] Thus, in certain embodiments, in step (ii), bisulfite is provided to contact the nucleic acid molecule or its amplification product; optionally, cytidine deaminase and / or TET are also provided to contact the nucleic acid molecule or its amplification product.
[0119] In certain embodiments, after step (ii), the first strand of the nucleic acid molecule or its amplification product is enriched.
[0120] In certain embodiments, after step (ii), the second strand of the nucleic acid molecule or its amplification product is enriched.
[0121] In certain embodiments, after step (ii), the first strand and the second strand in the nucleic acid molecule or its amplification product are separated.
[0122] In certain embodiments, the first strand is used to construct a methylome library.
[0123] In certain embodiments, the second strand is used to construct a genomic library.
[0124] Sequencing primer
[0125] In certain embodiments, the method further includes: providing a sequencing primer and contacting it with the product of step (ii) under conditions that permit nucleic acid ligation.
[0126] In certain embodiments, the sequencing primer comprises a first oligonucleotide strand and a second oligonucleotide strand; wherein the first oligonucleotide strand comprises a first hybridization sequence and a first template sequence; the second oligonucleotide strand comprises a second hybridization sequence and a second template sequence; wherein the second hybridization sequence is partially or fully complementary to the first hybridization sequence; the second template sequence is partially or fully non-complementary to the first template sequence.
[0127] In certain embodiments, in the first oligonucleotide strand, the first template sequence is located upstream of the first hybridization sequence. In certain embodiments, the first template sequence is a free 5' single-stranded arm; in certain embodiments, the hybridization sequence of the first oligonucleotide strand is linked to the first strand in a nucleic acid molecule or its amplification product.
[0128] In certain embodiments, in the second oligonucleotide strand, the second hybridization sequence is located upstream of the second template sequence. In certain embodiments, the second template sequence is a free 3' single-stranded arm; in certain embodiments, the hybridization sequence of the second oligonucleotide strand is linked to the second strand in a nucleic acid molecule or its amplification product.
[0129] In certain embodiments, the cytosine in the sequencing primer comprises or consists of a modified cytosine. In certain embodiments, the modified cytosine is selected from 5-propynylcytosine, 5-methylcytosine, 5-hydroxymethylcytosine, 5-carboxypyrimidine, 5-hydroxycytosine, 5-fluorocytosine, 5-chlorocytosine, 5-bromocytosine, 5-iodocytosine, or any combination thereof.
[0130] In certain embodiments, the cytosine in the sequencing primer comprises or consists of 5-propynylcytosine.
[0131] In certain embodiments, the specific sequence of the sequencing primer is adjusted according to the sequencing platform (e.g., BGI, Illumina).
[0132] In certain embodiments, the sequencing primer has the sequences shown in SEQ ID NO: 11 and 12.
[0133] In certain embodiments, the first oligonucleotide strand and / or the second oligonucleotide strand of the sequencing primer further comprises a unique molecular identifier (UMI).
[0134] In certain embodiments, the UMI is in the hybridization sequence of the first oligonucleotide strand and / or the second oligonucleotide strand, or the UMI is located at the end of the hybridization sequence of the first oligonucleotide strand and / or the second oligonucleotide strand.
[0135] In certain embodiments, the UMI is located at the 3' end of the first hybridization sequence. In certain embodiments, the UMI of the first oligonucleotide strand is linked to the first strand in the nucleic acid molecule or its amplification product.
[0136] In certain embodiments, the UMI is located at the 5' end of the second hybridization sequence. In certain embodiments, the UMI of the second oligonucleotide strand is linked to the second strand in the nucleic acid molecule or its amplification product.
[0137] In certain embodiments, the UMI is selected from random UMI, non-random UMI, or any combination thereof.
[0138] In certain embodiments, the random UMI contains N bases of 4 - 28 bp (e.g., 4 - 10 bp, 10 - 16 bp, 16 - 22 bp, 22 - 28 bp), where N is any one of A, T, C, G.
[0139] In certain embodiments, the random UMI contains N bases of 10 - 16 bp (e.g., 10 bp, 11 bp, 12 bp, 13 bp, 14 bp, 15 bp, 16 bp), where N is any one of A, T, C, G.
[0140] In certain embodiments, the non-random UMI contains multiple (e.g., 2 - 4, 4 - 8, 8 - 16, 16 - 32, 32 - 64, 64 - 96, or more) specific sequences of 4 - 28 bp (e.g., 4 - 10 bp, 10 - 16 bp, 16 - 22 bp, 22 - 28 bp).
[0141] In certain embodiments, the non-random UMI contains 2 - 8 (e.g., 2, 3, 4, 5, 6, 7, 8) specific sequences of 4 - 10 bp (e.g., 4 bp, 5 bp, 6 bp, 7 bp, 8 bp, 9 bp, 10 bp).
[0142] In certain embodiments, among the multiple specific sequences contained in the non-random UMI, each specific sequence has at least 1 (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10) nucleotide differences from other specific sequences at their corresponding nucleotide positions.
[0143] In certain embodiments, the specific sequence of the non-random UMI is as shown in any one of SEQ ID NO: 13 - 16.
[0144] In certain embodiments, the first oligonucleotide strand and / or the second oligonucleotide strand of the sequencing primer further comprises a localization tag, wherein the localization tag comprises at least 1 (e.g., 1, 2, 3, 4, 5, 6, 7, 8) specific sequences, and the specific sequences are composed of 3 - 8 (e.g., 3, 4, 5, 6, 7, 8) bases.
[0145] In certain embodiments, the bases are selected from A, T, C, G.
[0146] In certain embodiments, the specific sequences are composed of 3 - 8 identical or different bases.
[0147] In certain embodiments, the localization tag comprises 4 specific sequences, and the 4 specific sequences are all composed of identical or different bases. In certain embodiments, the lengths of the 4 specific sequences are different from each other (e.g., differ by 1 base, differ by 2 bases, differ by 3 bases).
[0148] In certain embodiments, the localization tag comprises a first specific sequence composed of 3 identical or different bases, a second specific sequence composed of 4 identical or different bases, a third specific sequence composed of 5 identical or different bases, and a fourth specific sequence composed of 6 identical or different bases.
[0149] In certain embodiments, the first specific sequence has the sequence shown in SEQ ID NO:17. In certain embodiments, the second specific sequence has the sequence shown in SEQ ID NO:18. In certain embodiments, the third specific sequence has the sequence shown in SEQ ID NO:19. In certain embodiments, the fourth specific sequence has the sequence shown in SEQ ID NO:20.
[0150] In certain embodiments, the localization tag further comprises at least 1 base G at the 5' end of the specific sequence.
[0151] In certain embodiments, the localization tag has the sequence shown in SEQ ID NO:21. In certain embodiments, the localization tag has the sequence shown in SEQ ID NO:4. In certain embodiments, the localization tag has the sequence shown in SEQ ID NO:5. In certain embodiments, the localization tag has the sequence shown in SEQ ID NO:6.
[0152] In certain embodiments, the localization tag is located in the hybridization sequence of the first oligonucleotide strand and / or the second oligonucleotide strand, or the localization tag is located at the end of the hybridization sequence of the first oligonucleotide strand and / or the second oligonucleotide strand.
[0153] In certain embodiments, the positioning tag is located at the 3'-end of the first hybridization sequence. In certain embodiments, the positioning tag of the first oligonucleotide strand is linked to the first strand in the nucleic acid molecule or its amplification product.
[0154] In certain embodiments, the positioning tag is located at the 5'-end of the second hybridization sequence. In certain embodiments, the positioning tag of the second oligonucleotide strand is linked to the second strand in the nucleic acid molecule or its amplification product.
[0155] DNA library
[0156] In another aspect, the present application provides a DNA library constructed by the method as described above in the claims.
[0157] In certain embodiments, the DNA library is selected from a genomic library, a methylome library, or any combination thereof.
[0158] In certain embodiments, a genomic library and a methylome library are simultaneously and separately prepared by the method as described above.
[0159] In certain embodiments, the genomic library comprises or consists of the second strand in the nucleic acid molecule or its amplification product.
[0160] In certain embodiments, the methylome library comprises or consists of the first strand in the nucleic acid molecule or its amplification product.
[0161] Method for analyzing epigenetic information and / or genetic information
[0162] In another aspect, the present application provides a method for analyzing epigenetic information and / or genetic information, the method comprising: detecting and / or analyzing the nucleic acid molecule or its amplification product as described above, or detecting and / or analyzing the DNA library as described above.
[0163] In certain embodiments, the first strand and the second strand in the nucleic acid molecule or its amplification product are separately detected and / or analyzed.
[0164] In certain embodiments, the genomic library and the methylome library are separately detected and / or analyzed.
[0165] In certain embodiments, the detection is sequencing. In certain embodiments, the sequencing is selected from next-generation sequencing, massively parallel sequencing, pyrosequencing, sequencing by synthesis, single molecule real-time sequencing, Polony sequencing, DNA nanoball sequencing, Sunlight microscope single molecule sequencing, nanopore sequencing, Sanger sequencing, Shotgun sequencing, Gilbert sequencing analysis, or any combination thereof.
[0166] In certain embodiments, the detection is selected from microarray, quantitative PCR (qPCR), digital droplet PCR (ddPCR), molecular inversion probe, or any combination thereof.
[0167] In certain embodiments, the analysis is a bioinformatics analysis. In certain embodiments, the bioinformatics analysis includes, but is not limited to: the bioinformatics analysis is selected from sequence alignment, genome analysis, transcriptome analysis, single nucleotide variant (SNV) analysis, gene copy number variant (CNV) analysis, measuring chromosome copy number, detecting genetic lesions.
[0168] In certain embodiments, the bioinformatics analysis can be used to quantify the number of genomic equivalents analyzed in a DNA library (e.g., cfDNA), or to detect the genetic status of a target locus, or to detect genetic lesions in a target locus, or to measure copy number fluctuations within a target locus.
[0169] In certain embodiments, sequence alignment can be performed between the sequence reads obtained by sequencing and one or more reference DNA sequences (e.g., the human genome sequence). In a specific embodiment, the sequencing alignment can be used to detect genetic lesions in a target locus, including but not limited to detecting nucleotide transitions or transversions, nucleotide insertions or deletions, genomic rearrangements, copy number changes, or gene fusions. In certain embodiments, the epigenetic information is DNA methylation information.
[0170] In certain embodiments, the genetic information is whole genome information and / or mutation information.
[0171] Method for distinguishing the first strand (methylation information) / second strand (genomic information)
[0172] In certain embodiments, the nucleic acid molecule or its amplification product is sequenced, and sequencing data containing the first strand and the second strand is obtained, or the sequencing data of the first strand and the second strand is obtained separately.
[0173] In certain embodiments, when the first oligonucleotide strand of the sequencing primer contains a UMI and the second oligonucleotide strand does not contain a UMI, the sequencing data of the first strand and the second strand are distinguished by identifying whether a UMI sequence exists in the sequencing data of the nucleic acid molecule or its amplification product.
[0174] In certain embodiments, the sequencing data in which a UMI is identified is the sequencing data of the first strand. In certain embodiments, the sequencing data in which a UMI is not identified is the sequencing data of the second strand.
[0175] In certain embodiments, when the second oligonucleotide strand of the sequencing primer contains a UMI and the first oligonucleotide strand does not contain a UMI, the sequencing data of the first strand and the second strand are distinguished by identifying whether a UMI sequence is present in the sequencing data.
[0176] In certain embodiments, the sequencing data in which a UMI is identified is the sequencing data of the second strand. In certain embodiments, the sequencing data in which a UMI is not identified is the sequencing data of the first strand.
[0177] In certain embodiments, when the first oligonucleotide strand of the sequencing primer contains a UMI and the second oligonucleotide strand also contains a UMI, the sequencing data of the first strand and the second strand are distinguished by identifying whether the UMI in the sequencing data is closer to the 5' single-stranded arm or the 3' single-stranded arm of the sequencing primer.
[0178] In certain embodiments, the sequencing data in which the identified UMI is closer to the 5' single-stranded arm is the sequencing data of the first strand. In certain embodiments, the sequencing data in which the identified UMI is closer to the 3' single-stranded arm is the sequencing data of the second strand.
[0179] In certain embodiments, the sequencing data of the first strand and the second strand are distinguished by separately identifying whether the cytosine in the UMI of the first oligonucleotide strand and the second oligonucleotide strand has undergone cytosine conversion.
[0180] In certain embodiments, the sequencing data in which the cytosine in the UMI has undergone conversion (e.g., converted to uracil) is the sequencing data of the first strand.
[0181] In certain embodiments, the sequencing data in which the cytosine in the UMI has not undergone cytosine conversion (e.g., remains cytosine) is the sequencing data of the second strand.
[0182] In certain embodiments, when the first oligonucleotide strand of the sequencing primer contains a positioning tag and the second oligonucleotide strand does not contain a positioning tag, the sequencing data of the first strand and the second strand are distinguished by identifying whether a sequence of the positioning tag is present in the sequencing data of the nucleic acid molecule or its amplification product.
[0183] In certain embodiments, the sequencing data in which a positioning tag is identified is the sequencing data of the first strand. In certain embodiments, the sequencing data in which a positioning tag is not identified is the sequencing data of the second strand.
[0184] In certain embodiments, when the second oligonucleotide strand of the sequencing primer contains a positioning tag and the first oligonucleotide strand does not contain a positioning tag, the sequencing data of the first strand and the second strand are distinguished by identifying whether a positioning tag sequence is present in the sequencing data.
[0185] In certain embodiments, the sequencing data in which the localization tag is identified is the sequencing data of the second strand. In certain embodiments, the sequencing data in which the localization tag is not identified is the sequencing data of the first strand.
[0186] In certain embodiments, when the first oligonucleotide strand of the sequencing primer contains a localization tag and the second oligonucleotide strand also contains a localization tag, the sequencing data of the first strand and the second strand are distinguished by identifying whether the localization tag in the sequencing data is closer to the 5' single-stranded arm or the 3' single-stranded arm of the sequencing primer.
[0187] In certain embodiments, the sequencing data in which the identified localization tag is closer to the 5' single-stranded arm is the sequencing data of the first strand. In certain embodiments, the sequencing data in which the identified localization tag is closer to the 3' single-stranded arm is the sequencing data of the second strand.
[0188] In certain embodiments, the sequencing data of the first strand and the second strand are distinguished by separately identifying whether the sequence of the localization tag in the first oligonucleotide strand and the second oligonucleotide strand has changed.
[0189] In certain embodiments, the sequencing data in which guanine (G) in the localization tag is converted to adenine (A) is the sequencing data of the first strand.
[0190] In certain embodiments, the sequencing data in which guanine (G) in the localization tag is not converted is the sequencing data of the second strand.
[0191] In certain embodiments, the sequencing data of the first strand is used for the analysis of the methylation information of the target nucleic acid.
[0192] In certain embodiments, the sequencing data of the second strand is used for the analysis of the genomic information of the target nucleic acid.
[0193] Use of DNA library
[0194] In another aspect, the present application provides the DNA library as described above for the analysis of epigenetic information and / or genetic information.
[0195] In certain embodiments, the DNA library is selected from a genomic library, a methylome library, or any combination thereof.
[0196] In certain embodiments, the epigenetic information is DNA methylation information.
[0197] In certain embodiments, the genetic information is whole-genome information and / or mutation information.
[0198] Use of kit
[0199] On the other hand, the present application provides the use of the nucleic acid molecule or its amplification product as described above, or the DNA library as described above, in the preparation of a kit for detecting whether a subject has a disease.
[0200] In certain embodiments, the disease causes changes in the epigenetic information and / or genetic information of the subject (e.g., nucleotide transitions or transversions, nucleotide insertions or deletions, genomic rearrangements, copy number variations, or gene fusions).
[0201] In certain embodiments, the disease is cancer. In certain embodiments, the disease is a genetic disease.
[0202] In certain embodiments, the kit is used for:
[0203] (1) Detecting whether a subject has cancer;
[0204] (2) Detecting whether cancer has recurred in a subject;
[0205] (3) Detecting the responsiveness of a subject with a disease to a therapy and / or drug received for the disease;
[0206] (4) Detecting whether a subject has a genetic disease.
[0207] Thus, in certain embodiments, the subject to be detected by the kit can be a subject who has received chemotherapy, radiotherapy, or immunotherapy. In certain embodiments, the subject to be detected by the kit can be a subject with cancer recurrence after failure or success of cancer treatment.
[0208] In certain embodiments, the cancer is selected from liver cancer, hepatocellular carcinoma, melanoma, pancreatic cancer, lung cancer, kidney cancer, gastric cancer, esophageal cancer, colon cancer, breast cancer, ovarian cancer, cervical cancer, testicular cancer, prostate cancer, lymphoma, B-cell lymphoma, diffuse large B-cell lymphoma, follicular lymphoma, mantle cell lymphoma, small lymphocytic lymphoma, splenic marginal zone B-cell lymphoma, extranodal marginal zone mucosa-associated lymphoid tissue B-cell lymphoma, nodal marginal zone B-cell lymphoma, lymphoplasmacytic lymphoma, primary effusion lymphoma, Burkitt lymphoma / Burkitt cell leukemia, T-cell lymphoma, anaplastic large cell lymphoma (primary cutaneous type), anaplastic large cell lymphoma, (systemic type), peripheral T-cell lymphoma, angioimmunoblastic T-cell lymphoma, adult T-cell lymphoma / leukemia (human T-cell lymphotropic virus type I positive), extranodal NK / T-cell lymphoma (nasal type), enteropathy-associated T-cell lymphoma, gamma / delta hepatosplenic T-cell lymphoma, subcutaneous panniculitis-like T-cell lymphoma, multiple myeloma, mycosis fungoides.
[0209] In certain embodiments, the genetic disease is selected from Alzheimer's disease (APOE1), Charcot-Marie-Tooth disease, Leber's hereditary optic neuropathy (LHON), Angelman syndrome (UBE3A, ubiquitin-protein ligase E3A), Prader-Willi syndrome (region in chromosome 15), β-thalassemia (HBB, β-globin), Gaucher disease (type I) (GBA, glucocerebrosidase), cystic fibrosis (CFTR epithelial chloride channel), sickle cell disease (HBB, β-globin), phenylketonuria (PAH, phenylalanine hydroxylase), familial hypercholesterolemia (LDLR, low density lipoprotein receptor), Huntington's disease (HDD, huntingtin), neurofibromatosis type I (NF1, NF1 tumor suppressor gene), myotonic dystrophy (DM, coix seed), tuberous sclerosis (TSC1, tuberin), achondroplasia (FGFR3, fibroblast growth factor receptor), fragile X syndrome (FMR1, RNA binding protein), Duchenne muscular dystrophy (DMD, dystrophin), hemophilia A (F8C, coagulation factor VIII), Lesch-Nyhan syndrome (HPRT1, hypoxanthine guanine phosphoribosyltransferase 1), and adrenoleukodystrophy (ABCD1).
[0210] Kit
[0211] In another aspect, the present application provides a kit, which comprises:
[0212] (1) Propargyl-modified cytosine; (2) cytosine conversion reagent (e.g., bisulfite, cytidine deaminase, and / or TET); and (3) single-stranded linker and / or double-stranded linker as described above.
[0213] In certain embodiments, the kit further comprises: (4) reagents for library construction.
[0214] In certain embodiments, the reagents for library construction are selected from enzymes, reagents for DNA end repair, RNase-free water, or any combination thereof.
[0215] In certain embodiments, the enzymes are selected from T4 DNA ligase, T3 DNA ligase, T7 DNA ligase, Escherichia coli DNA ligase, HIFI Taq DNA ligase, T4 RNA ligase, RTCB ligase, circular ligase, thermostable DNA ligase, or any combination thereof.
[0216] Term Definitions
[0217] In the present invention, unless otherwise specified, scientific and technical terms used herein have the meanings commonly understood by those skilled in the art. Also, the virological, biochemical, and immunological laboratory procedures used herein are all conventional procedures widely used in the relevant fields. Meanwhile, to better understand the present invention, definitions and explanations of relevant terms are provided below.
[0218] When the terms "for example", "such as", "such", "including", "comprising" or their variants are used herein, these terms will not be considered restrictive terms, but will be interpreted as meaning "but not limited to" or "not limited to".
[0219] Unless otherwise specified herein or clearly contradictory according to the context, the terms "a" and "an" and "the" and similar referents in the context of describing the present invention (especially in the context of the following claims) shall be construed to cover the singular and the plural.
[0220] As used herein, the terms "polynucleotide", "nucleic acid sequence", "nucleic acid fragment", "oligonucleotide", "nucleic acid" and "nucleic acid molecule" refer to polymeric forms of nucleotides of any length, such as deoxyribonucleotides (dNTPs) or ribonucleotides (rNTPs) or their analogs, and can be used interchangeably. Oligonucleotides are typically composed of a specific sequence of four nucleobases: adenine (A); cytosine (C); guanine (G); and thymine (T) (when the polynucleotide is RNA, thymine (T) is uracil (U)). Non-limiting examples of nucleic acids include DNA, RNA, genomic DNA (e.g., gDNA, such as sheared gDNA), cell-free DNA (e.g., cfDNA), synthetic DNA / RNA, coding or non-coding regions of genes or gene fragments, loci defined by linkage analysis, exons, introns, messenger RNA (mRNA), transfer RNA, ribosomal RNA, short interfering RNA (siRNA), short hairpin RNA (shRNA), microRNA (miRNA), ribozymes, complementary DNA (cDNA), recombinant nucleic acids, branched nucleic acids, plasmids, vectors, isolated DNA of any sequence, isolated RNA of any sequence, nucleic acid probes and primers. Nucleic acids can contain one or more modified nucleotides, such as methylated nucleotides and nucleotide analogs.
[0221] As used herein, the term "amplification" generally refers to the production of one or more copies or extension products of a nucleic acid molecule (e.g., the product of a primer extension reaction on a nucleic acid molecule). Amplification of a nucleic acid molecule can produce single-stranded or multiple copies of the nucleic acid molecule or its complementary sequence that hybridize to the nucleic acid molecule. The amplification product can be a single-stranded or double-stranded nucleic acid molecule generated from an initial template nucleic acid molecule by an amplification procedure. An "amplification product" can include a nucleic acid strand in which at least a portion thereof is substantially identical or substantially complementary to at least a portion of the initial template. In the case where the initial template is a double-stranded nucleic acid molecule, the amplification product can include a nucleic acid strand that is substantially identical to at least a portion of one strand and substantially complementary to at least a portion of either strand. In certain embodiments, the amplification product can be single-stranded or double-stranded. The amplification reaction can be, for example, a polymerase chain reaction (PCR), such as an emulsion polymerase chain reaction (ePCR; e.g., PCR conducted within a microreactor such as a well or droplet).
[0222] As used herein, the term "complementary" means that two nucleic acid sequences are capable of forming hydrogen bonds with each other according to the base pairing principle (Waston-Crick principle) and thereby form a duplex. In the present application, the term "complementary" includes "substantially complementary" and "fully complementary". As used herein, the term "fully complementary" means that each base in one nucleic acid sequence is capable of pairing with a base in the other nucleic acid strand without mismatch or gap. As used herein, the term "substantially complementary" means that most of the bases in one nucleic acid sequence are capable of pairing with the bases in the other nucleic acid strand, allowing for mismatch or gap (e.g., mismatch or gap of one or several nucleotides). Generally, under conditions that permit nucleic acid hybridization, annealing, or amplification, two nucleic acid sequences that are "complementary" (e.g., substantially complementary or fully complementary) will selectively / specifically hybridize or anneal and form a duplex. Accordingly, the term "non-complementary" means that two nucleic acid sequences cannot hybridize or anneal and cannot form a duplex under conditions that permit nucleic acid hybridization, annealing, or amplification. As used herein, the term "incompletely complementary" means that the bases in one nucleic acid sequence cannot pair completely with the bases in the other nucleic acid strand, with at least one mismatch or gap.
[0223] As used herein, the term "unmodified nucleotide" includes naturally occurring nucleotides such as nucleotides containing adenine (A), thymine (T), cytosine (C), uracil (U), or guanine (G).
[0224] As used herein, the term "modified nucleotide" includes, but is not limited to, analogs derived from naturally occurring nucleotides, or nucleotides that have a structural similarity to naturally occurring nucleotides but contain one or more differences. In certain embodiments, modified nucleotides include inosine, diaminopurine, 5-fluorouracil, 5-bromouracil, 5-chlorouracil, 5-iodouracil, hypoxanthine, xanthine, deazaxanthine, deazaguanine, isocytosine, isoguanine, 4-acetylcytosine, 5-(carboxyhydroxymethyl)uracil, 5-carboxymethylaminomethyl-2-thiouridine, 5-carboxymethylaminomethyluracil, dihydrouracil, β-D-galactosylqueosine, N6-isopentenyladenine, 1-methylguanine, 1-methylinosine, 2,2-dimethylguanine, 2-methyladenine, 2-methylguanine, 3-methylcytosine, 5-methylcytosine, N6-adenine, 7-methylguanine, 5-methylaminomethyluracil, 5-methoxyaminomethyl-2-thiouracil, β-D-mannosylqueosine, 5”-methoxycarboxymethyluracil, 5-methoxyuracil, 2-methylthio-D46-isopentenyladenine, uracil-5-oxyacetic acid (v), wybutoxosine, pseudouracil, queosine, 2-thiocytosine, 5-methyl-2-thiouracil, 2-thiouracil, 4-thiouracil, 5-methyluracil, methyl uracil-5-oxyacetate, uracil-5-oxyacetic acid (v), 5-methyl-2-thiouracil, 3-(3-amino-3-N-2-carboxypropyl)uracil, (acp3)w, 2,6-diaminopurine, ethynyl nucleobases, 1-propynyl nucleobases, azido nucleobases, selenophosphate nucleic acids and their modified forms (e.g., by oxidation, reduction, and / or addition of substituents such as alkyl, hydroxyalkyl, hydroxy or halogen moieties).
[0225] As used herein, the term "sequencing primer" generally refers to a molecule (e.g., a polynucleotide) that interacts with a target nucleic acid to facilitate sequencing (e.g., next-generation sequencing (NGS)). In the presence of a sequencing primer, the target nucleic acid can be sequenced by a sequencer. In certain embodiments, a sequencing primer can comprise a nucleotide sequence that hybridizes or binds to a captured target nucleic acid that is attached to a solid support of a sequencing system, such as a bead or a flow cell. In certain embodiments, a sequencing primer can comprise a nucleotide sequence that hybridizes or binds to a target nucleic acid to generate a hairpin loop, thereby allowing the target nucleic acid to be sequenced by a sequencing system. In certain embodiments, a sequencing primer can comprise a sequencer motif, which can be a nucleotide sequence that is complementary to a flow cell sequence of another molecule (e.g., a polynucleotide) and can be used by a sequencing system to sequence a target nucleic acid.
[0226] As used herein, the term "next-generation sequencing (NGS)" refers to a sequencing method that allows for massively parallel sequencing of clonally amplified molecules and individual nucleic acid molecules. Non-limiting examples of NGS include sequencing by synthesis using reversible dye terminators and sequencing by ligation.
[0227] As used herein, the term "Spacer" refers to an arm-like modification. Introducing a spacer into an oligonucleotide is generally to create a distance between oligonucleotides or between an oligonucleotide and other functional groups to avoid steric hindrance, reduce unfavorable interactions between groups, increase flexibility, etc. Alternatively, in cases where oligonucleotide extension is not required, the spacer can be used as a blocking group. Different spacers have different numbers of atoms, and the desired spatial distance can be achieved by adjusting the number and type of inserted spacers. Examples of spacers include, but are not limited to, hydrophobic spacers C3, C6, C12, hydrophilic spacers 9, 18, dSpacer, and PC linker.
[0228] As used herein, the term "UMI" or "unique molecular identifier" refers to one or more nucleotide sequences that can be used to identify one or more specific nucleic acids, which can be used to distinguish different nucleotide sequences from each other. For methods of using UMI to identify nucleotide sequences, see, for example, Kivioja, Nature Methods 9, 72-74 (2012). A UMI can comprise at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 or more nucleotides (e.g., consecutive nucleotides). In high-throughput analysis of multiple samples using next-generation sequencing technology, using UMI can distinguish samples. In certain embodiments, a UMI can be randomly generated (i.e., random UMI) or non-randomly generated (i.e., non-random UMI).
[0229] As used herein, the term "cleavable moiety" refers to any cleavable or excisable portion contained in a nucleic acid. For example, the cleavable moiety may include uracil, ribonucleotides, or other modified nucleotides that can be excised or cleaved using enzymes (e.g., UDG, ribonuclease, endonuclease, etc.). The cleavable moiety may include a spacer, such as a C3 spacer, hexanediol, triethylene glycol spacer (e.g., spacer 9), hexaethylene glycol spacer (e.g., spacer 18), or a combination or analog thereof. The cleavable moiety may include modified nucleotides, such as methylated nucleotides. The modified nucleotides can be specifically recognized by enzymes (e.g., methylated nucleotides can be recognized by MspJI). The cleavable moiety can be enzymatically cleaved (e.g., using enzymes such as UDG, ribonuclease, APE1, MspJI, etc.). The cleavable moiety can be cleaved using stimuli, such as light stimuli, chemical stimuli, thermal stimuli, etc.
[0230] It should be understood that the cleavable moiety can be linked to any position in the nucleic acid molecule. For example, in Figure 1 , a cleavable moiety (e.g., uracil deoxyribonucleic acid) is linked between the second complementary sequence and the UMI. Thus, the cleavable moiety (e.g., uracil, ribonucleotides, spacers, abasic sites, enzyme-specifically modified nucleotides, such as methylated nucleotides, etc.) and any combination thereof, can be linked to the 5'-end or 3'-end of any fragment of the nucleic acid molecule. Similarly, in the case where the nucleic acid molecule is double-stranded, the cleavable moiety can be disposed on different strands. For example, one strand can be linked with a cleavable moiety at the 5'-end, and the other strand can be linked with a cleavable moiety at the 3'-end.
[0231] As used herein, the term "sample" refers to a sample derived from a cell, tissue, organ, or organism, which contains a nucleic acid or a mixture of nucleic acids of at least one nucleic acid sequence suspected of having epigenetic information variation (e.g., methylation modification) and / or genetic information variation (e.g., but not limited to single nucleotide polymorphism, insertion, deletion, and structural variation). In certain embodiments, the sample contains at least one target nucleic acid, the epigenetic information and / or genetics of which are suspected of having undergone variation.
[0232] Advantages of the Invention
[0233] Through extensive research, the inventors of the present application have obtained a method capable of simultaneously collecting genetic information and epigenetic information using the same target nucleic acid. Further, a method for simultaneously performing genome sequencing and methylome sequencing using the same target nucleic acid has been obtained. This method can significantly reduce the probability of false positive events in genome sequencing results and greatly improve the accuracy of sequencing. Specifically, the error rate of genome C>T mutation events has been reduced by at least two orders of magnitude, and the correct splitting rate of UMI has also been significantly improved. Further, the present application has also made design improvements to the single linker / double linker and sequencing linker used in the method, thereby further improving the accuracy of the method.
[0234] The following will describe the implementation embodiments of the present invention in detail with reference to the accompanying drawings and examples. However, those skilled in the art will understand that the following drawings and examples are only for illustrating the present invention and not for limiting the scope of the present invention. According to the following detailed description of the drawings and preferred implementation embodiments, various objects and advantageous aspects of the present invention will become apparent to those skilled in the art. BRIEF DESCRIPTION OF THE DRAWINGS
[0235] Figure 1 FIG. is a schematic structural diagram of the single linker MMssAdapter1TU for single-stranded target nucleic acid. Among them, the single linker sequentially includes from the 5' to 3' direction: a first complementary sequence, a linking sequence connecting the first complementary sequence and the second complementary sequence, a second complementary sequence partially or fully complementary to the first complementary sequence, a cleavable part (e.g., uracil deoxyribonucleic acid) and UMI (e.g., random UMI); wherein, the linking sequence contains a Spacer and does not contain a cleavable part; and the 3'-end of the linker is closed.
[0236] Figure 2 FIG. shows the fragmentation analysis result of the MMLib0710_18 library.
[0237] Figure 3 FIG. shows the fragmentation analysis result of the MMLib0803_11 library.
[0238] Figure 4 FIG. is a schematic structural diagram of the single linker MMssAdapter1TU-4U for single-stranded target nucleic acid. Among them, the single linker sequentially includes from the 5' to 3' direction: a first complementary sequence, a linking sequence connecting the first complementary sequence and the second complementary sequence, a second complementary sequence partially or fully complementary to the first complementary sequence, a cleavable part (e.g., uracil deoxyribonucleic acid) and UMI (e.g., random UMI); wherein, the linking sequence contains a Spacer and contains 2 cleavable parts, and the 2 cleavable parts are respectively adjacent to both sides of the Spacer; and the 3'-end of the linker is closed.
[0239] Figure 5 It is the fragmentation analysis result of the MMLib0726_2 library.
[0240] Figure 6 It is a schematic diagram of the structure of the single-stranded linker MMssAdapter1rTrT for single-stranded target nucleic acid. Among them, the single-stranded linker sequentially includes from the 5'-end to the 3'-end: a first complementary sequence, a connecting sequence connecting the first complementary sequence and the second complementary sequence, a second complementary sequence partially or completely complementary to the first complementary sequence, a cleavable part (for example, thymidine ribonucleic acid), and a UMI (for example, a random UMI); wherein, a Spacer is included in the connecting sequence and no cleavable part is included; wherein, the 3'-end of the linker is blocked.
[0241] Figure 7 It is the fragmentation analysis result of the MMLib0824_5 library.
[0242] Figure 8 It is a schematic diagram of the structure of the double-stranded linker MMssAdapter1TU10 for single-stranded target nucleic acid; among them, the double-stranded linker includes: a first oligonucleotide chain and a second oligonucleotide chain; wherein, the first oligonucleotide chain includes a first hybridization sequence and a first template sequence; the second oligonucleotide chain includes a second hybridization sequence and a second template sequence; wherein, the second hybridization sequence is partially or completely complementary to the first hybridization sequence; the second template sequence is partially or completely non-complementary to the first template sequence;
[0243] Among them, the first oligonucleotide chain of the double-stranded linker sequentially includes from the 5'-end to the 3'-end: a first hybridization sequence, a first template sequence, wherein the first template sequence is a free 3'-single-stranded arm; the end of the 3'-single-stranded arm is blocked;
[0244] Among them, the second oligonucleotide chain of the double-stranded linker sequentially includes from the 5'-end to the 3'-end: a second template sequence, a second hybridization sequence, a cleavable part (for example, uridine deoxyribonucleic acid), and a UMI (for example, a random UMI), wherein the second template sequence is a free 5'-single-stranded arm; the end of the 5'-single-stranded arm is blocked.
[0245] Figure 9 It is the fragmentation analysis result of the MMLib1009_3 library.
[0246] Figure 10 It is the fragmentation analysis result of the MMLib0906_3 library.
[0247] Figure 11Schematic diagram of the structure after the first sequencing primer is ligated to the target nucleic acid. Among them, the first sequencing primer includes a first oligonucleotide chain and a second oligonucleotide chain; among them, the first oligonucleotide chain includes a first hybridization sequence and a first template sequence; the second oligonucleotide chain includes a second hybridization sequence and a second template sequence; among them, the second hybridization sequence is partially or completely complementary to the first hybridization sequence; the second template sequence is partially or completely non-complementary to the first template sequence;
[0248] Among them, the first oligonucleotide chain sequentially includes from the 5' to 3' direction: a template sequence, a hybridization sequence, a random UMI, and a positioning tag; the second oligonucleotide chain sequentially includes from the 5' to 3' direction: a positioning tag, a random UMI, a template sequence, and a hybridization sequence.
[0249] Figure 12 Schematic diagram of the structure after the second sequencing adapter is ligated to the target nucleic acid. Among them, the second sequencing primer includes a first oligonucleotide chain and a second oligonucleotide chain; among them, the first oligonucleotide chain includes a first hybridization sequence and a first template sequence; the second oligonucleotide chain includes a second hybridization sequence and a second template sequence; among them, the second hybridization sequence is partially or completely complementary to the first hybridization sequence; the second template sequence is partially or completely non-complementary to the first template sequence;
[0250] Among them, the first oligonucleotide chain sequentially includes from the 5' to 3' direction: a template sequence, a hybridization sequence, and a non-random UMI; the second oligonucleotide chain sequentially includes from the 5' to 3' direction: a non-random UMI, a template sequence, and a hybridization sequence.
[0251] Figure 13 Schematic diagram of the structure after the third sequencing adapter is ligated to the target nucleic acid. Among them, the third sequencing primer includes a first oligonucleotide chain and a second oligonucleotide chain; among them, the first oligonucleotide chain includes a first hybridization sequence and a first template sequence; the second oligonucleotide chain includes a second hybridization sequence and a second template sequence; among them, the second hybridization sequence is partially or completely complementary to the first hybridization sequence; the second template sequence is partially or completely non-complementary to the first template sequence;
[0252] Among them, the first oligonucleotide chain sequentially includes from the 5' to 3' direction: a template sequence, a hybridization sequence, a random UMI, and a positioning tag; the second oligonucleotide chain sequentially includes from the 5' to 3' direction: a positioning tag, a random UMI, a template sequence, and a hybridization sequence.
[0253] Sequence information
[0254] The description of the sequences involved in this application is provided in the following table.
[0255] Table 1: Sequence Information
[0256]
[0257] Detailed Embodiments
[0258] The present invention will now be described with reference to the following examples which are intended to illustrate, but not limit, the present invention.
[0259] Unless otherwise specified, the molecular biology experimental methods used in the present invention are basically carried out according to the methods described in J. Sambrook et al., Molecular Cloning: A Laboratory Manual, 2nd Edition, Cold Spring Harbor Laboratory Press, 1989, and F. M. Ausubel et al., Current Protocols in Molecular Biology, 3rd Edition, John Wiley & Sons, Inc., 1995; the use of restriction endonucleases is carried out according to the conditions recommended by the product manufacturers. Those skilled in the art will appreciate that the examples describe the present invention by way of illustration and are not intended to limit the scope of the present invention as claimed.
[0260] Example 1.5 Comparison of py-dCTP and 5m-dCTP
[0261] 1.1 5m-dCTP
[0262] Design and synthesize a single-stranded DNA ligation adapter MMssAdapter1TU as shown in SEQ ID NO: 1 (the structure is as Figure 1 shown).
[0263] The design and preparation of the double-stranded adapter with a positioning tag and UMI are carried out according to the content described in the patent "An Improved Method for Making Random Tag Adapters for Next-Generation Sequencing" CN107190067B (the entire content is incorporated herein by reference) and synthesized by Sangon Biotech (Shanghai) Co., Ltd. The specific synthesis is as follows:
[0264] The sequence of the adapter primer P5 is as shown in SEQ ID NO: 2.
[0265] The adapter primer P7-A is obtained by connecting FFFFF(Spacer)EEEEEJJJJJNNNNNNNNNNNN to the 5'-end of SEQ ID NO: 2; the adapter primer P7-B is obtained by connecting FFFFF(Spacer)EEEEEKKKKKNNNNNNNNNNNN to the 5'-end of SEQ ID NO: 2; the adapter primer P7-C is obtained by connecting FFFFF(Spacer)EEEEELLLLLNNNNNNNNNNNN to the 5'-end of SEQ ID NO: 2; SEQ ID NO: 3 is AGATCGGAAGAGCACACGTCTGAACTCCAGTCAC. All cytosines in SEQ ID NO: 2 and SEQ ID NO: 3 are methylated at the 5th carbon atom.
[0266] The adapter after preparation is named MM-UMI.
[0267] For library construction, 30 ng of fragmented DNA of the HCT116 cell line was mixed with 0.005 ng of fully methylated pUC19 and 0.1 ng of completely unmethylated Lambda DNA for library construction. The dephosphorylation treatment system of the template is as follows:
[0268] Number Volume μL DNA sample 12 H2O 16.6 T4 RNA ligation buffer (10x) 8 Tween20 (2%(vol / vol)) 2 rSAP (NEB) 1 Total volume 39.6
[0269] The dephosphorylation program is as follows:
[0270] Temperature Time 37℃ 10min 95℃ 2min Immediately ice-bath after the program ends
[0271] Thereafter, directly prepare the reaction system for Ligation 1 (ligating single-stranded adapter MMssAdapter1TU):
[0272] Number Volume μL Dephosphorylated DNA sample 39.6 PEG 8000 (wt / vol) 27 ATP (100mM) 0.4 T4 DNA Ligase 6 MMssAdapter1TU 1 ddH2O 6 Total volume 80
[0273] The reaction program for Ligation 1 is as follows:
[0274] Temperature Time 22℃ 60min 4℃ Maintain
[0275] Thereafter, add 160 μl of the fragment sorting magnetic beads prepared by our company to the reaction well, and after purification, elute with 25 μl of magnetic beads.
[0276] The eluted product is digested using the USER enzyme digestion system, and the system is as follows:
[0277]
[0278]
[0279] The reaction program of the digestion system is as follows:
[0280] Temperature Time 37℃ 30min 4℃ Maintain
[0281] Subsequently, 60 μl of the fragment sorting magnetic beads prepared by our company was added to the reaction wells. After purification, 6.5 μl of magnetic bead eluate was used.
[0282] Prepare the reaction system for second-strand synthesis:
[0283] Reagent Volume μL Digestion product 6 10×NEBuffer 2 2 dATP (2.5mM) 2 dGTP (2.5mM) 2 dTTP (2.5mM) 2 5m-dCTP (2.5mM) 2 Klenow (exo-) (5U / μL) 2 T4 PNK 2 TotalVolume(uL) 20
[0284] The reaction procedure for second-strand synthesis is as follows:
[0285] Temperature Time 37℃ 30min 95℃ 2min 4℃ 0.1℃ / s
[0286] Prepare the ligation 2 reaction system as follows:
[0287] Reagent Volume μL Second-strand synthesis product 20 MM-UMI (1uM) 1 Ligation buffer (NEB) 15 Ligation enhancer (NEB) 0.5 ddH2O 13.5 Total volume 50
[0288] The reaction procedure for ligation 2 is as follows:
[0289] Temperature Time 20℃ 30min 4℃ Maintain
[0290] Subsequently, 60 μl of the fragment sorting magnetic beads prepared by our company was added to the reaction wells. After purification, 22 μl of magnetic bead eluate was used.
[0291] For methylation conversion, the NEBNext Enzymatic Methyl-seq Conversion Module kit was used. The operation process was carried out according to the kit procedure. Since the operation is the same as the kit instruction manual, it will not be elaborated here.
[0292] The product after conversion was subjected to library pre-amplification. The pre-amplification system is as follows:
[0293]
[0294] The pre-amplification procedure is as follows:
[0295]
[0296] Add 40 μL of the fragment sorting magnetic beads prepared by our company, perform fragment sorting purification, and use 40 μL of ultrapure water for elution after purification.
[0297] The mixed library obtained after elution was quantified using the Qubit HS dsDNA Kit. The results are as follows:
[0298] Library name Library concentration Total library amount MMLib0710_18 44.1ng / μL 1323ng
[0299] Use Qsep for fragment analysis of the library. The results are as Figure 2 shown.
[0300] When analyzing data, the whole methylome data and whole genome data are distinguished by UMI at the R1 end or R2 end of the library data. The distinguishing results and quality control results are as follows:
[0301]
[0302] The evaluation of the conversion efficiency of the methylome data is as follows:
[0303]
[0304] 1.2 5py-dCTP
[0305] According to the single-stranded DNA ligation adapter described in 1.1 and the library construction method described in 1.1, the co-detection library was constructed, and only the following formula was modified when preparing the reaction system for the second-strand synthesis, replacing 5m-dCTP with 5py-dCTP:
[0306] Reagent Volume μL Digestion product 6 10×NEBuffer 2 2 dATP (2.5mM) 2 dGTP (2.5mM) 2 dTTP (2.5mM) 2 5py-dCTP (2.5mM) 2 Klenow (exo-) (5U / μL) 2 T4 PNK 2 TotalVolume(uL) 20
[0307] The MMLib0803_11 mixed library was obtained, and the mixed library was quantified using the Qubit HS dsDNA Kit. The results are as follows:
[0308] Library name Library concentration Total library amount MMLib0803_11 27.2ng / μL 1088ng
[0309] The Qsep was used for the fragmentation analysis of the library, and the results are as Figure 3 shown.
[0310] When analyzing data, the whole methylome data and whole genome data are distinguished by UMI at the R1 end or R2 end of the library data. The distinguishing results and quality control results are as follows:
[0311]
[0312] The evaluation of the conversion efficiency of the methylome data is as follows:
[0313]
[0314] In summary, by the method of this application, using 5py-dCTP when synthesizing the second strand (i.e., ligation two) can significantly reduce the error rate of genomic C>T mutation events, and reduce it by two orders of magnitude! Therefore, for the simultaneous genomic sequencing and methylome sequencing, the method of this application can significantly reduce the probability of false positive events in the genomic sequencing results and greatly improve the accuracy of sequencing.
[0315] Example 2. Comparison of different cytosine derivatives
[0316] This example uses a commercial xGen Prism Library Prep Kit (IDT, Cat. 10009821). The specific experimental operations are carried out as described in the instruction manual, except for corresponding modifications and changes in the adapter part. Briefly, the experiment is as follows:
[0317] The end repair module in the kit is used to perform end repair on the sample DNA, and the first-step ligation is carried out using the ligation-one module in the kit. Among them, the adapter used in ligation-one in the kit is appropriately modified. Specifically, cytosine other than the UMI part in the adapter is replaced with methylated cytosine, and the modified adapter is synthesized by Sangon Biotech. Subsequently, the second-step ligation is carried out on the ligation-one product using the ligation-two system. Among them, the adapter used in ligation-two in the kit is appropriately modified. Specifically, cytosine other than the UMI part in the adapter is replaced with methylated cytosine, and the modified adapter is synthesized by Sangon Biotech. After ligation-two is completed, purification is carried out, and the purified product is used for the synthesis of the protection strand.
[0318] Among them, when synthesizing the second strand (i.e., ligation-two), different cytosine derivatives (i.e., 5m-dCTP, 5hm-dCTP, 5ca-dCTP, or 5py-dCTP) are used in this application to explore their effects on the sequencing results. The specific formulations carried out in this example are shown in the following table:
[0319]
[0320] The reaction program for the synthesis of the protection strand is as follows:
[0321]
[0322]
[0323] Thereafter, 2.5x magnetic beads are used to purify the product of the protection PCR, and the purified product is tried Enzymatic Methyl-seq Conversion Module (NEB). After conversion, the product is subjected to 7 rounds of pre-amplification of library construction, and then sequencing is carried out. The results of the detection accuracy brought by different cytosine derivatives can be clearly seen in the sequencing results as follows:
[0324]
[0325] Meanwhile, the probability of false positive events occurring at the WGS level in the WGS test used as the standard control is 0.04%, which is in the same order of magnitude as the consistency rate when using 5-py-dCTP, and the errors of the remaining modifications are 1-2 orders of magnitude higher. Considering that there are approximately 700 million cytosines at the genomic level, such a high error rate is unacceptable for the detection of mutations and does not have practical clinical application value.
[0326] In summary, compared with other cytosine derivatives, using 5py-dCTP during the synthesis of the second strand (i.e., ligation two) can significantly reduce the probability of false positive events in the genomic sequencing results.
[0327] Example 3. Exploration of single-stranded adapters with different structures
[0328] Design and synthesize the single-stranded DNA ligation adapter MMssAdapter1TU-4U as shown in SEQ ID NO: 7 (the structure is as Figure 4 shown).
[0329] Construct the co-detection library according to the library construction method described in 1.1 of Example 1, and modify the formula as follows during the first ligation process:
[0330] Number Volume μL Dephosphorylated DNA sample 39.6 PEG 8000 (wt / vol) 27 ATP (100mM) 0.4 T4 DNA Ligase 6 MMssAdapter1TU-4U 1 ddH2O 6 Total volume 80
[0331] Obtain the MMLib0726_2 mixed library, and quantify the mixed library using the Qubit HS dsDNAKit. The results are as follows:
[0332] Library name Library concentration Total library amount MMLib0726_2 38.2ng / μL 1910ng
[0333] Use Qsep for fragment analysis of the library, and the results are as Figure 5 .
[0334] During data analysis, distinguish between the whole methylome data and the whole genome data through UMI at the R1 end or the R2 end of the library data. The discrimination results and quality control results are as follows:
[0335]
[0336] The evaluation of the conversion efficiency of the methylome data is as follows:
[0337]
[0338] In summary, using the single linker provided in this application can achieve simultaneous genomic sequencing and methylome sequencing with good accuracy.
[0339] Example 4. Exploration of single-stranded adapters with different structures
[0340] The single-stranded DNA ligation adapter MMssAdapter1rTrT of this embodiment is different in that thymidine ribonucleic acid is used instead of deoxyuridine, as specifically shown in SEQ ID NO: 8, (the structure is as Figure 6 shown).
[0341] The specific process is basically the same as 1.1 in Example 1, with the differences being:
[0342] After ligation 1, the eluted product is digested using an RNaseH digestion system, and the system is as follows:
[0343] Reagent Volume μL Ligation product 23.75 10×RNase H Reaction Buffer 3 RNase H 3.25 Total volume 30
[0344] The reaction program for digestion of the RNase H digestion system is as follows:
[0345] Temperature Time 37℃ 30 min 4℃ Maintain
[0346] The MMLib0824_5 mixed library is obtained, and the mixed library is quantified using the Qubit HS dsDNA Kit, and the results are as follows:
[0347] Library name Library concentration Total library amount MMLib0824_5 35.4 ng / μL 1416 ng
[0348] The fragmented analysis of the library is performed using Qsep, and the results are as Figure 7 shown.
[0349] When analyzing the data, the whole methylome data and the whole genome data are distinguished by whether the UMI is at the R1 end or the R2 end of the library data, and the distinguishing results and the quality control results are as follows:
[0350]
[0351] The evaluation of the conversion efficiency of the methylome data is as follows:
[0352]
[0353] In summary, using the single linker provided in this application, it is possible to simultaneously perform genome sequencing and methylome sequencing and have good accuracy.
[0354] Example 5. Exploration of double adapters
[0355] The double linker MMssAdapter1TU10 of this embodiment, specifically, uses a complementary paired Y-shaped linker instead of a circular linker as the linker used in the first step of ligation. The two strands of the linker are as shown in SEQ ID NO: 9 and 10, and the structure is as Figure 8 shown.
[0356] Mix the two oligonucleotide strands of the fourth linker in equimolar amounts before use.
[0357] The specific process is basically the same as 1.1 in Example 1, except that:
[0358] Remove the process of dephosphorylation using rSAP in the first step, and only retain the denaturation process:
[0359] For library construction, mix 30 ng of fragmented HCT116 cell line DNA, 0.005 ng of fully methylated pUC19, and 0.1 ng of fully unmethylated Lambda DNA for library construction. The dephosphorylation treatment system for the template is as follows:
[0360] Number Volume μL DNA sample 12 H2O 17.6 T4 RNA ligation buffer (10x) 8 Tween20 (2% (vol / vol)) 2 Total volume 39.6
[0361] The denaturation program is as follows:
[0362] Temperature Time 95℃ 2 min Immediately ice-bath after the program ends
[0363] Construct the co-detection library according to the library construction method described in 1.1 of Example 1, and modify the formula in the first step of ligation as follows:
[0364]
[0365]
[0366] Obtain the MMLib1009_3 mixed library, and quantify the mixed library using the Qubit HS dsDNAKit. The results are as follows:
[0367] Library name Library concentration Total library amount MMLib1009_3 36 ng / μL 1440 ng
[0368] Use Qsep for fragment analysis of the library, and the results are as shown in Figure 9.
[0369] When analyzing data, distinguish between the whole methylome data and the whole genome data by whether the UMI is at the R1 end or the R2 end of the library data. The differentiation results and quality control results are as follows:
[0370]
[0371] The evaluation of the conversion efficiency of the methylome data is as follows:
[0372]
[0373] In summary, using the dual linker provided in this application, it is possible to simultaneously perform genome sequencing and methylome sequencing with good accuracy.
[0374] Example 6. Exploration of sequencing adapters
[0375] The specific process is basically the same as 1.1 in Example 1. The difference is that in Ligation 2, a common Y-type adapter for next-generation sequencing is used. The Y-type adapter consists of two partially complementary sequences. Different from the adapters commonly recognized in the industry, the cytosines in this adapter are all methylated cytosines. The sequences of the two strands of the Y-type adapter, M-AdaptorU and M-AdaptorB, are shown in SEQ ID NO:11 and 12 respectively.
[0376] During the experiment, the system of Ligation 2 was modified as follows:
[0377]
[0378]
[0379] The reaction program of Ligation 2 remains unchanged.
[0380] The MMLib0906_3 mixed library was obtained, and the mixed library was quantified using the Qubit HS dsDNA Kit. The results are as follows:
[0381] Library name Library concentration Total library amount MMLib0906_3 26.4 ng / μL 1056 ng
[0382] The fragmented analysis of the library was performed using Qsep, and the results are as Figure 10 shown.
[0383] During data analysis, the full methylome data and the whole genome data were distinguished by whether the UMI was at the R1 end or the R2 end of the library data. The differentiation results and the quality control results are as follows:
[0384]
[0385] The evaluation of the conversion efficiency of the methylome data is as follows:
[0386]
[0387] In summary, using the common Y-type adapter for next-generation sequencing, the method of the present application can achieve simultaneous genome sequencing and methylome sequencing and has good accuracy.
[0388] Example 7. Exploration of sequencing adapters
[0389] The specific process is basically the same as 1.1 in Example 1. The difference is that in Ligation 2, a uniquely designed sequencing adapter is used. The sequencing adapter contains a random UMI and distinguishes the genome and the methylome by the presence of a positioning tag.
[0390] The schematic diagram of the structure after ligating the sequencing adapter to the target nucleic acid is as Figure 11As shown in the figure, a hybrid library of genome and methylome is constructed, where:
[0391] The UMI is 12 random bases NNNNNNNNNNNN, the positioning tags are JJJ, KKKK, LLLLL, EEEEEE, and all Cs are methylated at the 5th carbon atom. The sequencing adapter contains any one of the above four positioning tags, where the positioning tag contains non-random bases. For example, JJJ refers to any combination of three specific bases (ACT in this embodiment, SEQ ID NO:17), KKKK refers to any combination of four specific bases (GACT in this embodiment, SEQ ID NO:18), LLLLL refers to any combination of five specific bases (TGACT in this embodiment, SEQ ID NO:19), and EEEEEE refers to any combination of six specific bases (CTGACT in this embodiment, SEQ ID NO:20).
[0392] The obtained library is sequenced, and the off-machine data distinguishes the genome and methylome according to the method based on the UMI position, specifically as follows:
[0393] The off-machine data is quality-controlled, the sequencing adapter is cut off, and low-quality sequences and short sequences are filtered out.
[0394] The quality-controlled data matches the positioning tags at 12 bases from the starting position in R1 and R2 respectively. When any one of the sequences JJJ, KKKK, LLLLL, EEEEEE is matched, it is a valid match.
[0395] The matching results of all data are classified: both R1 and R2 match the positioning tag, only R1 matches the positioning tag, only R2 matches the positioning tag, and neither R1 nor R2 matches the positioning tag. The data where only R1 matches the positioning tag is methylome data, and the data where only R2 matches the positioning tag is genome data. The data splitting results are as follows:
[0396]
[0397] The split methylome data and genome data are respectively aligned to the reference genome for subsequent analysis. The methylation data evaluates the methylation conversion rate according to lambda DNA, and the analysis results are as follows:
[0398]
[0399] In summary, this embodiment proves that the uniquely designed sequencing adapter and method of this application can accurately distinguish methylome data and genome data.
[0400] Example 8. Exploration of sequencing adapters
[0401] The specific process is basically the same as 1.1 in Example 1. The difference lies in that a uniquely designed sequencing adapter is used in Linkage 2. The sequencing adapter contains a non-random UMI, and the genome and methylome are distinguished by the sequence of the UMI.
[0402] The schematic diagram of the structure after connecting the sequencing adapter to the target nucleic acid is as Figure 12 shown. A mixed library of the genome and methylome is constructed, where:
[0403] The UMI is four sets of fixed sequences ATCGAGTC, CCGTGGAA, ATCATGCG, TGTAGCGT (SEQ ID NO: 13 - 16) with a length of 8 bases.
[0404] The obtained library is sequenced, and the off-machine data distinguishes the genome and methylome according to the method based on the UMI sequence, specifically as follows:
[0405] The off-machine data is quality-controlled, the sequencing adapter is cut off, and low-quality sequences and short sequences are filtered out.
[0406] The first 8 bases of the sequence after quality control are intercepted as the UMI sequence.
[0407] The extracted UMI is aligned with the four sets of designed UMI. When the extracted UMI is consistent with the designed UMI, that is, when the extracted UMI is one of ATCGAGTC, CCGTGGAA, ATCATGCG, TGTAGCGT, the data is genomic data. When the extracted UMI is different from the designed UMI only at the base type at the C position in the designed UMI, and the base type at this position in the extracted UMI is T, that is, when the extracted UMI is one of ATTGAGTT, TTGTGGAA, ATTATGTG, TGTAGTGT, this part of the data is methylome data. The data splitting results are as follows:
[0408] The split methylome data and genomic data are respectively aligned to the reference genome for subsequent analysis. The methylation data evaluates the methylation conversion rate according to lambda DNA, and the analysis results are as follows:
[0409]
[0410] In summary, this example proves that the uniquely designed sequencing adapter and method of this application can accurately distinguish methylome data and genomic data.
[0411] Example 9. Exploration of sequencing adapters
[0412] The specific process is basically the same as 1.1 in Example 1. The difference lies in that a uniquely designed sequencing adapter is used in Linkage 2. The sequencing adapter contains a non-random UMI and a positioning tag, and the genomic and methylated groups are distinguished by the sequence of the positioning tag.
[0413] The schematic structural diagram after ligating the sequencing adapter to the target nucleic acid is as shown in Figure 13 shown. A mixed library of the genomic and methylated groups is constructed, where:
[0414] The UMI is 12 random bases NNNNNNNNNNNN, and the positioning tags are GJJJ, GKKKK, GLLLLL, GEEEEEE. Except for the C on the complementary strand at the first position of the positioning tag, all other Cs are methylated at the 5th carbon atom. The sequencing adapter contains any one of the above four positioning tags. Among them, the positioning tag contains non-random bases. For example, JJJ refers to any combination of three specific bases (ACT in this example). Therefore, the sequence of GJJJ is GACT (SEQ ID NO: 21); KKKK refers to any combination of four specific bases (GACT in this example). Therefore, the sequence of GKKKK is GGACT (SEQ ID NO: 4); LLLLL refers to any combination of five specific bases (TGACT in this example). Therefore, the sequence of GLLLLL is GTGACT (SEQ ID NO: 5); EEEEEE refers to any combination of six specific bases (CTGACT in this example). Therefore, the sequence of GEEEEEE is GCTGACT (SEQ ID NO: 6).
[0415] The obtained library is sequenced, and the off-machine data distinguishes the genomic and methylated groups according to the method based on the positioning tag sequence, specifically as follows:
[0416] The off-machine data is quality-controlled, the sequencing adapter is cut off, and low-quality sequences and short sequences are filtered out.
[0417] The quality-controlled data matches the positioning tag at 12 bases from the starting position in R1 and R2 respectively. When any one of the sequences GJJJ, GKKKK, GLLLLL, GEEEEEE or AJJJ, AKKKK, ALLLLL, AEEEEEE is matched, it is a valid match.
[0418] Classify the matching results of all data: Both R1 and R2 match the positioning label, only R1 matches the positioning label, only R2 matches the positioning label, neither R1 nor R2 matches the positioning label. Among them, both R1 and R2 matching the positioning label are valid data. Further, the part of R2 that matches the positioning label as one of GJJJ, GKKKK, GLLLLL, GEEEEEE is genomic data, and the part of R2 that matches the positioning label as one of AJJJ, AKKKK, ALLLLL, AEEEEEE is methylome data (the first base of the positioning label is glycine G, and its corresponding base in the complementary strand is unmodified cytosine C, which will be converted to U after transformation, so its corresponding complementary base is A). The data splitting results are as follows:
[0419]
[0420] The split methylome data and genomic data are respectively aligned to the reference genome and subsequent analyses are carried out. The methylation data evaluates the methylation conversion rate according to lambda DNA. The analysis results are as follows:
[0421] Methylation data alignment rate Genomic data alignment rate Methylation conversion rate 95.99% 99.62% 99.60%
[0422] In summary, this embodiment proves that the uniquely designed sequencing adapter and method of this application can accurately distinguish methylome data and genomic data.
[0423] Although the specific implementation manners of the present invention have been described in detail, those skilled in the art will understand that various modifications and changes can be made to the details according to all the teachings that have been published, and these changes are all within the protection scope of the present invention. The entire scope of the present invention is given by the appended claims and any equivalents thereof.
Claims
1. A method for preparing a double-stranded nucleic acid molecule, the method comprising: (a-1) providing at least one single-stranded target nucleic acid, and a nucleotide mixture comprising adenine (A), cytosine (C), guanine (G), and thymine (T), wherein the cytosine comprises or consists of propargyl-modified cytosine; (b-1) contacting the single-stranded target nucleic acid with the nucleotide mixture under conditions that permit synthesis of the complementary strand of the single-stranded target nucleic acid; Optionally, the cytosine further comprises other (e.g., methyl, hydroxymethyl, carboxyl, halogenated) modified cytosines, and / or unmodified cytosine; Optionally, the propargyl-modified cytosine further has one or more other (e.g., methyl, hydroxymethyl, carboxyl, halogenated) modifications.
2. The method according to claim 1, wherein, The method has one or more of the following characteristics: (1) The other modified cytosines are selected from 5-methylcytosine, 5-hydroxymethylcytosine, 5-carboxylpyrimidine, 5-hydroxycytosine, 5-fluorocytosine, 5-chlorocytosine, 5-bromocytosine, 5-iodocytosine, or any combination thereof; (2) The cytosine comprises or consists of 5-propargylcytosine; (3) The cytosine comprises or consists of modified cytosine, wherein the modified cytosine comprises at least 10% (e.g., at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 99%) of 5-propargylcytosine; Preferably, the modified cytosine comprises at least 50%, 60%, 70%, 80%, 90%, 95%, or 99% of 5-propargylcytosine; Preferably, the modified cytosine comprises or consists of 5-propargylcytosine and 5-hydroxymethylcytosine; (4) The adenine comprises modified adenine and / or unmodified adenine; (5) The thymine comprises modified thymine and / or unmodified thymine; (6) The guanine comprises modified guanine and / or unmodified guanine; (7) The complementary strand in the nucleic acid molecule is resistant to cytosine conversion; Preferably, the method has one or more of the following characteristics: (1) The target nucleic acid is DNA (e.g., genomic DNA, cfDNA) and / or RNA; (2) The single-stranded target nucleic acid is a naturally occurring single-stranded nucleic acid, or is derived from one nucleic acid strand in a double-stranded nucleic acid; (3) The method further comprises: obtaining the single-stranded target nucleic acid from a sample, or obtaining a double-stranded target nucleic acid from a sample and preparing it into a single-stranded target nucleic acid; Preferably, the sample or target nucleic acid is obtained from a prokaryote, eukaryote (e.g., protozoa, parasite, fungus, yeast, plant, animal including mammals and humans) or virus (e.g., Herpes virus, HIV, influenza virus, Epstein-Barr virus, hepatitis virus, poliovirus, etc.) or viroid; Preferably, the sample is a sample comprising cells and / or tissues; Preferably, the sample is selected from whole blood, serum, plasma, cerebrospinal fluid, sputum, feces, urine, saliva, or any combination thereof.
3. The method according to claim 1 or 2, wherein Before step (b-1), the method further comprises: (a-2) providing one or more adapters and a ligase, and contacting the single-stranded target nucleic acid with the adapter and the ligase under conditions permitting nucleic acid ligation; Preferably, the adapter is selected from single-stranded adapters, double-stranded adapters, or any combination thereof; Preferably, the ligase is selected from T4 DNA ligase, T3 DNA ligase, T7 DNA ligase, Escherichia coli DNA ligase, HIFI Taq DNA ligase, T4 RNA ligase, RTCB ligase, circular ligase, thermostable DNA ligase, or any combination thereof; Optionally, the 3'-end of the adapter is blocked; for example, by adding a chemical moiety (such as biotin or alkyl) to the 3'-OH of the last nucleotide of a single-stranded adapter, by removing the 3'-OH of the last nucleotide of a probe, or by replacing the last nucleotide with a dideoxynucleotide, thereby blocking the 3'-end of the adapter; Preferably, the method further comprises: (a-3) providing a nucleic acid cleavage agent and contacting the nucleic acid cleavage agent with the product of (a-2); Preferably, after contacting the nucleic acid cleavage agent with the product of (a-2), the nucleic acid cleavage agent is capable of cleaving a cleavable moiety in the adapter.
4. The method according to claim 3, wherein The single-stranded adapter comprises: a first complementary sequence, and a second complementary sequence that is partially or fully complementary to the first complementary sequence; Preferably, the adapter further comprises: a linking sequence linking the first complementary sequence and the second complementary sequence; Preferably, the linking sequence is located downstream of the first complementary sequence; Preferably, the second complementary sequence is located downstream of the linking sequence; Preferably, under conditions permitting nucleic acid hybridization or annealing, the adapter is in a stem-loop conformation; Preferably, the adapter contains one or more (such as 2, 3, 4, 5) Spacers; Preferably, the one or more Spacers are located in the linking sequence of the adapter; Preferably, the Spacer is selected from Spacer C3, Spacer C6, Spacer C12, Spacer9, Spacer 18, abasic linker (dSpacer), PC linker, or any combination thereof.
5. The method according to claim 3 or 4, wherein The single-stranded adapter further comprises one or more (such as 2, 3, 4, 5) cleavable moieties; Preferably, after the adapter contacts the nucleic acid cleavage agent, under conditions permitting the nucleic acid cleavage agent to cleave nucleic acid, the nucleic acid cleavage agent is capable of cleaving the cleavable moiety in the adapter; Preferably, the nucleic acid cleavage agent is selected from uracil DNA glycosylase (UDG), apurinic / apyrimidinic endonuclease (APE), endonuclease (such as endonuclease VIII (EndoVIII) or V (EndoV)), uracil-specific excision reagent (USER) enzyme, formamidopyrimidine DNA glycosylase (Fpg), 8-oxoguanine glycosylase (OGG1), ribonuclease, or any combination thereof Preferably, when the cleavable portion comprises ribonucleotides (e.g., ribothymidylic acid), the nucleic acid cleaving agent comprises an RNA enzyme (e.g., ribonuclease H (RNase H)); Preferably, when the cleavable portion comprises deoxyuridylic acid, the nucleic acid cleaving agent comprises UDG or USER enzyme; Preferably, the second complementary sequence comprises at least one cleavable portion, or the 3'-end of the second complementary sequence is linked to at least one cleavable portion; Preferably, the 3'-end of the second complementary sequence is linked to deoxyuridylic acid or ribonucleotides; Optionally, the linking sequence further comprises one or more (e.g., 2, 3, 4, 5) cleavable portions; Preferably, the multiple cleavable portions in the linking sequence are located on both sides of the Spacer; Preferably, the linking sequence comprises 2 cleavable portions, and the 2 cleavable portions are respectively adjacent to both sides of the Spacer.
6. The method according to any one of claims 3 to 5, wherein, The single-stranded linker further comprises: a unique molecular identifier (UMI); Preferably, the UMI is selected from random UMI, non-random UMI, or any combination thereof; Preferably, the random UMI comprises N bases of 4-28 bp (e.g., 4-10 bp, 10-16 bp, 16-22 bp, 22-28 bp), wherein N is any one of A, T, C, G; Preferably, the non-random UMI comprises multiple (e.g., 4, 8, 16, 32, 64, 96, or more) specific sequences of 4-28 bp (e.g., 4-10 bp, 10-16 bp, 16-22 bp, 22-28 bp); Preferably, each non-random UMI differs from other non-random UMIs by at least 1 (e.g., 1, 2, 3, 4) nucleotide at its corresponding sequence position; Preferably, the single-stranded linker comprises a random UMI of 4-10 bp (e.g., 4 bp, 5 bp, 6 bp, 7 bp, 8 bp, 9 bp, 10 bp); Preferably, the random UMI is located downstream of the second complementary sequence; Preferably, a cleavable portion is included between the random UMI and the second complementary sequence.
7. The method according to any one of claims 3-6, wherein, The single-stranded linker sequentially comprises from the 5'- to 3'-direction: a first complementary sequence, a linking sequence, a second complementary sequence, a cleavable portion, and a UMI; Wherein, the linking sequence may or may not comprise a cleavable portion; Preferably, the linking sequence does not comprise a cleavable portion; preferably, the single-stranded linker has the sequence shown in SEQ ID NO:1 or SEQ ID NO:8; Preferably, the linking sequence comprises a cleavable portion; preferably, the single-stranded linker has the sequence shown in SEQ ID NO:
7.
8. The method according to claim 3, wherein The double linker comprises: a first oligonucleotide chain, and a second oligonucleotide chain; wherein, the first oligonucleotide chain comprises a first hybridization sequence and a first template sequence; the second oligonucleotide chain comprises a second hybridization sequence and a second template sequence; wherein, the second hybridization sequence is partially or completely complementary to the first hybridization sequence; the second template sequence is partially or completely non-complementary to the first template sequence; Preferably, in the first oligonucleotide chain, the first template sequence is located downstream of the first hybridization sequence; preferably, the first template sequence is a free 3' single-stranded arm; Preferably, the end of the 3' single-stranded arm is blocked; for example, by adding a chemical moiety (such as biotin or an alkyl group) to the 3'-OH of the last nucleotide of the single linker, by removing the 3'-OH of the last nucleotide of the probe, or by replacing the last nucleotide with a dideoxynucleotide, thereby blocking the 3'-end of the linker; Preferably, in the second oligonucleotide chain, the second hybridization sequence is located downstream of the second template sequence; preferably, the second template sequence is a free 5' single-stranded arm; Preferably, the end of the 5' single-stranded arm is blocked; for example, by adding a chemical moiety (such as biotin or an alkyl group) to the 3'-OH of the last nucleotide of the single linker, by removing the 3'-OH of the last nucleotide of the probe, or by replacing the last nucleotide with a dideoxynucleotide, thereby blocking the 3'-end of the linker; Preferably, the double linker further comprises a cleavable moiety as claimed in claim 5; Preferably, the double linker further comprises a UMI as claimed in claim 6; Preferably, the first oligonucleotide chain of the double linker sequentially comprises, from the 5' to the 3' direction: a first hybridization sequence, a first template sequence, wherein the first template sequence is a free 3' single-stranded arm; Preferably, the second oligonucleotide chain of the double linker sequentially comprises, from the 5' to the 3' direction: a second template sequence, a second hybridization sequence, a cleavable moiety, a UMI, wherein the second template sequence is a free 5' single-stranded arm; Preferably, the 3'-end of the second oligonucleotide chain is blocked; for example, by adding a chemical moiety (such as biotin or an alkyl group) to the 3'-OH of the last nucleotide of the single linker, by removing the 3'-OH of the last nucleotide of the probe, or by replacing the last nucleotide with a dideoxynucleotide, thereby blocking the 3'-end of the linker; Preferably, the double linker has the sequences shown in SEQ ID NO: 9 and 10; 9. The method according to any one of claims 1-8, wherein, The method is achieved by the following steps (1) to (4): (1) Provide a single-stranded target nucleic acid; optionally, the 5'-end of the single-stranded target nucleic acid does not have a free phosphate group; (2) Provide one or more of the linkers and a ligase, and under conditions allowing nucleic acid ligation, bring the single-stranded target nucleic acid into contact with the linker and the ligase; (3) Provide the nucleic acid cleavage agent, and under conditions that allow the nucleic acid cleavage agent to cleave the nucleic acid, contact the product of step (2) with the nucleic acid cleavage agent; (4) Provide the nucleotide mixture, and under conditions that allow the complementary strand of the single-stranded target nucleic acid to be synthesized, contact the single-stranded target nucleic acid with the nucleotide mixture; Preferably, in step (1), contact the single-stranded target nucleic acid with alkaline phosphatase so that the 5'-end of the single-stranded target nucleic acid does not have a free phosphate group.
10. A nucleic acid molecule or its amplification product, which is prepared by the method according to any one of claims 1-9; Preferably, the amplification product is an amplification product of the first strand of the nucleic acid molecule, and the first strand has the same sequence as the single-stranded target nucleic acid sequence; Preferably, the amplification product is an amplification product of the second strand of the nucleic acid molecule, and the second strand has a sequence complementary to the single-stranded target nucleic acid sequence; Preferably, the amplification product is an amplification product of the first strand and the second strand of a nucleic acid molecule, wherein, The first strand has the same sequence as the single-stranded target nucleic acid sequence, and the second strand has a sequence complementary to the single-stranded target nucleic acid sequence.
11. Use of the nucleic acid molecule or its amplification product according to claim 10 for cytosine conversion; Preferably, the nucleic acid molecule or its amplification product according to claim 10 is used for epigenetic information (for example, DNA methylation, DNA mutation) analysis; Preferably, the nucleic acid molecule or its amplification product according to claim 10 is used for genetic information (for example, genome) analysis.
12. A method for preparing a DNA library, the method comprising: (i) Provide the nucleic acid molecule or its amplification product according to claim 10; (ii) Under conditions that allow unmodified cytosine to be converted into uracil, perform cytosine conversion treatment on the nucleic acid molecule or its amplification product; Preferably, the DNA library is selected from a genomic library, a methylome library, or any combination thereof; Preferably, the method is used for simultaneously preparing a genomic library and a methylome library respectively.
13. The method according to claim 12, wherein, The method has one or more of the following features: (1) In step (ii), provide bisulfite and contact it with the nucleic acid molecule or its amplification product; optionally, also provide cytidine deaminase and / or TET and contact them with the nucleic acid molecule or its amplification product; (2) After step (ii), enrich the first strand of the nucleic acid molecule or its amplification product; (3) After step (ii), enrich the second strand of the nucleic acid molecule or its amplification product; (4) After step (ii), separate the first strand and the second strand in the nucleic acid molecule or its amplification product; Preferably, the first strand is used to construct a methylome library; Preferably, the second strand is used to construct a genomic library.
14. The method according to claim 12 or 13, the method further comprising: Provide a sequencing primer, and under conditions that allow nucleic acid ligation, contact it with the product of step (ii); Preferably, the sequencing primer comprises a first oligonucleotide chain and a second oligonucleotide chain; wherein, the first oligonucleotide chain comprises a first hybridization sequence and a first template sequence; the second oligonucleotide chain comprises a second hybridization sequence and a second template sequence; wherein, the second hybridization sequence is partially or completely complementary to the first hybridization sequence; the second template sequence is partially or completely non-complementary to the first template sequence; Preferably, in the first oligonucleotide chain, the first template sequence is located upstream of the first hybridization sequence; preferably, the first template sequence is a free 5' single-stranded arm; preferably, the hybridization sequence of the first oligonucleotide chain is linked to the first strand in a nucleic acid molecule or its amplification product; Preferably, in the second oligonucleotide chain, the second hybridization sequence is located upstream of the second template sequence; preferably, the second template sequence is a free 3' single-stranded arm; preferably, the hybridization sequence of the second oligonucleotide chain is linked to the second strand in a nucleic acid molecule or its amplification product; Preferably, the cytosine in the sequencing primer comprises or consists of a modified cytosine; preferably, the modified cytosine is selected from 5-propynylcytosine, 5-methylcytosine, 5-hydroxymethylcytosine, 5-carboxypyrimidine, 5-hydroxycytosine, 5-fluorocytosine, 5-chlorocytosine, 5-bromocytosine, 5-iodocytosine, or any combination thereof; Preferably, the specific sequence of the sequencing primer is adjusted according to the sequencing platform (e.g., BGI, Illumina); Preferably, the sequencing primer has the sequences shown in SEQ ID NO: 11 and 12; 15. The method according to any one of claims 12-14, wherein, The first oligonucleotide chain and / or the second oligonucleotide chain of the sequencing primer further comprises a unique molecular identifier (UMI); Preferably, the UMI is located in the hybridization sequence of the first oligonucleotide chain and / or the second oligonucleotide chain, or the UMI is located at the end of the hybridization sequence of the first oligonucleotide chain and / or the second oligonucleotide chain; Preferably, the UMI is located at the 3' end of the first hybridization sequence; preferably, the UMI of the first oligonucleotide chain is linked to the first strand in a nucleic acid molecule or its amplification product; Preferably, the UMI is located at the 5' end of the second hybridization sequence; preferably, the UMI of the second oligonucleotide chain is linked to the second strand in a nucleic acid molecule or its amplification product; Preferably, the UMI is selected from random UMI, non-random UMI, or any combination thereof; 16. The method according to any one of claims 12 - 15, wherein, The random UMI comprises N bases of 4-28 bp (e.g., 4-10 bp, 10-16 bp, 16-22 bp, 22-28 bp), wherein N is any one of A, T, C, G; Preferably, the random UMI comprises N bases of 10-16 bp (e.g., 10 bp, 11 bp, 12 bp, 13 bp, 14 bp, 15 bp, 16 bp), wherein N is any one of A, T, C, G; Preferably, the non-random UMI contains multiple (e.g., 2-4, 4-8, 8-16, 16-32, 32-64, 64-96, or more) specific sequences of 4-28 bp (e.g., 4-10 bp, 10-16 bp, 16-22 bp, 22-28 bp); Preferably, the non-random UMI contains 2-8 (e.g., 2, 3, 4, 5, 6, 7, 8) specific sequences of 4-10 bp (e.g., 4 bp, 5 bp, 6 bp, 7 bp, 8 bp, 9 bp, 10 bp); Preferably, among the multiple specific sequences contained in the non-random UMI, each specific sequence has at least 1 (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10) nucleotide differences from other specific sequences at their corresponding nucleotide positions; Preferably, the specific sequence of the non-random UMI is as shown in any one of SEQ ID NO:13-16.
17. The method according to any one of claims 12 - 16, wherein, The first oligonucleotide chain and / or the second oligonucleotide chain of the sequencing primer further contains a positioning tag, wherein the positioning tag contains at least 1 (e.g., 1, 2, 3, 4, 5, 6, 7, 8) specific sequence, and the specific sequence is composed of 3-8 (e.g., 3, 4, 5, 6, 7, 8) bases; Preferably, the bases are selected from A, T, C, G; Preferably, the specific sequence is composed of 3-8 identical or different bases; Preferably, the positioning tag contains 4 specific sequences, and the 4 specific sequences are all composed of the same or different fixed bases; preferably, the lengths of the 4 specific sequences are different from each other (e.g., differ by 1 base, differ by 2 bases, differ by 3 bases); Preferably, the positioning tag contains a first specific sequence composed of 3 identical or different bases, a second specific sequence composed of 4 identical or different bases, a third specific sequence composed of 5 identical or different bases, and a fourth specific sequence composed of 6 identical or different bases; Preferably, the first specific sequence has the sequence as shown in SEQ ID NO:17; preferably, the second specific sequence has the sequence as shown in SEQ ID NO:18; preferably, the third specific sequence has the sequence as shown in SEQ ID NO:19; preferably, the fourth specific sequence has the sequence as shown in SEQ ID NO:20; Preferably, the positioning tag further contains at least 1 base G at the 5' end of the specific sequence; preferably, the positioning tag is located in the hybridization sequence of the first oligonucleotide chain and / or the second oligonucleotide chain, or the positioning tag is located at the end of the hybridization sequence of the first oligonucleotide chain and / or the second oligonucleotide chain; Preferably, the positioning tag is located at the 3' end of the first hybridization sequence; preferably, the positioning tag of the first oligonucleotide chain is linked to the first strand in the nucleic acid molecule or its amplification product; Preferably, the positioning tag is located at the 5' end of the second hybridization sequence; preferably, the positioning tag of the second oligonucleotide chain is linked to the second strand in the nucleic acid molecule or its amplification product.
18. A DNA library constructed by the method according to any one of claims 12-17; Preferably, the DNA library is selected from a genomic library, a methylome library, or any combination thereof; Preferably, a genomic library and a methylome library are simultaneously and separately prepared by the method according to any one of claims 12-17; Preferably, the genomic library comprises or consists of the second strand in the nucleic acid molecule or its amplification product; Preferably, the methylome library comprises or consists of the first strand in the nucleic acid molecule or its amplification product.
19. A method for analyzing epigenetic information and / or genetic information, the method comprising: Detecting and / or analyzing the nucleic acid molecule or its amplification product according to claim 10, or detecting and / or analyzing the DNA library according to claim 18; Preferably, the first strand and the second strand in the nucleic acid molecule or its amplification product are separately detected and / or analyzed; Preferably, the genomic library and the methylome library are separately detected and / or analyzed. The method according to claim 19, wherein The method has one or more of the following characteristics: (1) The detection is sequencing; preferably, the sequencing is selected from next-generation sequencing, massively parallel sequencing, pyrosequencing, sequencing by synthesis, single molecule real-time sequencing, Polony sequencing, DNA nanoball sequencing, SunTag single molecule sequencing, nanopore sequencing, Sanger sequencing, Shotgun sequencing, Gilbert sequencing analysis, or any combination thereof; (2) The detection is selected from microarray, quantitative PCR (qPCR), digital droplet PCR (ddPCR), molecular inversion probe, or any combination thereof; (3) The analysis is bioinformatics analysis; preferably, the bioinformatics analysis is selected from sequence alignment, genomic analysis, transcriptomic analysis, single nucleotide variant (SNV) analysis, gene copy number variant (CNV) analysis, measuring chromosome copy number, detecting genetic lesions, or any combination thereof; (4) The epigenetic information is DNA methylation information; (5) The genetic information is whole genome information and / or mutation information.
21. The method according to claim 20, wherein, Sequencing the nucleic acid molecule or its amplification product, and obtaining sequencing data containing the first strand and the second strand, or separately obtaining the sequencing data of the first strand and the second strand; Preferably, the method has one or more of the following characteristics: (1) When the first oligonucleotide chain of the sequencing primer contains a UMI and the second oligonucleotide chain does not contain a UMI, the sequencing data of the first strand and the second strand are distinguished by identifying whether a UMI sequence exists in the sequencing data of the nucleic acid molecule or its amplification product; Preferably, the sequencing data in which a UMI is identified is the sequencing data of the first strand; preferably, the sequencing data in which a UMI is not identified is the sequencing data of the second strand; (2) When the second oligonucleotide chain of the sequencing primer contains a UMI and the first oligonucleotide chain does not contain a UMI, the sequencing data of the first strand and the second strand are distinguished by identifying whether a UMI sequence exists in the sequencing data; Preferably, the sequencing data with UMI identified is the sequencing data of the second strand; preferably, the sequencing data without UMI identified is the sequencing data of the first strand; (3) When the first oligonucleotide strand of the sequencing primer contains UMI and the second oligonucleotide strand also contains UMI, the sequencing data of the first strand and the second strand are distinguished by identifying whether the UMI in the sequencing data is closer to the 5' single-stranded arm or the 3' single-stranded arm of the sequencing primer; Preferably, the sequencing data with UMI closer to the 5' single-stranded arm identified is the sequencing data of the first strand; preferably, the sequencing data with UMI closer to the 3' single-stranded arm identified is the sequencing data of the second strand; (4) The sequencing data of the first strand and the second strand are distinguished by separately identifying whether the cytosine in the UMI of the first oligonucleotide strand and the second oligonucleotide strand has undergone cytosine conversion; Preferably, the sequencing data with the cytosine in the UMI having undergone conversion (e.g., converted to uracil) is the sequencing data of the first strand; Preferably, the sequencing data with the cytosine in the UMI not having undergone cytosine conversion (e.g., remaining as cytosine) is the sequencing data of the second strand.
22. The method according to claim 20, wherein, The method has one or more of the following characteristics: (1) When the first oligonucleotide strand of the sequencing primer contains a positioning tag and the second oligonucleotide strand does not contain a positioning tag, the sequencing data of the first strand and the second strand are distinguished by identifying whether the sequence of the positioning tag exists in the sequencing data of the nucleic acid molecule or its amplification product; Preferably, the sequencing data with the positioning tag identified is the sequencing data of the first strand; preferably, the sequencing data without the positioning tag identified is the sequencing data of the second strand; (2) When the second oligonucleotide strand of the sequencing primer contains a positioning tag and the first oligonucleotide strand does not contain a positioning tag, the sequencing data of the first strand and the second strand are distinguished by identifying whether the positioning tag sequence exists in the sequencing data; Preferably, the sequencing data with the positioning tag identified is the sequencing data of the second strand; preferably, the sequencing data without the positioning tag identified is the sequencing data of the first strand; (3) When the first oligonucleotide strand of the sequencing primer contains a positioning tag and the second oligonucleotide strand also contains a positioning tag, the sequencing data of the first strand and the second strand are distinguished by identifying whether the positioning tag in the sequencing data is closer to the 5' single-stranded arm or the 3' single-stranded arm of the sequencing primer; Preferably, the sequencing data with the positioning tag closer to the 5' single-stranded arm identified is the sequencing data of the first strand; preferably, the sequencing data with the positioning tag closer to the 3' single-stranded arm identified is the sequencing data of the second strand; (4) The sequencing data of the first strand and the second strand are distinguished by separately identifying whether the sequence of the positioning tag in the first oligonucleotide strand and the second oligonucleotide strand has changed; Preferably, the sequencing data with guanine (G) in the positioning tag converted to adenine (A) is the sequencing data of the first strand; Preferably, the sequencing data with guanine (G) in the positioning tag not having undergone conversion is the sequencing data of the second strand.
23. The method according to claim 21 or 22, wherein, The sequencing data of the first strand is used for the analysis of the methylation information of the target nucleic acid; Preferably, the sequencing data of the second strand is used for the analysis of the genomic information of the target nucleic acid.
24. The DNA library according to claim 18 is used for analyzing epigenetic information and / or genetic information; Preferably, the DNA library is selected from a genomic library, a methylome library, or any combination thereof; Preferably, the epigenetic information is DNA methylation information; Preferably, the genetic information is whole genome information and / or mutation information.
25. Use of the nucleic acid molecule according to claim 10, or its amplification product, or the DNA library according to claim 18 in the preparation of a kit for detecting whether a subject has a disease; Preferably, the disease causes changes in the epigenetic information and / or genetic information of the subject (for example, nucleotide transitions or transversions, nucleotide insertions or deletions, genomic rearrangements, copy number variations or gene fusions); Preferably, the disease is cancer; Preferably, the disease is a genetic disease; Preferably, the kit is used for: (1) Detecting whether a subject has cancer; (2) Detecting whether cancer recurs in a subject; (3) Detecting the responsiveness of a subject with a disease to the therapy and / or drug received for the disease; or, (4) Detecting whether a subject has a genetic disease; Preferably, the cancer is selected from liver cancer, hepatocellular carcinoma, melanoma, pancreatic cancer, lung cancer, kidney cancer, gastric cancer, esophageal cancer, colon cancer, breast cancer, ovarian cancer, cervical cancer, testicular cancer, prostate cancer, lymphoma, B-cell lymphoma, diffuse large B-cell lymphoma, follicular lymphoma, mantle cell lymphoma, small lymphocytic lymphoma, splenic marginal zone B-cell lymphoma, extranodal marginal zone mucosa-associated lymphoid tissue B-cell lymphoma, nodal marginal zone B-cell lymphoma, lymphoplasmacytic lymphoma, primary effusion lymphoma, Burkitt lymphoma / Burkitt cell leukemia, T-cell lymphoma, anaplastic large cell lymphoma (primary cutaneous type), anaplastic large cell lymphoma, (systemic type), peripheral T-cell lymphoma, angioimmunoblastic T-cell lymphoma, adult T-cell lymphoma / leukemia (human T-cell lymphotropic virus type I positive), extranodal NK / T-cell lymphoma (nasal type), enteropathy-associated T-cell lymphoma, gamma / delta hepatosplenic T-cell lymphoma, subcutaneous panniculitis-like T-cell lymphoma, multiple myeloma, mycosis fungoides; Preferably, the genetic disease is selected from Alzheimer's disease (APOE1), Charcot-Marie-Tooth disease, Leber's hereditary optic neuropathy (LHON), Angelman syndrome (UBE3A, ubiquitin-protein ligase E3A), Prader-Willi syndrome (region in chromosome 15), β-thalassemia (HBB, β-globin), Gaucher disease (type I) (GBA, glucocerebrosidase), cystic fibrosis (CFTR epithelial chloride channel), sickle cell disease (HBB, β-globin), phenylketonuria (PAH, phenylalanine hydroxylase), familial hypercholesterolemia (LDLR, low-density lipoprotein receptor), Huntington's disease (HDD, huntingtin), neurofibromatosis type I (NF1, NF1 tumor suppressor gene), myotonic dystrophy (DM, coix seed), tuberous sclerosis (TSC1, tuberin), achondroplasia (FGFR3, fibroblast growth factor receptor), fragile X syndrome (FMR1, RNA-binding protein), Duchenne muscular dystrophy (DMD, dystrophin), hemophilia A (F8C, coagulation factor VIII), Lesch-Nyhan syndrome (HPRT1, hypoxanthine-guanine phosphoribosyltransferase 1), and adrenoleukodystrophy (ABCD1).
26. A kit, the kit comprising: (1) Propargyl-modified cytosine; (2) a cytosine conversion reagent (e.g., bisulfite, cytidine deaminase, and / or TET); and (3) the single-stranded linker and / or double-stranded linker according to any one of claims 4-8; optionally, the kit further comprises: (4) a reagent for library construction; Preferably, the reagent for library construction is selected from enzymes, reagents for DNA end repair, RNase-free water, or any combination thereof; Preferably, the enzyme is selected from T4 DNA ligase, T3 DNA ligase, T7 DNA ligase, Escherichia coli DNA ligase, HIFI Taq DNA ligase, T4 RNA ligase, RTCB ligase, circular ligase, heat-stable DNA ligase, or any combination thereof.
Citation Information
Patent Citations
An Improved Method for Fabricating Random Tag Adapters for Next-Generation Sequencing
CN107190067B
Methods and Compositions for Discrimination Between Cytosine and Modifications Thereof and for Methylome Analysis
US20130244237A1