Method for simultaneously analyzing genetic information and epigenetic information

By preparing double-stranded nucleic acid molecules and using propynyl-modified cytosine and improved linker design, the detection of gene mutations and methylation on the same target nucleic acid is achieved, solving the problems of complexity and high cost of detection in the prior art, and significantly improving sequencing accuracy.

WO2025140181A1PCT designated stage expired Publication Date: 2025-07-03AMOY DIAGNOSTICS CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/141792
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-28
Filing Date
2024-12-24
Publication Date
2025-07-03

AI Technical Summary

Technical Problem

The existing technology cannot efficiently detect gene mutations and gene methylation at the same time, resulting in high cost of sample detection and complex operation, making it difficult to promote clinically, especially in special application scenarios such as plasma free DNA, FFPE small sample DNA, and single-cell DNA.

Method used

By preparing double-stranded nucleic acid molecules and using the same target nucleic acid to collect genetic and epigenetic information simultaneously, using propynyl-modified cytosine and improved single-link head/double-link head design, the probability of false positive events is reduced and the sequencing accuracy is improved.

Benefits of technology

It significantly reduces the probability of false positive events in genome sequencing results and improves sequencing accuracy. In particular, the error rate of genome C>T mutation events is reduced by at least two orders of magnitude, and the accuracy rate of UMI splitting is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure PCTCN2024141792-FTAPPB-I100001
    Figure PCTCN2024141792-FTAPPB-I100001
  • Figure PCTCN2024141792-FTAPPB-I100002
    Figure PCTCN2024141792-FTAPPB-I100002
  • Figure PCTCN2024141792-FTAPPB-I100003
    Figure PCTCN2024141792-FTAPPB-I100003
Patent Text Reader

Abstract

The present invention relates to a method for preparing a double-stranded nucleic acid molecule, a method for preparing a DNA library, and a method for analyzing epigenetic information and / or genetic information. In addition, the present invention further relates to a double-stranded nucleic acid molecule and a DNA library respectively prepared by the methods.
Need to check novelty before this filing date? Find Prior Art

Description

Methods for simultaneous analysis of genetic and epigenetic information

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims priority to Chinese patent application No. 202311842967.X filed on December 28, 2023, the entire contents of which are incorporated herein by reference in their entirety. Technical Field

[0003] The present invention relates to a method for preparing a double-stranded nucleic acid molecule, a method for preparing a DNA library, and a method for analyzing epigenetic information and / or genetic information. Furthermore, the present invention also relates to the double-stranded nucleic acid molecule and DNA library prepared by the methods, respectively. Background Art

[0004] Gene mutation refers to changes in the DNA sequence, including point mutations, insertions, and deletions. Gene mutation is an important form of genetic variation that plays a vital role in the growth, development, and disease occurrence of organisms. Gene mutation detection is of great significance for studying the mechanisms of disease occurrence, diagnosing diseases, and formulating personalized treatment plans. For example, in tumor research, gene mutation detection can be used to screen mutations in tumor-related genes, evaluate the occurrence and development of tumors, and formulate personalized treatment plans for different mutation types. In genetic disease research, gene mutation detection can be used to identify the causative genes and mutation types of genetic diseases, providing an important basis for the prevention and treatment of familial genetic diseases. Modern gene mutation detection technologies include Sanger sequencing, PCR amplification sequencing, and Next-Generation Sequencing (NGS). These technologies are highly sensitive and specific, can detect mutations in multiple genes simultaneously, and can diagnose diseases quickly and accurately.

[0005] Gene methylation is an epigenetic modification that occurs when a methyl group (CH3) on a DNA molecule covalently bonds with the carbon atom at position 5 of the cytosine ring. Gene methylation plays a crucial role in gene expression, cell differentiation, development, and disease. It is particularly important for understanding gene regulation, studying disease mechanisms, assessing disease risk, and developing personalized treatment plans.

[0006] At present, gene mutation and gene methylation detection are widely used in the early diagnosis of tumors, tumor typing, prognosis assessment and treatment effect monitoring, and have brought significant clinical benefits. Gene methylation has received increasing attention due to its good diagnostic performance and tissue traceability performance. However, the detection products currently on the market are mainly divided into two categories. One is based on qPCR, Sanger and NGS technology to detect gene mutation information, and the other is based on the above technologies to detect gene methylation information. There are few products based on fluorescent quantitative PCR methodology that can perform gene mutation and gene methylation detection and analysis at the same time. The main reason for the above problems is that the existing detection technology cannot detect gene mutations and gene methylation in one sample at the same time. The current more common practice is to test gene mutations and methylation separately. This process consumes a large amount of template DNA, and the operation complexity and detection cost are significantly increased. It is often difficult to apply and promote in clinical practice.

[0007] Currently, simultaneous gene methylation and mutation detection requires the construction of separate methylation and mutation libraries, requiring more raw DNA samples and two wet lab processes: methylation library construction and mutation library construction. However, in specialized applications such as plasma cell-free DNA, FFPE small sample DNA, and single-cell DNA, the sample source DNA is often very limited, making it impossible to construct separate libraries.

[0008] Therefore, providing a method that can simultaneously detect genetic information and epigenetic information through one sample / one process will have broad application prospects. Summary of the Invention

[0009] After extensive research, the inventors of this application have obtained a method for simultaneously collecting genetic and epigenetic information using the same target nucleic acid, and further obtained a method for simultaneously performing genome sequencing and methylation group sequencing using the same target nucleic acid. This method can significantly reduce the probability of false positive events that occur in genome sequencing results, greatly improving the accuracy of sequencing. Specifically, the error rate of genome C>T mutation events is reduced by at least two orders of magnitude, and the accuracy of UMI splitting is also significantly improved. Furthermore, this application has also designed and improved the single-chain adapter / double-chain adapter and sequencing adapter used in the method, thereby further improving the accuracy of the method.

[0010] Method for preparing double-stranded nucleic acid molecules

[0011] Therefore, in a first aspect, the present application provides a method for preparing a double-stranded nucleic acid molecule, the method comprising:

[0012] (a-1) providing at least one single-stranded target nucleic acid and a nucleotide mixture comprising adenine (A), cytosine (C), guanine (G), and thymine (T), wherein the cytosine comprises or consists of a propynyl-modified cytosine;

[0013] (b-1) contacting the single-stranded target nucleic acid with the nucleotide mixture under conditions that allow synthesis of a complementary strand to the single-stranded target nucleic acid.

[0014] In the method of the present invention, under conditions allowing the synthesis of the complementary chain of the single-stranded target nucleic acid, after the single-stranded target nucleic acid is contacted with the nucleotide mixture, the single-stranded target nucleic acid (i.e., the first chain) is used as a template to synthesize the complementary chain of the single-stranded target nucleic acid (i.e., the second chain). Further, since the cytosine in the nucleotide mixture contains or consists of a propynyl-modified cytosine, the complementary chain of the single-stranded target nucleic acid (i.e., the second chain) contains a propynyl-modified cytosine. The cytosine contained in the single-stranded target nucleic acid (i.e., the first chain) is a cytosine without modification. Therefore, in certain embodiments, some or all of the cytosines of the complementary chain (i.e., the second chain) of the nucleic acid molecule are resistant to cytosine conversion. That is, during the cytosine conversion process, the propynyl-modified cytosine in the complementary chain (i.e., the second chain) of the nucleic acid molecule does not change.

[0015] As used herein, the expression "conditions that allow synthesis of the complementary strand of the single-stranded target nucleic acid" has the meaning generally understood by those skilled in the art, and refers to conditions that allow a nucleic acid polymerase (e.g., a DNA polymerase) to synthesize another nucleic acid strand using one nucleic acid strand (i.e., the single-stranded target nucleic acid) as a template and form a duplex. Such conditions are well known to those skilled in the art and may involve factors such as temperature, the pH value, composition, concentration, and ionic strength of the hybridization buffer. Suitable conditions can be determined by conventional methods (see, for example, Joseph Sambrook, et al., Molecular Cloning, A Laboratory Manual, Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY (2001)). In the methods of the present invention, "conditions that allow synthesis of the complementary strand of the single-stranded target nucleic acid" are preferably working conditions for a nucleic acid polymerase (e.g., a DNA polymerase).

[0016] In certain embodiments, the cytosine further comprises other (eg, methyl, hydroxymethyl, carboxyl, halo) modified cytosines, and / or unmodified cytosines.

[0017] In certain embodiments, the propynyl-modified cytosine further has one or more other (eg, methyl, hydroxymethyl, carboxyl, halo) modifications.

[0018] In certain embodiments, the other modified cytosine is selected from 5-methylcytosine, 5-hydroxymethylcytosine, 5-carboxypyrimidine, 5-hydroxycytosine, 5-fluorocytosine, 5-chlorocytosine, 5-bromocytosine, 5-iodocytosine, or any combination thereof.

[0019] In certain embodiments, the cytosine comprises or consists of 5-propynylcytosine.

[0020] In certain embodiments, the cytosine comprises or consists of a modified cytosine, wherein the modified cytosine comprises at least 10% (e.g., at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 99%) 5-propynylcytosine.

[0021] In certain embodiments, the modified cytosine comprises at least 50%, 60%, 70%, 80%, 90%, 95%, or 99% 5-propynylcytosine.

[0022] In certain embodiments, the cytosine comprises or consists of 5-propynylcytosine and 5-hydroxymethylcytosine.

[0023] In certain embodiments, the adenine comprises modified adenine and / or unmodified adenine.

[0024] In certain embodiments, the thymine comprises modified thymine and / or unmodified thymine.

[0025] In certain embodiments, the guanine comprises modified guanine and / or unmodified guanine.

[0026] Target nucleic acid

[0027] In certain embodiments, the target nucleic acid is DNA (e.g., genomic DNA, cfDNA) and / or RNA.

[0028] In certain embodiments, the single-stranded target nucleic acid is a naturally occurring single-stranded nucleic acid, or is derived from one nucleic acid strand of a double-stranded nucleic acid.

[0029] In certain embodiments, the method further comprises: obtaining a single-stranded target nucleic acid from the sample, or obtaining a double-stranded target nucleic acid from the sample and preparing the target nucleic acid into a single-stranded target nucleic acid.

[0030] Methods for extracting and / or purifying nucleic acids (eg, genomic DNA) from cells or from biological tissues composed of cells are well known in the art, and technicians can select appropriate methods or commercial kits based on the biological tissue.

[0031] In some embodiments, extraction of nucleic acids (e.g., genomic DNA) from tissues requires cell disruption or cell lysis. In such embodiments, chemical and physical methods can be used, such as mixing, grinding, or sonication of tissue samples; membrane lipids can also be removed by adding detergents or surfactants for cell lysis, such as by adding proteases to remove proteins; and RNA can be removed, such as by adding RNases. Nucleic acid (e.g., genomic DNA) purification can be performed by precipitation with ethanol or isopropanol, or by phenol-chloroform extraction.

[0032] sample

[0033] In certain embodiments, the sample or target nucleic acid is obtained from a prokaryotic organism, a eukaryotic organism (e.g., a protozoan, a parasite, a fungus, a yeast, a plant, an animal including a mammal and a human) or a virus (e.g., Herpes virus, HIV, influenza virus, Epstein-Barr virus, hepatitis virus, polio virus, etc.) or a viroid.

[0034] In certain embodiments, the sample is a sample comprising cells and / or tissue.

[0035] In certain embodiments, the sample is selected from whole blood, serum, plasma, cerebrospinal fluid, sputum, feces, urine, saliva, sweat, tears, ear discharge, lymph, lavage fluid, bone marrow suspension, vaginal discharge, transcervical lavage, ascites, breast milk, secretions of the respiratory, intestinal and urogenital tracts, amniotic fluid, or any combination thereof.

[0036] In certain embodiments, sample is a sample that can be easily obtained by non-invasive methods, such as whole blood, blood plasma, serum, sweat, tears, sputum, urine, ear discharge, saliva or excreta. In certain embodiments, sample is a peripheral blood sample or the plasma and / or serum fractions of a peripheral blood sample. In certain embodiments, sample is a swab or smear, biopsy specimen or cell culture. In certain embodiments, sample is a mixture of two or more samples, such as sample can include two or more of a biofluid sample, a tissue sample and a cell culture sample. As used herein, the terms "whole blood", "blood plasma" and "serum" clearly encompass its fraction or processed part. Similarly, when the sample is taken from a biopsy, swab, smear, etc., "sample" clearly encompasses the processed fraction or part derived from a biopsy, swab, smear, etc.

[0037] Although samples are typically obtained from a subject (e.g., a patient), they can also be obtained from any mammal (including but not limited to dogs, cats, horses, goats, sheep, cattle, pigs, etc.), as well as mixed populations (such as microbial populations from the wild or viral populations from patients).

[0038] The sample can be directly used after obtaining from a biological source, or can be used after pre-treatment to change the characteristics of the sample. For example, this type of pre-treatment can include preparing plasma from blood, diluting viscous fluids, etc. The pre-treated method can also include but is not limited to filtering, precipitation, dilution, distillation, mixing, centrifugation, freezing, lyophilization, concentration, amplification, nucleic acid fragmentation, the inactivation of interfering components, the addition of reagents, cracking, etc. If this type of pre-treatment method is adopted for sample, then this type of pre-treatment method usually makes the target nucleic acid remain in the test sample.

[0039] In certain embodiments, the method further comprises, before step (b-1): (a-2) providing one or more linkers and a ligase, and contacting the single-stranded target nucleic acid with the linkers and the ligase under conditions allowing nucleic acid ligation.

[0040] In the method of the present invention, the single-stranded target nucleic acid is contacted with the adaptor and a ligase under conditions that allow nucleic acid ligation, thereby ligating the adaptor to the single-stranded target nucleic acid.

[0041] As used herein, "conditions that allow nucleic acid ligation" have the meaning generally understood by those skilled in the art and can be determined by conventional methods. Such ligation conditions may involve factors such as temperature, pH, composition, and ionic strength of the hybridization buffer.

[0042] In certain embodiments, the ligase is selected from T4 DNA ligase, T3 DNA ligase, T7 DNA ligase, E. coli DNA ligase, HIFI Taq DNA ligase, T4 RNA ligase, RTCB ligase, ring ligase, thermostable DNA ligase, or any combination thereof.

[0043] Adapters for single-stranded target nucleic acids

[0044] In certain embodiments, the linker is selected from a single-stranded linker, a double-stranded linker, or any combination thereof.

[0045] In certain embodiments, the 3'-end of the linker is blocked; for example, by adding a chemical moiety (e.g., biotin or an alkyl group) to the 3'-OH of the last nucleotide of the single-stranded linker, by removing the 3'-OH of the last nucleotide of the probe, or replacing the last nucleotide with a dideoxynucleotide.

[0046] Single-link connector

[0047] In certain embodiments, the single-stranded adapter comprises: a first complementary sequence; and a second complementary sequence that is partially or fully complementary to the first complementary sequence.

[0048] In certain embodiments, the linker further comprises: a connecting sequence connecting the first complementary sequence and the second complementary sequence.

[0049] In certain embodiments, the linker sequence is located downstream of the first complementary sequence.

[0050] In certain embodiments, the second complementary sequence is located downstream of the linking sequence.

[0051] In certain embodiments, the linker is in the form of a stem-loop under conditions that allow nucleic acid hybridization or annealing.

[0052] In certain embodiments, the linker contains 1 or more (eg, 2, 3, 4, 5) spacers.

[0053] In certain embodiments, the one or more spacers are located in the linker sequence of the linker.

[0054] In certain embodiments, the Spacer is selected from Spacer C3, Spacer C6, Spacer C12, Spacer 9, Spacer 18, abasic linker (dSpacer), PC linker, or any combination thereof.

[0055] In certain embodiments, the single-stranded linker further comprises one or more (e.g., two, three, four, five) cleavable moieties.

[0056] As used herein, the term "cleavable moiety" refers to any cleavable or removable portion contained in a nucleic acid. For example, the cleavable moiety can comprise uracil, ribonucleotides, or other modified nucleotides that can be removed or cleaved using a nucleic acid cleavage agent (e.g., UDG, RNase, endonuclease, etc.).

[0057] In certain embodiments, the nucleic acid cleavage agent is capable of cleaving a cleavable moiety in a linker.

[0058] In certain embodiments, the nucleic acid cleavage agent is selected from uracil DNA glycosylase (UDG), apyrimidinic / apurinic endonuclease (APE), an endonuclease (e.g., endonuclease VIII (EndoVIII) or V (EndoV)), a uracil-specific excision reagent (USER) enzyme, formamidopyrimidine DNA glycosylase (Fpg), 8-oxoguanine glycosylase (OGG1), an RNase, or any combination thereof.

[0059] It is understood that when the linker contains different cleavable parts, multiple corresponding nucleic acid cleavage agents can be selected to perform corresponding cleavage.

[0060] In certain embodiments, when the cleavable moiety comprises a ribonucleotide (eg, a thymidine ribonucleotide), the nucleic acid cleavage agent comprises an RNase (eg, ribonuclease H (RNase H)).

[0061] In certain embodiments, when the cleavable moiety comprises uracil deoxyribonucleic acid, the nucleic acid cleavage agent comprises UDG or USER enzyme.

[0062] In certain embodiments, the second complementary sequence comprises at least one cleavable moiety, or the 3' end of the second complementary sequence is linked to at least one cleavable moiety.

[0063] In certain embodiments, the 3' end of the second complementary sequence is linked to a uracil deoxyribonucleotide or ribonucleotide.

[0064] In certain embodiments, the linker sequence further comprises one or more (eg, two, three, four, five) cleavable moieties.

[0065] In certain embodiments, the plurality of cleavable moieties in the linker sequence are located on both sides of the Spacer.

[0066] In certain embodiments, the linker sequence comprises two cleavable portions, and the two cleavable portions are adjacent to both sides of the Spacer, respectively.

[0067] In certain embodiments, the single-stranded adapter further comprises: a unique molecular identifier (UMI).

[0068] As used herein, the term "UMI" or "unique molecular identifier" refers to one or more nucleotide sequences that can be used to identify one or more specific nucleic acids and to distinguish different nucleotide sequences from one another. Methods for using UMIs to identify or distinguish nucleotide sequences can be found, for example, in Kivioja, Nature Methods 9, 72-74 (2012).

[0069] In some implementations, the UMI is selected from a random UMI, a nonrandom UMI, or any combination thereof.

[0070] In some implementations, the random UMI comprises N bases of 4-28 bp (e.g., 4-10 bp, 10-16 bp, 16-22 bp, 22-28 bp), where N is any one of A, T, C, and G.

[0071] In certain implementations, the nonrandom UMIs include a plurality (e.g., 4, 8, 16, 32, 64, 96, or more) of specific sequences of 4-28 bp (e.g., 4-10 bp, 10-16 bp, 16-22 bp, 22-28 bp).

[0072] In certain implementations, each nonrandom UMI differs from other nonrandom UMIs by at least 1 (e.g., 1, 2, 3, 4) nucleotides at their corresponding sequence positions.

[0073] In certain implementations, the single-stranded adapter comprises random UMIs of 4-10 bp (e.g., 4 bp, 5 bp, 6 bp, 7 bp, 8 bp, 9 bp, 10 bp).

[0074] In certain implementations, the random UMI is located downstream of the second complementary sequence.

[0075] In certain implementations, a cleavable portion is included between the random UMI and the second complementary sequence.

[0076] In some embodiments, the single-stranded adapter comprises, in order from 5' to 3' direction: a first complementary sequence, a linker sequence, a second complementary sequence, a cleavable portion, and a UMI.

[0077] Wherein, the linker sequence contains or does not contain a cleavable portion.

[0078] In certain embodiments, the linker sequence does not comprise a cleavable portion. In certain embodiments, the single-stranded linker has a sequence as shown in SEQ ID NO: 1 or SEQ ID NO: 8.

[0079] In certain embodiments, the linker sequence comprises a cleavable portion. In certain embodiments, the single-stranded linker has a sequence as shown in SEQ ID NO:7.

[0080] Dual-link connector

[0081] In certain embodiments, the double-stranded adapter comprises: a first oligonucleotide chain, and a second oligonucleotide chain; wherein the first oligonucleotide chain comprises a first hybridization sequence and a first template sequence; the second oligonucleotide chain comprises a second hybridization sequence and a second template sequence; wherein the second hybridization sequence is partially or fully complementary to the first hybridization sequence; and the second template sequence is partially or fully non-complementary to the first template sequence.

[0082] In certain embodiments, in the first oligonucleotide chain, the first template sequence is located downstream of the first hybridization sequence; preferably, the first template sequence is a free 3' single-stranded arm.

[0083] In certain embodiments, the end of the 3' single-stranded arm is blocked; for example, by adding a chemical moiety (e.g., biotin or an alkyl group) to the 3'-OH of the last nucleotide of the single-stranded linker, by removing the 3'-OH of the last nucleotide of the probe, or replacing the last nucleotide with a dideoxynucleotide, thereby blocking the 3'-end of the linker.

[0084] In certain embodiments, in the second oligonucleotide chain, the second hybridization sequence is located downstream of the second template sequence. In certain embodiments, the second template sequence is a free 5' single-stranded arm.

[0085] In certain embodiments, the end of the 5' single-stranded arm is blocked; for example, by adding a chemical moiety (e.g., biotin or an alkyl group) to the 3'-OH of the last nucleotide of the single-stranded linker, by removing the 3'-OH of the last nucleotide of the probe, or replacing the last nucleotide with a dideoxynucleotide, thereby blocking the 3'-end of the linker.

[0086] In certain embodiments, the double-stranded linker further comprises a cleavable moiety as previously described.

[0087] In certain implementations, the dual-strand connector further comprises a UMI as previously described.

[0088] In certain embodiments, the first oligonucleotide strand of the double-stranded adapter comprises, in order from 5' to 3' direction: a first hybridization sequence, a first template sequence, wherein the first template sequence is a free 3' single-stranded arm.

[0089] In certain embodiments, the second oligonucleotide strand of the double-stranded adapter comprises, in order from 5' to 3' direction: a second template sequence, a second hybridization sequence, a cleavable portion, and a UMI, wherein the second template sequence is a free 5' single-stranded arm.

[0090] In certain embodiments, the 3'-end of the second oligonucleotide strand is blocked; for example, by adding a chemical moiety (e.g., biotin or an alkyl group) to the 3'-OH of the last nucleotide of the single-stranded linker, by removing the 3'-OH of the last nucleotide of the probe, or replacing the last nucleotide with a dideoxynucleotide, thereby blocking the 3'-end of the linker.

[0091] In certain embodiments, the double-stranded linker has the sequence shown in SEQ ID NOs: 9 and 10.

[0092] Therefore, in certain embodiments, the method further comprises: (a-3) providing a nucleic acid cleavage agent, and contacting the nucleic acid cleavage agent with the product of (a-2).

[0093] In certain embodiments, upon contacting the nucleic acid cleavage agent with the product of (a-2), the nucleic acid cleavage agent is capable of cleaving the cleavable portion of the linker.

[0094] In the methods of the present invention, after the nucleic acid cleavage agent cleaves the cleavable portion of the adapter, the adapter generates a broken nucleotide chain at the cleavable portion. As a result, the UMI downstream of the cleavable portion and the blocked 3' end are subsequently detached from the nucleotide chain of the adapter.

[0095] In such embodiments, after cleavage, since there is no closed 3' end at the break of the adapter, after contact with the nucleotide mixture, the complementary chain of the single-stranded target nucleic acid (i.e., the second chain) will be synthesized downward from the break using the single-stranded target nucleic acid (i.e., the first chain) as a template.

[0096] In certain embodiments, the method is achieved by the following steps (1) to (4):

[0097] (1) providing a single-stranded target nucleic acid; optionally, the 5' end of the single-stranded target nucleic acid does not have a free phosphate group;

[0098] (2) providing one or more linkers and a ligase, and contacting the single-stranded target nucleic acid with the linker and the ligase under conditions that allow nucleic acid ligation;

[0099] (3) providing the nucleic acid cleavage agent, and contacting the product of step (2) with the nucleic acid cleavage agent under conditions that allow the nucleic acid cleavage agent to cleave nucleic acid;

[0100] (4) providing the nucleotide mixture, and contacting the single-stranded target nucleic acid with the nucleotide mixture under conditions that allow synthesis of a complementary strand to the single-stranded target nucleic acid.

[0101] In certain embodiments, in step (1), the single-stranded target nucleic acid is contacted with alkaline phosphatase so that the 5' end of the single-stranded target nucleic acid does not have a free phosphate group.

[0102] Nucleic acid molecules or their amplified products

[0103] In a second aspect, the present application provides a nucleic acid molecule or an amplified product thereof, which is prepared by the method described in the first aspect.

[0104] In certain embodiments, the amplification product is an amplification product of a first strand of a nucleic acid molecule, the first strand having a sequence identical to a single-stranded target nucleic acid sequence.

[0105] In certain embodiments, the amplification product is an amplification product of a second strand of a nucleic acid molecule, the second strand having a sequence that is complementary to the single-stranded target nucleic acid sequence.

[0106] In certain embodiments, the amplification product is the amplification product of a first strand and a second strand of a nucleic acid molecule, wherein the first strand has a sequence identical to the single-stranded target nucleic acid sequence and the second strand has a sequence complementary to the single-stranded target nucleic acid sequence.

[0107] Use of nucleic acid molecules or their amplified products

[0108] In a third aspect, the present application provides the nucleic acid molecule or its amplified product described in the second aspect for use in cytosine conversion.

[0109] In certain embodiments, the nucleic acid molecule or its amplified product described in the second aspect is used for epigenetic information (eg, DNA methylation, DNA mutation) analysis.

[0110] In certain embodiments, the nucleic acid molecule or its amplified product described in the second aspect is used for genetic information (eg, genome) analysis.

[0111] Method for preparing DNA library

[0112] In a fourth aspect, the present application provides a method for preparing a DNA library, the method comprising:

[0113] (i) providing the nucleic acid molecule or its amplified product according to the third aspect;

[0114] (ii) subjecting the nucleic acid molecule or its amplified product to a cytosine conversion treatment under conditions that allow unmodified cytosine to be converted to uracil.

[0115] In certain embodiments, the DNA library is selected from a genomic library, a methylome library, or any combination thereof.

[0116] In certain embodiments, the method is used to prepare genomic and methylomic libraries separately and simultaneously.

[0117] In some embodiments, the above method further comprises: fragmenting the nucleic acid molecule or its amplified product. A person skilled in the art will be able to select an appropriate DNA fragmentation method to make the DNA library suitable for sequencing (including but not limited to next-generation sequencing). In certain embodiments, the DNA is fragmented by physical fragmentation, enzymatic fragmentation, and chemical shearing to generate fragments of suitable and / or sufficient length for the DNA library.

[0118] In this article, cytosine conversion can be carried out using standard experimental procedures or commercially available test kits. In certain embodiments, NaOH is used to denature nucleic acid molecules or their amplification products. After nucleic acid molecules or their amplification products are denatured, sodium bisulfite or sodium metabisulfite can be used to process 4-16 hours at 55°C with a final concentration of, for example, 2M (pH between about 5 and 6). The step is covalently modified with sulfite to form a cytosine that does not have modification. After conversion, nucleic acid molecules or their amplification products are desalted, and then desulfonated by cultivating nucleic acid molecules or their amplification products at alkaline pH and room temperature, which causes deamination and conversion into uracil. In one embodiment, commercially available test kits such as EZ DNA Methylation-Gold, EZ DNA Methylation-Direct or EZ DNA Methylation-Lightning test kits (purchased from Zymo Research Corp (Irvine, California)) are used for bisulfite conversion.

[0119] In certain embodiments, cytosine can also be converted to uracil by methods that do not use bisulfite ions. For example, cytidine deaminase is used to catalyze the irreversible hydrolytic deamination of cytidine and deoxycytidine to uridine and deoxyuridine, respectively. In some embodiments, the conversion of unmodified cytosine to uracil is accomplished by an enzymatic reaction. In some embodiments, the nucleic acid molecule is incubated with a cytidine deaminase. In some embodiments, the cytidine deaminase includes activation-induced cytidine deaminase (AID) and apolipoprotein B mRNA editing enzyme, catalytic polypeptide-like protein (APOBEC). In some embodiments, the APOBEC enzyme is selected from the human APOBEC family, which is composed of the following groups: APOBEC-1 (Apo1), APOBEC-2 (Apo2), AID, APOBEC-3A, -3B, -3C, -3DE, -3F, -3G, -3H, and APOBEC-4 (Apo4). In some embodiments, the enzyme is a variant of APOBEC (US 20130244237). In some embodiments, the conversion uses a commercially available kit. In one example, a kit such as APOBEC-Seq (NEBiolabs; US20130244237) is used.

[0120] Therefore, in certain embodiments, in step (ii), bisulfite is provided and contacted with the nucleic acid molecule or an amplified product thereof; optionally, cytidine deaminase and / or TET are also provided and contacted with the nucleic acid molecule or an amplified product thereof.

[0121] In certain embodiments, after step (ii), the first strand of the nucleic acid molecule or amplified product thereof is enriched.

[0122] In certain embodiments, after step (ii), the second strand of the nucleic acid molecule or amplified product thereof is enriched.

[0123] In certain embodiments, after step (ii), the first and second strands of the nucleic acid molecule or its amplified product are separated.

[0124] In certain embodiments, the first strand is used to construct a methylome library.

[0125] In certain embodiments, the second strand is used to construct a genomic library.

[0126] Sequencing primers

[0127] In certain embodiments, the method further comprises providing a sequencing primer and contacting it with the product of step (ii) under conditions that allow nucleic acid ligation.

[0128] In certain embodiments, the sequencing primer comprises a first oligonucleotide chain and a second oligonucleotide chain; wherein the first oligonucleotide chain comprises a first hybridization sequence and a first template sequence; the second oligonucleotide chain comprises a second hybridization sequence and a second template sequence; wherein the second hybridization sequence is partially or fully complementary to the first hybridization sequence; and the second template sequence is partially or fully non-complementary to the first template sequence.

[0129] In certain embodiments, the first template sequence is located upstream of the first hybridization sequence in the first oligonucleotide strand. In certain embodiments, the first template sequence is a free 5' single-stranded arm; and in certain embodiments, the hybridization sequence of the first oligonucleotide strand is connected to the first strand of a nucleic acid molecule or an amplified product thereof.

[0130] In certain embodiments, the second hybridization sequence is located upstream of the second template sequence in the second oligonucleotide strand. In certain embodiments, the second template sequence is a free 3' single-stranded arm; and in certain embodiments, the hybridization sequence of the second oligonucleotide strand is connected to the second strand of the nucleic acid molecule or its amplified product.

[0131] In certain embodiments, the cytosine in the sequencing primer comprises or consists of a modified cytosine. In certain embodiments, the modified cytosine is selected from 5-propynylcytosine, 5-methylcytosine, 5-hydroxymethylcytosine, 5-carboxypyrimidine, 5-hydroxycytosine, 5-fluorocytosine, 5-chlorocytosine, 5-bromocytosine, 5-iodocytosine, or any combination thereof.

[0132] In certain embodiments, the cytosine in the sequencing primer comprises or consists of 5-propynylcytosine.

[0133] In certain embodiments, the specific sequence of the sequencing primer is adjusted according to the sequencing platform (eg, BGI, Illumina).

[0134] In certain embodiments, the sequencing primers have sequences as shown in SEQ ID NOs: 11 and 12.

[0135] In certain embodiments, the first oligonucleotide strand and / or the second oligonucleotide strand of the sequencing primer further comprises a unique molecular identifier (UMI).

[0136] In some embodiments, the UMI is within the hybridization sequence of the first oligonucleotide strand and / or the second oligonucleotide strand, or the UMI is located at the end of the hybridization sequence of the first oligonucleotide strand and / or the second oligonucleotide strand.

[0137] In some embodiments, the UMI is located 3' to the first hybridization sequence. In some embodiments, the UMI of the first oligonucleotide strand is attached to the first strand of a nucleic acid molecule or an amplified product thereof.

[0138] In some embodiments, the UMI is located 5' to the second hybridization sequence. In some embodiments, the UMI of the second oligonucleotide strand is attached to the second strand of the nucleic acid molecule or an amplified product thereof.

[0139] In some implementations, the UMI is selected from a random UMI, a nonrandom UMI, or any combination thereof.

[0140] In some implementations, the random UMI comprises N bases of 4-28 bp (e.g., 4-10 bp, 10-16 bp, 16-22 bp, 22-28 bp), where N is any one of A, T, C, and G.

[0141] In some implementations, the random UMI comprises N bases of 10-16 bp (e.g., 10 bp, 11 bp, 12 bp, 13 bp, 14 bp, 15 bp, 16 bp), where N is any one of A, T, C, and G.

[0142] In certain implementations, the nonrandom UMIs include a plurality (e.g., 2-4, 4-8, 8-16, 16-32, 32-64, 64-96, or more) of specific sequences of 4-28 bp (e.g., 4-10 bp, 10-16 bp, 16-22 bp, 22-28 bp).

[0143] In certain implementations, the nonrandom UMIs include 2-8 (e.g., 2, 3, 4, 5, 6, 7, 8) specific sequences of 4-10 bp (e.g., 4 bp, 5 bp, 6 bp, 7 bp, 8 bp, 9 bp, 10 bp).

[0144] In certain implementations, each of the multiple specific sequences included in the nonrandom UMI differs from the other specific sequences by at least 1 (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10) nucleotide at corresponding nucleotide positions.

[0145] In certain implementations, the specific sequence of the nonrandom UMI is set forth in any one of SEQ ID NOs: 13-16.

[0146] In certain embodiments, the first oligonucleotide chain and / or the second oligonucleotide chain of the sequencing primer further comprises a positioning tag, wherein the positioning tag comprises at least one (e.g., one, two, three, four, five, six, seven, eight) specific sequence, and the specific sequence consists of 3-8 (e.g., three, four, five, six, seven, eight) bases.

[0147] In certain embodiments, the base is selected from A, T, C, G.

[0148] In certain embodiments, the specific sequence consists of 3-8 identical or different bases.

[0149] In certain embodiments, the positioning tag comprises four specific sequences, each of which is composed of the same or different bases. In certain embodiments, the four specific sequences are of different lengths (e.g., differ by 1 base, differ by 2 bases, or differ by 3 bases).

[0150] In certain embodiments, the positioning tag comprises a first specific sequence consisting of 3 identical or different bases, a second specific sequence consisting of 4 identical or different bases, a third specific sequence consisting of 5 identical or different bases, and a fourth specific sequence consisting of 6 identical or different bases.

[0151] In certain embodiments, the first specific sequence has the sequence shown in SEQ ID NO: 17. In certain embodiments, the second specific sequence has the sequence shown in SEQ ID NO: 18. In certain embodiments, the third specific sequence has the sequence shown in SEQ ID NO: 19. In certain embodiments, the fourth specific sequence has the sequence shown in SEQ ID NO: 20.

[0152] In certain embodiments, the positioning tag further comprises at least one base G at the 5' end of the specific sequence.

[0153] In certain embodiments, the positioning tag comprises the sequence set forth in SEQ ID NO: 21. In certain embodiments, the positioning tag comprises the sequence set forth in SEQ ID NO: 4. In certain embodiments, the positioning tag comprises the sequence set forth in SEQ ID NO: 5. In certain embodiments, the positioning tag comprises the sequence set forth in SEQ ID NO: 6.

[0154] In certain embodiments, the positioning tag is located in the hybridization sequence of the first oligonucleotide chain and / or the second oligonucleotide chain, or the positioning tag is located at the end of the hybridization sequence of the first oligonucleotide chain and / or the second oligonucleotide chain.

[0155] In certain embodiments, the positioning tag is located at the 3' end of the first hybridization sequence. In certain embodiments, the positioning tag of the first oligonucleotide strand is attached to the first strand of the nucleic acid molecule or its amplified product.

[0156] In certain embodiments, the positioning tag is located at the 5' end of the second hybridization sequence. In certain embodiments, the positioning tag of the second oligonucleotide strand is attached to the second strand of the nucleic acid molecule or its amplified product.

[0157] DNA library

[0158] In another aspect, the present application provides a DNA library constructed by the method as claimed in the preceding claims.

[0159] In certain embodiments, the DNA library is selected from a genomic library, a methylome library, or any combination thereof.

[0160] In certain embodiments, the genomic library and the methylome library are prepared separately and simultaneously by the methods described above.

[0161] In certain embodiments, the genomic library comprises or consists of the second strand of a nucleic acid molecule or an amplification product thereof.

[0162] In certain embodiments, the methylome library comprises or consists of a first strand of a nucleic acid molecule or an amplified product thereof.

[0163] Methods for analyzing epigenetic and / or genetic information

[0164] On the other hand, the present application provides a method for analyzing epigenetic information and / or genetic information, the method comprising: detecting and / or analyzing the nucleic acid molecule or its amplified product as described above, or detecting and / or analyzing the DNA library as described above.

[0165] In certain embodiments, the first strand and the second strand of the nucleic acid molecule or its amplified product are detected and / or analyzed separately.

[0166] In certain embodiments, the genomic library and the methylome library are detected and / or analyzed separately.

[0167] In certain embodiments, the detection is sequencing. In certain embodiments, the sequencing is selected from the group consisting of next generation sequencing, massively parallel sequencing, pyrosequencing, sequencing by synthesis, single molecule real-time sequencing, polony sequencing, DNA nanoball sequencing, helioscopic single molecule sequencing, nanopore sequencing, Sanger sequencing, shotgun sequencing, Gilbert sequencing analysis, or any combination thereof.

[0168] In certain embodiments, the detecting is selected from microarray, quantitative PCR (qPCR), digital droplet PCR (ddPCR), molecular inversion probes, or any combination thereof.

[0169] In certain embodiments, the analysis is a bioinformatics analysis. In certain embodiments, the bioinformatics analysis includes, but is not limited to, bioinformatics analysis selected from sequence alignment, genomic analysis, transcriptome analysis, single nucleotide variation (SNV) analysis, gene copy number variation (CNV) analysis, measurement of chromosome copy number, and detection of genetic lesions.

[0170] In certain embodiments, bioinformatics analysis can be used to quantify the number of genome equivalents analyzed in a DNA library (e.g., cfDNA), or to detect the genetic state of a target locus, or to detect genetic lesions in a target locus, or to measure copy number fluctuations within a target locus.

[0171] In certain embodiments, sequence reads obtained by sequencing can be compared with one or more reference DNA sequences (e.g., human genome sequences). In specific embodiments, sequencing comparisons can be used to detect genetic lesions in target loci, including but not limited to detecting nucleotide transitions or transversions, nucleotide insertions or deletions, genomic rearrangements, copy number changes, or gene fusions. In certain embodiments, the epigenetic information is DNA methylation information.

[0172] In certain embodiments, the genetic information is whole genome information and / or mutation information.

[0173] Method for distinguishing first strand (methylation information) / second strand (genomic information)

[0174] In certain embodiments, the nucleic acid molecule or an amplified product thereof is sequenced, and sequencing data comprising the first strand and the second strand are obtained, or sequencing data of the first strand and the second strand are obtained separately.

[0175] In certain embodiments, when the first oligonucleotide strand of a sequencing primer includes a UMI and the second oligonucleotide strand does not include a UMI, sequencing data of the first strand and the second strand are distinguished by identifying the presence of the UMI sequence in sequencing data of the nucleic acid molecule or its amplified product.

[0176] In some embodiments, the sequencing data for which a UMI is identified is the sequencing data for the first strand. In some embodiments, the sequencing data for which a UMI is not identified is the sequencing data for the second strand.

[0177] In certain embodiments, when the second oligonucleotide strand of the sequencing primer comprises a UMI and the first oligonucleotide strand does not comprise a UMI, the sequencing data of the first strand and the second strand are distinguished by identifying the presence of the UMI sequence in the sequencing data.

[0178] In some embodiments, the sequencing data for which a UMI is identified is the sequencing data for the second strand. In some embodiments, the sequencing data for which a UMI is not identified is the sequencing data for the first strand.

[0179] In certain embodiments, when a first oligonucleotide strand of a sequencing primer includes a UMI and a second oligonucleotide strand also includes a UMI, sequencing data for the first strand and the second strand are distinguished by identifying whether the UMI in the sequencing data is closer to the 5' single-stranded arm or the 3' single-stranded arm of the sequencing primer.

[0180] In some embodiments, the sequencing data for which the UMI is identified closer to the 5' single-stranded arm is the sequencing data for the first strand. In some embodiments, the sequencing data for which the UMI is identified closer to the 3' single-stranded arm is the sequencing data for the second strand.

[0181] In certain embodiments, sequencing data of the first and second strands are distinguished by separately identifying whether cytosine conversion has occurred in the cytosine of the UMI in the first and second oligonucleotide strands.

[0182] In some embodiments, sequencing data in which cytosine in a UMI has been converted (e.g., to uracil) is the sequencing data of the first strand.

[0183] In certain implementations, sequencing data in which the cytosine in the UMI has not undergone cytosine conversion (e.g., remains as cytosine) is sequencing data for the second strand.

[0184] In certain embodiments, when the first oligonucleotide chain of the sequencing primer contains a positioning tag and the second oligonucleotide chain does not contain a positioning tag, the sequencing data of the first chain and the second chain are distinguished by identifying whether the sequence of the positioning tag exists in the sequencing data of the nucleic acid molecule or its amplified product.

[0185] In certain embodiments, the sequencing data in which the positioning tag is identified is the sequencing data of the first strand. In certain embodiments, the sequencing data in which the positioning tag is not identified is the sequencing data of the second strand.

[0186] In certain embodiments, when the second oligonucleotide strand of the sequencing primer contains a positioning tag and the first oligonucleotide strand does not contain a positioning tag, the sequencing data of the first strand and the second strand are distinguished by identifying whether the positioning tag sequence exists in the sequencing data.

[0187] In certain embodiments, the sequencing data in which the positioning tag is identified is the sequencing data of the second strand. In certain embodiments, the sequencing data in which the positioning tag is not identified is the sequencing data of the first strand.

[0188] In certain embodiments, when the first oligonucleotide chain of the sequencing primer includes a positioning tag and the second oligonucleotide chain also includes a positioning tag, the sequencing data of the first chain and the second chain are distinguished by identifying whether the positioning tag in the sequencing data is closer to the 5' single-stranded arm or the 3' single-stranded arm of the sequencing primer.

[0189] In certain embodiments, the sequencing data in which the positioning tag is identified to be closer to the 5' single-stranded arm is the sequencing data of the first strand. In certain embodiments, the sequencing data in which the positioning tag is identified to be closer to the 3' single-stranded arm is the sequencing data of the second strand.

[0190] In certain embodiments, sequencing data of the first and second oligonucleotide chains are distinguished by separately identifying whether the sequences of the positioning tags in the first and second oligonucleotide chains are changed.

[0191] In certain embodiments, the sequencing data for converting guanine (G) in the positioning tag to adenine (A) is the sequencing data of the first strand.

[0192] In certain embodiments, the sequencing data in which the guanine (G) in the positioning tag is not converted is the sequencing data of the second strand.

[0193] In certain embodiments, the sequencing data of the first strand is used to analyze the methylation information of the target nucleic acid.

[0194] In certain embodiments, the sequencing data of the second strand is used to analyze the genomic information of the target nucleic acid.

[0195] Uses of DNA Libraries

[0196] In another aspect, the present application provides a DNA library as described above for analyzing epigenetic information and / or genetic information.

[0197] In certain embodiments, the DNA library is selected from a genomic library, a methylome library, or any combination thereof.

[0198] In certain embodiments, the epigenetic information is DNA methylation information.

[0199] In certain embodiments, the genetic information is whole genome information and / or mutation information.

[0200] Purpose of the kit

[0201] In another aspect, the present application provides use of the aforementioned nucleic acid molecule or its amplified product, or the aforementioned DNA library, in preparing a kit for detecting whether a subject has a disease.

[0202] In certain embodiments, the disease results in a change in the subject's epigenetic and / or genetic information (eg, a nucleotide transition or transversion, a nucleotide insertion or deletion, a genomic rearrangement, a copy number change, or a gene fusion).

[0203] In certain embodiments, the disease is cancer.In certain embodiments, the disease is a genetic disease.

[0204] In certain embodiments, the kit is used to:

[0205] (1) Detect whether the subject has cancer;

[0206] (2) Detect whether the subject has cancer recurrence;

[0207] (3) detecting the responsiveness of a subject with a disease to a therapy and / or drug received for the disease;

[0208] (4) Detect whether the subject has a genetic disease.

[0209] Therefore, in certain embodiments, the detection subject of the kit can be a subject who has received chemotherapy, radiotherapy or immunotherapy. In certain embodiments, the detection subject of the kit can be a subject who has undergone cancer recurrence after failure or success of cancer treatment.

[0210] In certain embodiments, the cancer is selected from the group consisting of liver cancer, hepatocellular carcinoma, melanoma, pancreatic cancer, lung cancer, kidney cancer, stomach cancer, esophageal cancer, colon cancer, breast cancer, ovarian cancer, cervical cancer, testicular cancer, prostate cancer, lymphoma, B-cell lymphoma, diffuse large B-cell lymphoma, follicular lymphoma, mantle zone cell lymphoma, small lymphocytic lymphoma, splenic marginal zone B-cell lymphoma, extranodal marginal zone mucosa-associated lymphoid tissue B-cell lymphoma, nodal marginal zone B-cell lymphoma, Tumor, lymphoplasmacytic lymphoma, primary effusion lymphoma, Burkitt lymphoma / Burkitt cell leukemia, T-cell lymphoma, anaplastic large cell lymphoma (primary cutaneous type), anaplastic large cell lymphoma (systemic type), peripheral T-cell lymphoma, angioimmunoblastic T-cell lymphoma, adult T-cell lymphoma / leukemia (human T-cell lymphotropic virus type I positive), extranodal NK / T-cell lymphoma (nasal type), enteropathy-associated T-cell lymphoma, GAMMA / delta hepatosplenic T-cell lymphoma, subcutaneous panniculitis-like T-cell lymphoma, multiple myeloma, and mycosis fungoides.

[0211] In certain embodiments, the genetic disease is selected from Alzheimer's disease (APOE1), Charcot-Marie-Tooth disease, Leber hereditary optic neuropathy (LHON), Angelman syndrome (UBE3A, ubiquitin-protein ligase E3A), Prader-Willi syndrome (region in chromosome 15), beta-thalassemia (HBB, beta-globin), Gaucher disease (type I) (GBA, glucocerebrosidase), cystic fibrosis (CFTR epithelial chloride channel), sickle cell disease (HBB, beta-globin), phenylketonuria (PAH, phenylalanine hydrolase), familial hypercholesterolemia (LDLR, low-density lipoprotein cholesterol), phenylketonuria (PAH, phenylalanine hydrolase), familial hypercholesterolemia (LDLR, low-density lipoprotein cholesterol), phenylketonuria (PAH, phenylalanine hydrolase), phenylketonuria (PAH, phenylalanine hydrolase), phenylketonuria (PAH, phenylalanine hydrolase), phenylketonuria (PAH, phenylalanine hydrolase), phenylketonuria (PAH, phenylalanine hydrolase), phenylketonuria (PAH, phenylketonuria ... These include Huntington's disease (HDD, huntingtin), neurofibromatosis type 1 (NF1, NF1 tumor suppressor gene), myotonic dystrophy (DM, Job's tears), tuberous sclerosis complex (TSC1, potato globulin), achondroplasia (FGFR3, fibroblast growth factor receptor), fragile X syndrome (FMR1, RNA-binding protein), Duchenne muscular dystrophy (DMD, dystrophin), hemophilia A (F8C, coagulation factor VIII), Lesch-Nyhan syndrome (HPRT1, hypoxanthine guanine ribosyltransferase 1), and adrenoleukodystrophy (ABCD1).

[0212] Reagent test kit

[0213] In another aspect, the present application provides a kit comprising:

[0214] (1) a propynyl-modified cytosine; (2) a cytosine conversion reagent (e.g., bisulfite, cytidine deaminase and / or TET); and (3) a single-stranded linker and / or a double-stranded linker as described above.

[0215] In certain embodiments, the kit further comprises: (4) reagents for library construction.

[0216] In certain embodiments, the reagents for library construction are selected from enzymes, reagents for DNA end repair, RNase-free water, or any combination thereof.

[0217] In certain embodiments, the enzyme is selected from T4 DNA ligase, T3 DNA ligase, T7 DNA ligase, E. coli DNA ligase, HIFI Taq DNA ligase, T4 RNA ligase, RTCB ligase, ring ligase, thermostable DNA ligase, or any combination thereof.

[0218] Definition of terms

[0219] Unless otherwise indicated, scientific and technical terms used herein have the meanings commonly understood by those skilled in the art. Furthermore, the virology, biochemistry, and immunology laboratory procedures used herein are conventional procedures widely used in the respective fields. To facilitate a better understanding of the present invention, definitions and explanations of relevant terms are provided below.

[0220] When the terms "for example," "such as," "including," "including," "comprising," or variations thereof are used herein, these terms will not be considered as limiting terms, but will be interpreted to mean "but not limited to" or "not limited to."

[0221] The terms "a" and "an" and "the" and similar referents in the context of describing the invention (especially in the context of the following claims) are to be construed to cover both the singular and the plural, unless otherwise indicated herein or clearly contradicted by context.

[0222] As used herein, the terms "polynucleotide," "nucleic acid sequence," "nucleic acid fragment," "oligonucleotide," "nucleic acid," and "nucleic acid molecule" refer to a polymeric form of nucleotides of any length, such as deoxyribonucleotides (dNTPs) or ribonucleotides (rNTPs) or their analogs, and are used interchangeably. Oligonucleotides are typically composed of a specific sequence of four nucleotide bases: adenine (A); cytosine (C); guanine (G); and thymine (T) (when the polynucleotide is RNA, thymine (T) is uracil (U)). Non-limiting examples of nucleic acids include DNA, RNA, genomic DNA (e.g., gDNA, such as sheared gDNA), cell-free DNA (e.g., cfDNA), synthetic DNA / RNA, coding or non-coding regions of a gene or gene fragment, loci defined according to linkage analysis, exons, introns, messenger RNA (mRNA), transfer RNA, ribosomal RNA, short interfering RNA (siRNA), short hairpin RNA (shRNA), microRNA (miRNA), ribozymes, complementary DNA (cDNA), recombinant nucleic acids, branched nucleic acids, plasmids, vectors, isolated DNA of any sequence, isolated RNA of any sequence, nucleic acid probes, and primers. Nucleic acids may comprise one or more nucleotides with modifications, such as methylated nucleotides and nucleotide analogs.

[0223] As used herein, the term "amplification" generally refers to the production of one or more copies or extension products of a nucleic acid molecule (e.g., the product of a primer extension reaction on a nucleic acid molecule). The amplification of a nucleic acid molecule can produce a single strand hybridized with a nucleic acid molecule or multiple copies of a nucleic acid molecule or its complementary sequence. An amplification product can be a single-stranded or double-stranded nucleic acid molecule generated from a starting template nucleic acid molecule by an amplification program. An "amplification product" can include a nucleic acid chain that is substantially identical or substantially complementary to at least a portion of the starting template. In the case where the starting template is a double-stranded nucleic acid molecule, the amplification product can include a nucleic acid chain that is substantially identical to at least a portion of a chain and is substantially complementary to at least a portion of any chain. In certain embodiments, the amplification product can be single-stranded or double-stranded. The amplification reaction can be, for example, polymerase chain reaction (PCR), such as emulsion polymerase chain reaction (ePCR; for example, the PCR performed in a microreactor such as a hole or droplet).

[0224] As used herein, the term "complementary" means that two nucleic acid sequences can form hydrogen bonds between each other according to the base pairing principle (Waston-Crick principle), and thus form a duplex. In the present application, the term "complementary" includes "substantially complementary" and "completely complementary". As used herein, the term "completely complementary" means that each base in a nucleic acid sequence can be paired with the base in another nucleic acid chain without mismatching or gaps. As used herein, the term "substantially complementary" means that most of the bases in a nucleic acid sequence can be paired with the base in another nucleic acid chain, which allows the presence of mismatches or gaps (e.g., mismatches or gaps of one or several nucleotides). Typically, under conditions that allow nucleic acid hybridization, annealing or amplification, two nucleic acid sequences that are "complementary" (e.g., substantially complementary or completely complementary) will selectively / specifically hybridize or anneal and form a duplex. Accordingly, the term "non-complementary" means that two nucleic acid sequences cannot hybridize or anneal under conditions that allow nucleic acid hybridization, annealing or amplification, and cannot form a duplex. As used herein, the term "incomplete complementarity" means that the bases in one nucleic acid sequence cannot be completely paired with the bases in another nucleic acid chain, and there is at least one mismatch or gap.

[0225] As used herein, the term "unmodified nucleotides" includes naturally occurring nucleotides, such as nucleotides containing adenine (A), thymine (T), cytosine (C), uracil (U), or guanine (G).

[0226] As used herein, the term "modified nucleotides" includes, but is not limited to, analogs derived from naturally occurring nucleotides, or nucleotides that have similar structures to naturally occurring nucleotides but contain one or more differences. In certain embodiments, modified nucleotides include inosine, diaminopurine, 5-fluorouracil, 5-bromouracil, 5-chlorouracil, 5-iodouracil, hypoxanthine, xanthine, deazaxanthine, deazaguanine, isocytosine, isoguanine, 4-acetylcytosine, 5-(carboxyhydroxymethyl)uracil, 5-carboxymethylaminomethyl-2-thiouridine, 5-carboxymethylaminomethyluracil, dihydrouracil, β-D-galactosyl Q Nucleosides (galactosylqueosine), N6-isopentenyl adenine, 1-methylguanine, 1-methylinosine, 2,2-dimethylguanine, 2-methyladenine, 2-methylguanine, 3-methylcytosine, 5-methylcytosine, N6-adenine, 7-methylguanine, 5-methylaminomethyluracil, 5-methoxyaminomethyl-2-thiouracil, β-D-mannosyl Q nucleoside (mannosylqueosine), ylqueosine), 5"-methoxycarboxymethyluracil, 5-methoxyuracil, 2-methylthio-D46-isopentenyl adenine, uracil-5-oxyacetic acid (v), wybutoxosine, pseudouracil, Q nucleoside (queosine), 2-thiocytosine, 5-methyl-2-thiouracil, 2-thiouracil, 4-thiouracil, 5-methyluracil, uracil-5 -oxyacetate, methyl uracil-5-oxyacetate (v), 5-methyl-2-thiouracil, 3-(3-amino-3-N-2-carboxypropyl)uracil, (acp3)w, 2,6-diaminopurine, ethynyl nucleotide bases, 1-propynyl nucleotide bases, azido nucleotide bases, selenophosphate nucleic acids, and modified versions thereof (e.g., by oxidation, reduction, and / or addition of substituents such as alkyl, hydroxyalkyl, hydroxyl, or halogen moieties).

[0227] As used herein, the term "sequencing primer" generally refers to a molecule (e.g., polynucleotide) that interacts with a target nucleic acid to facilitate sequencing (e.g., next generation sequencing (NGS)). In the presence of a sequencing primer, the target nucleic acid can be sequenced by a sequencer. In certain embodiments, a sequencing primer can include a nucleotide sequence that hybridizes or binds to the capture target nucleic acid, and the capture target nucleic acid is connected to a solid support, such as a bead or a flow cell, of a sequencing system. In certain embodiments, a sequencing primer can include a nucleotide sequence that hybridizes or binds to the target nucleic acid to generate a hairpin loop, thereby allowing the target nucleic acid to be sequenced by a sequencing system. In certain embodiments, a sequencing primer can include a sequencer motif, which can be a nucleotide sequence complementary to the flow cell sequence of another molecule (e.g., polynucleotide) and can be used by a sequencing system to sequence the target nucleic acid.

[0228] As used herein, the term "next generation sequencing (NGS)" refers to sequencing methods that allow for massive parallel sequencing of clonally amplified molecules and single nucleic acid molecules. Non-limiting examples of NGS include synthesis sequencing using reversible dye terminators and ligation sequencing.

[0229] As used herein, the term "spacer" refers to a spacer-type modification. Spacer is introduced into oligonucleotides, usually to establish a distance between oligonucleotides or between oligonucleotides and other functional groups, to avoid steric hindrance, reduce adverse interactions between groups, increase flexibility and other effects. Alternatively, in the case where the oligonucleotide is not required to be extended, spacer can be used as a blocking group. The number of atoms in different spacers is different, and the required spatial distance can be reached by adjusting the number and type of inserted spacers. The types of spacers include, but are not limited to, hydrophobic Spacer C3, C6, C12, hydrophilic Spacer 9, Spacer 18, dSpacer, PC linker.

[0230] As used herein, the term "UMI" or "unique molecular identifier" refers to one or more nucleotide sequences that can be used to identify one or more specific nucleic acids, which can be used to distinguish different nucleotide sequences from each other. Methods for using UMIs to identify nucleotide sequences can be found, for example, in Kivioja, Nature Methods 9, 72-74 (2012). A UMI can include at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 or more nucleotides (e.g., consecutive nucleotides). In high-throughput analysis of multiple samples using next-generation sequencing technology, UMIs can be used to distinguish samples. In certain embodiments, UMIs can be randomly generated (i.e., random UMIs) or non-randomly generated (i.e., non-random UMIs).

[0231] As used herein, the term "cleavable portion" refers to any cleavable or removable portion contained in a nucleic acid. For example, the cleavable portion may comprise uracil, ribonucleotides, or other modified nucleotides that can be excised or cut using an enzyme (e.g., UDG, RNase, endonuclease, etc.). The cleavable portion may comprise a spacer, such as a C3 spacer, hexanediol, a triethylene glycol spacer (e.g., spacer 9), a hexaethylene glycol spacer (e.g., spacer 18), or a combination or analog thereof. The cleavable portion may comprise modified nucleotides, such as methylated nucleotides. The modified nucleotides can be specifically recognized by an enzyme (e.g., methylated nucleotides can be recognized by MspJI). The cleavable portion can be enzymatically cut (e.g., using enzymes such as UDG, RNase, APE1, MspJI, etc.). The cleavable portion can be cut using stimulation, such as light stimulation, chemical stimulation, thermal stimulation, etc.

[0232] It should be understood that the cleavable moiety can be attached to any position in the nucleic acid molecule. For example, in Figure 1, a cleavable moiety (e.g., uracil deoxyribonucleic acid) is attached between the second complementary sequence and the UMI. Therefore, the cleavable moiety (e.g., uracil, ribonucleotide, spacer, abasic site, enzyme-specific modified nucleotides such as methylated nucleotides, etc.) and any combination thereof can be attached to the 5' end or 3' end of any fragment of the nucleic acid molecule. Similarly, in the case where the nucleic acid molecule is double-stranded, the cleavable moiety can be provided on different strands. For example, one strand can have a cleavable moiety attached to its 5' end, and the other strand can have a cleavable moiety attached to its 3' end.

[0233] As used herein, the term "sample" refers to a sample derived from a cell, tissue, organ, or organism, comprising a nucleic acid or a mixture of nucleic acids of at least one nucleic acid sequence suspected of having epigenetic information variation (e.g., methylation modification) and / or genetic information variation (e.g., but not limited to single nucleotide polymorphisms, insertions, deletions, and structural variations). In certain embodiments, the sample comprises at least one target nucleic acid whose epigenetic information and / or genetic information is suspected of having undergone variation.

[0234] Advantageous Effects of the Invention

[0235] After extensive research, the inventors of this application have obtained a method for simultaneously collecting genetic and epigenetic information using the same target nucleic acid, and further obtained a method for simultaneously performing genome sequencing and methylation group sequencing using the same target nucleic acid. This method can significantly reduce the probability of false positive events that occur in genome sequencing results, greatly improving the accuracy of sequencing. Specifically, the error rate of genome C>T mutation events is reduced by at least two orders of magnitude, and the accuracy of UMI splitting is also significantly improved. Furthermore, this application has also designed and improved the single-chain adapter / double-chain adapter and sequencing adapter used in the method, thereby further improving the accuracy of the method.

[0236] The embodiments of the present invention will be described in detail below with reference to the accompanying drawings and examples, but it will be understood by those skilled in the art that the following drawings and examples are intended only to illustrate the present invention and are not intended to limit the scope of the invention. Various objects and advantages of the present invention will become apparent to those skilled in the art based on the following detailed description of the accompanying drawings and preferred embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0237] Figure 1 is a schematic diagram of the structure of a single-stranded adapter MMssAdapter1TU for a single-stranded target nucleic acid. The single-stranded adapter comprises, from the 5' to the 3' direction, a first complementary sequence, a linker sequence connecting the first complementary sequence and the second complementary sequence, a second complementary sequence partially or fully complementary to the first complementary sequence, a cleavable portion (e.g., uracil deoxyribonucleic acid), and a UMI (e.g., a random UMI). The linker sequence contains a spacer but does not contain a cleavable portion. Furthermore, the 3' end of the adapter is closed.

[0238] Figure 2 shows the fragmentation analysis results of the MMLib0710_18 library.

[0239] FIG3 shows the fragmentation analysis results of the MMLib0803_11 library.

[0240] Figure 4 is a schematic diagram of the structure of the single-stranded adapter MMssAdapter1TU-4U for a single-stranded target nucleic acid, wherein the single-stranded adapter comprises, from the 5' to the 3' direction, a first complementary sequence, a connecting sequence connecting the first complementary sequence and the second complementary sequence, a second complementary sequence partially or completely complementary to the first complementary sequence, a cleavable portion (e.g., uracil deoxyribonucleic acid), and a UMI (e.g., a random UMI); wherein the connecting sequence comprises a spacer and two cleavable portions, and the two cleavable portions are respectively adjacent to both sides of the spacer; wherein the 3' end of the adapter is closed.

[0241] FIG5 shows the fragmentation analysis results of the MMLib0726_2 library.

[0242] Figure 6 is a schematic diagram of the structure of a single-stranded adapter MMssAdapter1rTrT for a single-stranded target nucleic acid, wherein the single-stranded adapter comprises, from the 5' to the 3' direction, a first complementary sequence, a connecting sequence connecting the first complementary sequence and the second complementary sequence, a second complementary sequence partially or completely complementary to the first complementary sequence, a cleavable portion (e.g., thymidine ribonucleic acid), and a UMI (e.g., a random UMI); wherein the connecting sequence contains a spacer and does not contain a cleavable portion; wherein the 3'-end of the adapter is closed.

[0243] FIG7 shows the fragmentation analysis results of the MMLib0824_5 library.

[0244] Figure 8 is a schematic structural diagram of a double-stranded adapter MMssAdapter1TU10 for a single-stranded target nucleic acid; wherein the double-stranded adapter comprises: a first oligonucleotide chain and a second oligonucleotide chain; wherein the first oligonucleotide chain comprises a first hybridization sequence and a first template sequence; wherein the second oligonucleotide chain comprises a second hybridization sequence and a second template sequence; wherein the second hybridization sequence is partially or fully complementary to the first hybridization sequence; and wherein the second template sequence is partially or fully non-complementary to the first template sequence;

[0245] The first oligonucleotide chain of the double-stranded adapter comprises, from 5' to 3' direction, a first hybridization sequence and a first template sequence, wherein the first template sequence is a free 3' single-stranded arm; the end of the 3' single-stranded arm is closed;

[0246] The second oligonucleotide strand of the double-stranded adapter comprises, in order from the 5' to the 3' direction: a second template sequence, a second hybridization sequence, a cleavable portion (e.g., uracil deoxyribonucleic acid), and a UMI (e.g., a random UMI). The second template sequence is a free 5' single-stranded arm; the end of the 5' single-stranded arm is closed.

[0247] FIG9 shows the fragmentation analysis results of the MMLib1009_3 library.

[0248] FIG10 shows the fragmentation analysis results of the MMLib0906_3 library.

[0249] Figure 11 is a schematic diagram of the structure of a first sequencing primer after ligation with a target nucleic acid, wherein the first sequencing primer comprises a first oligonucleotide chain and a second oligonucleotide chain; wherein the first oligonucleotide chain comprises a first hybridization sequence and a first template sequence; and the second oligonucleotide chain comprises a second hybridization sequence and a second template sequence; wherein the second hybridization sequence is partially or fully complementary to the first hybridization sequence; and the second template sequence is partially or fully non-complementary to the first template sequence;

[0250] The first oligonucleotide chain comprises, from the 5' to the 3' direction, a template sequence, a hybridization sequence, a random UMI, and a positioning tag; the second oligonucleotide chain comprises, from the 5' to the 3' direction, a positioning tag, a random UMI, a template sequence, and a hybridization sequence.

[0251] Figure 12 is a schematic diagram of the structure of the second sequencing adapter after connection to the target nucleic acid, wherein the second sequencing primer comprises a first oligonucleotide chain and a second oligonucleotide chain; wherein the first oligonucleotide chain comprises a first hybridization sequence and a first template sequence; the second oligonucleotide chain comprises a second hybridization sequence and a second template sequence; wherein the second hybridization sequence is partially or fully complementary to the first hybridization sequence; and the second template sequence is partially or fully non-complementary to the first template sequence;

[0252] The first oligonucleotide chain comprises, from the 5' to the 3' direction, a template sequence, a hybridization sequence, and a non-random UMI; the second oligonucleotide chain comprises, from the 5' to the 3' direction, a non-random UMI, a template sequence, and a hybridization sequence.

[0253] Figure 13 is a schematic diagram of the structure of the third sequencing adapter after connection to the target nucleic acid, wherein the third sequencing primer comprises a first oligonucleotide chain and a second oligonucleotide chain; wherein the first oligonucleotide chain comprises a first hybridization sequence and a first template sequence; the second oligonucleotide chain comprises a second hybridization sequence and a second template sequence; wherein the second hybridization sequence is partially or fully complementary to the first hybridization sequence; and the second template sequence is partially or fully non-complementary to the first template sequence;

[0254] The first oligonucleotide chain comprises, from the 5' to the 3' direction, a template sequence, a hybridization sequence, a random UMI, and a positioning tag; the second oligonucleotide chain comprises, from the 5' to the 3' direction, a positioning tag, a random UMI, a template sequence, and a hybridization sequence.

[0255] Sequence information

[0256] A description of the sequences involved in this application is provided in the table below.

[0257] Table 1: Sequence information DETAILED DESCRIPTION

[0258] The invention will now be described with reference to the following examples which are intended to illustrate the invention but not to limit it.

[0259] Unless otherwise specified, the molecular biology experimental methods used in the present invention are basically based on the methods described in J. Sambrook et al., Molecular Cloning: A Laboratory Manual, 2nd edition, Cold Spring Harbor Laboratory Press, 1989, and F.M. Ausubel et al., Molecular Biology: A Laboratory Manual, 3rd edition, John Wiley & Sons, Inc., 1995. Restriction endonucleases were used according to the conditions recommended by the product manufacturers. It will be appreciated by those skilled in the art that the examples illustrate the present invention by way of example and are not intended to limit the scope of the present invention.

[0260] Example 1. Comparison of 5py-dCTP and 5m-dCTP

[0261] 1.1 5m-dCTP

[0262] The single-stranded DNA ligation adapter MMssAdapter1TU shown in SEQ ID NO: 1 was designed and synthesized (the structure is shown in FIG1 ).

[0263] The double-stranded adapters with location tags and UMIs were designed and prepared according to the patent "An Improved Method for Preparing Random-Tag Adapters for Next-Generation Sequencing" CN107190067B (the entire contents of which are incorporated herein by reference) and synthesized by Sangon Biotech (Shanghai) Co., Ltd. The specific synthesis is as follows:

[0264] The sequence of the adapter primer P5 is shown in SEQ ID NO:2.

[0265] Adapter primer P7-A is obtained by ligating FFFFF (Spacer) EEEEEJJJJJNNNNNNNNNN to the 5' end of SEQ ID NO: 2; adapter primer P7-B is obtained by ligating FFFFF (Spacer) EEEEEKKKKKNNNNNNNNNN to the 5' end of SEQ ID NO: 2; adapter primer P7-C is obtained by ligating FFFFF (Spacer) EEEEELLLLLNNNNNNNNNNNN to the 5' end of SEQ ID NO: 2; SEQ ID NO: 3 is AGATCGGAAGAGCACACGTCTGAACTCCAGTCAC. All cytosines in SEQ ID NO: 2 and SEQ ID NO: 3 are methylated at the carbon atom at position 5.

[0266] The prepared connector was named MM-UMI.

[0267] The library was constructed by mixing 30 ng of sheared HCT116 cell line DNA with 0.005 ng of fully methylated pUC19 and 0.1 ng of completely unmethylated Lambda DNA. The template dephosphorylation treatment system was as follows:

[0268] The dephosphorylation procedure is as follows:

[0269] Then directly prepare the reaction system of connection 1 (connection single-chain connector MMssAdapter1TU):

[0270] The reaction procedure for ligation 1 is as follows:

[0271] Then, 160 μl of our own fragment sorting magnetic beads were added to the reaction wells, and after purification, 25 μl of magnetic beads were used for elution.

[0272] The eluted product was digested using the USER enzyme digestion system, as follows:

[0273] The reaction process of the digestion system is as follows:

[0274] Then, 60 μl of our own fragment sorting magnetic beads were added to the reaction wells, and after purification, 6.5 μl of magnetic beads were used for elution.

[0275] Prepare the reaction system for second-strand synthesis:

[0276] The reaction procedure for the second-strand synthesis is as follows:

[0277] The ligation 2 reaction system was prepared as follows:

[0278] The ligation 2 reaction procedure is as follows:

[0279] Then, 60 μl of our own fragment sorting magnetic beads were added to the reaction wells, and after purification, 22 μl of magnetic beads were used for elution.

[0280] The methylation conversion was performed using the NEBNext Enzymatic Methyl-seq Conversion Module kit. The operation procedure was performed according to the kit procedure. Since the operation is consistent with the kit instructions, it will not be repeated here.

[0281] After the conversion, the product is pre-amplified. The pre-amplification system is as follows:

[0282] The pre-amplification procedure is as follows:

[0283] Add 40 μL of our self-prepared fragment sorting magnetic beads for fragment sorting and purification. After purification, use 40 μL of ultrapure water for elution.

[0284] The pooled library obtained after elution was quantified using the Qubit HS dsDNA Kit, and the results are as follows:

[0285] The library was fragmented and analyzed using Qsep. The results are shown in Figure 2.

[0286] During data analysis, the full methylation group data and the whole genome data are distinguished by whether the UMI is at the R1 end or the R2 end of the library data. The distinction results and quality control results are as follows:

[0287] The conversion efficiency of methylome data was evaluated as follows:

[0288] 1.2 5py-dCTP

[0289] The co-detection library was constructed using the single-stranded DNA adapters and library construction methods described in 1.1. The only modification was to use the following formula in the second-strand synthesis reaction system, replacing 5m-dCTP with 5py-dCTP:

[0290] The MMLib0803_11 mixed library was obtained and quantified using the Qubit HS dsDNA Kit. The results are as follows:

[0291] The library was fragmented and analyzed using Qsep. The results are shown in FIG3 .

[0292] During data analysis, the full methylation group data and the whole genome data are distinguished by whether the UMI is at the R1 end or the R2 end of the library data. The distinction results and quality control results are as follows:

[0293] The conversion efficiency of methylome data was evaluated as follows:

[0294] In summary, the use of 5py-dCTP during second-strand synthesis (i.e., ligation 2) in this application's method significantly reduces the error rate of genomic C>T mutation events by two orders of magnitude! Therefore, for simultaneous genome and methylome sequencing, this method can significantly reduce the probability of false positive events in genome sequencing results, greatly improving sequencing accuracy.

[0295] Example 2. Comparison of different cytopyrimidine derivatives

[0296] This example uses the commercial xGen Prism Library Prep Kit (IDT, Cat. 10009821). The specific experimental procedures were performed as described in the instructions, except for modifications and changes in the linker. In brief, the experiment is as follows:

[0297] The sample DNA was end-repaired using the end-repair module in the kit, and the first ligation step was performed using the ligation module in the kit. The adapter used in the first ligation step was modified, specifically replacing cytosines outside the UMI portion of the adapter with methylated cytosines, and the modified adapter was synthesized by Bio-Tech. The product of the first ligation step was then ligated using the ligation module in the second ligation step. The adapter used in the second ligation step was modified, specifically replacing cytosines outside the UMI portion of the adapter with methylated cytosines, and the modified adapter was synthesized by Bio-Tech. After the second ligation step was completed, the product was purified, and the protected chain was synthesized.

[0298] In the synthesis of the second strand (i.e., ligation 2), different cytosine derivatives (i.e., 5m-dCTP, 5hm-dCTP, 5ca-dCTP, or 5py-dCTP) were used to explore their effects on sequencing results. The specific formulations used in this example are shown in the following table:

[0299] The reaction procedure for protection chain synthesis is as follows:

[0300] Afterwards, 2.5x magnetic beads were used to purify the protected PCR product. Enzymatic Methyl-seq Conversion Module (NEB) was used to perform seven rounds of library construction and pre-amplification of the converted products, followed by sequencing. The sequencing results clearly show the detection accuracy of different cytosine derivatives, as shown below:

[0301] The WGS test, used as a standard control, also showed a false-positive rate of 0.04%, which is on the same order of magnitude as the concordance rate using 5-py-dCTP. The error rates for other modifications were 1-2 orders of magnitude higher. Considering that there are approximately 700 million cytosines in the genome, such a high error rate is unacceptable for mutation detection and has no practical clinical application value.

[0302] In summary, compared with other cytosine derivatives, the use of 5py-dCTP in the synthesis of the second chain (i.e., ligation 2) can significantly reduce the probability of false positive events in genome sequencing results.

[0303] Example 3. Exploration of single-stranded joints with different structures

[0304] The single-stranded DNA ligation adapter MMssAdapter1TU-4U shown in SEQ ID NO: 7 was designed and synthesized (the structure is shown in FIG4 ).

[0305] The co-detection library was constructed according to the library construction method described in 1.1 of Example 1, and the formula was modified as follows during the first ligation step:

[0306] The MMLib0726_2 mixed library was obtained and quantified using the Qubit HS dsDNA Kit. The results are as follows:

[0307] The library was fragmented and analyzed using Qsep. The results are shown in Figure 5.

[0308] During data analysis, the full methylation group data and the whole genome data are distinguished by whether the UMI is at the R1 end or the R2 end of the library data. The distinction results and quality control results are as follows:

[0309] The conversion efficiency of methylome data was evaluated as follows:

[0310] In summary, the use of the single-chain linker provided in this application can achieve simultaneous genome sequencing and methylome sequencing with good accuracy.

[0311] Example 4. Exploration of single-stranded joints with different structures

[0312] The difference of the single-stranded DNA ligation adapter MMssAdapter1rTrT of this embodiment is that thymidine ribonucleic acid is used instead of uracil deoxyribonucleic acid, as shown in SEQ ID NO: 8 (the structure is shown in FIG6 ).

[0313] The specific process is basically the same as 1.1 of Example 1, except that:

[0314] After ligation 1, the eluted product was digested using the RNaseH enzyme digestion system as follows:

[0315] The reaction procedure of the RNase H enzyme digestion system is as follows:

[0316] The MMLib0824_5 mixed library was obtained and quantified using the Qubit HS dsDNA Kit. The results are as follows:

[0317] The library was fragmented and analyzed using Qsep. The results are shown in FIG7 .

[0318] During data analysis, the full methylation group data and the whole genome data are distinguished by whether the UMI is at the R1 end or the R2 end of the library data. The distinction results and quality control results are as follows:

[0319] The conversion efficiency of methylome data was evaluated as follows:

[0320] In summary, the use of the single-chain linker provided in this application can achieve simultaneous genome sequencing and methylome sequencing with good accuracy.

[0321] Example 5. Exploration of Dual-Strand Joints

[0322] Specifically, the double-stranded adapter MMssAdapter1TU10 of this embodiment uses a complementary Y-shaped adapter instead of a circular adapter as the adapter used for the first step of connection. The two chains of the adapter are shown in SEQ ID NOs: 9 and 10, and the structure is shown in FIG8 .

[0323] Before use, the two oligonucleotide chains of the fourth adapter were mixed in equal amounts.

[0324] The specific process is basically the same as 1.1 of Example 1, except that:

[0325] Remove the first step of rSAP dephosphorylation and only keep the denaturation process:

[0326] The library was constructed by mixing 30 ng of sheared HCT116 cell line DNA with 0.005 ng of fully methylated pUC19 and 0.1 ng of completely unmethylated Lambda DNA. The template dephosphorylation treatment system was as follows:

[0327] The denaturation procedure is as follows:

[0328] The co-detection library was constructed according to the library construction method described in 1.1 of Example 1, and the formula was modified as follows during the first ligation step:

[0329] The MMLib1009_3 mixed library was obtained and quantified using the Qubit HS dsDNA Kit. The results are as follows:

[0330] The library was fragmented and analyzed using Qsep. The results are shown in Figure 9.

[0331] During data analysis, the full methylation group data and the whole genome data are distinguished by whether the UMI is at the R1 end or the R2 end of the library data. The distinction results and quality control results are as follows:

[0332] The conversion efficiency of methylome data was evaluated as follows:

[0333] In summary, the use of the double-stranded linker provided in this application can achieve simultaneous genome sequencing and methylome sequencing with good accuracy.

[0334] Example 6. Exploration of Sequencing Adapters

[0335] The specific process is basically the same as 1.1 of Example 1. The difference is that a Y-shaped adapter commonly used in common second-generation sequencing is used in connection 2. The Y-shaped adapter consists of two partially complementary paired sequences. Unlike the adapters commonly known in the industry, the cytosines in this adapter are all methylated cytosines. The sequences of the two chains of the Y-shaped adapter, M-AdaptorU and M-AdaptorB, are shown in SEQ ID NOs: 11 and 12, respectively.

[0336] During the experiment, the system of connection 2 was modified as follows:

[0337] The Ligation 2 reaction procedure was unchanged.

[0338] The MMLib0906_3 mixed library was obtained and quantified using the Qubit HS dsDNA Kit. The results are as follows:

[0339] The library was fragmented and analyzed using Qsep. The results are shown in FIG10 .

[0340] During data analysis, the full methylation group data and the whole genome data are distinguished by whether the UMI is at the R1 end or the R2 end of the library data. The distinction results and quality control results are as follows:

[0341] The conversion efficiency of methylome data was evaluated as follows:

[0342] In summary, using the Y-type adapter commonly used in general second-generation sequencing, the method of the present application can achieve simultaneous genome sequencing and methylation group sequencing with good accuracy.

[0343] Example 7. Exploration of Sequencing Adapters

[0344] The specific process is basically the same as 1.1 of Example 1. The difference is that a uniquely designed sequencing adapter is used in connection 2. The sequencing adapter contains random UMIs and distinguishes the genome and methylation groups by the presence of positioning tags.

[0345] The schematic diagram of the structure after connecting the sequencing adapter to the target nucleic acid is shown in Figure 11, which constructs a mixed library of the genome and methylation group, wherein:

[0346] The UMI is 12 random bases NNNNNNNNNNNN, and the positioning tags are JJJ, KKKK, LLLLL, EEEEEE, and all Cs are methylated at carbon atom 5. The sequencing adapter contains any one of the four positioning tags mentioned above, wherein the positioning tag contains non-random bases, for example, JJJ refers to any combination of three specific bases (ACT in this embodiment, SEQ ID NO: 17), KKKK refers to any combination of four specific bases (GACT in this embodiment, SEQ ID NO: 18), LLLLL refers to any combination of five specific bases (TGACT in this embodiment, SEQ ID NO: 19), and EEEEEE refers to any combination of six specific bases (CTGACT in this embodiment, SEQ ID NO: 20).

[0347] The resulting library was sequenced, and the data was separated into genome and methylome using a method based on UMI position, as follows:

[0348] The data was quality controlled, the sequencing adapters were cut off, and low-quality and short sequences were filtered out.

[0349] The data after quality control were matched with the positioning tags at 12 bases away from the starting position of R1 and R2 respectively. A valid match was considered when it matched any of the sequences JJJ, KKKK, LLLLL, and EEEEEE.

[0350] The matching results of all data are classified into the following categories: both R1 and R2 match the positioning tag, only R1 matches the positioning tag, only R2 matches the positioning tag, and neither R1 nor R2 matches the positioning tag. The data where only R1 matches the positioning tag is the methylation group data, and the data where only R2 matches the positioning tag is the genomic data. The data split results are as follows:

[0351] The split methylome data and genomic data were aligned to the reference genome and subsequently analyzed. The methylation conversion rate of the methylation data was evaluated based on lambda DNA. The analysis results are as follows:

[0352] In summary, this example demonstrates that the uniquely designed sequencing adapter and method of the present application can accurately distinguish methylome data from genomic data.

[0353] Example 8. Exploration of Sequencing Adapters

[0354] The specific process is basically the same as 1.1 of Example 1. The difference is that a uniquely designed sequencing adapter is used in connection 2. The sequencing adapter contains non-random UMIs and uses the sequence of the UMI to distinguish the genome and methylation group.

[0355] The schematic diagram of the structure after connecting the sequencing adapter to the target nucleic acid is shown in Figure 12, which constructs a mixed library of the genome and methylation group, wherein:

[0356] The UMIs were four fixed sequences of 8 bases in length: ATCGAGTC, CCGTGGAA, ATCATGCG, and TGTAGCGT (SEQ ID NOs: 13-16).

[0357] The resulting library was sequenced, and the data was separated into genome and methylome using a UMI sequence-based method, as follows:

[0358] The data was quality controlled, the sequencing adapters were cut off, and low-quality and short sequences were filtered out.

[0359] After quality control, the first 8 bases of the sequence are truncated as the UMI sequence.

[0360] The extracted UMIs were aligned with the four designed UMIs. If the extracted UMIs were consistent with the designed UMIs, that is, if the extracted UMIs were one of ATCGAGTC, CCGTGGAA, ATCATGCG, or TGTAGCGT, the data was genomic data. If the extracted UMIs differed from the designed UMIs only at the C position in the designed UMI, and the base type at that position in the extracted UMI was T, that is, if the extracted UMIs were one of ATTGAGTT, TTGTGGAA, ATTATGTG, or TGTAGTGT, the data was methylome data. The data splitting results are as follows:

[0361] The split methylome data and genomic data were aligned to the reference genome and subsequently analyzed. The methylation conversion rate of the methylation data was evaluated based on lambda DNA. The analysis results are as follows:

[0362] In summary, this example demonstrates that the uniquely designed sequencing adapter and method of the present application can accurately distinguish methylome data from genomic data.

[0363] Example 9. Exploration of Sequencing Adapters

[0364] The specific process is basically the same as 1.1 of Example 1. The difference is that a uniquely designed sequencing adapter is used in connection 2. The sequencing adapter contains non-random UMIs and positioning tags, and the sequence of the positioning tag is used to distinguish the genome and methylation group.

[0365] The schematic diagram of the structure after connecting the sequencing adapter to the target nucleic acid is shown in Figure 13, to construct a mixed library of genome and methylation group, wherein:

[0366] The UMI is 12 random bases NNNNNNNNNNNN, and the positioning tags are GJJJ, GKKKK, GLLLLL, and GEEEEEE. Except for the C of the first complementary chain of the positioning tag, all other Cs are methylated at position 5. The sequencing adapter contains any one of the four positioning tags mentioned above, wherein the positioning tag contains non-random bases, for example, JJJ refers to any combination of three specific bases (ACT in this embodiment), so the sequence of GJJJ is GACT (SEQ ID NO: 21); KKKK refers to any combination of four specific bases (GACT in this embodiment), so the sequence of GKKKK is GGACT (SEQ ID NO: 4); LLLLL refers to any combination of five specific bases (TGACT in this embodiment), so the sequence of GLLLLL is GTGACT (SEQ ID NO: 5); EEEEEE refers to any combination of six specific bases (CTGACT in this embodiment), so the sequence of GEEEEEE is GCTGACT (SEQ ID NO: 6).

[0367] The resulting library was sequenced, and the data was separated into genome and methylome using a method based on the location tag sequence, as follows:

[0368] The data was quality controlled, the sequencing adapters were cut off, and low-quality and short sequences were filtered out.

[0369] The data after quality control were matched with the positioning tags at 12 bases away from the starting position in R1 and R2 respectively. A valid match was considered when it matched any of the sequences GJJJ, GKKKK, GLLLLL, GEEEEEE or AJJJ, AKKKK, ALLLLL, AEEEEEE.

[0370] The matching results of all data are classified as follows: both R1 and R2 match the positioning tag, only R1 matches the positioning tag, only R2 matches the positioning tag, and neither R1 nor R2 matches the positioning tag. Among them, the data that both R1 and R2 match the positioning tags are valid data. Furthermore, the part where R2 matches the positioning tags GJJJ, GKKKK, GLLLLL, GEEEEEE is the genomic data, and the part where R2 matches the positioning tags AJJJ, AKKKK, ALLLLL, AEEEEEE is the methylation group data (the first base of the positioning tag is glycine G, and its corresponding base in the complementary chain is unmodified cytosine C, which will be converted to U after conversion, so its corresponding complementary base is A). The data splitting results are as follows:

[0371] The split methylome data and genomic data were aligned to the reference genome and subsequently analyzed. The methylation conversion rate of the methylation data was evaluated based on lambda DNA. The analysis results are as follows:

[0372] In summary, this example demonstrates that the uniquely designed sequencing adapter and method of the present application can accurately distinguish methylome data from genomic data.

[0373] Although the specific embodiments of the present invention have been described in detail, those skilled in the art will understand that various modifications and changes can be made to the details based on all the teachings published, and these changes are all within the scope of protection of the present invention. The entire invention is given by the appended claims and any equivalents thereof.

Claims

1. A method for preparing a double-stranded nucleic acid molecule, the method comprising: (a-1) providing at least one single-stranded target nucleic acid, and a nucleotide mixture comprising adenine (A), cytosine (C), guanine (G), and thymine (T), wherein the cytosine comprises or consists of propargyl-modified cytosine; (b-1) contacting the single-stranded target nucleic acid with the nucleotide mixture under conditions that permit synthesis of the complementary strand of the single-stranded target nucleic acid; Optionally, the cytosine further comprises other (e.g., methyl, hydroxymethyl, carboxyl, halogenated) modified cytosines, and / or unmodified cytosine; Optionally, the propargyl-modified cytosine further has one or more other (e.g., methyl, hydroxymethyl, carboxyl, halogenated) modifications.

2. The method according to claim 1, wherein, The method has one or more of the following characteristics: (1) The other modified cytosines are selected from 5-methylcytosine, 5-hydroxymethylcytosine, 5-carboxypyrimidine, 5-hydroxycytosine, 5-fluorocytosine, 5-chlorocytosine, 5-bromocytosine, 5-iodocytosine, or any combination thereof; (2) The cytosine comprises or consists of 5-propargylcytosine; (3) The cytosine comprises or consists of modified cytosine, wherein the modified cytosine comprises at least 10% (e.g., at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 99%) of 5-propargylcytosine; Preferably, the modified cytosine comprises at least 50%, 60%, 70%, 80%, 90%, 95%, or 99% of 5-propargylcytosine; Preferably, the modified cytosine comprises or consists of 5-propargylcytosine and 5-hydroxymethylcytosine; (4) The adenine comprises modified adenine and / or unmodified adenine; (5) The thymine comprises modified thymine and / or unmodified thymine; (6) The guanine comprises modified guanine and / or unmodified guanine; (7) The complementary strand in the nucleic acid molecule is resistant to cytosine conversion; Preferably, the method has one or more of the following characteristics: (1) The target nucleic acid is DNA (e.g., genomic DNA, cfDNA) and / or RNA; (2) The single-stranded target nucleic acid is a naturally occurring single-stranded nucleic acid, or is derived from one nucleic acid strand in a double-stranded nucleic acid; (3) The method further comprises: obtaining the single-stranded target nucleic acid from a sample, or obtaining a double-stranded target nucleic acid from a sample and preparing it into a single-stranded target nucleic acid; Preferably, the sample or target nucleic acid is obtained from a prokaryote, a eukaryote (e.g., protozoa, parasite, fungus, yeast, plant, animal including mammals and humans) or a virus (e.g., Herpes virus, HIV, influenza virus, Epstein-Barr virus, hepatitis virus, poliovirus, etc.) or a viroid; Preferably, the sample is a sample comprising cells and / or tissues; Preferably, the sample is selected from whole blood, serum, plasma, cerebrospinal fluid, sputum, feces, urine, saliva, or any combination thereof.

3. The method according to claim 1 or 2, wherein Before step (b-1), the method further comprises: (a-2) providing one or more adapters and a ligase, and contacting the single-stranded target nucleic acid with the adapter and the ligase under conditions permitting nucleic acid ligation; Preferably, the adapter is selected from single-stranded adapters, double-stranded adapters, or any combination thereof; Preferably, the ligase is selected from T4 DNA ligase, T3 DNA ligase, T7 DNA ligase, Escherichia coli DNA ligase, HIFI Taq DNA ligase, T4 RNA ligase, RTCB ligase, circular ligase, thermostable DNA ligase, or any combination thereof; Optionally, the 3'-end of the adapter is blocked; for example, by adding a chemical moiety (such as biotin or alkyl) to the 3'-OH of the last nucleotide of the single-stranded adapter, by removing the 3'-OH of the last nucleotide of the probe, or by replacing the last nucleotide with a dideoxynucleotide, thereby blocking the 3'-end of the adapter; Preferably, the method further comprises: (a-3) providing a nucleic acid cleavage agent and contacting the nucleic acid cleavage agent with the product of (a-2); Preferably, after contacting the nucleic acid cleavage agent with the product of (a-2), the nucleic acid cleavage agent is capable of cleaving a cleavable portion in the adapter.

4. The method according to claim 3, wherein, The single-stranded adapter comprises: a first complementary sequence, and a second complementary sequence that is partially or fully complementary to the first complementary sequence; Preferably, the adapter further comprises: a linking sequence connecting the first complementary sequence and the second complementary sequence; Preferably, the linking sequence is located downstream of the first complementary sequence; Preferably, the second complementary sequence is located downstream of the linking sequence; Preferably, under conditions permitting nucleic acid hybridization or annealing, the adapter is in a stem-loop structure; Preferably, the adapter contains one or more (such as 2, 3, 4, 5) Spacers; Preferably, the one or more Spacers are located in the linking sequence of the adapter; Preferably, the Spacer is selected from Spacer C3, Spacer C6, Spacer C12, Spacer 9, Spacer 18, abasic linker (dSpacer), PC linker, or any combination thereof.

5. The method according to claim 3 or 4, wherein, The single-stranded adapter further comprises one or more (such as 2, 3, 4, 5) cleavable portions; Preferably, after contacting the adapter with the nucleic acid cleavage agent, under conditions permitting the nucleic acid cleavage agent to cleave nucleic acid, the nucleic acid cleavage agent is capable of cleaving the cleavable portion in the adapter; Preferably, the nucleic acid cleavage agent is selected from uracil DNA glycosylase (UDG), apurinic / apyrimidinic endonuclease (APE), endonuclease (such as endonuclease VIII (EndoVIII) or V (EndoV)), uracil-specific excision reagent (USER) enzyme, formamidopyrimidine DNA glycosylase (Fpg), 8-oxoguanine glycosylase (OGG1), ribonuclease, or any combination thereof Preferably, when the cleavable portion contains ribonucleotides (e.g., ribothymidylic acid), the nucleic acid cleaving agent comprises an RNA enzyme (e.g., ribonuclease H (RNase H)); Preferably, when the cleavable portion contains deoxyuridylic acid, the nucleic acid cleaving agent comprises UDG or USER enzyme; Preferably, the second complementary sequence contains at least one cleavable portion, or the 3'-end of the second complementary sequence is linked to at least one cleavable portion; Preferably, the 3'-end of the second complementary sequence is linked to deoxyuridylic acid or ribonucleotides; Optionally, the linking sequence further contains one or more (e.g., 2, 3, 4, 5) cleavable portions; Preferably, multiple cleavable portions in the linking sequence are located on both sides of the Spacer; Preferably, the linking sequence contains 2 cleavable portions, and the 2 cleavable portions are adjacent to both sides of the Spacer respectively.

6. The method according to any one of claims 3 to 5, wherein The single-stranded linker further comprises: a unique molecular identifier (UMI); Preferably, the UMI is selected from random UMI, non-random UMI, or any combination thereof; Preferably, the random UMI contains N bases of 4-28 bp (e.g., 4-10 bp, 10-16 bp, 16-22 bp, 22-28 bp), where N is any one of A, T, C, G; Preferably, the non-random UMI contains multiple (e.g., 4, 8, 16, 32, 64, 96, or more) specific sequences of 4-28 bp (e.g., 4-10 bp, 10-16 bp, 16-22 bp, 22-28 bp); Preferably, each non-random UMI differs from other non-random UMIs by at least 1 (e.g., 1, 2, 3, 4) nucleotide at their corresponding sequence positions; Preferably, the single-stranded linker contains a random UMI of 4-10 bp (e.g., 4 bp, 5 bp, 6 bp, 7 bp, 8 bp, 9 bp, 10 bp); Preferably, the random UMI is located downstream of the second complementary sequence; Preferably, a cleavable portion is included between the random UMI and the second complementary sequence.

7. The method according to any one of claims 3 to 6, wherein The single-stranded linker sequentially comprises, from the 5'- to 3'-direction: a first complementary sequence, a linking sequence, a second complementary sequence, a cleavable portion, and a UMI; Wherein, the linking sequence contains or does not contain a cleavable portion; Preferably, the linking sequence does not contain a cleavable portion; preferably, the single-stranded linker has the sequence shown in SEQ ID NO:1 or SEQ ID NO:8; Preferably, the linking sequence contains a cleavable portion; preferably, the single-stranded linker has the sequence shown in SEQ ID NO:

7.

8. The method according to claim 3, wherein, The double linker comprises: a first oligonucleotide chain, and a second oligonucleotide chain; wherein, the first oligonucleotide chain comprises a first hybridization sequence and a first template sequence; the second oligonucleotide chain comprises a second hybridization sequence and a second template sequence; wherein, the second hybridization sequence is partially or completely complementary to the first hybridization sequence; the second template sequence is partially or completely non-complementary to the first template sequence; Preferably, in the first oligonucleotide chain, the first template sequence is located downstream of the first hybridization sequence; preferably, the first template sequence is a free 3'-single-stranded arm; Preferably, the end of the 3'-single-stranded arm is blocked; for example, by adding a chemical moiety (such as biotin or an alkyl group) to the 3'-OH of the last nucleotide of the single linker, by removing the 3'-OH of the last nucleotide of the probe, or by replacing the last nucleotide with a dideoxynucleotide, thereby blocking the 3'-end of the linker; Preferably, in the second oligonucleotide chain, the second hybridization sequence is located downstream of the second template sequence; preferably, the second template sequence is a free 5'-single-stranded arm; Preferably, the end of the 5'-single-stranded arm is blocked; for example, by adding a chemical moiety (such as biotin or an alkyl group) to the 3'-OH of the last nucleotide of the single linker, by removing the 3'-OH of the last nucleotide of the probe, or by replacing the last nucleotide with a dideoxynucleotide, thereby blocking the 3'-end of the linker; Preferably, the double linker further comprises a cleavable moiety as claimed in claim 5; Preferably, the double linker further comprises a UMI as claimed in claim 6; Preferably, the first oligonucleotide chain of the double linker sequentially comprises, from the 5'- to 3'-direction: a first hybridization sequence, a first template sequence, wherein the first template sequence is a free 3'-single-stranded arm; Preferably, the second oligonucleotide chain of the double linker sequentially comprises, from the 5'- to 3'-direction: a second template sequence, a second hybridization sequence, a cleavable moiety, a UMI, wherein the second template sequence is a free 5'-single-stranded arm; Preferably, the 3'-end of the second oligonucleotide chain is blocked; for example, by adding a chemical moiety (such as biotin or an alkyl group) to the 3'-OH of the last nucleotide of the single linker, by removing the 3'-OH of the last nucleotide of the probe, or by replacing the last nucleotide with a dideoxynucleotide, thereby blocking the 3'-end of the linker; Preferably, the double linker has the sequences shown in SEQ ID NO: 9 and 10; 9. The method according to any one of claims 1-8, wherein, The method is achieved by the following steps (1) to (4): (1) Provide a single-stranded target nucleic acid; optionally, the 5'-end of the single-stranded target nucleic acid does not have a free phosphate group; (2) Provide one or more of the linkers and a ligase, and under conditions permitting nucleic acid ligation, contact the single-stranded target nucleic acid with the linker and the ligase; (3) Provide the nucleic acid cleavage agent, and under conditions allowing the nucleic acid cleavage agent to cleave nucleic acids, contact the product of step (2) with the nucleic acid cleavage agent; (4) Provide the nucleotide mixture, and under conditions allowing the complementary strand of the single-stranded target nucleic acid to be synthesized, contact the single-stranded target nucleic acid with the nucleotide mixture; Preferably, in step (1), contact the single-stranded target nucleic acid with alkaline phosphatase so that the 5'-end of the single-stranded target nucleic acid does not have a free phosphate group.

10. A nucleic acid molecule or its amplification product, which is prepared by the method according to any one of claims 1-9; Preferably, the amplification product is an amplification product of the first strand of the nucleic acid molecule, and the first strand has the same sequence as the single-stranded target nucleic acid sequence; Preferably, the amplification product is an amplification product of the second strand of the nucleic acid molecule, and the second strand has a sequence complementary to the single-stranded target nucleic acid sequence; Preferably, the amplification product is an amplification product of the first strand and the second strand of a nucleic acid molecule, wherein, The first strand has the same sequence as the single-stranded target nucleic acid sequence, and the second strand has a sequence complementary to the single-stranded target nucleic acid sequence.

11. Use of the nucleic acid molecule or its amplification product according to claim 10 for cytosine conversion; Preferably, the nucleic acid molecule or its amplification product according to claim 10 is used for analysis of epigenetic information (for example, DNA methylation, DNA mutation); Preferably, the nucleic acid molecule or its amplification product according to claim 10 is used for analysis of genetic information (for example, genome).

12. A method for preparing a DNA library, the method comprising: (i) Provide the nucleic acid molecule or its amplification product according to claim 10; (ii) Under conditions allowing unmodified cytosine to be converted into uracil, perform cytosine conversion treatment on the nucleic acid molecule or its amplification product; Preferably, the DNA library is selected from a genomic library, a methylome library, or any combination thereof; Preferably, the method is used for simultaneously preparing a genomic library and a methylome library respectively.

13. The method according to claim 12, wherein, The method has one or more of the following characteristics: (1) In step (ii), provide bisulfite and contact it with the nucleic acid molecule or its amplification product; optionally, also provide cytidine deaminase and / or TET and contact them with the nucleic acid molecule or its amplification product; (2) After step (ii), enrich the first strand of the nucleic acid molecule or its amplification product; (3) After step (ii), enrich the second strand of the nucleic acid molecule or its amplification product; (4) After step (ii), separate the first strand and the second strand in the nucleic acid molecule or its amplification product; Preferably, the first strand is used to construct a methylome library; Preferably, the second strand is used to construct a genomic library.

14. The method according to claim 12 or 13, the method further comprising: Provide a sequencing primer and contact it with the product of step (ii) under conditions allowing nucleic acid ligation; Preferably, the sequencing primer comprises a first oligonucleotide chain and a second oligonucleotide chain; wherein, the first oligonucleotide chain comprises a first hybridization sequence and a first template sequence; the second oligonucleotide chain comprises a second hybridization sequence and a second template sequence; wherein, the second hybridization sequence is partially or completely complementary to the first hybridization sequence; the second template sequence is partially or completely non-complementary to the first template sequence; Preferably, in the first oligonucleotide chain, the first template sequence is located upstream of the first hybridization sequence; preferably, the first template sequence is a free 5' single-stranded arm; preferably, the hybridization sequence of the first oligonucleotide chain is linked to the first strand in a nucleic acid molecule or its amplification product; Preferably, in the second oligonucleotide chain, the second hybridization sequence is located upstream of the second template sequence; preferably, the second template sequence is a free 3' single-stranded arm; preferably, the hybridization sequence of the second oligonucleotide chain is linked to the second strand in a nucleic acid molecule or its amplification product; Preferably, the cytosine in the sequencing primer comprises or consists of a modified cytosine; preferably, the modified cytosine is selected from 5-propynylcytosine, 5-methylcytosine, 5-hydroxymethylcytosine, 5-carboxypyrimidine, 5-hydroxycytosine, 5-fluorocytosine, 5-chlorocytosine, 5-bromocytosine, 5-iodocytosine, or any combination thereof; Preferably, the specific sequence of the sequencing primer is adjusted according to the sequencing platform (e.g., BGI, Illumina); Preferably, the sequencing primer has the sequences shown in SEQ ID NO: 11 and 12; 15. The method according to any one of claims 12-14, wherein, The first oligonucleotide chain and / or the second oligonucleotide chain of the sequencing primer further comprises a unique molecular identifier (UMI); Preferably, the UMI is located in the hybridization sequence of the first oligonucleotide chain and / or the second oligonucleotide chain, or the UMI is located at the end of the hybridization sequence of the first oligonucleotide chain and / or the second oligonucleotide chain; Preferably, the UMI is located at the 3' end of the first hybridization sequence; preferably, the UMI of the first oligonucleotide chain is linked to the first strand in a nucleic acid molecule or its amplification product; Preferably, the UMI is located at the 5' end of the second hybridization sequence; preferably, the UMI of the second oligonucleotide chain is linked to the second strand in a nucleic acid molecule or its amplification product; Preferably, the UMI is selected from random UMI, non-random UMI, or any combination thereof; 16. The method according to any one of claims 12-15, wherein, The random UMI comprises N bases of 4-28 bp (e.g., 4-10 bp, 10-16 bp, 16-22 bp, 22-28 bp), wherein N is any one of A, T, C, G; Preferably, the random UMI comprises N bases of 10-16 bp (e.g., 10 bp, 11 bp, 12 bp, 13 bp, 14 bp, 15 bp, 16 bp), wherein N is any one of A, T, C, G; Preferably, the non-random UMI contains multiple (e.g., 2 - 4, 4 - 8, 8 - 16, 16 - 32, 32 - 64, 64 - 96, or more) specific sequences of 4 - 28 bp (e.g., 4 - 10 bp, 10 - 16 bp, 16 - 22 bp, 22 - 28 bp); Preferably, the non-random UMI contains 2 - 8 (e.g., 2, 3, 4, 5, 6, 7, 8) specific sequences of 4 - 10 bp (e.g., 4 bp, 5 bp, 6 bp, 7 bp, 8 bp, 9 bp, 10 bp); Preferably, among the multiple specific sequences contained in the non-random UMI, each specific sequence has at least 1 (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10) nucleotide differences from other specific sequences at their corresponding nucleotide positions; Preferably, the specific sequence of the non-random UMI is as shown in any one of SEQ ID NO:13 - 16.

17. The method according to any one of claims 12 - 16, wherein, The first oligonucleotide chain and / or the second oligonucleotide chain of the sequencing primer further contains a positioning tag, wherein the positioning tag contains at least 1 (e.g., 1, 2, 3, 4, 5, 6, 7, 8) specific sequence, and the specific sequence is composed of 3 - 8 (e.g., 3, 4, 5, 6, 7, 8) bases; Preferably, the bases are selected from A, T, C, G; Preferably, the specific sequence is composed of 3 - 8 identical or different bases; Preferably, the positioning tag contains 4 specific sequences, and the 4 specific sequences are all composed of identical or different fixed bases; preferably, the lengths of the 4 specific sequences are different from each other (e.g., differ by 1 base, differ by 2 bases, differ by 3 bases); Preferably, the positioning tag contains a first specific sequence composed of 3 identical or different bases, a second specific sequence composed of 4 identical or different bases, a third specific sequence composed of 5 identical or different bases, and a fourth specific sequence composed of 6 identical or different bases; Preferably, the first specific sequence has the sequence as shown in SEQ ID NO:17; preferably, the second specific sequence has the sequence as shown in SEQ ID NO:18; preferably, the third specific sequence has the sequence as shown in SEQ ID NO:19; preferably, the fourth specific sequence has the sequence as shown in SEQ ID NO:20; Preferably, the positioning tag further contains at least 1 base G at the 5' end of the specific sequence; preferably, the positioning tag is located in the hybridization sequence of the first oligonucleotide chain and / or the second oligonucleotide chain, or the positioning tag is located at the end of the hybridization sequence of the first oligonucleotide chain and / or the second oligonucleotide chain; Preferably, the positioning tag is located at the 3' end of the first hybridization sequence; preferably, the positioning tag of the first oligonucleotide chain is linked to the first strand in the nucleic acid molecule or its amplification product; Preferably, the positioning tag is located at the 5' end of the second hybridization sequence; preferably, the positioning tag of the second oligonucleotide chain is linked to the second strand in the nucleic acid molecule or its amplification product.

18. A DNA library constructed by the method according to any one of claims 12-17; Preferably, the DNA library is selected from a genomic library, a methylome library, or any combination thereof; Preferably, a genomic library and a methylome library are simultaneously and separately prepared by the method according to any one of claims 12-17; Preferably, the genomic library comprises or consists of the second strand in the nucleic acid molecule or its amplification product; Preferably, the methylome library comprises or consists of the first strand in the nucleic acid molecule or its amplification product.

19. A method for analyzing epigenetic information and / or genetic information, the method comprising: Detecting and / or analyzing the nucleic acid molecule or its amplification product according to claim 10, or detecting and / or analyzing the DNA library according to claim 18; Preferably, the first strand and the second strand in the nucleic acid molecule or its amplification product are separately detected and / or analyzed; Preferably, the genomic library and the methylome library are separately detected and / or analyzed. The method according to claim 19, wherein, The method has one or more of the following characteristics: (1) The detection is sequencing; preferably, the sequencing is selected from next-generation sequencing, massively parallel sequencing, pyrosequencing, sequencing by synthesis, single molecule real-time sequencing, Polony sequencing, DNA nanoball sequencing, SunTag single molecule sequencing, nanopore sequencing, Sanger sequencing, Shotgun sequencing, Gilbert sequencing analysis, or any combination thereof; (2) The detection is selected from microarray, quantitative PCR (qPCR), digital droplet PCR (ddPCR), molecular inversion probe, or any combination thereof; (3) The analysis is bioinformatics analysis; preferably, the bioinformatics analysis is selected from sequence alignment, genomic analysis, transcriptomic analysis, single nucleotide variant (SNV) analysis, gene copy number variant (CNV) analysis, measuring chromosome copy number, detecting genetic lesions, or any combination thereof; (4) The epigenetic information is DNA methylation information; (5) The genetic information is whole genome information and / or mutation information.

21. The method according to claim 20, wherein, Sequencing the nucleic acid molecule or its amplification product and obtaining sequencing data containing the first strand and the second strand, or separately obtaining the sequencing data of the first strand and the second strand; Preferably, the method has one or more of the following characteristics: (1) When the first oligonucleotide chain of the sequencing primer contains UMI and the second oligonucleotide chain does not contain UMI, the sequencing data of the first strand and the second strand are distinguished by identifying whether there is a UMI sequence in the sequencing data of the nucleic acid molecule or its amplification product; Preferably, the sequencing data with UMI identified is the sequencing data of the first strand; preferably, the sequencing data without UMI identified is the sequencing data of the second strand; (2) When the second oligonucleotide chain of the sequencing primer contains UMI and the first oligonucleotide chain does not contain UMI, the sequencing data of the first strand and the second strand are distinguished by identifying whether there is a UMI sequence in the sequencing data; Preferably, the sequencing data with UMI identified is the sequencing data of the second strand; preferably, the sequencing data without UMI identified is the sequencing data of the first strand; (3) When the first oligonucleotide strand of the sequencing primer contains UMI and the second oligonucleotide strand also contains UMI, the sequencing data of the first strand and the second strand are distinguished by identifying whether the UMI in the sequencing data is closer to the 5' single-stranded arm or the 3' single-stranded arm of the sequencing primer; Preferably, the sequencing data with UMI closer to the 5' single-stranded arm identified is the sequencing data of the first strand; preferably, the sequencing data with UMI closer to the 3' single-stranded arm identified is the sequencing data of the second strand; (4) The sequencing data of the first strand and the second strand are distinguished by separately identifying whether the cytosine of UMI in the first oligonucleotide strand and the second oligonucleotide strand has undergone cytosine conversion; Preferably, the sequencing data with the cytosine in UMI having undergone conversion (e.g., converted to uracil) is the sequencing data of the first strand; Preferably, the sequencing data with the cytosine in UMI not having undergone cytosine conversion (e.g., remaining as cytosine) is the sequencing data of the second strand.

22. The method according to claim 20, wherein, The method has one or more of the following characteristics: (1) When the first oligonucleotide strand of the sequencing primer contains a positioning tag and the second oligonucleotide strand does not contain a positioning tag, the sequencing data of the first strand and the second strand are distinguished by identifying whether the sequence of the positioning tag exists in the sequencing data of the nucleic acid molecule or its amplification product; Preferably, the sequencing data with the positioning tag identified is the sequencing data of the first strand; preferably, the sequencing data without the positioning tag identified is the sequencing data of the second strand; (2) When the second oligonucleotide strand of the sequencing primer contains a positioning tag and the first oligonucleotide strand does not contain a positioning tag, the sequencing data of the first strand and the second strand are distinguished by identifying whether the positioning tag sequence exists in the sequencing data; Preferably, the sequencing data with the positioning tag identified is the sequencing data of the second strand; preferably, the sequencing data without the positioning tag identified is the sequencing data of the first strand; (3) When the first oligonucleotide strand of the sequencing primer contains a positioning tag and the second oligonucleotide strand also contains a positioning tag, the sequencing data of the first strand and the second strand are distinguished by identifying whether the positioning tag in the sequencing data is closer to the 5' single-stranded arm or the 3' single-stranded arm of the sequencing primer; Preferably, the sequencing data with the positioning tag closer to the 5' single-stranded arm identified is the sequencing data of the first strand; preferably, the sequencing data with the positioning tag closer to the 3' single-stranded arm identified is the sequencing data of the second strand; (4) The sequencing data of the first strand and the second strand are distinguished by separately identifying whether the sequence of the positioning tag in the first oligonucleotide strand and the second oligonucleotide strand has changed; Preferably, the sequencing data with guanine (G) in the positioning tag converted to adenine (A) is the sequencing data of the first strand; Preferably, the sequencing data with guanine (G) in the positioning tag not having undergone conversion is the sequencing data of the second strand.

23. The method according to claim 21 or 22, wherein, The sequencing data of the first strand is used for the analysis of the methylation information of the target nucleic acid; Preferably, the sequencing data of the second strand is used for the analysis of the genomic information of the target nucleic acid.

24. The DNA library according to claim 18 is used for analyzing epigenetic information and / or genetic information; Preferably, the DNA library is selected from a genomic library, a methylome library, or any combination thereof; Preferably, the epigenetic information is DNA methylation information; Preferably, the genetic information is whole genome information and / or mutation information.

25. Use of the nucleic acid molecule according to claim 10 or its amplification product, or the DNA library according to claim 18 in the preparation of a kit for detecting whether a subject has a disease; Preferably, the disease causes changes in the epigenetic information and / or genetic information of the subject (for example, nucleotide transitions or transversions, nucleotide insertions or deletions, genomic rearrangements, copy number variations or gene fusions); Preferably, the disease is cancer; Preferably, the disease is a genetic disease; Preferably, the kit is used for: (1) Detecting whether a subject has cancer; (2) Detecting whether a subject has cancer recurrence; (3) Detecting the responsiveness of a subject with a disease to the therapy and / or drug received for the disease; or, (4) Detecting whether a subject has a genetic disease; Preferably, the cancer is selected from liver cancer, hepatocellular carcinoma, melanoma, pancreatic cancer, lung cancer, kidney cancer, gastric cancer, esophageal cancer, colon cancer, breast cancer, ovarian cancer, cervical cancer, testicular cancer, prostate cancer, lymphoma, B-cell lymphoma, diffuse large B-cell lymphoma, follicular lymphoma, mantle cell lymphoma, small lymphocytic lymphoma, splenic marginal zone B-cell lymphoma, extranodal marginal zone mucosa-associated lymphoid tissue B-cell lymphoma, nodal marginal zone B-cell lymphoma, lymphoplasmacytic lymphoma, primary effusion lymphoma, Burkitt lymphoma / Burkitt cell leukemia, T-cell lymphoma, anaplastic large cell lymphoma (primary cutaneous type), anaplastic large cell lymphoma, (systemic type), peripheral T-cell lymphoma, angioimmunoblastic T-cell lymphoma, adult T-cell lymphoma / leukemia (human T-cell lymphotropic virus type I positive), extranodal NK / T-cell lymphoma (nasal type), enteropathy-associated T-cell lymphoma, gamma / delta hepatosplenic T-cell lymphoma, subcutaneous panniculitis-like T-cell lymphoma, multiple myeloma, mycosis fungoides; Preferably, the genetic disease is selected from Alzheimer's disease (APOE1), Charcot-Marie-Tooth disease, Leber hereditary optic neuropathy (LHON), Angelman syndrome (UBE3A, ubiquitin-protein ligase E3A), Prader-Willi syndrome (region in chromosome 15), β-thalassemia (HBB, β-globin), Gaucher disease (type I) (GBA, glucocerebrosidase), cystic fibrosis (CFTR epithelial chloride channel), sickle cell disease (HBB, β-globin), phenylketonuria (PAH, phenylalanine hydroxylase), familial hypercholesterolemia (LDLR, low density lipoprotein receptor), Huntington's disease (HDD, huntingtin), neurofibromatosis type I (NF1, NF1 tumor suppressor gene), myotonic dystrophy (DM, coix seed), tuberous sclerosis (TSC1, tuberin), achondroplasia (FGFR3, fibroblast growth factor receptor), fragile X syndrome (FMR1, RNA binding protein), Duchenne muscular dystrophy (DMD, dystrophin), hemophilia A (F8C, coagulation factor VIII), Lesch-Nyhan syndrome (HPRT1, hypoxanthine-guanine phosphoribosyltransferase 1), and adrenoleukodystrophy (ABCD1).

26. A kit, the kit comprising: (1) Propargyl-modified cytosine; (2) a cytosine conversion reagent (e.g., bisulfite, cytidine deaminase, and / or TET); and (3) the single-stranded linker and / or double-stranded linker according to any one of claims 4-8; optionally, the kit further comprises: (4) a reagent for library construction; Preferably, the reagent for library construction is selected from an enzyme, a reagent for DNA end repair, RNase-free water, or any combination thereof; Preferably, the enzyme is selected from T4 DNA ligase, T3 DNA ligase, T7 DNA ligase, Escherichia coli DNA ligase, HIFI Taq DNA ligase, T4 RNA ligase, RTCB ligase, circular ligase, heat-stable DNA ligase, or any combination thereof.

Citation Information

Patent Citations

  • Novel construction method for genome methylation library and application thereof

    CN110195095A

  • Methods of determining the number of copies or sequence of one or more RNA molecules

    WO2023012065A1