Tail sequence directed linker addition

CN122535709APending Publication Date: 2026-08-07富国良 +1
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
富国良
Filing Date
2024-12-16
Publication Date
2026-08-07

Smart Images

  • Figure FT_1
    Figure FT_1
  • Figure FT_2
    Figure FT_2
  • Figure FT_3
    Figure FT_3
Patent Text Reader

Abstract

The present invention relates to methods, compositions and kits for constructing a polynucleotide sequencing library using a first primer and utilizing the 3' end or a copy thereof to direct the addition of an adaptor sequence for replication of a target polynucleotide. The sequencing library is suitable for massively parallel sequencing and comprises a plurality of double-stranded nucleic acid molecules.
Need to check novelty before this filing date? Find Prior Art

Description

Background of the Invention

[0002] This invention relates to methods and compositions for amplifying target polynucleotide populations, wherein the amplified portion is processed to generate epigenetic and / or genetic information. Copies of the target polynucleotide are generated using oligonucleotides, these copies being enriched with regions containing epigenetic information, which are then fitted with adapters.

[0003] Next-generation DNA sequencing (NGS) continues to revolutionize clinical medicine and basic research, particularly in the rapidly evolving field of liquid biopsy testing. Genetic and epigenetic biomarkers detected from cell-free DNA extracted from liquid biopsies have the ability to identify the presence of diseases such as cancer. Among these biomarkers, aberrant DNA methylation has been shown to be associated with multiple disease processes, including cancer—altered methylation patterns lead to abnormal gene expression regulation, which in turn causes dysregulation of normal cellular processes.

[0004] DNA methylation profiling using methylation sequencing (e.g., whole-genome sulfite sequencing, WGBS) is considered an important diagnostic tool for detecting, diagnosing, and / or monitoring cancer. For example, specific patterns in differentially methylated regions can serve as molecular biomarkers for various diseases and disease stages. DNA methylation is enriched in regions of the human genome known as CpG islands, which have high GC content but account for only about 1% of the human genome. Therefore, "whole-genome" methods are costly and inefficient, generating large amounts of data that are of little informational value—these data either lack differential methylation in cancer or have localized CpG densities that are too low to provide reliable signals. Given this limitation, methylation target enrichment methods will help improve the sensitivity of disease biomarkers while reducing detection costs.

[0005] Detailed description

[0006] While various embodiments of compositions and methods have been shown and described herein, those skilled in the art will understand that these embodiments are provided by way of example only. Many variations, modifications, and substitutions can be made by those skilled in the art without departing from the compositions and methods of the invention. It should be understood that various alternatives may be adopted for the embodiments described herein.

[0007] To facilitate understanding of this invention, several terms are defined below.

[0008] Terminology Definition

[0009] As used herein, "sample" means any substance that contains or may contain nucleic acids, including tissue or liquid samples isolated from one or more individuals.

[0010] As used herein, "nucleotide sequence" refers to a homopolymer or hybrid of deoxyribonucleotides, ribonucleotides or other nucleic acids, or any combination of nucleic acids.

[0011] As used herein, "nucleotide" generally refers to the monomeric component of a nucleotide sequence, although these monomers may be nucleosides and / or nucleotide analogs, and / or modified nucleosides (such as amino-modified nucleosides), as well as nucleotides. Furthermore, "nucleotide" also includes "nucleoside triphosphates" and naturally occurring or non-natural analog structures developed through selective / targeted approaches.

[0012] As used herein, "nucleic acid" refers to at least two nucleotides covalently linked together. The nucleic acids of this invention typically contain phosphodiester bonds, but in some cases also contain nucleic acid analogs with an alternative backbone. Nucleic acids can be single-stranded or double-stranded, or contain both single-stranded and double-stranded sequences. Nucleic acids can be DNA (genomic DNA and cDNA), RNA, a mixture of DNA and RNA, or a DNA-RNA hybrid, wherein the nucleic acid contains any combination of deoxyribonucleotides and ribonucleotides, and any combination of bases including uracil, adenine, thymine, cytosine, guanine, hypoxanthine, xanthine, and so on. Both "DNA sequence" and "RNA sequence" can include single-stranded and double-stranded DNA or RNA. Unless the context otherwise requires, a specific sequence refers to the single-stranded DNA or RNA of that sequence, the double-stranded form of that sequence with its complementary strand (double-stranded DNA or RNA), and / or the complementary strand of that sequence.

[0013] As used herein, "polynucleotide" and "oligonucleotide" are types of "nucleic acids," typically indicators, or oligomeric fragments to be detected. There is no specific distinction in length between "nucleic acid," "polynucleotide," and "oligonucleotide," and these terms are used interchangeably. "Nucleic acid," "DNA," and similar terms also include nucleic acid analogs. Oligonucleotides are not necessarily physically derived from any existing or natural sequence and can be generated in any way, including chemical synthesis, enzymatic synthesis, DNA replication, reverse transcription, or any combination thereof.

[0014] As used herein, the terms "original target polynucleotide," "target sequence," "target nucleic acid," "target nucleic acid sequence," and "target nucleic acid" are used interchangeably to refer to the target region to be amplified, detected, or both, or the object that hybridizes with complementary oligonucleotides, polynucleotides (e.g., blocking oligomers), or the object in the primer extension process. The target sequence may consist of DNA, RNA, analogues, or any combination thereof, and may be single-stranded or double-stranded. During primer extension, the target nucleic acid that forms a double strand with the primer is also called the "template." The template serves as a pattern for synthesizing the complementary polynucleotide. This invention may use target sequences derived from any living organism or formerly existing organism, including but not limited to prokaryotes, eukaryotes, plants, animals, and viruses, as well as synthetic and / or recombinant target sequences; it may also be a mixture of nucleic acids, wherein the target nucleic acid is a subset of the total nucleic acids.

[0015] As used herein, the term "primer" is used interchangeably to describe one or more primers or a set of primers, referring to naturally occurring or synthetically produced oligonucleotides. Multiple primers in a set may have different sequences and hybridize at multiple different positions. "First primer," "set of first primers," and "set of first primers" are interchangeable, as is "second primer." Functionally, a "primer" is a molecule that, under appropriate conditions, can act as a synthesis initiation site, is complementary to the nucleic acid strand, and can initiate the synthesis of primer extension products, i.e., in the presence of nucleotides and a polymerizing agent (such as DNA polymerase), at appropriate temperatures and buffer solutions. Such conditions include one, two, three, or four different deoxyribonucleoside triphosphates, including but not limited to deoxyadenosine triphosphate (dATP), deoxythymidine triphosphate (dTTP), deoxyguanosine triphosphate (dGTP), and deoxycytidine triphosphate (dCTP), or suitable additional or substituted nucleotides, unconventional nucleotides, and polymerization inducers (such as DNA polymerase and / or RNA polymerase and / or reverse transcriptase), in a suitable buffer (the "buffer" includes cofactors or components that affect pH, ionic strength, etc.), and at a suitable temperature. Primers are preferably single-stranded to achieve maximum efficiency in amplification. Primers selected herein are substantially complementary to one strand of each specific sequence to be amplified. This means that primers must be sufficiently complementary to their respective strands to hybridize. One or more non-complementary sequence regions may be attached to the 5' end (5' tail portion) of the primer or to the interior of the primer (convex loop portion), with the remainder of the primer sequence complementary to the target fragment of the target base sequence. Typically, primers are complementary unless non-complementary nucleotides are present in the predetermined primer terminal or middle regions. In other words, the primers in this paper are chosen to be substantially identical to one strand of each specific sequence to be amplified. This means that the primers must be sufficiently identical to one strand to be able to hybridize with the other strand of their respective sequence.

[0016] As used herein, a "linker" refers to an oligonucleotide designed to serve as a substrate for ligases or a template for polymerase extension in a reaction. The linker sequence may be complementary to or identical to the oligonucleotide sequence. Linkers may consist of functional components. The term "functional component" is used interchangeably to describe any position or nucleotide within the linker.

[0017] As used herein, "complementarity" refers to the ability of two nucleotide sequences (random or designed) to bind together in a sequence-complementary manner according to the Watson-Crick rule via hydrogen bonds between purine and / or pyrimidine bases. Alternatively, it can refer to the ability of nucleotide sequences containing modified nucleotides or deoxyribonucleotides, ribonucleotide analogs, or combinations thereof to bind together in a sequence-specific manner via unconventional Watson-Crick rules to form alternative nucleic acid double-stranded structures.

[0018] As used in this article, "hybridization" and "annealing" are interchangeable, referring to the process by which two partially or completely complementary nucleotide sequences combine to form a double-stranded sequence or fragment.

[0019] "Double helix" and "double strand" are interchangeable, both referring to structures formed by the hybridization of two complementary nucleic acid sequences. Such double helixes can be formed by the complementary binding of two DNA fragments, two RNA fragments, one DNA fragment and one RNA fragment, or two fragments composed of a mixture of RNA and DNA, the latter being called a hybrid double helix. Either or both members of a double helix may contain modified nucleotides and / or nucleotide analogs, as well as nucleoside analogs. As described herein, such double helixes can be formed by the binding of one or more blocking oligonucleotides to a sample sequence. Double helixes can be partially or completely complementary, or partially or completely double-stranded.

[0020] As used in this article, "wild-type nucleic acid", "normal nucleic acid", "nucleic acid containing normal nucleotides", "wild-type", "normal", "wild-type DNA" and "wild-type template" are used interchangeably to refer to polynucleotides that have a nucleotide sequence that is considered normal or unchanged.

[0021] As used herein, "mutated polynucleotide," "mutated nucleic acid," "variant nucleic acid," and "nucleic acid containing a variant nucleotide" refer to polynucleotides whose nucleotide sequence differs from the expected nucleotide sequence of the corresponding wild-type polynucleotide. The difference in nucleotide sequence between a mutant polynucleotide and a wild-type polynucleotide is referred to as a nucleotide "mutation," "variant nucleotide," "variation," or "change." A "variant nucleotide" also refers to a substitution, deletion, insertion, methylation, and / or modification of one or more nucleotides.

[0022] As used in this article, "amplification" refers to increasing the concentration or copy number of a specific nucleic acid sequence in a mixture of nucleic acid sequences using any amplification procedure. Amplification can be one or more rounds of linear amplification, one or more rounds of exponential amplification, or a combination of both.

[0023] As used in this article, "replication" or "replicate" refers to the process of creating complementary copies using polynucleotides as templates for polymerase extension. Multiple rounds of replication can achieve amplification.

[0024] As used herein, "reaction mixture," "amplification mixture," or "PCR mixture" refers to a mixture of components required to amplify at least one product, which may contain one or more nucleotides (dNTPs), polymerase (thermostable or non-thermostable), primers, various nucleic acid templates, and other specific nucleotides required by the present invention. The mixture may further contain Tris buffer, monovalent salts, and Mg²⁺. Except for the specific nucleotides required by the present invention, the concentrations of the components are well known in the art and can be further optimized by those skilled in the art.

[0025] "Amplification product" or "amplifier" refers to a DNA or RNA fragment amplified by amplification methods using polymerase and a primer, primer pool, a pair of primers, multiple primer pools, or any combination thereof.

[0026] "Primer extension products" refer to DNA or RNA fragments extended by a reaction using a polymerase and one or a pair of primers. This can involve single extension (such as first-strand cDNA synthesis), double extension (such as double-strand cDNA synthesis), or multiple-cycle extension, including PCR or isothermal amplification using polymerases with strand displacement activity.

[0027] "Compatible" means that the primer sequence or part thereof is the same as, substantially the same as, complementary to, substantially complementary to or similar to the PCR primer / sequencing primer sequence used in a massively parallel sequencing platform.

[0028] Unless otherwise stated, the present invention may be practiced using conventional techniques familiar to those skilled in the art, including molecular biology, microbiology, recombinant DNA, and next-generation sequencing technologies. All patents, patent applications, and publications mentioned herein (whether prior or subsequent) are incorporated herein by reference.

[0029] This article discloses methods, systems, and compositions that can significantly increase the amount of information obtained from a single patient sample, especially superior to other techniques that employ a whole-sample non-enrichment workflow.

[0030] Cytosine methylation forms 5-methylcytosine (5mC or mC), for example at the cytosine-phosphate-guanine motif (CpG), which can serve as an epigenetic marker, playing important roles in development and tissue specificity, genomic imprinting, and environmental responses. Dysregulation of 5mC can lead to abnormal gene expression and, in some cases, affect cancer risk, progression, or treatment response. 5-hydroxymethylcytosine (5hmC or hmC) can act as an intermediate in the active DNA demethylation pathway, exhibiting tissue-specific distribution and influencing gene expression and carcinogenesis.

[0031] DNA methylation (in this context, methylation can refer to the addition of or presence of methyl groups in nucleic acid bases; methyl groups can be oxidized or non-oxidized; non-oxidized methyl groups are, for example, methyl groups; oxidized methyl groups can be hydroxymethyl, formyl, carboxylic acid, or carboxylates) plays a role in the silencing of repetitive DNA elements and alternative splicing. DNA methylation is associated with a variety of biological processes, such as genomic imprinting, transposon inactivation, stem cell differentiation, transcriptional repression, and inflammation. DNA methylation maps can be inherited through cell division in some cases and sometimes across generations. Since methyl markers play important roles under both physiological and pathological conditions, mapping DNA methylation to answer biological questions may be of great significance. Furthermore, revealing genomic regions with DNA methylation is attractive for translational research because methylation sites can be altered through pharmacological intervention.

[0032] Early detection of cancer in patients is crucial because it often allows for earlier intervention, thereby improving survival rates. During cancer development, genomic methylation patterns often change while retaining characteristics of their cellular origin. These changes can also be identified by examining cell-free DNA (cfDNA) fragments, providing information relevant to cancer identification and classification in a low-cost, non-invasive manner, enabling early cancer detection. By employing an enrichment method targeting any methylated genomic region, rather than sequencing all nucleic acids in the test sample (i.e., "whole-genome sequencing"), the method disclosed herein increases the sequencing depth of the target region and reduces costs compared to whole-genome sequencing (WGS) or whole-genome sulfite sequencing (WGBS).

[0033] This invention, used for cancer detection, can further provide information related to the identification of cancer presence, cancer staging assessment, and determination of cancer tissue origin. This specification also provides a method for diagnosing cancer, wherein the cancer diagnosis further includes cancer type and / or cancer stage. This document further provides methods for identifying genomic sites with cancer- or various cancer type-specific methylation patterns.

[0034] This paper discloses a detection method for enriching cfDNA molecules for cancer diagnosis. It also discloses a detection method for generating sequencing libraries. These sequencing libraries are enriched specific sequence populations, such as libraries enriched for CG. Alternatively, the sequencing library can also be a library of random global sequence populations using random primers.

[0035] The methods described herein can be used to identify and recognize genetic alterations and / or epigenetic modifications present in the original target polynucleotide. In some cases, the target polynucleotide is treated with reagents to differentially transform cytosine and methylated cytosine, producing transformed target polynucleotides. These transformed target polynucleotides are then amplified to generate amplified transformed polynucleotides, which are further processed through a workflow to generate an NGS library. The final library derived from the original target polynucleotide can be sequenced using any suitable method, including but not limited to Pacific Biosciences, Ion Torrent sequencing, Illumina, Nanopore, GenapSys, MGI, or Element Biosciences instruments. Sequencing data can be analyzed to reveal some or all of the following disease-related alterations, including mutations, hypomethylation, and / or hypermethylation, which, individually or in combination, can characterize the disease and / or serve as epigenetic changes, and can be used to identify tissue origin.

[0036] In some cases, epigenetic markers can be determined using a computer program (e.g., containing instructions for analyzing sequencing data and / or performing one or more steps of the methods described herein). In some cases, such a computer program can be stored in the computer's memory.

[0037] This invention provides a method for adding adapter sequences to target polynucleotides in a sample, the method comprising:

[0038] a) Provides a reaction mixture comprising at least one first primer (short-tailed primer), the short-tailed primer comprising a 3' initiating portion and a 5' universal tail portion, wherein the 3' initiating portion is capable of hybridizing with a target polynucleotide and initiating extension, and wherein the 5' universal tail portion or its complementary strand is capable of mediating the addition of a guiding linker sequence to the primer extension product;

[0039] b) Perform at least one round of extension reaction, wherein the extension reaction includes hybridization of the primer with the target polynucleotide and extension under extension conditions to produce a primer extension product, wherein the primer comprises a short-tailed primer; and

[0040] c) Adding adapter sequences to the extension product, mediated by hybridizing a specific oligonucleotide to the 3' and / or 5' end of the extension product, wherein the specific oligonucleotide contains a sequence that is identical to or complements the 5' universal tail portion of the short-tail primer.

[0041] At least one round of extension can be two or more rounds, wherein the first round produces an extension product, and the second or more rounds produce copies of the extension product, wherein a specific oligonucleotide is capable of hybridizing with the 3' end sequence of the extension product, which is the complementary strand of the 5' universal tail sequence of the short-tailed primer in the extension product.

[0042] At least one round of extension reaction can be a multi-round extension of isothermal amplification using a polymerase with strand displacement activity. The polymerase is selected from Klenow, Bst polymerase, and φ29 DNA polymerase group. At least one round of extension reaction can also be a multi-round extension of a PCR reaction.

[0043] Specific oligonucleotides can be adapter template oligonucleotides (ATO). Adding adapter sequences to the extension product is done by extending the 3' end of the extension product, where the 3' end of the extension product serves as a primer and the ATO serves as a template.

[0044] The specific oligonucleotide can be a splice oligonucleotide that can hybridize with the 3' end sequence of the extension product and the linker oligonucleotide. Adding a linker to the extension product includes hybridizing the splice oligonucleotide to the 3' or 5' end of the extension product and the linker oligonucleotide, and linking the 3' or 5' end of the extension product to the 5' or 3' end of the linker oligonucleotide.

[0045] The length of the 5' universal tail portion can be 2 to 12 nucleotides, 2 to 9 nucleotides, 2 to 6 nucleotides, or 2 to 4 nucleotides.

[0046] Short-tailed primers may contain degradable nucleotides. Degradable nucleotides may include uracil nucleotides. The 3' initiation region may contain a target-specific sequence. The 3' initiation region may contain a random sequence. The 3' initiation region may contain both a random sequence and a target-specific sequence.

[0047] The target polynucleotide can be sulfite or enzymatically converted DNA, in which cytosine in the original DNA has been converted into uracil.

[0048] The method may further include step d) amplification using primers capable of hybridizing with adapter sequences.

[0049] A kit for preparing sequencing libraries, comprising:

[0050] a) At least one primer (short-tailed primer) having a 3' initiating portion and a 5' universal tail portion, wherein the 3' initiating portion is capable of hybridizing with a target polynucleotide and initiating extension, wherein the 5' universal portion of the extension product or the 3' end of the extension product derived from the 5' universal tail portion of the primer is capable of mediating the addition of a guiding linker sequence to the extension product; and

[0051] b) A linker template oligonucleotide (ATO) or splice oligonucleotide capable of hybridizing with the 3' end of the extension product, wherein the ATO or splice oligonucleotide contains a sequence that is identical or complementary to the 5' universal tail portion of the short-tail primer.

[0052] To enrich specific target sequences, extension of a specific primer (the first short-tailed primer) generates a copy of the target sequence. However, the nucleic acid population remains a mixture of the original and target sequences. The original sequence may be majority-proportionate in the nucleic acid population, especially in single or few extension runs. For selective analysis of the target sequence (e.g., sequencing), it is necessary to add adapters only to the target sequence, not to the original sequence. The 5' tail sequence of the primer or its complementary sequence at the 3' end of the extension product provides guidance for the targeted addition of adapter sequences to the target sequence.

[0053] This invention provides a method for guiding the addition of adapters to target polynucleotides in a sample, comprising:

[0054] a) Provides a reaction mixture comprising at least one primer (short-tailed primer) having a 3' initiating portion and a 5' universal tail portion, wherein the 3' initiating portion is capable of hybridizing with a target polynucleotide and initiating extension, and wherein the 5' universal tail portion or its complementary strand is capable of mediating the addition of a guiding linker sequence to the primer extension product;

[0055] b) Perform at least one round of extension reaction, wherein the extension reaction includes hybridization of the primer with the target polynucleotide (which may be the original polynucleotide or a copy generated from the previous round of primer extension), and extension under extension conditions to produce the primer extension product (initial extension product); and

[0056] c) Add adapters to the extension product, mediated by hybridizing specific oligonucleotides to the 3' and / or 5' ends of the extension product.

[0057] At least one round of extension can be two or more rounds, wherein the first round produces an extension product, and the second or more rounds produce copies of the extension product, wherein a specific oligonucleotide is capable of hybridizing with the 3' end of the extension product, which can be the complementary strand of the 5' universal tail in the extension product.

[0058] At least one round of extension reaction can be a multi-round extension of isothermal amplification using a polymerase with chain displacement activity.

[0059] The polymerase can be any polynucleotide polymerase, selected from Klenow, Bst polymerase, or φ29 DNA polymerase group.

[0060] Multiple-round extension reactions can also be PCR reactions.

[0061] In one embodiment, the specific oligonucleotide is an adapter template oligonucleotide (ATO), and adding an adapter to the extension product may include: hybridizing the adapter template oligonucleotide (ATO) to the 3' end of the extension product, and extending the product using the 3' end as a primer and the ATO as a template.

[0062] Adding a linker to the extension product may also include hybridizing a splice oligonucleotide to the 3' or 5' end of the extension product and linking the 3' or 5' end of the extension product to the 5' or 3' end of the linker oligonucleotide (both hybridizing to the splice oligonucleotide).

[0063] The 5' universal tail portion of a short-tailed primer can be 12 to 2 nucleotides, 9 to 2 nucleotides, 6 to 2 nucleotides, or 4 to 2 nucleotides in length.

[0064] The 5' universal tail portion may contain degradable nucleotides, such as uracil nucleotides, inosine nucleotides, or RNA.

[0065] The 3' initiation portion may contain a target-specific sequence, a random sequence, or a combination of a random sequence and a target-specific sequence.

[0066] The target polynucleotide can be sulfite or enzymatically converted DNA, in which cytosine in the original DNA has been converted into uracil.

[0067] The target polynucleotide can be partially or completely processed through other workflows. The target polynucleotide can be processed in workflows such as ATO reactions or ligation workflows to extend its 3' end, and the extended portion may contain nucleotides that can be selectively destroyed. The extended target polynucleotide can be amplified to generate copies through single or multiple linear or PCR amplifications, wherein the incorporated nucleotides can subsequently allow selective destruction of the amplification product, for example, by incorporating uracil nucleotides and treating with UDG. A portion of the amplified extended target polynucleotide can be extracted for other workflows, such as targeted or whole-sample enrichment by PCR, probe capture, or other enrichment methods. The remaining amplified extended target polynucleotide can be selectively destroyed or inactivated (e.g., by UDG treatment), and the remaining target polynucleotide and / or extended target polynucleotide is subsequently used in this invention.

[0068] The present invention also provides a method for preparing a sequencing library, comprising:

[0069] a) Provides a reaction mixture comprising at least one primer (short-tailed primer) having a 3' initiating portion and a 5' universal tail portion, wherein the 3' initiating portion is capable of hybridizing with a target polynucleotide and initiating extension, and wherein the 5' universal tail portion or its complementary strand is capable of mediating the addition of a guiding linker sequence to the extension product;

[0070] b) Perform at least one round of extension reaction, wherein the extension reaction includes hybridization of the primer with the target polynucleotide and extension under extension conditions to produce primer extension product.

[0071] c) Adding linkers to the extension product, mediated by hybridizing specific oligonucleotides to the 3' and / or 5' ends of the extension product; and

[0072] d) Amplification is performed using primers that can hybridize with the adapter sequence.

[0073] This invention also provides a kit for preparing sequencing libraries, comprising:

[0074] a) At least one primer (short-tailed primer) having a 3' initiating portion and a 5' universal tail portion, wherein the 3' initiating portion is capable of hybridizing with a target polynucleotide and initiating extension, wherein the 5' universal portion of the extension product (equivalent to the 5' universal portion of the primer) or the 3' end of the extension product (equivalent to the 5' universal portion) is capable of mediating the addition of a guiding linker sequence to the extension product; and

[0075] b) Linker template oligonucleotides (ATO) or splice oligonucleotides capable of hybridizing with the 3' end of the extension product.

[0076] The methods and systems described herein may include providing and / or processing (e.g., chemical or enzymatic treatment) single-stranded or double-stranded DNA polynucleotides. The single-stranded or double-stranded DNA polynucleotides may contain the original target polynucleotides described herein. In many cases, the original target polynucleotide may be transformed with reagents to produce a transformed polynucleotide. In many cases, the original or transformed polynucleotide is hybridized and extended or amplified using short-tailed primers to produce an extension product (which may be the initial extension product). In many cases, the 3' end of the extension product is labeled with adapters by hybridizing a specific oligonucleotide to the 3' and / or 5' end of the extension product.

[0077] Preferably, the method for generating the 3' extension is the ATO reaction provided in WO2018 / 193233 A1, the entire contents of which are incorporated herein by reference.

[0078] In some embodiments, the target polynucleotide is modified by a reagent to produce a transformed target polynucleotide. This reagent chemically and / or enzymatically modifies the target polynucleotide to achieve a distinction between methylated and unmethylated cytosine. The reagent can be a sulfite treatment, which converts cytosine to uracil but not methylated cytosine (i.e., 5-methylcytosine, which is resistant to this treatment and remains cytosine). The reagent can also be an enzymatic treatment, such as a combination of a TET family member and APOBEC (or an equivalent enzyme), which converts unmethylated C to U but not methylated cytosine. The reagent can also be a "TAPS chemistry" chemical transformation, involving the conversion of 5-methylcytosine (and 5-hydroxymethylcytosine) to 5-carboxycytosine, followed by the selective conversion of the 5-carboxycytosine residues to dihydrouracil using pyridineborane.

[0079] The target polynucleotide is preferably naturally or artificially fragmented. The target polynucleotide can be any nucleic acid, such as DNA, cDNA, RNA, mRNA, small RNA, or microRNA, or any combination thereof. The target polynucleotide may contain multiple target polynucleotides. Each target polynucleotide may contain different or the same sequence. One or more target polynucleotides may contain variant sequences.

[0080] Target polynucleotides can be obtained from natural sources or synthesized. Natural sources include prokaryotic or eukaryotic RNA and / or genomic DNA. For example, but not limited to, sources can be animals including humans or mice, viruses, plants, or bacteria. In various ways, target polynucleotides are tagged or extended with adapter sequences at the 3' end for microarray detection and the creation of next-generation nucleic acid sequencing libraries.

[0081] If the target polynucleotide source is genomic DNA or RNA, or both, in some embodiments, the genomic DNA or RNA, or both, are fragmented before elongation. Genomic DNA / RNA fragmentation is a routine procedure well known to those skilled in the art, including, but not limited to, in vitro operations, such as DNA / RNA shearing (nebulization), cleavage with endonucleases, sonication, heating, irradiation with α, β, γ rays or other radiation sources, light exposure, chemical cleavage in the presence of metal ions, free radical cleavage, and combinations thereof. Genomic DNA / RNA fragmentation can also occur in vivo, for example, but not limited to, by apoptosis, radiation, and / or asbestos exposure. According to the methods described herein, the target polynucleotide population is not required to be uniform in size. Therefore, the methods of the present invention are applicable to populations of target polynucleotide fragments of varying sizes.

[0082] In one embodiment, generating the initial extension product involves using a first primer (also known as a short-tailed primer) and a target polynucleotide (transformed or original), wherein the first primer hybridizes to the target polynucleotide and is extended by polymerase. The first primer anneals to the target polynucleotide. If the first primer contains a 3' random sequence, the first primer can anneal to any position of any polynucleotide in the sample. The first primer may also contain additional sequences compatible with the NGS platform, such as a 5' tail containing the necessary sequence. The first primer may also contain unconventional nucleotides, such as uracil nucleotides or inosine (bisinosine) nucleotides, making all or part of the first primer unreplicable by polymerase or degradable by reagents such as glycosidases.

[0083] In some cases, the first oligonucleotide is a target-specific primer pool, complementary to the forward or reverse strand of the target polynucleotide or the transformed target polynucleotide. In some embodiments, the first primer may be a short-tailed primer having a 3' initiating portion and a 5' universal tail portion, wherein the 3' initiating portion is capable of hybridizing with the target polynucleotide and initiating extension, and wherein the 5' universal tail portion or its complementary strand is capable of mediating the addition of a guiding linker sequence to the primer extension product.

[0084] In some cases, the 3' end of the first primer can be target-specific. In some cases, the 3' end of the first primer can be randomly selected from A, T, C, and / or G nucleotides. In some cases, the 3' end of the first primer can be biased or contain degenerate bases. In some cases, the content of G and C biased nucleotides is higher. In some cases, the content of G and C biased, as well as A or T nucleotides, is higher. In some cases, the 3' end can be a combination of fixed bases and random or degenerate bases. In some cases, the terminal end can be a target-specific sequence, such as a CG dinucleotide. In some cases, the 3' end can have 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more consecutive or discontinuous CG dinucleotides. In some cases, the 5' end of the first primer can be a conserved or universal sequence. In some cases, a degradable or non-polymerase-template nucleotide may be present between the 3' and 5' ends of the first primer. In some cases, the non-polymerase-template nucleotide can be inosine or uracil nucleotide. In some cases, the first primer may also contain a sample barcode (SBC) sequence and additional universal sequences required for compatibility with the NGS platform.

[0085] In some cases, the initial extension product is generated using the first primer in a single (one-round) extension. Amplification can be performed in 1, 2, 3, 4, 5, 1-30, 2-25, 3-24, 4-23, 5-22, 6-21, 7-20, 8-19, 9-18, 10-17, 31-40, 41-50, 51-100 or more cycles. In some cases, the initial extension product is generated by isothermal amplification. In some cases, the initial extension product is generated by a PCR reaction. In some cases, multiple extension cycles can be performed by repeated temperature cycling (annealing, extension, and denaturation).

[0086] In some cases, the initial extension product can be generated using any polymerase and / or any reverse transcriptase, or a mixture of different polymerases. In many cases, the polymerase is a DNA polymerase. In some cases, the polymerase may have 3' to 5' exonuclease activity. In some cases, the DNA polymerase may be active at low temperatures. The polymerase may comprise a mixture of different polymerases, which may have 3' to 5' exonuclease activity, 5' to 3' exonuclease activity, and / or strand displacement activity.

[0087] Polymerases that can be used in the methods described herein include, but are not limited to: Deep VentR™ DNA polymerase, LongAmp™ Taq DNA polymerase, Phusion™ High Fidelity DNA Polymerase, Phusion™ Hot-Start High Fidelity DNA Polymerase, VentR® DNA Polymerase, DyNAzyme™ II Hot-Start DNA Polymerase, Phire™ Hot-Start DNA Polymerase, CrimsonLongAmp™ Taq DNA Polymerase, DyNAzyme™ EXT DNA Polymerase, LongAmp™ Taq DNA Polymerase, Standard Taq Buffer (Mg-Free) Taq DNA Polymerase, Standard Taq Buffer Taq DNA Polymerase, ThermoPol II Buffer (Mg-Free) Taq DNA Polymerase, ThermoPol Buffer Taq DNA Polymerase, Crimson Taq™ DNA Polymerase, Mg-Free Buffer Crimson Taq™ DNA Polymerase, VentR® (exo-) DNA Polymerase, Hemo KlenTaq™, DeepVentR™ (exo-) DNA Polymerase, and ProtoScript®. AMV First-Strand cDNA Synthesis Kit, ProtoScript® M-MuLV First-Strand cDNA Synthesis Kit, Bst DNA Polymerase (Full-Length), Bst DNA Polymerase (Large Fragment), ThermoPol Buffered Taq DNA Polymerase, 9°Nm DNA Polymerase, Crimson Taq™ DNA Polymerase, Mg-Free Buffered Crimson Taq™ DNA Polymerase, Deep VentR™ (exo-) DNA Polymerase, Deep VentR™ DNA Polymerase, DyNAzyme™ EXT DNA Polymerase, DyNAzyme™ II Hot-Start DNA Polymerase, Hemo KlenTaq™, Phusion™ High-Fidelity DNA Polymerase, Phusion™ Hot-Start High-Fidelity DNA Polymerase, Sulfolobus DNA Polymerase IV, Therminator™ γ DNA Polymerase, Therminator™ DNA Polymerase, Therminator™ II DNA Polymerase, Therminator™ III DNA Polymerase, VentR® DNA Polymerase, VentR® (exo-) DNA Polymerase, Bsu DNA polymerase (large fragment), Bst DNA polymerase (large fragment), DNA polymerase I (E. coli), DNA polymerase I large fragment (Klenow fragment), Klenow fragment (3'→5' exo-), phi29 DNA polymerase, T4 DNA polymerase, T7DNA polymerase (unmodified), reverse transcriptase and RNA polymerase, AMV reverse transcriptase, M-MuLV reverse transcriptase, phi6 RNA polymerase (RdRP), SP6 RNA polymerase and T7 RNA polymerase.

[0088] In some cases, once the initial extension product is generated, it can itself serve as a template for the first primer for additional amplification, generating copies of the initial extension product. These copies, along with all subsequent copies, can then be used as templates for the first primer for further amplification, generating even more copies of the initial extension product.

[0089] In many cases, the 3' end of the initial extension product may contain a sequence equivalent to that of the first primer (which may be a short-tailed primer), and its 5' end may be complementary to the template used to generate the initial extension product. The 3' ends of copies of the initial extension product and subsequent copies may contain sequences equivalent to those of the first or second primer (which may be a short-tailed primer), and their 5' ends may be complementary to either the first or second primer. The sequences between these primers may be identical to or complementary to the target polynucleotide.

[0090] In many cases, a single target polynucleotide can generate multiple initial extension products. When the first primer extension produces an initial extension product, a second first primer can bind upstream; when a polymerase with strand displacement activity extends, the first initial extension product can be replaced by the second first initial extension product; this process can continue until a primer anneals and extends at the 5' end of the transformed target polynucleotide, at which point the primer itself can no longer be replaced.

[0091] In some cases, the 3' and / or 5' ends of any initial extension product can be modified by adding adapters using a ligation-independent method. The specific oligonucleotide can be an adapter template oligonucleotide (ATO). Adapters can be generated via a 3' extension ATO reaction, using any ATO design and ATO reaction method in WO2018 / 193233 A1, or any combination thereof. Adapter generation can be replaced by or supplemented with uracil-DNA glycosidase (UDG), using an endonuclease capable of recognizing and removing uracil or other modified nucleotides, including 2'-deoxyxanthine nucleotides, such as thermostable endonuclease Q.

[0092] In many cases, the 3' extension reaction may contain one or more specific nucleotides, which may include any combination or mixture of the following: 5-methyl-2'-deoxycytidine-5'-triphosphate, 5-hydroxymethyl-2'-deoxycytidine-5'-triphosphate, 5-formyl-2'-deoxycytidine-5'-triphosphate, 5-carboxyl-2'-deoxycytidine-5'-triphosphate, 2'-deoxyuridine-5'-triphosphate, and / or 2'-deoxyinosine-5'-triphosphate. In some cases, the specific nucleotide may contain α-thiotriphosphate instead of triphosphate, thereby giving the polymerase extension product exonuclease resistance by replacing the oxygen atom with a sulfur atom.

[0093] In many cases, ATO can have the following design:

[0094] (a) The 3' end has a blocking group, which makes ATO non-extendable;

[0095] (b) The 5' side (a) contains one or any combination of the following (5' to 3' directions):

[0096] (i) A random sequence of 3 to 36 "N" bases;

[0097] (ii) A target-specific region of 4 to 10 bases, followed by a random region of 3 to 32 bases;

[0098] (c) The general sequence at the 5' side (b);

[0099] (d) One or more portions that make ATO degradable.

[0100] In one embodiment, the portion is uracil nucleotide, wherein the reagent is dU-glycosidase, or dU-glycosidase with apurinol / pyrimidine-free site endonuclease, capable of digesting / removing ATO after the first extension reaction.

[0101] In some cases, methods for extending the extended product include:

[0102] (i) The elongation product is incubated with an adapter template oligonucleotide (ATO) having:

[0103] (a) Contains a 3' portion of a random and / or specific motif sequence;

[0104] (b) The 3' end has a blocking group, making ATO non-extending; and

[0105] (c) The general sequence at the 3' portion on the 5' side.

[0106] The incubation buffer contains dNTPs in which the target polynucleotide hybridizes with the 3' portion of ATO;

[0107] (ii) Using ATO as a template, polymerase extension of the initial extension product is performed and dNTPs are incorporated to generate an initial extension product with a 3' universal sequence.

[0108] (iii) Treat with a reagent capable of digesting ATO, such as an enzyme with dU-glycosidase or 3' exonuclease activity, ribonuclease or endonuclease V activity, to digest ATO; and

[0109] (iv) Generate the initial extension product of amplification, wherein the generation of the first amplified polynucleotide includes polymerase extension of the primer from hybridization to the 3' universal sequence, using the target polynucleotide as a template.

[0110] In another embodiment, the portion is a ribonucleotide, wherein the ribonucleotide is incorporated with ATO during oligonucleotide synthesis, replacing any or all nucleotides; wherein the reagent is a ribonuclease capable of digesting / removing ATO after the first extension reaction.

[0111] In another embodiment, the portion is deoxyinosine, wherein deoxyinosine is incorporated into ATO during oligonucleotide synthesis, replacing any or all nucleotides; wherein the reagent is an enzyme capable of digesting / removing ATO after the first extension reaction. The enzyme may be an endonuclease, such as endonuclease V; or it may be a glycosidase, such as human alkyladenine DNA glycosidase.

[0112] ATOs can be RNA oligonucleotides, DNA oligonucleotides, or a combination of DNA and RNA oligonucleotides.

[0113] An ATO can be a combination of one or more different ATOs. Combined ATOs may differ in sequence; combined ATOs may differ in design; combined ATOs may differ in function. In this document, "ATO" can refer to a combination of one or more ATOs, an ATO of any sequence, an ATO of any design with any combination of ATO design features, and an ATO of any combination of functions. When using a combination of one or more ATOs, the general sequence of the ATOs used may vary; in this case, the term "general sequence" is still used.

[0114] The universal sequence of the ATO can be double-stranded or partially double-stranded. Protecting the universal sequence as a double-stranded region prevents randomization of the target polynucleotide or initial elongation product, or sequence-specific hybridization of the 3' end with the ATO. In one embodiment, the ATO includes a 5' stem sequence that is wholly or partially complementary to the universal sequence, capable of forming a stem-loop structure or a split stem-loop structure. Alternatively, the loop portion may not contain a non-replicating link. Alternatively, the stem portion may contain a non-replicating link. If the 5' stem portion contains an additional sequence, a non-replicating link may exist between the stem portion and the additional sequence. Alternatively, the stem portion may also not contain a non-replicating link. In another embodiment, the 5' stem portion contains a non-replicating link. The non-replicating link may be selected from, but is not limited to: C3 spacer phosphoramide, triethylene glycol spacer, 18-atom hexaethylene glycol spacer, or 1',2'-dideoxyribose (dSpacer).

[0115] The stem portion of the double-stranded sequence may contain non-complementary regions, where the non-complementary regions in the universal sequence strand contain random sequences, degenerate sequences, or specially designed mismatches. The stem portion may form two or more split segments separated by one or more non-replicable links. The stem portion may also form two or more split segments separated by one or more mismatched base pair regions.

[0116] ATO may further include a specific sequence at the 5' and / or 3' end of the random sequence, wherein the specific sequence is capable of hybridizing to a specific position of the initial extension product, or is a sequence not designed for a specific target, and a portion of the 3' random / degenerate sequence serves as a template for polymerase to extend the polynucleotide.

[0117] In some cases, the annealing of the 3' end of the target polynucleotide or the initial extension product to ATO can be directly linked to the 5' or 3' end of ATO without an extension reaction. In some cases, the 5' end of ATO may contain a phosphate group, and the 3' end of the initial extension product chain may contain an OH group, wherein ATO has a 3' dangling end containing a sequence that guides annealing to either the 3' or 5' end of the initial extension product. In some cases, the 3' end of ATO may contain an OH group, and the 5' end of the initial extension product chain may contain a phosphate group, wherein ATO has a 5' dangling end containing a sequence that guides annealing to either the 3' or 5' end of the initial extension product. In extension embodiments without ligation, the 5' end of the initial extension product chain does not contain a phosphate group, and the 3' end of the upper independent chain does not contain biotin, but the upper independent chain may contain degradable nucleotides, such as uracil nucleotides. Any ATO reaction may involve ligation, whether or not it is accompanied by extension; if the 3' end is extended, the polymerase extends the 3' end of the initial extension product, and the ligase ligates the extended target sequence to the 5' stem of the ATO or to the ATO strand.

[0118] In one embodiment, the first ATO reaction is a primer extension reaction, wherein the transformed target polynucleotide or the initial extension product serves as a primer, and DNA polymerase extends the ATO template. The DNA polymerase may have strand displacement activity or 5' to 3' exonuclease activity, during which the stem-loop structure opens, or the ATO strand is replaced or digested. Any polymerase can be used, such as Klenow exo-, Bst polymerase, or T4 DNA polymerase.

[0119] In another embodiment, the first reaction is an extension-ligation reaction, wherein a DNA polymerase extends the target, and a DNA ligase ligates the extended target sequence to the 5' stem portion of ATO or the upper strand of ATO. Any DNA polymerase and DNA ligase can be used, such as Klenow large fragment or T4 DNA ligase.

[0120] In another embodiment, the first reaction is a ligation reaction, in which a DNA ligase ligates the target polynucleotide to the 5' stem portion of the ATO or the upper strand of the ATO. Any DNA ligase can be used, such as T4 DNA ligase.

[0121] In some implementations, the specific oligonucleotide is a splice or bridging oligonucleotide, i.e., the linker comprises two strands: a top linker strand attached to the single-stranded target sequence, and a splice strand bridging the top linker strand and the single-stranded target sequence. The ligation reaction uses the splice or bridging oligonucleotide to guide the ligation of the single-stranded linker to the single-stranded target polynucleotide, which may be the initial extension product (containing a copy of the short primer extension product and / or the complementary strand). The splice oligonucleotide may contain a sequence at its 5' or 3' end that anneals to the 3' or 5' end of the target polynucleotide, and at the other end (3' or 5' end) a sequence at the 5' or 3' end of the linker that anneals to it. When ligation occurs at the 5' end of the target polynucleotide, its 5' end may contain a phosphate group, which may be present at the 5' end of the first primer or enzymatically added. When ligation occurs at the 3' end of the target polynucleotide, its 3' end may contain an OH group, which may be present on the initial extension product. When ligation occurs at the 5' end of the target polynucleotide, the 3' end of the linker may contain an OH group, which can be added during oligonucleotide manufacturing. When ligation occurs at the 3' end of the target polynucleotide, the 5' end of the linker may contain a phosphate group, which can be added during oligonucleotide manufacturing or enzymatically.

[0122] The method may further include: a second primer that extends and hybridizes with the initial extension product to generate a copy extension initial extension product, wherein the second primer contains a target-specific portion or a universal sequence, or contains both a 3' target-specific sequence and a 5' universal sequence.

[0123] The method may further include: generating a labeled copy of the initial extension product, wherein the 3' end of the initial extension product is annealed to the 3' sequence of the adapter template oligonucleotide (second ATO), and extending the 3' end of the initial extension product using ATO as a template in an enzymatic second ATO reaction, wherein the second ATO contains a 5' universal sequence different from the first ATO, so that each end has a unique sequence.

[0124] The method may further include: exponential amplification using a first primer and a second primer. The first primer may be a universal primer annealed to a sequence (or a copy thereof) added to the initial extension product at the 3' end; the second primer may be a universal primer annealed to a sequence (or a copy thereof) added to the initial extension product at the 3' end. Alternatively, the second primer may be a target-specific primer annealed to a specific region of the amplified product. The second primer may be a combination of multiple primers targeting multiple target sequence regions. When the second primer is a target-specific primer, after linear or exponential amplification using the second primer, further amplification is performed using a nested target-specific third primer.

[0125] The first primer may contain a sample barcode (SBC) sequence and an additional 5' universal sequence compatible with the NGS platform.

[0126] The kit contains the above-described composition.

[0127] The kit for generating a multinucleotide library contains short-tailed primers, the aforementioned adapter template oligonucleotides (ATO) or splice oligonucleotides, polymerase, and primers compatible with the NGS platform.

[0128] The method of this invention can significantly improve the accuracy and sensitivity of testing individual patient samples. This method can enrich methylated DNA molecules from a DNA template population. Brief description of the attached diagram

[0130] Figure 1 This diagram illustrates an illustrative implementation. A first primer (which may be a short-tailed primer) hybridizes with a target polynucleotide (which may be original or transformed), and is then extended by a polymerase to produce an initial extension product. The first or second primer subsequently hybridizes with and extends the initial extension product, producing a copy of the initial extension product. An adapter is then added to the 3' end of the copied initial extension product. The 3' end sequence of the extension product, corresponding to the 5' universal tail sequence of the short-tailed primer, serves as a guide to mediate the addition of the adapter sequence (either through ligation or extension using an ATO template).

[0131] Figure 2This diagram illustrates an illustrative implementation. Several designs for the first primer are shown. The first primer may be a short-tailed primer comprising a 5' universal tail portion and a 3' initiation portion. The 3' initiation portion may be a target-specific sequence, a random sequence, or a combination of a random sequence and a target-specific sequence. The first primer may contain uracil or inosine between the 5' universal tail portion and the 3' initiation portion.

[0132] Figure 3 : A schematic diagram illustrating an illustrative implementation. A) An example of guiding the addition of a linker to the initial extension product or a copy of the initial extension product using a 3' extension reaction (ATOM-Seq reaction, which can be performed by any detailed method). B) An example of guiding the addition of a linker using a splint or bridging oligonucleotide and a single linker (by linking to the 5' or 3' end of the target polynucleotide, which can be the initial extension product or a copy of the initial extension product).

[0133] Figure 4 This demonstrates the results of implementing the present invention. As described in Example 1, three different masses of starting material were used: 50 ng, 100 ng, and 200 ng.

[0134] Figure 5 This tool displays the global methylation patterns of 12 healthy samples and 5 lung cancer samples, and visualizes them as a cluster dendrogram by clustering all methylation patterns. Example

[0135] Table 1: Details of all oligonucleotides

[0136]

[0137] Example 1

[0138] Deoxyribonucleic acid (DNA) was used as the target polynucleotide, and after reagent conversion, it was used as a template to generate the initial extension product. The initial extension product was then subjected to a 3' extension ATO reaction for guide adapter addition. A copy of the initial extension product was then generated, and its 3' end was also subjected to a 3' extension ATO reaction for guide adapter addition. Finally, a sequencing library for NGS was generated through full-sample amplification.

[0139] Material

[0140] EpiTect Rapid Sulfite Kit (Qiagen)

[0141] Target polynucleotides, human gDNA

[0142] Klenow segment (NEB)

[0143] dNTP solution (NEB)

[0144] NEbuffer 2 (NEB)

[0145] Oligonucleotides: Table 1

[0146] AMPure XP magnetic beads (Beckman Coulter)

[0147] NEBNext Q5 Master Mix (NEB)

[0148] method

[0149] 1. Sulfite Conversion

[0150] Take 50 ng, 100 ng, or 200 ng of DNA and perform sulfite conversion using the EpiTect Rapid Sulfite Kit according to the manufacturer's recommended protocol. Elute the final product in 19.5 µl of molecular biology grade water.

[0151] 2. Generate the initial extension product

[0152] The transformation target polynucleotide was used as a template for targeted enrichment of CG-containing DNA. The mixture contained 17.5 µl of the transformation target polynucleotide, oligonucleotide 1-001 or 1-007, and buffer. The mixture was heated to 98°C for 2 minutes, then to 10°C for 1 minute and maintained at 10°C. The Klenow fragment was added. The following thermal cycling was then performed: 10°C for 1 minute, 26°C for 6 minutes, 30°C for 10 minutes, and 65°C for 1 minute. 28 µl of water was added to the product to increase the volume to 50 µl, and then purified using AMPure XP magnetic beads (2x volume beads) according to the manufacturer's recommendations, eluting in 18.5 µl of molecular biology grade water.

[0153] 3. Add guiding connectors to the initial elongation product.

[0154] The purified initial extension product was used for initiator adapter addition: 16.5 µl of the purified initial extension product was mixed with 2.0 µl of 1-003 and 1.5 µl of Step 2 buffer, treated according to the thermal cycling program of Step 2, then the enzyme was added and further cycles were performed according to Step 2. Then, 2.0 µl of the treatment enzyme mixture was added to the reaction mixture and incubated at 37°C for 15 min.

[0155] 4. Generate the initial copy extension product.

[0156] Combine the entire volume with 2.0 µl of oligonucleotide 1-002 and 26 µl of Q5 Master Mix. After preheating at 98°C for 30 seconds, perform 6 cycles (98°C for 10 seconds, 65°C for 75 seconds, 72°C for 2 minutes), and finally at 65°C for 2 minutes. Purify using AMPure XP magnetic beads (2 volumes) as per manufacturer's recommendations, eluting in 15 µl of molecular biology grade water.

[0157] 5. Add guiding connectors to the initial extension product of the copy.

[0158] The purified initial copy extension product was used for guide adapter addition. First, 13 µl of the purified initial copy extension product was mixed with oligonucleotide 1-006 and incubated at 65°C for 2.5 min, followed by incubation at 10°C for 1 min. Then, 0.35 µl of ldNTP, 2 µl of NEbuffer 2, and 1.0 µl of the Klenow fragment were added. Incubation was performed at the following conditions: 10°C for 1 min, 26°C for 6 min, 30°C for 10 min, 65°C for 1 min, 10°C for 1 min, 26°C for 6 min, and 30°C for 10 min. Then, 2.0 µl of the treatment enzyme mixture was added to the reaction mixture, and incubation was performed at 37°C for 15 min.

[0159] 6. Whole sample PCR

[0160] The entire product from step 5 was mixed with 25 µl of Q5 Master Mix and 1.5 µl each of oligonucleotides 1-004 and 1-005, and incubated at 98°C for 30 seconds, followed by 15 cycles (98°C for 10 seconds, 60°C for 30 seconds, 65°C for 75 seconds), with a final extension at 65°C for 2 minutes. Purification was performed using AMPure XP magnetic beads (2 volumes) as per manufacturer's recommendations, eluting in 30 µl of molecular biology grade water. The final magnetic bead purified product was visualized on a Bioanalyzer high-sensitivity DNA chip, as shown below. Figure 4 As shown.

[0161] Example 2

[0162] Deoxyribonucleic acid (DNA) was used as the target polynucleotide, which was then converted with reagents and used as a template to generate the initial extension product. The initial extension product was then guided to add an oligonucleotide ligation linker via a splint.

[0163] Material

[0164] EpiTect Rapid Sulfite Kit (Qiagen)

[0165] Target polynucleotides, human gDNA

[0166] Klenow segment (NEB)

[0167] dNTP solution (NEB)

[0168] NEbuffer 2 (NEB)

[0169] DNA ligase

[0170] Oligonucleotides: Table 1

[0171] AMPure XP magnetic beads (Beckman Coulter)

[0172] NEBNext Q5 Master Mix (NEB)

[0173] method

[0174] 1. Sulfite Conversion

[0175] Same as method 1.

[0176] 2. Generate the initial extension product

[0177] Same as method 1.

[0178] 3. Add guiding connectors to the extended product.

[0179] The purified extension product was used for guiding adapter ligation: 16.5 µl of the purified extension product was mixed with 2.0 µl of splint oligonucleotide 2-008 and single-linker oligonucleotide 2-009 (5' phosphorylated). Enzymatic ligation was performed under standard conditions. The manufacturer's 1× T4 DNA ligase buffer ([MgCl2] = 10 mM, [ATP] = 500 μM, [DTT] = 10 mM, [Tris-HCl] = 40 mM) was used directly. Ligation was carried out at 20°C for 12 hours, followed by termination at 65°C for 10 minutes.

[0180] 4. Amplification of adapter ligation products

[0181] The adapter ligation product was mixed with primers capable of hybridizing with the adapter sequence, and primer extension and amplification were performed using standard methods. The amplified product was then prepared for sequencing.

[0182] Example 3

[0183] Deoxyribonucleic acid (DNA) from clinical samples was used as the target polynucleotide. After reagent conversion, it was used as a template to generate the initial extension product. The initial extension product was then guided by a 3' extension ATO reaction to add a 3' adapter. A copy of the initial extension product was then generated, and its 3' end was also guided by a 3' extension ATO reaction to add an adapter. Finally, a sequencing library for NGS was generated through whole-sample amplification, followed by sequencing and analysis to enrich CpG dinucleotides.

[0184] Material

[0185] Same as Example 1.

[0186] method

[0187] 1. Sulfite Conversion

[0188] Take 200 ng of DNA from non-cancer or cancer samples and perform sulfite conversion as in Example 1.

[0189] 2. Generate the initial extension product

[0190] Same as Example 1.

[0191] 3. Add guiding connectors to the initial elongation product.

[0192] Same as Example 1.

[0193] 4. Generate the initial copy extension product.

[0194] Same as Example 1.

[0195] 5. Add guiding connectors to the initial extension product of the copy.

[0196] Same as Example 1.

[0197] 6. Whole sample PCR

[0198] Same as Example 1.

[0199] 7. Sequencing and Data Analysis

[0200] Sequencing data were spliced ​​using FASTP, and then aligned to the human genome using BISMark. Alignment rates were determined using SAMTools, revealing that 66.11% to 85.48% of reads aligned to the human genome. After read deduplication, 67.01% to 91.91% of reads were unique, with a total of 6.9 million to 46.3 million reads per sample. Intersection of reads with CpG island genomic locations revealed that 1.69% to 3.94% of unique reads aligned to CpG islands, resulting in CpG island coverage of 91.29% to 98.6%. The CpG dinucleotide content of the reads ranged from 1.6% to 3.9%, representing a 1.6- to 2.9-fold enrichment compared to randomly captured DNA, given that CpG accounts for approximately 1% of the human genome. Table 2 summarizes the data from 23 samples. Global methylation patterns of 12 healthy samples and 5 lung cancer samples were visualized as a clustering dendrogram by clustering all methylation patterns. Figure 5 This indicates a clear separation between healthy and cancerous samples.

[0201] Table 2: Sequencing results (showing the enriched CG content)

[0202]

[0203]

Claims

1. A method for adding an adapter sequence to a target polynucleotide in a sample, the method comprising: a) Provides a reaction mixture comprising at least one primer (short-tailed primer), the short-tailed primer comprising a 3' initiating portion and a 5' universal tail portion, wherein the 3' initiating portion is capable of hybridizing with a target polynucleotide and initiating extension, and wherein the 5' universal tail portion or its complementary strand is capable of mediating the addition of a guiding linker sequence to the primer extension product; b) Perform at least one round of extension reaction, wherein the extension reaction includes hybridization of the primer with the target polynucleotide and extension under extension conditions to produce a primer extension product, wherein the primers include short-tailed primers. as well as c) Adding adapter sequences to the extension product, mediated by hybridizing a specific oligonucleotide to the 3' and / or 5' end of the extension product, wherein the specific oligonucleotide contains a sequence that is identical to or complementary to the 5' universal tail portion of the short-tail primer.

2. The method of claim 1, wherein at least one round of extension is two or more rounds, wherein the first round produces an extension product, and the second or more rounds produce copies of the extension product, wherein the specific oligonucleotide is capable of hybridizing with the 3' end sequence of the extension product, the 3' end sequence being the complementary strand to the 5' universal tail sequence of the short-tailed primer in the extension product.

3. The method according to claim 1, wherein at least one round of extension reaction is a multi-round extension of isothermal amplification using a polymerase with chain displacement activity.

4. The method according to claim 3, wherein the polymerase is selected from Klenow, Bst polymerase, and φ29 DNA polymerase group.

5. The method according to claim 1, wherein at least one round of extension reaction is a multi-round extension of a PCR reaction.

6. The method according to claim 1 or 2, wherein the specific oligonucleotide is an adapter template oligonucleotide (ATO), and the addition of an adapter sequence to the extension product is performed by extending the 3' end of the extension product, wherein the 3' end of the extension product serves as a primer and the ATO serves as a template.

7. The method according to claim 1 or 2, wherein the specific oligonucleotide is a splice oligonucleotide capable of hybridizing with the 3' end sequence of the extension product and the linker oligonucleotide, and adding a linker to the extension product comprises: The splice oligonucleotide is hybridized to the 3' or 5' end of the extension product and the linker oligonucleotide, and then ligated between the 3' or 5' end of the extension product and the 5' or 3' end of the linker oligonucleotide.

8. The method of claim 1, wherein the 5' universal tail portion is 2 to 12 nucleotides in length.

9. The method of claim 1, wherein the 5' universal tail portion is 2 to 9 nucleotides in length.

10. The method of claim 1, wherein the 5' universal tail portion is 2 to 6 nucleotides in length.

11. The method of claim 1, wherein the 5' universal tail portion is 2 to 4 nucleotides in length.

12. The method of claim 1, wherein the short-tailed primer comprises a degradable nucleotide.

13. The method of claim 12, wherein the degradable nucleotide is uracil nucleotide.

14. The method of claim 1, wherein the 3' initiation portion comprises a target-specific sequence.

15. The method of claim 1, wherein the 3' triggering portion comprises a random sequence.

16. The method of claim 1, wherein the 3' initiation portion comprises a random sequence and a target-specific sequence.

17. The method of claim 1, wherein the target polynucleotide is sulfite-converted or enzymatically converted DNA, wherein cytosine in the original DNA has been converted to uracil.

18. The method according to claim 1, further comprising: d) Amplification is performed using primers that can hybridize with the adapter sequence.

19. A kit for preparing sequencing libraries, comprising: a) At least one primer (short-tailed primer) having a 3' initiating portion and a 5' universal tail portion, wherein the 3' initiating portion is capable of hybridizing with a target polynucleotide and initiating extension, wherein the 5' universal portion of the extension product or the 3' end of the extension product derived from the 5' universal tail portion of the primer is capable of mediating the addition of a guiding linker sequence to the extension product; and b) A linker template oligonucleotide (ATO) or splice oligonucleotide capable of hybridizing with the 3' end of the extension product, wherein the ATO or splice oligonucleotide contains a sequence that is identical or complementary to the 5' universal tail portion of the short-tail primer.

Citation Information

Patent Citations

  • Methods, compositions, and kits for preparing nucleic acid libraries

    WO2018193233A1