A method for constructing a multiplex PCR library for high-throughput targeted sequencing
By introducing MoCODE barcode and corresponding sequencing linker in multiple PCR reactions, combined with specific endonuclease digestion technology, the non-specific amplification problem in the existing technology is solved, efficient and accurate targeted enrichment library construction is achieved, and the sequencing quality and efficiency are improved.
Patent Information
- Application Number
- CN202180088322.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-12-31
- Filing Date
- 2021-12-31
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2041-12-31
AI Technical Summary
The existing PCR-based targeted enrichment library construction methods have non-specific amplification problems in targeted methylation sequencing, making it difficult to remove non-specific amplification products, affecting the sequencing quality.
Using multi-base MoCODE barcode technology, specific MoCODE barcodes were added to the multiple PCR reaction, and efficient ligation was used to build a library using a complementary sequencing linker, and the non-specific amplification products were removed through specific endonuclease digestion.
It effectively reduces the generation of non-specific products in multiple PCR amplification, improves the targeting rate and sequencing depth of sequencing data, simplifies the operation process, and reduces contamination and manual operation time.
Smart Images

Figure CN116888276B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of biomedicine, and more specifically, to a method for constructing a DNA library, particularly to a method for constructing a multiplex PCR library for high-throughput targeted sequencing. Background Technology
[0002] This disclosure relates to the field of library construction technology, specifically to a targeted high-throughput DNA library construction method. Over the past decade, with the continuous advancement of next-generation sequencing technology, its application in life science research has been expanding. Different nucleic acid preparation methods and sequencing library construction techniques have also become more efficient.
[0003] High-throughput sequencing (NGS), also known as next-generation sequencing, is a technology that enables massively parallel sequencing on high-density biochips. It boasts high data output and low cost per unit of data. However, its drawback lies in its short read lengths, typically 2x300bp or 2x150bp. Assembling these short reads without a reference genome or with genomes containing highly complex structures presents significant challenges. In such cases, large-scale metapair libraries can assist in the assembly of short reads. Furthermore, analysis of these large-scale libraries using the Link algorithm can detect structural variations in large chromosomal segments, such as insertions, deletions, inversions, and translocations.
[0004] High-throughput targeted sequencing is a cost-effective and highly sensitive detection method, with the key step being the targeted enrichment of the target gene. Currently, the main methods for achieving targeted enrichment include hybridization capture and PCR-based library construction. Generally speaking, hybridization capture methods are expensive and cumbersome due to the need for streptavidin-coated magnetic beads, and require larger DNA samples. While PCR-based targeted enrichment using unique molecular identifiers (UMI) has made significant progress compared to hybridization capture in recent years, overcoming the previous difficulty in removing PCR repetitive sequences, errors in UMI remain difficult to eliminate and the operation is cumbersome. Therefore, it is necessary to provide a precise, efficient, and simple method for constructing multiplex PCR targeted enrichment libraries.
[0005] Existing PCR-based targeted enrichment library construction methods mainly include AmpliSeq (thermo), SLIMAmplification, and Relay PCR. These methods all involve two PCR reactions: the first step is targeted amplification of the target fragment, and the second step is PCR enrichment after adapter ligation. However, these methods all use traditional TA ligation or blunt-end ligation, and the overall library construction process does not include a step to control non-specific amplification, nor can it effectively remove non-specific amplification products. This situation is particularly prominent in targeted methylation sequencing. Because bisulfite treatment of DNA converts most cytosine to thymine, primer dimers or non-specific amplification are more likely to form between multiple primers. Summary of the Invention
[0006] The purpose of this disclosure is to provide a method for constructing multiplex PCR libraries for high-throughput targeted sequencing.
[0007] To achieve the above objectives, the present disclosure employs the following technical means:
[0008] This disclosure relates to a method for constructing a multiplex PCR library for high-throughput targeted sequencing. The method involves adding a polybase MoCODE barcode to specific amplification products and using the MoCODE barcode to efficiently ligate the amplification products to sequencing adapters containing the MoCODE barcode decoding sequence. The MoCODE barcode refers to the two sticky, protruding single-stranded nucleotide sequences that form the two sticky ends of the obtained PCR product after digestion with a specific endonuclease. The MoCODE barcode decoding sequence is a nucleotide sequence complementary to the MoCODE barcode.
[0009] Preferably, the MoCODE barcode is generated by one or more of the following methods: modified nucleotides, nicking enzymes, endonucleases, chemical modifications, and photolyzable bases; preferably, the modified nucleotides include one or more of dUTP, dITP, and RNA bases.
[0010] Preferably, the MoCODE barcodes can be the same or different within the molecule.
[0011] Preferably, the MoCODE barcode is a non-random, specific barcode.
[0012] Preferably, the length of the MoCODE barcode is 2-20 nt.
[0013] Preferably, the MoCODE barcode decoding sequence and the MoCODE barcode sequence are complementary sequences with a length of 2-20 nt.
[0014] Preferably, the sequencing adapter can be artificially designed and synthesized, or matched with the sequence of the target region itself.
[0015] Preferably, the sequencing adapter can be a single adapter or a bidirectional adapter.
[0016] Preferably, the enrichment of each specific segment can be achieved through single-connector decoding, dual-connector decoding, or automatic loop decoding.
[0017] This disclosure also relates to primers for multiplex PCR for high-throughput targeted sequencing, the primers comprising MoCODE barcode generation sequences, preferably, the primer sequences comprising the sequences shown in Seq ID Nos: 1-22, 27-52, 53, 55, 57-104, 109, 111.
[0018] Accordingly, this disclosure also relates to a sequencing adapter for multiplex PCR for high-throughput targeted sequencing, the sequencing adapter comprising a MoCODE barcode decoding sequence. Preferably, the sequencing adapter further comprises one or more of a sequencing adapter for a sequencing platform and an index tag. Preferably, the sequencing adapter comprises a high-throughput sequencing universal sequence, an index tag, and the MoCODE barcode decoding sequence. The sequence of the sequencing adapter comprises the sequences shown in Seq ID Nos: 23-26, 54, 56, 105-108, 110, and 112.
[0019] This disclosure discloses a method for constructing multiplex PCR libraries for high-throughput targeted sequencing, the method comprising the following steps:
[0020] 1) Extract DNA from the sample to be tested;
[0021] 2) Perform multiplex PCR reaction, wherein each primer participating in the multiplex PCR reaction contains a specific MoCODE barcode generation sequence, preferably, the primer also contains a gene-specific sequence;
[0022] 3) Purify the PCR product obtained in step 2) using magnetic beads;
[0023] 4) Generate 5' and 3' sticky ends in the purified PCR product obtained in step 3), and generate MoCODE barcodes at the 5' and / or 3' sticky ends, respectively;
[0024] 5) Purify the PCR product containing the MoCODE barcode from step 4) using the magnetic bead method;
[0025] 6) Connect the purified PCR product containing the MoCODE barcode obtained in step 5) to the sequencing adapter, wherein the sequencing adapter contains a MoCODE barcode decoding sequence complementary to MoCODE.
[0026] 7) Purify the ligation product obtained in step 6) with magnetic beads to complete the construction of a multiplex PCR library for high-throughput targeted sequencing.
[0027] Preferably, the method of generating the MoCODE barcode in step 4) includes one or more of the following: modified nucleotides, nicking enzymes, endonucleases, chemical modifications, and photolyzable bases; preferably, the modified nucleotides include one or more of dUTP, dITP, and RNA bases; more preferably, the method of generating the MoCODE barcode is by enzymatic digestion using a specific endonuclease.
[0028] Preferably, in step 4), a MoCODE barcode is generated at each of the 5' and 3' adhesive ends, wherein the MoCODE barcodes at the 5' and 3' adhesive ends may be the same or different.
[0029] Preferably, the sequencing adapter in step 6) can be a single adapter, a bidirectional adapter, or a circular adapter.
[0030] Compared with the prior art, this disclosure has the following advantages:
[0031] (1) Reduce nonspecific products in multiplex PCR amplification
[0032] While current PCR-based targeted enrichment library construction methods incorporate UMIs (Unique Motion Analyzers), which can filter out errors during library construction and sequencing to some extent, random errors are not solely due to template fragment sequences but can also originate from the UMIs themselves. If errors occur within the UMIs, PCR repetitive sequences will be incorrectly identified as unique molecules from the UMI identifier, leading to overestimation of sequencing depth and impacting sequencing quality. Furthermore, UMIs themselves are random sequences and cannot remove non-specific amplification products, primer dimers, or more complex single- or double-stranded multimers from multiplex PCR.
[0033] By designing highly specific multiplex PCR primer sets and adding specific restriction enzyme sites and a unique sequence to each primer set, only correctly amplified PCR products can be ligated to specifically paired adapters after enzymatic digestion, thus completing the construction of sequencing libraries. Dimers and multimers generated during amplification are removed by digestion with specific endonucleases. Non-specific amplification products cannot correctly combine with decoding adapters, and the final ligation products cannot be amplified and recognized during high-throughput sequencing. The resulting sequencing data consists entirely or mostly of specific target fragments, greatly improving the target rate of sequencing data and ensuring sequencing depth.
[0034] (2) High efficiency and reduced pollution
[0035] By designing sticky-end adapters for ligation, compared to blunt-end ligation which only involves the ligase, the complementary base interaction is highlighted, and the affinity between the enzyme and the substrate is increased, resulting in a significant improvement in ligation efficiency. Compared to other companies' PCR-based targeted enrichment library construction methods that involve two PCR steps, the entire library construction process requires only one PCR reaction, reducing contamination and providing better resistance to contamination.
[0036] (3) Easy to operate and quick
[0037] By designing highly specific multiplex PCR primer sets and increasing adapter ligation efficiency, the library construction process is made more efficient. Compared with other companies' PCR-based targeted enrichment library construction methods, manual operation time is reduced by 40-50% and the overall library construction time is shortened by 30-40%. Attached Figure Description
[0038] Figure 1 The process of constructing a library using MoCODE differs from the method disclosed herein;
[0039] Figure 2 This is a schematic diagram of the upstream and downstream primer structures for multiplex PCR disclosed in this paper;
[0040] Figure 3 This is a schematic diagram of the upstream and downstream connector structure disclosed in this publication;
[0041] Figure 4A This is a schematic diagram of the MoCODE (different) double-stranded structure at both ends of the PCR product in Example 3 of this disclosure;
[0042] Figure 4B This is a schematic diagram of the upstream connector double-chain structure in Embodiment 3 of this disclosure;
[0043] Figure 4C This is a schematic diagram of the downstream connector double-chain structure in Embodiment 3 of this disclosure;
[0044] Figure 5A This is a schematic diagram of the identical double-stranded structure of the PCR product at both ends in Example 4 of this disclosure;
[0045] Figure 5B This is a schematic diagram of the upstream connector double-chain structure in Embodiment 4 of this disclosure;
[0046] Figure 5C This is a schematic diagram of the downstream connector double-chain structure in Embodiment 4 of this disclosure;
[0047] Figure 6A This is a schematic diagram of the primers used in this disclosure to generate MoCODE barcodes using the MoCODE generation sequence contained in the amplified target region itself;
[0048] Figure 6BThis is a schematic diagram of the target fragment of PCR amplification containing the MoCODE generation sequence when generating a MoCODE barcode using the MoCODE generation sequence contained in the target amplification region itself in this disclosure.
[0049] Figure 6C This is a schematic diagram of the PCR product that generates a MoCODE barcode when the MoCODE generation sequence contained in the amplified target region is used to generate a MoCODE barcode.
[0050] Figure 7 The results of agarose gel electrophoresis of the PCR amplification products in Example 1 of this disclosure are shown.
[0051] Figure 8 The agarose gel electrophoresis results of the sequencing adapter ligation product in Example 2 of this disclosure are shown. Detailed Implementation
[0052] Based on the above content of this disclosure, and in accordance with common technical knowledge and practices in the field, various other modifications, substitutions, or alterations can be made without departing from the basic technical ideas of this disclosure.
[0053] I. Definition
[0054] The term "sample" includes samples or cultures containing nucleic acids (e.g., microbial cultures), and is also intended to include biological samples and environmental samples. Samples may include samples of synthetic origin. Biological samples include whole blood, serum, plasma, cord blood, chorionic villus, amniotic fluid, cerebrospinal fluid, cerebrospinal fluid, lavage fluid (e.g., bronchoalveolar, gastric, peritoneal, ductal, ear, arthroscopic lavage fluid), biopsy samples, urine, feces, sputum, saliva, nasal mucus, prostatic fluid, semen, lymph, bile, tears, sweat, breast milk, mammary fluid, embryonic cells, and fetal cells. In a preferred embodiment, the biological sample is blood, and more preferably plasma. As used herein, the term "blood" includes whole blood or any blood fraction, such as serum and plasma as conventionally defined. Blood plasma refers to a whole blood fraction produced by centrifugation of blood treated with an anticoagulant. Blood serum refers to the fluid, watery portion remaining after a blood sample has coagulated. Environmental samples include environmental materials such as surface substances, soil, water, and industrial samples, as well as samples obtained from food and dairy processing facilities, instruments, equipment, appliances, and disposable and non-disposable items. These examples should not be construed as limiting the types of samples that can be used in this invention.
[0055] The terms “target,” “target nucleic acid,” and “target gene” are intended to refer to any molecule whose presence is to be detected or measured, or whose function, interaction, or properties are to be studied.
[0056] The terms “nucleic acid” and “nucleic acid molecule” are used interchangeably throughout this disclosure. The term refers to oligonucleotides, oligomers, polynucleotides, deoxyribonucleotides (DNA), genomic DNA, mitochondrial DNA (mtDNA), complementary DNA (cDNA), bacterial DNA, viral DNA, viral RNA, RNA, messenger RNA (mRNA), transfer RNA (tRNA), ribosomal RNA (rRNA), siRNA, catalytic RNA, clones, plasmids, M13, P1, granules, bacterial artificial chromosomes (BAC), yeast artificial chromosomes (YAC), amplified nucleic acids, amplicones, PCR products and other types of amplified nucleic acids, RNA / DNA hybrids, and polyamide nucleic acids (PNA), all of which may be in single-stranded or double-stranded form, and unless otherwise limited, will include known analogs of natural nucleotides that can function in a similar manner to naturally occurring nucleotides, and combinations and / or mixtures thereof. Therefore, the term “nucleotide” refers to naturally occurring and modified / non-naturally occurring nucleotides, including tris, dis, and monophosphate nucleotides, as well as monophosphate monomers present within polynucleotides or oligonucleotides. Nucleotides can also be ribose; 2'-deoxy; 2',3'-deoxy, and a large number of other nucleotide mimics well known in the art. Mimics include chain-terminated nucleotides such as 3'-O-methyl, halobase, or sugar substitution; substituted sugar structures, including non-sugar, alkyl ring structures; substituted bases, including inosine; denitrogenated; chi and psi, linker modified; mass-marked modified; phosphodiester modified or substituted, including thiophosphates, methylphosphonates, boranophosphates, amides, esters, ethers; and essentially or completely internucleotide substitutions, including cleavage linkages, such as optically cleavable nitrophenyl moieties.
[0057] The term "amplification reaction" refers to any in vitro method used to amplify copies of a target nucleic acid sequence. "Amplification" refers to the step of bringing a solution to conditions sufficient to allow amplification. Components of an amplification reaction may include, but are not limited to, primers, polynucleotide templates, polymerases, nucleotides, dNTPs, etc. The term "amplification" typically refers to an "exponential" increase in the number of target nucleic acids. However, as used herein, "amplification" can also refer to a linear increase in the number of selected target nucleic acid sequences, but this is distinct from a one-time, single-primer extension step.
[0058] The term "polymerase chain reaction" or "PCR" refers to a method used to amplify a specific segment or subsequence of a target double-stranded DNA in geometric progression. PCR is well known to those skilled in the art.
[0059] The term "oligonucleotide" refers to a linear oligomer of natural or modified nucleoside monomers linked by phosphodiester bonds or similar bonds. Oligonucleotides include deoxyribonucleosides, ribonucleosides, their terminal isomers, peptide nucleic acids (PNAs), etc., capable of specifically binding to target nucleic acids. Typically, monomers are linked by phosphodiester bonds or similar bonds to form oligonucleotides, which range in size from a few monomer units (e.g., 3-4) to several tens of monomer units (e.g., 40-60). Whenever an oligonucleotide is represented by a sequence of letters (such as "ATGCCTG"), it should be understood that, unless otherwise specified, the nucleotides are in a 5'-3' sequence from left to right, and "A" refers to deoxyadenosine, "C" to deoxycytidine, "G" to deoxyguanosine, "T" to deoxythymidine, and "U" to ribonucleoside, uridine. Oligonucleotides typically contain four natural deoxynucleotides; however, they may also contain ribonucleosides or non-natural nucleotide analogs. When an enzyme has specific oligonucleotide or polynucleotide substrate requirements for its activity (e.g., single-stranded DNA, RNA / DNA duplex, etc.), the selection of an appropriate composition of the oligonucleotide or polynucleotide substrate is entirely within the knowledge of a person skilled in the art.
[0060] The term "primer," or "oligonucleotide primer," refers to a polynucleotide sequence that hybridizes with a sequence on a target nucleic acid template and facilitates the detection of the oligonucleotide probe. In the amplification embodiments of the present invention, the oligonucleotide primer acts as the starting point for nucleic acid synthesis. In non-amplification embodiments, the oligonucleotide primer can be used to establish structures that can be cleaved by cleavage reagents. Primers can have various lengths and are typically less than 50 nucleotides in length. The length and sequence of primers used in PCR can be designed based on principles known to those skilled in the art.
[0061] A “mismatched nucleotide” or “mismatch” refers to a nucleotide that is not complementary to the target sequence at one or more positions. Oligonucleotide probes may have at least one mismatch, but may also have 2, 3, 4, 5, 6, or 7 or more mismatched nucleotides.
[0062] The term "specific" or "specific" in relation to the binding of one molecule to another (such as a probe for a target polynucleotide) refers to the recognition, contact, and formation of a stable complex between the two molecules, as well as a significantly reduced recognition, contact, or complex formation between the molecule and other molecules. The term "annealing," as used herein, refers to the formation of a stable complex between two molecules.
[0063] The term "cleavage reagent" refers to any tool, including but not limited to enzymes, capable of cleaving oligonucleotides to produce fragments. For methods in which amplification does not occur, the cleavage reagent may be used solely to cleave, degrade, or otherwise separate a second portion or fragment of the oligonucleotide probe. The cleavage reagent may be an enzyme. The cleavage reagent may be natural, synthetic, unmodified, or modified.
[0064] For methods involving amplification, the cleavage reagent is preferably an enzyme with both synthetic (or polymeric) and nuclease activities. Such enzymes are typically nucleic acid amplification enzymes. Examples of nucleic acid amplification enzymes are nucleic acid polymerases, such as those found in *Thermus aquaticus* (Taq) and DNA polymerase. Or E. coli DNA polymerase I. The enzyme may be naturally occurring, unmodified, or modified.
[0065] The term "nucleic acid polymerase" refers to an enzyme that catalyzes the incorporation of nucleotides into nucleic acids. Exemplary nucleic acid polymerases include DNA polymerase, RNA polymerase, terminal transferase, reverse transcriptase, telomerase, etc.
[0066] "Thermostable DNA polymerase" refers to a DNA polymerase that is stable (i.e., resistant to degradation or denaturation) and retains sufficient catalytic activity when subjected to high temperatures for a selected period of time. For example, a thermostable DNA polymerase retains sufficient activity to perform subsequent primer extension reactions when subjected to high temperatures for the time necessary for double-stranded nucleic acid denaturation. The heating conditions necessary for nucleic acid denaturation are well known in the art and are exemplified in U.S. Patent Nos. 4,683,202 and 4,683,195. Thermostable polymerases, as used herein, are generally suitable for temperature-cyclic reactions such as polymerase chain reaction ("PCR"). Examples of thermostable nucleic acid polymerases include Taq DNA polymerase, Z05 polymerase, Thermus flavus polymerase, Thermotoga maritima polymerase, such as TMA-25 and TMA-30 polymerases, Tth DNA polymerase, etc.
[0067] "Modified polymerase" refers to a polymerase in which at least one monomer differs from a reference sequence, such as a natural or wild-type form of the polymerase or another modified form of the polymerase. Exemplary modifications include monomer insertion, deletion, and substitution. Modified polymerases also include chimeric polymerases having identifiable component sequences (e.g., structural or functional domains) derived from two or more parents. The definition of modified polymerases also includes those chemically modified polymerases that incorporate a reference sequence. Examples of modified polymerases include G46E E678G CS5 DNA polymerase, G46EL329A E678G CS5 DNA polymerase, G46E L329A D640G S671F CS5 DNA polymerase, G46EL329AD640G S671F E678G CS5 DNA polymerase, G46E E678G CS6 DNA polymerase, Z05 DNA polymerase, ΔZ05 polymerase, ΔZ05-Gold polymerase, ΔZ05R polymerase, E615G Taq DNA polymerase, E678G TMA-25 polymerase, and E678G TMA-30 polymerase.
[0068] The term "5'-3' nuclease activity" or "5'-3' nuclease activity" refers to the activity of a nucleic acid polymerase, typically associated with nucleic acid chain synthesis, which involves the removal of nucleotides from the 5' end of the nucleic acid chain. For example, *E. coli* DNA polymerase I possesses this activity, while the Klenow fragment does not. Some enzymes with 5'-3' nuclease activity are 5'-3' exonucleases. Examples of such 5'-3' exonucleases include: exonucleases from *Bacillus subtilis*, phosphodiesterases from spleen, λ exonucleases, exonuclease II from yeast, exonuclease V from yeast, and exonucleases from *Neurospora crassa*.
[0069] As used in this disclosure, the terms "MoCODE barcode," "Molecular Code," and "specific molecular barcode" refer to the protruding single-stranded sequences at the two sticky ends of the PCR product obtained after digestion of multiplex PCR products with specific endonucleases.
[0070] The term “MoCODE barcode decoding sequence” or “molecular barcode decoding sequence” as used in this disclosure refers to a nucleotide sequence complementary to the “MoCODE barcode”, “Molecular Code”, or “specific molecular barcode”.
[0071] II. Implementation Method
[0072] The principle underlying this method for constructing multiplex PCR targeted enrichment libraries for high-throughput sequencing is:
[0073] 1. Introduce MoCODE (Molecular Code) into the primers for each amplification segment.
[0074] 2. The MoCODE barcodes for each pair of amplification primers can be different or the same.
[0075] Specific amplification products are selected by matching the ligation of the adapters during the later stages. The length of the MoCODE barcode can range from 2nt to 20nt or longer.
[0076] 3. Non-specific fragments cannot form an effective match with the adapter, thus failing to form the correct structure required for sequencing. They cannot be amplified in the sequencing reaction system and are therefore removed from the reaction system.
[0077] 4. The matching connection between the MoCODE barcode and the connector is an adhesive end connection. Compared with the current TA connection or flat end connection used for database construction, this method can improve the connection efficiency and the final detection sensitivity.
[0078] 5. Amplification: Gene-specific and universal amplification, along with the introduction of MoCODE barcodes, can be achieved in the same PCR reaction, shortening operation steps and manual operation time, avoiding cross-contamination during library construction, reducing costs, and improving clinical applicability.
[0079] 6. MoCODE barcodes can be used in conjunction with UMI to further improve the accuracy of mutation detection in targeted sequencing through error correction.
[0080] This disclosure discloses a method for constructing multiplex PCR libraries for high-throughput targeted sequencing, which involves adding MoCODE barcodes to specific amplification products and using sequencing adapters containing matching MoCODE barcode decoding sequences for efficient library construction.
[0081] In some embodiments of this disclosure, the sample sources of the specific amplification products include, but are not limited to, genomic DNA, cell-free DNA, cell-free cells, and cDNA generated by reverse transcription of RNA samples.
[0082] In some embodiments of this disclosure, the template DNA for the multiplex PCR reaction may be DNA, bisulfite-converted DNA, cDNA, etc.
[0083] In some embodiments of this disclosure, the template DNA for the multiplex PCR reaction can be extracted using methods such as column extraction, magnetic bead extraction, and phenol-chloroform extraction followed by ethanol or isopropanol precipitation.
[0084] In some embodiments of this disclosure, the primers participating in the multiplex PCR reaction contain a specific MoCODE barcode generation sequence, and preferably, the primers also contain a gene-specific sequence;
[0085] In some embodiments of this disclosure, the generation of the MoCODE barcode includes: modified nucleotides (dUTP, dITP, RNA base), nicking enzymes, endonucleases, chemical modifications, photolyzable bases, etc. The purpose is to create identifiable cleavage sites at the ends of the PCR product, thereby cleaving out sticky ends containing the MoCODE barcode.
[0086] In a specific embodiment of this disclosure, the MoCODE barcode is generated by including a gene-specific sequence in the primers for multiplex PCR, and at the 5' end a recognition site for a specific endonuclease that is universal between primers. The purified PCR product is then digested with one or two specific endonucleases. The enzyme-digested PCR product will contain two sticky ends. The protruding single-stranded sequence at each sticky end forms a specific molecular barcode, namely the Molecular CODE (MoCODE) barcode.
[0087] In some embodiments of this disclosure, the primer sequences comprise the sequences shown in Seq ID Nos: 1-22, 27-52, 53, 55, 57-104, 109, 111, where n represents a nucleotide dITP or dUTP.
[0088] In a specific embodiment of this disclosure, the MoCODE barcode is generated by including a dITP site in each primer of a multiplex PCR reaction, in addition to a gene-specific sequence. This site, after being recognized by a specific enzyme digestion, forms a sticky 6-base end, thus generating the MoCODE barcode sequence.
[0089] In some embodiments of this disclosure, the MoCODE barcodes may be the same or different within the molecule. For example, "the same" means that the MoCODE barcodes at both ends of the same PCR product molecule are formed by being recognized and cut by one restriction enzyme, and "different" means that the MoCODE barcodes at both ends of the same PCR product molecule are formed by being recognized and cut by two different restriction enzymes.
[0090] In some embodiments of this disclosure, the same nucleotide molecule contains a MoCODE barcode, for example, the same MoCODE barcode generated at the 5' and 3' sticky ends of a PCR product molecule.
[0091] In some embodiments of this disclosure, the same nucleotide molecule contains two MoCODE barcodes, for example, the MoCODE barcodes generated at the 5' and 3' sticky ends of a PCR product molecule are different.
[0092] In some embodiments of this disclosure, the MoCODE barcode is a non-random, specific barcode.
[0093] In some embodiments of this disclosure, the length of the MoCODE barcode is 2-20 nt.
[0094] In some embodiments of this disclosure, the MoCODE barcode sequence includes the sequence shown in Seq ID No: 53, 59, 109, 111.
[0095] In some embodiments of this disclosure, the MoCODE barcode decoding sequence is a complementary sequence to the MoCODE barcode sequence, with a length of 2-20 nt.
[0096] In some embodiments of this disclosure, the MoCODE barcode decoding sequence includes the sequence shown in Seq ID No: 54, 56, 110, 112.
[0097] In some embodiments of this disclosure, the sequencing adapter containing the MoCODE barcode decoding sequence may be artificially designed and synthesized, or may be matched with the sequence of the target region itself.
[0098] The sequencing adapter containing the MoCODE barcode decoding sequence can be, for example, a sequence matching the target region itself. This means that if the target region for PCR amplification contains its own MoCODE generation sequence, and this inherent MoCODE generation sequence will be used to generate a 5' MoCODE barcode, then the 5' primer for PCR does not need to contain the MoCODE generation sequence; if the target region for amplification contains MoCODE that will be used to generate a 3' MoCODE barcode, then the 3' primer for PCR does not need to contain the MoCODE generation sequence. Figure 6A ).
[0099] In some embodiments of this disclosure, the sequencing adapter comprises sequences as shown in Seq ID Nos. 23-26, 105-108, where “nnnnnnnn”, [i5], or [i7] represents an index tag, such as an 8-nt Illumina Index tag sequence. As is known in the art, the 5' end for the cohesive link may be phosphorylated.
[0100] In some embodiments of this disclosure, the 5th position of the primer sequence Seq ID No: 57-104, where "n" or "I" is "dITP".
[0101] In some embodiments of this disclosure, the target fragment for PCR amplification may contain one or two self-generating MoCODE sequences. Figure 6B Accordingly, the MoCODE generation sequence can be used to generate MoCODE bars at one or both ends of a DNA molecule. Digestion with a nuclease corresponding to the MoCODE generation sequence can generate the corresponding MoCODE bar at one or both ends of the PCR product. Figure 6C ).
[0102] In some embodiments of this disclosure, the sequencing adapter containing the MoCODE barcode decoding sequence can be a single adapter or a bidirectional adapter, and enrichment of each specific region can be achieved through single adapter decoding, bidirectional adapter decoding, or automatic circularization decoding. The use of a "single adapter" occurs when the MoCODE barcodes at both ends of the PCR product are "the same"; the use of a "bidirectional adapter" occurs when the barcodes at both ends of the PCR product are "different". Understandably, when using different adapters, the adapters on both sides of the nonspecific product are the same, which cannot form the correct test product and is thus eliminated in the sequencing process.
[0103] In some embodiments of this disclosure, the "circularization" can use various different MoCODE barcodes, with a structure of MoCODE + common sequence binding to sequencing primers + gene-specific sequence. The circularization decoding steps are: PCR, digestion, circularization, exonuclease digestion, and add-on PCR (adding complete sequencing primer binding sites + library index + sequence adapter), which can be used to form various amplicones.
[0104] In some embodiments of this disclosure, the sequencing adapter containing the MoCODE barcode decoding sequence includes an upstream sequencing adapter and a downstream sequencing adapter, wherein the upstream sequencing adapter contains a MoCODE barcode decoding sequence complementary to the MoCODE barcode at the 5' end of the digested PCR product, and the downstream sequencing adapter contains a MoCODE barcode decoding sequence complementary to the MoCODE barcode at the 3' end of the digested PCR product.
[0105] Furthermore, the upstream and downstream sequencing adapters each comprise an upper adapter strand and a lower adapter strand, wherein the upper adapter strand is a sense strand and the lower adapter strand is an antisense strand. The MoCODE barcode decoding sequence can be located at the 3' end of the upper adapter strand of the upstream sequencing adapter or at the 5' end of the lower adapter strand of the upstream sequencing adapter, or it can be located at the 5' end of the upper adapter strand of the downstream sequencing adapter or at the 3' end of the lower adapter strand of the downstream sequencing adapter. Figure 3).
[0106] In some embodiments of this disclosure, multiple amplification of 2-1000 destination segments can be achieved. Each destination segment may have its own unique barcode, or multiple destination segments may share the same barcode.
[0107] In some embodiments of this disclosure, the MoCODE barcode is a non-random, specific barcode and can also be used for multi-purpose cancatmerization.
[0108] In some embodiments of this disclosure, the DNA polymerase used in the multiplex PCR can be a commercially available enzyme such as Taq polymerase, PFx, KOD, Pfu, Q5, Bst, or Phusion.
[0109] In some embodiments of this disclosure, the ligase used for the multiplex PCR may be T4 DNA ligase, 9NTM DNA ligase, Taq DNA ligase, Tth DNA ligase, Tfi DNA ligase, AmpligaseR, etc.
[0110] In some embodiments of this disclosure, excess removal of the sequencing adapter can be achieved using methods such as magnetic beading, column extraction, ethanol precipitation, agarose or polyacrylamide gel recovery.
[0111] In some embodiments of this disclosure, the constructed libraries are applicable to high-throughput sequencing platforms such as Illumina, Roche, ThermoFisher, Pacific Biosciences, BGI Genomics, Oxford Nanopore Technologies, Huayinkang, and HanHai Genomics.
[0112] Specifically, in some embodiments of this disclosure, the method for constructing a multiplex PCR library for high-throughput targeted sequencing includes the following steps (exemplary library construction process as follows): Figure 1 As shown):
[0113] Step 1: Prepare the sample to be tested and extract DNA. If it is a methylated sequencing library construction, bisulfite conversion is required afterwards.
[0114] Step 2: Using the DNA sample obtained in Step 1 as a template, use a high-fidelity PCR enzyme and multiple pairs of primers ( Figure 2 Multiplex PCR reactions are performed; each primer pair involved in the multiplex PCR reaction contains a gene-specific sequence and a specific molecular barcode generation sequence that is universal between primers at its 5' end.
[0115] Step 3: Purify the PCR product from Step 2 using magnetic beads;
[0116] Step 4: Digest the purified product from Step 3 using a specific endonuclease. The 3' and 5' ends of a correctly amplified multiplex PCR product should contain a specific barcode generation site. Digestion with a specific endonuclease will form sticky ends, generating the MoCODE barcode sequence, which mediates the ligation in Step 5. There are various ways to generate the barcode, including: modified nucleotides, dUTP, dITP, RNA base, cleavage enzymes, endonucleases, chemical modifications, and photolyzable bases, etc.
[0117] Step 5: Purify the enzyme digestion products from Step 4 using magnetic beads;
[0118] Step Six: For the purified enzyme digestion product obtained in Step Five, introduce upstream and downstream sequencing adapters using a ligase that catalyzes the connection between sticky ends. The introduced upstream sequencing adapter contains a universal high-throughput sequencing sequence (which may include an index tag sequence) and a MoCODE barcode decoding sequence complementary to the 5' MOCODE of the digested PCR product obtained in Step Four. The introduced downstream sequencing adapter contains a universal high-throughput sequencing sequence (including an index tag sequence) and a MoCODE barcode decoding sequence complementary to the 3' MOCODE of the digested PCR product obtained in Step Four. Figure 3 );
[0119] Step 7: Purify the ligation product from Step 6 using magnetic beads and complete the construction of the sequencing library.
[0120] III. Examples
[0121] The present invention will be further described below with reference to specific examples, and the advantages and features of the present invention will become clearer as a result. However, these examples are merely exemplary and do not constitute any limitation on the scope of the present invention. Those skilled in the art should understand that modifications or substitutions can be made to the details and form of the technical solutions of the present invention without departing from the spirit and scope of the present invention, but all such modifications and substitutions fall within the protection scope of the present invention.
[0122] Example 1: Enrichment and elimination of non-specific PCR products using MoCODE-targeted methylation multiplex PCR
[0123] In this embodiment, two sets of 10 pairs of bisulfite sequencing primers (BSPs) were designed, with each primer in both sets containing the same gene-specific sequence. In the experimental group, each BSP pair, in addition to the gene-specific sequence, included a MoCODE barcode generation sequence common to both primers at its 5' end; the control group's BSP pairs contained only the gene-specific sequence and did not include the MoCODE barcode generation sequence at their 5' ends. The two MoCODE barcode sequences were generated by digesting the PCR product with two restriction endonucleases. The enrichment effect of the products from both groups was then observed by agarose gel electrophoresis.
[0124] 1) PCR template preparation
[0125] a) HeLa cell genomic DNA (NEB, USA) was converted to bisulfite using the EZ DNA Methylation-Gold Kit (ZYMO, USA).
[0126] b) Measure the concentration of the obtained transformed DNA using a Qubit fluorometer.
[0127] c) Adjust the concentration of bisulfite-converted DNA to 50 ng / μl with water.
[0128] 2) Multiplex PCR
[0129] a) PCR reaction system
[0130]
[0131] b) PCR procedure
[0132] Step 1: 94℃, 2 minutes.
[0133] Step 2: 6 cycles (98℃, 10 seconds; 59℃, 5 seconds; 68℃, 5 seconds).
[0134] Step 3: 35 cycles (98℃, 10 seconds; 68℃, 10 seconds).
[0135] Step 4: 68℃, 1 minute.
[0136] Step 5: Keep it at 8℃.
[0137] 3) Purify multiplex PCR products using HiPrep PCR magnetic beads (MAGBIO, USA).
[0138] a) Purify the PCR product using 60 μl of magnetic beads (1.2 times).
[0139] b) The purified product was eluted in 15 μl of water.
[0140] c) Measure the concentration of purified PCR products using a Qubit fluorometer.
[0141] d) Adjust the concentration of the product to 10 ng / μl with water.
[0142] 4) Treat the purified PCR product with restriction endonucleases Bbvl and Earl (the structural diagram of the generated product is shown in the figure). Figure 5A (As shown)
[0143] Components volume 10x Cutsmart Buffer (NEB) 2μl BbvI (NEB, 2U / μl) 1μl EarI (NEB, 20 U / μl) 0.5μl Purification of PCR products 5μl 50ng Nuclease-free water 11.5μl Total volume 20μl
[0144] Incubate at 37°C for 30 minutes on a heat circulator.
[0145] Incubate at 65°C for 20 minutes to deactivate the enzyme.
[0146] The reaction mixture was purified using HiPrep PCR magnetic beads (1.2x) and eluted in 15 μl of water.
[0147] 5) Agarose gel electrophoresis
[0148] a) Prepare a 2% agarose gel using 0.5×TBE and add nucleic acid dye (GelSafe) (1 μl of dye per 10 ml system).
[0149] b) Add 5 μL of the purified PCR product and treat it with restriction endonuclease.
[0150] c) Electrophoresis at 150V for 30 minutes, followed by photographic observation using a gel imaging system.
[0151] 6) Agarose gel electrophoresis results
[0152] In the experimental group, the PCR amplification products of 10 primer pairs showed clear bands with no primer dimers; in the control group, the PCR products were diffuse bands with obvious primer dimers. Figure 7 ).
[0153] 7) PCR primer sequences used in this embodiment
[0154] The following is an example of the universal specific molecular barcode generation sequences for the upstream and downstream primers, which are Seq ID Nos. 1 and 12, respectively. The upstream primer sequence for Moko1-10 is Seq ID No. 2-11, and the downstream primer sequence for Moko1-10 is Seq ID No. 13-22.
[0155]
[0156] Example 2: Ligation of sequencing adapters after targeted methylation multiplex PCR enrichment using MoCODE
[0157] In this embodiment, sequencing adapters were ligated into the PCR products purified by restriction endonuclease treatment in the experimental group of Example 1. The sequencing adapter ligation effect was then observed by agarose gel electrophoresis.
[0158] 1) Connector connection (connector structure diagram as shown) Figure 5B -C (as shown)
[0159] a) Joint preparation
[0160]
[0161] Incubate at 82°C in a heat circulator for 2 minutes.
[0162] Cool to 25°C at a rate of 0.1°C / 3 seconds.
[0163] Annealing procedure: 82℃, 2 minutes; 570 x {82℃, 3 seconds, -0.1℃ / cycle}; 4℃ hold.
[0164] b) Connection reaction
[0165] Components capacity 10x T4 DNA ligase buffer (NEB) 2μl Purified enzyme digestion PCR products 15μl Upstream connector (10μM) 1μl Downstream connector (10μM) 1μl T4 DNA ligase (NEB, 200 U / μl) 1μl Total volume 20μl
[0166] The reaction mixture was gently mixed up and down with a pipette and then briefly centrifuged.
[0167] Incubate at room temperature for 15 minutes.
[0168] 2) Agarose gel electrophoresis
[0169] a) Prepare a 2% agarose gel using 0.5×TBE and add nucleic acid dye (GelSafe) (1 μl of dye per 10 ml system).
[0170] b) Add 5 μl of the purified PCR product and treat it with restriction endonuclease.
[0171] c) Electrophoresis at 150V for 30 minutes, followed by photographic observation using a gel imaging system.
[0172] 3) Agarose gel electrophoresis results
[0173] The electrophoresis results clearly show that the size of the products that underwent sequencing adapter ligation increased by approximately 100 bp, indicating successful adapter ligation. Figure 8 ).
[0174] 4) Connector sequence used in this embodiment
[0175]
[0176] [i5] / [i7] represents an 8nt Illumina Index tag sequence.
[0177] Example 3: Method 1 for constructing NGS libraries using MoCODE
[0178] In this embodiment, two different adapters were used for library construction. Two MoCODE barcode sequences were generated by digesting the PCR product with two restriction endonucleases.
[0179] 1) PCR template preparation
[0180] a) HeLa cell genomic DNA (NEB, USA) was processed using the EZ DNA Methylation-Gold Kit (USA).
[0181] ZYMO Corporation conducts bisulfite conversion.
[0182] b) Measure the concentration of the obtained transformed DNA using a Qubit fluorometer.
[0183] c) Adjust the concentration of bisulfite-converted DNA to 50 ng / μl with water.
[0184] 2) Multiplex PCR
[0185] a) PCR reaction system.
[0186]
[0187] b) PCR procedure
[0188] Step 1: 94℃, 2 minutes.
[0189] Step 2: 6 cycles (98℃, 10 seconds; 59℃, 5 seconds; 68℃, 5 seconds).
[0190] Step 3: 35 cycles (98℃, 10 seconds; 68℃, 10 seconds).
[0191] Step 4: 68℃, 1 minute.
[0192] Step 5: Keep it at 8℃.
[0193] 3) Purify multiplex PCR products using HiPrep PCR magnetic beads (MAGBIO, USA).
[0194] a) Purify the PCR product using 60 μl of magnetic beads (1.2 times).
[0195] b) The purified product was eluted in 15 μl of water.
[0196] c) Measure the concentration of purified PCR products using a Qubit fluorometer.
[0197] d) Adjust the concentration of the product to 10 ng / μl with water.
[0198] 4) Treat the purified PCR product with restriction endonucleases Bbvl and Earl (the structural diagram of the generated product is shown in the figure). Figure 4A (As shown)
[0199] Components volume 10x Cutsmart Buffer (NEB) 2μl BbvI (NEB, 2U / μl) 1μl EarI (NEB, 20 U / μl) 0.5μl Purification of PCR products 5μl 50ng Nuclease-free water 11.5μl Total volume 20μl
[0200] Incubate at 37°C for 30 minutes on a heat circulator.
[0201] Incubate at 65°C for 20 minutes to deactivate the enzyme.
[0202] The reaction mixture was purified using HiPrep PCR magnetic beads (1.2x) and eluted in 15 μl of water.
[0203] 5) Connector connection (connector structure diagram as shown) Figure 4B -C (as shown)
[0204] a) Joint preparation
[0205]
[0206] Incubate at 82°C in a heat circulator for 2 minutes.
[0207] Cool to 25°C at a rate of 0.1°C / 3 seconds.
[0208] Annealing procedure: 82℃, 2 minutes; 570 x {82℃, 3 seconds, -0.1℃ / cycle}; 4℃ hold.
[0209] b) Connection reaction
[0210]
[0211]
[0212] The reaction mixture was gently mixed up and down with a pipette and then briefly centrifuged.
[0213] Incubate at room temperature for 15 minutes.
[0214] The ligation mixture was purified using HiPrep PCR magnetic beads (1x) and eluted in 10 μl of water.
[0215] 6) Measure library concentration
[0216] Take 1 μl of the purified ligation product and prepare a series of 10-fold dilutions (1:10 to 1:10,000).
[0217] The concentration of the 1:10,000 dilution was determined using the Kapa Library Quantitative Reagent Kit.
[0218] Adjust the concentration of the library to 4 nM with water.
[0219] Sequencing was performed on the Illumina sequencing platform.
[0220] 7) Sequencing results
[0221] The raw .fastq files from Illumina paired-end sequencing were assembled into complete sequencing regions using PEAR software. Each assembled sequencing result was compared with the target region sequence. Sequences that met the expected read length generated by correctly paired primers were identified as on-target sequences. The on-target rate was the percentage of on-target sequences in the total number of reads.
[0222] Total reads: 554,265; hit rate: 97.0%.
[0223] 8) PCR primer sequences used in this embodiment
[0224] As shown below, the upstream and downstream universal specific molecular barcode generation sequences and the upstream and downstream primers described in Moko1-10 are the same as in Example 1. The upstream primer sequences of Moko11-23 are Seq ID Nos: 27, 29, 31, 33, 35, 37, 39, 41, 43, 45, 47, 49, 51, and the downstream primer sequences of Moko11-23 are Seq ID Nos: 28, 30, 32, 34, 36, 38, 40, 42, 44, 46, 48, 50, 52.
[0225]
[0226]
[0227] The underlined part indicates the specific target gene sequence.
[0228] 9) Connector sequence used in this embodiment
[0229] The connector sequence used below is the same as that used in Example 2 (Seq ID No: 23-26).
[0230]
[0231] [i5] / [i7] represents an 8nt Illumina Index tag sequence.
[0232] 10) The MoCODE barcode sequence and MoCODE barcode decoding sequence used in this embodiment
[0233] MoCODE barcode sequence (5'>3') MoCODE barcode decoding sequence (5'>3') upstream connector TGTA (Seq ID No: 53) TACA (Seq ID No: 54) Downstream connector GAT (Seq ID No: 55) ATC (Seq ID No: 56)
[0234] Example 4: Method 2 for constructing NGS libraries using MoCODE
[0235] In this embodiment, two different adapters were used for library construction. The two MoCODE barcode sequences were generated by digesting the PCR product with a single endonuclease.
[0236] 1) PCR template preparation
[0237] a) Take 1-1.5 ml of the TCT / LCT (Thin-Cytologic Test / Liquid-based cytologic test) cell preservation solution, centrifuge and remove the supernatant, then add 200 ml of PBS to resuspend, and extract DNA using the DNeasy Blood & Tissue Kit (QIAGEN, Germany).
[0238] b) Measure the concentration of the obtained DNA using a Qubit fluorometer.
[0239] c) The obtained DNA was subjected to bisulfite conversion using the EZ DNA Methylation-Gold Kit (ZYMO, USA).
[0240] e) Measure the concentration of the obtained transformed DNA using a Qubit fluorometer.
[0241] d) Adjust the concentration of bisulfite-converted DNA to 10 ng / μl with water.
[0242] 2) Multiplex PCR
[0243] a) PCR reaction system
[0244] Components volume Nuclease-free water 17.5μl 2x KOD-Multi Epi PCR Premix (TOYOBO) 25μl Primer mixture (10 μM) 1.5μl sulfite-treated genomic DNA 5μl (50ng) KOD-Multi&Ep (TOYOBO) 1μl Total volume 50μl
[0245] b) PCR procedure:
[0246] Step 1: 94℃, 2 minutes;
[0247] Step 2: 6 cycles (98℃, 10 seconds; 59℃, 5 seconds; 68℃, 5 seconds);
[0248] Step 3: 35 cycles (98℃, 10 seconds; 64℃, 5 seconds; 68℃, 5 seconds);
[0249] Step 4: 68℃, 1 minute;
[0250] Step 5: Keep it at 8℃.
[0251] 3) Purify multiplex PCR products using AMPure XP magnetic beads (Beckman Coulter, USA).
[0252] a) Purify the PCR product using 75 μl of magnetic beads (1.5 times).
[0253] b) The purified product was eluted in 15 μl of water.
[0254] c) Measure the concentration of purified PCR products using a Qubit fluorometer.
[0255] d) Adjust the concentration of the product to 20 ng / μl with water.
[0256] 4) Treat the purified PCR product with endonuclease V (NEB, USA) (the structural diagram of the generated product is shown in Figure 1). Figure 5A (As shown)
[0257]
[0258]
[0259] Incubate at 37°C for 30 minutes on a heat circulator.
[0260] Incubate at 65°C for 20 minutes to deactivate the enzyme.
[0261] The reaction mixture was purified using AMPure XP magnetic beads (1.5x) and eluted in 13 μl of water.
[0262] 5) Connector connection
[0263] a) Joint preparation (see schematic diagram of joint structure as follows) Figure 5B -C (as shown)
[0264]
[0265] Incubate at 82°C in a heat circulator for 2 minutes.
[0266] Cool to 25°C at a rate of 0.1°C / 3 seconds.
[0267] Annealing procedure: 82℃, 2 minutes; 570 x {82℃, 3 seconds, -0.1℃ / cycle}; 4℃ hold.
[0268] b) Connection reaction
[0269] Components capacity 10x T4 DNA ligase buffer (NEB) 2μl Purified enzyme digestion PCR products 13μl Upstream connector (10μM) 2μl Downstream connector (10μM) 2μl T4 DNA ligase (NEB, 200 U / μl) 1μl Total volume 20μl
[0270] The reaction mixture was gently mixed up and down with a pipette and then briefly centrifuged.
[0271] Incubate at room temperature for 15 minutes.
[0272] The ligation mixture was purified using AMPure XP magnetic beads (1.2x) and eluted in 10 μl of water.
[0273] 6) Measure library concentration
[0274] a) Take 1 μl of the purified ligation product and prepare a series of 10-fold dilutions (1:10 to 1:10,000).
[0275] b) The concentration of the 1:10,000 dilution was determined using the Kapa Library Quantitative Reagent Kit.
[0276] c) Adjust the concentration of the library to 4 nM with water.
[0277] d) Sequencing was performed on the Illumina sequencing platform.
[0278] 7) Sequencing results
[0279] The raw .fastq files from Illumina paired-end sequencing were assembled into complete sequencing regions using PEAR software. Each assembled sequencing result was compared with the target region sequence. Sequences that met the expected read length generated by correctly paired primers were identified as on-target sequences. The on-target rate was the percentage of on-target sequences in the total number of reads.
[0280] Sample 1 Sample 2 Total reads 1225399 1143004 Hit rate 98.0% 98.2%
[0281] 8) PCR primer sequences used in this embodiment
[0282] As shown below, from left to right and from top to bottom, the Seq ID No is 57-104.
[0283]
[0284]
[0285]
[0286] I:dITP
[0287] The underlined sequence fragment is a specific target gene sequence.
[0288] 9) Connector sequence used in this embodiment
[0289] As shown below, their Seq ID Nos are 105-108 in sequence.
[0290]
[0291] [i5] / [i7] represents an 8nt Illumina Index tag sequence.
[0292] 10) The MoCODE barcode sequence and MoCODE barcode decoding sequence used in this embodiment are as follows, which are Seq ID No: 109-112 in sequence.
[0293] MoCODE barcode sequence (5'>3') MoCODE barcode decoding sequence 5'>3') upstream connector CACAT (Seq ID No: 109) ATGTG (Seq ID No: 110) Downstream connector CGGAA (Seq ID No: 111) TTCCG (Seq ID No: 112) sequence list <110> Beijing Meikang Gene Science Co., Ltd.; Nanjing Dongke Zhisheng Gene Technology Co., Ltd. <120> A method for constructing multiplex PCR libraries for high-throughput targeted sequencing <130> MTPICN20259 <150> 2020116282342 <151> 2020-12-31 <150> PCT / CN2021 / 143948 <151> 2021-12-31 <160> 112 <170> SIPOSequenceListing 1.0 <210> 1 <211> 33 <212> DNA <213> (Artificial Sequence) <220> <223> <400> 1 agatcggcag cgtcagatgt gtataagaga cag 33 <210> 2 <211> 55 <212> DNA <213> (Artificial Sequence) <220> <223> <400> 2 agatcggcag cgtcagatgt gtataagaga caggagtagt tgggattata ggtgt 55 <210> 3 <211> 57 <212> DNA <213> (Artificial Sequence) <220> <223> <400> 3 agatcggcag cgtcagatgt gtataagaga cagttagaaa tttagttgta gaggggg 57 <210> 4 <211> 56 <212> DNA <213>(Artificial Sequence) <220> <223> <400> 4 agatcggcag cgtcagatgt gtataagaga caggaggtta gggttttaga ttggga 56 <210> 5 <211> 55 <212> DNA <213>(Artificial Sequence) <220> <223> <400> 5 agatcggcag cgtcagatgt gtataagaga caggtaayga attggtagag tttta 55 <210> 6 <211> 57 <212> DNA <213>(Artificial Sequence) <220> <223> <400> 6 agatcggcag cgtcagatgt gtataagaga cagtgagggt aagaattatt tagaggt 57 <210> 7 <211> 58 <212> DNA <213>(Artificial Sequence) <220> <223> <400> 7 agatcggcag cgtcagatgt gtataagaga cagagggtta aagaagagaa tgatttat 58 <210> 8 <211> 61 <212> DNA <213>(Artificial Sequence) <220> <223> <400> 8 agatcggcag cgtcagatgt gtataagaga caggagggtt gaatattaaa aatagtaggg 60 t 61 <210> 9 <211> 62 <212> DNA <213>(Artificial Sequence) <220> <223> <400> 9 agatcggcag cgtcagatgt gtataagaga cagggataat tataagaatt gtaaaggagg 60 at 62 <210> 10 <211> 58 <212> DNA <213>(Artificial Sequence) <220> <223> <400> 10 agatcggcag cgtcagatgt gtataagaga cagggtagtt ggaaatggta aatttgag 58 <210> 11 <211> 58 <212> DNA <213>(Artificial Sequence) <220> <223> <400> 11 agatcggcag cgtcagatgt gtataagaga caggagttat gttatgggag taagtggg 58 <210> 12 <211> 18 <212> DNA <213>(Artificial Sequence) <220> <223> <400> 12 agatcgctct tccgatct 18 <210> 13 <211> 45 <212> DNA <213>(Artificial Sequence) <220> <223> <400> 13 agatcgctct tccgatctat atatatcaaa cactrgactt aaaat 45 <210> 14 <211> 42 <212> DNA <213>(Artificial Sequence) <220> <223> <400> 14 agatcgctct tccgatctcc ttaaaacaaa cttatcttct cc 42 <210> 15 <211> 46 <212> DNA <213>(Artificial Sequence) <220> <223> <400> 15 agatcgctct tccgatctca ccttaacaaa taaaataata attcac 46 <210> 16 <211> 42 <212> DNA <213>(Artificial Sequence) <220> <223> <400> 16 agatcgctct tccgatctta tactaactcc cttcaaccat ta 42 <210> 17 <211> 41 <212> DNA <213>(Artificial Sequence) <220> <223> <400> 17 agatcgctct tccgatctct acccacacct accaaaccta a 41 <210> 18 <211> 44 <212> DNA <213>(Artificial Sequence) <220> <223> <400> 18 agatcgctct tccgatctat caaaaataat tctaaaaata taca 44 <210> 19 <211> 48 <212> DNA <213>(Artificial Sequence) <220> <223> <400> 19 agatcgctct tccgatctac caacttctat ataactaata aatacaca 48 <210> 20 <211> 43 <212> DNA <213>(Artificial Sequence) <220> <223> <400> 20 agatcgctct tccgatctaa aattcacttc taaatttaaa cca 43 <210> 21 <211> 48 <212> DNA <213> (Artificial Sequence) <220> <223> <400> 21 agatcgctct tccgatctaa aataatcttc atcaaattaa taaaaaca 48 <210> 22 <211> 42 <212> DNA <213> (Artificial Sequence) <220> <223> <400> 22 agatcgctct tccgatctac accaaaaaca atttaataaa ca 42 <210> 23 <211> 56 <212> DNA <213> (Artificial Sequence) <220> <221> modified_base <222> (29)..(36) <223> 8nt Illumina Index <400> 23 aatgatacgg cgaccaccga gatctacacn nnnnnnntcg tcggcagcgt cagatg 56 <210> 24 <211> 23 <212> DNA <213> (Artificial Sequence) <220> <223> <400> 24 tacacatctg acgctgccga cga 23 <210> 25 <211> 64 <212> DNA <213>(Artificial Sequence) <220> <221> modified_base <222> (33)..(40) <223> 8nt Illumina Index <400> 25 atcggaagag cacacgtctg aactccagtc acnnnnnnnn atctcgtatg ccgtcttctg 60 cttg 64 <210> 26 <211> 29 <212> DNA <213>(Artificial Sequence) <220> <223> <400> 26 gtgactggag ttcagacgtg tgctcttcc 29 <210> 27 <211> 54 <212> DNA <213>(Artificial Sequence) <220> <223> <400> 27 agatcggcag cgtcagatgt gtataagaga cagttagggt tttagattgg gagg 54 <210> 28 <211> 44 <212> DNA <213>(Artificial Sequence) <220> <223> <400> 28 agatcgctct tccgatctt taccaaac tatactaac aact 44 <210> 29 <211> 58 <212> DNA <213> (Artificial Sequence) <220> <223> <400> 29 agatcggcag cgtcagatgt gtataagaga caggttaggg aagttgatgt class 58 <210> 30 <211> 41 <212> DNA <213> (Artificial Sequence) <220> <223> <400> 30 agatcgctct tccgatcta tcaatctc taaaccaaaa a 41 <210> 31 <211> 62 <212> DNA <213> (Artificial Sequence) <220> <223> <400> 31 agatcggcag cgtcagatgt gtataagaga cagtagttat atggaagtt gagatagaag 60 ga 62 <210> 32 <211> 47 <212> DNA <213> (Artificial Sequence) <220> <223> <400> 32 agatcgctct tccgatctaa taaatttaca taaaaaa 47 <210> 33 <211> 62 <212> DNA <213>(Artificial Sequence) <220> <223> <400> 33 agatcggcag cgtcagatgt gtataagaga cagaagaata atttaatagg attggaagga 60 at 62 <210> 34 <211> 47 <212> DNA <213>(Artificial Sequence) <220> <223> <400> 34 agatcgctct tccgatctac ataataaaac cctatctcta ctaaaaa 47 <210> 35 <211> 56 <212> DNA <213>(Artificial Sequence) <220> <223> <400> 35 agatcggcag cgtcagatgt gtataagaga cagtataggt gattttaggg gtgaga 56 <210> 36 <211> 51 <212> DNA <213>(Artificial Sequence) <220> <223> <400> 36 agatcgctct tccgatctta aatccttaaa taaactacat aaaaattttc c 51 <210> 37 <211> 62 <212> DNA <213>(Artificial Sequence) <220> <223> <400> 37 agatcggcag cgtcagatgt gtataagaga caggaggtag taatagggaa aatagttatt 60 gg 62 <210> 38 <211> 42 <212> DNA <213>(Artificial Sequence) <220> <223> <400> 38 agatcgctct tccgatctac caactatacc tctacatcaa aa 42 <210> 39 <211> 57 <212> DNA <213>(Artificial Sequence) <220> <223> <400> 39 agatcggcag cgtcagatgt gtataagaga cagaaggggg aattttagtt ttaggaa 57 <210> 40 <211> 48 <212> DNA <213>(Artificial Sequence) <220> <223> <400> 40 agatcgctct tccgatctaa aacctatatc tctaataaaa actcaata 48 <210> 41 <211> 54 <212> DNA <213> (Artificial Sequence) <220> <223> <400> 41 agatcggcag cgtcagatgt gtataagaga cagtttgttt taggaaagag gtgg 54 <210> 42 <211> 42 <212> DNA <213> (Artificial Sequence) <220> <223> <400> 42 agatcgctct tccgatctaa aaccccaaca ttcaattaaa aa 42 <210> 43 <211> 63 <212> DNA <213> (Artificial Sequence) <220> <223> <400> 43 agatcggcag cgtcagatgt gtataagaga cagaataatg taataagaat aaaaggtaag 60 gtt 63 <210> 44 <211> 44 <212> DNA <213> (Artificial Sequence) <220> <223> <400> 44 agatcgctct tccgatctaa caccatctca actcactaca aact 44 <210> 45 <211> 53 <212> DNA <213> (Artificial Sequence) <220> <223> <400> 45 agatcggcag cgtcagatgt gtataaga caggagtatt ggggatttag ggg 53 <210> 46 <211> 44 <212> DNA <213> (Artificial Sequence) <220> <223> <400> 46 agatcgctct tccgatctcc caacctcta atatatatac ccaa 44 <210> 47 <211> 62 <212> DNA <213> (Artificial Sequence) <220> <223> <400> 47 agatcggcag cgtcagatgt gtataagaga cagggataaa gtaaaggaga tattgtatgg 60 aa 62 <210> 48 <211> 48 <212> DNA <213> (Artificial Sequence) <220> <223> <400> 48 agatcgctct tccgatctaa ccacaaataa aatataaata ctcataaa 48 <210> 49 <211> 59 <212> DNA <213> (Artificial Sequence) <220> <223> <400> 49 agatcggcag cgtcagatgt gtataagaga cagggaggaa agagaatatt tgatatttg 59 <210> 50 <211> 43 <212> DNA <213>(Artificial Sequence) <220> <223> <400> 50 agatcgctct tccgatctaa cctctttatt tacaaaccta aac 43 <210> 51 <211> 59 <212> DNA <213>(Artificial Sequence) <220> <223> <400> 51 agatcggcag cgtcagatgt gtataagaga cagtatttta atctcctcac caacaaaaa 59 <210> 52 <211> 43 <212> DNA <213>(Artificial Sequence) <220> <223> <400> 52 agatcgctct tccgatctca cttcctaaaa crgaaaaatt cta 43 <210> 53 <211> 4 <212> DNA <213>(Artificial Sequence) <220> <223> <400> 53 tgta 4 <210> 54 <211> 4 <212> DNA <213>(Artificial Sequence) <220> <223> <400> 54 taca 4 <210> 55 <211> 3 <212> DNA <213>(Artificial Sequence) <220> <223> <400> 55 gat 3 <210> 56 <211> 3 <212> DNA <213>(Artificial Sequence) <220> <223> <400> 56 atc 3 <210> 57 <211> 17 <212> DNA <213>(Artificial Sequence) <220> <221> modified_base <222> (5)..(5) <223> n=dITP <400> 57 atgtntataa gagacag 17 <210> 58 <211> 8 <212> DNA <213>(Artificial Sequence) <220> <221> modified_base <222> (5)..(5) <223> n=dITP <400> 58 ttccnatc 8 <210> 59 <211> 39 <212> DNA <213>(Artificial Sequence) <220> <221> modified_base <222> (5)..(5) <223> n=dITP <400> 59 atgtntataa gagacaggag tagttgggat tataggtgt 39 <210> 60 <211> 35 <212> DNA <213>(Artificial Sequence) <220> <221> modified_base <222> (5)..(5) <223> n=dITP <400> 60 ttccnatcat atatatcaaa cactrgactt aaaat 35 <210> 61 <211> 41 <212> DNA <213>(Artificial Sequence) <220> <221> modified_base <222> (5)..(5) <223> n=dITP <400> 61 atgtntataa gagacagtta gaaatttagt tgtagagggg g 41 <210> 62 <211> 32 <212> DNA <213>(Artificial Sequence) <220> <221> modified_base <222> (5)..(5) <223> n=dITP <400> 62 ttccnatccc ttaaaacaaa cttatcttct cc 32 <210> 63 <211> 40 <212> DNA <213> (Artificial Sequence) <220> <221> modified_base <222> (5)..(5) <223> n=dITP <400> 63 atgtntataa gagacaggag gttagggttt tagattggga 40 <210> 64 <211> 36 <212> DNA <213> (Artificial Sequence) <220> <221> modified_base <222> (5)..(5) <223> n=dITP <400> 64 ttccnatcca ccttaacaaa taaaataata attcac 36 <210> 65 <211> 39 <212> DNA <213> (Artificial Sequence) <220> <221> modified_base <222> (5)..(5) <223> n=dITP <400> 65 atgtntataa gagacaggta aygaattggt agagtttta 39 <210> 66 <211> 32 <212> DNA <213>(Artificial Sequence) <220> <221> modified_base <222> (5)..(5) <223> n=dITP <400> 66 ttccnatcta tactaactcc cttcaaccat ta 32 <210> 67 <211> 41 <212> DNA <213>(Artificial Sequence) <220> <221> modified_base <222> (5)..(5) <223> n=dITP <400> 67 atgtntataa gagacagtga gggtaagaat tatttagagg t 41 <210> 68 <211> 31 <212> DNA <213>(Artificial Sequence) <220> <221> modified_base <222> (5)..(5) <223> n=dITP <400> 68 ttccnatcct acccacacct accaaaccta a 31 <210> 69 <211> 42 <212> DNA <213>(Artificial Sequence) <220> <221> modified_base <222> (5)..(5) <223> n=dITP <400> 69 atgtntata gagacagagg gttaaagag agaatgattt at 42 <210> 70 <211> 34 <212> DNA <213> (Artificial Sequence) <220> <221> modified_base <222> (5)…(5) <223> n=dITP <400> 70 ttccnatcat caaaataat tcaaaaata taca 34 <210> 71 <211> 45 <212> DNA <213> (Artificial Sequence) <220> <221> modified_base <222> (5)…(5) <223> n=dITP <400> 71 atgtntata gagacaggag ggttgaat taaaatagt agggt 45 <210> 72 <211> 38 <212> DNA <213> (Artificial Sequence) <220> <221> modified_base <222> (5)…(5) <223> n=dITP <400> 72 ttccnatcac caacttctat atacataata atacaca 38 <210> 73 <211> 46 <212> DNA <213>(Artificial Sequence) <220> <221> modified_base <222> (5)..(5) <223> n=dITP <400> 73 atgtntataa gagacaggga taattataag aattgtaaag gaggat 46 <210> 74 <211> 33 <212> DNA <213>(Artificial Sequence) <220> <221> modified_base <222> (5)..(5) <223> n=dITP <400> 74 ttccnatcaa aattcacttc taaatttaaa cca 33 <210> 75 <211> 42 <212> DNA <213>(Artificial Sequence) <220> <221> modified_base <222> (5)..(5) <223> n=dITP <400> 75 atgtntataa gagacagggt agttggaaat ggtaaatttg ag 42 <210> 76 <211> 38 <212> DNA <213>(Artificial Sequence) <220> <221> modified_base <222> (5)..(5) <223> n=dITP <400> 76 ttccnatcaa aataatcttc atcaaattaa taaaaaca 38 <210> 77 <211> 42 <212> DNA <213> (Artificial Sequence) <220> <221> modified_base <222> (5)..(5) <223> n=dITP <400> 77 atgtntataa gagacaggag ttatgttatg ggagtaagtg gg 42 <210> 78 <211> 32 <212> DNA <213> (Artificial Sequence) <220> <221> modified_base <222> (5)..(5) <223> n=dITP <400> 78 ttccnatcac accaaaaaca atttaataaa ca 32 <210> 79 <211> 38 <212> DNA <213> (Artificial Sequence) <220> <221> modified_base <222> (5)..(5) <223> n=dITP <400> 79 atgtntataa gagacagtta gggttttaga ttgggagg 38 <210> 80 <211> 34 <212> DNA <213>(Artificial Sequence) <220> <221> modified_base <222> (5)..(5) <223> n=dITP <400> 80 ttccnatctt ttaccaaaac taatactaac aact 34 <210> 81 <211> 42 <212> DNA <213>(Artificial Sequence) <220> <221> modified_base <222> (5)..(5) <223> n=dITP <400> 81 atgtntataa gagacaggtt agggaagttg atgttaggaa at 42 <210> 82 <211> 31 <212> DNA <213>(Artificial Sequence) <220> <221> modified_base <222> (5)..(5) <223> n=dITP <400> 82 ttccnatcaa tcaatctctc taaaccaaaa a 31 <210> 83 <211> 46 <212> DNA <213>(Artificial Sequence) <220> <221> modified_base <222> (5)..(5) <223> n=dITP <400> 83 atgtntata gagacagtag tttatgga agttgagata gaga 46 <210> 84 <211> 37 <212> DNA <213> (Artificial Sequence) <220> <221> modified_base <222> (5)…(5) <223> n=dITP <400> 84 ttccnatca tachaatca ttccnatcaa 37 <210> 85 <211> 46 <212> DNA <213> (Artificial Sequence) <220> <221> modified_base <222> (5)…(5) <223> n=dITP <400> 85 atgtntata gagacagaag attaatttaa taggattgga aggaat 46 <210> 86 <211> 37 <212> DNA <213> (Artificial Sequence) <220> <221> modified_base <222> (5)…(5) <223> n=dITP <400> 86 ttccnatcac atataaaac cctatctctta ctaaaaa <210> 87 <211> 40 <212> DNA <213>(Artificial Sequence) <220> <221> modified_base <222> (5)..(5) <223> n=dITP <400> 87 atgtntataa gagacagtat aggtgatttt aggggtgaga 40 <210> 88 <211> 45 <212> DNA <213>(Artificial Sequence) <220> <221> modified_base <222> (5)..(5) <223> n=dITP <400> 88 agatcgctct tccgatctta aatccttaaa taaactacat aaaaa 45 <210> 89 <211> 46 <212> DNA <213>(Artificial Sequence) <220> <221> modified_base <222> (5)..(5) <223> n=dITP <400> 89 atgtntataa gagacaggag gtagtaatag ggaaaatagt tattgg 46 <210> 90 <211> 32 <212> DNA <213>(Artificial Sequence) <220> <221> modified_base <222> (5)..(5) <223> n=dITP <400> 90 ttccnatcac caactatacc tctacatcaa aa 32 <210> 91 <211> 41 <212> DNA <213>(Artificial Sequence) <220> <221> modified_base <222> (5)..(5) <223> n=dITP <400> 91 atgtntataa gagacagaag ggggaatttt agttttagga a 41 <210> 92 <211> 38 <212> DNA <213>(Artificial Sequence) <220> <221> modified_base <222> (5)..(5) <223> n=dITP <400> 92 ttccnatcaa aacctatatc tctaataaaa actcaata 38 <210> 93 <211> 38 <212> DNA <213>(Artificial Sequence) <220> <221> modified_base <222> (5)..(5) <223> n=dITP <400> 93 atgtntataa gagacagttt gttttaggaa agaggtgg 38 <210> 94 <211> 32 <212> DNA <213>(Artificial Sequence) <220> <221> modified_base <222> (5)..(5) <223> n=dITP <400> 94 ttccnatcaa aaccccaaca ttcaattaaa aa 32 <210> 95 <211> 47 <212> DNA <213>(Artificial Sequence) <220> <221> modified_base <222> (5)..(5) <223> n=dITP <400> 95 atgtntataa gagacagaat aatgtaataa gaataaaagg taaggtt 47 <210> 96 <211> 34 <212> DNA <213>(Artificial Sequence) <220> <221> modified_base <222> (5)..(5) <223> n=dITP <400> 96 ttccnatcaa caccatctca actcactaca aact 34 <210> 97 <211> 37 <212> DNA <213>(Artificial Sequence) <220> <221> modified_base <222> (5)..(5) <223> n=dITP <400> 97 atgtntataa gagacaggag tattggggat ttagggg <210> 98 <211> 34 <212> DNA <213> (Artificial Sequence) <220> <221> modified_base <222> (5)..(5) <223> n=dITP <400> 98 ttccnatccc ccaacctcta 34. ttccnatccc ccaacctcta <210> 99 <211> 46 <212> DNA <213> (Artificial Sequence) <220> <221> modified_base <222> (5)..(5) <223> n=dITP <400> 99 atgtntataa gagacaggga taagtaag gagatattgt atggaa <210> 100 <211> 38 <212> DNA <213> (Artificial Sequence) <220> <221> modified_base <222> (5)..(5) <223> n=dITP <400> 100 ttccnatcaa ccacaata aatataaata ctcataa <210> 101 <211> 43 <212> DNA <213>(Artificial Sequence) <220> <221> modified_base <222> (5)..(5) <223> n=dITP <400> 101 atgtntataa gagacaggga ggaaagagaa tatttgatat ttg 43 <210> 102 <211> 33 <212> DNA <213>(Artificial Sequence) <220> <221> modified_base <222> (5)..(5) <223> n=dITP <400> 102 ttccnatcaa cctctttatt tacaaaccta aac 33 <210> 103 <211> 43 <212> DNA <213>(Artificial Sequence) <220> <221> modified_base <222> (5)..(5) <223> n=dITP <400> 103 atgtntataa gagacagtat tttaatctcc tcaccaacaa aaa 43 <210> 104 <211> 33 <212> DNA <213>(Artificial Sequence) <220> <221> modified_base <222> (5)..(5) <223> n=dITP <400> 104 ttccnatcca cttcctaaaa crgaaaaatt cta 33 <210> 105 <211> 58 <212> DNA <213>(Artificial Sequence) <220> <221> modified_base <222> (29)..(36) <223> 8nt Illumina Index <400> 105 aatgatacgg cgaccaccga gatctacacn nnnnnnntcg tcggcagcgt cagatgtg 58 <210> 106 <211> 16 <212> DNA <213>(Artificial Sequence) <220> <223> <400> 106 ctgacgctgc cgacga 16 <210> 107 <211> 57 <212> DNA <213>(Artificial Sequence) <220> <221> modified_base <222> (26)..(33) <223> 8nt Illumina Index <400> 107 gagcacacgt ctgaactcca gtcacnnnnn nnnatctcgt atgccgtctt ctgcttg 57 <210> 108 <211> 30 <212> DNA <213> (Artificial Sequence) <220> <223> <400> 108 gtgactggag ttcagacgtg tgctcttccg 30 <210> 109 <211> 5 <212> DNA <213> (Artificial Sequence) <220> <223> <400> 109 cacat 5 <210> 110 <211> 5 <212> DNA <213> (Artificial Sequence) <220> <223> <400> 110 atgtg 5 <210> 111 <211> 5 <212> DNA <213> (Artificial Sequence) <220> <223> <400> 111 cggaa 5 <210> 112 <211> 5 <212> DNA <213> (Artificial Sequence) <220> <223> <400> 112 ttccg 5
Claims
1. A method for constructing a multiplex PCR library for high-throughput targeted methylation sequencing, characterized in that, The template DNA is subjected to bisulfite conversion, and then a multi-base MoCODE barcode is introduced into the specific amplification product by using highly specific multiplex PCR primers containing the MoCODE barcode generation sequence. The amplification product is efficiently ligated to the sequencing adapter containing the MoCODE barcode decoding sequence to construct a library. The MoCODE barcode refers to the protruding single-stranded nucleotide sequences that form the two sticky ends of the obtained PCR product after digesting the multiplex PCR product with a specific endonuclease. The MoCODE barcode decoding sequence is a nucleotide sequence complementary to the MoCODE barcode. The multiplex PCR primers include primers with the nucleotide sequences shown below: (1) The sequences shown in Seq ID No: 2-11, 13-22, 27-52; or (2) The sequences shown in Seq ID No: 59-104, where n represents dITP.
2. The method according to claim 1, wherein The MoCODE barcode is generated by enzymatic digestion using a specific endonuclease.
3. The method according to claim 1 or 2, wherein the sequencing adapter is artificially designed and synthesized.
4. The method according to claim 3, wherein, The sequencing adapter is a dual adapter.
5. The method according to claim 4, wherein, Each specific region enrichment is decoded by a dual adapter.
6. The method according to claim 3, wherein The sequencing adapter contains the MoCODE barcode decoding sequence.
7. The method according to claim 6, wherein, The sequencing adapter further contains one or more of the sequencing adapters of the sequencing platform and index tags.
8. The method according to claim 6, wherein, The sequencing adapter contains the high-throughput sequencing universal sequence, index tag, and the MoCODE barcode decoding sequence.
9. The method according to claim 1, wherein The sequence of the MoCODE barcode includes the following nucleotide sequences: (1) The sequences shown in Seq ID No: 53, 55; or (2) The sequences shown in Seq ID No: 109, 111.
10. The method according to any one of claims 1-2 or 4-9, characterized in that, The sequence of the sequencing adapter includes the following nucleotide sequences: (1) The sequences shown in Seq ID No: 23-26; or (2) The sequences shown in Seq ID No: 105-108.
11. A method for constructing a multiplex PCR library for high-throughput targeted methylation sequencing, characterized in that, The method includes the following steps: 1) Extract template DNA from the sample to be tested, and the template DNA is subjected to bisulfite conversion; 2) Perform multiplex PCR reaction. Each primer participating in the multiplex PCR reaction contains a specific MoCODE barcode generation sequence, and the primer also contains a gene-specific sequence; 3) Purify the PCR product obtained in step 2) by the magnetic bead method; 4) Make the purified PCR product obtained in step 3) generate 5' and 3' sticky ends, and generate MoCODE barcodes at the 5' and 3' sticky ends respectively; 5) Purify the PCR product containing the MoCODE barcode in step 4) by the magnetic bead method; 6) Ligate the purified PCR product containing the MoCODE barcode obtained in step 5) and the sequencing adapter, and the sequencing adapter contains the MoCODE barcode decoding sequence complementary to the MoCODE; 7) Purify the ligation product obtained in step 6) by magnetic beads to complete the construction of the multiplex PCR library for high-throughput targeted sequencing; The multiplex PCR primers include primers with the nucleotide sequences shown below: (1) The sequences shown in Seq ID No: 2-11, 13-22, 27-52; or (2) The sequences shown in Seq ID No: 59-104, where n represents dITP.
12. The method according to claim 11, wherein The MoCODE barcode is generated by enzymatic digestion using a specific endonuclease.
13. The method of claim 11, wherein The sequencing adapter described in step 6) is a bidirectional adapter.
Citation Information
Patent Citations
Process for amplifying, detecting, and / or-cloning nucleic acid sequences
US4683195A
Process for amplifying nucleic acid sequences
US4683202A
Selective restriction fragment amplification: fingerprinting
US6045994A
Method for building library and method for SNP typing
WO2018040961A1