Construction method of methylation library and kit

Through the method of combining the Y-type linker element with single enzyme digestion and T7 ligase, the methylation library construction process is simplified, the problems of cumbersome and high cost in the existing technology are solved, efficient and low-cost methylation library construction is achieved, and the pass rate of sequencing data is improved.

CN119932729APending Publication Date: 2025-05-06CHINESE ACAD OF FISHERY SCI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311466708.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-06
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

In the prior art, the construction process of methylation library is cumbersome and expensive, especially the reagents that convert non-methylated C bases into U are expensive and cumbersome, and the operation is lacking an effective simplified solution.

Method used

The DNA sample was enzymatically cleaved by single enzyme digestion to obtain enzyme fragments with sticky ends, and then the T7 ligase and the Y-type linker element with complementary sticky ends were ligated for double-terminal linkages. After C-T base conversion and PCR amplification, a methylation library was finally constructed.

Benefits of technology

The methylation library construction process is simplified, the reagent cost is reduced, the two-end sequencing requirements of the methylation library are realized, and the qualified data rate of the library construction is improved to 100%.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119932729A_ABST
    Figure CN119932729A_ABST
Patent Text Reader

Abstract

The invention provides a construction method of a methylation library and a kit. The construction method comprises the following steps: carrying out single enzyme digestion on a DNA sample to obtain an enzyme digestion fragment with a cohesive end; under the action of T7 ligase, connecting a linker element with the enzyme digestion fragment to obtain a linker connection fragment; c-U base conversion is carried out on the linker connection fragment to obtain a base conversion fragment, C in the linker element is modified, and C-U conversion is not carried out; carrying out PCR (Polymerase Chain Reaction) amplification on the base transformation fragment to obtain a methylated library; wherein the linker element has a cohesive end complementary to the cohesive end of the digestion fragment. According to the method, enzyme digestion library building is utilized, the operation of tedious step of adding A at the tail end and increasing cost is not needed, and the requirement of the methylation library on double-end sequencing is met. Meanwhile, the effect that the qualified data rate of the constructed library reaches 100% is also unexpectedly achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of high-throughput sequencing library construction, and in particular to a method and a kit for constructing a methylation library. Background Art

[0002] DNA high-throughput sequencing and DNA methylation high-throughput sequencing are widely used in medicine, agriculture and other biological research fields. Among them, the commonly used methods for constructing DNA libraries and methylation libraries are:

[0003] DNA library: DNA shearing -> end repair -> add A tail to 3′ end -> adapter ligation -> PCR amplification using primers containing index.

[0004] DNA methylation library: DNA shearing -> end repair -> add A tail to 3′ end -> adapter ligation -> convert non-methylated C base to U -> PCR amplification using primers containing index.

[0005] In these library construction processes, especially DNA methylation library construction, the reagents for converting non-methylated C bases to U are expensive and the operation process is cumbersome. How to simplify the workload becomes the key to reducing costs and improving efficiency.

[0006] However, there is currently no effective solution. Summary of the invention

[0007] The main purpose of the present invention is to provide a method and a kit for constructing a methylation library, so as to solve the problems of complicated process and high cost of constructing a methylation library in the prior art.

[0008] In order to achieve the above-mentioned purpose, according to one aspect of the present invention, a method for constructing a methylation library is provided, and the method comprises: performing single enzyme digestion on a DNA sample to obtain enzyme-digested fragments with sticky ends; connecting a linker element to the enzyme-digested fragments under the action of T7 ligase to obtain a linker-connected fragment; performing CU base conversion on the linker-connected fragment to obtain a base conversion fragment, wherein C in the linker element is modified and is not subjected to CU conversion; performing PCR amplification on the base conversion fragment to obtain a methylation library; wherein the linker element has a sticky end complementary to the sticky end of the enzyme-digested fragment.

[0009] Furthermore, the endonuclease is Msp I.

[0010] Furthermore, the linker element is a Y-shaped linker, and the ends of the Y-shaped linker are sticky ends complementary to the sticky ends of the enzyme-cleaved fragments.

[0011] Furthermore, the Y-shaped connector includes a portion matching the PCR amplification primer, a portion matching the target sequence, and a sticky end complementary to the sticky end of the enzyme-cut fragment, which are sequentially connected. The Y-shaped connector also includes a molecular tag UMI, and the molecular tag UMI is located between the target sequence matching portion and the sticky end complementary to the sticky end of the enzyme-cut fragment.

[0012] Further, the linker element has the following sequence and structure:

[0013]

[0014] Furthermore, there are multiple DNA samples, and the construction method includes: performing enzyme digestion on the multiple DNA samples to obtain enzyme-digested fragments with sticky ends corresponding to the multiple DNA samples; under the action of T7 ligase, connecting the adapter element with the enzyme-digested fragments corresponding to each DNA sample to obtain adapter-connected fragments; mixing the adapter-connected fragments from the multiple DNA samples to obtain mixed adapter-connected fragments; performing CT base conversion on the mixed adapter-connected fragments to obtain base conversion fragments, wherein the C in the adapter element is modified and is not subjected to CU conversion; performing PCR amplification on the base conversion fragments to obtain a methylation library.

[0015] Furthermore, the molecular tag UMI in the adapter element for the multiple DNA sample pairs is selected from any one of the following:

[0016] TAGC, CCTA, CTTG, GAAC, CCAA, TATC, CACA, ACTC, TAGA, CGAA, AAGA, ACTA, GGTA, TTGA, ACTG, GACA, AACG, TCTA, TGTG, CATC, TGTC, ATGC, GGAA, AAGG, C TTA, TTCA, GTAA, CGTA, CAAG, GTAG, AGGA, TTAG, AGCA, CTGA, AGTC, GTTC, TCCA, CTTC, GTCA, ATCG, TACG, TCAA, AGAC, ACGA, CAGA, ATAC, AATG, CTAC, AA GC, TGGA, AATC, TATG, TCTG, GTGA, TCAG, ATTG, CAAC, TTGC, TACC, TTCG, AGTA, ATCC, ACAG, AGAA, TAAC, ATCA, ATAG, ACAA, TGTA, TCAC, TTCC, ATGA, TTG G, CTAA, TTAC, GATG, GTTG, ACCA, CATA, GAAG, GCAA, TCTC, GATA, ACAC, TAGG, AACC, GAGA, TGAA, TACA, AACA, TAAG, AGAG, TGAG, ATTC, AGTG, GTTA and CTCA.

[0017] According to the second aspect of the present application, a methylation library construction kit is provided, which comprises: T7 ligase, an endonuclease for enzymatically cleaving a DNA sample, and a linker element, wherein the linker element has sticky ends with complementary ends, and the sticky ends of the linker element are complementary to the sticky ends of the cleavage fragments produced by the endonuclease.

[0018] Further, the endonuclease is Msp I; preferably, the linker element is a Y-shaped linker, and the end of the Y-shaped linker is a sticky end complementary to the sticky end of the enzyme-cut fragment; more preferably, the Y-shaped linker includes a sequentially connected portion matching the PCR amplification primer, a portion matching the target sequence, and a sticky end complementary to the sticky end of the enzyme-cut fragment, and the Y-shaped linker also includes a molecular tag UMI, and the molecular tag UMI is located between the target sequence matching portion and the sticky end complementary to the sticky end of the enzyme-cut fragment; further preferably, the linker element has the following sequence and structure:

[0019]

[0020] Preferably, the molecular tag UMI is selected from any one of the following:

[0021] TAGC, CCTA, CTTG, GAAC, CCAA, TATC, CACA, ACTC, TAGA, CGAA, AAGA, ACTA, GGTA, TTGA, ACTG, GACA, AACG, TCTA, TGTG, CATC, TGTC, ATGC, GGAA, AAGG, C TTA, TTCA, GTAA, CGTA, CAAG, GTAG, AGGA, TTAG, AGCA, CTGA, AGTC, GTTC, TCCA, CTTC, GTCA, ATCG, TACG, TCAA, AGAC, ACGA, CAGA, ATAC, AATG, CTAC, AA GC, TGGA, AATC, TATG, TCTG, GTGA, TCAG, ATTG, CAAC, TTGC, TACC, TTCG, AGTA, ATCC, ACAG, AGAA, TAAC, ATCA, ATAG, ACAA, TGTA, TCAC, TTCC, ATGA, TTG G, CTAA, TTAC, GATG, GTTG, ACCA, CATA, GAAG, GCAA, TCTC, GATA, ACAC, TAGG, AACC, GAGA, TGAA, TACA, AACA, TAAG, AGAG, TGAG, ATTC, AGTG, GTTA and CTCA.

[0022] The technical scheme of the present invention is applied to digest the sample DNA with a single enzyme to generate a digested fragment with a sticky end, and then use T7 ligase and a connector element with a sticky end complementary to the sticky end of the digested fragment to connect the double-end connector to obtain a connection fragment, and then through base conversion, finally amplify the connection fragment through a PCR amplification primer universal to the sequencing platform to obtain a methylation library. This method not only utilizes enzyme digestion to build a library, but also does not require the cumbersome step of end repair and A addition, which increases the cost, and meets the requirements of the methylation library for double-end sequencing. At the same time, it also unexpectedly achieves the effect of 100% qualified data rate of the constructed library. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] The drawings constituting a part of the present application are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:

[0024] Figure 1 A schematic diagram of the structure of a connector element based on TA connection in the prior art is shown;

[0025] Figure 2 The result diagram of partial sequencing data of the library obtained by the methylation library construction method in Comparative Example 1 is shown.

[0026] Figure 3 The result diagram shows partial sequencing data of the library obtained by the methylation library construction method in Example 1.

[0027] Figure 4 A schematic diagram showing the process of the methylation library construction method in the prior art and the improved methylation library construction method of the present application is shown;

[0028] Figure 5 The library test results of the library obtained by the methylation library construction method in Example 1 are shown;

[0029] Figure 6 The library test results of the library obtained by the methylation library construction method in Comparative Example 1 are shown. DETAILED DESCRIPTION

[0030] It should be noted that, in the absence of conflict, the embodiments and features in the embodiments of the present application can be combined with each other. The present invention will be described in detail below in conjunction with the embodiments.

[0031] As mentioned in the background technology, in DNA methylation research, since the C nucleotides in the adapter sequence need to be modified, the cost of adapter synthesis is more expensive than that of commonly used adapters. What is more troublesome is that the library construction step usually requires the step of end repair and A addition, which further increases the cost and makes the process cumbersome. In addition, the efficiency of library construction will also be affected by the efficiency of repair and A addition.

[0032] To improve this situation, the inventors conducted a detailed analysis and comparison of existing DNA library construction methods and linker elements used in library construction, and found the following:

[0033] Different DNA library construction methods are often related to the type of adapter. The adapter elements for constructing DNA libraries are mainly divided into two types according to the adapter ends. One is the blunt-end T adapter, which is used to bind the A base at the 3' end of the target DNA fragment. The other is the adapter of a specific restriction endonuclease, which cuts DNA with a specific restriction endonuclease that produces sticky ends, and produces corresponding ends that match the sticky end sequence of the adapter. Under the action of DNAT4 ligase, the double-stranded DNA cut between the adapter and the target DNA fragment is repaired.

[0034] Taking the blunt-ended T-junction as an example, according to the DNA structure of the junction, it can be divided into: linear junction ( Figure 1 a), hairpin connector ( Figure 1 b), Y-type connector ( Figure 1 c) and bridge joint ( Figure 1 (d)

[0035] The steps for constructing a DNA library with blunt-ended T-type adapters are as follows: a) using ultrasound or enzyme digestion to break the DNA, b) performing end repair on the broken DNA and adding an A base to the 3′ end, c) connecting the DNA to the adapter, and d) using primers containing an index to perform PCR amplification on the connection product.

[0036] The general steps for constructing a DNA library based on linkers designed with specific restriction endonucleases are: a) using two enzymes to cut DNA, b) connecting DNA to linkers, and c) PCR amplification of the ligation products using primers containing indexes.

[0037] The difference between using a blunt-end T-type connector and using a specific endonuclease sticky end to build a library is that when using a blunt-end T-type connector, the DNA end needs to be repaired and an A tail needs to be added to the 3′ end, so there is one more step than building a library based on enzyme digestion, and the cost of DNA end repair and 3′ end A tail addition is relatively high. When building a library based on a specific endonuclease connector, double enzyme digestion is generally used. In order to reduce the self-ligation of the connector and improve the amplification efficiency, usually, one of the enzymes uses a Y-type connector and the other uses a linear connector design. The hairpin connector is a special connector for NEB, which generally requires a special enzyme for NEB. The bridge connector is a special connector designed by BGI, which is specially adapted to the BGI sequencing platform.

[0038] The above description is applicable to the construction of all DNA libraries, but not to the construction of all DNA methylation libraries. Although the principles and steps of DNA methylation library construction are generally similar to those of DNA library construction, since DNA methylation library construction involves the step of converting unmethylated C bases to U bases using sulfite, the adapter element of the DNA library cannot be directly used for the construction of the DNA methylation library. Only after the C in the adapter is modified can it be used for the construction of the methylation library.

[0039] In addition, there is another very important difference between DNA methylation library and DNA library, that is, during sequencing, DNA library usually only needs to measure one of the DNA chains to meet the requirements of genotyping. For DNA methylation library, when C base is converted to U base, in fact, some bases in the original double-stranded DNA no longer match the bases of the original complementary chain, especially for the area with high CG content, which has actually become a single strand, so it is necessary to measure both chains as much as possible during sequencing. Taking the commonly used PE150 sequencing as an example, to achieve this, the length of the inserted target DNA fragment should be controlled as much as possible to 50-200bp, while the DNA library is usually best with an inserted fragment of 300-500bp.

[0040] Currently, looking at all the products on the market, only BGI, NEB and Qiagen can provide a complete set of kits for DNA methylation library construction, while Takara does not provide a complete set of kits. The kits it provides are for enrichment first, and then other kits are used to build the library. BGI and NEB use their proprietary adapters and specifically modify the C base to avoid the impact of subsequent steps.

[0041] In addition, in order to reduce sequencing costs, some people have proposed simplified methylation sequencing (RRBS), which uses MspI digestion, DNA end repair and A tail addition at the 3′ end, so that the subsequent steps can be completed using NEB or BGI kits. The library construction process of BGI is similar to that of NEB except for the final circularization step.

[0042] However, in terms of cost, the existing RRBS technology still requires end repair and 3′ end A tailing, which makes it difficult to reduce costs. One way to avoid or skip this step is to design a linker that matches the corresponding restriction site. If the cost of subsequent steps is to be further reduced, UMI technology can be used and added when designing the linker. Since the recognition site of MspI is CCGG, it is easy to appear in areas with high GC content or areas where methylation modification occurs. Therefore, the linker design is based on the MspI enzyme recognition site, combined with UMI, and domestic reagents can be used in all steps of library construction, without relying on expensive full kits from foreign companies, which simplifies experimental operations and greatly reduces library construction costs.

[0043] On this basis, the inventors of the present application optimized and designed a linker element specifically suitable for the construction of a methylation library, the specific structure of which is as follows:

[0044]

[0045] Different from conventional DNA methylation library construction kits, the improved design of the adapter adds a unique molecular identifier (UMI) before the T base at the 3′ end of the commonly used adapter. The UMI can be composed of 2 to 8 (preferably 2 to 4) bases. After the samples are connected to the adapter, they can be mixed, thereby reducing subsequent costs. At the same time, the samples can be easily distinguished in the sequencing data after library construction. All C bases in the adapter need to be methylated to avoid affecting the subsequent library construction.

[0046] Further analysis found that the existing DNA methylation library constructed based on the enzyme digestion method all uses double linear joints, or linear joints + Y-type joints, and no DNA methylation library construction method with double Y-type joints has been found. As analyzed above, double Y-type joints must be used to construct the DNA methylation library based on enzyme digestion. When constructing a conventional DNA library, double digestion is generally adopted, so double joints are used, but it is not possible to construct the DNA methylation library using the conventional DNA library construction double joints, because it is difficult to obtain DNA fragments with high GC content, so a large number of practices have found that MspI single digestion can obtain DNA fragments with high GC content at a high ratio, so it is easy to obtain related DNA methylation sites, but the problem brought about by this is that the same joints must be used at both ends, because the C bases of the positive and negative chains of the DNA may be methylated, so the DNA methylation library should ensure that the positive and negative chains of the same section of DNA fragment can be detected, and when only single joints can be used, only Y-type joints can be used, and linear joints cannot.

[0047] Based on this idea, the inventors first tried to construct a library by enzyme digestion of double Y-type adapters, and used T4 DNA ligase, which is used in all existing DNA library construction methods, for adapter ligation. The results showed that this method can construct a library, but a large proportion (90% on average) of the resulting library is useless data (such as Figure 2 ), among the 8 randomly selected valid reads, only the third read is correct, and the other reads are all wrong.

[0048] Therefore, the inventors thought it was DNA contamination, so they repeated the experiment using freshly extracted DNA from different species (including Cynoglossus semilaevis, Japanese shrimp, and humans), and obtained the same results, thus ruling out DNA contamination. After further careful analysis of the cause, it was speculated that it was caused by the specificity of DNA ligase. According to the information consulted, the commonly used T4 DNA ligase was replaced with T7 DNA ligase, which has better specificity and is specially used to connect sticky ends. The same DNA was used for library construction, and the sequencing results showed that all the measured data were qualified data (such as Figure 3 ). All 9 randomly selected clean reads were qualified, and no unqualified data were found in the batch test. The proportion of qualified data increased from the original 10% to 100%.

[0049] This proves that the present invention avoids the steps of DNA end repair and 3′ end A addition in the conventional DNA methylation library construction step by making full use of the characteristics of T7 DNA ligase in specifically connecting sticky ends, thereby simplifying the process and saving reagent costs. Further, by adding a molecular tag UMI to the connector (when performing low-frequency mutation detection or polymorphism detection on multiple samples, a UMI of 4 to 8 bases is sometimes added before the sticky end base of the connector to identify different individuals or tissues, cells and other different parts. Different combinations are formed with existing index-containing primers to distinguish these individuals, tissues or cells. In the step of library construction PCR amplification, it is equivalent to doing multiple PCR. Since the cost of PCR amplification accounts for a very low proportion of the entire library construction cost, the use of a connector containing UMI is equivalent to reducing the number of required index primers and reducing the cost of index primer synthesis), the connection products can be mixed and then the base conversion step is performed, so the cost of base conversion of a single sample is also reduced. Based on the current market price, the cost of building a single sample library of the present invention is compared with that of other commonly used kits (the price is subject to the results of the query on October 20, 2023), and the results are as follows:

[0050] Table 1:

[0051]

[0052] Based on the above research results, the applicant proposed the technical solution of the present application. In a typical embodiment, a method for constructing a methylation library is provided, which comprises: performing a single enzyme digestion on a DNA sample to obtain an enzyme-cut fragment with a sticky end; connecting a linker element to the enzyme-cut fragment under the action of T7 ligase to obtain a linker-connected fragment; performing CT base conversion on the linker-connected fragment to obtain a base conversion fragment, wherein the C in the linker element is modified and no CU conversion is performed; performing PCR amplification on the base conversion fragment to obtain a methylation library; wherein the linker element has a sticky end complementary to the sticky end of the enzyme-cut fragment.

[0053] The methylation library construction method of the present application uses a single enzyme to digest the sample DNA to produce a digested fragment with a sticky end, and then uses T7 ligase and a connector element with a sticky end complementary to the sticky end of the digested fragment to connect the double-end connector to obtain a connection fragment, and then through base conversion, finally amplify the connection fragment through a PCR amplification primer universal to the sequencing platform to obtain a methylation library. This method not only utilizes enzyme digestion to build a library, but also does not require the cumbersome step of end repair and A addition, which increases the cost, and also meets the requirements of the methylation library for double-end sequencing. At the same time, it also unexpectedly achieves the effect of 100% qualified data rate of the constructed library.

[0054] The similarities and differences between the existing methylation library construction process and the library construction process of this application are as follows Figure 4 shown. Figure 4 In the figure, a shows the schematic diagram of the DNA methylation library construction process of the present application, b shows the existing RRBS methylation library construction process; c shows the library construction process of the Qiagen DNA methylation kit, and d shows the library construction process of the NEB DNA methylation kit. It can be seen that although there are many existing methylation library construction methods, none of them can truly achieve low-cost, simple, fast and accurate library construction. The improved methylation library construction method of the present application can have the advantages of low cost and simple process, fast and accurate library construction.

[0055] The above-mentioned endonuclease can be any single endonuclease, and preferably the endonuclease MspI that recognizes high GC content regions is used to help enrich and detect methylated fragments.

[0056] As mentioned above, for the construction of methylation library, it is necessary to use a linker capable of double-end sequencing, and the double-end sequencing linker currently available is the Y-type linker. Therefore, in a preferred embodiment of the present application, the linker element is a Y-type linker, and in order to match the library construction method of enzyme digestion, the end of the Y-type linker is a sticky end complementary to the sticky end of the enzyme digestion fragment.

[0057] In some preferred embodiments, the Y-shaped adapter of the present application includes a portion matching a PCR amplification primer, a portion matching a target sequence, and a sticky end complementary to a sticky end of an enzyme-cut fragment, which are sequentially connected, and the Y-shaped adapter also includes a molecular tag UMI, and the molecular tag UMI is located between the target sequence matching portion and the sticky end complementary to the sticky end of the enzyme-cut fragment.

[0058] The above-mentioned linker element has the following sequence and structure:

[0059]

[0060] By adding a 2-4 base unique molecular identifier (UMI) inside the connector, multiple samples with different UMIs can be mixed after the connector is connected without affecting the subsequent steps. This can greatly reduce the workload, reduce the amount of reagents used, and improve efficiency, especially for methylation libraries, which can significantly reduce the reagent cost. This method can be used for Illumina sequencing platforms and BGI sequencing platforms, solving the problem that traditional methods cannot mix samples after connection and are costly.

[0061] In some preferred embodiments, there are multiple DNA samples, and the above-mentioned construction method includes: performing enzyme digestion on the multiple DNA samples to obtain enzyme-digested fragments with sticky ends corresponding to the multiple DNA samples respectively; under the action of T7 ligase, connecting the linker element with the enzyme-digested fragments corresponding to each DNA sample to obtain linker-connected fragments; mixing the linker-connected fragments from the multiple DNA samples to obtain mixed linker-connected fragments; performing CT base conversion on the mixed linker-connected fragments to obtain base conversion fragments, wherein the C in the linker element is modified and is not subjected to CU conversion; performing PCR amplification on the base conversion fragments to obtain a methylation library.

[0062] When the adapter contains UMI, multiple samples are labeled with unique molecular tags through the adapter connection step, so that the samples can be mixed before the base conversion step. Therefore, not only the amount of samples after mixing is reduced, reducing the workload, but also the cost of subsequent steps is greatly reduced.

[0063] UMIs of 2-4 bases can label different DNA sample molecules. In order to further facilitate the construction of large-scale sample libraries, in a preferred embodiment of the present application, the molecular label UMI (5' to 3' direction) in the adapter element used for multiple DNA samples is selected from any of the following:

[0064] TAGC, CCTA, CTTG, GAAC, CCAA, TATC, CACA, ACTC, TAGA, CGAA, AAGA, ACTA, GGTA, TTGA, ACTG, GACA, AACG, TCTA, TGTG, CATC, TGTC, ATGC, GGAA, AAGG, C TTA, TTCA, GTAA, CGTA, CAAG, GTAG, AGGA, TTAG, AGCA, CTGA, AGTC, GTTC, TCCA, CTTC, GTCA, ATCG, TACG, TCAA, AGAC, ACGA, CAGA, ATAC, AATG, CTAC, AA GC, TGGA, AATC, TATG, TCTG, GTGA, TCAG, ATTG, CAAC, TTGC, TACC, TTCG, AGTA, ATCC, ACAG, AGAA, TAAC, ATCA, ATAG, ACAA, TGTA, TCAC, TTCC, ATGA, TTG G, CTAA, TTAC, GATG, GTTG, ACCA, CATA, GAAG, GCAA, TCTC, GATA, ACAC, TAGG, AACC, GAGA, TGAA, TACA, AACA, TAAG, AGAG, TGAG, ATTC, AGTG, GTTA and CTCA.

[0065] In a second typical embodiment, a methylation library construction kit is provided, which comprises: T7 ligase, an endonuclease for enzymatically cleaving a DNA sample, and a linker element, wherein the linker element has sticky ends with complementary ends, and the sticky ends of the linker element are complementary to the sticky ends of the cleavage fragments generated by the endonuclease.

[0066] In a preferred embodiment, the endonuclease is Msp I. The recognition site of Msp I endonuclease is CCGG, so using this enzyme for single digestion helps to obtain DNA fragments with a high proportion of GC content, so it is easy to obtain relevant DNA methylation sites.

[0067] In a preferred embodiment, the above-mentioned connector element is a Y-shaped connector, and the end of the Y-shaped connector is a sticky end complementary to the sticky end of the enzyme-cut fragment; more preferably, the Y-shaped connector includes a sequentially connected portion matching the PCR amplification primer, a portion matching the target sequence, and a sticky end complementary to the sticky end of the enzyme-cut fragment, and the Y-shaped connector also includes a molecular tag UMI, and the molecular tag UMI is located between the target sequence matching portion and the sticky end complementary to the sticky end of the enzyme-cut fragment; further preferably, the connector element has the following sequence and structure:

[0068]

[0069] Preferably, the molecular tag UMI is selected from any one of the following:

[0070] TAGC, CCTA, CTTG, GAAC, CCAA, TATC, CACA, ACTC, TAGA, CGAA, AAGA, ACTA, GGTA, TTGA, ACTG, GACA, AACG, TCTA, TGTG, CATC, TGTC, ATGC, GGAA, AAGG, C TTA, TTCA, GTAA, CGTA, CAAG, GTAG, AGGA, TTAG, AGCA, CTGA, AGTC, GTTC, TCCA, CTTC, GTCA, ATCG, TACG, TCAA, AGAC, ACGA, CAGA, ATAC, AATG, CTAC, AA GC, TGGA, AATC, TATG, TCTG, GTGA, TCAG, ATTG, CAAC, TTGC, TACC, TTCG, AGTA, ATCC, ACAG, AGAA, TAAC, ATCA, ATAG, ACAA, TGTA, TCAC, TTCC, ATGA, TTG G, CTAA, TTAC, GATG, GTTG, ACCA, CATA, GAAG, GCAA, TCTC, GATA, ACAC, TAGG, AACC, GAGA, TGAA, TACA, AACA, TAAG, AGAG, TGAG, ATTC, AGTG, GTTA and CTCA.

[0071] The use of a Y-type linker element containing the above structure and T7 ligase for the construction of a methylation library can not only simplify the existing library construction process, but also enable the construction of mixed methylation libraries for large quantities of samples when the Y-type linker contains UMI.

[0072] It should be noted that, in the Y-type connector used in this application, except for the base sequences of the UMI and the sticky ends, the remaining sequences are the sequences in the universal Y-type connector of the Illumina sequencing platform. Accordingly, in the last PCR amplification step, the universal amplification primers containing the P5-end and P7-end indexes (library tags) of the Illumina sequencing platform are also used.

[0073] The beneficial effects of the present application will be further illustrated below with reference to specific embodiments.

[0074] Embodiment 1:

[0075] 1. Specific connector design

[0076] 1. Specific linker sequence structure

[0077]

[0078] All C bases in the linker are methylated. The sequence used for the UMI in SEQ ID NO: 1 in this example is: 5'-TAGC-3', and its complementary sequence in SEQ ID NO: 2 is 3'-ATCG-5'.

[0079] 2. Dilute the synthesized linker to a concentration of 10uM.

[0080] 2. Database construction steps:

[0081] 1. Perform quality test on gonadal DNA samples from Cynoglossus semilaevis.

[0082] Specifically, the DNA quality is tested by electrophoresis, UV spectrophotometer or Qubit and other instruments.

[0083] 2. Restriction endonuclease (RE) cleavage. Select a methylase whose recognition site does not contain C bases or is insensitive to 5m C to interrupt the DNA. It is best to use a single enzyme, preferably Msp I.

[0084] The Msp I single enzyme digestion 20uL reaction system is as follows:

[0085] Table 2:

[0086]

[0087] Reaction conditions: 37℃1hours; 4℃∞; (if a restriction endonuclease that cannot be heat-inactivated is used, DNA purification can be performed).

[0088] 3. Connector connection

[0089] The above-mentioned adapter with specific UMI is ligated with the above-mentioned DNA solution using T4 DNA ligase.

[0090] Reaction system:

[0091] Table 3:

[0092]

[0093] Reaction conditions: 20-22℃ for 2 hours; 65℃ for 30 min; 4℃∞ or 16℃ overnight; 65℃ for 30 min; 4℃∞.

[0094] 4. Sample mixing

[0095] Since the molecular tag UMI is connected, samples containing different UMIs can be mixed to reduce subsequent costs.

[0096] 5. Base conversion

[0097] The conversion was performed using a bisulfite conversion kit (Tiangen kit: DP215-02) or enzymes (such as NEB kit: E7125) according to standard corresponding procedures.

[0098] 6. PCR Amplification

[0099] PCR amplification is performed using primers with library indices. In this application, universal amplification primers for the Illumina sequencing platform are used.

[0100] The PCR primer structures are as follows:

[0101] P5 index primer:

[0102] 5′AATGATACGGCGACCACCGAGATCTACAC[i5:AGTCGCTT]TCGTCGGCAGCGTC3′ (SEQ IDNO: 3);

[0103] P7 index primer:

[0104] 5′-CAAGCAGAAGACGGCATACGAGAT[i7:CACTGTAG]GTCTCGTGGGCTCGG-3′ (SEQ ID NO: 4);

[0105] Among them, i5 and i7 are sequences containing indexes, and the index is generally 4 to 10 bases. In this embodiment, 8 bases are used.

[0106] Amplification requires the use of a polymerase that can amplify templates with U bases, such as NEB's Q5U Master Mix or Novozymes' EpiTaq enzyme.

[0107] Reaction system (taking Q5U as an example):

[0108] Table 4:

[0109]

[0110] Reaction conditions: 98°C for 30 sec; 12 cycles {98°C for 10 sec; 62°C for 30 sec; 65°C for 30 sec}; 65°C for 5 min; 4°C∞.

[0111] 7. Library detection

[0112] The library detection results are shown in Figure 5 , it can be seen that the library size is consistent with the size of the methylation library.

[0113] Comparative Example 1

[0114] The adapter element and library construction process of this application were used, with the only difference being that T4 DNA ligase was used in the adapter ligation step. The results of the constructed library test are shown in Figure 6 , it can be seen that the library size is consistent with the size of the methylation library.

[0115] Comparative Example 2

[0116] The same library construction method as in Comparative Example 1 was used, except that the DNA samples were replaced with Japanese shrimp and human DNA samples.

[0117] The methylation libraries constructed in the above examples and comparative examples were sequenced, and the sequencing data were sorted and analyzed. Some of the results are shown in Figure 3 and Figure 2 .from Figure 2 It can be seen that a large proportion (90% on average) of the obtained library is useless data. As shown in the 8 randomly selected valid reads, only the third read is correct, and the other reads are all wrong. Figure 3 The sequencing results show that all the measured data are qualified data. In the 9 randomly selected valid reads, it can be seen that the obtained data are all qualified data, and no unqualified data was found in the batch test. It can be seen that the improved method of the present application can increase the proportion of qualified data from the original 10% to 100%.

[0118] From the above description, it can be seen that the above embodiments of the present invention achieve the following technical effects: for DNA methylation research, since the subsequent steps after conventional linker connection are costly, complicated, time-consuming and labor-intensive, the use of the linker designed by the present invention and the use of T7 DNA ligase for linker connection can not only mix multiple samples, greatly reducing the cost of subsequent steps, but also simplify the library construction process and reduce the workload due to the reduction in sample size.

[0119] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A method for constructing a methylation library, characterized in that: The construction method comprises: Perform single enzyme digestion on DNA samples to obtain enzyme-digested fragments with sticky ends; Under the action of T7 ligase, the linker element is connected to the enzyme-cut fragment to obtain a linker-connected fragment; Performing CU base conversion on the linker-connected fragment to obtain a base conversion fragment, wherein the C in the linker element is modified and is not subjected to CU conversion; Performing PCR amplification on the base conversion fragments to obtain the methylation library; Wherein, the linker element has a sticky end complementary to the sticky end of the enzyme-cleaved fragment.

2. The construction method according to claim 1, characterized in that: The endonuclease was MspI.

3. The construction method according to claim 1, characterized in that: The linker element is a Y-shaped linker, and the end of the Y-shaped linker is a sticky end complementary to the sticky end of the enzyme-cut fragment.

4. The construction method according to claim 3, characterized in that: The Y-shaped connector includes a portion matching a PCR amplification primer, a portion matching a target sequence, and a sticky end complementary to the sticky end of the enzyme-cut fragment, which are connected in sequence. The Y-shaped connector also includes a molecular tag UMI, and the molecular tag UMI is located between the target sequence matching portion and the sticky end complementary to the sticky end of the enzyme-cut fragment.

5. The construction method according to claim 4, characterized in that: The linker element has the following sequence and structure:

6. The construction method according to claim 5, characterized in that: There are multiple DNA samples, and the construction method includes: Performing enzyme digestion on the multiple DNA samples to obtain enzyme-digested fragments with sticky ends corresponding to the multiple DNA samples; Under the action of T7 ligase, the adapter element is connected to the enzyme-cut fragments corresponding to each of the DNA samples to obtain adapter-connected fragments; Mixing the adapter-ligated fragments from the plurality of DNA samples to obtain mixed adapter-ligated fragments; Performing CT base conversion on the mixed adapter-connected fragments to obtain base conversion fragments, wherein the C in the adapter element is modified and is not subjected to CU conversion; The base conversion fragments are amplified by PCR to obtain the methylation library.

7. The construction method according to claim 6, characterized in that: The molecular tag UMI in the adapter element used for the multiple DNA sample pairs is selected from any one of the following: TAGC, CCTA, CTTG, GAAC, CCAA, TATC, CACA, ACTC, TAGA,CGAA,AAGA,ACTA,GGTA,TTGA,ACTG,GACA, AACG, TCTA, TGTG, CATC, TGTC, ATGC, GGAA, AAGG, CTTA, TTCA, GTAA, CGTA, CAAG, GTAG, AGGA, TTAG, AGCA, CTGA, AGTC, GTTC, TCCA, CTTC, GTCA, ATCG, TACG, TCAA, AGAC, ACGA, CAGA, ATAC, AATG, CTAC, AAGC, TGGA, AATC, TATG, TCTG, GTGA, TCAG, ATTG, CAAC, TTGC, TACC, TTCG, AGTA, ATCC, ACAG, AGAA, TAAC, ATCA, ATAG, ACAA, TGTA, TCAC, TTCC, ATGA, TTGG, CTAA, TTAC, GATG, GTTG, ACCA, CATA, GAAG, GCAA, TCTC, GATA, ACAC, TAGG, AACC, GAGA, TGAA, TACA, AACA, TAAG, AGAG, TGAG, ATTC, AGTG, GTTA and CTCA.

8. A methylation library construction kit, characterized in that: The methylation library construction kit comprises: T7 ligase, an endonuclease for enzymatically cleaving a DNA sample, and a linker element, wherein the linker element has a sticky end with complementary ends, and the sticky end of the linker element is complementary to the sticky end of the enzyme-cleaved fragment generated by the endonuclease.

9. The kit according to claim 8, characterized in that The endonuclease is Msp I; Preferably, the linker element is a Y-shaped linker, and the end of the Y-shaped linker is the sticky end complementary to the sticky end of the enzyme-cleaved fragment; More preferably, the Y-shaped adapter includes a portion matching a PCR amplification primer, a portion matching a target sequence, and a sticky end complementary to the sticky end of the enzyme-cut fragment, which are sequentially connected, and the Y-shaped adapter further includes a molecular tag UMI, and the molecular tag UMI is located between the target sequence matching portion and the sticky end complementary to the sticky end of the enzyme-cut fragment; Further preferably, the linker element has the following sequence and structure: Preferably, the molecular tag UMI is selected from any one of the following: TAGC, CCTA, CTTG, GAAC, CCAA, TATC, CACA, ACTC, TAGA, CGAA, AAGA, ACTA, GGTA, TTGA, ACTG, GACA, AACG, TCTA, TGTG, CATC, TGTC, ATGC, GGAA, AAGG, C TTA, TTCA, GTAA, CGTA, CAAG, GTAG, AGGA, TTAG, AGCA, CTGA, AGTC, GTTC, TCCA, CTTC, GTCA, ATCG, TACG, TCAA, AGAC, ACGA, CAGA, ATAC, AATG, CTAC, AA GC, TGGA, AATC, TATG, TCTG, GTGA, TCAG, ATTG, CAAC, TTGC, TACC, TTCG, AGTA, ATCC, ACAG, AGAA, TAAC, ATCA, ATAG, ACAA, TGTA, TCAC, TTCC, ATGA, TTG G, CTAA, TTAC, GATG, GTTG, ACCA, CATA, GAAG, GCAA, TCTC, GATA, ACAC, TAGG, AACC, GAGA, TGAA, TACA, AACA, TAAG, AGAG, TGAG, ATTC, AGTG, GTTA and CTCA.