Detection of dna methylation in targeted segments of the maize genome
By constructing a maize genomic DNA methylation detection library containing target regions and utilizing PCR amplification technology and recognition sequences, the problem of high cost and low efficiency in maize genomic DNA methylation detection using whole-genome sequencing methods was solved, achieving efficient and accurate DNA methylation detection of multiple target regions in the maize genome.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- HUAZHONG AGRI UNIV
- Filing Date
- 2022-05-25
- Publication Date
- 2026-06-02
Smart Images

Figure CN115161408B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of maize genomic DNA methylation detection technology, and in particular to the detection of DNA methylation in target regions of the maize genome. Specifically, it relates to the construction method, detection method, kit, and system of maize genomic DNA methylation detection library containing the target region. Background Technology
[0002] DNA methylation is a form of DNA epigenetic modification that can influence the inheritance of phenotypes without altering the DNA sequence. The level of DNA methylation in the maize genome is dynamic, influenced by the maize's growth and development stages and various environmental factors, such as salt stress and drought stress. Therefore, detecting maize DNA methylation is an important means of investigating how DNA methylation affects the regulation of maize growth and development and its response to abiotic stresses.
[0003] Whole-genome sequencing is a common method for detecting DNA methylation in plant genomes, allowing for the detection of methylation across the entire genome. However, DNA methylation that regulates gene expression is typically located in specific regions. For detecting methylation in these specific regions, whole-genome DNA methylation sequencing is not only costly, but even with increased sequencing depth, it often fails to adequately cover these regions. Furthermore, detecting methylation in multiple genomes and multiple specific regions requires multiple sequencing runs using whole-genome DNA methylation sequencing, increasing both cost and efficiency. Summary of the Invention
[0004] In view of this, the purpose of this application is to provide at least a method for constructing a library, a detection method, a kit, and a system for detecting DNA methylation in target regions of the maize genome, so as to solve one of the above-mentioned technical problems to a certain extent.
[0005] In a first aspect, embodiments of this application disclose a method for constructing a maize genomic DNA methylation detection library containing a target region, wherein the target region contains at least one or more SNP sites capable of distinguishing different maize genomes, and the DNA methylation detection library is used to detect a mixed sample of DNA from at least two maize genomes containing the target region, wherein the construction method includes:
[0006] The unmethylated cytosine in the DNA molecules of the DNA mixture sample is converted into uracil;
[0007] The first amplification is performed, which includes amplifying the DNA mixture sample in the same PCR reaction system to obtain a first amplification product. The first amplification product contains a DNA molecule having the target segment and a first recognition sequence and a bridge sequence attached to the 3' end of the DNA molecule. The first recognition sequence is used to identify each DNA molecule of the target segment in the DNA mixture sample.
[0008] A second amplification is performed, comprising the step of PCR amplification of the first amplification product to obtain a second amplification product, the second amplification product comprising the first amplification product and a second recognition sequence attached to the 3' end of the first amplification product, the second recognition sequence being used to identify the first amplification product; and
[0009] The second amplification product was used to construct the maize genome DNA methylation detection library containing the target region.
[0010] In this embodiment, the first amplification uses an upstream primer and a first downstream primer. The C base in the upstream primer is designed as a degenerate base Y; the G base in the first downstream primer is designed as a degenerate base R. The first downstream primer contains the first recognition sequence and a bridge sequence connected to the 3' end of the first recognition sequence. The bridge sequence is used for homologous matching with the second bridge sequence. The SNP site in the first amplification product is located within 150 bp of the binding start position of the DNA molecule having the target region with the upstream primer or with the binding start position of the first downstream primer.
[0011] In the application embodiments, the first amplification reaction step includes:
[0012] Pre-denaturation was performed at a temperature of 95℃ for 5 minutes.
[0013] A first amplification cycle is performed, which includes sequentially performing a first unwinding process, a first annealing process, and a first extension process.
[0014] A second amplification cycle is performed, comprising sequentially performing a second melting process, a second annealing process, and a second extension process; with treatment times of 95°C for 30 s, 60°C for 30 s, and 70°C for 30 s; and...
[0015] Treat at 72℃ for 5 min and at 12℃ for 1 s;
[0016] The first amplification cycle includes at least 8 to 12 cycles, and the annealing temperature of the first annealing process decreases gradually with each of the first amplification cycles.
[0017] In this embodiment of the application, the processing temperature of the first annealing treatment is 68-55℃, 65-53℃, 65-55℃ or 63-55℃.
[0018] In this embodiment of the application, the second amplification uses the upstream primer and the second downstream primer, the second downstream primer including the bridge sequence and the second recognition sequence, and the reaction steps of the second amplification include:
[0019] Pre-denaturation was performed at a temperature of 95℃ for 5 minutes.
[0020] An amplification cycle is performed, comprising sequentially performing melting, annealing, and extension processes; and
[0021] Treat at 72℃ for 5 min and at 12℃ for 1 s.
[0022] In this embodiment of the application, the second amplification product further includes a bridge sequence connecting the first identification sequence and the second identification sequence at the 3' end of the first amplification product, the bridge sequence being used to connect the first identification sequence and the second identification sequence.
[0023] Secondly, a method for detecting DNA methylation in a target region of the maize genome, wherein the target region contains at least one or more SNP sites capable of distinguishing different maize genomes, and the method for detecting DNA methylation in the target region is used to detect a mixed sample of DNA from at least two maize genomes containing the target region, wherein the detection method includes:
[0024] The first aspect of the construction method involves constructing a DNA methylation detection library of the maize genome containing the target region;
[0025] The DNA methylation detection library was sequenced to obtain a first reads library;
[0026] The first reads library is subjected to quality control, splitting, and reassembly to obtain the second reads library;
[0027] By comparing, deduplicating, and calculating the second reads library, the DNA methylation information of each maize genome in the DNA mixture sample is obtained.
[0028] In this embodiment of the application, the first reads library includes a number of first reads corresponding to the DNA methylation detection library, and the second reads library includes second reads composed of the same second recognition sequence, wherein the second reads are reads with the first recognition sequence, the second recognition sequence and the bridge sequence removed.
[0029] Thirdly, embodiments of this application disclose a kit for constructing a maize genomic DNA methylation detection library containing a target region, wherein the target region contains at least one or more SNP sites capable of distinguishing different maize genomes, and the DNA methylation detection library is used to detect mixed samples of DNA from at least two maize genomes containing the target region, wherein the kit comprises:
[0030] A maize genomic DNA processing reagent for extracting the maize genomic DNA and converting unmethylated cytosine into uracil;
[0031] The upstream primer, wherein the C base in the upstream primer sequence is replaced with a degenerate Y base sequence;
[0032] A first downstream primer, wherein the G base in the downstream primer sequence is replaced with a degenerate R base, and the first downstream primer includes a first recognition sequence and a bridging sequence; and
[0033] The second downstream primer comprises the bridge sequence and the second recognition sequence;
[0034] The upstream primer and the first downstream primer are used to perform a first amplification to obtain a first amplification product. The first amplification product contains a DNA molecule having the target segment and a first recognition sequence and a bridge sequence attached to the 3' end of the DNA molecule.
[0035] The upstream primer and the second downstream primer can be used to obtain a second amplification product in a second amplification. The second amplification product includes the first amplification product and a second recognition sequence attached to the 3' end of the first amplification product to construct the DNA methylation detection library.
[0036] The first identification sequence is used to identify each DNA molecule in the target region of the DNA mixture sample, and the second identification sequence is used to identify the first amplification product.
[0037] Fourthly, embodiments of this application disclose a system for detecting maize genomic DNA methylation containing a target region, wherein the target region contains at least one or more SNP sites capable of distinguishing different maize genomes, and the DNA methylation detection library is used to detect mixed samples of DNA from at least two maize genomes containing the target region, wherein the system includes:
[0038] The third aspect involves a kit for obtaining DNA methylation information for each maize genome in the said DNA mixture sample; and
[0039] A processing device, wherein the processing device is configured to run a program, wherein the program, when running, performs the detection method described in the third aspect.
[0040] "The DNA methylation detection library was sequenced to obtain a first reads library;"
[0041] The first reads library is subjected to quality control, splitting, and reassembly to obtain the second reads library;
[0042] The step is to compare, remove duplicates, and perform calculations on the second reads library to obtain the DNA methylation information of each maize genome in the DNA mixture sample.
[0043] Compared with the prior art, this application has at least one of the following beneficial effects:
[0044] The methods, detection methods, kits, and systems for constructing maize genomic DNA methylation detection libraries containing target regions involved in the embodiments of this application, for target regions containing SNP sites, can not only detect multiple maize genomic samples containing target regions at once, with high coverage, eliminating the need for multiple detections and improving detection efficiency, but also improve the reliability and accuracy of DNA methylation detection. Attached Figure Description
[0045] Figure 1 A schematic diagram illustrating the principle of maize genomic DNA methylation detection, including the target region, provided for embodiments of this application.
[0046] Figure 2 The primer and product structure diagrams for the first and second amplifications in the construction of maize genomic DNA methylation containing the target region provided in the embodiments of this application.
[0047] Figure 3 These are the target segments of the four maize inbred lines involved in the embodiments of this application, where SNP sites are marked in red.
[0048] Figure 4 The mixed samples of the genomes of the four maize inbred lines provided in this application were detected using the methylation detection method or system provided in this application; CHH\CHG\CG represent three sequence environments of methylated C bases.
[0049] Figure 5 The DNA methylation results of the two target segments provided in this application embodiment were obtained by whole-genome DNA methylation sequencing; CHH, CHG, and CG represent three sequence environments for methylated C bases, respectively.
[0050] Figure 6This is a gel electrophoresis diagram of the purified first amplification product provided in the embodiments of this application. Lane 1 is the target region, and lane 2 is the DNA molecular weight marker.
[0051] Figure 7 This is a gel electrophoresis diagram of the purified second amplification product provided in the embodiments of this application. Lane 1 is the target region, and lane 2 is the DNA molecular weight marker. Detailed Implementation
[0052] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application. Reagents not specifically described in detail herein are all conventional reagents and are commercially available; methods not specifically described in detail are all conventional experimental methods and can be learned from the prior art.
[0053] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence, nor do they substantially limit the technical features thereafter. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0054] Methods for constructing maize genomic DNA methylation detection libraries
[0055] When using specific regions of maize genomic DNA as target segments for DNA methylation detection, whole-genome DNA methylation sequencing, a method well-known to those skilled in the art, suffers from limited coverage and high costs, resulting in less than ideal detection performance. Even increasing sequencing depth does not adequately cover the target regions. Furthermore, for the detection of DNA methylation in multiple genomes and multiple specific regions, whole-genome DNA methylation sequencing can only detect one genome at a time, making it extremely expensive.
[0056] Therefore, this application discloses a method for constructing a DNA methylation detection library of a maize genome containing a target region. The target region contains at least one or more SNP sites capable of distinguishing different maize genomes. This DNA methylation detection library is used to detect mixed DNA samples containing at least two maize genomes with the target region. This application utilizes SNP sites to locate the target region. The sequentially constructed methylation library can be used to simultaneously detect the DNA methylation information of at least one target region in multiple maize genomes, or it can be used to detect the DNA methylation information of one target region in one maize genome, significantly improving the coverage of the target region and the detection efficiency.
[0057] Specifically, such as Figures 1-2 As shown, the method for constructing a maize genome DNA methylation detection library containing the target region disclosed in this application includes:
[0058] S1. Convert unmethylated cytosine in DNA molecules in a mixed DNA sample to uracil;
[0059] S2. Perform the first amplification, which includes amplifying the DNA mixture sample in the same PCR reaction system to obtain the first amplification product. The first amplification product contains DNA molecules with the target segment and a first recognition sequence and a bridge sequence attached to the 3' end of the DNA molecule. The first recognition sequence is used to identify each DNA molecule of the target segment in the DNA mixture sample.
[0060] S3. Perform a second amplification, which includes the step of PCR amplification of the first amplification product to obtain a second amplification product. The second amplification product includes the first amplification product and a second recognition sequence linked to the 3' end of the first amplification product. The second recognition sequence is used to identify the first amplification product.
[0061] S4. Construct a DNA methylation detection library of the maize genome containing the target region using the second amplification product.
[0062] Therefore, steps S1 to S4 above will be explained in more detail below.
[0063] S1. Obtaining a mixed DNA sample of the maize genome.
[0064] In the embodiments of step S1 of this application, the DNA mixed sample is obtained by mixing the extracted maize genomic DNA samples. In more embodiments, there may be two or more maize genomic DNA samples.
[0065] Specifically, the "conversion" can be induced by physical conditions such as chemical mutagen or ultraviolet mutagenesis, as long as the unmethylated deaminated cytosine is converted into uracil. There are no restrictions on the specific "conversion" methods, such as using bisulfite for treatment.
[0066] A specific implementation process for step S1 is as follows:
[0067] (1) Extraction of maize genomic DNA
[0068] In this embodiment, four common maize inbred lines (B73, Mo17, W22, and SK) were selected as materials at the third leaf stage (two leaves and one heart). Genomic DNA was extracted using the 2% CTAB method. The DNA from the four inbred lines was mixed together in equal mass to prepare a mixed DNA sample for testing.
[0069] The specific steps include: taking an appropriate amount of leaves and placing them into a 1.5 mL sterile centrifuge tube, grinding them in a mortar and pestle under liquid nitrogen freezing conditions; quickly adding 750 μL of 2% CTAB extraction solution and rapidly shaking to mix; incubating in a 65℃ water bath for 45 min, gently shaking several times every 10 min during the water bath; removing the centrifuge tube, adding 750 μL of chloroform:isoamyl alcohol (24:1), and gently shaking to mix for a moment; centrifuging at 8000 r / min for 10 min at room temperature, and transferring approximately 500 μL of supernatant to a new 1.5 mL centrifuge tube; adding 500 μL of pre-cooled isopropanol (-20℃), gently shaking to mix, and letting stand for a moment; centrifuging at 12000 r / min for 2 min at room temperature, and discarding the supernatant; washing with 500 μL of 75% alcohol, centrifuging at 12000 r / min for 1 min at room temperature, and discarding the supernatant. Repeat once; after air drying, add 100 μL ddH2O to fully dissolve the DNA; mix the DNA from the four inbred lines together in equal mass to prepare a mixed DNA sample for testing.
[0070] (2) Bisulfite treatment.
[0071] The mixed DNA sample to be tested was purified by bisulfite treatment using the EZ DNA Methylation-Lightning Kit (Zymo, D5031) to obtain the purified mixed DNA sample. The treatment steps included:
[0072] 1) Add 130 μL of Zymo Lightning Conversion Reagent to 20 μL (approximately 1 μg) of the DNA sample to be tested. Mix thoroughly by pipetting up and down at least 10 times.
[0073] 2) Transfer 75 μL of the total 150 μL sample from step 1) above into two 200 μL PCR tubes, and perform the PCR in the PCR instrument according to the procedure shown in Table 1:
[0074] Table 1
[0075] temperature time 98℃ 8min 54℃ 60min 4℃ hold
[0076] 3) Add 600 μL of M-Binding Buffer to the Zymo Spin IC column. Transfer the two 75 μL samples from step 2) to the Zymo Spin IC column containing M-Binding Buffer and mix by inverting the container at least 10 times. Centrifuge at 13000 rpm for 30 seconds and discard the flow-through liquid.
[0077] 4) Add 100 μL of M-Wash buffer and centrifuge at 13000 rpm for 30 seconds.
[0078] 5) Add 200 μL L-DesμLfonation Buffer and let stand at room temperature for 20 min. Centrifuge at 13000 rpm for 30 s and discard the flow-through liquid.
[0079] 6) Add 200 μL of M-Wash buffer. Centrifuge at 13000 rpm for 30 seconds and discard the flow-through fluid.
[0080] 7) Repeat step 6) once.
[0081] 8) Place the Zymo Spin IC column onto a new 1.5 mL centrifuge tube, add 21 μL ddH2O, incubate at room temperature for 2 min, centrifuge at 13000 rpm for 30 s, and collect the flow-through liquid, which is the purified maize genome sample. Mix the maize genome samples from four common inbred lines (B73, Mo17, W22, and SK) to obtain the DNA mixed sample.
[0082] S2, First Amplification
[0083] In this embodiment, the first amplification uses upstream and first downstream primers to perform PCR amplification of the target segment to obtain a first amplification product. The first amplification product has a first recognition sequence and a bridging sequence to identify each DNA molecule of the target segment in the mixed DNA sample. For example... Figure 2 As shown, the SNP site in the first amplification product obtained by first amplification using upstream primer and first downstream primer is located within 150 bp at its 3' end.
[0084] A specific example of step S2 is as follows:
[0085] (1) Selection of target section
[0086] The target region applicable in this application should contain at least one or more SNP sites that can distinguish different maize genomes. The number of SNP sites in the target region is related to the number of maize genomes that can be mixed into a DNA mixture sample at one time by the method provided in the embodiments of this application. For example, if there is one SNP site in the target region and the polymorphism of the SNP site is two, then two maize genome extraction samples can be mixed to make a DNA mixture sample, and so on.
[0087] In a specific embodiment, regions 1 (1:83554963-83555226) and 2 (7:46311503-46311767) of the maize B73 RefGen_v4 genome were used as target regions. Both regions contain multiple SNP sites that can distinguish the four genomic sequence information. For example... Figure 3 As shown, the four common maize inbred lines B73, Mo17, W22, and SK, which are to be tested for DNA methylation, all contain these two target regions. The red dots in the figure represent SNP sites.
[0088] (2) Primer design
[0089] For region 1 (1:83554963-83555226) and region 2 (7:46311503-46311767), upstream primers and a first downstream primer were designed. The C base in the upstream primer sequence was designed as a degenerate Y base; the G base in the first downstream primer sequence was designed as a degenerate R base. The first downstream primer contains a first recognition sequence and a bridging sequence.
[0090] In some embodiments, the first identification sequence is a 6-base sequence, such as CCCCC, TTTTTT, GGGGGG, AAAAAA, AGAGGG, CCCGGG, etc., as shown in the underlined sequences in Table 2. The first identification sequence is used to identify each DNA molecule in the target region of the DNA mixture sample.
[0091] In some embodiments, the bridging sequence serves to connect the first and second recognition sequences in the final second amplification product. Specifically, the length of the bridging sequence is not limited; it only needs to connect and reduce interference with the target region sequence. Extending the bridging sequence can increase the stability of the second PCR amplification experiment, while shortening the bridging sequence can reduce primer synthesis costs. For example, it can be a short sequence of 9–18 bp (such as the bolded sequence in Table 2). In some embodiments, the nucleotide sequence of the bridging sequence is shown in Table 2.
[0092] Table 2
[0093]
[0094] (3) First amplification reaction
[0095] In some embodiments, multiple target segments are set, and primers can be designed for each of the multiple target segments. Each target segment can be used with a pair of primers for a first amplification reaction; alternatively, primers for multiple target segments can be mixed and only a first amplification reaction can be performed.
[0096] In one specific embodiment, the upstream primer and the first downstream primer designed for segment 1 and segment 2 are mixed and simultaneously added to the DNA mixed sample "transformed" in step S1 as a template, and a first amplification PCR amplification reaction is performed to obtain the first amplification product.
[0097] In some embodiments, the reaction system for the first amplification reaction is shown in Table 3. In Table 3, the upstream primer can be a mixture of upstream primers designed for multiple target segments, and the first downstream primer can be a mixture of first downstream primers designed for multiple target segments.
[0098] Table 3 First Amplification PCR Reaction System
[0099]
[0100] To increase the specificity of amplification, some embodiments provide a first amplification reaction step including:
[0101] S21. Perform pre-denaturation at a temperature of 95℃ for 5 minutes.
[0102] S22. Perform the first amplification cycle, which includes sequentially performing the first melting process, the first annealing process, and the first extension process.
[0103] S23. Perform the second amplification cycle, which includes performing the second melting process, the second annealing process, and the second extension process in sequence.
[0104] S24, 72℃ treatment for 5 min and 12℃ treatment for 1 s.
[0105] In some embodiments of step S22, the first amplification cycle includes at least 8 to 12 cycles in order to sufficiently increase the concentration of the substrate chain and improve the synthesis efficiency of the first amplification product.
[0106] In some embodiments of step S22, the conditions for the first unwinding process are 95°C and a processing time of 30 seconds.
[0107] In some embodiments of step S22, the processing temperature in the first annealing treatment gradually decreases with each increase in the number of cycles of the first amplification cycle. This not only sufficiently increases the concentration of the substrate chain and improves the synthesis efficiency of the first amplification product, but also enhances the specificity of the amplification and reduces the probability that the first amplification product does not carry the target fragment of the first recognition sequence.
[0108] In some embodiments of step S22, the processing temperature of the first annealing treatment is 68-55°C, 65-53°C, 65-55°C, or 63-55°C. For example, if the processing temperature of the first annealing treatment is 65-55°C, then the first amplification cycle is performed 11 times. The processing temperature in the first annealing treatment decreases gradually with the increase of the number of cycles in the first amplification cycle. For example, the processing temperature of the first annealing treatment in the first amplification cycle is 65°C, and the processing temperature of the first annealing treatment in the second amplification cycle is 64°C.
[0109] In some embodiments of step S23, the second amplification cycle includes at least 20 to 24 cycles to increase the concentration of the first amplification product.
[0110] In a specific embodiment of step S23, 24 second amplification cycles are performed, each second amplification cycle including sequential treatment at 95°C for 30 seconds, treatment at 60°C for 30 seconds, and treatment at 70°C for 30 seconds.
[0111] Table 4 shows a specific first amplification reaction step for S21 to S24.
[0112]
[0113]
[0114] (4) Purification of the first amplification product
[0115] In this step, well-known techniques can be used to purify the first amplification product. For example, such as... Figure 6 As shown, the first amplification product was purified by magnetic beads. 2.2 times the volume of the PCR product was purified by Beckman magnetic beads, and finally eluted with 30 μL ddH2O to obtain the first amplification product.
[0116] S3, Second Amplification
[0117] In this embodiment, the upstream and downstream primers are used to amplify the first amplification product as a template to obtain a second amplification product. The second amplification product includes the first amplification product and a second recognition sequence attached to the 3' end of the first amplification product. The second recognition sequence is used to identify the first amplification product.
[0118] In some embodiments, the first downstream primer includes a bridge sequence to form a structure with a first recognition sequence, a bridge sequence, and a second recognition sequence at the 3' end of the second amplification product, and the second recognition sequence is used to identify the first amplification product during sequencing. For example, the nucleotide sequence of the second downstream primer is shown in Table 5. In Table 5, the bolded sequence is the second recognition sequence, and the remaining part is the bridge sequence.
[0119] Table 5 Artificial Sequences
[0120]
[0121]
[0122] In some embodiments, if multiple first amplification reactions are performed on multiple maize genomes to obtain multiple first amplification products, then a second recognition sequence is designed for each amplification product, thereby obtaining a corresponding number of second amplification products.
[0123] The reaction system for a specific second amplification reaction is shown in Tables 6 and 7.
[0124] Some embodiments provide a step in a second amplification reaction, including:
[0125] S31. Perform pre-denaturation at a temperature of 95℃ for 5 minutes.
[0126] S32. Perform an amplification cycle, wherein the amplification cycle includes sequentially performing melting, annealing, and extension processes; and
[0127] S33, 72℃ treatment for 5 min and 12℃ treatment for 1 s.
[0128] In the embodiment of step S32, the conditions for the melting treatment are 95°C for 30 seconds, the conditions for the annealing treatment are 60°C for 30 seconds, and the conditions for the extension treatment are 72°C for 30 seconds. In this way, the specific recognition of the second amplification product by the first downstream primer can be guaranteed, thus providing the specificity of the amplification.
[0129] Table 6 Second Amplification PCR Reaction System
[0130]
[0131] Table 7. A specific S31-S33 second amplification PCR reaction procedure.
[0132]
[0133]
[0134] (4) Purification of the second amplification product
[0135] In this step, well-known techniques can be used to purify the second amplification product. For example, such as... Figure 7 As shown, the product of the second amplification was purified by magnetic beads. 2.2 times the volume of the PCR product was purified by Beckman magnetic beads, and finally eluted with 30 μL ddH2O to obtain the second amplification product.
[0136] Therefore, this application also discloses a kit for constructing a maize genomic DNA methylation detection library containing a target region, wherein the target region contains at least one or more SNP sites capable of distinguishing different maize genomes, and the DNA methylation detection library is used to detect mixed samples of DNA from at least two maize genomes containing the target region, wherein the kit comprises:
[0137] A maize genomic DNA processing reagent for extracting the maize genomic DNA and converting unmethylated cytosine into uracil;
[0138] An upstream primer, wherein the C base in the upstream primer sequence is replaced with a degenerate Y base sequence;
[0139] A first downstream primer, wherein the G base in the downstream primer sequence is replaced with a degenerate R base, and the first downstream primer includes a first recognition sequence and a bridging sequence; and
[0140] The second downstream primer comprises the bridge sequence and the second recognition sequence;
[0141] The upstream primer and the first downstream primer are used to perform a first amplification to obtain a first amplification product. The first amplification product contains a DNA molecule having the target segment and a first recognition sequence and a bridge sequence attached to the 3' end of the DNA molecule.
[0142] The upstream primer and the second downstream primer can be used to obtain a second amplification product in a second amplification. The second amplification product includes the first amplification product and a second recognition sequence attached to the 3' end of the first amplification product to construct the DNA methylation detection library.
[0143] The first identification sequence is used to identify each DNA molecule in the target region of the DNA mixture sample, and the second identification sequence is used to identify the first amplification product.
[0144] Detection method for maize genomic DNA methylation containing target regions
[0145] In addition, embodiments of this application also disclose a method for detecting maize genomic DNA methylation including a target region, comprising:
[0146] Steps S1 to S3 are used to construct the maize genomic DNA methylation detection library containing the target region;
[0147] S5. Sequencing the DNA methylation detection library to obtain a first reads library;
[0148] S6. Perform quality control, splitting and reassembling on the first reads library to obtain the second reads library;
[0149] S7. The second reads library is compared, deduplicated, and calculated to obtain the DNA methylation information of each maize genome in the DNA mixed sample.
[0150] In the embodiment of step S5, second-generation high-throughput sequencing technology is used to sequence the constructed DNA methylation detection library. The first reads library includes a number of first reads corresponding to the DNA methylation detection library, where "reads" refers to a short sequencing fragment, which is the raw sequencing data generated by the second-generation high-throughput sequencer. The second reads library includes second reads based on the same second identification sequence, wherein the second reads are reads with the first identification sequence, the second identification sequence, and the bridge sequence removed.
[0151] In an embodiment of step S6, it specifically includes S61 to S63.
[0152] S61. Obtain the reference sequence of the target region. For example, based on the above-mentioned regions 1 and 2, obtain the sequence information in the genomes of four maize inbred lines B73, Mo17, W22, and SK.
[0153] S62. Perform quality control on the first reads library, removing low-quality first reads. Specifically, this can be done based on the commonly used sequencing quality threshold Q20.
[0154] S63. The first reads in the first reads library after quality control are split and grouped to obtain the second reads library. The term "split" means separating the first reads with different second identification sequences, and the term "group" means combining the first reads with the same second identification sequence; for example, if there are multiple second identification sequences, a corresponding number of first read groups can be obtained.
[0155] In some embodiments of step S63, as shown in Table 5, a total of 17 second identification sequences were designed, thereby generating 17 first reads groups. Each specific first reads group is compiled into a group fq file, and the first identification sequence, bridge sequence and second identification sequence of each first read are removed to obtain the second reads library; at the same time, the data of the first identification sequence corresponding to each first read is output.
[0156] In a specific embodiment of step S7, aligning the second reads library includes aligning each of the second reads with a reference sequence, allowing for mismatches at all sites except the C site (where mismatches between C and T are permitted). The reference sequence is a maize genome sequence containing the target region. For example, using sequences 1:83554963-83555226 and 7:46311503-46311767 from four maize inbred lines (B73, Mo17, W22, SK) as reference sequences, the reads are organized into fa format. The Bsmap alignment software is used to perform a strict, mismatch-free alignment (-v 0) between the split reads and the organized reference sequence, resulting in an aligned bam file. An index is then built using samtools software. Note that this mismatch-free alignment method is mandatory. For the new fq format file after processing the second identification sequence, which contains reads from a mixture of four genome samples, the reads can be assigned to the corresponding genomes during the alignment process based on SNPs between different genomes.
[0157] In a specific embodiment of step S7, deduplication of the second reads library includes removing duplicate reads caused by PCR amplification from the aligned reads. For example, based on the first recognition sequence, duplicate data caused by PCR amplification in the preliminary DNA methylation alignment information is removed to obtain deduplicated DNA methylation alignment information.
[0158] In a specific embodiment of step S7, the calculation of the second reads library includes comparing the number of C and T bases paired at each C site in the target region of each sample to count the number of duplicate reads on the pair, and calculating the DNA methylation level at each C site: number of C bases at each C site / (number of C bases at each C site + number of T bases at each C site).
[0159] A specific result obtained through the detection methods in steps S1 to S7 is as follows: Figure 4As shown. Additionally, whole-genome DNA methylation sequencing data of leaf tissues from known maize inbred lines B73 and Mo17 were obtained from the NCBI public database (https: / / www.ncbi.nlm.nih.gov / geo / query / acc.cgi?acc=GSE128859). The DNA methylation levels of target regions 1 and 2 obtained by the conventional whole-genome DNA methylation sequencing method were statistically analyzed, as shown... Figure 5 .contrast Figure 4 , 5 It is evident that the DNA methylation levels of target regions 1 and 2 of B73 and Mo17, as determined by the method of this application, are consistent with the methylation levels of target regions detected by conventional methods.
[0160] As described above, this application also substantially discloses a system for detecting maize genomic DNA methylation containing a target region. The target region contains at least one or more SNP sites capable of distinguishing different maize genomes. A DNA methylation detection library is used to detect mixed DNA samples containing at least two maize genomes. The system includes:
[0161] The kit provided in the foregoing embodiments is used to obtain DNA methylation information for each maize genome in a mixed DNA sample; and
[0162] A processing device is used to run a program, which, when running, executes the detection method provided in the above embodiments, "sequencing the DNA methylation detection library to obtain a first reads library."
[0163] The first reads library is subjected to quality control, splitting, and reassembly to obtain the second reads library;
[0164] The step is to compare, remove duplicates, and perform calculations on the second reads library to obtain the DNA methylation information of each maize genome in the DNA mixture sample.
[0165] The above description is merely a preferred embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. sequence list <110> Huazhong Agricultural University <120> DNA methylation detection of target regions in the maize genome <160> twenty one <170> SIPOSequenceListing 1.0 <210> 1 <211> 21 <212> DNA <213> Artificial Sequence <400> 1 gtatyggtgg ygtgtggaat g <210> 2 <211> 21 <212> DNA <213> Artificial Sequence <400> 2 gaagtggtga yyagyagtgt g <210> 3 <211> 39 <212> DNA <213> Artificial Sequence <220> <221> misc_feature <222> (10)..(15) <223> n is a, c, g, or t <400> 3 atagcgacgn nnnnncccca attractive carcaactc <210> 4 <211> 38 <212> DNA <213> Artificial Sequence <220> <221> misc_feature <222> (10)..(15) <223> n is a, c, g, or t <400> 4 atagcgn nnnnncttca ctraccttcc aartcctc <210> 5 <211> 17 <212> DNA <213> Artificial Sequence <400> 5 atcacgttat agcgacg 17 <210> 6 <211> 17 <212> DNA <213> Artificial Sequence <400> 6 cgatgtttat agcgacg 17 <210> 7 <211> 17 <212> DNA <213> Artificial Sequence <400> 7 ttaggcatat agcgacg 17 <210> 8 <211> 17 <212> DNA <213> Artificial Sequence <400> 8 tgaccactat agcgacg 17 <210> 9 <211> 17 <212> DNA <213> Artificial Sequence <400> 9 acagtggtat agcgacg 17 <210> 10 <211> 17 <212> DNA <213> Artificial Sequence <400> 10 gccaatgtat agcgacg 17 <210> 11 <211> 17 <212> DNA <213> Artificial Sequence <400> 11 cagatctgat agcgacg 17 <210> 12 <211> 17 <212> DNA <213> Artificial Sequence <400> 12 acttgatgat agcgacg 17 <210> 13 <211> 17 <212> DNA <213> Artificial Sequence <400> 13 gatcagcgat agcgacg 17 <210> 14 <211> 17 <212> DNA <213> Artificial Sequence <400> 14 tagcttgtat agcgacg 17 <210> 15 <211> 17 <212> DNA <213> Artificial Sequence <400> 15 ggctacagat agcgacg 17 <210> 16 <211> 17 <212> DNA <213> Artificial Sequence <400> 16 cttgtactat agcgacg 17 <210> 17 <211> 17 <212> DNA <213> Artificial Sequence <400> 17 tggttgttat agcgacg 17 <210> 18 <211> 17 <212> DNA <213> Artificial Sequence <400> 18 tctcggttat agcgacg 17 <210> 19 <211> 17 <212> DNA <213> Artificial Sequence <400> 19 taagcgttat agcgacg 17 <210> 20 <211> 17 <212> DNA <213> Artificial Sequence <400> 20 tccgtcttat agcgacg 17 <210> twenty one <211> 17 <212> DNA <213> Artificial Sequence <400> twenty one ttctgtgtat agcgacg 17
Claims
1. A method for detecting DNA methylation of a target region in a maize genome, wherein the target region contains at least one or more SNP sites capable of distinguishing different maize genomes, and the method for detecting DNA methylation of the target region is used to detect a mixed sample of DNA from at least two maize genomes containing the target region, wherein... The detection method includes the following steps: The unmethylated cytosine in the DNA molecules of the DNA mixture sample is converted into uracil; The first amplification is performed using an upstream primer and a first downstream primer. The C base in the upstream primer is designed as a degenerate base Y; the G base in the first downstream primer is designed as a degenerate base R. The first downstream primer contains a first recognition sequence and a bridge sequence. The first amplification includes amplifying the DNA mixture sample in the same PCR reaction system to obtain a first amplification product; the first amplification product contains a DNA molecule having the target region and a first recognition sequence and a bridging sequence attached to the 3' end of the DNA molecule; the first recognition sequence is used to identify each DNA molecule of the target region in the DNA mixture sample; The SNP site in the first amplification product is located within 150 bp of the binding start position of the DNA molecule of the target region with the upstream primer or with the first downstream primer. A second amplification is performed, which includes the step of performing PCR amplification on the first amplification product to obtain a second amplification product. The second amplification product includes the first amplification product and a second recognition sequence attached to the 3' end of the first amplification product. The second recognition sequence is used to identify the first amplification product. The second amplification uses the upstream primer and the second downstream primer, the second downstream primer containing the second recognition sequence and the bridge sequence; A DNA methylation detection library containing the target region of the maize genome was constructed using the second amplification product; The DNA methylation detection library was subjected to second-generation high-throughput sequencing to obtain a first reads library; The first reads library is subjected to quality control, splitting, and reassembly to obtain the second reads library; By comparing, deduplicating, and calculating the second reads library, the methylation level information of C bases in the three sequence environments of CG, CHG, and CHH for each maize genome in the DNA mixture sample is obtained.
2. The detection method according to claim 1, wherein, The first reads library includes a number of first reads corresponding to the DNA methylation detection library, and the second reads library includes second reads composed of the same second recognition sequence, wherein the second reads are reads in which the first recognition sequence, the second recognition sequence and the bridge sequence have been removed.
3. Application of a maize genome target region DNA methylation detection kit The target region contains at least one or more SNP sites that can distinguish different maize genomes. The kit is used to detect DNA mixture samples containing at least two maize genomes. The kit is used to obtain information on the methylation level of C bases in the CG, CHG, and CHH sequence environments of each maize genome in the DNA mixture sample; in, The kit includes: A maize genomic DNA processing reagent for extracting the maize genomic DNA and converting unmethylated cytosine into uracil; An upstream primer, wherein the C base in the upstream primer sequence is replaced with a degenerate Y base sequence; A first downstream primer, wherein the G base in the first downstream primer sequence is replaced with a degenerate R base, the first downstream primer comprising a first recognition sequence and a bridge sequence; and a second downstream primer, the second downstream primer comprising the bridge sequence and the second recognition sequence; The upstream primer and the first downstream primer are used to perform a first amplification to obtain a first amplification product. The first amplification product contains a DNA molecule having the target segment and a first recognition sequence and a bridge sequence attached to the 3' end of the DNA molecule. The first identification sequence is used to identify each DNA molecule in the target region of the DNA mixture sample; the SNP site in the first amplification product is located within 150 bp of the binding start position of the DNA molecule with the target region with the upstream primer or with the binding start position of the first downstream primer; The upstream primer and the second downstream primer are used for the second amplification, which includes the step of performing PCR amplification on the first amplification product to obtain the second amplification product. The second amplification product contains the first amplification product and a second recognition sequence linked to the 3' end of the first amplification product to construct a DNA methylation detection library. The second identification sequence is used to identify the first amplification product.