Simplified genome sequencing library construction method and kit and related application thereof

By using restriction endonuclease enzyme fragments themselves as barcode sequences in simplified genome sequencing technology to distinguish samples (biRAD-seq), the problems of inconsistent sample distinction, sequencing depth and cumbersome experimental operation procedures in the prior art are solved, and more uniform sequencing depth and more efficient reagent use are achieved.

CN120210964APending Publication Date: 2025-06-27HUAZHONG AGRI UNIV +1
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510269622.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-07
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The existing simplified genome sequencing technology has shortcomings in sample distinction, inconsistent sequencing depth caused by PCR amplification, and cumbersome experimental operation procedures, which affect the output quality and stability of sequencing data.

Method used

A simplified genome sequencing method (biRAD-seq) based on the restriction enzyme fragment itself as the barcode sequence to distinguish different samples. This method does not require PCR amplification. DNA samples are enzymatically cut, mixed and sorted fragments to construct a PCR free sequencing library.

Benefits of technology

Under the same data volume, more RAD-seq tags with more uniform depth are obtained, which solves the problems of sample distinction, inconsistent sequencing depth and cumbersome experimental operation procedures, and does not require design of any connectors, saving reagent costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0005302276830000161
    Figure BDA0005302276830000161
  • Figure BDA0005302276830000171
    Figure BDA0005302276830000171
  • Figure HDA0005302276840000011
    Figure HDA0005302276840000011
Patent Text Reader

Abstract

The invention provides a simplified genome sequencing library construction method and a kit and related application thereof. The construction method of the simplified genome sequencing library comprises the following steps: (1) respectively carrying out enzyme digestion on a plurality of DNA (Deoxyribonucleic Acid) samples by using restriction endonuclease; wherein restriction endonucleases adopted by all the samples are different from one another, and the tail ends of specific DNA fragments of the samples are generated to serve as sample bar codes of the samples; (2) mixing the enzyme digestion products of the samples, constructing a DNA library, carrying out fragment sorting on the DNA library, and taking the selected library as a simplified genome sequencing library; or carrying out fragment sorting on the enzyme digestion product of each sample, and carrying out DNA library construction by using the selected fragments to obtain the simplified genome sequencing library. According to the method, the experimental operation process is simplified, the library building efficiency is improved, and meanwhile, the feasibility of the restriction endonuclease for recognizing 6 basic groups based on a third-generation long-read-long sequencing platform in the simplified genome sequencing technology is also explored.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for constructing a reduced-representation genome sequencing library, a sequencing method, a kit used, and related applications, belonging to the technical field of genetic engineering. Background Art

[0002] Whole Genome Genotyping (WGG) is an important technology in biological research, which is of great significance for analyzing the molecular mechanisms of biological phenomena. At the same time, it has wide applications in research fields such as genetic breeding of agricultural animals and plants, disease treatment, biological evolution, biodiversity conservation, and inspection and quarantine.

[0003] Currently, in addition to the WGG method based on DNA chips (DNA array; Rasheed et al., 2017), there are also various WGG methods based on next-generation sequencing (NGS) platforms, such as whole genome re-sequencing, different types of reduced-representation genome sequencing (RRGS; or restriction-site associated DNA sequencing, RAD-seq, Andrews et al., 2016), genome targeted capture (Kozarewa et al., 2015), multiplex PCR (Onda et al., 2018), MRA-seq (Multiplex Restriction Amplicon Sequencing, Bernardo et al., 2020), FBI-seq (Zhao et al., 2023), TAIL-peq (Zhao et al., 2024), etc.

[0004] Compared with whole-genome resequencing, RAD-seq can obtain genome-wide genetic information at a lower sequencing cost. Since it was first proposed in 2008, RAD-seq has been widely used in the fields of biology, medicine, and agronomy. At the same time, people have been developing simpler and faster RAD-seq library construction methods to simplify the experimental process of library construction, improve the throughput of library construction, and thus reduce the cost of library construction. Currently, commonly used reduced-representation genome sequencing technologies include RAD-seq (Restriction-site-associated DNA sequencing), GBS-seq (Genotyping by sequencing), ddRAD-seq (Doubledigest restriction-site associated DNA sequencing), and 2bRAD-seq, etc. Among them, in the RAD-seq technology, the genome is first digested by a restriction endonuclease to break the DNA strand and generate sticky ends; subsequently, an artificially synthesized DNA adapter is ligated, and the adapter carries three DNA elements: the element at the 5' end is used for subsequent sequencing; the middle element is a barcode to distinguish different samples; the element at the 3' end pairs with the sticky end of the restriction endonuclease. Since each sample is ligated with a different barcode, 8 to 12 samples can be mixed together for subsequent processing to improve the throughput of library construction and save reagent costs; then, the mixed samples are randomly fragmented by sonication; subsequently, another artificially synthesized DNA adapter is ligated using T4 DNA ligase for subsequent sequencing; finally, the lengths of the DNA fragments are screened by agarose gel electrophoresis technology, and suitable fragments are selected for PCR amplification and high-throughput sequencing. In the GBS-seq technology, the genome is first fragmented by a restriction endonuclease, and then two different sequences of adapters are added using T4 DNA ligase. One adapter contains three elements similar to those in RAD-seq. The sequence of the 3' end of the other adapter that pairs with the sticky end is the same as that of the first adapter, but it does not contain the barcode element, and the sequence of the 5' end is also different from that of the first adapter. Only the DNA fragments ligated with different adapters at both ends can be PCR amplified and sequenced. The ddRAD technology uses two restriction endonucleases to digest genomic DNA. Subsequently, two artificially synthesized DNA adapters are ligated, and the element composition of the adapters is similar to that of RAD-seq; next, the lengths of the DNA fragments are screened by agarose gel electrophoresis technology, and DNA fragments of appropriate lengths are selected and PCR amplified. Only the DNA fragments ligated with different adapters at both ends can be effectively amplified and sequenced.The 2bRAD technique uses type IIB restriction endonucleases to cut out DNA fragments of consistent length from the genome, and T4 DNA ligase is used to add synthetic DNA adapters to both ends of the fragments. Subsequently, PCR amplification and sequencing are carried out. Since the cleavage sites of the type IIB restriction endonucleases used in 2bRAD-seq are located approximately 10 bp on both sides of the recognition site, unlike conventional restriction endonucleases that produce sticky ends with fixed sequences, the sticky end sequences produced by type IIB restriction endonucleases are not fixed. Correspondingly, the sequence element that pairs with the sticky end at the 3' end of the synthetic adapter added is composed of NNN, and the 5' end and the middle elements are similar to those in conventional RAD-seq.

[0005] The main differences among these reduced-representation genome sequencing techniques lie in the types of restriction endonucleases selected and the time points for adding sequencing adapters and sorting the fragment lengths during the library construction process. However, the principles of reducing the genome are all based on the enzymatic cleavage of restriction endonucleases, removing some genomic sequences, ligating the sequence elements required during the sequencing of the remaining fragments, and then enriching fragments approximately 300 bp near the cleavage sites for sequencing. To distinguish different samples and improve the throughput of library construction, barcode sequences are introduced through synthetic DNA adapters in these methods.

[0006] The above reduced-representation genome sequencing techniques still have deficiencies: First, DNA adapters with different barcodes need to be added through DNA ligase to distinguish different samples, and the experimental operation steps are cumbersome, affecting the efficiency of library construction; Second, due to random errors or the nature of the DNA fragments themselves, the cleavage efficiencies of restriction endonucleases at different sites on the genome and the ligation efficiencies of the adapters are different, resulting in inconsistent effective template numbers at different sites before PCR amplification. This inconsistency is amplified again after a dozen PCR cycles, causing a huge difference in the sequencing depths at different sites. Some sites have too high a sequencing depth, resulting in a waste of data, while some sites are not sequenced, resulting in data loss at these sites, affecting the output quality and stability of the sequencing data. Summary of the Invention

[0007] An object of the present invention is to provide a new method for constructing a reduced-representation genome sequencing library to simplify the operation and improve the library construction efficiency.

[0008] Another object of the present invention is to provide related applications of the method for constructing a reduced-representation genome sequencing library.

[0009] Another object of the present invention is to provide a reduced-representation genome sequencing method.

[0010] Another object of the present invention is to provide a kit for constructing a reduced representation genomic sequencing library and / or for reduced representation genomic sequencing.

[0011] Another object of the present invention is to provide related applications of the kit.

[0012] On the one hand, the present invention provides a new reduced representation genomic sequencing technology. The reduced representation genomic sequencing technology of the present invention is a reduced representation genomic sequencing method (Barcode itself restriction-site-associated DNA sequencing, biRAD-seq) based on the restriction enzyme digested fragments themselves as barcode sequences to distinguish different samples without PCR (PCR free). By digesting different DNA samples with different restriction enzymes, then mixing and sorting the fragments, and then constructing a PCR free sequencing library. The method of the present invention can obtain more and more evenly deep RAD-seq tags (tags) under the same amount of data. The biRAD-seq of the present invention can solve the defects of the existing RAD-seq that sample differentiation depends on adding additional Barcodes, the inconsistent sequencing depth at different sites caused by PCR amplification, and the cumbersome experimental operation process. In addition, the biRAD-seq of the present invention does not require the design of any adapters, which also saves reagent costs.

[0013] Specifically, the present invention provides a method for constructing a reduced representation genomic sequencing library, the method comprising:

[0014] (1) Digesting multiple DNA samples with restriction enzymes respectively; wherein, different restriction enzymes are used for each sample to generate sample-specific DNA fragment ends as its own sample barcode;

[0015] (2) Mixing the digestion products of each sample, constructing a DNA sequencing library, sorting the fragments of the DNA sequencing library, and selecting the library as the reduced representation genomic sequencing library; or,

[0016] Sorting the fragments of the digestion products of each sample, constructing a DNA sequencing library with the selected fragments, and obtaining the reduced representation genomic sequencing library.

[0017] In the method for constructing a reduced representation genomic sequencing library of the present invention, since different restriction enzymes are used for each sample, sample-specific DNA fragment ends can be generated as its own sample barcode to distinguish itself from other samples. The method of the present invention can simplify the operation process and improve the library construction efficiency.

[0018] According to a specific embodiment of the present invention, in the method for constructing a reduced-representation genomic sequencing library of the present invention, the multiple DNA samples are plant or animal samples from different individuals.

[0019] According to a specific embodiment of the present invention, in the method for constructing a reduced-representation genomic sequencing library of the present invention, the number of the multiple DNA samples is 2 - 24, preferably 4 - 18, and more preferably 6 - 16. In some specific embodiments of the present invention, the number of the multiple DNA samples is 6, 7, 8, 9, 10, 11, 12, 13, 14, 15 or 16 samples, and each sample is a plant or animal sample from a different individual.

[0020] According to a specific embodiment of the present invention, in the method for constructing a reduced-representation genomic sequencing library of the present invention, in a single DNA sample, the DNA concentration is greater than or equal to 40 ng / μL.

[0021] According to a specific embodiment of the present invention, in the method for constructing a reduced-representation genomic sequencing library of the present invention, in a single DNA sample, the total amount of DNA is greater than or equal to 1 μg.

[0022] According to a specific embodiment of the present invention, in the method for constructing a reduced-representation genomic sequencing library of the present invention, the use of restriction endonucleases for DNA samples is mainly to simplify the genome. In the present invention, generally, digestion is carried out to reach a simplification level of 65% - 97%, preferably reaching a simplification level of 75% - 96%.

[0023] In the present invention, the simplification level refers to the proportion of the genomic part that is not sequenced in the total genome, and its definition is: Simplification level (P 简化水平 ) = (1 - total length of the target fragments selected for sequencing (bp) / total length of the genome (bp)) × 100%.

[0024] According to a specific embodiment of the present invention, in the method for constructing a reduced-representation genomic sequencing library of the present invention, any feasible restriction endonuclease can be used to digest the genomic DNA of the resequencing library to achieve the expected simplification level. According to a specific embodiment of the present invention, the simplification level can be controlled by selecting the specific type of restriction endonuclease and controlling the digestion conditions.

[0025] According to a specific embodiment of the present invention, in the method for constructing a reduced-representation genomic sequencing library of the present invention, the restriction endonucleases used include enzymes that produce 5'-sticky ends and / or enzymes that produce 3'-sticky ends. The present invention can distinguish samples based on the recognition site sequence of the restriction endonuclease (applicable to enzymes that produce 5'-sticky ends) or its adjacent sequence (applicable to enzymes that produce 3'-sticky ends).

[0026] According to the specific embodiments of the present invention, in the method for constructing a reduced-representation genomic sequencing library of the present invention, the restriction endonuclease used is a restriction endonuclease that recognizes a 4- to 6-base sequence.

[0027] The method of the present invention can be used to construct second-generation sequencing libraries or third-generation sequencing libraries. When used to construct second-generation sequencing libraries, restriction endonucleases that recognize 4 bases are usually used. After the genomic DNA is digested by restriction endonucleases that recognize 4 bases, the generated fragments are relatively short, generally about 100 bp to 1000 bp; about 10% of the target-length fragments suitable for constructing second-generation sequencing libraries can be sorted out. When used to construct third-generation sequencing libraries, restriction endonucleases that recognize 5 or 6 bases are usually used for digestion. After the genomic DNA is digested by restriction endonucleases that recognize 5 or 6 bases, the generated fragments are relatively long, generally about 1 to 20 kb, and about 10% of the target-length fragments suitable for constructing third-generation sequencing libraries can be sorted out.

[0028] According to the specific embodiments of the present invention, in the method for constructing a reduced-representation genomic sequencing library of the present invention, the restriction endonucleases used include, but are not limited to, at least two of Aci I, Bfa I, CviQ I, Fat I, HinP1 I, HpyCH4 IV, MboI, MluC I, Mse I, Msp I, Taq I, Alu I, Dde I, Fnu4H I, Hinf I, HpyCH4 III, HpyCH4 V, RsaI, Sau96 I, StyD4 I, Hae III, Nla III, Hpy188 I, HpyCH4 III, Acl I, Afl II, Age I, ApaLI, BamH I, Bcl I, BseY I, BsiW I, BspD I, BspH I, BsrG I, BssS I, Eag I, EcoR I, HindIII, Kas I, Mfe I, Mlu I, Nco I, Nde I, Nhe I, Sal I, Spe I, Xma I.

[0029] According to the specific embodiments of the present invention, in the method for constructing a reduced representation genomic sequencing library of the present invention, the restriction endonucleases used include, but are not limited to, one or more of Bfa I, CviQ I, Fat I, HinP1 I, HpyCH4 IV, MboI, MluC I, Mse I, Msp I, Taq I, Alu I, Dde I, Fnu4H I, Hinf I, HpyCH4 III, HpyCH4 V, Rsa I. Preferably, the restriction endonucleases used include at least 6 of Bfa I, CviQ I, Fat I, HinP1 I, HpyCH4 IV, MboI, MluC I, Mse I, Msp I, Taq I, Alu I, Dde I, Fnu4H I, Hinf I, HpyCH4III, HpyCH4 V, Rsa I.

[0030] According to the specific embodiments of the present invention, in the restriction endonucleases used in the method for constructing a reduced representation genomic sequencing library of the present invention, the restriction endonucleases that recognize 4-base sequences include, but are not limited to, Aci I, Bfa I, CviQI, Fat I, HinP1 I, HpyCH4 IV, MboI, MluC I, Mse I, Msp I, Taq I, Alu I, Dde I, Fnu4HI, Hinf I, HpyCH4 III, HpyCH4 V, Rsa I, Sau96 I, StyD4 I, Hae III, Nla III, Hpy188I, HpyCH4 III; the restriction endonucleases that recognize 6-base sequences include, but are not limited to: Acl I, Afl II, Age I, ApaLI, BamH I, Bcl I, BseY I, BsiW I, BspD I, BspH I, BsrG I, BssS I, Eag I, EcoR I, HindIII, Kas I, Mfe I, Mlu I, Nco I, Nde I, Nhe I, Sal I, Spe I, Xma I.

[0031] In some specific embodiments of the present invention, the method for constructing a reduced-representation genomic sequencing library of the present invention is used to construct a next-generation sequencing library, and the restriction endonucleases used include at least 2, at least 4, at least 6, at least 8, at least 10, at least 12 or all of Aci I, Bfa I, CviQ I, Fat I, HinP1 I, HpyCH4 IV, MboI, MluC I, Mse I, Msp I, Taq I, Alu I, Dde I, Fnu4H I, Hinf I, HpyCH4III, HpyCH4 V, Rsa I, Sau96 I, StyD4 I, Hae III, Nla III, Hpy188 I, HpyCH4 III.

[0032] In some specific embodiments of the present invention, the method for constructing a reduced-representation genomic sequencing library of the present invention is used to construct a third-generation sequencing library, and the restriction endonucleases used include at least 2, at least 4, at least 6, at least 8, at least 10, at least 12 or all of Acl I, Afl II, Age I, ApaL I, BamH I, Bcl I, BseY I, BsiW I, BspD I, BspH I, BsrG I, BssS I, Eag I, EcoR I, Hind III, Kas I, Mfe I, Mlu I, Nco I, Nde I, Nhe I, Sal I, Spe I, Xma I. According to some specific embodiments of the present invention, in the method for constructing a reduced-representation genomic sequencing library of the present invention, for each DNA sample, the restriction endonuclease is selected from one of Bfa I, CviQ I, Fat I, HinP1 I, HpyCH4 IV, MboI, MluC I, Mse I, Msp I, Taq I, Alu I, Dde I, Fnu4H I, Hinf I, HpyCH4 III, HpyCH4 V, Rsa I.

[0033] According to a specific embodiment of the present invention, in the method for constructing a reduced-representation genomic sequencing library of the present invention, the fragment sorting is to select target fragments of 100-600 bp, preferably 250-400 bp for next-generation sequencing, or select fragments of 5-15 kb, preferably 7-11 kb for third-generation sequencing.

[0034] According to a specific embodiment of the present invention, in the method for constructing a reduced-representation genomic sequencing library of the present invention, the total length (bp) of the target fragments sorted for sequencing accounts for about 4%-20% of the total length of the genome, preferably about 5%-15% of the total length of the genome, and more preferably about 8%-12% of the total length of the genome.

[0035] According to specific embodiments of the present invention, in the method for constructing a reduced-representation genomic sequencing library of the present invention, when performing fragment sorting, any feasible existing technology can be used to sort fragments of the target length. For example, magnetic beads or agarose gel excision can be used to select the target DNA fragments, or an agarose gel electrophoresis system (such as the agarose gel electrophoresis system developed by Lifetechnology) or a series of instruments developed by Sage Science (SageBluePippin, Sage ELF, or Sage PippinHT) can be used for fragment selection.

[0036] According to specific embodiments of the present invention, in the method for constructing a reduced-representation genomic sequencing library of the present invention, the enzymatic digestion products of each sample are mixed in equal amounts.

[0037] According to specific embodiments of the present invention, the method for constructing a reduced-representation genomic sequencing library of the present invention further includes:

[0038] Detecting the concentration and fragment length distribution of the fragments of the DNA enzymatic digestion products;

[0039] Calculating the mixing ratio of the DNA enzymatic digestion fragments of each sample: According to the concentration of the enzymatic digestion fragments of each sample and the proportion of the target fragments in this sample, the concentration of the target fragments in the library is obtained, and this concentration information is the basis for mixing different DNA enzymatic digestion fragments; according to the concentration of the target fragments, the same amount of target DNA fragments is taken from the DNA enzymatic digestion fragments of each sample for mixing to obtain the finally mixed library.

[0040] According to specific embodiments of the present invention, in the method for constructing a reduced-representation genomic sequencing library of the present invention, the concentration of genomic DNA can be measured by a DNA concentration measurement method based on a fluorescent dye, and the DNA concentration measurement method based on a fluorescent dye is

[0041] According to specific embodiments of the present invention, in the method for constructing a reduced-representation genomic sequencing library of the present invention, the concentration of the DNA library is measured by a method; the DNA fragment length distribution is detected by Qsep100.

[0042] According to specific embodiments of the present invention, the method for constructing a reduced-representation genomic sequencing library of the present invention further includes the process of detecting and quantifying the DNA library. Specifically, this process includes: using Qsep100 to detect the actual fragment size of the selected library; the quantification of the library is performed by or Q-PCR for measurement.

[0043] According to a specific embodiment of the present invention, in the method for constructing a reduced-representation genomic sequencing library of the present invention, the sequencing library is a PCR-free library.

[0044] According to a specific embodiment of the present invention, for the method for constructing a reduced-representation genomic sequencing library of the present invention, the construction of a PCR-free Truseq library includes: adding DNA adapter sequences required for sequencing to both ends of the digested DNA fragments of the DNA to be tested to complete the construction of the DNA library. More specifically, the method of adding DNA adapter sequences required for sequencing to both ends of the DNA fragment to be tested is: performing end repair and adding A to the digested and fragmented genomic DNA, and then using a ligase to ligate the DNA adapter to the genomic DNA fragment to obtain the library.

[0045] On the other hand, the present invention also provides related applications of the method for constructing a reduced-representation genomic sequencing library described above. Specifically, it can be used for reduced-representation genomic sequencing.

[0046] On the other hand, the present invention also provides a reduced-representation genomic sequencing method, which includes:

[0047] Constructing a reduced-representation genomic sequencing library by using the method described above in the present invention;

[0048] Performing sequencing on a machine.

[0049] According to a specific embodiment of the present invention, in the reduced-representation genomic sequencing method of the present invention, according to the concentration of the library, sequencing is performed according to the requirements of relevant instruments such as Illumina or MGI.

[0050] In some specific embodiments of the present invention, the method of the present invention is used for reduced-representation genomic sequencing of plant varieties or animal varieties. Preferably, the plants are crop plants or flower plants, such as gramineous plants (such as rice, corn, wheat, etc.), leguminous plants (such as soybeans, etc.), peppers, cotton, chrysanthemums, etc. Preferably, the animals are livestock, such as cattle, sheep, pigs, etc.

[0051] On the other hand, the present invention also provides a kit for constructing a reduced-representation genomic sequencing library and / or for reduced-representation genomic sequencing. The kit includes:

[0052] (1) 2 to 24 kinds (preferably at least 4 kinds, more preferably at least 6 kinds) of restriction endonucleases and their Buffers;

[0053] The restriction endonuclease is selected from Aci I, Bfa I, CviQ I, Fat I, HinP1 I, HpyCH4 IV, MboI, MluC I, Mse I, Msp I, Taq I, Alu I, Dde I, Fnu4H I, Hinf I, HpyCH4 III, HpyCH4 V, RsaI, Sau96 I, StyD4 I, Hae III, Nla III, Hpy188 I, HpyCH4 III, Acl I, Afl II, Age I, ApaLI, BamH I, Bcl I, BseY I, BsiW I, BspD I, BspH I, BsrG I, BssS I, Eag I, EcoR I, HindIII, Kas I, Mfe I, Mlu I, Nco I, Nde I, Nhe I, Sal I, Spe I, Xma I;

[0054] Preferably, the restriction endonuclease includes at least 6 of Bfa I, CviQ I, Fat I, HinP1 I, HpyCH4 IV, MboI, MluC I, Mse I, Msp I, Taq I, Alu I, Dde I, Fnu4H I, Hinf I, HpyCH4 III, HpyCH4 V, Rsa I;

[0055] More preferably, the restriction endonuclease includes Bfa I, Dde I, Hae III, MboI, MluC I and Mse I.

[0056] The reagent components of the kit described in the present invention may optionally include one or more of the following:

[0057] (2) End repair Enzyme and its buffer;

[0058] (3) A-Tailing Enzyme and its Buffer;

[0059] (4) T4 DNALigase and its buffer;

[0060] (5) DNA adapter.

[0061] In some specific embodiments of the present invention, the kit of the present invention is used for constructing a next-generation sequencing library of reduced-representation genome, wherein the restriction enzymes include at least 2, at least 4, at least 6, at least 8, at least 10, at least 12 or all of Aci I, Bfa I, CviQ I, Fat I, HinP1 I, HpyCH4 IV, MboI, MluC I, Mse I, Msp I, Taq I, Alu I, Dde I, Fnu4H I, Hinf I, HpyCH4 III, HpyCH4V, Rsa I, Sau96 I, StyD4 I, Hae III, Nla III, Hpy188 I, HpyCH4 III. In some more specific embodiments of the present invention, the restriction enzymes include Bfa I, CviQ I, Fat I, HinP1 I, HpyCH4 IV, MboI, MluC I, Mse I, Msp I, Taq I, Alu I, Dde I, Fnu4H I, Hinf I, HpyCH4 III, HpyCH4 V, Rsa I.

[0062] In some specific embodiments of the present invention, the kit of the present invention is used for constructing a third-generation sequencing library of reduced-representation genome, wherein the restriction enzymes include at least 2, at least 4, at least 6, at least 8, at least 10, at least 12 or all of Acl I, Afl II, Age I, ApaL I, BamH I, Bcl I, BseYI, BsiW I, BspD I, BspH I, BsrG I, BssS I, Eag I, EcoR I, Hind III, Kas I, Mfe I, Mlu I, Nco I, Nde I, Nhe I, Sal I, Spe I, Xma I.

[0063] Another object of the present invention is to provide related applications of the kit. Specifically, the applications include being used for constructing a reduced-representation genome sequencing library and / or for reduced-representation genome sequencing. For example, it can be used for performing reduced-representation genome sequencing on plants. Preferably, the plants are crop plants or flower plants, such as gramineous plants (such as rice, corn, wheat, etc.), leguminous plants (such as soybean, etc.), chili peppers, cotton, chrysanthemums, etc. Preferably, the animals are livestock, such as cattle, sheep, pigs, etc.

[0064] The biRAD-seq library construction technology developed in the present invention mainly innovates in using the specificity of the cleavage recognition sequences of different restriction endonucleases to replace the Barcode added in the library construction process in the traditional method for distinguishing samples, so as to simplify the library construction experiment and improve the sequencing efficiency. Compared with the traditional method, the method of the present invention omits the step of adding an extra barcode to distinguish samples, simplifies the experimental operation process, improves the library construction efficiency, and at the same time, explores the feasibility of more types of restriction endonucleases in the reduced-representation sequencing technology. BRIEF DESCRIPTION OF THE DRAWINGS

[0065] Figure 1 It is a schematic diagram of the principle of the present invention (biRAD-seq).

[0066] Figure 2A and Figure 2B respectively show the proportions of the digestion fragments of 300 - 400 bp by the 4-base recognition restriction endonuclease mimic in the total genome length of multiple species and the proportions of the digestion fragments of 9 - 11 kb by the 6-base recognition restriction endonuclease mimic in the total genome length of multiple species.

[0067] Figures 3A to 3J It is a distribution map of the 300 - 400 bp fragments obtained by electronically simulating the digestion of rice genomic DNA with different restriction endonucleases on different chromosomes.

[0068] Figure 4 It shows the experimental results of the agarose gel electrophoresis of the Nipponbare rice DNA sample.

[0069] Figure 5 It is a schematic diagram of data error splitting.

[0070] Figure 6 It shows the coverage statistics of 1.0 Gb data with an inserted fragment size of 300 - 400 bp on the rice genome.

[0071] Figure 7 It shows the distribution results of DNA fragments detected by Qsep100.

[0072] Figures 8A to 8J It shows the inserted fragment size distribution of the sequencing data of 10 samples.

[0073] Figure 9A and Figure 9B It shows the coverage statistics of randomly selected 1.0 Gb and 2.0 Gb data on the rice genome respectively.

[0074] Figure 10 It is a distribution map of the sequencing sequences of the digestion fragments of Bfa I, CviQ I, MluC I, and Mse I enzymes in IGV in biRAD-seq. Detailed Embodiments

[0075] For a clearer understanding of the technical features, objectives, and beneficial effects of the present invention, the technical solutions of the present invention will be described in detail below in conjunction with specific embodiments. It should be understood that these examples are only used to illustrate the present invention and not to limit the scope of the present invention. In the examples, all original reagent materials can be obtained commercially. The experimental methods without specific conditions are conventional methods and conventional conditions well-known in the art, or the conditions recommended by the instrument manufacturers.

[0076] Unless otherwise specifically defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the relevant art.

[0077] The principle and operation process of the method for reduced-representation genome sequencing of the present invention are as Figure 1 shown.

[0078] In some specific embodiments of the present invention, the method for reduced-representation genome sequencing of the present invention includes:

[0079] Using different restriction endonucleases to digest different sample DNAs to obtain digestion products;

[0080] Measuring the concentration of the digested sample DNA and equally mixing the digestion products of different samples;

[0081] Adding the DNA adapter sequences required for sequencing to both ends of the DNA digestion fragments after mixing different samples through the Truseq library construction method to obtain a DNA library;

[0082] DNA library fragment selection: Performing fragment selection on the DNA library constructed from the DNA digestion fragments after mixing different samples, screening the DNA library fragment lengths using a series of instruments developed by Sage Science (Sage BluePippin, Sage ELF, or SagePippinHT), and recovering DNA library fragments with a length of 420 - 520 bp (the sequencing adapter is about 120 bp long) using a 2% agarose gel DNA recovery cassette (Cassette);

[0083] Detection and quantification of DNA library fragments: Using Qsep100 to detect the actual fragment size of the selected library fragments; the quantification of library fragments is performed using a 3.0 fluorometer for measurement

[0084] Sequencing of the DNA library: According to the concentration of the library, sequencing is performed on the machine according to the requirements of Illumina-related instruments.

[0085] Preferably, in the method for reduced-representation genome sequencing of the present invention, when digesting genomic DNA, each sample is digested with 1 restriction endonuclease, preferably with 2-24 restriction endonucleases that recognize different sequences, and each enzyme digests one sample to generate specific end sequences.

[0086] Preferably, in the method for reduced-representation genome sequencing of the present invention, when digesting genomic DNA, a reduction level of 65%-97%, preferably 75%-96%, more preferably 80%-90% is achieved.

[0087] Preferably, in the method for reduced-representation genome sequencing of the present invention, a PCR-free library is constructed using the True-seq library construction method or other methods for constructing sequencing libraries.

[0088] Preferably, in the method for reduced-representation genome sequencing of the present invention, the in silico digestion results of a predetermined restriction endonuclease can be evaluated for a specific species first to screen the restriction endonucleases for the specific species. Specifically, the in silico digestion of the reference genome sequence of the species to be sequenced can be performed using the fa2cmap_multi_color.pl script in the Solve software (https: / / bionanogenomics.com / ), and then the find_enzyme.pl script can be used to count the digested fragments of different lengths, and analyze the length distribution of the DNA fragments after in silico digestion of the genome with the predetermined combination of restriction endonucleases, so as to select appropriate (i.e., those that can achieve the predetermined reduction level) restriction endonucleases. The dosage of each restriction endonuclease can refer to the dosage in conventional molecular biology experiments. Usually, the dosage of each enzyme is 10 U for 1 μg of DNA.

[0089] In some specific embodiments of the present invention, the method for reduced-representation genome sequencing of the present invention is used for sequencing the genomes of rice, soybean, corn or wheat.

[0090] Preferably, the restriction endonuclease includes, but is not limited to, one or more of Bfa I, CviQ I, Fat I, HinP1 I, HpyCH4IV, MboI, MluC I, Mse I, Msp I, Taq I, Alu I, Dde I, Fnu4H I, Hinf I, HpyCH4 III, HpyCH4 V, Rsa I.

[0091] In some more specific embodiments of the present invention, the restriction endonuclease is:

[0092] A combination of Bfa I, CviQ I, Fat I, HinP1 I, HpyCH4 IV, MboI, MluC I, Mse I, Msp I, Taq I; or

[0093] A combination of Bfa I, HpyCH4 IV, MboI, MluC I, Mse I, Alu I, Dde I, Fnu4H I, Hinf I, HpyCH4III, HpyCH4 V, Rsa I, etc.

[0094] Preferably, in the method for reduced-representation genome sequencing of the present invention, the method for sorting fragments of a target length includes, but is not limited to: sorting using an ELF or Sage HT instrument from Sage Science, sorting by gel cutting, or sorting using magnetic beads.

[0095] Preferably, in the method for reduced-representation genome sequencing of the present invention, the method for sequencing the sorted fragments includes, but is not limited to: sequencing using a BGI sequencing platform, an Illumina sequencing platform, a Life Technology sequencing platform, a PacBio sequencing platform, an ONT sequencing platform, or other sequencing platforms.

[0096] Example 1: Simulating genome digestion using Perl scripts

[0097] This example evaluated the feasibility of the biRAD-seq technical solution of the present invention at the theoretical level, selected suitable restriction endonucleases for rice samples, and verified the feasibility of this solution in other common crops and livestock.

[0098] The reference genome sequence of Nipponbare rice (genome version information: T2T-NIP AGIS-1.0) was simulated for digestion using the fa2cmap_multi_color.pl script in Solve software (https: / / bionanogenomics.com / ), and then the number and proportion of fragments of different lengths after simulated digestion were counted using the find_enzyme.pl script. According to the proportion of fragments with a length range of 300 - 400 bp in the reference genome, suitable restriction endonucleases (i.e., those that can achieve a predetermined reduction level) were selected.

[0099] The results of simulated enzyme digestion show (here, Nipponbare rice is used as an example, but this method is not limited to rice): First, after the rice genome is digested by the restriction endonuclease that recognizes 6 base sequences, the total length of the enzyme-cut fragments of 300 to 400 bp accounts for less than 1% of the rice genome, and the total length of the enzyme-cut fragments of 9000 to 11000 bp accounts for about 10% of the rice genome; after the rice genome is digested by the restriction endonuclease that recognizes 4 base sequences, the total length of the enzyme-cut fragments of 300 to 400 bp accounts for about 10% of the rice genome. Considering that the proportion of the simplified genome to be tested should account for 2% to 10% of the whole genome, the restriction endonuclease that recognizes 4 base sequences will be selected for the construction of the second-generation sequencing library. Secondly, considering that the recognition site sequence of the restriction endonuclease is used as a marker to distinguish samples, the overhang sequences generated after the selected restriction endonuclease digestion should be different. Therefore, based on the above two considerations and the results of simulated enzyme digestion, 10 restriction endonucleases (Mse I, Msp I, Bfa I, HpyCH4IV, CviQ I, Taq I, HinP1I, MluC I, FatI, MboI) were initially selected for biRAD-seq. The distribution of the lengths of the simulated enzyme digestion fragments of these 10 restriction endonucleases on the rice genome is shown in Table 1.

[0100] Table 1: Proportion of different lengths of enzyme fragments of 10 restriction endonucleases in the rice genome

[0101] Serial number Restriction enzyme Restriction sequence Below 250bp 250 - 350bp 350 - 450bp 450 - 550bp 1 Bfa I C↓TA↑G 16.75% 9.75% 9.32% 8.65% 2 CviQ I G↓TA↑C 14.46% 9.05% 8.52% 7.66% 3 Fat I ↓CATG↑ 33.73% 14.62% 11.51% 9.36% 4 HinP1 I G↓CG↑C 12.82% 4.10% 3.65% 3.86% 5 HpyCH4 IV A↓CG↑T 9.39% 5.75% 5.86% 5.67% 6 MboI ↓GATC↑ 20.18% 11.09% 10.37% 9.37% 7 MluC I ↓AATT↑ 48.47% 12.11% 8.13% 6.09% 8 Mse I T↓TA↑A 36.09% 11.12% 9.01% 7.04% 9 Msp I C↓CG↑G 15.61% 5.66% 4.84% 5.14% 10 Taq I T↓CG↑A 15.31% 8.11% 7.81% 7.37%

[0102] In addition, the present invention also found that the length of the fragment produced by the restriction endonuclease digestion of the six base sequences is 9000-11000bp and is very suitable for the construction of the three-generation sequencing library. Therefore, the present embodiment also selected 12 kinds of restriction endonucleases (Afl II, Age I, ApaL I, BamH I, BseY I, BssS I, Eag I, EcoR I, Hind III, Nco I, Sal I, Spe I) that recognize the six base sequences and can be used for the construction of the three-generation sequencing library. The length distribution of the simulated enzyme digestion fragments of these 12 kinds of restriction endonucleases on the rice genome is shown in Table 2.

[0103] Table 2: Proportion of different lengths of restriction fragments of 12 restriction endonucleases in the rice genome

[0104] Serial number Restriction enzyme Restriction sequence Less than 1kb 5 - 7kb 7 - 9kb 9 - 11kb 1 Afl II C↓TTAA↑G 1.10% 10.76% 11.68% 9.43% 2 Age I A↓CCGG↑T 1.26% 9.40% 9.90% 7.94% 3 ApaL I G↓TGCA↑C 2.01% 13.64% 12.59% 10.33% 4 BamH I G↓GATC↑C 1.37% 11.10% 11.08% 9.90% 5 BseY I C↓CCAG↑C 4.98% 16.20% 11.94% 8.86% 6 BssS I C↓ACGA↑G 4.61% 16.64% 12.84% 8.90% 7 Eag I C↓GGCC↑G 5.48% 13.64% 11.80% 9.11% 8 EcoR I G↓AATT↑C 2.62% 14.39% 13.80% 9.78% 9 Hind III A↓AGCT↑T 3.75% 16.89% 12.39% 9.29% 10 Nco I C↓CATG↑G 3.90% 16.53% 13.82% 9.05% 11 Sal I G↓TCGA↑C 1.81% 10.17% 9.90% 8.84% 12 Spe I A↓CTAG↑T 2.45% 15.16% 13.28% 10.00%

[0105] Based on the results of simulated enzymatic digestion of the rice genome described above, the present invention also used restriction endonucleases that recognize 4-base sequences and restriction endonucleases that recognize 6-base sequences to simulate the enzymatic digestion of the genomes of common crops and livestock such as rice, maize, wheat, pepper, cotton, chrysanthemum, cattle, and pigs, and obtained the length ratios of the 300-400 bp enzymatic digestion fragments generated by the restriction endonucleases that recognize 4-base sequences and the 9-11 kb enzymatic digestion fragments generated by the restriction endonucleases that recognize 6-base sequences on the genomes of each species. The results are as Figure 2A and Figure 2B shown. The results indicate that for most restriction endonucleases, the target length fragments (300-400 bp or 9-11 kb) generated in these species account for approximately 10% of the genome length. Performing biRAD-seq library construction and sequencing on these fragments can achieve a simplification level of 80%-90%, thus demonstrating that the biRAD-seq technical solution of the present invention is theoretically feasible in multiple species.

[0106] In this example, the proportion of the total length of fragments with different lengths obtained by simulating the enzymatic digestion of rice genomic DNA with the above 10 restriction endonucleases that recognize 4-base sequences was predicted, reflecting the size and simplification degree of the measured genomic regions. The results showed that for the rice genome, after electronically digesting genomic DNA with different restriction endonucleases, the density of the Tags distributed on each chromosome was as Figures 3A to 3J shown. It can be seen from this that these 10 restriction endonucleases are predicted to be able to effectively process rice DNA. The Tag of the enzymatic digestion fragments with a length of 300-400 bp is evenly distributed on the genome and only covers approximately 10% of the entire genome, meeting the requirements for reduced-representation genome sequencing.

[0107] Example 2: Construction of sequencing libraries for 10 samples individually to perform biRAD-seq

[0108] In this example, after enzymatic digestion of each sample, a sequencing library was constructed and sequenced individually to examine the performance of each restriction endonuclease in biRAD-seq, including enzymatic digestion efficiency, characteristics of the RAD-seq tags generated, and the impact of un-digested enzymatic digestion sites on subsequent bioinformatics analysis processes. By randomly extracting quantitative data from each sample and mixing them to simulate the sequencing data generated by mixing and constructing a sequencing library after enzymatic digestion of each sample during the actual use of biRAD-seq, a corresponding bioinformatics analysis process was established, and the accuracy of splitting the sequencing data from different samples was evaluated and examined.

[0109] 1. DNA quantification

[0110] After obtaining the gDNA of Nipponbare rice samples, the concentration was first detected using a NanoDrop 2000 spectrophotometer (NanoDrop 2000 Spectrophotometer, Thermo) and a 3.0 fluorometer (3.0 Fluorometer, Invitrogen) (Table 3). Then, the DNA integrity was detected. The quality integrity of the extracted rice gDNA (part) was detected using 1% agarose gel electrophoresis. The results are shown in Figure 4 , and it was found that all DNA bands were intact, indicating good DNA integrity and meeting the requirements for library construction.

[0111] Table 3: DNA Concentration Detection

[0112] Sample number Concentration (ng / μL) 1 50.4 2 54.6 3 52.0 4 49.2 5 48.4 6 49.2 7 55.1 8 49.6 9 51.2 10 49.5

[0113] 2. Digestion and fragmentation of rice gDNA

[0114] Ten samples were digested with ten restriction endonucleases (see Table 4), and all the enzymes used were purchased from NEB.

[0115] Table 4: Restriction Endonucleases Used and Their Digestion Sequences

[0116] Sample number Restriction endonuclease Restriction sequence 1 Bfa I C↓TA↑G 2 CviQ I G↓TA↑C 3 Fat I ↓CATG↑ 4 HinP1 I G↓CG↑C 5 HpyCH4 IV A↓CG↑T 6 Mbo I ↓GATC↑ 7 MluC I ↓AATT↑ 8 Mse I T↓TA↑A 9 Msp I C↓CG↑G 10 Taq I T↓CG↑A

[0117] Among the ten restriction endonucleases, CviQ I uses NEBuffer TM r3.1, and Fat I uses NEBuffer TM r2.1). The digestion reaction system and reaction conditions are shown in Tables 5 and 6.

[0118] Table 5: Digestion Reaction System

[0119] Reagent Volume gDNA XμL (1μg) rCutSmart or r3.1 or r2.1 Buffer 5μL Restriction endonuclease 1μL <![CDATA[ddH2O]]> (44 - X)μL

[0120] Table 6: Digestion Reaction Conditions

[0121] Sample number Restriction endonuclease Reaction temperature and time Heat inactivation temperature and time 1 Bfa I 37℃,3h 80℃, 20min 2 CviQ I 25℃,3h 65℃, 20min 3 Fat I 55℃,3h 65℃, 20min 4 HinP1 I 37℃,3h 65℃, 20min 5 HpyCH4 IV 37℃,3h 65℃, 20min 6 Mbo I 37℃,3h 65℃, 20min 7 MluC I 37℃,3h 65℃, 20min 8 Mse I 37℃,3h 65℃, 20min 9 Msp I 37℃,3h / 10 Taq I 65℃,3h /

[0122] The digested DNA fragments were purified using the Select-A-Size DNA Clean&Concentrator from ZYMO TM by column.

[0123] 3. Truseq Library Construction

[0124] In this embodiment, the Truseq library construction was carried out using the Hieff NGS Uitima Pro Free DNA Library Prep Kit V2 all-in-one DNA library construction kit (PCR-Free) V2 (Cat#12196) and referring to its user manual. The main operation steps are as follows.

[0125] (1) End repair of DNA fragments, and the reaction system is shown in Table 7.

[0126] Table 7: End Repair reaction system

[0127] Reagent Volume Fragmented DNA XμL (1μg) Endprep Buffer 2.0 6μL Endprep Enzyme 2.0 4μL <![CDATA[ddH2O]]> (50 - X)μL

[0128] The end repair was carried out in a PCR instrument, and the reaction program was: hot lid at 105°C; react at 30°C for 20 min; react at 72°C for 20 min; hold at 4°C.

[0129] (2) Adding A to the ends of DNA fragments, and the reaction system is shown in Table 8.

[0130] Table 8: dA-Tailing reaction system

[0131] Reagent Volume dA-tailed DNA 60μL Ligation Enhancer 2.0 30μL DNAAdapter 5μL Rapid T4 DNALigase 2.0 5μL

[0132] Adding A was carried out in a PCR instrument, and the reaction program was: hot lid off; react at 20°C for 15 min; hold at 4°C.

[0133] (3) Purify the reaction product using 0.8× magnetic beads, and elute the purified product with ddH2O.

[0134] 4. Fragment selection

[0135] Fragment selection of the DNA library was carried out using the Sage HT instrument from Sage Science, and fragments with a size of approximately 300 - 400 bp were recovered using a 2% agarose gel DNA recovery cassette.

[0136] 5. Sequencing

[0137] According to the concentration of the recovered library, sequencing was performed on the relevant Illumina sequencing instrument according to the sample requirements.

[0138] 6. Data analysis

[0139] Software such as FastQC, Fastp, BWA, Samtools, Bedtools, and Qualimap were used to analyze and process the sequencing data downloaded from the instrument; Excel, R language, and Adobe Illustrator were used for plotting.

[0140] First, the raw data of the off-machine data is preprocessed by software such as FastQC and Fastp to obtain clean data. Then, the clean data is aligned to the AGIS-1.0 reference genome of rice through BWA. After that, personalized analyses such as coverage statistics can be performed using tools such as Samtools, Bedtools, and Qualimap.

[0141] (1) Estimation of digestion efficiency

[0142] For the method of using restriction enzymes to break DNA fragments for constructing sequencing libraries, whether the fragments are completely broken (the level of digestion efficiency) determines the quality of the constructed sequencing library. Therefore, in this example, the proportion of reads containing the digestion recognition sequence in the data of samples treated with different restriction enzymes is statistically analyzed to estimate the digestion efficiency of the restriction enzymes. The results are shown in Table 9.

[0143] Table 9: Statistical results of the number of reads containing digestion sites

[0144]

[0145]

[0146] It can be seen from the results in Table 9 that among the above 10 restriction enzymes, only Fat I and Msp I have relatively low digestion efficiency, and the other restriction enzymes can show good digestion effects.

[0147] (2) Simulation and verification of splitting different sample libraries

[0148] In this example, PCR-free Truseq libraries were constructed and sequenced for 10 samples respectively, and it was not necessary to split the sequencing data. To simulate the actual application scenario of mixing the digested products of different samples, constructing a PCR-free Truseq library and sequencing, 1.0 Gb of data was randomly extracted from the 10 sample fastq files in this example to form 10.0 Gb of mixed data. Then, the mixed data was split by the script split_multi_sample.py according to the different digestion recognition sequences of these 10 restriction endonucleases. The splitting results are shown in Table 10, indicating that the splitting rate of this data splitting method is about 85% and the splitting accuracy can reach 99%, which can be used for actual applications.

[0149] Table 10

[0150] Serial number Restriction enzyme Restriction sequence Original data volume (Gb) Split data volume (Gb) Accuracy rate 1 Bfa I C↓TA↑G 1.0 0.88 98.49% 2 CviQ I G↓TA↑C 1.0 0.88 98.91% 3 Fat I ↓CATG↑ 1.0 0.88 98.93% 4 HinP1 I G↓CG↑C 1.0 0.67 99.99% 5 HpyCH4 IV A↓CG↑T 1.0 0.77 99.01% 6 MboI ↓GATC↑ 1.0 0.85 98.95% 7 MluC I ↓AATT↑ 1.0 0.88 99.98% 8 Mse I T↓TA↑A 1.0 0.84 98.94% 9 Msp I C↓CG↑G 1.0 0.81 98.89% 10 Taq I T↓CG↑A 1.0 0.85 98.86%

[0151] For the data that fails to be successfully split and the data with splitting errors, this embodiment analyzes the reasons for their generation. There are mainly two reasons why the data cannot be split: one is that the restriction endonuclease has star activity and will recognize random sites and break them, and such data cannot be split according to the generated overhang; the other is the experimental error introduced during library construction, which will lead to the contamination of other restriction fragments. The reason for the data splitting error may also come from the contamination of other restriction fragments. By artificially mixing the sample data treated with different restriction enzymes, then splitting them according to the method of the present invention, and finally tracing the actual origin of the reads in the sequencing data through the indexes added during the library construction of different samples. It is found that in the sample data of Mbo I (index: CACTTCGA) after splitting, there are a few reads with an overhang sequence of GATC but the index belongs to HpyCH4 IV (index: ACCACTGT), as Figure 5 shown, it is judged that the reason for the generation of these reads is that the sample of Mbo I was contaminated by the sample of HpyCH4 IV due to mistakes in the experiment, resulting in flaws in data splitting.

[0152] (3) Statistics of genome coverage

[0153] This embodiment compares the coverage of the sequencing data of each restriction enzyme-digested sample (insert size is 300 - 400bp) on the whole genome under the condition of 1.0 Gb data volume. The results are as Figure 6 shown, and the results show that the data measured from the samples treated with these 10 restriction enzymes can basically cover 1% - 10% of the regions on the whole genome at a certain sequencing depth. Compared with the simulated digestion results, it is slightly lower than the theoretical coverage, but it can also meet the needs of subsequent data analysis.

[0154] Example 3: Implement biRAD-seq on 10 samples

[0155] This embodiment verifies the actual application effect of the biRAD-seq method. The library construction process is to first digest 10 samples with restriction endonucleases, and then mix the 10 samples to construct a PCR-free Truseq library and sequence it.

[0156] 1. DNA quantification

[0157] After obtaining the gDNA of the Nipponbare rice sample, first use the NanoDrop 2000 spectrophotometer (NanoDrop2000 Spectrophotometer, Thermo) and 3.0 Fluorometer (3.0 Fluorometer, Invitrogen) was used to detect its concentration (Table 3). Then, DNA integrity was detected. The quality and integrity of the extracted rice gDNA (part) were detected by 1% agarose gel electrophoresis. The results are shown in Figure 4 , and it was found that all DNA bands were intact, indicating good DNA integrity and meeting the requirements for library construction.

[0158] 2. Digest and fragment rice gDNA

[0159] Ten samples were digested with ten restriction endonucleases (see Table 4), and all the enzymes were purchased from NEB.

[0160] Among the ten restriction endonucleases, CviQ I uses NEBuffer TM r3.1, and Fat I uses NEBuffer TM r2.1). The digestion reaction systems and reaction conditions are shown in Table 5 and Table 6.

[0161] The digested DNA fragments were column-purified using ZYMO's Select-A-Size DNA Clean&Concentrator TM .

[0162] 3. Fragment selection

[0163] The DNA library was fragment-selected using Sage Science's Sage HT instrument, and fragments with a size of approximately 300 - 400 bp were recovered using a 2% agarose gel DNA recovery cassette. The Qsep100 was used to view the fragment sorting effect, and the results are as Figure 7 shown.

[0164] 4. Truseq library construction

[0165] In this example, the Hieff NGS Uitima Pro Free DNA Library Prep Kit V2 all-in-one DNA library construction kit (PCR-Free) V2 (Cat#12196) was used and its user manual was referred to for Truseq library construction. The main operating steps are as follows.

[0166] (1) End repair of DNA fragments. The reaction system is shown in Table 7.

[0167] The end repair was carried out in a PCR instrument. The reaction program was: hot lid at 105°C; reaction at 30°C for 20 min; reaction at 72°C for 20 min; hold at 4°C.

[0168] (2) Add A to the ends of the DNA fragments. The reaction system is shown in Table 8.

[0169] The addition of A is carried out in a PCR instrument. The reaction program is as follows: hot lid off; react at 20 °C for 15 min; hold at 4 °C.

[0170] (3) Purify the reaction products using 0.8× magnetic beads, and elute the purified products with ddH2O.

[0171] 5. Sequencing

[0172] According to the concentration of the recovered library, perform on-machine sequencing according to the sample requirements of the relevant Illumina sequencing instruments.

[0173] 6. Data splitting

[0174] The off-machine data is mixed data of 10 samples, and data splitting needs to be carried out according to the different 5’ overhang sequences of the restriction endonucleases used. Run the python script split_multi_sample.py to split the sequencing data into paired-end sequencing fastq files of 10 samples. This script uses multi-threading technology and regular expression matching to identify the enzyme digestion recognition sequences, and efficiently splits the reads starting with specific enzyme digestion recognition sequences from the original FASTQ file to generate the corresponding sample data files processed by restriction endonucleases.

[0175] 7. Data analysis

[0176] Use software such as FastQC, Fastp, BWA, Samtools, Bedtools, and Qualimap to analyze and process the off-machine sequencing data; use Excel, R language, and AI for plotting.

[0177] (1) Through sequencing and data preprocessing, 64.9 Gb of clean data was obtained. After being split by the python script split_multi_sample.py, a total of 50.2 Gb of available sequencing data was obtained, and the data splitting rate was 77.34%. It has the same effect as the traditional barcode data splitting method, but the method of the present invention relatively simplifies the experimental process. Therefore, it is proved that the method of the present invention is feasible.

[0178] (2) By analyzing the insert fragment length distribution of the sequencing data ( Figures 8A to 8J ) and 1.0 Gb ( Figure 9A )、2.0 Gb ( Figure 9B)The coverage of data on the rice genome shows that the vast majority of the inserted fragment lengths are distributed between 300 and 400 bp. 1.0 Gb of data covers approximately 3% - 10% of the rice genome at a certain sequencing depth. And by viewing the reads alignment using igv, it can be seen that ( Figure 10 ), biRAD-seq has achieved the purpose of reduced-representation genome sequencing by sequencing partial restriction enzyme digestion interval sequences. Even within a 10 Kb window of the rice genome, partial genomic sequences can still be well measured. This shows that the actual test effect is similar to the simulated restriction enzyme digestion effect. This method has achieved consistent results with expectations and can be used for reduced-representation genome sequencing.

[0179] Example 4: Application of 3'-sticky-end restriction enzymes in biRAD-seq

[0180] In this example, a new splitting method was developed for sample data treated with 3'-sticky-end restriction enzymes and extended to 5'-sticky-end restriction enzymes.

[0181] Since the 3'-sticky ends will be degraded after end repair, data cannot be directly split by the overhang sequences. However, the sequences at both ends of the fragments digested by 3'-sticky-end restriction enzymes can still be retrieved through the reference genome sequence and can be used as barcodes to distinguish different samples.

[0182] In this example, with the help of the reference genome sequence, the first 8 bp base sequence of the simulated restriction enzyme digestion fragments was extracted as the barcode. Each restriction enzyme generates approximately 50,000 barcodes for 300 - 400 bp fragments. Using the software Seqkit (https: / / bioinf.shenwei.me / seqkit / ), reads belonging to the sample data treated with a certain restriction enzyme can be retrieved and matched in the mixed data through these barcodes. Using this method, the sample data treated with 3'-sticky-end restriction enzymes was successfully split in this example. Then this method was extended to 5'-sticky-end restriction enzymes, and their splitting rates and accuracies were explored. The results are shown in Table 11.

[0183] Table 11

[0184] Serial number Restriction enzyme Digestion sequence Split rate Accuracy rate 1 Bfa I C↓TA↑G 93.16% 97.37% 2 CviQ I G↓TA↑C 93.33% 97.67% 3 Fat I ↓CATG↑ 92.34% 98.18% 4 HinP1 I G↓CG↑C 70.94% 98.71% 5 HpyCH4 IV A↓CG↑T 84.18% 97.79% 6 MboI ↓GATC↑ 95.17% 98.25% 7 MluC I ↓AATT↑ 92.42% 99.27% 8 Mse I T↓TA↑A 89.24% 97.26% 9 Msp I C↓CG↑G 88.64% 97.98% 10 Taq I T↓CG↑A 90.12% 97.79%

[0185] It can be seen from this that the method of the present invention is applicable to both 3'-sticky-end restriction enzymes and 5'-sticky-end restriction enzymes, further expanding the selection of restriction enzymes for biRAD-seq. And combined with the results of the previous splitting method, they can be mutually verified to further improve the accuracy of data splitting.

[0186] The above are only specific embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.

Claims

1. A method for constructing a simplified genome sequencing library, the method comprising: (1) using restriction endonucleases to digest multiple DNA samples respectively; wherein the restriction endonucleases used for each sample are different, so as to generate sample-specific DNA fragment ends as their own sample barcodes; (2) mixing the enzyme digestion products of each sample, constructing a DNA sequencing library, performing fragment sorting on the DNA sequencing library, and using the selected library as a simplified genome sequencing library; or, The enzyme digestion products of each sample were fragmented and sorted, and the selected fragments were used to construct a DNA sequencing library to obtain a simplified genome sequencing library.

2. The method according to claim 1, wherein: The multiple DNA samples are samples of plants or animals from different individuals; Preferably, the number of the multiple DNA samples is 2-24, preferably 4-18, and more preferably 6-16; Preferably, in a single DNA sample, the DNA concentration is greater than or equal to 40 ng / μL, and the total DNA amount is greater than or equal to 1 μg.

3. The method according to claim 1, wherein: The DNA sample is digested with restriction endonucleases to achieve a simplification level of 65% to 97%, preferably a simplification level of 75% to 96%; Preferably, the restriction endonuclease comprises an enzyme that produces a 5' sticky end and / or an enzyme that produces a 3' sticky end; Preferably, the restriction endonuclease is a restriction endonuclease that recognizes a 4-6 base sequence; Further preferably, the restriction endonuclease includes but is not limited to one or more of Aci I, Bfa I, CviQ I, Fat I, HinP1I, HpyCH4 IV, MboI, MluC I, Mse I, Msp I, Taq I, Alu I, Dde I, Fnu4H I, Hinf I, HpyCH4III, HpyCH4 V, Rsa I, Sau96 I, StyD4 I, Hae III, Nla III, Hpy188 I, HpyCH4III, Acl I, Afl II, Age I, ApaL I, BamH I, Bcl I, BseY I, BsiW I, BspD I, BspH I, BsrG I, BssS I, Eag I, EcoR I, Hind III, Kas I, Mfe I, Mlu I, Nco I, Nde I, Nhe I, Sal I, Spe I, and Xma I; Further preferably, for each DNA sample, the restriction endonuclease is selected from one of Bfa I, CviQ I, Fat I, HinP1I, HpyCH4 IV, MboI, MluC I, Mse I, Msp I, Taq I, Alu I, Dde I, Fnu4H I, Hinf I, HpyCH4III, HpyCH4 V, and Rsa I.

4. The method according to claim 1, wherein: The fragment sorting is to select target fragments of 100 to 600 bp, preferably 250 to 400 bp for second-generation sequencing, or to select fragments of 5 to 15 kb, preferably 7 to 11 kb for third-generation sequencing; Preferably, when performing fragment sorting, the target DNA fragments are selected by using magnetic beads or agarose gel excision, or fragments are selected using an agarose gel electrophoresis system or Sage BluePippin, Sage ELF or Sage PippinHT.

5. The method according to claim 1, wherein: The digestion products of each sample were mixed in equal amounts; Preferably, the method for constructing the simplified genome sequencing library further comprises: Detection of the concentration and length distribution of DNA fragments from enzyme-digested products; Calculate the mixing ratio of the DNA digested fragments of each sample: according to the concentration of the digested fragments of each sample and the ratio of the target fragments in the sample, obtain the concentration of the target fragments in the library; according to the concentration of the target fragments, take the same amount of target DNA fragments from the DNA digested fragments of each sample and mix them to obtain the final mixed library.

6. The method according to claim 1, wherein: The construction process of the DNA sequencing library includes: adding the universal DNA adapter sequence required for sequencing to the two segments of the DNA enzyme-cut fragment to be tested to complete the construction of the DNA library; Preferably, the sequencing library is a PCR free library.

7. Use of the method for constructing a simplified genome sequencing library according to any one of claims 1 to 6 in simplified genome sequencing.

8. A simplified genome sequencing method, the method comprising: Constructing a simplified genome sequencing library using the method described in any one of claims 1 to 6; On-line sequencing.

9. The method according to claim 8, which is to perform simplified genome sequencing on plant species or animal species; Preferably, the plant is a crop or flower plant, such as rice, corn, wheat, soybean, pepper, cotton, chrysanthemum, etc.; Preferably, the animal is livestock, such as cattle, horses, sheep, pigs, etc.

10. A kit for constructing a simplified genome sequencing library and / or for simplifying genome sequencing, the reagent components of the kit comprising: (1) 2 to 24 (preferably at least 4, more preferably at least 6) restriction endonucleases and their buffers; the restriction endonucleases are selected from Aci I, Bfa I, CviQ I, Fat I, HinP1 I, HpyCH4 IV, MboI, MluC I, Mse I, Msp I, Taq I, Alu I, Dde I, Fnu4H I, Hinf I, HpyCH4 III, HpyCH4 V, Rsa I, Sau96 I, StyD4 I, Hae III, Nla III, Hpy188 I, HpyCH4 III, Acl I, Afl II, Age I, ApaL I, BamH I, Bcl I, BseY I, BsiW I, BspD I, BspH I, BsrG I, BssS I, Eag I, EcoR I, Hind III, Kas I, Mfe I, Mlu I, Nco I, Nde I, Nhe I, Sal I, Spe I, Xma I; Preferably, the restriction endonuclease comprises at least 6 of Bfa I, CviQ I, Fat I, HinP1 I, HpyCH4 IV, MboI, MluCI, Mse I, Msp I, Taq I, Alu I, Dde I, Fnu4H I, Hinf I, HpyCH4 III, HpyCH4 V, and Rsa I; More preferably, the restriction endonuclease comprises Bfa I, Dde I, Hae III, Mbo I, MluC I and Mse I; The reagent components of the kit may also selectively include one or more of the following: (2) End repair enzyme and its buffer; (3) A-Tailing Enzyme and its Buffer; (4) T4 DNA Ligase and its buffer; (5) DNA linker.

Citation Information

Cited By

  • Library contamination specific degradation method and kit based on methylation sensitive restriction enzyme

    CN120738328A