Sequencing library construction method and application thereof in genome typing and high-density physical map construction
By constructing genome sequencing libraries using the BacPhase method and employing BstB I digestion and PacBio HiFi sequencing technology, the challenges of genome typing and assembly in homologous polyploid species have been solved, achieving efficient and low-cost genome typing and assembly, and promoting the breeding and genome analysis of polyploid crops.
Patent Information
- Application Number
- CN202511458929.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-13
- Publication Date
- 2026-01-23
AI Technical Summary
Genome typing and assembly of autopolyploid species such as potatoes are difficult, and existing methods have significant gaps in accuracy and cost, especially in achieving accurate genome typing in autopolyploid species.
The BacPhase method was used to construct a genome sequencing library, digest it with restriction endonucleases such as BstB I, and combine it with PacBio HiFi sequencing technology. This method omits the steps of BAC library construction and single clone selection, and selects restriction fragments of appropriate length for PCR amplification to generate a high-efficiency and low-cost long insert paired-end sequencing library for genome typing and physical map construction.
It significantly improved genome assembly quality and marker generation efficiency, successfully identified a large number of bin markers, and improved the accuracy of genome typing. It is applicable to breeding and genome analysis of polyploid crops, especially trait mapping and genome selection.
Smart Images

Figure FT_1 
Figure FT_2 
Figure FT_3
Abstract
Description
Technical Field
[0001] This invention belongs to the field of bioinformatics technology, specifically relating to a sequencing library construction method and its application in genome typing and high-density physical map construction. Background Technology
[0002] Polyploidy is widespread in both prokaryotes and eukaryotes. Most crops are polyploid, posing challenges to genome assembly and genotyping. This problem is particularly pronounced in species such as potato, wheat, banana, sugarcane, and strawberry. While genotyping has been successfully achieved in diploid and allopolyploid species using algorithms such as Hifiasm (Cheng et al., 2021), autopolyploid species (such as potato) still face challenges. The high sequence similarity across large regions of their genome complicates accurate genotyping. Traditional genotyping methods, based on pedigree information or population inference, require additional information and can only genotype partial variations. Despite progress in chromosome-level assembly of polyploid genomes, significant gaps remain in achieving complete and accurate genome assembly, both in terms of genome quality and cost.
[0003] potato( Solanum tuberosum Potatoes (L.) are the world's most important tuber crop and the third most important food crop. In 2022, the global potato planting area was approximately 17.79 million hectares, with an annual yield of 374.78 million tons (FAOSTAT, 2024). Due to its widespread availability, the potato makes a significant contribution to food security, livelihoods, and income, and is one of the cornerstones of the agricultural and economic system. Published potato genomes include the initial haploid genome (PGSC, 2011), diploid genome (Feng et al., 2024; Zhou et al., 2020), and tetraploid genome (Bao et al., 2022; Hoopes et al., 2022; Serra Mari et al., 2024), as well as the T2T haploid genome completed by our laboratory team in this invention (Yang et al., 2023), etc. However, genomic typing of autotetraploid potato cultivars remains challenging, and genome assembly remains difficult, which complicates genome analysis, basic research, and breeding efforts.
[0004] In autopolyploid species (such as potato), due to high sequence similarity within the genome, genome phasing faces great challenges. In previous work, our laboratory developed a BAC anchor-based method (BAC end analysis protocol, related literature: Yang, X., Yang, Y., Ling, J., Guan, J., Guo, X., Dong, D., Jin, L., Huang, S., Liu, J., and Li, G. (2019). A high-throughput BAC end analysis protocol (BAC-anchor) for profiling genome assembly and physical mapping. Plant Biotechnology Journal 18: 364-372. 10.1111 / pbi.13203) for genome assembly and physical mapping using next-generation sequencing technology. Although this method improves efficiency and is more efficient and less costly than Sanger sequencing, it relies on library construction and has difficulties in accurate genome phasing (such as autopolyploidy), and is not suitable for accurate genome phasing of autopolyploid species. SUMMARY
[0005] The technical problem to be solved by the present application is how to accurately genotype autopolyploid species and / or how to accurately genotype potato and / or how to assemble or physically map the genome of autopolyploid species and / or how to accurately genotype species containing a large number of repetitive sequences.
[0006] To solve the above technical problems, the present application first provides a method for constructing a genomic sequencing library, which can comprise the following steps: A1) connecting genomic DNA fragments covering the whole genome of the target species to a cloning vector to obtain a mixture of recombinant cloning vectors; A2) using a restriction endonuclease to cut the mixture of recombinant cloning vectors to obtain cut fragments, and screening fragments longer than the length of the cloning vector from the cut fragments to obtain target cut fragments; A3) adding ligase to the target cut fragments to catalyze ligation to obtain self-ligation products, using a primer pair to perform PCR amplification with the self-ligation products as templates to obtain PCR amplification products, and the PCR products are the genomic sequencing library of the target species; The primer sequences in the primer pair contain sequences on the cloning vector. The above method does not include the steps of BAC library construction and picking single clones. The BAC library construction includes the step of transforming BAC vectors into bacterial cells.
[0007] In the above method, the restriction endonuclease can be greater than or equal to one.
[0008] In the above method, the restriction endonuclease can be R1 and / or R2, R1 is BstB I, and R2 is Cla I.
[0009] In the above method, A1) the genomic DNA fragments can be obtained by using a restriction endonuclease 2 to digest the whole genome of the target species; the restriction endonuclease 2 is different from the restriction endonuclease in A2). The genomic DNA fragments can also be obtained by other methods.
[0010] In one specific embodiment of the present application, the restriction endonuclease 2 is Hind III, and the restriction endonuclease 2 can also be EcoR I, BamH I, or other restriction endonucleases.
[0011] In the above method, the cloning vector can be a bacterial artificial chromosome.
[0012] In the above method, the target species can be a homologous polyploid organism. The homologous polyploid organism can be a plant, an animal, or / and a microorganism. In one specific embodiment of the present application, the homologous polyploid organism is potato.
[0013] To solve the above technical problems, the present application also provides a method for performing genome typing on a polyploid species genome, which can include: B1) obtaining a genomic sequencing library of a target polyploid species using the method in any one of claims 1-5; B2) sequencing the genomic sequencing library using a long-read sequencing platform to obtain raw sequencing reads; screening sequencing reads containing sequences of the cloning vector and containing recognition site sequences of the restriction endonuclease in A2) from the raw sequencing reads to obtain target sequencing reads; B3) splitting each of the target sequencing reads based on the recognition site sequences to obtain paired-end sequences; aligning the paired-end sequences to a reference genome of the target polyploid species to obtain alignment results; B4) based on the alignment results, retaining paired-end reads that both end sequences of the paired-end reads align to the same haplotype on the same chromosome of the reference genome of the target polyploid species as target paired-end reads; performing genome typing of the target polyploid species genome based on the target paired-end reads.
[0014] In the above method, the polyploid species can be a species that contains more than two sets of chromosomes in a cell.
[0015] In the above method, the target polyploid species can be a homologous polyploid species.
[0016] The homologous polyploid species can be a plant, an animal, or / and a microorganism. In one specific embodiment of the present application, the homologous polyploid species is potato.
[0017] The genomic sequencing library described above also falls within the protection scope of the present application 。
[0018] The length of the sequencing reads of the long-length sequencing platform can be greater than or equal to 8 kb. In one specific embodiment of the present application, the long-length sequencing platform is a PacBio HiFi sequencing platform, i.e., a long-read-based SMRT sequencing technology.
[0019] To solve the above technical problems, the present application also provides the following any one of the applications of the above method: C1) application in homologous polyploid species genome assembly and / or physical map construction; C2) application in detecting chromosome blocks in the genome of a polyploid species that have not undergone recombination.
[0020] The homologous polyploid species described above can be a plant, an animal, or / and a microorganism. In one specific embodiment of the present application, the homologous polyploid species is potato.
[0021] The present application provides BacPhase: a high-efficiency and low-cost long-insert paired-end sequencing method for constructing bin markers and promoting genome typing.
[0022] To address these limitations, the present application made two key improvements to the BAC end sequencing pipeline in the prior art. First, the enzyme selection was optimized, and BstB I was determined to be the most effective enzyme because it has more and more evenly distributed enzyme sites in the potato genome, which is crucial for the efficiency of the BacPhase method. Second, the present application adopted the PacBio HiFi sequencing technology, which has long read characteristics that are particularly suitable for assembling complex genomes, such as the potato genome containing 62% repetitive sequences (PGSC, 2011). These improvements significantly enhanced the scaffolding and bin marker distribution of the commercial potato variety C88 genome. The method of the present application provides a solid foundation for more efficient assembly of complex genomes and development of polyploid crop markers, and has great potential in promoting research and breeding in the field of polyploidy.
[0023] The present application develops a new strategy that omits the time-consuming and laborious process of BAC library construction and single clone picking; at the same time, by using the PacBio HiFi sequencing technology, the marker generation efficiency and genome typing accuracy are significantly improved. Specifically, the present application identified 39,484 bin markers in the C88 genome, highlighting the efficiency of the method. In addition, 59.58% of the scaffolds in the C88 genome were successfully anchored to the chromosomes, significantly improving the quality of genome assembly. This method not only promotes the development of high-resolution bin markers, but also provides a powerful tool for genome typing, especially for polyploid species. By advancing sequencing technology and marker construction, this method establishes a robust platform for assembling complex genomes and supports breeding applications of polyploid crops, such as trait mapping, genome selection, and accelerating crop improvement. BRIEF DESCRIPTION OF DRAWINGS
[0024] Figure 1 Workflow of BacPhase. A. Construction and sequence acquisition of BAC end library. B. HiFi sequencing, alignment and distribution of four mapping types on C88 genome. C. Define possible distances using reads mapping to the same haplotype. FP represents forward primer, and RP represents reverse primer.
[0025] Figure 2 Align the interval distance of BAC end sequences with enzyme sites in the same haplotype to the reference genome, and the alignment results of enzymes BstB I or Cla I. A is the alignment result of BstB I; B is the alignment result of Cla I. The horizontal axis is the interval distance on the genome, with each unit being 10 kb; the vertical axis is the number of BAC end sequences.
[0026] Figure 3Figure 6. Distribution of marker number in 60-100 kb intervals. A Distribution of marker number in 60-100 kb intervals in BAC end sequences generated by enzyme BstB I. B Distribution of marker number in 60-100 kb intervals in BAC end sequences generated by enzyme Cla I. The horizontal axis is the chromosome haplotype; the vertical axis is the number of markers of different types.
[0027] Figure 4 Figure 7. Marker density distribution. A Number of markers in 1 Mb windows in the C88 genome, different densities are indicated by different colors. B Enlarged view of three randomly selected regions (chr1_1, chr4_1 and chr5_1), red circles indicate fragments generated by enzyme BstB I, blue triangles indicate fragments generated by enzyme Cla I.
[0028] Figure 5 Figure 8. Average (A) and median (B) length of restriction fragments generated by different restriction enzymes in other polyploid crops.
[0029] Figure 6 Figure 9. Length frequency distribution of BstBI restriction fragments.
[0030] Figure 7 Figure 10. Distribution of BstBI (A) and ClaI (B) markers in the genome. DETAILED DESCRIPTION
[0031] The application will be further described in conjunction with the specific embodiments, and the examples given are only to illustrate the application, and are not intended to limit the scope of the application. The examples provided below can serve as a guide for further improvement by those of ordinary skill in the art, and do not in any way constitute a limitation on the application.
[0032] In the following examples, the experimental methods are routine methods, and are performed according to the techniques or conditions described in the literature in the art or according to the product instructions, unless otherwise specified. The materials, reagents, etc. used in the following examples can be obtained commercially, unless otherwise specified.
[0033] The plant materials used in the application are as follows: The potato used in the present application is the homologous tetraploid potato variety Cooperation-88 (C88) preserved in the laboratory. Related literature: Yang, X., Yang, Y., Ling, J., Guan, J., Guo, X., Dong, D., Jin, L., Huang, S., Liu, J., and Li, G. (2019). A high-throughput BAC end analysis protocol (BAC-anchor) for profiling genome assembly and physical mapping. Plant Biotechnology Journal 18:364-372. 10.1111 / pbi.13203.
[0034] The pCC1 BAC cloning vector used in the embodiments of the present application is preserved in the laboratory and is available to the public from the applicant for the purpose of repeating the present application only. Related literature: Yang, X., Yang, Y., Ling, J., Guan, J., Guo, X., Dong, D., Jin, L., Huang, S., Liu, J., and Li, G. (2019). A high-throughput BAC end analysis protocol (BAC-anchor) for profiling genome assembly and physical mapping. Plant Biotechnology Journal 18:364-372. 10.1111 / pbi.13203.
[0035] Example 1. Workflow of sequencing library construction and sequencing analysis (BacPhase) 1. Sequencing library construction 1.1 Obtaining genomic DNA fragments Genomic DNA was extracted from two-week-old seedlings of C88 tetraploid potato, and restriction enzyme Hind III was used for enzyme digestion, obtaining Hind III enzyme digestion products. The enzyme digestion system was 3 plugs containing DNA, 1 mL enzyme digestion buffer, 10-20 u Hind III. The enzyme digestion conditions were: 37°C enzyme digestion for 1 h, 0.5 M EDTA PH 8.0 was added, and the reaction was terminated after 1 h.
[0036] Select 100-300 kb through rubber cutting and recycling. Hind The product fragment was digested with enzyme III to obtain C88 digested genomic DNA.
[0037] 1.2 BAC Library Construction The digested genomic DNA of C88 obtained in step 1.1 and the processed... Hind The pCC1 BAC cloning vector purified by enzyme III was ligated to obtain a mixture of recombinant BAC cloning vectors containing digested genomic DNA fragments of all C88.
[0038] The ligation system (100 µL) consisted of: pCC1 BAC vector, 5 ng; digested genomic DNA fragment, 50 ng; 10 X ligase buffer, 10 µL; T4 DNA ligase (NEB), 20 U; and deionized water to a final volume of 100 µL. The ligation conditions were: 16 °C for 14 h. 1.3 Digestion of recombinant BAC cloning vector mixture using restriction endonucleases 1.3.1 Restriction endonuclease screening The effectiveness of BAC end sequencing largely depends on the number and uniformity of restriction enzyme sites in the genome. Selecting a suitable restriction endonuclease is crucial to achieving this goal. An ideal enzyme must meet specific criteria, such as high efficiency, specificity, and the ability to generate a large number of uniformly distributed restriction sites. Based on these criteria, this invention screened several candidate enzymes (BstBI, AatII, BfrBI, and NheI), and ultimately selected BstBI for further analysis because it exhibited superior performance compared to other enzymes.
[0039] This invention also analyzed the performance of different restriction endonucleases in other polyploid crops. Figure 5 The results showed that the length distributions of BstBI and ClaI restriction fragments were consistent, and both the median and mean data confirmed this pattern. Figure 5A and B, respectively. This indicates that BstBI and ClaI exhibit comparable cleavage efficiency across these diverse genomes, with no significant outliers in any species indicating abnormal distribution of enzyme recognition sites. Other commonly used restriction enzymes such as NsiI, BfrBI, BspDI, AvrII, Bsu36I, PmlI, and NheI also exhibit similar trends, maintaining consistent fragment length distributions across the tested genomes, further supporting their reliability for cross-species genomic studies. The median and average fragment lengths for RsrII, PacI, and AatII exhibit slight variations across different species, which can be attributed to differences in GC content, enzyme cleavage site distribution, or recognition site frequency, but the overall cleavage patterns remain comparable. MluI, PmeI, and BbvCI generate relatively longer fragments across different species, but the differences do not reach a significant level, insufficient to indicate low cleavage efficiency.
[0040] These findings suggest that multiple restriction enzymes, including BstBI, ClaI, BspDI, and AvrII, are effective in diverse plant genomes, supporting their broad applicability in genomic studies. Although individual optimization may be necessary for specific species, the data demonstrate that these enzymes are generally suitable for comparative studies. Further validation in more species will contribute to optimizing their use in specialized applications.
[0041] For the target species (potato genome), we further evaluated these enzymes by analyzing the fragment size characteristics (maximum, minimum, and standard deviation; Table S2). Based on the number of fragments, we prefer to select restriction enzymes that generate evenly distributed and moderately sized fragments. Referring to the fragment size criteria, we set the following thresholds: we prefer to select restriction enzymes with average / median fragment sizes between 3-5 kb matching the PacBio HiFi sequencing read length, and exclude fragments with total length exceeding 10 kb. We also exclude ultra-small fragments that would result in excessively low PacBio sequencing efficiency.
[0042] To evaluate the cleavage efficiency of BstB I, we first performed an electronic simulation analysis of the cleavage sites in the entire genome of C88, which showed a total of 168,248 cleavage sites, more than 21% than Cla I (138,535) and about 11 times more than Mul I (15,390) Figure 6(Mull enzyme site data can be found in the literature: Yang, X., Yang, Y., Ling, J., Guan, J., Guo, X., Dong, D., Jin, L., Huang, S., Liu, J., and Li, G. (2019). A high-throughput BAC end analysis protocol (BAC-anchor) for profiling genome assembly and physical mapping. Plant Biotechnology Journal 18:364-372. 10.1111 / pbi.13203). The average interval between adjacent enzyme cutting sites was 4,591 bp, which was significantly shorter than that of Cla I (5,849 bp) and Mui I (15,390 bp), and the interval ranged from a minimum of 6 bp to a maximum of 186,808 bp, which was less than the observed values of Cla I (236,389 bp) and Mui I (1,254,638 bp). The median interval of BstB I was 2,857 bp, which was much lower than that of Cla I (3,178 bp) and Mui I (29,680 bp) Figure 6 ). Compared with Cla I and Mui I, the enzyme cutting site distribution of BstB I was more uniform, making it an ideal candidate restriction enzyme for BAC end sequencing.
[0043] 1.3.2 Restriction enzyme digestion of the recombined BAC cloning vector mixture Restriction endonucleases (BstB I or Cla I) (New England Biolabs, Ipswich MA) were used to digest the recombined BAC cloning vector mixture obtained from step 1.2 to obtain initial enzyme cutting fragments. The target enzyme cutting fragments were obtained by screening fragments with a length greater than the length of the pCC1 BAC cloning vector (8128 bp) from the initial enzyme cutting fragments; the target enzyme cutting fragments include two types of sequences: BAC end sequences with pCC1 BAC cloning vectors (a represents Figure 1 in A of the middle, where the red section represents the end sequence of the genomic insert, and the black section represents the BAC vector sequence) and inserts without pCC1 BAC cloning vectors (b represents Figure 1 in A of the middle).
[0044] 1.4 Enzyme cutting fragment self-ligation and PCR amplification 1.4.1 Enzyme cutting fragment self-ligation Add T4 ligase to the target enzyme-digested fragments, and allow self-ligation of both types of sequences to obtain mixed ligaments. The ligation system is as follows: 50-100 ng DNA, 10 μL 10 X ligase buffer, 20 U T4 DNA ligase (NEB), X μL water, add water to 100 μL, 16 °C, 14-18 h.
[0045] 1.4.2 PCR amplification After self-ligation, the mixed ligaments are subjected to selective PCR amplification using the flanking primers FR / RP on the pCC1 BAC cloning vector (PCR amplification system 50 μL (Buffer 25 μL, dNTPs 10 μL, 5 uM primers 1.5 μL each, KOD FX Neo 1 μL), the reaction system is 94 °C for 2 min, 98 °C for 10 s, 58 °C for 30 s, 68 °C for 5 min, 30 cycles.), to obtain the C88 (BAC end) sequencing library (BstB I enzyme-digested sequencing library obtained using BstB I or Cla I enzyme-digested sequencing library obtained using Cla I), which is used for subsequent sequencing.
[0046] At the same time, the recombined BAC cloning vector mixture obtained in step 1.2 is subjected to PCR amplification using the flanking primers FR / RP on the pCC1 BAC cloning vector to obtain the C88 genomic control library, which is subjected to sequencing.
[0047] The sequences of the flanking primers FR / RP are as follows: FP (Foward Sequencing Primer): 5'-GGATGTGCTGCAAGGCGATTAAGTTGG-3'; RP (Reverse Sequencing Primer): 5'-CTCGTATGTTGTGTGGAATTGTGAGC-3'.
[0048] 1.5 Pacbio Hifi sequencing and BacPhase analysis Based on the C88 sequencing library and the C88 genomic control library obtained in step 1.4, Pabio Hifi CCS sequencing is performed on the PacBio platform (Anovo Biological Technology Co., Ltd. PacBio Sequel / Sequel II sequencing platform) Figure 1B). C88 genome itself as reference genome for analyzing sequencing reads (related literature: Bao, Z., Li, C, Li, G., Wang, P., Peng, Z., Cheng, L., Li, H., Zhang, Z., Li, Y., Huang, W., et al. (2022). Genome architecture and tetrasomic inheritance of autotetraploid potato. Molecular Plant 15:1211-1226. 10.1016 / j.molp.2022.06.009).
[0049] The analysis process mainly includes: The original Pacbio h Hifi sequencing reads were counted using Seqtk (GitHub - lh3 / seqtk: Toolkit for processing sequences in FASTA / Q formats). Sequences containing BstB I or Cla I restriction enzyme sites were extracted, and self-redundancy was removed by cd-hit default parameters to remove redundant sequences produced during PCR amplification (step reference: Li, W., and Godzik, A. (2006). Cd-hit: a fast program for clustering and comparing large sets of protein or nucleotide sequences. Bioinformatics 22:1658-1659. 10.1093 / bioinformatics / btl158). Next, control sequences from the genome rather than BAC ends were further removed by cd-hit default parameters.
[0050] After completing the above two steps, the reads obtained by the corresponding enzyme were split using the enzyme digestion recognition site sequence of the enzyme BstB I or Cla I, respectively, to obtain the corresponding paired ends.
[0051] and the end of the last cleavage site were used for alignment. The alignment types were classified as (paired-end) both ends uniquely aligned, one end uniquely aligned and the other end multiple aligned, both ends multiple aligned, and unaligned according to the situation of one end. In each alignment type, the cases of one enzyme site and multiple enzyme sites were further distinguished. Statistical analysis included alignment rate, alignment position, gap distance, and chromosome distribution parameters. Python scripts for statistics are available on GitHub (https: / / github.com / Jianyq-1 / BacPhase). R scripts were used to draw graphs (https: / / www.r-project.org / ).
[0052] Each sequencing read was divided into a pair according to the enzyme cleavage site (BstB I or Cla I) and aligned to the homologous tetraploid potato genome C88 (related literature: Bao, Z., Li, C., Li, G., Wang, P., Peng, Z., Cheng, L., Li, H., Zhang, Z., Li, Y., Huang, W., et al. (2022). Genome architecture and tetrasomic inheritance of autotetraploid potato. Molecular Plant 15:1211-1226. 10.1016 / j.molp.2022.06.009), and the target reads were screened according to the alignment results. The aligned paired reads were divided into three types: same haplotype, same chromosome, and different chromosomes. Based on the average insert size of about 90 kb evaluated from the initial BAC library, the reads mapped to the same haplotype and within the expected distance were further counted Figure 1 C).
[0053] After cleavage with BstB I or Cla I, sequencing reads contain different marker sequences Figure 1(B) For example, (on the cloning vector) forward sequencing primer sequence (FP), (on the cloning vector) reverse sequencing primer sequence (RP), Hind III, and restriction enzyme site (BstBI or ClaI) sequence. Based on the order of different markers, this invention classifies these reads into four types: reads containing restriction enzyme sites and forward primer to Hind III (FH), reads containing restriction enzyme sites and reverse primer to Hind III (HR), reads containing restriction enzyme sites, FH, and HR (FHR), and reads containing only restriction enzyme sites (RS). Within each type except FHR, reads containing restriction enzyme sites are further classified as internal reads or terminal reads.
[0054] For FH, the BstBI and ClaI restriction sites were primarily located in the inner regions of the (C88 genomic DNA fragment), with 29,462 and 15,473 reads, respectively, while the number of reads at the ends of the (C88 genomic DNA fragment) was lower (37 and 6, respectively). Similarly, for HR, the number of reads with restriction sites in the inner regions was 26,033 (BstBI) and 12,534 (ClaI), while the number of reads at the ends was 44 and 7. In FHR, the number of reads with restriction sites in the inner regions was lower, at 6,494 (BstBI) and 7,391 (ClaI). Furthermore, for RS, the number of reads with restriction sites in the inner regions was significantly higher than those with restriction sites at the ends (32,939 for BstBI and 24,928 for ClaI, while 59 and 23 were at the ends, respectively).
[0055] Furthermore, the control group sequences of this invention (i.e., the C88 genome control library obtained in step 1.2, which is directly amplified by PCR using FP / RP primers without BstBI or ClaI digestion) are genomic sequences. Even if they contain restriction enzyme sites, there are no gaps between the restriction enzyme sites. The control group sequences are used to filter out genomic DNA sequences obtained from sequencing BAC-terminated libraries that are identical to the control group. These results highlight the distribution bias of restriction enzyme sites.
[0056] The specific analysis results are shown below: A total of 280,442 reads were obtained for the control group, 197,442 reads for the BstBI enzyme digestion group (using BstBI to digest the C88 genomic BAC library), and 240,684 reads for the ClaI enzyme digestion group (using ClaI to digest the C88 genomic BAC library). The total read lengths were 735 Mb for the control group, 353 Mb for the BstBI enzyme digestion group, and 425 Mb for the ClaI enzyme digestion group. The average read lengths were 2,623 bp for the control group, 1,788 bp for the BstBI enzyme digestion group, and 1,763 bp for the ClaI enzyme digestion group (Supplemental Table 2).
[0057] Next, sequences containing the enzyme digestion sites were selected from each group of sequencing reads. In the ClaI enzyme digestion group, only 60,362 reads contained the ClaI enzyme digestion site, accounting for 25.08% of the original reads. The newly introduced BstBI enzyme digestion group performed better, generating 95,068 reads containing the BstBI enzyme digestion site, accounting for 48.15% of the original reads, almost twice as many as the ClaI enzyme digestion site. To remove the PCR redundancy introduced before sequencing, the present application performed self-redundancy filtering based on sequences containing enzyme digestion sites. After filtering, the BstBI enzyme digestion group obtained 64,779 reads, and the ClaI enzyme digestion group obtained 45,123 reads, of which 68.15% and 74.75% of reads, respectively, contained enzyme digestion sites. After removing the sequences of the control group (i.e., sequences directly from the genome rather than BAC ends), the BstBI enzyme digestion group finally obtained 64,144 reads, and the ClaI enzyme digestion group finally obtained 44,146 reads. These reads can be used as the final sequences for subsequent BAC end analysis. The above results show that BstBI generated more usable BAC end sequences than ClaI.
[0058] Finally, the present application evaluated whether the enzyme digestion with BstB I or Cla I was thorough during the process of constructing the BAC end library. The results showed that BstB I performed better than Cla I, with 86.30% of reads in the BstB I-digested group containing only one BstB I site, while 72.59% of reads in the Cla I-digested group contained only one Cla I site. The proportion of reads containing two enzyme sites was 11.63% for BstB I and 21.12% for Cla I. In addition, the maximum number of BstB I sites detected in BstB I reads was six, while the maximum number of Cla I sites detected in Cla I reads was 12. Therefore, BstB I has more and evenly distributed enzyme sites in the C88 genome, and the enzyme digestion is more thorough.
[0059] The selection of BstB I as the restriction enzyme for BAC end library construction significantly improved the efficiency of BAC end sequencing, both increasing the number of available reads and optimizing the uniform distribution of enzyme sites. Therefore, the use of BstB I enzyme for BAC library digestion of the C88 genome in the sequencing library construction process of the present application has greater advantages in uniformity, number of enzyme sites, etc. compared to the original enzyme combination (Cla I and Mlu) used by the laboratory in the previous period (related literature: Yang, X., Yang, Y., Ling, J., Guan, J., Guo, X., Dong, D., Jin, L., Huang, S., Liu, J., and Li, G. (2019). A high-throughput BAC end analysis protocol (BAC-anchor) for profiling genome assembly and physical mapping. Plant Biotechnology Journal 18:364-372. 10.1111 / pbi.13203).
[0060] These findings provide a solid foundation for BAC end sequencing of complex genomes (such as the potato genome) and great potential for future applications in marker development and genome assembly in polyploid crops.
[0061] Example 2. Analysis of the application of BacPhase of the present application 1. BacPhase can effectively assist the assembly of complex genomes The alignment results of BacPhase in Example 1 show that BacPhase can effectively support the assembly of complex genomes, which is verified from the results of the final sequence alignment to the C88 genome from the BstB I or Cla I enzyme cutting sites. The final sequence from the enzyme BstB I or Cla I is split into paired ends according to the enzyme cutting site and aligned to the C88 genome.
[0062] For BstB I, the sequences containing one recognition site are mostly in the "multiple alignment" category (56.02%, 62,016), followed by the "unique alignment" category (36.10%, 39,970), and the proportion of the "unaligned" category is the smallest (7.88%) (Table 1). In the sequences containing multiple recognition sites, the same trend also occurs, and the proportion of the "multiple alignment" category is the largest (54.00%, 9,493), followed by the "unique alignment" category (31.43%, 5,525), and the proportion of the "unaligned" category is the smallest (14.57%) (Table 1). For Cla I, the distribution trend is similar. In the sequences containing one recognition site and multiple recognition sites, the proportion of the "multiple alignment" category is the highest (58.35%, 37,395 and 58.09%, 14,059, respectively), and the "unique alignment" category ranks second, with 32.32% (20,712) for sequences containing one recognition site and 28.58% (6,916) for sequences containing multiple recognition sites. Regardless of whether the sequences contain one recognition site or multiple recognition sites, the proportion of the "unaligned" category is the smallest, which is 9.34% and 13.33%, respectively (Table 1). Overall, for the two enzymes, the sequences in the "unaligned" category are the least, which highlights the effectiveness of the BAC end library construction and sequencing analysis in Example 1 of the present application. These alignment distributions are closely related to the proportion of non-repetitive sequences and repetitive sequences in the potato genome (PGSC, 2011), further confirming the targeting and successful alignment of the BAC end sequences. In addition, the proportion of sequences containing one or more enzyme cutting sites in the "unaligned" category is always low, indicating that the BAC end sequences in the BAC end library obtained by the present application are effective.
[0063] The alignment results show that BacPhase combined with BstB I or Cla I enzyme cutting provides an effective method for assembling complex genomes, generating high-quality and evenly distributed BAC end sequences that are highly consistent with the inherent complexity of polyploid genomes.
[0064] Table 1. Statistics of data aligned to C88
[0065] 2. Application of BacPhase in complex genome typing Based on the alignment results of BAC end sequences in the BAC end library in Example 1 to the C88 genome, the present application analyzed the alignment information according to chromosome and haplotype positions. Considering the composition of non-repetitive and repetitive sequences in the genome, the present application classified the alignments into three categories: (1) unique alignment (both ends uniquely aligned); (2) one end uniquely aligned and the other end multiply aligned; (3) multiply aligned (both ends multiply aligned). For each category, the present application examined the chromosome and haplotype positions to assess the distribution and potential regularities (Table 2).
[0066] Among the unique alignment category, there were the most sequences aligned to the same haplotype, a total of 10,815, which was consistent with the expectation. A large number of alignments occurred on different chromosomes (6,040 sequences), which could be due to the complexity of the genome and the challenges inherent to genome assembly. In contrast, there were relatively fewer sequences aligned to the same chromosome, only 526 appeared in both enzymes (Table 2). These results reflect the inherent complexity of the potato genome and suggest that the assembly can need further optimization. In the one end uniquely aligned and the other end multiply aligned category, the number of alignments to the same haplotype decreased slightly to 8,690. This decrease could be due to a reduction in regions where one end is non-repetitive and the other end is repetitive. A large number of alignments occurred on different chromosomes (21,158). This indicates a higher frequency of repetitive sequences, explaining why more sequences were aligned to different chromosomes. In the multiply aligned (multiple alignments) category, the present application focused on sequences aligned to the same haplotype, obtaining a total of 6,908 alignments (Table 2). This further emphasizes the complementarity of BstB I or Cla I in providing more accurate and consistent alignments in complex genomic regions.
[0067] Finally, the present application used sequences aligned to the same haplotype for the next round of inter-distance analysis (Table 2). This further validated the application of BacPhase in complex genome typing and highlighted the significant advantage of using BstB I or Cla I for more accurate and comprehensive genome assembly in polyploid crops such as potato.
[0068] Table 2. Alignment information on chromosomes and haplotypes
[0069] To assess the distance of BstB I and Cla I sequences in the same haplotype, the present application analyzed them according to the three mapping categories. Since the fragments linked to the vector were mainly selected in the range of 60-100 kb, the present application focused on this distance for further analysis (Table 2). Figure 2 .
[0070] For both enzymes, most distances fell within the 60-100 kb range. Specifically, for BstBI, 95.21% of distances in the unique mapping category fell within this range; for the one-end-unique, other-end-multiple mapping category, this proportion was 93.90%; and for the two-end-multiple mapping category, this proportion was 78.23% (considering only one enzyme site). For sequences with multiple enzyme sites, the proportion of 60-100 kb distances decreased slightly, but remained above 80% for the "one-end-unique, other-end-multiple mapping" category Figure 2 A). Overall, BstBI showed a very prominent proportion of 60-100 kb gaps, as expected, indicating that BstBI effectively captured these internal distance regions.
[0071] For Cla I, the proportion of 60-100 kb distances was 49.39% in the unique mapping category; 46.97% in the "one-end-unique, other-end-multiple mapping" category; and 31.75% in the "two-end-multiple mapping" category. When considering sequences with multiple enzyme sites, the proportion of 60-100 kb distances was always lower than for single enzyme site mapping Figure 2 B). Both enzymes showed a prominent peak in the 76-80 kb distance range Figure 2 C and D). Sequences mapping to the same haplotype and their internal distances within the 60-100 kb range can serve as valuable bin markers in genomic analyses. In subsequent analyses, the inventors refer to these distance regions as bin markers, which are crucial for further polyploid crop (e.g., potato) genome assembly and marker development.
[0072] Overall, the inventors' BacPhase performed excellently in complex genome typing, as seen from the high proportion of unique alignments, particularly within the 60-90 kb gap region. This consistent alignment pattern indicates a high accuracy of the original genome assembly, validating the integrity of the sequencing data.
[0073] 3. High-density bin marker physical map constructed by BacPhase Based on the alignment results of the BAC end sequences in the BAC end library in Example 1 to the C88 genome, to examine the distribution of each type of marker from the alignment results on chromosomes and haplotypes, the inventors performed visualization Figure 3 ). In the BstBI enzyme site category, most unique alignment markers Figure 3The markers for chromosome A (represented by dark blue / orange) are distributed across different chromosomes, with a higher number found on specific haplotypes (such as chr1_1, chr4_1, and chr8_1). These unique markers constitute the largest proportion, followed by markers with unique alignments at one end and multiple alignments at the other end. Figure 3 (The blue / orange color representation of A) contributes nearly half of the markers. Multiple alignment categories ( Figure 3 The light blue / orange markers (Chr5) constitute the smallest proportion, and these markers are more frequently found on chromosomes with fewer markers or shorter lengths, such as chr5_3, chr5_4, chr10_2, and chr12_1. This pattern reflects the inherent complexity of polyploid genomes, where the presence of highly similar sequences in pairs, triplets, and quadruplets on these chromosomes leads to variations in marker distribution (Bao et al., 2022). If a chromosome has fewer unique markers, the total number of markers on that haplotype will also be smaller, indicating a close correlation between the number of unique markers and the total number of markers on a haplotype.
[0074] A similar distribution pattern was observed for Cla I. Figure 3 In B, the color of the comparison marker indicates the same... Figure 3 A) Unique alignment markers are distributed across different chromosomes, with a higher number of markers on specific haplotypes (such as chr1_1, chr1_3, and chr4_2). Although there are fewer markers in the categories of unique alignment at one end, multiple alignment at the other end, and multiple alignment, chromosomes like chr7_2 and chr10_1 have a higher number of markers in these categories. Figure 3 (B)
[0075] Overall, the total number of markers was considerable, approximately one marker per 81 kb, highlighting the effectiveness of the sequencing method. BstBI generated 17,692 marker pairs across all chromosomes, while Cla I generated 2,050, providing valuable information on the number and location of markers. Figure 7 (A and B in the original text). Furthermore, the markers generated by multiple restriction sites in the Cla I sequence complement those generated by a single restriction site in the BstBI sequence, further highlighting the complementarity of these two enzymes in generating comprehensive and robust marker sets for complex genome analysis. BAC end sequencing is expected to be used for genome assembly and marker construction; therefore, uniform distribution of markers across different chromosomes and haplotypes is necessary. Thus, this invention constructs a marker density map for chromosomes and haplotypes. The results show that the markers are generally uniformly distributed, covering every haplotype on every chromosome (…). Figure 3 , Figure 4). The low density or blank marker regions correspond to two, three and four sets of highly similar regions in the C88 genomic sequence, further confirming the accuracy of the alignment method of the present application. It is worth noting that the number of markers on the four haplotypes varies (Figure 1 Figure 3 , Figure 4 ). A higher marker density was observed in the chr7_3 region, which can be affected by the specific recognition sequences of the restriction enzymes used, suggesting that this region can have more restriction enzyme specific recognition sequences (Figure 1 Figure 4 A). Nevertheless, the overall distribution on the chromosomes is still uniform.
[0076] To further investigate the number and location of the two source markers from enzymes BstB I and Cla I, the present application randomly selected three regions from Chr1_1, Chr4_1 and Chr5_1, each region 5 Mb (Figure 1 Figure 4 B). Each panel represents the marker location on the genomic coordinates, and different densities are observed in different regions. It is worth noting that there are obvious differences in the patterns of marker distribution, and in some regions red and blue markers often cluster together, while in other regions they are sparsely distributed. Overall, the two sources of markers are complementary in number and location (Figure 1 Figure 7 A and B).
[0077] In terms of enzyme performance, Cla I complements BstB I under both single restriction site and multiple restriction site conditions. In terms of the number of enzyme cleavage sites, the performance of a single enzyme cleavage site is always better than that of multiple enzyme cleavage sites.
[0078] 4. BacPhase anchors C88 unmounted scaffold to chromosomes Based on the alignment results of BAC end sequences in the BAC end library in Example 1 to the C88 genome, the BacPhase effectively anchored the un-mapped scaffolds of the C88 genome to the corresponding chromosomes. The results showed that multiple sequences could be mapped to the scaffolds by unique alignment or multiple alignment. In the case of unique alignment, paired end sequences were mapped to the chromosomes and scaffolds, facilitating the anchoring of these scaffolds to specific chromosomes. There were three paired end sequences in BstB I and two paired end sequences in Cla I that successfully mapped to the scaffolds (Table 3). For multiple alignment, since haplotype assignment had already been performed on each chromosome and scaffold, and there were four haplotypes, the present application could more accurately resolve the positions of the scaffolds. For example, when a sequence was mapped to chr9_1, chr9_2, chr9_4, and scaffold2099_3, all alignment positions were 996 base matches, which inferred that scaffold2099_3 was located on chr9_3 and its approximate position was inferred. By this method, 272-662 unequal scaffolds in BstB I and Cla I were successfully anchored to the chromosomes (Table 3). In addition, for paired end sequences with multiple alignment, according to different alignment types, 55-242 scaffolds were anchored to the C88 genome (Table 3). After merging and removing redundancies, the present application anchored a total of 815 scaffolds to the chromosomes, accounting for 59.58% of the total number of 1368 scaffolds. This method provides a solid framework for future research and application of complex polyploid genomes.
[0079] Table 3. Number of scaffolds anchored to chromosomes
[0080] Regarding the application scope of BAC end sequences, there are mainly two aspects, which are bin markers and genome assembly. As an electronic physical map, the present application proposes a new method for efficiently constructing a physical map. In the present application, a more efficient enzyme, BstB I, is introduced, which is nine times more efficient than the previously used enzyme Cla I (Table 2). The BstB I enzyme is used to digest the BAC library, and the BAC end sequences are obtained by sequencing. The BAC end sequences are then aligned to the C88 genome, and the results are used to anchor the un-mapped scaffolds to the corresponding chromosomes. The results show that multiple sequences can be mapped to the scaffolds by unique alignment or multiple alignment. In the case of unique alignment, paired end sequences are mapped to the chromosomes and scaffolds, facilitating the anchoring of these scaffolds to specific chromosomes. There are three paired end sequences in BstB I and two paired end sequences in Cla I that successfully map to the scaffolds (Table 3). For multiple alignment, since haplotype assignment has already been performed on each chromosome and scaffold, and there are four haplotypes, the present application can more accurately resolve the positions of the scaffolds. For example, when a sequence is mapped to chr9_1, chr9_2, chr9_4, and scaffold2099_3, all alignment positions are 996 base matches, which infers that scaffold2099_3 is located on chr9_3 and its approximate position is inferred. By this method, 272-662 unequal scaffolds in BstB I and Cla I are successfully anchored to the chromosomes (Table 3). In addition, for paired end sequences with multiple alignment, according to different alignment types, 55-242 scaffolds are anchored to the C88 genome (Table 3). After merging and removing redundancies, the present application anchors a total of 815 scaffolds to the chromosomes, accounting for 59.58% of the total number of 1368 scaffolds. This method provides a solid framework for future research and application of complex polyploid genomes. Figure 3In addition, the present invention also anchors some scaffolds to chromosomes. When developing markers, if located on these anchored scaffolds, their approximate positions can be determined. When developing markers using BacPhase, a lower sequencing depth is recommended, which can ensure both the density and uniform distribution, while also having a low cost. When assembling genomes or constructing T2T genomes, it is recommended to construct higher content of BAC end sequences and increase the sequencing depth. BAC end sequences can also be used for genome assembly and anchoring the position of sequences, such as anchoring scaffolds to chromosomes. For example, in the assembly of the C88 genome, 1,368 scaffolds were not assigned to chromosomes, while in the present invention, 59.58% of these scaffolds have already been positioned to chromosomes.
[0081] The sequence length between BAC end sequences in the present invention mainly ranges from 60-100 kb, although only the two ends are sequenced, but due to the similar long sequence length as Nanopore, it can be a good substitute for long sequence. However, the accuracy of Nanopore is much lower than PacBio HiFi, especially in the assembly of polyploid genomes with complex structures (such as potato), this low accuracy can be fatal, because these genomes contain a large number of homologous regions, such as diplotigs, triplotigs and tetralotigs. It is these complex regions that make genome assembly very difficult. In the C88 genome, 1034 progeny were used for haplotype assignment, which consumed a lot of manpower, material resources and financial resources (Bao et al., 2022). In another tetraploid commercial potato variety Otava, 717 pollen cells were used for haplotype assignment (Sun et al., 2022), and it is very difficult to obtain pollen single cells.
[0082] Table 4. The number of bin markers with paired BAC end sequences with a distance size of 60-100 kb in all haplotypes
[0083] BAC end sequencing with PacBio HiFi has an exceptionally high accuracy, usually around 99.999%, far exceeding the accuracy of Nanopore, which is around 95.64%-99.75% (https: / / nanoporetech.com / platform / accuracy). Nanopore sequencing errors, in addition to indels and mismatches, are mainly caused by homopolymer regions and tandem repeats, especially the high frequency of homopolymer deletion (Wick et al., 2019). These errors are usually confined to specific genomic sequences or regions, and are difficult to correct even with internal error correction or second-generation sequencing data, thus affecting the accuracy of genome assembly (Chen et al., 2021; Hotaling et al., 2023). In addition, studies have shown that inverted repeat sequences in the genome can reduce the quality of Nanopore sequencing, affecting the final sequence accuracy (Spealman et al., 2019). Therefore, in species with high repeat content or polyploidy, the accuracy of the repeat region may not be reliable. However, combined with BAC end sequences, different enzyme combinations, and library capacity of HiFi sequencing, these complex genomic structures can be accurately resolved. The key advantage of Nanopore sequencing is its ultra-long read length, ranging from 20 bp to 4 Mb (https: / / nanoporetech.com). This makes Nanopore sequencing suitable for assembling large genomes, even at the expense of lower accuracy. However, as the demand for genome assembly increases, especially for both length and quality, Nanopore sequencing is difficult to meet the requirements of complex genomic regions. In contrast, PacBio HiFi reads, although shorter (usually 15-30 kb), provide higher accuracy. If the BacPhase workflow and sequencing technology are improved in the future to extend the read length to more than 300 kb, these longer and high-precision sequences will surpass the ultra-long reads of Nanopore in both length and accuracy. BAC end sequencing with ultra-long read length capability and high accuracy can solve the length and quality problems in high-quality assembly of complex genomes and variant analysis.
[0084] The above describes the present application in detail. For those skilled in the art, without departing from the spirit and scope of the present application, and without unnecessary experiments, the present application can be implemented within a wider range of equivalent parameters, concentrations and conditions. Although the present application gives a special example, it should be understood that further improvements can be made to the present application. In summary, according to the principles of the present application, this application intends to include any changes, uses or improvements of the present application, including changes made by conventional techniques known in the art, which are outside the scope disclosed in this application.
Claims
1. A method of constructing a genomic sequencing library, characterized by: The method comprises the following steps: A1) connecting genomic DNA fragments covering the whole genome of a target species to a cloning vector to obtain a mixture of recombined cloning vectors; A2) using a restriction endonuclease to cut the mixture of recombined cloning vectors to obtain cut fragments, and screening fragments longer than the cloning vector from the cut fragments to obtain target cut fragments; A3) adding ligase to the target cut fragments to catalyze ligation to obtain self-ligation products, using a primer pair to perform PCR amplification with the self-ligation products as templates to obtain PCR products, which are genomic sequencing libraries of the target species; The primer sequences in the primer pair contain sequences on the cloning vector.
2. The method of claim 1, wherein: The restriction endonuclease is greater than or equal to one.
3. The method according to claim 1 or 2, characterized in that: The restriction endonuclease is R1 and / or R2, the R1 is BstB I, and the R2 is Cla I.
4. The method according to any one of claims 1-3, characterized by: A1) The genomic DNA fragments are obtained by using a restriction endonuclease 2 to cut the whole genome of the target species; the restriction endonuclease 2 is a restriction endonuclease different from the restriction endonuclease in A2).
5. The method according to any one of claims 1-4, characterized by: The cloning vector is a bacterial artificial chromosome.
6. A method of performing genome typing on a genome of a polyploid species, characterized by: The method comprises B1) using the method in any one of claims 1-5 to obtain a genomic sequencing library of a target polyploid species; B2) using a long-read sequencing platform to sequence the genomic sequencing library to obtain raw sequencing reads, and screening sequencing reads containing sequences of the cloning vector and containing recognition site sequences of the restriction endonuclease in A2) from the raw sequencing reads to obtain target sequencing reads; B3) splitting each of the target sequencing reads based on the recognition site sequences to obtain paired-end sequences, and aligning the paired-end sequences to a reference genome of the target polyploid species to obtain alignment results; B4) according to the alignment results, retaining paired-ends in which both of the end sequences are aligned to the same haplotype on the same chromosome of the reference genome of the target polyploid species to obtain target paired-ends, and performing genomic typing on the genome of the target polyploid species based on the target paired-ends.
7. The method of claim 6, wherein: The target polyploid species is a homologous polyploid species.
8. Constructing a genomic sequencing library using the method in any one of claims 1-5.
9. Any one of the following applications of the method in any one of claims 1-7: C1) application in assembling and / or constructing a physical map of a homologous polyploid species genome; C2) application in detecting chromosome blocks in which no recombination has occurred in a polyploid species genome.