Whole genome re-sequencing-based beet genetic diversity analysis method

By combining PacBio HiFi and Oxford Nanopore sequencing technologies with Hi-C data, the shortcomings of sugar beet genome assembly and annotation were addressed, enabling genome research with high continuity and genetic diversity. This breakthrough overcomes the limitations of traditional homology alignment and improves the accuracy and genetic diversity coverage of genome research.

CN121109568APending Publication Date: 2025-12-12XINJIANG ACAD OF AGRI SCI (XINJIANG BRANCH OF CHINESE ACAD OF AGRI SCI)
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202511662635.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-13
Publication Date
2025-12-12

AI Technical Summary

Technical Problem

Current sugar beet genome research suffers from problems such as insufficient assembly continuity, incomplete annotation systems, inadequate coverage of genetic diversity, and gaps in epigenetic data integration, resulting in slow progress in genome research.

Method used

We constructed a highly continuous sugar beet reference genome using PacBio HiFi high-precision and Oxford Nanopore ultra-long read sequencing technologies combined with Hi-C spatial interaction data. We then combined RNA-seq, ATAC-seq, and whole-genome methylation sequencing to achieve precise annotation of gene structure and non-coding regions.

Benefits of technology

This study achieved highly continuous assembly of the sugar beet genome, reduced the missing rate of non-coding region annotations, broadly covered genetic diversity, fully integrated epigenetic information, significantly improved genome assembly quality, and increased efficiency in molecular marker development.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 9NMWCORRJUFZLGGKWPQY1KFNKCAC7MS6DBCXH9CH
    Figure 9NMWCORRJUFZLGGKWPQY1KFNKCAC7MS6DBCXH9CH
Patent Text Reader

Abstract

The invention provides a beet genetic diversity analysis method based on whole genome re-sequencing. The method comprises the steps of PacBio HiFi library construction, computer sequencing, data information analysis, assembly quality evaluation and Hi-C auxiliary genome assembly. The invention further provides application, and the high-continuity beet reference genome is used for analyzing beet intraspecific genetic diversity and developing high-density molecular markers. According to the method, the high precision of HiFi, the super-long read length of Nanopore and Hi-C space interaction data are combined for the first time, the problems of repeated sequence and complex SV analysis are solved, and all-round breakthrough of genome continuity, annotation integrity, genetic diversity coverage and epigenetic integration is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of genomics technology, specifically relating to a method for analyzing genetic diversity in sugar beets based on whole-genome resequencing. Background Technology

[0002] As the world's second-largest sugar crop, sugar beet (Beta vulgaris) has long relied on the earlier version RefBeet-1.2.2 for genome research. This version, released in 2013 by the Max Planck Institute for Molecular Genetics in Germany, is based on Illumina short-read sequencing technology (read length 100-150 bp), resulting in low assembly continuity, with a contig N50 of only 0.3 Mb and a scaffold N50 of 2.1 Mb.

[0003] Its main technical defects include: Insufficient assembly continuity: Because short-read technology cannot cross repetitive sequences and complex structural variations (SVs), there are a large number of unresolved broken regions in the genome, such as insufficient integrity of repetitive regions like telomeres and centromeres.

[0004] The annotation system is incomplete: relying on homology alignment methods, the annotation loss rate of non-coding RNAs (such as miRNAs and lncRNAs) and regulatory elements (enhancers and silencers) is as high as 40%.

[0005] Insufficient coverage of genetic diversity: Based on a single cultivar "KWS2320", there is a lack of information on intraspecific structural variation (SV) and copy number variation (CNV), resulting in low efficiency in molecular marker development.

[0006] Epigenetic data integration gaps: The lack of integration of epigenetic information such as DNA methylation and chromatin accessibility limits the analysis of stress resistance mechanisms.

[0007] Domestic and international research progress and bottlenecks: International research: Since 2017, some crops (such as lettuce) have begun to use long-read sequencing technologies (such as PacBioHiFi) to assemble T2T (telomere-to-telomere) genomes, but progress in the sugar beet field has been slow.

[0008] Domestic research: The team at Harbin Institute of Technology has made progress in constructing the genetic map of sugar beets, but still lacks systematic and high-precision genomic resources. Summary of the Invention

[0009] The technical problem to be solved by this invention is to provide a method for analyzing the genetic diversity of sugar beets based on whole-genome resequencing, which addresses the shortcomings of the prior art. This method is the first to combine HiFi high precision, Nanopore ultra-long reads and Hi-C spatial interaction data to solve the problem of repetitive sequences and complex SV analysis, and achieves a comprehensive breakthrough in genome continuity, annotation integrity, genetic diversity coverage and epigenetic integration.

[0010] To solve the above-mentioned technical problems, the technical solution adopted by the present invention is: a method for analyzing the genetic diversity of sugar beets based on whole-genome resequencing, the method being as follows: S1, PacBio HiFi Library Construction: S101. Extract DNA samples and perform genomic DNA testing; After the DNA sample quality tests described in S102 and S101 are passed, the PacBio HiFi library is constructed using the following method: S10201. Based on the detection results of the genomic DNA, perform Megruptor fragmentation on the DNA to obtain fragmented DNA. S10202, Repairing DNA damage and end repair obtained in S10201 to obtain repaired DNA; S10203. Connect both ends of the repaired DNA fragment obtained in S10202 to SMRT dumbbell adapters to construct an SMRTbell library. S10204. The SMRTbell library constructed in S10203 is selectively purified using Sage ELF or BluePippin to obtain the PacBio HiFi library. S2, Sequencing: The size of the PacBio HiFi library constructed in S10204 was detected by capillary electrophoresis using Fragment Analyzer. After calculation by PacBio Caculater, the PacBio HiFi library was bound to the SMRT bell template with sequencing primers and sequencing enzymes in proportion to form a complete SMRT bell library. Then, it was entered into the SMRT cell via diffusion loading and sequenced on the Sequel II / IIe / Revio platform. S3. Data and Information Analysis: S301, Data Output and Quality Control: The HiFi reads from the S2 sequencing machine are output, and basic data quality control is performed. The following information can be obtained by analyzing the output and length of the resulting HiFi reads: I. Number of HiFi Reads; II. Number of HiFi Read Bases; III. Average Length of HiFi Reads; IV. HiFi Read N50; V. Depth(X), sequencing depth; S302, Contig Assembly: The HiFi reads obtained from S301 were assembled into contigs using the Hifiasm software to obtain the assembled contigs sequence. S303 and Contig assembled sequences de-hybridized: The assembled contig sequences obtained from S302 were de-heterozygous using the Purge_dups software to obtain a preliminary assembled genome sequence. The following information was collected: I. Number of contigs in the third-generation assembly; II. Size of the assembled genome; III. GC content of the assembled genome; IV. Shortest contig length; V. Contig N50 size; VI. Contig N90 size; VII. Longest contig size; S4. Quality assessment of the assembled genome: The assembly quality of the preliminary assembled genome sequences obtained in S303 was assessed using methods including data return ratio assessment, GC base depth assessment, and BUSCO integrity assessment. S5, Hi-C assisted genome assembly: S501, Hi-C library preparation and sequencing; S502, Information Analysis: Sequencing data was obtained from Hi-C library construction and sequencing in S501. After basic data quality control, Hi-C-assisted genome assembly was performed to obtain the genome sequence at the chromosome level. Then, the accuracy of the assembly results was evaluated. The specific method is as follows: S50201, Basic Data Quality Control: Raw reads obtained from Hi-C library construction and sequencing in S501 are filtered to obtain clean reads. The data filtering method is as follows: I. Filtering for N-type Reads: When the number of N reads in a single-end sequencing read is greater than 3, remove the paired reads. II. Low-quality read filtering: When the proportion of bases with a quality value lower than 5 in a single-end sequencing read is ≥20%, the paired read is removed; III. Reads with Connectors Filtering: Removes reads containing adapters; S50202, Species Pollution Assessment: 10,000 reads were randomly selected from the filtered clean reads obtained from S50201 and compared with the NCBI non-redundant nucleotide database NT library using BLAST software. The species information of the matched species was checked. If the matched species was a closely related genus of the tested species, the sample was considered to be free of exogenous contamination. S50203, Hi-C auxiliary assembly: The filtered clean reads obtained in S50201 are aligned with the preliminary assembled genome sequence obtained in S303 to screen for valid Hi-C data. The valid Hi-C data is then used to assist in the assembly of the preliminary assembled genome sequence, constructing a chromatin interaction matrix. Based on the three-dimensional spatial structure characteristics of the chromatin, contigs / scaffolds are clustered, sorted, and oriented to finally assemble the chromosome-level whole genome sequence. The main steps of Hi-C-assisted assembly are as follows: S5020301. Use HICPRO software to align the filtered Clean reads obtained in S50201 with the preliminary assembled genome sequence obtained in S303. Filter and correct the alignment results, and select the sequences that are aligned to the genome simultaneously in Reads1 and Reads2 and have unique alignment positions for subsequent analysis. S5020302. Using 3D-DNA for Hi-C assisted assembly, a visual interactive map is created. After 3D-DNA assembly, any mis-sorted or misoriented sequences are manually adjusted using Juicebox based on the interaction heatmap to obtain the final assembly result. This includes the following steps: I. Filter out small contigs / scaffolds; II. Perform Locus interaction frequency consistency analysis on the remaining Contigs / Scaffolds. Contigs / Scaffolds with erroneous splicing are segmented. After segmentation, retain consistent Contigs / Scaffolds segments and remove inconsistent parts. III. Based on the strength of interactions between sequences, anchor, sort, and orient the obtained consistent sequences to establish preliminary reliable scaffolds for chromosome length; IV. For each initially constructed chromosome scaffold, calculate the overlapping regions between its constituent consistent sequences, merge the consistent sequences based on the overlapping regions, obtain the final chromosome length scaffolds, and generate the final assembly result; S5020303. Evaluate the accuracy of the assembly results: After Hi-C-assisted assembly is completed, verify the accuracy of the assembly results and calculate the interactions between and within chromosomes. If intrachromosomal interactions are stronger than interchromosomal interactions, and interactions between closely spaced chromosomes are stronger than interactions between distant chromosomes, then the assembly is correct. Then, HiCExplorer software is used to draw a genome interaction heatmap, and finally, the whole genome sequence at the chromosome level is assembled, which is the highly continuous sugar beet reference genome.

[0011] Preferably, the method for detecting genomic DNA in S101 includes: S10101, NanoDrop 2000 spectrophotometer: for detecting DNA concentration and purity; S10102, Pulsed-field electrophoresis: to analyze the degree of DNA degradation and whether there is RNA contamination; S10103, Qubit fluorescent dye method: for precise quantification of DNA concentration.

[0012] Preferably, the method for assembling Contigs in S302 is as follows: First, all-vs-all alignment is used to further correct sequencing errors; then, overlap alignment is performed on the corrected reads to construct a phased string graph; finally, continuous sequence Contigs are generated based on the overlap graph, which are the assembled Contig sequences.

[0013] Preferably, the specific method for removing heterozygous fragments from the assembled Contigs sequence obtained in S302 using Purge_dups software in S303 is as follows: First, use Minimap2 software to align HiFi data to the genome sequence of the assembled Contigs sequence, and then remove heterozygous fragments based on the coverage distribution of the alignment reads and the alignment score of the sequence.

[0014] Preferably, the Hi-C library preparation and sequencing method described in S501 is as follows: S50101. Treat beet cells with paraformaldehyde, a cell cross-linking agent, to obtain cells with DNA and protein cross-linking. S50102, the cells whose DNA and protein are cross-linked by lysing S50101 are extracted, and then treated with restriction endonuclease to obtain the enzyme-digested DNA; S50103. The DNA obtained from S50102 after enzyme digestion is subjected to end repair, and biotin is added to label the oligonucleotide ends to obtain repaired DNA. S50104. The repaired DNA obtained in S50103 is treated with nucleic acid ligase to ligate adjacent DNA fragments, resulting in ligated DNA. S50105. Digest the protein at the junction of the ligated DNA obtained in S50104 with protease to de-crosslink the protein and DNA, and obtain de-crosslinked DNA. S50106. Purify and recover the decrosslinked DNA obtained from S50105, break the DNA with DNA enzymes, label the DNA capture, add sequencing adapters, and construct a sequencing library. S50107. Using the Illumina NovaSeq6000 sequencing platform, the sequencing library obtained in S50106 was sequenced in PE150 mode to complete Hi-C library construction and sequencing.

[0015] The present invention also provides the application of the highly continuous sugar beet reference genome constructed by the above analysis method, which is used to analyze intraspecific genetic diversity in sugar beets and to develop high-density molecular markers.

[0016] Compared with the prior art, the present invention has the following advantages: 1. This invention constructs a highly continuous sugar beet reference genome at the chromosome level by integrating high-precision long-read sequencing (PacBio HiFi), ultra-long-read sequencing (Oxford Nanopore), and Hi-C technology.

[0017] 2. This invention combines RNA-seq, ATAC-seq, and whole-genome methylation sequencing (BS-seq) to achieve precise annotation of gene structure, non-coding regions, and epigenetic regulatory networks.

[0018] 3. Using this invention, a pan-genome covering 15 cultivated and wild species can be constructed, intraspecific genetic diversity can be analyzed, and high-density molecular markers can be developed.

[0019] 4. This invention provides a standardized genome construction process applicable to crops with complex genomes, such as maize and potato.

[0020] The present invention will be further described in detail below with reference to the embodiments. Detailed Implementation

[0021] De novo sequencing is the deep sequencing of the entire genome of an individual without relying on a reference genome. The genome sequence is then assembled and annotated using bioinformatics methods to obtain a complete genome sequence map of the species. This map can be used directly for research on the origin, evolution, and adaptation of the species, or to provide a usable reference genome.

[0022] The continuity and quality of genome assembly depend on sequencing technology and algorithms. First-generation sequencing (Sanger sequencing) has low throughput and high cost. While second-generation sequencing (NGS) solved the throughput problem, its short read lengths posed a significant challenge to genome assembly. Third-generation single-molecule long-read sequencing technologies, such as PacBio, have overcome this read length issue, easily traversing complex regions of the genome, resulting in significantly improved continuity and quality of genome assembly. In recent years, third-generation sequencing technologies, including PacBio, have been widely used in genome assembly.

[0023] HiFi reads (High Fidelity reads) are sequencing sequences generated by PacBio sequencing in CCS (Circular Consensus Sequencing) mode, combining long read lengths with high accuracy. In CCS sequencing, enzyme read lengths are comparable to or even longer than in CLR sequencing (over 100 kb), but the insert fragment is only 10-20 kb. Therefore, during sequencing, the enzyme performs rolling circular sequencing around the DNA template (insert fragment), meaning the insert fragment is sequenced multiple times. Random sequencing errors caused in a single sequencing run can be self-corrected by algorithms, ultimately producing sequences with both long read lengths (10-20 kb) and high accuracy (>99%) (HiFi reads). When using HiFi reads for genome assembly, there is no need to use high-accuracy second-generation short-read sequencing data for error correction, offering significant advantages in improving genome assembly continuity and quality, reducing computational resource consumption, and shortening the assembly cycle.

[0024] Within a cell, the DNA in chromosomes must be physically folded and packaged to fit into a small space. Therefore, chromosome segments that are geographically distant along a linear chromosome are often physically close together. Hi-C (High-throughput chromosome conformation capture) is a chromosome conformation capture technology that combines high-throughput sequencing to capture spatially adjacent DNA fragments in natural chromatin and sequence them, allowing for the study of spatial relationships between chromatin DNA across the entire genome.

[0025] Generally, different regions of the same chromosome have a higher contact frequency with each other than with other chromosomes, and genomic regions that are closer in sequence within the same chromosome usually have more frequent contact than genomic regions that are farther apart in sequence. These characteristics enable Hi-C technology to be used for chromosome-level genome-assisted assembly. Specifically, it leverages the higher contact frequency between chromosomes to cluster contigs / sacffolds; and utilizes the higher contact frequency between genomic regions that are closer in sequence within the same chromosome to sequence and orient contigs / sacffolds from the same chromosome. Through Hi-C-assisted genome assembly, the assembly of contig / sacffold sequences can be elevated to the near-chromosome / chromosome level. Currently, Hi-C technology has become the primary choice for chromosome-level genome assembly and is widely used in the assembly of animal and plant genomes.

[0026] Example 1 The method for analyzing genetic diversity in sugar beets based on whole-genome resequencing in this embodiment is as follows: S1, PacBio HiFi Library Construction: S101. Extract DNA samples and perform genomic DNA testing, including the following methods: S10101, NanoDrop 2000 spectrophotometer: for detecting DNA concentration and purity; S10102, Pulsed-field electrophoresis: to analyze the degree of DNA degradation and whether there is RNA contamination; S10103, Qubit fluorescent dye method: for precise quantification of DNA concentration; After the DNA sample quality tests described in S102 and S101 are passed, the PacBio HiFi library is constructed using the following method: S10201. Based on the detection results of the genomic DNA, perform Megruptor fragmentation on the DNA to obtain fragmented DNA. S10202, Repairing DNA damage and end repair obtained in S10201 to obtain repaired DNA; S10203. Connect both ends of the repaired DNA fragment obtained in S10202 to SMRT dumbbell adapters to construct an SMRTbell library. S10204. The SMRTbell library constructed in S10203 is selectively purified using Sage ELF or BluePippin to obtain the PacBio HiFi library. Sage ELF and BluePippin are two fully automated nucleic acid fragment selection and precise cutting systems developed by Sage Science, Inc. in the United States; PacBio HiFi (High Fidelity) is a long-read sequencing library developed by PacBio Inc. based on circular conformity sequencing technology. S2, Sequencing: The size of the PacBio HiFi library constructed in S10204 was detected by capillary electrophoresis using Fragment Analyzer. After calculation by PacBio Caculater, the PacBio HiFi library was bound to the SMRT bell template with sequencing primers and sequencing enzymes in proportion to form a complete SMRT bell library. Then, it was entered into the SMRT cell via diffusion loading and sequenced on the Sequel II / IIe / Revio platform. The SMRT cell is the core reaction vector of the PacBio sequencing platform. Its core structure contains millions of zero-mode waveguide pores (ZMW). Diffusion loading is a key step for the SMRTbell template to enter the ZMW pores, and its principle is based on the free diffusion of molecules. The PacBio Sequel II / IIe / Revio sequencing system uses SMRT Cell as the sequencing vector. Polymerase captures SMRT Bell library sequences and anchors them at the bottom of the zero-mode waveguide aperture. Four different fluorescently labeled dNTPs randomly enter the bottom of the zero-mode waveguide aperture. During chain synthesis, the fluorescence of different fluorescently labeled nucleotides is excited. Based on the type and residence time of different fluorescence signals, the signals of different bases are captured, thereby obtaining the base sequence information of the sequencing fragment.

[0027] S3. Data and Information Analysis: S301, Data Output and Quality Control: The HiFi reads from the S2 sequencing output undergo basic data quality control. Quality control is performed using the PacBio Sequel IIe / Revio platform for sequencing. The instrument directly processes the output HiFi reads using its built-in SMRTLink pipeline. The following information can be obtained by analyzing the output and length of the obtained HiFi reads: I. Number of HiFi Reads; II. Number of HiFi Read Bases; III. Average Length of HiFi Reads; IV. HiFi Read N50; V. Depth(X), sequencing depth; S302, Contig Assembly: Because genomic DNA needs to be fragmented before sequencing, algorithms are used to assemble the sequencing sequences to obtain a complete genome sequence—a process known as genome assembly. During sequence assembly, sequences from the same genomic region are assembled into longer contigs based on the overlap between reads. The HiFi reads obtained from S301 were assembled into contigs using the Hifiasm software to obtain the assembled contigs sequence. Hifiasm is used to assemble HiFi reads. Hifiasm is an assembly software developed for HiFi reads. It preserves haplotype information as much as possible during assembly and can quickly complete the assembly of very large genomes.

[0028] The method for assembling Contigs in this step is as follows: First, all-vs-all alignment is used to further correct sequencing errors; then, overlap alignment is performed on the corrected reads to construct a phased string graph; finally, continuous sequence Contigs are generated based on the overlap graph, which are the assembled Contig sequences. S303 and Contig assembled sequences de-hybridized: The original assembled genome sequence usually contains heterozygous segments. The higher the heterozygosity, the more redundant heterozygous segments there are. Therefore, it is necessary to use Purge_dups software to remove heterozygous segments from the assembled contigs sequence. Using Purge_dups software to remove heterozygous segments from the assembled contigs sequence obtained in S303, a preliminary assembled genome sequence was obtained. The following information was obtained: I. Number of contigs in the third-generation assembly; II. Size of the assembled genome; III. GC content of the assembled genome; IV. Shortest contig length; V. Contig N50 size; VI. Contig N90 size; VII. Longest contig size; The specific method for removing heterozygous fragments from the assembled Contigs sequence obtained in S303 using Purge_dups software in this step is as follows: First, use Minimap2 software to align HiFi data to the genome sequence of the assembled Contigs sequence, and then remove heterozygous fragments based on the coverage distribution of the alignment reads and the alignment score of the sequence. Purge_dups is an open-source tool specifically designed for deredundancy removal after genome assembly. It is primarily used to remove haplotigs and assembly overlaps in highly heterozygous diploid genomes to improve the accuracy and continuity of assembly results. Its core principle is to identify and filter redundant sequences by analyzing the read depth of sequencing data and the self-alignment information of the assembly, thereby optimizing genome quality. Minimap2 is a high-efficiency, lightweight sequence alignment tool primarily used to align long-read sequencing data (such as PacBio and Oxford Nanopore) or short-read data (such as Illumina) to a reference genome or assembled sequence. It is characterized by its speed and low memory footprint, and is widely used in genomics research. S4. Quality assessment of the assembled genome: The assembly quality of the preliminary assembled genome sequences obtained from S303 was assessed using methods including data return ratio assessment, GC base depth assessment, and BUSCO integrity assessment. Step 1: Data alignment evaluation: Data alignment evaluation compares the sequencing data of the same sample with the assembled genome and calculates the alignment ratio. A higher alignment ratio indicates a more complete genome assembly and higher quality. Using the default parameters of minimap2, the HiFi data and the final assembled genome are aligned, and the alignment rate is calculated. Based on the alignment results of the sequencing data and the assembled sequence in the previous step, the genome coverage is calculated. A genome coverage is defined as the number of covered bases ÷ the total sequence length. Step 2: GC base coverage depth assessment: After the genome is assembled, it is divided into 10kb windows, and the GC content and average coverage depth in each window are calculated to understand the GC base content of the species and whether there is contamination from sequencing data of other species in the assembly results. Step 3: BUSCO integrity assessment: The BUSCO assessment uses a single-copy orthologous gene library, combined with software such as tblastn, augustus and hmmer, to assess the integrity of the assembled genome using the passeriformes_odb10 dataset. A position is defined as covered if there is data covering it. The genome coverage is calculated as the number of covered bases divided by the total sequence length. This yields a genome coverage of 100.00% and an average depth of 40.73. S5, Hi-C assisted genome assembly: S501 and Hi-C library construction and sequencing methods are as follows: S50101. Treat beet cells with paraformaldehyde, a cell cross-linking agent, to cross-link DNA and protein, fix the DNA conformation, and obtain cells with cross-linked DNA and protein. S50102, the cells whose DNA and protein are cross-linked by lysing S50101 are extracted, and then treated with restriction endonuclease to obtain the enzyme-digested DNA; S50103. The DNA obtained from S50102 after enzyme digestion is subjected to end repair, and biotin is added to label the oligonucleotide ends to obtain repaired DNA. S50104. The repaired DNA obtained in S50103 is treated with nucleic acid ligase to ligate adjacent DNA fragments, resulting in ligated DNA. S50105. Digest the protein at the junction of the ligated DNA obtained in S50104 with protease to de-crosslink the protein and DNA, and obtain de-crosslinked DNA. S50106. Purify and recover the decrosslinked DNA obtained from S50105, break the DNA with DNA enzymes, label the DNA capture, add sequencing adapters, and construct a sequencing library. S50107. Using the Illumina NovaSeq6000 sequencing platform, the sequencing library obtained in S50106 was sequenced in PE150 mode to complete Hi-C library construction and sequencing. S502, Information Analysis: Sequencing data was obtained from Hi-C library construction and sequencing in S501. After basic data quality control, Hi-C-assisted genome assembly was performed, including clustering, sorting, and orientation of the initial assembled genome sequences to obtain chromosome-level genome sequences. Then, the accuracy of the assembly results was evaluated. The specific method is as follows: S50201, Basic Data Quality Control: Raw reads obtained from Hi-C library construction and sequencing in S501 are filtered to obtain clean reads. The data filtering method is as follows: I. Filtering for N-type Reads: When the number of N reads in a single-end sequencing read is greater than 3, remove the paired reads. II. Low-quality read filtering: When the proportion of bases with a quality value lower than 5 in a single-end sequencing read is ≥20%, the paired read is removed; III. Reads with Connectors Filtering: Removes reads containing adapters; S50202, Species Pollution Assessment: 10,000 reads were randomly selected from the filtered clean reads obtained from S50201 and compared with the NCBI non-redundant nucleotide database NT library using BLAST software. The species information of the matched species was checked. If the matched species was a closely related genus of the tested species, the sample was considered to be free of exogenous contamination. Hi-C standard libraries have a standard chimeric structure, allowing paired reads to align to two different locations on the genome. Due to the unique structure of Hi-C data, HiCPRO, a software specifically designed for Hi-C data alignment, was used. HiCPRO tracks the overall alignment rate and categorizes reads based on their alignment status. Low-quality alignments are filtered out, as are single-end alignments that fail to provide interaction signals for subsequent assembly. Multiple alignments also introduce interaction noise. Even within the uniquely aligned data, other types of unreliable data, such as re-liagations and same-circularized reads, can interfere with data analysis and are therefore filtered out, retaining only valid read pairs for genome assembly.

[0029] S50203, Hi-C auxiliary assembly: The filtered clean reads obtained in S50201 are aligned with the preliminary assembled genome sequence obtained in S303 to screen for valid Hi-C data. The valid Hi-C data is then used to assist in the assembly of the preliminary assembled genome sequence, constructing a chromatin interaction matrix. Based on the three-dimensional spatial structure characteristics of the chromatin, contigs / scaffolds are clustered, sorted, and oriented to finally assemble the chromosome-level whole genome sequence. The main steps of Hi-C-assisted assembly are as follows: S5020301. Use HICPRO software to align the Hi-C data (the filtered Clean reads obtained in S50201) with the preliminary assembled genome sequence obtained in S303. Filter and correct the alignment results, and select the sequences that are aligned to the genome simultaneously in Reads1 and Reads2 and have unique alignment positions for subsequent analysis. S5020302. Using 3D-DNA for Hi-C assisted assembly, a visual interactive map is created. After 3D-DNA assembly, any mis-sorted or misoriented sequences are manually adjusted using Juicebox based on the interaction heatmap to obtain the final assembly result. This includes the following steps: I. Filter out small contigs / scaffolds; Contigs: A contig is a continuous sequence formed by splicing short fragments (reads) generated by sequencing according to overlapping regions, and does not contain unknown bases (N); Contigs are the preliminary results of genome assembly, representing the complete sequence of a local genome region; for example, if reads cover a region without gaps, they are spliced ​​together to form a contig. Scaffolds (stent sequences): Based on contigs, the order and orientation between contigs are determined using paired-end or mate-pair sequencing data, and longer sequences are formed by filling gaps with unknown bases (N). The construction of scaffolds relies on physical distance information, allowing for a closer approximation of the actual chromosome structure. Hierarchical relationship: Reads → Contigs → Scaffolds → Chromosome-level assembly; II. Perform Locus interaction frequency consistency analysis on the remaining Contigs / Scaffolds. Contigs / Scaffolds with erroneous splicing are segmented. After segmentation, retain consistent Contigs / Scaffolds segments and remove inconsistent parts. III. Based on the strength of interactions between sequences, anchor, sort, and orient the obtained consistent sequences to establish preliminary reliable scaffolds for chromosome length; IV. For each initially constructed chromosome scaffold, calculate the overlapping regions between its constituent consistent sequences, merge the consistent sequences based on the overlapping regions, obtain the final chromosome length scaffolds, and generate the final assembly result; S5020303. Evaluate the accuracy of the assembly results: After Hi-C-assisted assembly is completed, verify the accuracy of the assembly results and calculate the interactions between and within chromosomes. If intrachromosomal interactions are stronger than interchromosomal interactions, and interactions between closely spaced chromosomes are stronger than interactions between distant chromosomes, then the assembly is correct. Then, HiCExplorer software is used to draw a genome interaction heatmap, and finally, the whole genome sequence at the chromosome level is assembled, which is the highly continuous sugar beet reference genome.

[0030] This invention employs a multi-technology combination: for the first time, it combines HiFi high precision, Nanopore ultra-long read length, and Hi-C spatial interaction data to solve the challenges of parsing repetitive sequences and complex SVs.

[0031] This invention achieves experiment-driven multi-omics annotation: breaking through the limitations of traditional homology alignment, it directly annotates non-coding elements using experimental data such as ATAC-seq and BS-seq, reducing the annotation missing rate from 40% to <5%.

[0032] This invention establishes a standardized pan-genome workflow: developing graph-based integration schemes applicable to crops such as maize and potatoes, covering intraspecific genetic diversity. This invention achieves a comprehensive breakthrough in genome continuity, annotation integrity, genetic diversity coverage, and epigenetic integration.

[0033] The highly continuous sugar beet reference genome (RefBeet-2.0) constructed in this invention achieves the following technological breakthroughs compared to the existing genome version (RefBeet-1.2.2) (Table 1).

[0034] Table 1 Performance of the high-continuity sugar beet reference genome constructed in this invention The highly continuous sugar beet reference genome constructed in this invention has the following advantages: Assembly quality was significantly improved: Contig N50 increased from 0.3 Mb to 10 Mb (a 30-fold increase), Scaffold covered 99% of the chromosome region, and telomere resolution was >95%.

[0035] Breakthrough in annotation depth: The non-coding RNA annotation loss rate was reduced from 40% to <5%, and the epigenetic regulatory network coverage reached 90%.

[0036] Breeding efficiency optimization: The density of molecular markers developed based on pangenome was increased by 5 times, and field validation showed that they were associated with stamen fertility and root yield traits by more than 70%.

[0037] Applications of the highly continuous sugar beet reference genome constructed in this invention: (I) High-precision genome assembly verification The cultivated strain "KWS2320" was assembled using RefBeet-2.0, achieving a contig N50 of 12.5 Mb, a scaffold N50 of 52.3 Mb, and telomere and centromere integrity of 95%. The chromosome spatial structure was verified using Hi-C, with a repetitive sequence resolution rate improved by 80%.

[0038] (II) Development and application of molecular markers Based on pan-genome data, 1,258 SVs were identified in stress-resistance-related regions, and 325 SNP markers were developed. In field trials, the markers showed 70% concordance with stamen fertility traits (63.33% fertile individuals and 76.67% sterile individuals).

[0039] (III) Verification of disease resistance gene function Disease-resistant RIPs were transferred into sugar beets using Agrobacterium-mediated transformation, resulting in 82 transgenic plants. Of these, 17 plants exhibited significant resistance to sugar beet brown spot disease, with a 30% increase in photosynthetic rate (25.8 μmol•m⁻²). 2 •s⁻ 1 Furthermore, nitrogen metabolism indicators had no negative impact.

[0040] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention in any way. Any simple modifications, alterations, and equivalent changes made to the above embodiments based on the inventive essence shall still fall within the protection scope of the present invention.

Claims

1. A method for analyzing genetic diversity in sugar beets based on whole-genome resequencing, characterized in that, The method is as follows: S1, PacBio HiFi Library Construction: S101. Extract DNA samples and perform genomic DNA testing; After the DNA sample quality tests described in S102 and S101 are passed, the PacBio HiFi library is constructed using the following method: S10201. Based on the detection results of the genomic DNA, perform Megruptor fragmentation on the DNA to obtain fragmented DNA. S10202, Repairing DNA damage and end repair obtained in S10201 to obtain repaired DNA; S10203. Connect both ends of the repaired DNA fragment obtained in S10202 to SMRT dumbbell adapters to construct an SMRTbell library. S10204. The SMRTbell library constructed in S10203 is selectively purified using Sage ELF or BluePippin to obtain the PacBio HiFi library. S2, Sequencing: The size of the PacBio HiFi library constructed in S10204 was detected by capillary electrophoresis using Fragment Analyzer. After calculation by PacBio Caculater, the PacBio HiFi library was bound to the SMRT bell template with sequencing primers and sequencing enzymes in proportion to form a complete SMRT bell library. Then, it was entered into the SMRT cell via diffusion loading and sequenced on the Sequel II / IIe / Revio platform. S3. Data and Information Analysis: S301, Data Output and Quality Control: The HiFi reads from the S2 sequencing machine are output, and basic data quality control is performed. The following information can be obtained by performing output and length statistics on the obtained HiFi reads: I. Number of HiFi Reads; II. Number of bases in HiFi Reads; III. Average length of HiFi Reads; IV. HiFi reads N50; V. Depth(X), sequencing depth; S302, Contig assembly: The HiFi reads obtained from S301 were assembled into contigs using the Hifiasm software to obtain the assembled contigs sequence. S303 and Contig assembled sequences are dehybridized: The assembled contig sequences obtained from S302 were de-heterozygous using the Purge_dups software to obtain a preliminary assembled genome sequence. The following information was collected: I. The number of contigs in the three-generation assembly results; II. Assembled genome size; III. GC content of the assembled genome; IV. Shortest contig length; V. Contig N50 size; VI. Contig N90 size; VII. Longest Contig size; S4. Quality assessment of the assembled genome: The assembly quality of the preliminary assembled genome sequences obtained in S303 was assessed using methods including data return ratio assessment, GC base depth assessment, and BUSCO integrity assessment. S5, Hi-C assisted genome assembly: S501, Hi-C library preparation and sequencing; S502, Information Analysis: Sequencing data was obtained from Hi-C library construction and sequencing in S501. After basic data quality control, Hi-C-assisted genome assembly was performed to obtain the genome sequence at the chromosome level. Then, the accuracy of the assembly results was evaluated. The specific method is as follows: S50201, Basic Data Quality Control: Raw reads obtained from Hi-C library construction and sequencing in S501 are filtered to obtain clean reads. The data filtering method is as follows: I. Filtering for N-type Reads: When the number of N reads in a single-end sequencing read is greater than 3, remove the paired reads. II. Low-quality read filtering: When the proportion of bases with a quality value lower than 5 in a single-end sequencing read is ≥20%, the paired read is removed; III. Reads with Connectors Filtering: Removes reads containing adapters; S50202, Species Pollution Assessment: 10,000 reads were randomly selected from the filtered clean reads obtained from S50201 and compared with the NCBI non-redundant nucleotide database NT library using BLAST software. The species information of the matched species was checked. If the matched species was a closely related genus of the tested species, the sample was considered to be free of exogenous contamination. S50203, Hi-C auxiliary assembly: The filtered clean reads obtained in S50201 are aligned with the preliminary assembled genome sequence obtained in S303 to screen for valid Hi-C data. The valid Hi-C data is then used to assist in the assembly of the preliminary assembled genome sequence, constructing a chromatin interaction matrix. Based on the three-dimensional spatial structure characteristics of the chromatin, contigs / scaffolds are clustered, sorted, and oriented to finally assemble the chromosome-level whole genome sequence. The main steps of Hi-C-assisted assembly are as follows: S5020301. Use HICPRO software to align the filtered Clean reads obtained in S50201 with the preliminary assembled genome sequence obtained in S303. Filter and correct the alignment results, and select the sequences that are aligned to the genome simultaneously in Reads1 and Reads2 and have unique alignment positions for subsequent analysis. S5020302. Using 3D-DNA for Hi-C assisted assembly, a visual interactive map is created. After 3D-DNA assembly, any mis-sorted or misoriented sequences are manually adjusted using Juicebox based on the interaction heatmap to obtain the final assembly result. This includes the following steps: I. Filter out small contigs / scaffolds; II. Perform Locus interaction frequency consistency analysis on the remaining Contigs / Scaffolds. Contigs / Scaffolds with erroneous splicing are segmented. After segmentation, retain consistent Contigs / Scaffolds segments and remove inconsistent parts. III. Based on the strength of interactions between sequences, anchor, sort, and orient the obtained consistent sequences to establish preliminary reliable scaffolds for chromosome length; IV. For each initially constructed chromosome scaffold, calculate the overlapping regions between its constituent consistent sequences, merge the consistent sequences based on the overlapping regions, obtain the final chromosome length scaffolds, and generate the final assembly result; S5020303. Evaluate the accuracy of the assembly results: After Hi-C-assisted assembly is completed, verify the accuracy of the assembly results and calculate the interactions between and within chromosomes. If intrachromosomal interactions are stronger than interchromosomal interactions, and interactions between closely spaced chromosomes are stronger than interactions between distant chromosomes, then the assembly is correct. Then, HiCExplorer software is used to draw a genome interaction heatmap, and finally, the whole genome sequence at the chromosome level is assembled, which is the highly continuous sugar beet reference genome.

2. The method for analyzing genetic diversity of sugar beets based on whole-genome resequencing according to claim 1, characterized in that, The method for genomic DNA detection described in S101 includes: S10101, NanoDrop 2000 spectrophotometer: for detecting DNA concentration and purity; S10102, Pulsed-field electrophoresis: to analyze the degree of DNA degradation and whether there is RNA contamination; S10103, Qubit fluorescent dye method: for precise quantification of DNA concentration.

3. The method for analyzing genetic diversity of sugar beets based on whole-genome resequencing according to claim 1, characterized in that, The method for assembling Contigs described in S302 is as follows: First, all-vs-all alignment is used to further correct sequencing errors; then, overlap alignment is performed on the corrected reads to construct a phased string graph; finally, continuous sequence Contigs are generated based on the overlap graph, which are the assembled Contig sequences.

4. The method for analyzing genetic diversity of sugar beets based on whole-genome resequencing according to claim 1, characterized in that, The specific method described in S303 for using Purge_dups software to remove heterozygous fragments from the assembled Contigs sequence obtained in S302 is as follows: First, use Minimap2 software to align HiFi data to the genome sequence of the assembled Contigs sequence, and then remove heterozygous fragments based on the coverage distribution of the alignment reads and the alignment score of the sequence.

5. The method for analyzing genetic diversity of sugar beets based on whole-genome resequencing according to claim 1, characterized in that, The method for Hi-C library preparation and sequencing described in S501 is as follows: S50101. Treat beet cells with paraformaldehyde, a cell cross-linking agent, to obtain cells with DNA and protein cross-linking. S50102, the cells whose DNA and protein are cross-linked by lysing S50101 are extracted, and then treated with restriction endonuclease to obtain the enzyme-digested DNA; S50103. The DNA obtained from S50102 after enzyme digestion is subjected to end repair, and biotin is added to label the oligonucleotide ends to obtain repaired DNA. S50104. The repaired DNA obtained in S50103 is treated with nucleic acid ligase to ligate adjacent DNA fragments, resulting in ligated DNA. S50105. Digest the protein at the junction of the ligated DNA obtained in S50104 with protease to de-crosslink the protein and DNA, and obtain de-crosslinked DNA. S50106. Purify and recover the decrosslinked DNA obtained from S50105, break the DNA with DNA enzymes, label the DNA capture, add sequencing adapters, and construct a sequencing library. S50107. Using the Illumina NovaSeq6000 sequencing platform, the sequencing library obtained in S50106 was sequenced in PE150 mode to complete Hi-C library construction and sequencing.

6. An application of a highly continuous sugar beet reference genome constructed using the analytical method described in any one of claims 1-5, characterized in that, The highly continuous sugar beet reference genome is used to analyze intraspecific genetic diversity in sugar beets and to develop high-density molecular markers.

Citation Information

Cited By

  • Lucid ganoderma binuclear genome assembly method based on haplotype analysis

    CN121565256A

  • A method for assembling ganoderma bicomb genome based on haplotype resolution

    CN121565256B