RAD-seq low-cost genetic diversity evaluation system for endangered species
By using self-made reagent combinations and simplifying the library construction process, the problems of RAD-seq library construction complexity and low-quality DNA samples have been solved, enabling low-cost and efficient genetic diversity assessment, which is suitable for genetic diversity assessment of endangered species.
Patent Information
- Application Number
- CN202511689489.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-18
- Publication Date
- 2026-02-06
AI Technical Summary
The existing RAD-seq library preparation process is complex, requires expensive commercial kits and magnetic beads, making it difficult to apply to large-scale endangered species research. Furthermore, low-quality DNA samples can easily lead to library preparation failure, affecting the accuracy and reliability of genetic diversity assessment.
It employs a self-made reagent combination and a simplified library construction process, including a single-step enzyme digestion-ligation reaction, combined with DNA repair and concentration standardization, uses a universal adapter design, supports multiplex sample pooling sequencing, and performs SNP marker development and diversity index calculation through an automated process.
It significantly reduces reagent and operational costs, improves library construction success rate, ensures the accuracy and reliability of genetic diversity assessment, is suitable for large-scale endangered species research, and provides one-click genetic diversity assessment reports.
Smart Images

Figure CN121472391A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of biotechnology, in particular to a RAD-seq low-cost genetic diversity evaluation system for endangered species. BACKGROUND
[0002] Biodiversity is the core of the health and function of the Earth's ecosystem, which covers multiple levels such as species diversity, genetic diversity and ecosystem diversity. Among them, genetic diversity, as an important part of biodiversity, is the basis for species to adapt to environmental changes, resist diseases and maintain population survival and reproduction. For endangered species, accurate assessment of the level of genetic diversity not only can in-depth understand the evolutionary potential and survival status of the species, but also can provide key basis for formulating scientific and effective protection strategies, so as to reduce the risk of species extinction and maintain the balance and stability of the ecosystem. In order to balance the cost and throughput, the genomic sequencing technology (such as RAD-seq) is born. RAD-seq cuts the genome by a specific restriction enzyme, and sequences the DNA fragments on both sides of the enzyme cutting site, thereby realizing the development of multi-site single nucleotide polymorphism (SNP) markers without reference genome, which has been widely used in population genetics research.
[0003] The existing RAD-seq library construction process is complex, involving multiple steps of enzyme cutting, ligation, purification and fragment screening, which requires the use of a large number of expensive commercial library construction kits and magnetic beads. For an endangered species research project containing dozens to hundreds of valuable individuals, it seriously limits its large-scale and normalized application in protection practice, and the research samples of endangered species are usually obtained by non-damaging or non-invasive sampling, such as hair, feces, old leather, museum specimens, etc. These samples usually have very low DNA content and are severely degraded. The conventional RAD-seq technology has higher requirements for the integrity and concentration of DNA, and the use of such low-quality DNA templates is easy to cause library construction failure, low library complexity or dramatic data fluctuations, which ultimately affects the accuracy and reliability of genetic diversity evaluation. Therefore, the present application provides a RAD-seq low-cost genetic diversity evaluation system for endangered species. SUMMARY
[0004] In view of the deficiencies of the prior art, the present application provides a RAD-seq low-cost genetic diversity evaluation system for endangered species, which solves the problems of the existing RAD-seq library construction process, which is complex, involves multiple steps of enzyme digestion, ligation, purification and fragment screening, requires the use of a large number of expensive commercial library construction kits and magnetic beads, and seriously limits the large-scale and normalized application of the system in protection practice for an endangered species research project involving dozens to hundreds of valuable individuals, and the research samples of the endangered species are mostly obtained by non-invasive or non-invasive sampling, such as hair, feces, old leather, museum specimens, etc., which are usually extremely low in DNA content and severely degraded, and the conventional RAD-seq technology has higher requirements for the integrity and concentration of DNA, and the use of such low-quality DNA templates can easily lead to library construction failure, low library complexity or dramatic fluctuations in data volume, ultimately affecting the accuracy and reliability of genetic diversity evaluation.
[0005] To achieve the above object, the present application is implemented by the following technical scheme: a RAD-seq low-cost genetic diversity evaluation system for endangered species, comprising an evaluation system, the evaluation system comprising a DNA extraction and pretreatment module, a simplified RAD-seq library construction module, a high-throughput sequencing module and a genetic diversity analysis module, the DNA extraction and pretreatment module being used to extract low-quality DNA from non-invasive samples of endangered species and perform DNA repair and concentration standardization, the simplified RAD-seq library construction module using a low-cost self-made reagent combination to replace a commercial kit, reducing the steps of purification and fragment screening through a single-step enzyme digestion-ligation reaction and a universal adapter design, realizing the simplification of the library construction process, the high-throughput sequencing module sequencing the DNA fragments after library construction, and the genetic diversity analysis module developing SNP markers and calculating diversity indices through an automated process based on sequencing data, and outputting a genetic diversity evaluation report.
[0006] Preferably, the DNA extraction and pretreatment module comprises a DNA quality evaluation unit and a repair unit, and the repair unit uses an enzyme combination to perform end repair and gap filling on degraded DNA to improve DNA integrity and library construction success rate.
[0007] Preferably, the simplified RAD-seq library construction module uses a low-cost self-made reagent combination including a self-made restriction enzyme buffer, a ligase and a universal adapter, and the universal adapter contains a sample-specific barcode, allowing multiple samples to be mixed for sequencing and reducing reagent costs.
[0008] Preferably, the simplified RAD-seq library construction module uses a hot start PCR technique for library amplification, and reduces non-specific amplification and improves library complexity by optimizing annealing temperature and cycle number.
[0009] Preferably, the genetic diversity analysis module is integrated with a reference genome alignment unit and a de novo assembly unit, supports SNP calling without a reference genome, and uses a machine learning algorithm to filter low-quality data to ensure evaluation accuracy.
[0010] Preferably, the evaluation system further comprises a user interaction interface for inputting sample information, setting parameters, and visualizing results to realize one-key genetic diversity evaluation.
[0011] Preferably, in the simplified RAD-seq library construction module, the enzyme cutting-linking reaction is carried out in a single tube, the reaction time is shortened to within 30 minutes, and the operation steps and the risk of contamination are reduced.
[0012] Preferably, the genetic diversity analysis module can calculate at least three core genetic diversity indicators, including expected heterozygosity (He), observed heterozygosity (Ho), and nucleotide diversity (π), and can generate a genetic structure map of endangered species populations based on principal component analysis (PCA) and neighbor-joining tree (NJ tree) to intuitively present the degree of genetic differentiation within and between populations.
[0013] Advantages The application provides a RAD-seq low-cost genetic diversity evaluation system for endangered species. (1) The RAD-seq low-cost genetic diversity evaluation system for endangered species significantly reduces reagent and operation costs by using self-made reagent combinations and simplified library construction processes such as single-step enzyme cutting-linking reaction, shortens library construction time, is suitable for large-scale endangered species research, and improves DNA integrity and library construction success rate through the repair unit of the DNA extraction and pretreatment module for end repair and gap filling of degraded DNA, especially for non-invasive samples.
[0014] (2) The RAD-seq low-cost genetic diversity evaluation system for endangered species supports multiplex sample mixed sequencing through universal adapter design, improves sample processing efficiency, ensures the accuracy of SNP calling and diversity calculation through the genetic diversity analysis module through an automated process and a machine learning algorithm, reduces human intervention, realizes one-key operation and result visualization through the user interaction interface, reduces the use threshold, enables protection personnel to quickly obtain genetic diversity evaluation reports and population structure maps, supports de novo assembly without a reference genome, expands the application range, calculates various genetic diversity indicators, and provides comprehensive population genetic structure evaluation by combining PCA and NJ tree analysis. BRIEF DESCRIPTION OF DRAWINGS
[0015] Figure 1 It is the principle block diagram of the evaluation system of the application. Figure 2 The principle block diagram of the DNA extraction and pretreatment module of the present application is shown in FIG. 1. Figure 3 The principle block diagram of the simplified RAD-seq library construction module of the present application is shown in FIG. 2. Figure 4 The principle block diagram of the high-throughput sequencing module of the present application is shown in FIG. 3. Figure 5 The principle block diagram of the genetic diversity analysis module of the present application is shown in FIG. 4. DETAILED DESCRIPTION
[0016] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative labor fall within the scope of protection of the present application.
[0017] Please refer to Figure 1 and Figure 5 The present application provides a technical solution: a RAD-seq low-cost genetic diversity evaluation system for endangered species, which comprises an evaluation system, the evaluation system comprising a DNA extraction and pretreatment module, a simplified RAD-seq library construction module, a high-throughput sequencing module and a genetic diversity analysis module. The DNA extraction and pretreatment module is used to extract low-quality DNA from non-invasive samples of endangered species and perform DNA repair and concentration standardization. The simplified RAD-seq library construction module uses a low-cost self-made reagent combination to replace a commercial reagent kit, reduces the purification and fragment screening steps through a single-step enzyme cutting-linking reaction and universal adapter design, and realizes the simplification of the library construction process. The high-throughput sequencing module performs sequencing on the DNA fragments after library construction. The genetic diversity analysis module performs SNP marker development and diversity index calculation through an automated process based on the sequencing data, and outputs a genetic diversity evaluation report.
[0018] Specifically, the DNA extraction and pretreatment module: extract DNA from non-invasive samples (such as hair, feces, and old skin) of endangered species, use a commercial extraction reagent kit (such as Qiagen DNeasy Blood & Tissue Kit) or self-made reagents for extraction, and perform quality assessment (such as Nanodrop or Qubit detection) on the extracted DNA. For low-quality DNA (degradation or low concentration), enter the repair unit. The repair unit uses an enzyme combination (such as T4 DNA polymerase, Klenow fragment and T4 polynucleotide kinase) for end repair and gap filling, and the reaction conditions are incubation at 25°C for 30 minutes, followed by concentration standardization (such as dilution to 10 ng / μL); Simplified RAD-seq library preparation module: Utilizes a low-cost, self-made reagent combination, including self-made restriction endonuclease buffers (such as Tris-HCl, MgCl2, DTT), ligases (such as T4 DNA ligase), and universal adapters (containing an 8bp sample-specific barcode). The enzyme digestion-ligation reaction is performed in a single tube: The repaired DNA is mixed with restriction endonucleases (such as EcoRI), ligases, and universal adapters, and reacted at 25°C for 30 minutes without intermediate purification steps. Subsequently, library amplification is performed: Hot-start PCR technology (such as KAPAHiFiHotStartReadyMix) is used, with optimized annealing temperature (62°C) and cycle number (12 cycles) to reduce non-specific amplification. After amplification, the library is purified by magnetic beads, eliminating the need for fragment screening. High-throughput sequencing module: Use the Illumina platform (such as NovaSeq6000) to sequence the DNA fragments after library preparation. The recommended sequencing depth is 10-20×, generating raw data in FASTQ format. Genetic diversity analysis module: The analysis process includes: First, data quality control is performed using FastQC; second, SNP calling is performed using a reference genome alignment unit (such as BWA software) or a Denovo assembly unit (such as Stacks software); then, low-quality SNPs are filtered using machine learning algorithms (such as Python-based random forest models) (based on sequencing depth, quality value, and heterozygosity); finally, genetic diversity indices (He, Ho, π) are calculated and PCA and NJ tree analysis are performed using R software to generate a population genetic structure map. Users can input sample information, set parameters (such as minimum allele frequency), and view the visualization results through an interactive interface (such as the Shiny application).
[0019] By using a self-made reagent kit instead of a commercial kit, reagent costs are significantly reduced. The single-step enzyme digestion-ligation reaction reduces operational steps and improves library construction efficiency. It is suitable for low-quality DNA samples (such as non-destructive samples). DNA repair improves the success rate of library construction. Automated SNP development and diversity calculation reduce human error and improve evaluation efficiency.
[0020] Automated SNP development and diversity calculation include: The system automatically assesses the quality of FASTQ raw data generated by the high-throughput sequencing module. Using tools such as FastQC, it generates reports including indicators such as Phred quality value, GC content, and adapter contamination. Based on the quality control report, it automatically calls trimming tools (such as Trimmomatic or FastP) to remove low-quality bases, remove adapters from the ends of sequences, and discard sequences that are too short. The system automatically runs the quality control tools and determines the filtering parameters based on preset quality thresholds (which can be set by the user through the interface or by using the system default values). There is no need for manual review of the report before executing the trimming command, ensuring the accuracy and reliability of downstream analysis and eliminating the impact of low-quality data from the source. The system automatically invokes alignment software (such as BWA or Bowtie2) to align high-quality sequences after quality control with user-uploaded or system-embedded reference genomes, generating BAM / SAM files. It then automatically runs the Denovo assembly process using software such as Stacks or ipyrad. This process constructs a "virtual reference genome" for analysis by aligning sequences from all individuals and clustering them into "tags" (loci) based on sequence similarity. Based on the aligned or assembled loci, SNP calling tools (such as GATK's UnifiedGenotyper or SAMtools / bcftools) are used for preliminary SNP identification, generating a VCF file containing all potential SNP loci. The system automatically selects the analysis path based on the availability of a reference genome and automatically executes a series of commands, including sequence alignment, sorting, deduplication, and initial SNP calling, achieving a "dual-mode driven" analysis. This greatly expands the system's application scope, making it independent of hard-to-obtain reference genomes of endangered species. Multiple features are extracted from each SNP locus, such as average sequencing depth, genotype quality, allele balance, and flanking sequence complexity. A model pre-trained on a high-quality dataset is used to score each SNP locus and determine its probability of being a "true positive". Based on the probability score output by the model, a threshold (e.g., probability > 0.95) is automatically set to filter out false positive loci that, although they pass the traditional rules, are likely caused by sequencing or alignment errors. The system automatically inputs the high-quality SNP dataset (final VCF file) into population genetics analysis software (such as PLINK, vcftools, or the R adegenet package) to calculate a series of core genetic diversity indicators.
[0021] In a preferred embodiment, the enzyme digestion-ligation reaction time of the simplified RAD-seq library preparation module can be further shortened to 20 minutes to improve efficiency. The genetic diversity analysis module can also calculate additional indicators such as allele frequency and inbreeding coefficient to enhance the depth of evaluation.
[0022] In a preferred embodiment, the DNA extraction and pretreatment module includes a DNA quality assessment unit and a repair unit. The repair unit uses an enzyme combination to repair the ends and fill gaps in degraded DNA to improve DNA integrity and library construction success rate. By repairing degraded DNA with enzymes, the quality of the DNA template is enhanced, the success rate of subsequent library construction is improved, and the accuracy of genetic diversity assessment is ensured.
[0023] In a preferred embodiment, the simplified RAD-seq library preparation module uses a low-cost, self-made reagent combination including self-made restriction endonuclease buffer, ligase, and universal adapter. The universal adapter contains a sample-specific barcode, allowing multiplexed sequencing of samples and reducing reagent costs. The self-made reagents further reduce costs, and the barcode design supports multiplexed sequencing of samples, reducing sequencing costs and improving sample processing efficiency, making it suitable for large-scale endangered species research.
[0024] In a preferred embodiment, the simplified RAD-seq library preparation module uses hot-start PCR technology for library amplification, and reduces non-specific amplification and improves library complexity by optimizing annealing temperature and cycle number.
[0025] Specifically, the simplified RAD-seq library preparation module uses hot-start PCR technology for library amplification and reduces non-specific amplification by optimizing the annealing temperature and cycle number (e.g., annealing temperature 60-65°C, cycle number 10-15). Hot-start PCR and optimized parameters reduce non-specific amplification, ensure library diversity, improve sequencing data quality, reduce amplification bias, and make genetic diversity assessment more accurate.
[0026] In a preferred embodiment, the genetic diversity analysis module integrates a reference genome alignment unit and a Denovo assembly unit, supports SNP calling without a reference genome, and uses machine learning algorithms to filter low-quality data to ensure evaluation accuracy.
[0027] Specifically, the genetic diversity analysis module integrates a reference genome alignment unit and a Denovo assembly unit, supports SNP calling without a reference genome, and uses machine learning algorithms (such as random forest or support vector machine) to filter low-quality data, expanding its application scope to endangered species lacking a reference genome. The machine learning algorithm automatically filters low-quality SNPs, improving the accuracy of the assessment.
[0028] In a preferred embodiment, the evaluation system further includes a user interface for inputting sample information, setting parameters, and visualizing results, enabling one-click genetic diversity assessment.
[0029] Specifically, the evaluation system also includes a user interface (such as a web interface or desktop application) for inputting sample information, setting parameters (such as sequencing depth and diversity indicators), and visualizing results (such as charts and graphs). One-click operation simplifies the process, lowers the barrier to entry, and allows non-professionals to operate easily, intuitively displaying genetic diversity results for quick interpretation and decision-making.
[0030] In a preferred embodiment, the simplified RAD-seq library preparation module performs the enzyme digestion-ligation reaction in a single tube, reducing the reaction time to less than 30 minutes, reducing operational steps and contamination risks. The single-tube reaction reduces operational steps and contamination risks, shortens library preparation time, improves experimental efficiency, simplifies the process, reduces human error, and ensures library preparation consistency.
[0031] In a preferred embodiment, the genetic diversity analysis module can calculate at least three core genetic diversity indicators, including expected heterozygosity (He), observed heterozygosity (Ho), and nucleotide diversity (π). It can also generate genetic structure maps of endangered species populations based on principal component analysis (PCA) and neighbor-joining trees (NJ trees), which visually present the degree of genetic differentiation within and between populations.
[0032] Specifically, the genetic diversity analysis module can calculate at least three core genetic diversity indicators (expected heterozygosity He, observed heterozygosity Ho, and nucleotide diversity π), and can generate genetic structure maps of endangered species populations based on principal component analysis (PCA) and neighbor-joining trees (NJ trees). The multi-indicator calculation provides comprehensive genetic diversity analysis, and PCA and NJ trees visualize the population genetic structure, helping to identify population differentiation and conservation priorities.
[0033] Furthermore, any content not described in detail in this specification is existing technology known to those skilled in the art.
[0034] During operation, users first upload sample information through the interactive interface, and the system automatically starts the DNA extraction and preprocessing process; after library construction, the high-throughput sequencing module generates data; the analysis module automatically processes the data and outputs an evaluation report in PDF format, including a diversity index table and a genetic structure map.
[0035] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.
[0036] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A low-cost RAD-seq genetic diversity assessment system for endangered species, comprising an assessment system, characterized in that: The evaluation system includes a DNA extraction and pretreatment module, a simplified RAD-seq library preparation module, a high-throughput sequencing module, and a genetic diversity analysis module. The DNA extraction and pretreatment module is used to extract low-quality DNA from non-invasive samples of endangered species and perform DNA repair and concentration standardization. The simplified RAD-seq library preparation module uses a low-cost, self-made reagent combination to replace commercial kits. Through a single-step enzyme digestion-ligation reaction and universal adapter design, it reduces purification and fragment screening steps, simplifying the library preparation process. The high-throughput sequencing module sequences the DNA fragments after library preparation. Based on the sequencing data, the genetic diversity analysis module performs SNP marker development and diversity index calculation through an automated process, and outputs a genetic diversity assessment report.
2. The RAD-seq low-cost genetic diversity assessment system for endangered species according to claim 1, characterized in that: The DNA extraction and pretreatment module includes a DNA quality assessment unit and a repair unit. The repair unit uses an enzyme combination to repair the ends and fill gaps in degraded DNA to improve DNA integrity and library construction success rate.
3. The RAD-seq low-cost genetic diversity assessment system for endangered species according to claim 1, characterized in that: The simplified RAD-seq library preparation module uses a low-cost, self-made reagent combination, including self-made restriction endonuclease buffer, ligase, and universal adapter. The universal adapter contains a sample-specific barcode, allowing multiple sample mixing sequencing and reducing reagent costs.
4. The RAD-seq low-cost genetic diversity assessment system for endangered species according to claim 1, characterized in that: The simplified RAD-seq library preparation module uses hot-start PCR technology for library amplification and reduces non-specific amplification and improves library complexity by optimizing annealing temperature and cycle number.
5. The RAD-seq low-cost genetic diversity assessment system for endangered species according to claim 1, characterized in that: The genetic diversity analysis module integrates a reference genome alignment unit and a Denovo assembly unit, supports SNP calling without a reference genome, and uses machine learning algorithms to filter low-quality data to ensure evaluation accuracy.
6. The RAD-seq low-cost genetic diversity assessment system for endangered species according to claim 1, characterized in that: The assessment system also includes a user interface for inputting sample information, setting parameters, and visualizing results, enabling one-click genetic diversity assessment.
7. The RAD-seq low-cost genetic diversity assessment system for endangered species according to claim 1, characterized in that: In the simplified RAD-seq library construction module, the enzyme digestion-ligation reaction is carried out in a single tube, reducing the reaction time to less than 30 minutes, thus reducing the number of operation steps and the risk of contamination.
8. The RAD-seq low-cost genetic diversity assessment system for endangered species according to claim 1, characterized in that: The genetic diversity analysis module can calculate at least three core genetic diversity indicators, including expected heterozygosity (He), observed heterozygosity (Ho), and nucleotide diversity (π). It can also generate genetic structure maps of endangered species populations based on principal component analysis (PCA) and neighbor-joining trees (NJ trees), which visually present the degree of genetic differentiation within and between populations.