SNP molecular marker for macrobrachium rosenbergii germplasm identification and application thereof

CN119220694BActive Publication Date: 2026-09-22HUZHOU UNIVERSITY +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411279123.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-12
Publication Date
2026-09-22
Estimated Expiration
2044-09-12

AI Technical Summary

Technical Problem

目前我国主要养殖的罗氏沼虾品种/系较多,且多数为派生品系,这些不同的罗氏沼虾品种/系均没有典型的外观特征,因此在外观上无法准确区分

Benefits of technology

[0014]1)本发明提供的40个SNP位点是从10个群体共203份罗氏沼虾的全基因组重测序数据中筛选出来的,群体覆盖面广,具有代表性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure QLYQS_1
    Figure QLYQS_1
  • Figure QLYQS_2
    Figure QLYQS_2
  • Figure QLYQS_3
    Figure QLYQS_3
Patent Text Reader

Abstract

The application discloses a kind of SNP molecular markers for Macrobrachium rosenbergii germplasm resource identification and application thereof, the SNP molecular markers combination selected from 40 SNP molecular markers, 40 pairs of primers are developed based on the above 40 SNP molecular markers;Random forest classification model is constructed based on the 40 SNP sites screened in the application, reaches more than 90% accurate identification rate in Florida, Myanmar, Thailand and "Shu Feng No.1" Macrobrachium rosenbergii these four varieties / lines, solves the problem that it is difficult to carry out variety / line identification in production practice, provides strong technical support for subsequent germplasm identification and management.The application develops 40 pairs of primers or detection kit, which can quickly and low-costly complete Macrobrachium rosenbergii germplasm identification;Solve the problem that it is difficult to carry out variety / line identification in production practice, provide strong technical support for subsequent germplasm identification and management.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of molecular marker technology for giant freshwater prawns, specifically to an SNP molecular marker for the identification of giant freshwater prawn germplasm resources and its application. Background Technology

[0002] The giant freshwater prawn (Macrobrachium rosenbergii), belonging to the class Crustacea, order Decapoda, family Palaemonidae, and genus Macrobrachium, is one of my country's important freshwater economic shrimp species. It is characterized by its rapid growth, strong adaptability, and high nutritional value, making it popular with consumers. In aquaculture, the giant freshwater prawn has a short growth cycle and fast growth rate, offering broad prospects for cultivation. Currently, it is widely farmed in provinces such as Guangdong, Zhejiang, and Jiangsu. As of 2022, my country's aquaculture production had reached nearly 180,000 tons, accounting for more than half of global aquaculture production. Currently, there are many major aquaculture species / strains of giant freshwater prawns in my country, most of which are derived strains. These different species / strains lack typical morphological characteristics, making accurate differentiation based on appearance impossible. Without accurate species information, the identification of giant freshwater prawn germplasm resources and the determination of variety rights cannot be carried out, and the creation of new giant freshwater prawn varieties will inevitably be affected. Therefore, establishing an efficient giant freshwater prawn germplasm resource identification platform is particularly important.

[0003] With the continuous development of genotyping technology, molecular marker-assisted breeding is increasingly being applied to aquatic animal genetic breeding. Among them, SNP molecular markers, due to their high throughput, high stability, high compatibility, and low cost, have gradually replaced other molecular markers and become the main molecular markers currently used. For example, Zhang et al. established an identification platform based on 26 SNP loci that can distinguish mantis shrimp populations in the South China Sea and the Bohai Sea. Wang Yanyun obtained a marker combination composed of 27 SNP loci through screening, which can meet the identification needs of more than 80% of the main production areas of Yellow River carp in my country. Wang Qi et al. obtained a marker combination composed of 144 SNP loci through screening, which can be used to distinguish 16 carp varieties, including Yellow River carp, Super carp, and Fu'an carp, providing an effective technical means for carp variety rights identification and germplasm resource protection. Therefore, applying molecular markers to the genetic background survey of Macrobrachium rosenbergii is expected to identify different strains at the molecular level, which is helpful for the identification of variety authenticity and purity.

[0004] In conclusion, there is an urgent need to establish a reliable and efficient DNA fingerprinting system for giant freshwater prawns based on SNP molecular markers, so that it can play a more important role in variety selection and germplasm protection, anti-counterfeiting, and variety identification, thereby promoting the rapid development of my country's high-quality giant freshwater prawn industry. Summary of the Invention

[0005] The purpose of this invention is to overcome the shortcomings of the prior art and provide an application of whole-genome SNP molecular marker combinations in the identification of giant freshwater prawn varieties / lines. This invention constructs a random forest classifier based on 40 SNP molecular marker combinations through genotyping and core site screening, which can be used for the identification of giant freshwater prawn varieties / lines from Florida, New Myanmar, Thailand, and "Shufeng No. 1".

[0006] To achieve the above objectives, the technical solution designed by the present invention is as follows:

[0007] This invention provides an SNP molecular marker combination for the identification of giant freshwater prawn germplasm resources. The SNP molecular marker combination is selected from 40 SNP molecular markers. The physical location information and nucleoside sequence of the SNP molecular markers in the SNP molecular marker combination are shown in Table 1 below.

[0008] This invention also provides an application of the above-mentioned SNP site combinations in the identification of giant freshwater prawn varieties / lines.

[0009] The present invention also provides a primer combination for amplifying the above-mentioned SNP molecular marker using PCR technology, wherein the primer combination includes at least one of 40 primer pairs, and the nucleotide sequences of the 40 primer pairs are shown in Table 5 below.

[0010] The present invention also provides an application of the above-mentioned primer combination in the preparation of a detection kit.

[0011] The present invention also provides a detection kit, characterized in that: the detection kit includes the above-described primer combination.

[0012] The present invention also provides the application of the above-described primer combination or the above-described detection kit in the identification of giant freshwater prawn species / strains.

[0013] The beneficial effects of this invention are:

[0014] 1) The 40 SNP loci provided by this invention were screened from 203 whole genome resequencing data of giant freshwater prawns from 10 populations. The population coverage is broad and representative.

[0015] 2) Compared with germplasm identification methods based on second-generation sequencing data, this invention utilizes fixed and limited SNP sites to quickly and cost-effectively complete the germplasm identification of giant freshwater prawns.

[0016] 3) The random forest classification model constructed based on the 40 SNP loci screened in this invention achieved an accuracy identification rate of over 90% in four varieties / lines of giant freshwater prawns: Florida, New Myanmar, Thailand, and “Shufeng No. 1”. This solved the problem of difficulty in variety identification in production practice and provided strong technical support for subsequent germplasm identification and management.

[0017] 4) This invention has developed 40 primer pairs or detection kits that can quickly and cost-effectively complete the germplasm identification of giant freshwater prawns; it solves the problem of difficulty in carrying out variety identification in production practice and provides strong technical support for subsequent germplasm identification and management. Attached Figure Description

[0018] Figure 1 The simulated identification rate of different SNP site combinations in Example 1

[0019] Figure 2 The cross-validation curves based on different SNP site combinations in Example 1.

[0020] Figure 3 This is a graph showing the relationship between model error and the number of features and decision trees in Example 1.

[0021] Figure 4 Figure 1 shows the population structure analysis results based on 40 SNP loci in Example 2, and the SNP locus combination screening used for identification of Giant freshwater prawn germplasm resources in Example 1.

[0022] I. Genome resequencing and SNP site screening

[0023] 1. Sample collection and DNA extraction

[0024] Barbelic tissue samples were collected from 203 giant freshwater prawns from 10 cultured populations: Florida (FD, n=20), Guangxi (GX, n=20), Hefu (HF, n=19), Myanmar (MD, n=20), New Myanmar (NMD, n=27), “Nan Taihu No. 2” (NT, n=20), Thailand (TG, n=19), Taiwan (TW, n=20), “Shufeng No. 1” (SF, n=18), and Chia Tai (ZD, n=20). The barbel tissues were preserved in 95% ethanol. Genomic DNA was extracted using a DNA extraction kit (Aidlab, Beijing). The quality and integrity of the DNA samples were determined using a NanoDrop 2000 spectrophotometer (Thermo Scientific, USA) and 1% agarose gel electrophoresis.

[0025] 2. Library construction and genome resequencing

[0026] Genomic DNA samples were biodigested using restriction endonucleases, and sequencing libraries were constructed using a paired-end DNA sample preparation kit (Illumina Inc., San Diego, CA, USA). The constructed libraries were sequenced using the Illumina NovaSeq 6000 platform (Illumina, San Diego, CA) to obtain 150 bp (PE150) paired-end reads.

[0027] 3. Data filtering and comparison

[0028] The raw data was filtered using the software fastp (v 0.20.0) to remove adapter sequences, low-quality sequences, and excessively short sequences, resulting in a total of 6,430,808,165 clean reads. Subsequently, the quality-controlled data were aligned to the *Macrobrachium rosenbergii* reference genome using the software BWA-MEM (v 0.7.1563), and the deduplicated BAM file was output and indexed using the software smtools (v 1.10).

[0029] 4. SNP site search and filtering

[0030] Mutation detection and initial site screening were performed using FreeBayes (v 1.2.0) and vcflib software, respectively, resulting in a total of 704,709 SNP sites. Further quality control of the obtained SNP sites was performed using VCFtools (v 0.1.13) software with the following parameters: (1) max-missing 0.8; (2) maf 0.05; (3) min-alleles 2; (4) max-alleles2, resulting in 121,971 SNP sites selected for subsequent analysis. SNP sites with a deletion rate greater than 5%, a minimum allele frequency less than 0.2, and a polymorphism information content less than 0.3 were removed using plink (v1.90) software, ultimately yielding 6,178 SNP sites for subsequent core site screening.

[0031] II. SNP locus combination screening

[0032] 1. Selection of core SNP combinations and construction of random forest classifier

[0033] Since there is almost no genetic differentiation among artificially bred giant freshwater prawn (Macrobrachium rosenbergii) populations in my country, the "Shufeng No. 1" cultivar was selected as a representative of the artificially bred population. Together with giant freshwater prawn populations introduced in recent years from Florida, New Myanmar, and Thailand, this cultivar was used as the target population for SNP locus combination screening and the construction of a random forest classifier. The specific steps are as follows:

[0034] (1) Genotype coding

[0035] Genetic coding was performed on 6,178 SNP loci obtained after quality control for 84 individuals from 4 populations. Among them, the homozygous genotypes that are the same as the reference giant freshwater prawn genome are denoted as 0 / 0, 0 / 1 as 1, 1 / 0 as 2, and 1 / 1 as 3.

[0036] (2) Simulated classification analysis of different site combinations based on random forest model

[0037] Using R scripts, a concept of conditional random selection (CRS) algorithm was used to stratify and select site combinations from SNP sites. Initial screening showed that 20 core sites were sufficient for identifying four target populations. Subsequently, a decision tree was constructed using the random forest algorithm with these 20 core SNP sites, and identification was performed using a training set with random classification (0.75% of the total data). The results showed an identification rate of 69.84% (Table 1), where the simulated classification of the Thai population in the training set was misclassified as the Florida population. Therefore, the CRS algorithm was used to increase the number of core SNP sites to improve the identification rate. When the number of sites was less than 100, 5 sites were added each time; when the number exceeded 100, 20 sites were added each time. The trend of the identification rate with the increase of core SNP sites was observed. The results are as follows: Figure 1 As shown in Table 2, the simulated identification rate reached its highest value of 79.37% when the number of loci increased to 40. However, as the number of loci continued to increase, the identification rate did not increase; instead, it showed a downward trend.

[0038] Table 1. Simulated classification of four population training sets by combinations of 20 core SNP loci.

[0039]

[0040] Table 2. Simulated classification of four population training sets by combinations of 40 core SNP loci.

[0041]

[0042] To determine the final number and information of SNP loci to be used, we used two important metrics from the random forest model, "Mean Decrease Accuracy" and "Mean Decrease Gini," to sort the 100 loci selected by the CRS algorithm from highest to lowest. Then, we performed 10-fold cross-validation five times for different numbers of loci. The cross-validation curve results show that the model error initially decreases with the increase of SNP loci, with a large decrease at first. However, after reaching a certain range, the curve first rises and then falls, with the rate of decrease gradually decreasing. Figure 2 A). According to the cross-validation curve results, the model error remained within 0.1 when the number of SNP sites was in the range of 25-100. Adhering to the principle of simplicity, we selected the top 40 variables for cross-validation again. The results showed that the model error was minimized when the number of SNPs was 40. Figure 2 B).

[0043] 2. Optimization of the Random Forest Model

[0044] After determining to select 40 core SNP loci, we further reduced the model's misclassification rate by adjusting two important parameters in the `randomForest()` function: `ntree` and `mtry`. `ntree` is the number of base classifiers included, defaulting to 500; `mtry` is the number of variables included in each decision tree, defaulting to logN (where N is the number of objects). The results show that when `mtry` is set to 30, `rate` reaches its minimum value, indicating that `mtry` equal to 30 is the optimal solution. Figure 3 A). We used the optimal solution of mtry to simulate the optimal number of base classes for the decision tree. The results show that when the number of trees exceeds 270, the model's misclassification rate remains at its lowest value (error < 0.2). Figure 3 B). Considering the model's generalization ability and the risk of overfitting, we ultimately chose 500 trees. After optimization, the model's out-of-bag error (OOB error) decreased from 22.22% to 19.05%. Finally, the top 40 SNPs ranked by "Mean Decrease Accuracy" were determined as the core SNP combinations to distinguish between the Florida, New Myanmar, Thailand, and "Shufeng 1" giant freshwater prawn populations. Based on this, an optimized random forest classifier was constructed. Detailed information on the 40 SNPs is shown in Table 3.

[0045] Table 3 Location information of 40 SNP sites

[0046]

[0047]

[0048]

[0049] III. Evaluation of Actual Appraisal Results

[0050] To evaluate the identification capability of the random forest classifier built based on 40 core SNP loci, we performed germplasm identification on 84 individuals from Florida, New Myanmar, Thailand, and the "Shufeng 1" giant freshwater prawn population. The genotypic data of these 84 individuals at the 40 SNP loci were directly imported into the constructed random forest classification model to identify the species / line of the test samples. The identification rate was calculated based on the classification prediction results and the true values ​​of each sample. The results showed that 83 samples were correctly identified as species / lines, accounting for approximately 98.81% (Table 4). The classification errors originated from samples from the TG population being misclassified as belonging to the FD population.

[0051] Table 4. Identification analysis of 40 core SNP loci in 84 individuals from 4 populations.

[0052]

[0053] Example 2: Genotyping and Germplasm Identification Based on PCR Technology

[0054] Based on the 40 core SNP loci in Example 1, genotyping and germplasm identification of the giant freshwater prawn were performed using PCR technology. The specific methods are as follows:

[0055] 1. PCR amplification primers were designed based on the flanking sequences of 40 SNP sites. Detailed primer information is shown in Table 5.

[0056] Table 5. Amplification primer information for 40 SNP sites

[0057]

[0058]

[0059]

[0060] 2. Muscle samples were collected from *Macrobrachium rosenbergii*, specifically from 7 individuals of the New Myanmar population, 4 individuals of the "Shufeng 1" population, 5 individuals of the Thai population, and 5 individuals of the Florida population. Genomic DNA was extracted, and PCR amplification was performed. The amplified products were then directly sequenced to obtain the genotype data for each individual. The genotyping data for the four populations are shown in the table below:

[0061] Table 6. PCR detection results of 40 SNP loci in 4 populations

[0062]

[0063]

[0064] 3. The genotyping data was imported into the optimized random forest classifier from Example 1. The results showed that out of 21 samples from 4 populations, 20 samples could be correctly identified as varieties / lines, accounting for approximately 95.24% (Table 4). Only one sample had a classification prediction error, mainly due to the TG population being easily misclassified as the FD population.

[0065] 4. Phylogenetic tree construction and principal component analysis (PCA) were performed using the genotyping data of 21 individuals. The phylogenetic tree showed that the New Myanmar population and the "Shufeng 1" population each clustered into a separate clade, while the Thai population and the Florida population clustered into another clade. Figure 4 A), PCA analysis results show similar results ( Figure 4 B) Overall, the population structure analysis results are consistent with the classification results of the random forest model.

[0066] All other parts not described in detail are existing technologies. Although the above embodiments have provided a detailed description of the present invention, they are only some embodiments of the present invention, not all embodiments. People can obtain other embodiments based on these embodiments without creative effort, and these embodiments all fall within the protection scope of the present invention.

Claims

1. A combination of SNP molecular markers for identifying Giant freshwater prawn germplasm resources, characterized in that: The SNP molecular marker assemblage consists of 40 SNP molecular markers, and the nucleotide sequences of the SNP molecular markers in the assemblage are as follows:

2. The application of the SNP molecular marker combination as described in claim 1 in the identification of Giant freshwater prawn varieties / lines, characterized in that: The species / strains of giant freshwater prawns are the New Myanmar population, the "Shufeng No. 1" population, the Thai population, and the Florida population.

3. A primer combination for amplifying the SNP molecular marker combination of claim 1 using PCR technology, characterized in that: The primer combination comprises 40 primer pairs, wherein the nucleotide sequences of the 40 primer pairs are as follows:

4. The application of the primer combination according to claim 3 in the preparation of a reagent kit for identifying species / strains of Macrobrachium rosenbergii, characterized in that: The species / strains of giant freshwater prawns are the New Myanmar population, the "Shufeng No. 1" population, the Thai population, and the Florida population.

5. A test kit, characterized in that: The detection kit includes the primer combination as described in claim 3.

6. The application of the primer combination of claim 3 or the detection kit of claim 5 in the identification of Giant freshwater prawn species / strains, characterized in that: The application is for identifying giant freshwater prawn species / lines, which are the New Myanmar population, the "Shufeng No. 1" population, the Thai population, and the Florida population.

Citation Information

Patent Citations

  • Molecular marker for macrobrachium rosenbergii genetic breeding

    CN110016510A

  • SNP (Single Nucleotide Polymorphism) molecular marker related to low-temperature resistance character of juvenile macrobrachium rosenbergii and application

    CN118516462A