SNP molecular marker combination and application thereof in identification of orf fish varieties / strains

By constructing SNP molecular marker combinations and machine learning algorithms, the problem of tilapia variety/strain identification has been solved, achieving efficient and accurate germplasm resource identification and meeting the needs of germplasm resource protection and utilization.

CN120945062BActive Publication Date: 2026-06-16PEARL RIVER FISHERY RES INST CHINESE ACAD OF FISHERY SCI
View PDF -1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
PEARL RIVER FISHERY RES INST CHINESE ACAD OF FISHERY SCI
Filing Date
2025-08-04
Publication Date
2026-06-16

AI Technical Summary

Technical Problem

Existing technologies are insufficient for efficiently identifying multiple varieties/strains of tilapia. Traditional morphological methods are inadequate for differentiation, molecular marker technology has limited application scenarios, and machine learning has not yet formed an efficient molecular identification system for aquatic animal germplasm resource identification.

Method used

We constructed SNP molecular marker combinations, including SNP1-SNP40, and combined them with machine learning algorithms such as LightGBM and multinomial logistic regression to screen for germplasm-specific SNPs. We developed KASP primers and kits, optimized the genotyping method, and constructed a tilapia variety/strain classification model.

Benefits of technology

It has achieved high-precision identification of tilapia varieties/strains, with a classification model accuracy of 99.61% and a typing success rate of 95.08%~98.17%, improving the accuracy and operational efficiency of germplasm resource identification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120945062B_ABST
    Figure CN120945062B_ABST
Patent Text Reader

Abstract

The present application relates to the field of biotechnology, and particularly relates to a SNP molecular marker combination and its application in identifying tilapia varieties / strains. The present application constructs a whole genome SNP variation map covering 7 tilapia varieties / strains, and 40 core SNP molecular markers for germplasm identification are screened out through a machine learning algorithm, and a classification model with a high accuracy of 99.61% is constructed. On this basis, 24 KASP molecular marker sites are further optimized and a matching typing method is developed, and an average typing success rate of 95.08% to 98.17% is achieved in the target population, and the variety / strain identification accuracy reaches 98.6%. The present application applies machine learning algorithm to tilapia germplasm identification, supports dynamic expansion of SNP panels to include new varieties / strains, and provides an efficient and standardized tool for tilapia conservation breeding, fry purity control and digital management.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of biotechnology, and in particular to SNP molecular marker combinations and their application in the identification of tilapia varieties / strains. Background Technology

[0002] Because different tilapia germplasm exhibit unique advantages in terms of stress resistance, body color, or flesh quality, and because reproductive isolation between varieties is relatively weak, numerous superior varieties / strains have been rapidly obtained in recent years through interspecific hybridization. While these germplasm resources and physiological characteristics have maintained the vitality of tilapia germplasm innovation, they also show extensive infiltration of ancestral genes, increasing the difficulty of germplasm preservation and reuse.

[0003] Traditional morphological identification methods are insufficient to meet the needs of breed preservation and breeding due to the high degree of overlap in phenotypic characteristics among different varieties. Although molecular marker technologies such as random amplified polymorphic DNA, restriction fragment length polymorphism, mitochondrial DNA, and D-loop sequence analysis have been gradually applied to germplasm identification, these methods generally suffer from limited identification range, complex operation procedures, and reliance on professional personnel. While third-generation single nucleotide polymorphism (SNP) markers have shown the potential to construct high-density fingerprint maps, in tilapia, existing SNP sets can only identify each other among a maximum of four specific varieties, limiting their application scenarios. When more varieties are involved, system compatibility decreases significantly. Machine learning algorithms have made breakthrough progress in livestock breed identification. Models such as support vector machines and random forests have achieved identification accuracy of over 99% in species such as pigs, cattle, and chickens by optimizing SNP marker combinations and genetic differentiation indicators. However, the application of such technologies in aquatic animal germplasm resource identification is still in the technical verification stage. Existing research is mostly focused on the identification of a few specific varieties and has not yet formed an efficient molecular identification system that can be extended to multiple varieties. Summary of the Invention

[0004] In view of this, the present invention proposes SNP molecular marker combinations and their application in the identification of tilapia varieties / strains.

[0005] The technical solution of this invention is implemented as follows:

[0006] In a first aspect, the present invention provides an application of SNP molecular marker combinations in the identification of tilapia varieties / strains, wherein the SNP molecular marker combinations include SNP1-SNP40, and the locus information of SNP1-SNP40 is shown in Table 1 below:

[0007]

[0008]

[0009]

[0010] The NCBI accession number for the Nile tilapia reference genome sequence is GCF_001858045.2.

[0011] Secondly, this invention provides a method for identifying tilapia varieties / strains, comprising the following steps:

[0012] S1. Extract genomic DNA from the tilapia to be tested and perform whole-genome resequencing;

[0013] S2. Genotyping of the 40 SNP loci in the SNP molecular marker combination described in claim 1 to obtain the genotypes of the 40 SNP loci;

[0014] S3. Import the genotype results into the tilapia variety / strain classification model, and identify the variety / strain of the tilapia to be tested based on the classification results.

[0015] Furthermore, in step S3, the method for constructing the tilapia variety / strain classification model includes the following steps:

[0016] S3-1. Convert the tilapia species / strain classification label to be identified into binary code format;

[0017] S3-2. Identify SNP genotypes based on whole-genome sequencing data, and quantify the feature importance score of each SNP locus using LightGBM; use the SelectFromModel program of scikit-learn to screen and obtain a germplasm-specific SNP set (a set of key SNPs related to germplasm specificity) based on the feature importance score.

[0018] S3-3. Train a classification model using machine learning algorithms, identify the optimal set of germplasm-specific SNPs, and construct a mapping relationship between the set of germplasm-specific SNPs and the classification categories of tilapia varieties / strains. The machine learning algorithm is at least one of the following: multinomial logistic regression (MLR), support vector machine (SVM), k-nearest neighbors (KNN), Naïve Bayes (NB), decision tree (DT), random forest (RF), backpropagation neural network (BPNN), and gradient boosting decision tree (GBDT).

[0019] Thirdly, the present invention provides KASP primers for identifying tilapia varieties / strains, the sequences of which are shown in SEQ ID NO: 1-72.

[0020] Fourthly, this invention provides the application of the KASP primers in identifying tilapia varieties / strains.

[0021] Fifthly, the present invention provides reagents or kits containing the KASP primers.

[0022] Sixthly, the present invention provides a method for identifying tilapia varieties / strains, comprising: extracting genomic DNA of the tilapia to be tested; performing genotyping using the KASP primers described in claim 5; comparing the genotypes at the 24 SNP loci obtained with the genetic variation map of known varieties / strains, thereby determining the variety / strain of the tilapia to be tested.

[0023] The beneficial effects of the present invention include at least the following:

[0024] This invention constructed a genome-wide SNP variation map covering seven tilapia varieties / strains. Forty core SNP markers for germplasm identification were screened using machine learning algorithms, and the constructed classification model achieved an accuracy of 99.61%. Based on this, 24 KASP molecular marker loci were further optimized, and a corresponding genotyping method was developed, achieving an average genotyping success rate of 95.08%–98.17% in the target population, with a variety / strain identification accuracy of 98.6%.

[0025] This invention significantly improves the accuracy and operational efficiency of tilapia germplasm resource identification, providing standardized technical tools for seedling purity control, superior breed selection, and digital management of germplasm resources. Through iterative optimization of machine learning models, this invention allows for the dynamic expansion of SNP marker combinations based on resequencing data of new varieties, providing a technical interface for the rapid inclusion of new tilapia germplasm in the future and meeting the long-term needs of germplasm resource protection and utilization. Attached Figure Description

[0026] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0027] Figure 1 The tilapia whole-genome SNP variation map provided by this invention includes: A, which is the whole-genome SNP variation Circos map of 5 tilapia varieties / strains resequencing by this invention; and B, which is the whole-genome SNP variation Circos map of 2 tilapia varieties from the NCBI database, including 30 Nile tilapia (BioProject accession number: PRJCA040532) and 32 Mozambique tilapia (BioProject accession number: PRJCA004934).

[0028] Figure 2 The population structure analysis results among tilapia varieties / strains provided by this invention include: A) phylogenetic analysis results of 7 varieties / strains; B) principal component analysis (PCA) results of 7 varieties / strains: principal component 1 (PC1) and principal component 2 (PC2) can cluster 6 of the varieties / strains respectively; C) cross-validation error (K value) corresponding to different maximum likelihood values ​​in the ADMIXTURE analysis; D) ADMIXTURE population structure analysis: when K = 10, the estimated number of ancestors of the 7 varieties / strains; (Note: ten colors represent 10 different putative ancestors, and each vertical bar represents a tested individual)

[0029] Figure 3-6 The results of the analysis of tilapia varieties / strains based on the machine learning model provided by this invention include:

[0030] Figure 3 The accuracy of different machine learning models varies with the number of germplasm-specific SNPs; the accuracy increases rapidly when the number of SNPs increases from 20 to 40.

[0031] Figure 4 The estimated values ​​of eight machine learning models across four model accuracy evaluation metrics are given when the number of germplasm-specific SNPs is 40; the MLR model has the highest accuracy when there are 40 SNPs.

[0032] Figure 5 For the MLR model, the correlation heatmap of the confusion matrix when the number of germplasm-specific SNPs is 40;

[0033] Figure 6 This is a bar chart of regression coefficients for the MLR model; the horizontal axis represents the contribution of SNPs to variety / strain identification, and the vertical axis lists the top 20 SNPs with the highest contribution. Detailed Implementation

[0034] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below. Obviously, the described embodiments are only some embodiments of this invention, not all embodiments. Based on the embodiments of this invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this invention. Where specific conditions are not specified in the embodiments, conventional conditions or conditions recommended by the manufacturer shall be followed. Reagents or instruments whose manufacturers are not specified are all conventional products that can be purchased commercially.

[0035] The seven types of tilapia germplasm resources in this application embodiment include the following:

[0036]

[0037] Note: Both the Rainbow Tilapia and the Leopard Red Tilapia strain originated from the targeted breeding of red tilapia in Taiwan, China. Early breeding focused on improving body color, thus cultivating Rainbow Tilapia with an iridescent sheen on its body surface. In recent years, through continuous breeding, the Rainbow Tilapia strain has been further bred to produce the Leopard Red Tilapia strain, which has more vivid red markings and higher color saturation.

[0038] This invention first constructs whole-genome variation maps for Nile, ALY, SL, Moz, YSL, CHD, and BWH, analyzing the population structure of these seven tilapia varieties / strains. Then, machine learning algorithms are used to identify germplasm-specific SNPs (SNP sets) after rigorous filtering, and a classification model for these seven tilapia varieties / strains is constructed based on these germplasm-specific SNPs. Next, this invention develops a KASP genotyping method based on germplasm-specific SNPs and applies it to the genotyping of individual samples from these seven tilapia varieties / strains collected independently from other sources. Finally, these KASP genotyping data are used as a test set and input into the constructed classification model. A double-blind comparison is performed with the generated genetic variation map datasets of the seven varieties / strains to verify the reliability of the classification model based on germplasm-specific SNPs.

[0039] Example 1

[0040] 1. DNA extraction, library construction, and resequencing

[0041] In this embodiment, Nile, ALY, and SL were provided by the Gaoyao Base of the Pearl River Fisheries Research Institute, Chinese Academy of Fishery Sciences. CHD, YSL, and BWH were provided by Bailian Aquatic Seedling Co., Ltd. Each variety / strain was self-bred annually to ensure the stability of the pureline genetic structure.

[0042] First, 42 ALY, 48 CHD, 32 YSL, 50 BWH, and 48 SL (a total of 220 fish) were selected, and their tail fins (approximately 1.5 mm) were collected. 2 All sampling was conducted at the place of origin. Before sampling, the fish were anesthetized for 30 seconds with 90-100 mg / L ethyl 3-aminobenzoate methanesulfonate (Sigma, A5040, MS-222). The fins were then quickly cut off, completely immersed in anhydrous ethanol, and stored at -20°C for subsequent experiments.

[0043] Resequencing genomic data for 30 Nile (Bioproject accession number: PRJDB1657) and 32 Moz (Bioproject accession numbers: PRJDB1657, PRJCA004934) samples were downloaded from the NCBI database. Additionally, resequencing genomic data for 494 Nile samples from our laboratory were included in the analysis (Bioproject accession number: PRJCA040532).

[0044] Genomic DNA was extracted from the tail fin using the TIANamp Genomic DNA Kit (Tiangen, Beijing, China) according to the manufacturer's instructions. DNA quality and concentration were determined using a Nanodrop 2000 nucleic acid concentration meter (spectrophotometry) and 1.5% agarose gel electrophoresis. The purified DNA was then used to prepare sequencing libraries using the NEBNext Ultra II DNA Library Preparation Kit (New England Biolabs, Beijing, China). These libraries were then sequenced using the BGISEQ500 platform (150 bp paired-end reads) from Beijing Biomarker Technology Co., Ltd. (China). Raw reads were saved in FASTQ format.

[0045] Whole-genome resequencing was performed on 220 tilapia, generating a net read length of 3338.18 GB. After quality control and SNP identification, a total of 3,658,935 SNPs were identified, with a GC content of 40.33% and Q30 > 98.00%. Figure 1 (A). In addition, a total of 2,851,322 SNPs were obtained from publicly available Nile and Moz resequencing data. Figure 1 (Middle B). Meanwhile, a total of 4,058,381 SNPs were obtained from the Nile resequencing data in our laboratory. After data merging and filtering, a final dataset of 18,060,993 SNPs from 776 individual genomes of 7 varieties / lines was obtained.

[0046] 2. SNP identification, genotyping, and annotation

[0047] First, the raw data underwent quality control using the FASTP (v107) program with default parameters. In this step, adapters, reads containing more than 10 unknown bases (N), and low-quality reads with a phred value > 20 (Q20) and accounting for more than 50% of the total read length were removed. The filtered reads were defined as net reads and used for downstream analysis. Then, the net reads were aligned to the Nile tilapia reference genome sequence (NCBI accession number: GCF_001858045.2) using the -index and -mem options of the bwa-mem2 software (v 2.2.1). The alignment results were sorted, repetitive sequences were removed, and a SAM file was generated. Subsequently, the SNP was invoked using the HaplotypeCaller module in GATK (v3.8), and filtered using the following parameters: QD < 2.0 || MQ < 40.0 || FS > 60.0 || QUAL < 30.0 || MQrankSum < -12.5 || ReadPosRankSum < -8.0, -clusterSize 2, -clusterWindowSize 5. Finally, all VCF files were merged using GATKGenomicsDBImport, and the merged VCF data was genotyped using GATK GenotypeGVCFs to generate a single VCF file containing all 776 samples.

[0048] 3. Group Structure Analysis

[0049] Population structure analysis included phylogenetic analysis, principal component analysis (PCA), and ADMIXTURE analysis. To analyze phylogenetic relationships, a rootless phylogenetic tree was constructed using the neighbor-joining (NJ) method, and a Kimura two-parameter / p-distance model was used in MEGA-CC software (v 10.0.5) with 1000 guided replicates. PCA analysis was performed using the smartPCA program in the EIGENSOFT software package (v.7.2.1) with default parameters. The sample groups were visualized using the first two principal components (PCs). ADMIXTURE analysis was performed using ADMIXTURE software (v1.3.0). The cross-validation (CV) error for each K value was calculated, and the K value with the smallest CV error (K=10) was considered the optimal value for assessing the mixing degree of each sample.

[0050] The analysis showed that the rootless phylogenetic tree divided each variety / line into five main branches. Individuals from Nile, SL, and ALY clustered into three independent branches. Individuals from Moz clustered separately, but this branch, along with individuals from YSL, formed a mixed population branch (Mixed population 1, MP1). Similarly, CHD individuals and BHW individuals formed a second mixed population branch (Mixed population 2, MP2). Figure 2 (A). For ADMIXTURE analysis, the cross-validation (CV) error is minimized when the clustering is set to K=10. Figure 2 (C). When K=10, the population genetic differentiation is consistent with the PCA (principal component analysis) results. Figure 2 (B), and each variety / line was clearly isolated and supported by a phylogenetic tree. When all varieties / lines converged to a cluster with a mean likelihood of up to 10, Nile showed the highest mixing and diversity among the 7 varieties / lines. Figure 2 (D). CHD and BWH have a conserved and highly mixed ancestral lineage ( Figure 2 (D). SL, ALY, and Moz, however, show three single, distinct ancestral lineages. YSL exhibits low mixing, with one ancestral lineage sharing the same lineage as Moz ( Figure 2 (D).

[0051] 4. Identifying germplasm-specific SNPs and constructing classification models based on machine learning.

[0052] 4.1 SNP Screening

[0053] To select high-quality SNPs for variety / strain identification, this application rigorously screened the VCF files according to the following criteria: (1) markers were evenly distributed across the genome; (2) marker integrity was 100%; (3) markers with a minor allele frequency (MAF) of less than 20% were removed; (4) markers with a polymorphism information content (PIC) of less than 0.35 were removed; (5) markers with a p-value greater than 0.01 after the Harvin test were retained; and (6) no other SNPs were found within a 100 bp range before and after the selected marker. After the rigorous screening, a total of 203 high-quality SNPs were obtained. If the number of markers after screening exceeded 350, 350 were randomly selected; if there were fewer than 350, all were retained for downstream analysis. The final 203 retained SNPs were used for subsequent machine learning modeling.

[0054] 4.2 Identification of germplasm-specific SNPs

[0055] The methodology for identifying germplasm-specific SNPs using machine learning is referenced in Zhi et al., Advanced molecular system for accurate identification of chicken genetic resources[J]. Computers and Electronics in Agriculture, 2025, 231: 109989. The specific analysis was performed using a Python (v3.6) environment with pandas (v2.0.3), numpy (v1.24.3), and scikit-learn (v1.3.2). Pandas was used for data preprocessing, numpy for data computation, and scikit-learn for building the machine learning model.

[0056] 1) Data preprocessing: Since the varieties / strains of tilapia belong to unordered categorical data, before analyzing SNPs, the OneHotEncoder program of scikit-learn is first used to perform one-hot encoding on these nominal categorical variables, converting the nominal categorical variables into numerical forms that can be recognized by machine learning.

[0057] 2) Feature Selection: First, LightGBM (v4.5.0) is used to calculate the importance score of each SNP feature, quantifying the contribution of each SNP to the classification. Second, scikit-learn's SelectFromModel program is used to filter out key SNPs based on feature importance scores, reducing the interference of redundant data on model training.

[0058] 3) Model Construction and Optimization: Eight machine learning algorithms were used to construct a classification model. The optimal germplasm-specific SNPs and the mapping relationship between the optimal germplasm-specific SNPs and the classification model were determined through multi-model comparison. The machine learning algorithms included: MLR, SVM, KNN, NB, DT, RF, BPNN, and GBDT.

[0059] 5. Machine Learning Accuracy Assessment

[0060] The model accuracy evaluation employs a two-stage validation strategy. In the first part, the SNP dataset is randomly stratified at a ratio of 70% (training):30% (test) to initially evaluate the model's performance on the independent test set. In the second part, this application performs 10-fold cross-validation to improve evaluation robustness and minimize the risk of overfitting. The training data is divided into ten equal parts for iterative validation. In each iteration, nine parts are used for model development, and one part is used for validation. This process iterates through all parts to ensure that each data point contributes to both the training and validation phases, thus providing a comprehensive evaluation of the model's generalization ability. The results are presented using Matplotlib and Seaborn to generate feature importance bar charts to show the SNP contribution, and a confusion matrix heatmap to present classification accuracy. Finally, the model accuracy is quantitatively evaluated using the following four metrics: accuracy (the proportion of correctly predicted classes across all classes), precision (the reliability of correctly predicted classes), recall (the completeness of correctly identified classes), and F1 score (the average accuracy evaluated based on recall).

[0061] result The above eight machine learning models were used to attempt to distinguish varieties / strains using SNP datasets with 10-200 SNPs. The recognition accuracy of all models increased with the number of SNPs. When using the top 20 SNPs with the largest effects, the recognition accuracy of all models, except the DT model, rose to over 95%. The recognition accuracy of these models peaked when the number of large-effect SNPs increased to 40. Among them, the MLR model achieved a recognition accuracy as high as 99.61%. Figure 3 However, as the number of SNPs continued to increase, the recognition accuracy of all models did not change significantly. The results showed that the MLR model achieved the best results when using 40 SNPs. Figure 4 Table 1 shows the SNP set obtained by the MLR model for the identification of 7 tilapia varieties / strains.

[0062] Furthermore, in a double-blind test, a confusion model was used to verify the accuracy of the MLR model's identification. The confusion matrix showed that the classification model based on 40 SNP markers completely matched the actual varieties for all individuals. Only one individual had insufficient confidence in classification between Nile and BWH and was not clearly assigned; the rest were accurately identified. Figure 5To assess the taxonomic contribution of each SNP, this application plotted the regression coefficients of the MLR model. The SNP located at LG6: 27987220 had the highest regression coefficient (0.6749). The confidence scores of the top 9 SNPs were all above 0.5, while the confidence scores of the remaining SNPs were all below 0.3, confirming that the top 9 SNPs are the core markers for variety / strain identification. Figure 6 ).

[0063] Example 2

[0064] In this embodiment, the tilapia varieties / strains Nile, Moz, and ALY were provided by the Fangcun Base of the Pearl River Fisheries Research Institute, Chinese Academy of Fishery Sciences. SL, CHD, YSL, and BWH were provided by the Guangdong Tilapia Breeding Farm. Each variety / strain was individually cultured in a 4 m × 3 m × 2 m recirculating aquaculture system. They were fed twice daily (7:00 AM and 7:00 PM) with commercial feed at 3% of their body weight. The rearing conditions were identical for all varieties / strains: water temperature maintained at 28℃±2℃, dissolved oxygen maintained between 6.0 and 7.5 mg / L, pH stable between 7.5 and 8.5, and total ammonia nitrogen level maintained at <0.1 mg / L. Fin collection and storage methods were the same as in Example 1. Specifically, caudal fin samples were taken from 48 CHD fish, and fin samples were taken from 50 tilapia fish from each of the other six varieties / strains (a total of 348 samples). DNA was extracted for fingerprint verification.

[0065] First, a set of SNPs suitable for KASP identification was screened. The screening principle was that the specificity of the upstream and downstream 50 bp of the SNP in the genome was greater than 80%. Sequence specificity was analyzed by NCBI blast. A SNP genotyping set consisting of 24 SNPs (Table 2) was successfully developed from 40 SNPs for germplasm identification of the above 7 varieties / lines. Two allele-specific forward primers and one reverse universal primer were designed for the candidate SNPs using Primer 5.0. As shown in Table 2, the FX primer was modified with a FAM fluorescent group at the 5' end, and the 3' end base was complementary to the reference base. The FY primer was modified with a HEX fluorescent group at the 5' end, and the 3' end base was complementary to the mutant base. The primers were synthesized by Beijing Sangon Biotech Co., Ltd. All three SNP primers were diluted to 10 μmol and mixed at a volume ratio of 12:12:30.

[0066]

[0067]

[0068]

[0069]

[0070]

[0071]

[0072]

[0073] KASP genotyping was used to genotype 24 SNPs (as shown in Table 2) in the fin DNA samples of 348 tilapia. The genotyping experiments were performed on the high-throughput genotyping platform of LGC, Teddington, UK. The PCR reaction system is shown in Table 3. The amplification reaction was carried out in a high-throughput Hydrocycler water bath system. The PCR program was set as follows: first, pre-denaturation at 94 °C for 15 minutes; then 10 cycles of Touchdown program (denaturation at 94 °C for 20 seconds, annealing and extension between 61 °C and 55 °C, decreasing by 0.6 °C per cycle, for 1 minute); followed by 26 standard cycles (denaturation at 94 °C for 20 seconds, extension at 55 °C for 60 seconds). After amplification, the fluorescence signal was detected and genotyping was performed using a BMG PHERAstar multi-functional microplate reader (Oersberg, Germany).

[0074] Table 3 PCR amplification reaction system

[0075]

[0076] The genotyping set consisting of 24 SNP loci shown in Table 1 achieved an average genotyping success rate of 95.08%–98.17% for different varieties / strains (Table 4). The genotyping results were generated into VCF files using VCFtools (default parameters), and then double-blindly compared with the genetic variation maps (VCF files) of the seven varieties / strains. The comparison model was the best model evaluated by machine learning. Based on the comparison results, the variety / strain to which each individual belonged was determined and verified against its actual variety / strain to ultimately determine the reliability of the molecular identification system. Results from the MLR model showed that the MLR model achieved an accuracy of 98.2%, precision of 97.9%, recall of 97.6%, and an F1 score of 97.7% for the tested individuals using the 24 SNP loci, indicating that these 24 SNP loci have accurate germplasm identification capabilities (Table 5).

[0077] Table 4 KASP typing success rate

[0078]

[0079]

[0080] Table 5. Accuracy assessment of MLR models after genotyping of 24 germplasm-specific SNPs

[0081] Evaluation indicators Result (×100%) Accuracy 0.982 Precision 0.979 Recall 0.976 F1 score 0.977

[0082] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. The application of SNP molecular marker combination in the identification of Oreochromis niloticus varieties / strains, characterized in that, The SNP molecular marker combination includes SNP1-SNP40, and the site information of SNP1-SNP40 is as follows: , The NCBI accession number for the Nile tilapia reference genome sequence is GCF_001858045.2; the tilapia species / strains mentioned are Nile tilapia, Saro tilapia, Oreo tilapia, Mozambique tilapia, Israeli red tilapia, rainbow tilapia, and leopard red tilapia.

2. A method for identifying a variety / line of Oreochromis niloticus, characterized by, The tilapia species / strains are Nile tilapia, Salo tilapia, Oreo tilapia, Mozambique tilapia, Israeli red tilapia, rainbow tilapia, and leopard red tilapia. The method includes the following steps: S1. Extract genomic DNA from the tilapia to be tested and perform whole-genome resequencing; S2. Genotyping of the 40 SNP loci in the SNP molecular marker combination described in claim 1 to obtain the genotypes of the 40 SNP loci; S3. Import the genotype results into the tilapia variety / strain classification model, and identify the variety / strain of the tilapia to be tested based on the classification results.

3. The method of claim 2, wherein, In step S3, the method for constructing the tilapia variety / strain classification model includes the following steps: S3-1. Convert the tilapia species / strain classification label to be identified into binary code format; S3-2. Identify SNP genotypes based on whole-genome sequencing data and quantify the characteristic importance score of each SNP locus using LightGBM; The SelectFromModel program of scikit-learn was used to filter and obtain a set of germplasm-specific SNPs based on feature importance scores; S3-3. Train the classification model using machine learning algorithms, establish the optimal germplasm-specific SNP set, and construct the mapping relationship between the germplasm-specific SNP set and the tilapia variety / strain classification category. The machine learning algorithm is at least one of the following: multinomial logistic regression, support vector machine, k-nearest neighbor, naive Bayes, decision tree, random forest, backpropagation neural network, and gradient boosting decision tree.

4. A KASP primer combination for the identification of a Tilapia breed / line, characterized in that, The tilapia species / strains are Nile tilapia, Salo tilapia, Oreo tilapia, Mozambique tilapia, Israeli red tilapia, rainbow tilapia, and leopard red tilapia. The primer combination consists of primers with sequences as shown in SEQ ID NO: 1-72.

5. Use of the KASP primer combination according to claim 4 for the identification of a Tilapia variety / line, characterized in that, The tilapia species / strains mentioned are Nile tilapia, Salo tilapia, Oria tilapia, Mozambique tilapia, Israeli red tilapia, rainbow tilapia, and leopard red tilapia.

6. A reagent or kit containing the KASP primer combination of claim 4.

7. A method of identifying a variety / line of Oreochromis niloticus, characterized in that, include: Extract genomic DNA from the tilapia to be tested; Genotyping was performed using the KASP primer combination described in claim 4. The genotypes at the 24 SNP loci were compared with the genetic variation maps of known varieties / strains to determine the variety / strain of the tilapia to be tested. The tilapia varieties / strains are Nile tilapia, Salo tilapia, Oreo tilapia, Mozambique tilapia, Israeli red tilapia, rainbow tilapia, and leopard red tilapia.

8. The method of claim 7, characterized in that, A multinomial logistic regression model was used to compare the genotypes at the 24 SNP loci with the genetic variation maps of known varieties / strains.