Acquisition method of whole-genome SNP (Single Nucleotide Polymorphism) site combination of fruit fly guavas
Forty core SNP loci were selected through whole-genome sequencing and random forest algorithm, and a classification model was constructed to solve the problem of identifying the geographical origin of guava fruit flies, thus achieving efficient and accurate population tracing and control.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHINA AGRI UNIV SANYA RES INST
- Filing Date
- 2026-04-01
- Publication Date
- 2026-04-24
AI Technical Summary
Existing technologies make it difficult to accurately identify the geographical origin of guava fruit flies through genetic information, leading to difficulties in cross-border and cross-regional control. Furthermore, whole-genome SNP locus analysis is time-consuming, labor-intensive, and costly.
SNP loci were obtained by whole-genome sequencing technology. Forty core SNP loci were selected by using random forest combined with recursive feature elimination algorithm. A random forest classification model was constructed for tracing the origin of the guava fruit fly population.
It significantly reduces the number of detection sites, lowers costs, shortens the detection cycle, and achieves high-accuracy population tracing. The model has noise resistance and generalization ability, making it suitable for rapid and automated identification of massive samples.
Smart Images

Figure CN121922210A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of bioinformatics and pest control, and in particular to a method for obtaining the SNP locus combination of the whole genome of the guava fruit fly. Background Technology
[0002] Guava fruit flies are a serious pest of fruits and vegetables, and accurately identifying their geographical origin (such as determining whether they originated in Hainan) is crucial for developing precise quarantine strategies. However, accurate identification based solely on geographical location or morphological characteristics is no longer sufficient, and the lack of genetic information has become an obstacle to cross-border and cross-regional control. With the development of genome sequencing technology, numerous SNP loci have been obtained, and analyzing all loci is not only time-consuming and labor-intensive but also expensive.
[0003] In view of this, we propose a method for obtaining the combination of SNP sites in the whole genome of the guava fruit fly to solve the existing problems. Summary of the Invention
[0004] The purpose of this invention is to provide a method for obtaining the SNP locus combination of the whole genome of the guava fruit fly, so as to solve the problems mentioned in the background art.
[0005] To achieve the above objectives, the present invention provides the following technical solution: a method for obtaining SNP locus combinations from the whole genome of *F. guava*, comprising: obtaining the DNA of individual *F. guava*, performing genotyping using whole-genome sequencing technology, performing quality control on the obtained SNPs, and performing feature selection using a random forest combined with a recursive feature elimination algorithm; sorting the SNPs in descending order of importance value, and retaining the core SNP loci where the classification accuracy reaches its peak, i.e., the *F. guava* population-specific combination loci.
[0006] Furthermore, the SNP locus combination of the whole genome of the fruit fly guava includes 40 core SNP loci, and the variation information of the 40 SNP loci is represented by chromosome numbers connected by underscores to indicate their physical locations.
[0007] Furthermore, in the random forest classification model used for molecular identification of guava fruit fly populations: for the genotype of each guava fruit fly individual's SNP locus, 0 is used to replace the wild genotype, 1 to replace the heterozygous genotype, and 2 to replace the mutant genotype, to obtain a digital dataset; the dataset is divided into a training set and a validation set, and the random forest model is trained using the training set; the prediction accuracy of the trained model is evaluated using the validation set, to obtain a classification model for identifying whether guava fruit flies originate from a specific region.
[0008] Furthermore, a 10-fold cross-validation screening was performed using a random forest combined with a recursive feature elimination algorithm.
[0009] Furthermore, during the feature selection process, by setting the subset step size, the model accuracy under different numbers of SNPs is calculated, and the 40 core variables at which the accuracy reaches its peak are retained.
[0010] Furthermore, the digitized genotypes are converted into Factor types, and the model is trained using a parallel computing framework.
[0011] Furthermore, the genotype information of the individuals to be tested is preprocessed and then input into the model to obtain source information.
[0012] Furthermore, the preprocessing includes 0 / 1 / 2 encoding and factorization.
[0013] Furthermore, the prediction accuracy was evaluated using 10-fold cross-validation, with accuracy and Kappa value used to assess the validation data.
[0014] Furthermore, it can be applied to marker-assisted population geographic tracing, quarantine monitoring, species evolution analysis, or germplasm resource identification.
[0015] Compared with the prior art, the beneficial effects of the present invention are: 1. This invention significantly reduces the number of detection sites, from 15,000 initial sites to 40, greatly reducing the cost of subsequent PCR amplification or gene chip detection and shortening the detection cycle.
[0016] 2. The marker sites selected by the present invention using machine learning are all sites that are significantly different between a specific population and other populations. The validation set accuracy is over 90.95%, which is far superior to the whole site analysis and has high classification accuracy.
[0017] 3. Based on the ideas of parallel computing and random forest ensemble, the present invention has strong noise resistance and generalization ability, is suitable for rapid and automated source tracing of massive samples, and has model robustness. Attached Figure Description
[0018] Figure 1 The importance ranking of the 40 SNPs in this invention; Figure 2 This is the accuracy curve of the model in this invention. Detailed Implementation
[0019] The technical solution of the present invention will be further described below with reference to the accompanying drawings and specific embodiments. Example 1
[0020] The whole-genome SNP locus assemblage of *F. guava* includes 40 core SNP loci for detecting allelic variations at these loci. The physical locations of these 40 SNP loci are determined based on the *F. guava* whole-genome reference version. Variation information for these 40 SNP loci is represented in the format of "chromosome number_physical location," and the 40 core SNP loci are shown below as BC0001-0040: BC0001:Chr03_24541573;BC0002:Chr01_135702962;BC0003:Chr01_166829815;BC0004:Chr01_160512719;BC0005:Chr03_1459438 6;BC0006:Chr01_137943379;BC0007:Chr01_144720904;BC0008:Chr03_5928598;BC0009:Chr01_144071968;BC0010:Chr01_162674 957;BC0011:Chr04_117913466;BC0012:Chr01_129458107;BC0013:Chr01_144582908;BC0014:Chr03_4105222;BC0015:Chr04_5779 412;BC0016:Chr01_129391099;BC0017:Chr04_5609399;BC0018:Chr03_157341613;BC0019:Chr03_7326333;BC0020:Chr01_1627238 27;BC0021:Chr04_71887113;BC0022:Chr03_40729220;BC0023:Chr04_81060136;BC0024:Chr04_115347547;BC0025:Chr01_147406 8;BC0026:Chr03_26767843;BC0027:Chr03_7162168;BC0028:Chr03_138439283;BC0029:Chr04_84803439;BC0030:Chr01_10992390 ;BC0031:Chr01_115453810;BC0032:Chr03_6092298;BC0033:Chr03_39610115;BC0034:Chr03_39440286;BC0035:Chr04_100099295 ;BC0036:Chr01_7432835;BC0037:Chr01_143956323;BC0038:Chr01_155252379;BC0039:Chr03_24425928;BC0040:Chr03_24535899.
[0021] In this embodiment, the SNP marker combination for guava fruit fly origination is derived from feature screening of a genome-wide SNP dataset, and the method for obtaining it includes the following steps: (1) Sample collection and DNA extraction: Guava fruit fly individuals from different geographical populations were collected, as shown in Table 1, and divided into Hainan population (experimental group) and non-Hainan population (control group); tissue DNA was extracted using a genomic DNA extraction kit.
[0022] Table 1 Sample Collection Information
[0023] (2) SNP detection and quality control: Whole genome sequencing was performed, and genotype data was read using snpgdsOpen; the obtained SNPs were quality controlled, and sites with high deletion rates and abnormal heterozygosity were removed.
[0024] (3) Digital encoding and preprocessing: The genotypes were digitized using snpgdsGetGeno, with 0 replacing wild genotypes, 1 replacing heterozygous genotypes, and 2 replacing mutant genotypes; the missing values were filled using the five-neighbor missing value imputation method (KNN), that is, the average of the five values before and after the missing value was used for filling.
[0025] (4) Recursive Feature Elimination (RFE) to select core sites: The RFE algorithm in the caret package of R language is used for feature selection; the size of the recursive feature subset is set to c(5, 10, 15, 20, 30, 40, 50, 80, 100, 130, 150, 180, 200); rfFuncs is used as the underlying random forest function to perform 10-fold cross-validation (10-fold CV).
[0026] (5) Determination of 40 core sites: According to the results of RFE, when the number of variables increased to 40, the model classification accuracy reached a peak of 90.95% and the Kappa coefficient reached 0.69675; at this time, the top 40 sites with the highest model fit importance score (MeanDecrease Gini) were extracted as the core SNP site combination.
[0027] (6) Construction of training set: The samples are randomly divided into training set and validation set in a 7:3 ratio; the training set data is converted into Factor type and the model is fitted using a parallel computing environment (doParallel).
[0028] (7) Construction of Random Forest Model: Random Forest Model is an ensemble algorithm composed of multiple classification decision trees; for the training set data as input samples, the final classification result is determined by the mode generated by voting of all decision trees; the random forest algorithm effectively reduces the variance of a single tree by constructing multiple decision trees and summarizing the results; through the recursive elimination of the RFE algorithm, the model can eliminate redundant information in 15,000 initial sites and achieve accurate discrimination by using only 40 of the most discriminative sites. Example 2
[0029] An application of SNP site combinations in the molecular identification of geographical origin in guava fruit fly includes the following steps: (1) Extraction of target sites: Genotypes of 40 core SNP sites from the whole genome data of the sample to be tested in Example 1 were extracted.
[0030] (2) Data preprocessing: The locus genotypes of the individuals to be tested are digitized (0 / 1 / 2); the dataset is traversed, and for individual missing loci, the five-neighbor mean imputation method is used to fill them.
[0031] (3) Model performance evaluation: The classifier performance was evaluated by 10-fold cross-validation; the scores of each evaluation index of the model in this example are shown in Table 2, which shows that the random forest model based on 40 SNP sites proposed in this application has extremely high source tracing accuracy.
[0032] Table 2: Comparison of model evaluation metrics for different feature subset sizes
[0033] By inputting the pre-processed test samples into the trained random forest classification model, high-precision automatic identification of the Hainan population of guava fruit flies can be achieved.
[0034] like Figure 1 and Figure 2 As shown, the importance of 40 SNPs was ranked and the model accuracy curve was plotted.
[0035] The above specific embodiments are merely several preferred embodiments of the present invention. Based on the technical solutions of the present invention and the relevant teachings of the above embodiments, those skilled in the art can make various alternative improvements and combinations to the above specific embodiments.
Claims
1. A method for obtaining the SNP locus combination of the whole genome of the guava fruit fly, characterized in that, include: DNA was obtained from individual guava fruit flies, and genotyping was performed using whole-genome sequencing technology. After quality control of the obtained SNPs, feature selection was performed using a random forest combined with a recursive feature elimination algorithm. Sorting by feature importance value from largest to smallest, the core SNP loci at which the classification accuracy reaches its peak are retained, namely the guava fruit fly population-specific combination loci.
2. The method for obtaining the SNP locus combination of the whole genome of guava fruit fly according to claim 1, characterized in that: The genome-wide SNP locus combination of the guava fruit fly includes 40 core SNP loci. The variation information of these 40 SNP loci is represented by chromosome numbers connected by underscores to indicate their physical locations.
3. The method for obtaining the SNP locus combination of the whole genome of the guava fruit fly according to claim 1, characterized in that, In the random forest classification model used for molecular identification of guava fruit fly populations: for the genotype of each guava fruit fly individual's SNP locus, 0 is used to replace the wild genotype, 1 is used to replace the heterozygous genotype, and 2 is used to replace the mutant genotype, thus obtaining a digital dataset. The dataset is divided into a training set and a validation set. A random forest model is trained using the training set. The training model is then evaluated for prediction accuracy using the validation set to obtain a classification model for identifying whether guava fruit flies originate from a specific region.
4. The method for obtaining the whole genome SNP locus combination of guava fruit fly according to claim 1, characterized in that: We used a random forest combined with a recursive feature elimination algorithm for 10-fold cross-validation screening.
5. The method for obtaining the whole genome SNP locus combination of guava fruit fly according to claim 1, characterized in that: During the feature selection process, the model accuracy under different numbers of SNPs was calculated by setting the subset step size, and the 40 core variables at which the accuracy reached its peak were retained.
6. The method for obtaining the SNP locus combination of the whole genome of the guava fruit fly according to claim 3, characterized in that: The digitized genotypes are converted into Factor types, and the model is trained using a parallel computing framework.
7. The method for obtaining the SNP locus combination of the whole genome of guava fruit fly according to claim 3, characterized in that: The genotype information of the individuals to be tested is preprocessed and then input into the model to obtain source information.
8. The method for obtaining the whole genome SNP locus combination of guava fruit fly according to claim 7, characterized in that: Preprocessing includes 0 / 1 / 2 encoding and factorization.
9. The method for obtaining the SNP locus combination of the whole genome of guava fruit fly according to claim 3, characterized in that: Prediction accuracy was evaluated using 10-fold cross-validation, with accuracy and Kappa value used to assess the validation data.
10. A method for obtaining the whole genome SNP locus combination of guava fruit fly according to any one of claims 1-9, characterized in that: It is applied to marker-assisted population geographic tracing, quarantine monitoring, species evolution analysis, or germplasm resource identification.