Application of a single-SNP and multi-SNP marker combination in the identification of seven grass carp geographic populations

Through the association of grass carp SNP loci genotype and geographical information, the combination of single SNP and multiple SNP markers and random forest learning model are used to solve the accuracy of grass carp geographical population identification, and the effective identification and utilization of new grass carp germplasm is achieved.

CN119464506BActive Publication Date: 2025-09-02CHINESE ACAD OF FISHERY SCI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411591700.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-08
Publication Date
2025-09-02
Estimated Expiration
2044-11-08

AI Technical Summary

Technical Problem

The existing technology cannot accurately distinguish the different geographical populations of grass carp, which has led to the hindrance of the creation of new grass carp germplasm and the inability to effectively protect and utilize grass carp germplasm resources.

Method used

Through the association of grass carp SNP loci genotype and geographical information, the identification of grass carp geographical populations is performed using a combination of single SNP and multiple SNP markers, and the population classification is performed in combination with a random forest learning model to achieve accurate identification of grass carp geographical populations.

Benefits of technology

Accurate and efficient identification of seven grass carp geographical populations has been achieved, supporting the determination of grass carp geographical population rights, protection and utilization of germplasm resources, and genetic evolution research.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119464506B_ABST
    Figure CN119464506B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of identification of geographical populations of grass carp, and specifically to the application of a combination of single SNP and multi-SNP markers in the identification of seven geographical populations of grass carp. The present application associates the SNP site genotype of grass carp with its geographical information, judges its geographical information through the SNP site genotype, and determines its geographical population. The method provided in the present application can accurately and efficiently identify seven geographical populations of grass carp. The method provided in the present application provides a new technical means for identification of new grass carp germplasm for the needs of identification of geographical population rights of grass carp, protection and utilization of germplasm resources, genetic evolution research, etc.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of identification of grass carp geographic populations, and in particular to the application of a single SNP and multi-SNP marker combination in the identification of seven grass carp geographic populations. Background Art

[0002] Grass carp (Ctenopharyngodon idella) is farmed in 30 provinces (municipalities, and autonomous regions) in my country, with the Yangtze River, Pearl River, and Heilongjiang River systems being the main distribution areas for grass carp in China. Due to environmental selection, grass carp populations in different river sections differ in growth performance, nutritional quality, and other traits, providing abundant breeding material for the creation of new grass carp germplasm. However, different geographical populations of grass carp lack typical appearance characteristics and cannot be distinguished intuitively and accurately. Without accurate geographical population information, the identification of grass carp germplasm resources and the recognition of geographical population rights cannot be carried out at all, and the creation of new grass carp germplasm will also be hindered. Therefore, the development of precise identification methods for geographical populations of grass carp is of great significance to the protection and utilization of seedlings. Summary of the Invention

[0003] The inventors of this application have creatively linked grass carp SNP genotypes with their geographic information, using SNP genotypes to determine geographic information and identify their geographic populations. The method provided in this application is capable of accurately and efficiently identifying seven grass carp geographic populations. This method provides a new technical means for identifying new grass carp germplasm for purposes such as identifying grass carp geographic populations, protecting and utilizing germplasm resources, and studying genetic evolution.

[0004] To this end, the embodiments of the present application disclose at least the following technical solutions:

[0005] The embodiments disclose a method for identifying geographic populations of grass carp. The method comprises: obtaining a genotype at a first SNP site of a grass carp sample to be tested; dividing the grass carp population into a first population and a second population based on the genotype at the first SNP site; obtaining a genotype at a second SNP site combination and a genotype at a third SNP site combination of the grass carp sample to be tested; identifying a North China population, a Huaihe River population, and an upper Yangtze River population from the first population based on the genotype at the second SNP site combination; and identifying a Xiangjiang River population, a middle Yangtze River population, a lower Yangtze River population, and a Pearl River population from the second population based on the genotype at the third SNP site combination. BRIEF DESCRIPTION OF THE DRAWINGS

[0006] Figure 1 This is the Manhattan plot of the genome-wide association analysis of the first population provided in the embodiment, where the red line represents the molecular marker screening threshold P<10E-05.

[0007] Figure 2This is the Manhattan plot of the genome-wide association analysis of the second group provided in the example, where the red line represents the molecular marker screening threshold P<10E-08.

[0008] Figure 3 Schematic diagram of the process of identifying geographical populations of grass carp provided in the embodiment.

[0009] Figure 4 This is a schematic diagram of the method flow of step S40 provided in an embodiment.

[0010] Figure 5 A flow chart of the method for obtaining the optimal splitting point provided in the embodiment.

[0011] Figure 6 This is a schematic diagram of the method flow of step S42 provided in the embodiment.

[0012] Figure 7 A schematic diagram of the training steps of the random forest learning model provided in the embodiment. DETAILED DESCRIPTION

[0013] In order to make the purpose, technical solutions and advantages of this application more clearly understood, the present application is further described in detail below with reference to the following examples. It should be understood that the specific examples described herein are merely for the purpose of explaining this application and are not intended to limit this application. Reagents not described in detail in this application are all conventional reagents and can be obtained from commercial channels; methods not specifically described in detail are all conventional experimental methods and can be obtained from the prior art.

[0014] To identify the geographic populations of grass carp, the present invention uses the SNP genotypes of grass carp to correlate with their geographic information, using the SNP genotypes to determine their geographic information and identify their geographic population. The method provided in this application can accurately and efficiently identify seven grass carp geographic populations. This method provides a new technical means for identifying new grass carp germplasm for purposes such as identifying grass carp geographic population rights, protecting and utilizing germplasm resources, and studying genetic evolution.

[0015] The embodiment discloses a method for identifying a geographical population of grass carp. Figure 3 As shown, the method includes:

[0016] S10, obtaining the genotype of the grass carp sample to be tested at the first SNP site;

[0017] S20, dividing the grass carp population into a first population and a second population according to the genotype of the first SNP site;

[0018] S30, obtaining the genotype of the second SNP site combination and the genotype of the third SNP site combination of the grass carp sample to be tested;

[0019] S40. According to the genotype of the second SNP site combination, the North China population, the Huaihe River population, and the upper Yangtze River population are identified from the first population; according to the genotype of the third SNP site combination, the Xiangjiang River population, the middle Yangtze River population, the lower Yangtze River population, and the Pearl River population are identified from the second population.

[0020] In some embodiments, the first SNP site is a single nucleotide mutation formed by A>T found at position 17682028 nt of chromosome 1 of the grass carp genome.

[0021] In some embodiments, the second SNP site combination is as shown in Table 1.

[0022] Table 1 Second SNP site combination

[0023]

[0024]

[0025]

[0026]

[0027] In some embodiments, the third SNP site combination is as shown in Table 2.

[0028] Table 2 Combinations of the third SNP loci

[0029]

[0030]

[0031]

[0032]

[0033] The registration accession number of the grass carp genome in the NCBI database is HZGC01, reference Wu, CS; ZYMa; GDZheng; SMZou; XJZhangY.A.Zhang Chromosome-level genome assembly of grass carp (Ctenopharyngodon idella) provides insights into its genome evolution. BMC Genomics, 2022, 23, 271.10.1186 / s12864-022-08503-x.

[0034] In addition, the embodiment also discloses the development process of the first SNP site, the second SNP site combination, and the third SNP site combination. Specifically including:

[0035] (1) Combining the collected grass carp samples, a total of 726 samples were constructed. This group includes seven grass carp geographical populations: North China, Huaihe River, Xiangjiang River, middle Yangtze River, lower Yangtze River, Pearl River, and upper Yangtze River.

[0036] (2) For the collected grass carp samples, genomic DNA was extracted using the conventional CTAB method. After library construction, whole genome resequencing was performed using the DNBSEQ-T7 platform. The main band of genomic DNA gel electrophoresis showed no obvious degradation, the concentration should be greater than 70 ng / μL, the optical absorbance ratio at OD260 / 280 should be in the range of 1.7 to 2.0, and the optical absorbance ratio at OD260 / 230 should be in the range of 1.8 to 2.2. After filtering low-quality sequences, sequences containing N or adapters, the sequencing reads were aligned to the grass carp reference genome (NCBI accession number: HZGC01), and genomic variations were detected using GATK. Sites with minor allele frequencies below 0.01 and SNP sites with 20% sample deletions were filtered, and the remaining SNP molecular markers were used for subsequent association analysis.

[0037] (3) The seven grass carp geographic populations were divided into the first and second populations by the first SNP locus (1:17682028, A / T). The first population included the North China population, the Huaihe River population, and the upper Yangtze River population, and the genotypes of the first SNP locus were all AA homozygous. The second population included the Xiangjiang River population, the middle Yangtze River population, the lower Yangtze River population, and the Pearl River population, and the genotypes of the first SNP locus were mostly AT heterozygous.

[0038] Table 1 Genotyping statistics of the first SNP locus in seven different grass carp geographical populations

[0039] Grass carp geographical populations 1:17682028(A / T) genotype Genotype ratio North China AA 157 / 157(100%) Huaihe River AA 51 / 51(100%) Upper Yangtze River AA 11 / 11(100%) Xiangjiang River AT 25 / 26(96%) middle reaches of the Yangtze River AT 194 / 207(94%) Lower Yangtze River AT 157 / 169(93%) Pearl River AT 24 / 28(86%)

[0040] (4) Each grass carp geographical population was distinguished from the first group of AA genotype and the second group of AT genotype, and the geographical information of each sample in the two groups was converted into a continuous natural number as the phenotype. The genetic loci after quality control filtering were used as genotypes, and the mixed linear model of GEMMA was used to perform genome-wide association analysis. Figure 1 As shown in Figure 2, the genome-wide association analysis SNP marker screening threshold P<10E-05 for the geographical coding values ​​of each grass carp individual in the first population (North China, Huaihe River, and upper Yangtze River are coded as natural numbers 1, 2, and 3, respectively) obtained 158 SNP markers. Figure 2As shown, a genome-wide association analysis of the geographic coding values ​​of grass carp individuals in the second population (the Xiangjiang River, the middle Yangtze River, the lower Yangtze River, and the Pearl River are coded as natural numbers 1, 2, 3, and 4, respectively) screened for significant SNP markers at a threshold of P < 10E-08, resulting in 150 SNP markers. The different natural numbers converted from the geographic information of the two populations were independent of each other.

[0041] In some embodiments, as Figure 4 As shown, step S20 includes: S21. If the grass carp sample to be tested has an AA genotype at position 17682028 nt of chromosome 1, the grass carp sample to be tested belongs to the first group.

[0042] In some embodiments, as Figure 4 As shown, step S20 further includes: S22, if the grass carp sample to be tested has an AT genotype at position 17682028 nt of chromosome 1, then the grass carp sample to be tested belongs to the second population.

[0043] In some embodiments, as Figure 4 As shown, step S30 includes: S31, obtaining the genotype of the second SNP site combination of the grass carp sample to be tested.

[0044] In some embodiments, as Figure 4 As shown, step S30 includes: S32, obtaining the genotype of the third SNP site combination of the grass carp sample to be tested.

[0045] In some embodiments, as Figure 4 As shown, step S40 includes: S41, identifying the North China population, Huaihe population, and upper Yangtze River population from the first population based on the genotype of the second SNP site combination.

[0046] In some embodiments, as Figure 4 As shown, step S40 includes: S42, identifying the Xiangjiang population, the middle Yangtze River population, the lower Yangtze River population, and the Pearl River population from the second population based on the genotype of the third SNP site combination.

[0047] In some embodiments, as Figure 5 As shown, step S41 specifically includes:

[0048] S411, obtaining a first training set of the first population, wherein the first training set includes second SNP site combination genotype information and a first geocoding value of a plurality of grass carp samples from the first population;

[0049] S412: Perform bootstrap sampling with replacement from the first training set to construct multiple first sub-datasets, configure a decision tree for each of the first sub-datasets, and construct a random forest learning model using the decision trees;

[0050] S413, obtaining the second SNP site combination genotype information and the first geocode value of each of the first sub-datasets to train a random forest learning model to obtain the optimal splitting point of the multiple feature splits of each decision tree, thereby obtaining a trained random forest learning model;

[0051] S414, obtaining the second SNP site combination genotype information of the grass carp sample to be tested;

[0052] S415, inputting the second SNP site combination genotype information into the trained random forest learning model to obtain a first predicted value of the grass carp sample to be tested;

[0053] S416. Identify the geographical information of the grass carp population to be tested in the first population according to the first prediction value.

[0054] In some embodiments, in step S411, a total of 220 grass carp samples are obtained in the first training set.

[0055] In some embodiments, the second SNP site combined genotype information is the record information of the genotype of each SNP site as shown in Table 1. For example, for the G>A mutation, the homozygous genotype such as GG is recorded as "1", the homozygous mutant genotype such as AA is recorded as "-1", and the heterozygous genotype such as GA is recorded as "0".

[0056] In some embodiments, the first geocoding value is a natural number of geocoding values ​​corresponding to three grass carp geographical populations in North China, Huaihe River, and upper Yangtze River: 1, 2, 3.

[0057] In step S413 provided in some embodiments, the second SNP site combined genotype information is used as the input of each decision tree, and the first geocode value is used as the output of each decision tree to train each decision tree.

[0058] In step S416 provided in some embodiments, the first predicted value is a value that is relatively close to any of the first geocoding values. The higher the degree of closeness, the more it represents that the grass carp sample comes from a geographical location in North China, Huaihe River and the upper reaches of the Yangtze River in the first group.

[0059] In some embodiments, as Figure 6 As shown, step S42 specifically includes:

[0060] S421, obtaining a second training set of the second population, wherein the second training set includes third SNP site combination genotype information and second geocoding values ​​of a plurality of grass carp samples from the second population;

[0061] S422: Perform bootstrap sampling with replacement from the second training set to construct multiple second sub-datasets, configure a decision tree for each of the second sub-datasets, and construct a random forest learning model using the decision trees;

[0062] S423, obtaining the third SNP site combination genotype information and the second geocode value of each second sub-dataset to train the random forest learning model to obtain the optimal splitting point of the multiple feature splits of each decision tree, thereby obtaining a trained random forest learning model;

[0063] S424, obtaining the third SNP site combination genotype information of the grass carp sample to be tested;

[0064] S425, inputting the third SNP site combined genotype information into the trained random forest learning model to obtain a second predicted value of the grass carp sample to be tested;

[0065] S426. Identify the geographical information of the grass carp population to be tested in the second population according to the second prediction value.

[0066] In some embodiments, in step S411, a total of 506 grass carp samples are obtained in the first training set.

[0067] In some embodiments, the third SNP site combined genotype information is the record information of the genotype of each SNP site as shown in Table 2. For example, for the G>A mutation, the homozygous genotype such as GG is recorded as "1", the homozygous mutant genotype such as AA is recorded as "-1", and the heterozygous genotype such as GA is recorded as "0".

[0068] In some embodiments, the second geocoding value is a natural number of geocoding values ​​corresponding to the Xiangjiang River group, the middle reaches of the Yangtze River group, the lower reaches of the Yangtze River group, and the Pearl River group: 1, 2, 3, 4.

[0069] In step S423 provided in some embodiments, the third SNP site combined genotype information is used as the input of each decision tree, and the second geocoding value is used as the output of each decision tree, so as to train each decision tree.

[0070] In step S426 provided in some embodiments, the first prediction value is a value that is relatively close to any of the first geocoding values. The higher the degree of closeness, the more it represents that the grass carp sample comes from a geographical location in the Xiangjiang River, the middle reaches of the Yangtze River, the lower reaches of the Yangtze River and the Pearl River in the first group.

[0071] In some embodiments, the random forest model consists of multiple decision trees (n t ), each decision tree (T b ) are trained independently. For a given input sample x, each tree will give a corresponding prediction value

[0072] In some embodiments, such as Figure 7 As shown, the training steps of the random forest learning model include:

[0073] S501. Randomly obtain any sub-dataset from the training set, construct a decision tree, and for each node of the decision tree, randomly select one or more features from the sub-dataset for splitting;

[0074] S502, obtaining a first mean square error before splitting and a second mean square error after splitting for each feature at each node of each decision tree;

[0075] S503: taking the difference between the first mean square error and the second mean square error as a gain;

[0076] S504: The detected splitting point with the maximum gain number is used as the optimal splitting point, and the corresponding splitting feature and splitting threshold are configured at the optimal splitting point;

[0077] S505, repeating steps S501 to S504, training each decision tree of the random forest model separately to obtain the best splitting point for each decision tree, and configuring the trained splitting feature and splitting threshold for each best splitting point;

[0078] S506: Determine the training degree of the random forest learning model according to the training error of the random forest learning model.

[0079] Steps S501 to S506 are applicable to the training process of the random forest learning model in steps S41 and S42 and are not described in detail here.

[0080] In step S415 provided in some embodiments, the recorded value of the second SNP site combined genotype information of the grass carp sample to be tested is compared with the splitting threshold. If the recorded value of the second SNP site combined genotype information of the sample to be tested is less than or equal to the splitting threshold, the search continues using the left child node. If it is greater than the splitting threshold, the search continues using the right child node until a leaf node is reached. If the splitting threshold is "NA", the splitting threshold of the corresponding node is directly read as the predicted value.

[0081] In step S425 provided in some embodiments, the recorded value of the third SNP site combined genotype information of the grass carp sample to be tested is compared with the splitting threshold. If the recorded value of the third SNP site combined genotype information of the sample to be tested is less than or equal to the splitting threshold, the search continues using the left child node. If it is greater than the splitting threshold, the search continues using the right child node until a leaf node is reached. If the splitting threshold is "NA", the splitting threshold of the corresponding node is directly read as the predicted value.

[0082] Among them, the first mean square error can be expressed as: Among them, n t is the number of samples in the decision tree node t, y i is the target value of the i-th sample, is the average value of all sample targets in decision tree node t.

[0083] Among them, the second mean square error can be expressed as: Among them, n tL and n tR Decision tree nodes tL and t Number of samples, MSE in R tL and MSE tR is the mean square error of the nodes.

[0084] The gain can be expressed as: Gain = MSE t -MSE t,split

[0085] Among them, the error of the trained random forest learning model can be expressed as Where n is the number of samples, yi is the actual observation value of the i-th sample; is the predicted value of the i-th sample. The smaller the RMSE value, the higher the prediction accuracy of the model.

[0086] Table 2 shows the prediction results for the first and second grass carp populations using the method provided in this application. The 158 SNP markers in the first population and the corresponding random forest machine learning model predicted the geographic information of the North China population, the Huaihe River population, and the upper Yangtze River population with an accuracy of 0.842. The 150 SNP markers in the second population and the corresponding random forest machine learning model predicted the geographic information of the Xiangjiang River population, the middle Yangtze River population, the lower Yangtze River population, and the Pearl River population with an accuracy of 0.9327.

[0087] Table 2

[0088] Classification based on single marker genotype The first group The second group SNP markers 158 second SNP loci combination 150 third SNP loci combination Grass carp geographical populations North China, Huaihe River, and upper Yangtze River Xiangjiang River, middle Yangtze River, lower Yangtze River, Pearl River Prediction accuracy 0.842 0.9327 Standard deviation of forecast accuracy 0.0058 0.0018 Root mean square error 0.0001 0.0069

[0089] In some embodiments, 30 grass carp samples with 3 AA homozygous types (grass carp geographical populations in North China, Huaihe River, and upper Yangtze River) and 30 grass carp samples with 4 AT heterozygous types (grass carp geographical populations in Xiangjiang River, middle reaches of the Yangtze River, lower reaches of the Yangtze River, and Pearl River) were collected as independent test sets.

[0090] Genomic DNA was extracted separately using the conventional CTAB method, and whole-genome resequencing was performed using the DNBSEQ-T7 platform of Huazhi Biotechnology Co., Ltd. The sequencing amount of each sample was at least 10 times the genome coverage. For the three grass carp geographical populations of North China, Huaihe River, and upper Yangtze River with AA homozygous type, genotyping detection of the 158 SNP markers in Example 1 was performed; for the Xiangjiang River population, the middle Yangtze River population, the lower Yangtze River population, and the Pearl River population with AT heterozygous type, genotyping detection of the 150 SNP markers in Example 1 was performed. The genotypes obtained for each sample were directly introduced into the random forest regression model constructed in Example 1. Based on the relationship between the classification value of the classification variable and the sample genotype coding value, the predicted value was calculated to identify the geographical information of the test sample.

[0091] Table 3 Identification results of the second SNP locus combination in three grass carp geographical populations in North China, Huaihe River and upper Yangtze River

[0092]

[0093]

[0094] Table 4 Identification results of the third SNP locus combination in the Xiangjiang population, the middle Yangtze River population, the lower Yangtze River population, and the Pearl River population

[0095] Sample number Recording geographic information Second geocoded value Second prediction value Identification results XB-2342 Xiangjiang River 1 1.464 correct XB-2359 Xiangjiang River 1 1.39 correct XBR-10 Xiangjiang River 1 1.309 correct XBR-12 Xiangjiang River 1 1.092 correct XF2332 Xiangjiang River 1 1.574 mistake XM-12367 Xiangjiang River 1 1.551 mistake XM-22376 Xiangjiang River 1 1.291 correct 10121 middle reaches of the Yangtze River 2 2 correct 10148 middle reaches of the Yangtze River 2 2 correct 10173 middle reaches of the Yangtze River 2 2 correct 10193 middle reaches of the Yangtze River 2 2.008 correct CAS-13 middle reaches of the Yangtze River 2 2.031 correct CAS-38 middle reaches of the Yangtze River 2 2 correct CBR-20 middle reaches of the Yangtze River 2 1.992 correct CCS-29 middle reaches of the Yangtze River 2 2.068 correct 20231106-106 Lower Yangtze River 3 2.985 correct 20231106-118 Lower Yangtze River 3 2.99 correct 20231106-13 Lower Yangtze River 3 2.992 correct 20231106-159 Lower Yangtze River 3 2.975 correct 20231106-168 Lower Yangtze River 3 2.997 correct 20231106-176 Lower Yangtze River 3 2.951 correct 20231106-185 Lower Yangtze River 3 2.964 correct 20231106-194 Lower Yangtze River 3 2.993 correct ZhJ01 Pearl River 4 3.869 correct ZhJ03 Pearl River 4 3.965 correct ZhJ05 Pearl River 4 3.875 correct ZhJ07 Pearl River 4 3.824 correct ZhJ09 Pearl River 4 3.976 correct ZhJ11 Pearl River 4 3.909 correct ZhJ13 Pearl River 4 3.922 correct

[0096] The method for identifying a grass carp geographic population provided in the embodiments of the present application can be performed by a device for identifying a grass carp geographic population, or a control unit within the device for performing the method. In the embodiments of the present application, the method for identifying a grass carp geographic population is performed by the device for identifying a grass carp geographic population as an example to illustrate the device for identifying a grass carp geographic population provided in the embodiments of the present application.

[0097] Optionally, an embodiment of the present application also provides an electronic device, including a processor, a memory, and a program or instruction stored in the memory and executable on the processor. When the program or instruction is executed by the processor, each process of the above-mentioned method embodiment for identifying the geographical population of grass carp is implemented, and the same technical effect can be achieved. To avoid repetition, it will not be described here.

[0098] It should be noted that the electronic devices in the embodiments of the present application include the above-mentioned mobile electronic devices and non-mobile electronic devices.

[0099] An embodiment of the present application also provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, each process of the above-mentioned method embodiment for identifying the geographical population of grass carp is implemented, and the same technical effect can be achieved. To avoid repetition, it will not be repeated here.

[0100] The processor is the processor in the electronic device described in the above embodiment. The readable storage medium includes a computer-readable storage medium, such as a computer read-only memory (ROM), random access memory (RAM), a magnetic disk, or an optical disk.

[0101] An embodiment of the present application further provides a chip, which includes a processor and a communication interface, wherein the communication interface is coupled to the processor, and the processor is used to run programs or instructions to implement the various processes of the above-mentioned method embodiment for identifying the geographical population of grass carp, and can achieve the same technical effect. To avoid repetition, it will not be repeated here.

[0102] The above is only a preferred specific implementation method of the present application, but the scope of protection of the present application is not limited thereto. Any changes or replacements that can be easily thought of by any technician familiar with this technical field within the technical scope disclosed in this application should be covered by the scope of protection of the present application.

Claims

1. Methods for identifying geographical populations of grass carp, including: Obtaining the genotype of the grass carp sample to be tested at the first SNP site; If it is detected that the grass carp sample to be tested has an AA genotype at position 17682028 nt of chromosome 1, the grass carp sample to be tested is the first population; If it is detected that the grass carp sample to be tested has an AT genotype at position 17682028 nt of chromosome 1, the grass carp sample to be tested belongs to the second population; Obtaining the genotype of the second SNP site combination and the genotype of the third SNP site combination of the grass carp sample to be tested; Identifying the North China population, the Huaihe River population, and the upper Yangtze River population from the first population based on the genotype of the second SNP locus combination, and identifying the Xiangjiang River population, the middle Yangtze River population, the lower Yangtze River population, and the Pearl River population from the second population based on the genotype of the third SNP locus combination; Among them, the first SNP site is a single nucleotide mutation formed by A>T found at the 17682028nt position of chromosome 1 of the grass carp genome, the SNP sites in the second SNP site combination are shown in Table 1, and the third SNP site combination is shown in Table 2; “Identifying the North China population, the Huaihe River population, and the upper Yangtze River population from the first population based on the genotype of the second SNP site combination” includes: Obtaining a first training set of the first population, the first training set comprising genotype information and first geocoding values ​​of a second SNP site combination of a plurality of grass carp samples from the first population; uniformly sampling with replacement from the first training set to construct a plurality of first sub-datasets, configuring a decision tree for each of the first sub-datasets, and constructing a random forest learning model from the decision trees; Obtaining genotype information and a first geocode value of the second SNP site combination of each of the first sub-datasets to train a random forest learning model to obtain optimal splitting points for multiple feature splits of each of the decision trees, thereby obtaining a trained random forest learning model; Obtaining genotype information of a second SNP site combination of the grass carp sample to be tested; Inputting the genotype information of the second SNP site combination into the trained random forest learning model to obtain a first predicted value of the grass carp sample to be tested; Identifying the geographical information of the grass carp sample to be tested in the first population according to the first prediction value; “Identifying the Xiangjiang River population, the middle Yangtze River population, the lower Yangtze River population, and the Pearl River population from the second population based on the genotype of the third SNP locus combination” specifically includes: Obtaining a second training set of the second population, wherein the second training set includes genotype information of a third SNP site combination and a second geocoding value of a plurality of grass carp samples from the second population; uniformly sampling with replacement from the second training set to construct a plurality of second sub-datasets, configuring a decision tree for each of the second sub-datasets, and constructing a random forest learning model from the decision tree; Obtaining the genotype information of the third SNP site combination and the second geocode value of each second sub-dataset to train a random forest learning model to obtain an optimal splitting point for multiple feature splits of each decision tree, thereby obtaining a trained random forest learning model; Obtaining genotype information of the third SNP site combination of the grass carp sample to be tested; Inputting the genotype information of the third SNP site combination into the trained random forest learning model to obtain a second predicted value of the grass carp sample to be tested; The geographical information of the grass carp sample to be tested in the second population is identified according to the second prediction value.

2. The method according to claim 1, wherein the genotype information of the second SNP site combination is the record information of the genotype of each SNP site as shown in Table 1, the homozygous genotype information is recorded as "1", the homozygous mutant genotype information is recorded as "-1", and the heterozygous genotype information is recorded as "0".

3. The method according to claim 1, wherein the first geocoding value is a natural number of geocoding values ​​corresponding to three grass carp geographical populations in North China, Huaihe River, and upper reaches of the Yangtze River: 1, 2, 3.

4. The method according to claim 1, wherein the genotype information of the third SNP site combination is the record information of the genotype of each SNP site as shown in Table 2, the homozygous genotype information is recorded as "1", the homozygous mutant genotype information is recorded as "-1", and the heterozygous genotype information is recorded as "0".

5. According to the method described in claim 1, the second geocoding value is the natural number of the geocoding values ​​corresponding to the Xiangjiang River group, the middle reaches of the Yangtze River group, the lower reaches of the Yangtze River group, and the Pearl River group: 1, 2, 3, 4.

Citation Information

Patent Citations

  • Application of SNP (Single Nucleotide Polymorphism) molecular marker combination in identification of 16 carp varieties

    CN117210580A

  • Group of SNP (Single Nucleotide Polymorphism) markers capable of being applied to rapid identification of geographical population of eriocheir sinensis

    CN118600028A