Procambarus clarkii high-meat-content strain breeding method based on SNP (Single Nucleotide Polymorphism) marker
By integrating high-throughput sequencing, machine learning and molecular marker-assisted breeding, using the silica membrane column method to extract DNA, combined with XGBoost-SHAP multimodal association analysis and competitive allele-specific PCR, the problems of low efficiency and high cost of secondary verification in the breeding of high meat content traits in Procambarus clarkii were solved, and efficient and low-cost breeding of high meat content varieties was achieved.
Patent Information
- Application Number
- CN202510858820.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-25
- Publication Date
- 2025-09-23
AI Technical Summary
The existing technology has the problems of low efficiency and high cost of secondary verification in the breeding of high meat content traits in Procambarus clarkii, making it difficult to quickly screen out core SNP molecular markers that are highly correlated with meat content. In addition, the multi-gene regulation of the high meat content trait leads to insufficient QTL positioning resolution and robustness. The typing platform and standardized process have not been established, making it difficult to efficiently transform and promote in actual breeding and large-scale breeding.
DNA was extracted using the silica gel membrane column method, and whole-genome SNP typing was performed in combination with high-throughput sequencing. The XGBoost model and SHAP value were used to screen candidate SNP sites. A high-throughput typing platform was constructed through competitive allele-specific PCR. Multi-environment validation was performed and core SNP markers were identified. A high-throughput typing platform was constructed for typing screening of breeding populations.
The time from sample collection to core SNP marker screening has been significantly shortened, breeding efficiency and accuracy have been improved, costs have been controlled within 30% of traditional phenotypic breeding, the breeding cycle has been shortened by more than 50%, and rapid screening and large-scale breeding of high-meat content lines have been achieved.
Smart Images

Figure CN120683263A_ABST
Abstract
Description
Technical Field
[0001] The invention relates to the technical field of aquatic animal genetic breeding and molecular biology, and in particular to a method for breeding a high-meat content strain of Procambarus clarkii based on SNP markers. Background Art
[0002] As an important freshwater aquaculture species globally, traditional breeding of Procambarus clarkii relies heavily on phenotypic observations, such as body weight and growth rate, for generational selection. This results in a long breeding cycle, is susceptible to environmental influences, and has low breeding efficiency. With the development of molecular biology, markers such as SSR and RAPD have been used to study the genetic diversity and kinship of aquatic animals. However, these markers have limited throughput, poor reproducibility, and insufficient genome coverage. In recent years, high-throughput sequencing technology has enabled SNP markers to emerge in aquatic breeding: SNPs are distributed throughout the genome, with high detection throughput and gradually decreasing costs. They can be used to construct high-density genetic maps, precisely locate trait-related QTLs, and achieve high-throughput genotyping.
[0003] However, there are still many shortcomings in the breeding of high meat content traits in Procambarus clarkii: on the one hand, although a large number of candidate SNP site identifications have been reported, the secondary verification efficiency is low and the experimental cost is high, making it difficult to quickly screen out "core" markers that are highly correlated with meat content; on the other hand, complex traits such as high meat content are regulated by multiple genes, and the existing QTL positioning resolution and robustness cannot meet the needs of precise breeding; in addition, the SNP typing platform and standardized process for high meat content have not yet been established, resulting in the research results of molecular marker-assisted selection being difficult to efficiently transform and promote in actual breeding and large-scale breeding. Summary of the Invention
[0004] In response to the shortcomings of the existing technology, the present invention provides a method for breeding high-meat content crayfish strains based on SNP markers, which solves the problem of how to screen out core SNP molecular markers that are significantly associated with the high meat content trait in a short time, at low cost and with high accuracy.
[0005] To achieve the above objectives, the present invention is implemented through the following technical solutions: a method for breeding a high-meat content strain of Procambarus clarkii based on SNP markers, comprising:
[0006] S1. DNA was extracted from Procambarus clarkii individuals with significant differences in meat content phenotype using a silica gel column method.
[0007] S2. performing genome-wide SNP typing on the DNA using high-throughput sequencing;
[0008] S3. Constructing a training dataset based on the SNP typing results, and using the XGBoost model combined with the SHAP value to evaluate and screen candidate SNP sites significantly associated with high meat content through the training dataset;
[0009] S4. Typing and validating the candidate SNP sites in a multi-environment validation population, and determining several core SNP markers based on allele frequencies and genetic effects;
[0010] S5. A high-throughput typing platform is constructed based on the core SNP markers, and the breeding population is typed and screened by the high-throughput typing platform to retain individuals with excellent genotypes, and the individuals with excellent genotypes are bred in large quantities to obtain high meat content lines.
[0011] Preferably, the ultraviolet absorbance ratio of DNA extracted by the silica gel membrane column method ranges from 1.8 to 2.0, and the concentration is not less than 50 ng / μL.
[0012] Preferably, the high-throughput sequencing has a sequencing depth of no less than 10× and a coverage of ≥95%.
[0013] Preferably, the XGBoost model is specifically:
[0014]
[0015] Among them, y i is the observed meat content phenotypic value of the i-th sample, is the predicted value of the i-th sample in the t-1th iteration, f k is the decision function corresponding to the kth regression tree. The input is the feature vector x, and the output is the predicted increment of the target value by the tree. f t (x i ) is the t-th regression tree for sample x i The incremental prediction function, l(·,·) is a differentiable convex loss function, Ω(f) is the regularization term of the tree model, where: γ is the regularization coefficient for adjusting the number of leaf nodes T, T is the total number of leaf nodes in the regression tree, λ is the L2 regularization coefficient for the leaf weight, w j is the output weight of the jth leaf node.
[0016] Preferably, the SHAP value is obtained by the following model formula:
[0017]
[0018] Among them, F is the total feature set, |F| is the total number of features, S is any feature subset that does not contain the jth feature, |S| is the number of features in the subset S, and f S (x S ) represents the model output when only feature subset S is used for prediction, f S∪{j} (x) is the expected output of the model when only the feature subset S∪{j} is used for prediction, X S∪{j}It means that only the sample feature vectors of the latitude corresponding to the feature subset S∪{j} are retained.
[0019] Preferably, the minor allele frequency of the core SNP marker in the multi-environment validation population is ≥ 0.05, and the two-sided test P value corresponding to the single marker genetic effect is ≤ 1×10 -5 .
[0020] Preferably, the high-throughput typing platform adopts competitive allele-specific PCR technology, and the single-site typing call rate of the high-throughput typing platform is ≥95% and the typing accuracy is ≥98%.
[0021] Preferably, the favorable allele content in the core SNP markers carried by the superior genotype individual is ≥80%.
[0022] The present invention provides a method for breeding a high-meat content strain of Procambarus clarkii based on SNP markers. It has the following beneficial effects:
[0023] This SNP-based method for selecting high-meat content strains of Procambarus clarkii significantly shortens the time from sample collection to core SNP marker screening by integrating high-throughput sequencing, machine learning, and molecular marker-assisted selection. A silica gel column method was used to rapidly extract high-purity DNA, covering all major variant sites across the genome in one go. Subsequently, XGBoost-SHAP multimodal association analysis was used to accurately locate candidate loci based on multi-source data from meat content phenotypes, near-infrared spectroscopy, and morphological indicators, effectively avoiding the false positives and low resolution issues associated with single-phenotype analysis in traditional GWAS. Finally, the method was validated in multiple environmental populations using a competitive allele-specific PCR high-throughput typing platform, and based on MAFs ≥ 0.05 and P ≤ 1×10 -5 The genetic effect threshold was set to quickly lock the core SNP, achieving an efficient transition from massive candidates to precise markers.
[0024] This invention also establishes a process for early genotyping and high-frequency screening of superior genotypes, significantly improving breeding efficiency and accuracy. After core SNP marker verification, a genotype-trait prediction model is established based on data with a sequencing depth of at least 10x and a coverage of ≥95%, enabling genotyping and screening of breeding populations at the seedling stage. The entire process costs less than 30% of traditional phenotypic selection, with an accuracy rate of ≥98%, and shortens the average breeding cycle by over 50%. This not only meets the urgent demand for high-meat-content strains in large-scale farming, but also has great potential for industrialization and promotion. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] Figure 1 It is a schematic diagram of a process for realizing the invention;
[0026] Figure 2This is the correlation analysis chart between R1 and meat content;
[0027] Figure 3 This is the correlation analysis chart between R2 and meat content. DETAILED DESCRIPTION
[0028] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0029] Example 1
[0030] like Figure 1 As shown, the present invention provides a method for breeding a high-meat content crayfish strain based on SNP markers, comprising: S1. extracting DNA from crayfish individuals with significant differences in meat content phenotype using a silica gel column method. The DNA extracted using the silica gel column method has an ultraviolet absorbance ratio range of 1.8 to 2.0 and a concentration of no less than 50 ng / μL.
[0031] S2. Perform genome-wide SNP typing of DNA using high-throughput sequencing. High-throughput sequencing should have a minimum sequencing depth of 10× and a coverage of ≥95%.
[0032] S3. A training dataset was constructed based on the SNP typing results. The XGBoost model was then combined with the SHAP value to evaluate and screen candidate SNP sites that were significantly associated with high meat content. The XGBoost model is specifically:
[0033]
[0034] Among them, y i is the observed meat content phenotypic value of the i-th sample, is the predicted value of the i-th sample in the t-1th iteration, f k is the decision function corresponding to the kth regression tree. The input is the feature vector x, and the output is the predicted increment of the target value by the tree. f t (x i ) is the t-th regression tree for sample x i The incremental prediction function, l(·,·) is a differentiable convex loss function, Ω(f) is the regularization term of the tree model, where: γ is the regularization coefficient for adjusting the number of leaf nodes T, T is the total number of leaf nodes in the regression tree, λ is the L2 regularization coefficient for the leaf weight, w j is the output weight of the jth leaf node.
[0035] The SHAP value is derived from the following model formula:
[0036]
[0037] Among them, F is the total feature set, |F| is the total number of features, S is any feature subset that does not contain the jth feature, |S| is the number of features in the subset S, and f S (x S ) represents the model output when only feature subset S is used for prediction, f S∪{j} (x) is the expected output of the model when only the feature subset S∪{j} is used for prediction, X S∪{j} It means that only the sample feature vectors of the latitude corresponding to the feature subset S∪{j} are retained.
[0038] The specific implementation is as follows:
[0039] Sample and phenotypic data preparation:
[0040] A total of 200 adult Procambarus clarkii individuals were collected, and the meat content was measured to be 7.8% to 26.4%, with an average of 14.8% and a standard deviation of 4.0%.
[0041] The 200 samples were randomly divided into a training set (160 animals) and a validation set (40 animals) in a ratio of 8:2.
[0042] SNP typing and multimodal feature construction:
[0043] PE150 sequencing was performed, with an average sequencing depth of 12× for each individual in the training set and a genome coverage of 98.2%.
[0044] A total of 350,000 SNP sites were initially detected, and 180,000 remained after filtering out those with MAF < 0.05.
[0045] At the same time, muscle near-infrared spectra (700-1000 nm) were collected from individuals in the training set, and the first five principal components were extracted by PCA.
[0046] Four morphological indices, including body weight, body length, total length of abdominal segment and width of the second abdominal segment, were measured, and the abdominal muscle weight of each individual was measured using the anatomical method.
[0047] Meat content calculation:
[0048]
[0049] Among them, M t is the wet weight of the whole shrimp, M m It is the wet weight of the abdominal muscles.
[0050] Correlation between body shape data and meat content: Figure 2 、 3 As shown in the previous experiments of the research group, the morphological index body length3 / body weight (R1), total length of abdominal segment / width of the second abdominal segment (R2) can reflect the meat content of an individual to a certain extent. Incorporating these easily accessible external morphological features into the multimodal training set can compensate for the deficiency of near-infrared spectroscopy in capturing body shape differences and enhance the model's ability to explain meat content.
[0051] Finally, a training data set containing 180,007 features (180,000 SNPs + 5 spectral principal components + 2 morphological indicators) and 160 samples was formed.
[0052] XGBoost model training and performance:
[0053] The parameters used are: learning rate 0.1, maximum tree depth 6, row sampling rate 0.8, and regularization parameter set to 1.
[0054] Training results:
[0055] Training set: RMSE = 0.024, coefficient of determination R 2 =0.84.
[0056] Validation set: RMSE = 0.027, coefficient of determination R 2 =0.81.
[0057] SHAP value calculation and candidate SNP screening:
[0058] For the trained model, the SHAP value of each feature is calculated, and the average absolute SHAP value of the 180,007 columns of features is counted.
[0059] The top 1% (1800) SNPs were selected as candidate sites by sorting from high to low according to the average SHAP value; the top 200 sites were further selected for subsequent verification.
[0060] Results and efficiency:
[0061] In this embodiment, the above process is completed within 3 days from sequencing data acquisition, model training to candidate SNP screening.
[0062] The 200 candidate SNP sites in the training set cumulatively explained 78% of the meat content variation, meeting the needs of subsequent multi-environment verification.
[0063] S4. Perform typing verification on candidate SNPs in a multi-environment validation population and identify several core SNP markers based on allele frequencies and genetic effects. Core SNP markers should have a minor allele frequency ≥ 0.05 and a two-sided test P value ≤ 1 × 10- -5 .
[0064] The specific implementation is as follows:
[0065] Verification environment: South-North dual-region verification.
[0066] Validation group:
[0067] South (a pond breeding base in Hunan): 100 adults.
[0068] North (a rice-shrimp breeding base in Jiangsu): 100 adults.
[0069] Phenotypic determination: The meat content was measured separately, with an average of 15.1% (±3.8%) in the south and 14.2% (±4.3%) in the north.
[0070] Typing filtration: Competitive allele-specific PCR typing was performed on 200 candidate SNPs in 200 individuals, and alleles with MAF < 0.05 or single-marker two-sided test P > 1 × 10 in any population were eliminated. -5 's location.
[0071] Filter Results:
[0072] There are 192 southern transit points and 187 northern transit points.
[0073] Core SNPs that meet the criteria in both regions: 185.
[0074] Statistical characteristics:
[0075] The average MAF of core SNPs was 0.21 (0.05-0.42).
[0076] Average single marker effect = 0.35% meat yield gain.
[0077] Cumulatively explained ≈62% of phenotypic variance.
[0078] S5. Develop a high-throughput typing platform based on core SNP markers. This platform is used to screen and select individuals with superior genotypes in the breeding population. These individuals are then bred in large numbers to produce high-meat-yield lines. This high-throughput typing platform utilizes competitive allele-specific PCR technology, achieving a single-locus typing call rate ≥95% and a typing accuracy ≥98%. Individuals with superior genotypes must carry at least 80% of the favorable alleles in the core SNP markers.
[0079] The specific implementation is as follows:
[0080] 1. Construction and performance evaluation of a competitive allele-specific PCR high-throughput typing platform
[0081] A total of 150 core SNP loci, validated by the four combined environments in step S4, were selected. Competitive allele-specific PCR primers were designed for these 150 loci and pilot experiments were conducted. 145 loci were successfully amplified in one run, for a design success rate of 96.7%. Allele-specific PCR was used instead for the remaining five loci due to high GC content.
[0082] The platform was tested on two 96-well plates using "92 samples + 4 negative / positive controls" - the average typing call rate for a single site was 97.3%, and the typing accuracy after sequencing review was 99.1%, meeting the requirements of "call rate ≥ 95% and accuracy ≥ 98%", and was better than the values reported in conventional literature.
[0083] 2. Batch typing of breeding populations
[0084] During the seedling stage, 1,200 two-week-old juvenile shrimp were randomly sampled and subjected to competitive allele-specific PCR reactions using 384-well plates plated 10 at a time. The experimental cost was 0.72 yuan per site per sample, for a total of approximately 130,000 yuan. A typing matrix was generated within 24 hours. 87,000 genotypes were effectively called, with a call rate of 96.7%. 1,000 randomly sampled genotypes were retested using next-generation sequencing, achieving an accuracy rate of 98.9%.
[0085] 3. Screening strategies for individuals with superior genotypes
[0086] First, a hard threshold of “favorable allele ratio ≥ 80%” was set, that is, each individual must carry at least 120 favorable alleles among the 150 loci.
[0087] To avoid false positives due to “sufficient numbers but missing key sites”, the top 20% of the effect values (30 sites in total) were weighted 1.5 times to calculate the weighted core index.
[0088] In addition, COANCESTRY software was used to detect the consanguinity coefficient and exclude individuals with sibling or half-sibling relationship coefficients exceeding 0.125 to ensure genetic diversity.
[0089] Among the 1,200 seedlings, 284 individuals (accounting for 23.7%) met the 80% threshold in the initial selection; 37 were eliminated by weighted index screening, and then 9 close relatives were eliminated, and finally 247 high-quality parents were obtained, with a male-to-female ratio of about 1:4.
[0090] 4. Parental combination and propagation process
[0091] Based on the closeness of genomic breeding values, body shape complementarity and inbreeding coefficient, the 247 excellent parents were divided into three types of combination methods:
[0092] High-affinity combination: select 1 male and 3 females, and control the difference in genomic breeding values within 5%. It is used for large-scale commercial production, with a total of 30 groups.
[0093] Complementary combination of traits: 1 male with 2 females, using body length 3 / body weight (R1), total length of abdominal segment / width of second abdominal segment (R2) are complementary, with a total of 25 groups. Excellent rapid amplification combination: select male shrimp with genomic breeding values in the top 5%, and pair one male with four females, for a total of 10 groups.
[0094] like Figure 2 As shown, R1 is positively correlated with meat content:
[0095] Meaning of the indicator: R1 reflects the "length advantage" of body length relative to body weight - the larger the value, the more "slender" the shrimp.
[0096] Morphological reasons: Under the same body length conditions, individuals with high R1 values lose more weight due to the relative thinning of the exoskeleton and internal organs, while the reduction in the mass of the skeleton and internal organs does not affect the longitudinal extension of muscle fibers. Due to the decrease in the proportion of the exoskeleton and internal organs and the relative density of muscle fibers, when R1 increases, the proportion of muscle weight per unit body length to the total weight is higher, and thus the meat content increases accordingly.
[0097] like Figure 3 As shown, R2 is positively correlated with meat content:
[0098] Meaning of the indicator: R2 reflects the degree of elongation of the abdominal segments of the tail - the larger the value, the more slender the tail.
[0099] Morphological reasons: A slender tail has more longitudinal muscle extension, increasing the proportion of muscle volume after stripping. Wider tail width, on the other hand, may be more reflected in shell thickness and non-muscle tissue. Therefore, as R2 increases, the proportion of muscle per unit abdominal segment volume increases, and the meat content also increases.
[0100] Differential complementation: Utilize the extreme or opposite performance of parents in different traits so that the next generation can take into account the advantages of two or more target traits at the same time, achieving the effect of trait balance or optimization. It includes the following steps:
[0101] 4.1 Calculation of index ratios
[0102] The body length 3 / body weight (R1) and total length of abdominal segment / width of the second abdominal segment (R2) were measured for each shrimp individual.
[0103] 4.2 Screening out extreme groups
[0104] Sort all individuals by R1 and R2 and take the top 20%.
[0105] 4.3 Paired complementary combinations
[0106] After each male shrimp is selected, it is paired with a female shrimp with high R1 or high R2.
[0107] 5. Complete the number of sets
[0108] Repeat the above rules until the required 25 complementary pairs are obtained.
[0109] After artificial insemination, the overall fertilization rate was 87.5%, with the superior combination reaching a peak of 92.1%. The average hatchability rate was 82.8%. Six weeks later, data from the F1 generation at 60 days of age showed an average body length of 5.9 cm, compared to only 4.7 cm for the control group (unselected). The average meat content of the F1 generation was 18.8% ± 2.6%, compared to 14.3% ± 3.1% for the control group. Genotyping of F1 individuals revealed an average content of favorable alleles of 83.2% with a standard deviation of 2.5%, demonstrating the stable transmission of the trait. Feed efficiency also improved—the F1 feed-to-meat ratio was 2.10:1, compared to 2.36:1 for the control group, resulting in an approximately 11% feed saving.
[0110] 6. Cycles and economic advantages
[0111] It takes 4 weeks from typing to parent screening and then to preliminary performance verification of F1 offspring; compared with traditional phenotypic breeding that relies on growth tail testing and then screening (at least 6 months), the time is shortened by more than 85%.
[0112] In terms of cost, the cost of typing a single tail is about 11 yuan, which is lower than the labor and feed costs of relying on long-term breeding and tail testing; calculated based on the promotion of every 10,000 tails in the first year, the net profit increase is about 186,000 yuan, while the traditional method is only 54,000 yuan, and the economic value-added is more than doubled.
[0113] 7. Sustainable optimization plan
[0114] In the F1 generation, 100 loci will continue to be sampled for rolling monitoring to maintain the proportion of favorable alleles above 80%.
[0115] The SNP weights are adjusted every year based on the newly measured effect values, and the genomic breeding value model is updated.
[0116] About 10% of individuals with large genotype differences are retained to build a genetic safety pool to prevent the population bottleneck effect.
[0117] Example 2
[0118] The difference between this embodiment and the first embodiment is that the verification environment of this embodiment is a spring-summer-autumn seasonal verification.
[0119] Validation group:
[0120] Spring (open pond): 80.
[0121] Summer (greenhouse pond): 80.
[0122] Autumn (open pond): 80.
[0123] Phenotypic determination: The average meat content in the three seasons was 15.2% (±3.5%), 14.2% (±3.9%), and 14.8% (±4.1%), respectively.
[0124] Typing Filtering: 200 candidate SNPs were typed in 240 individuals, and only those with MAF ≥ 0.05 and P ≤ 1 × 10 were retained in all three seasons. -5 's location.
[0125] Filter Results:
[0126] The number of passing sites in spring, summer and autumn were 178, 184 and 176 respectively.
[0127] Three-quarter common core SNP: 172.
[0128] Statistical characteristics:
[0129] Average MAF = 0.19 (0.06 ~ 0.39).
[0130] Average single marker effect = 0.32%.
[0131] Cumulative explained variance ≈58%.
[0132] Example 3
[0133] The difference between this embodiment and the first embodiment is that the verification environment of this embodiment is family-environment joint verification.
[0134] Validation group:
[0135] Two full-sib families (60 each).
[0136] Commercial farm population (80).
[0137] Phenotypic determination: The average meat content within the family was 15.7% (±2.9%) and 14.2% (±3.2%), and the average meat content in the farm group was 15.1% (±3.7%).
[0138] Subtype filtering: a more stringent P value threshold of P≤5×10 -6 , and the MAF of the three groups ≥ 0.05.
[0139] Filter Results:
[0140] Family 1, family 2 and the farm passed through 160, 158 and 162 loci, respectively.
[0141] Common core SNPs among the three groups: 150.
[0142] Statistical characteristics:
[0143] Average MAF = 0.18 (0.05 ~ 0.36).
[0144] Average single marker effect = 0.38%.
[0145] Cumulative explained variance ≈55%.
[0146] Example 4
[0147] The difference between this embodiment and the first embodiment is that the verification environment of this embodiment is a breeding density comparison verification.
[0148] Validation group:
[0149] Low density (5 / m 2 ): 75 pieces.
[0150] Medium density (10 / m 2 ): 75 pieces.
[0151] High density (15 / m 2 ): 75 pieces.
[0152] Phenotypic determination: The average meat content of low, medium and high density meat was 15.4%, 14.5% and 14.1% respectively.
[0153] Typing Filtering: 200 candidate SNPs were typed in 225 samples. All three densities required MAF ≥ 0.05 and P ≤ 1 × 10 -6 .
[0154] Filter Results:
[0155] Low, medium, and high densities passed through 168, 162, and 165 sites, respectively.
[0156] Triple-density common core SNPs: 160.
[0157] Statistical characteristics:
[0158] Average MAF = 0.20 (0.05 ~ 0.38).
[0159] Average single marker effect = 0.37%.
[0160] Cumulative explained variance ≈60%.
[0161] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.
Claims
1. A method for breeding a high-meat content strain of Procambarus clarkii based on SNP markers, characterized in that: include: S1. DNA was extracted from Procambarus clarkii individuals with significant differences in meat content phenotype using a silica gel column method. S2. performing genome-wide SNP typing on the DNA using high-throughput sequencing; S3. Constructing a training dataset based on the SNP typing results, and using the XGBoost model combined with the SHAP value to evaluate and screen candidate SNP sites significantly associated with high meat content through the training dataset; S4. Typing and validating the candidate SNP sites in a multi-environment validation population, and determining several core SNP markers based on allele frequencies and genetic effects; S5. A high-throughput typing platform is constructed based on the core SNP markers, and the breeding population is typed and screened by the high-throughput typing platform to retain individuals with excellent genotypes, and the individuals with excellent genotypes are bred in large quantities to obtain high meat content lines.
2. The method for breeding a high-meat content strain of Procambarus clarkii based on SNP markers according to claim 1, characterized in that: The ultraviolet absorbance ratio range of DNA extracted by the silica gel membrane column method is 1.8-2.0, and the concentration is not less than 50 ng / μL.
3. The method for breeding a high-meat content strain of Procambarus clarkii based on SNP markers according to claim 1, characterized in that: The high-throughput sequencing has a sequencing depth of no less than 10× and a coverage of ≥95%.
4. The method for breeding a high-meat content crayfish strain based on SNP markers according to claim 1, characterized in that: The XGBoost model is specifically: Among them, y i is the observed meat content phenotypic value of the i-th sample, is the predicted value of the i-th sample in the t-1th iteration, f k is the decision function corresponding to the kth regression tree. The input is the feature vector x, and the output is the predicted increment of the target value by the tree. f t (x i ) is the t-th regression tree for sample x i The incremental prediction function, l(·,·) is a differentiable convex loss function, Ω(f) is the regularization term of the tree model, where: γ is the regularization coefficient for adjusting the number of leaf nodes T, T is the total number of leaf nodes in the regression tree, λ is the L2 regularization coefficient for the leaf weight, w j is the output weight of the jth leaf node.
5. The method for breeding a high-meat content strain of Procambarus clarkii based on SNP markers according to claim 1, characterized in that: The SHAP value is derived from the following model formula: Among them, F is the total feature set, |F| is the total number of features, S is any feature subset that does not contain the jth feature, |S| is the number of features in the subset S, and f S (x S ) represents the model output when only feature subset S is used for prediction, f S∪{j} (x) is the expected output of the model when only the feature subset S∪{j} is used for prediction, X S∪{j} It means that only the sample feature vectors of the latitude corresponding to the feature subset S∪{j} are retained.
6. The method for breeding a high-meat content strain of Procambarus clarkii based on SNP markers according to claim 1, characterized in that: The core SNP markers have a minor allele frequency ≥ 0.05 in the multi-environment validation population and a two-sided test P value corresponding to the single marker genetic effect ≤ 1×10 -5 .
7. The method for breeding a high-meat content strain of Procambarus clarkii based on SNP markers according to claim 1, characterized in that: The high-throughput typing platform adopts competitive allele-specific PCR technology, and the single-site typing call rate of the high-throughput typing platform is ≥95%, and the typing accuracy rate is ≥98%.
8. The method for breeding a high-meat content strain of Procambarus clarkii based on SNP markers according to claim 1, characterized in that: The favorable allele content in the core SNP marker carried by the superior genotype individual is ≥80%.