Alfalfa yield trait optimal prediction system and construction method and application thereof

By constructing an alfalfa yield trait prediction system based on whole-genome selection, and utilizing Ridge, GBLUP, or BayesA_Ridge models and a set of 5000 SNP loci, the problem of early prediction of yield traits in alfalfa breeding was solved, achieving efficient and accurate breeding prediction and shortening the breeding cycle.

CN121583323APending Publication Date: 2026-02-27INSTITUTE OF GRASSLAND RESEARCH OF CAAS
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511766555.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-27
Publication Date
2026-02-27

AI Technical Summary

Technical Problem

Existing technologies are insufficient for efficient and accurate early prediction and selection of alfalfa yield traits, resulting in long breeding cycles and low efficiency. This is especially true for alfalfa, a perennial herbaceous plant, where traditional breeding methods rely on phenotypic selection, which is inefficient, and molecular marker-assisted selection is difficult to apply given the complexity of tetraploid genomes.

Method used

Using the whole-genome selection statistical model Ridge, GBLUP, or BayesA_Ridge, combined with a set of 5000 high-quality SNP genotype loci, we constructed an optimal prediction system for alfalfa yield traits through genome-wide association analysis and 5-fold cross-validation. We used high-throughput sequencing and genotyping technologies to obtain high-quality SNP loci, established multiple prediction models, and determined the optimal model and locus set.

Benefits of technology

It enables rapid, efficient and accurate prediction of alfalfa yield traits with a prediction accuracy of up to 0.99, allowing for early prediction of yield traits, shortening the breeding cycle and increasing the intensity of breeding.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121583323A_ABST
    Figure CN121583323A_ABST
Patent Text Reader

Abstract

The invention discloses an optimal alfalfa yield trait prediction system as well as a construction method and application thereof. According to the invention, based on whole genome association analysis, SNP loci significantly related to alfalfa yield traits are identified, and on the basis, the influences of different statistical models and different numbers of SNP markers on whole genome selection prediction precision are compared; and an optimal whole genome selection prediction system based on a machine learning model and alfalfa yield trait significant association sites is established. The system has the capability of quickly, efficiently and accurately predicting the excellent yield of alfalfa, and the prediction accuracy is up to 0.99. The method provided by the invention can realize early prediction of alfalfa yield traits, can effectively shorten the alfalfa breeding period, improves the alfalfa breeding intensity, and accelerates the breeding process of excellent germplasm.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of alfalfa biological breeding, and relates to an optimal prediction system for alfalfa yield traits, its construction method, and its application. Background Technology

[0002] Alfalfa (*Medicago sativa*) is a perennial herbaceous plant belonging to the legume family, widely cultivated in temperate and cold-temperate regions, and is one of the world's most important forage and green manure crops. Its leaves are rich in protein, minerals, and vitamins, possessing high nutritional value and good palatability, making it an important source of high-quality feed for ruminants. Furthermore, alfalfa has a well-developed root system, which functions in nitrogen fixation, soil structure improvement, and ecological restoration. Due to its strong adaptability, long growth cycle, and high yield potential, alfalfa has become an indispensable part of the global agricultural system.

[0003] Alfalfa plays a crucial role not only in animal husbandry but is also widely used in ecological agriculture, soil and water conservation, and carbon sequestration. As a high-quality protein feed, the yield and quality of alfalfa hay directly affect livestock production performance and breeding efficiency. Particularly in dairy farming, the use of alfalfa silage can significantly improve milk yield and quality. Furthermore, alfalfa planting area and yield have become important indicators for measuring the level of regional agricultural sustainable development. However, due to alfalfa's long growth cycle, complex reproduction methods, and difficulty in genetic improvement, the genetic improvement of its yield traits has always been a key focus and challenge in breeding research.

[0004] Alfalfa yield traits (such as fresh weight, dry matter content, and seed yield) are key factors determining its forage and economic value. However, alfalfa has a long growth cycle and a complex genetic background (tetraploid, increasing the likelihood of key genetic traits), resulting in a long breeding cycle. Its yield traits are complex quantitative traits controlled by multiple genes, and their heritability is often influenced by gene-gene interactions and gene-environment interactions, posing a challenge to the efficient breeding of high-yielding alfalfa varieties. With the development of molecular biology, marker-assisted selection (MAS) and genome-wide association studies (GWAS) have been applied to alfalfa research. However, MAS struggles to simultaneously handle traits controlled by multiple genes with minor effects; while it can identify trait-related loci, its application in alfalfa is limited by the complexity of its tetraploid genome, making it difficult to systematically predict and select high-yielding individuals for early precision. Traditional breeding methods still largely rely on phenotypic selection, which is inefficient and time-consuming.

[0005] Genomic selection (GS) is a breeding strategy based on whole-genome molecular markers. By integrating individual genomic information, it predicts breeding values ​​and performs selection, thereby improving the selection efficiency of complex traits. This technology has achieved significant results in food crops such as wheat, rice, and maize, but its application in perennial herbaceous plants such as alfalfa is still in its early stages. GS technology can overcome the problems of long breeding cycles and low selection efficiency in traditional breeding methods and is expected to play an important role in the genetic improvement of yield traits in alfalfa. However, due to the complexity of the alfalfa genome and the diversity of its genetic background, how to construct a GS model suitable for its yield traits remains a key research focus and challenge. Summary of the Invention

[0006] To address the shortcomings of existing technologies in the efficient prediction of alfalfa yield traits, the first objective of this invention is to provide an optimal prediction system for alfalfa yield traits.

[0007] The second objective of this invention is to provide a method for constructing an optimal prediction system for alfalfa yield traits.

[0008] The third objective of this invention is to provide the application of the optimal prediction system for alfalfa yield traits in the breeding of superior alfalfa germplasm.

[0009] The first objective of this invention is achieved by the following technical solution: an optimal prediction system for alfalfa yield traits, wherein the whole genome selection statistical model is any one of the Ridge model, GBLUP model or BayesA_Ridge model; and the SNP genotype locus set is a set of 5000 SNP genotype loci.

[0010] The second objective of this invention is achieved by the following technical solution: a method for constructing an optimal prediction system for alfalfa yield traits, comprising the following steps: (1) The genomes of 300 alfalfa plants were resequencing and genotyping, and 3,425,009 high-quality SNP loci were obtained; (2) Based on the 3,425,009 high-quality SNP loci obtained, combined with alfalfa yield phenotypic data, genome-wide association analysis was performed. The top 200, 300, 500, 1,000, 2,000, 3,000, 4,000 and 5,000 SNPs were extracted according to the P value from small to large, forming 8 SNP genotype locus sets; (3) The alfalfa population was divided into a training set and a test set by using the 5-fold cross-validation method. In the training set, multiple alfalfa yield phenotypic data, 14 whole-genome selection statistical models and 8 different SNP genotype locus sets were used to establish alfalfa yield phenotypic whole-genome selection prediction models. (4) Using the average Pearson correlation coefficient between the breeding value and the measured value of the test population as the evaluation index, the optimal whole genome selection statistical model and the optimal SNP genotype locus set were determined after 100 repeated iterations, thereby constructing the optimal prediction system for alfalfa yield traits.

[0011] Furthermore, the specific method of step (1) includes: using WGS technology, performing paired-end PE150 sequencing on 300 alfalfa plants using the Illumina HiSeq6000 high-throughput sequencing platform; and using the BWA tool to align the sequencing data to Medicago. Alignment results in BAM format were obtained from the reference genome figshare_version_3 of sativa. To improve the accuracy of subsequent variant detection, the alignment results were preprocessed, including PCR repetitive sequence removal, quality control, local re-alignment, and base quality value correction. Variance detection was performed using the HaplotypeCaller tool in GATK, followed by single nucleotide variant and insertion / deletion detection using PLINK and VCFtools software. The variant results were first initially filtered based on quality and depth indicators using the VariantFiltration tool in GATK to remove false positives and spurious variants. Then, genotypes were rigorously filtered using PLINK and VCFtools software, with filtering criteria including integrity greater than 0.8, minimum allele frequency not less than 0.05, deletion rate less than 20%, compliance with the Hardy-Weinberg equilibrium law, and LD 50kb 0.2. These SNP sites were annotated and functionally predicted using SnpEff software. Finally, 3,425,009 high-quality SNP sites were obtained.

[0012] Further, genomic selection analysis was performed based on the 3,425,009 high-quality SNP loci and alfalfa yield phenotypic data obtained in step (1). In the data preprocessing stage, samples with missing phenotypes were first removed, and then outliers were identified and removed using the interquartile range (IQR) method. The SNP effect was estimated based on the mixed linear model `mixed.solve` in the `rrBLUP` package, and key SNPs that significantly contribute to yield traits were screened using a multi-iteration weighted strategy. In each iteration, 80% of the samples were randomly selected as the training set, and the remaining 20% ​​as the test set. Genotype missing values ​​were filled using `rrBLUP::A.mat`, and the SNP effect value was estimated using the `mixed.solve()` function. A mixed linear model was constructed on the training set data, and the absolute effect value of each SNP was extracted as its contribution measure. SNPs were sorted in descending order based on their absolute effect values. The following sets of SNP genotype loci were selected from the top 200, top 300, top 500, top 1000, top 2000, top 3000, top 4000, and top 5000 SNP genotype loci ranked by P-value from smallest to largest in genome-wide association analysis (GWAS). These sets constituted eight SNP genotype loci sets.

[0013] Furthermore, the 14 genome-wide selection statistical models are: Bayesian models: BayesA_Ridge, BayesB_Lasso, BayesC_ElasticNet, BayesLasso_Lasso; linear regression models: Ridge, LinearLasso, ElasticNet, LinearRegression; machine learning models: KernelRidge, PLSRegression, RandomForest, SVRlinear, SVRpoly; and the genome selection-specific model GBLUP.

[0014] Furthermore, the alfalfa yield phenotypic data are fresh weight, dry weight, and greening time.

[0015] Furthermore, the optimal genome-wide selection statistical model is Ridge, GBLUP, or BayesA_Ridge.

[0016] Furthermore, the optimal SNP genotype locus set is a set of 5000 SNP genotype loci.

[0017] The third objective of this invention is achieved by the following technical solution: the application of an optimal prediction system for alfalfa yield traits in the process of selecting superior alfalfa germplasm.

[0018] Advantages of this invention: (1) The optimal prediction system for alfalfa yield traits of the present invention has the ability to predict alfalfa yield quickly, efficiently and accurately, with a prediction accuracy of up to 0.99.

[0019] (2) The present invention can realize early prediction of alfalfa yield traits, effectively shorten the alfalfa breeding cycle, increase the intensity of alfalfa breeding, and accelerate the breeding process of superior germplasm. Attached image description: To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0020] Figure 1 This is a map showing the SNP density distribution across the entire alfalfa genome. Figure 2 SNP functional annotation classification (top) and genomic location distribution (bottom); Figure 3 Box plots comparing the prediction accuracy of different models for different SNP number locus sets; Figure 4 Heatmaps showing the prediction accuracy of each model for different SNP number locus sets; Figure 5 This is a line graph showing how the model's prediction accuracy changes with the size of the SNP genotype locus set. Detailed implementation method: The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0021] Unless otherwise specified, the technical means used in the following embodiments are all conventional means well known to those skilled in the art.

[0022] Thirty alfalfa plants from the alfalfa germplasm resource bank were selected as experimental materials. These alfalfa plants originated from different regions and were not directly related. A randomized block design was used for planting. A folded alfalfa yield measurement quadrat was used for quadrat selection. This quadrat consisted of two detachably connected single-sided folded frames. Each single-sided folded frame consisted of a first rod and a second rod. One end of the first rod and one end of the second rod were connected by a hinge. A connecting device was installed at the other end of the first rod. When the two single-sided folded frames were connected, the connecting device was detachably connected to the end of the second rod on the other side. In the alfalfa field, the alfalfa plants were separated from each other, and yield was measured on a per-plant basis. A belt-mounted manual mower was used for mowing. This mower included a body containing an engine. Batteries were installed on both sides of the engine. A grass collection device was installed at the rear of the body. The engine shaft was connected to a cutter head via a belt. The cutter head was located at the lower front of the body, and the stubble height could be adjusted by adjusting the height of the cutter head. The cutter head is a circular toothed cutter head with a diameter of 30cm, a tooth height of 2cm, and a tooth spacing of approximately 3cm. Harvesting is carried out from the initial flowering stage (when 15% of the total number of flowering plants per unit area) to the full flowering stage (when 80% of the total number of flowering plants per unit area). During harvesting, the height of the cutter head of the harvesting device is adjusted to a suitable stubble height, generally 4-6cm, which can be adjusted appropriately according to different regions and growing environments.

[0023] Example 1: 1. Yield Measurement Fresh weight determination: After cutting a single alfalfa plant, place it directly on an electronic scale with an accuracy of 0.1g and weigh the fresh weight of the single alfalfa plant. Record the data as W1.

[0024] Dry weight determination: From the fresh alfalfa hay after weighing, a single fresh hay sample was randomly selected, put into a mesh bag and weighed as W2 (1000g±10g), and the weight of the mesh bag was M1; then it was placed in an oven, blanched at 105℃ for 30min, and then dried at 65℃ to constant weight, which was recorded as W3.

[0025] Calculation of the dry-to-fresh ratio: Dry-to-fresh ratio T = (W3-M1) ÷ (W2-M1) × 100%.

[0026] Yield per plant: F = W1 × T.

[0027] 2. Phenotypic determination and analysis Statistical analysis of the phenotypic data (fresh weight, dry weight, and greening time) was performed using R 4.4.1 software, including calculation of the mean, minimum, maximum, standard deviation, and coefficient of variation. The skewness and kurtosis of the dataset were calculated using the moments package in R 4.4.1 software, and a normality test was performed.

[0028] Example 2: 1. Genome resequencing and genotyping We used WGS technology to perform paired-end PE150 sequencing on 300 alfalfa plants using the Illumina HiSeq 6000 high-throughput sequencing platform. The sequencing data were then aligned to the Medicago sativa reference genome (figshare_version_3) using the BWA tool to obtain alignment results in BAM format. To improve the accuracy of subsequent variant detection, the comparison results were preprocessed, including PCR repetitive sequence removal, quality control, local re-alignment, and base quality value correction. Variance detection was performed using the HaplotypeCaller tool in GATK, followed by single nucleotide variant and insertion / deletion detection using PLINK and VCFtools software. The variant results were first initially filtered using the VariantFiltration tool in GATK based on quality and depth indicators to remove false positives and spurious variants. Next, genotypes were rigorously filtered using PLINK and VCFtools software, with filtering criteria including integrity greater than 0.8, minimum allele frequency not less than 0.05, deletion rate less than 20%, compliance with the Hardy-Weinberg equilibrium, and an LD 50kb of 0.2. SnpEff software was used to annotate and predict the function of these SNP sites. Finally, 3,425,009 high-quality SNP sites were obtained for subsequent genetic analysis.

[0029] 2. SNP site detection and annotation The R language was used to display the density distribution of all loci at their locations on the chromosome, such as... Figure 1 As shown in the chromosome location density distribution map, the distribution of SNPs on chromosomes can be seen, with several concentrated regions on chromosomes 3 and 6. After detecting variant sites, the SnpEff software was used to annotate the variant sites. It annotates each SNP based on its specific location in the reference genome and the related gene information, such as... Figure 2 As shown.

[0030] 3. Genomic selection analysis of alfalfa yield traits Genomic selection analysis was performed based on the 3,425,009 high-quality SNP loci and alfalfa phenotypic data obtained above. SNP effects were estimated using a mixed linear model (mixed.solve) from the rrBLUP package, and key SNPs significantly contributing to yield traits were screened using a multi-iteration weighted strategy. In each iteration, 80% of the sample was randomly selected as the training set, and the remaining 20% ​​as the test set. Genotype missing values ​​were imputed using rrBLUP::A.mat, and SNP effect values ​​were estimated using the mixed.solve() function. A mixed linear model was constructed on the training set data, and the absolute effect value of each SNP was extracted as its contribution measure. SNPs were sorted in descending order based on their absolute effect values. Eight SNP genotype locus sets were selected from the following groups: the top 200 SNP genotype locus sets ranked by P-value in genome-wide association analysis (GWAS); the top 300 SNP genotype locus sets ranked by P-value in GWAS; the top 500 SNP genotype locus sets ranked by P-value in GWAS; the top 1000 SNP genotype locus sets ranked by P-value in GWAS; the top 2000 SNP genotype locus sets ranked by P-value in GWAS; the top 3000 SNP genotype locus sets ranked by P-value in GWAS; the top 4000 SNP genotype locus sets ranked by P-value in GWAS; and the top 5000 SNP genotype locus sets ranked by P-value in GWAS. These sets were used as different SNP genotype datasets for subsequent predictive analyses.

[0031] Example 3: 1. Organizing phenotypic and genotypic files Before performing genomic selection (GS) prediction, missing value imputation and formatting were performed on the phenotypic and genotypic data of the training population. Fresh weight, dry weight, and greening time of 300 alfalfa plants were used as phenotypic data for genomic selection prediction of alfalfa yield traits. PLINK software was used to convert the genotypic data to a uniform 0 / 1 / 2 format, where homozygous non-mutant genotypes were coded as 0, heterozygous genotypes as 1, homozygous mutant genotypes as 2, and -1 represented missing values.

[0032] 2. Parameter settings for the whole-genome selection model Fourteen different genome-wide selection statistical models were employed, including Bayesian models BayesA_Ridge, BayesB_Lasso, BayesC_ElasticNet, BayesLasso_Lasso, linear regression models Ridge, LinearLasso, ElasticNet, and LinearRegression, machine learning models KernelRidge, PLSRegression, RandomForest, SVRlinear, and SVRpoly, and the genome selection-specific model GBLUP. For each genome-wide selection statistical model, eight different sets of SNP genotype loci were defined to evaluate their impact on GS prediction accuracy. These eight different SNP genotype locus sets were derived from the following sets: the top 200 SNP genotype locus sets ranked by P-value in genome-wide association analysis (GWAS); the top 300 SNP genotype locus sets ranked by P-value in GWAS; the top 500 SNP genotype locus sets ranked by P-value in GWAS; the top 1000 SNP genotype locus sets ranked by P-value in GWAS; the top 2000 SNP genotype locus sets ranked by P-value in GWAS; the top 3000 SNP genotype locus sets ranked by P-value in GWAS; the top 4000 SNP genotype locus sets ranked by P-value in GWAS; and the top 5000 SNP genotype locus sets ranked by P-value in GWAS. These eight sets of SNP genotype locus sets served as different SNP genotype datasets for subsequent predictive analyses.

[0033] 3. Determine the optimal model for alfalfa yield traits through genome-wide selection. Using a 5-fold cross-validation method, 80% of the alfalfa population was used as the training population, and the remaining 20% ​​as the test population. In the training population, a genome-wide selection prediction model for alfalfa yield traits was established using phenotypic data of yield traits, 14 genome-wide selection statistical models, and data from 8 different SNP genotype locus sets. Subsequently, breeding values ​​for alfalfa yield traits were estimated in the validation population using genotype data and the prediction model. To eliminate sampling errors, this process was repeated 100 times, and the mean Pearson correlation coefficient (r) between the breeding values ​​of the test population and the actual observed values ​​was used as an indicator to evaluate the accuracy of genome-wide selection prediction. Finally, the optimal genome-wide selection statistical model and the optimal SNP genotype locus set were determined using this evaluation criterion. These two optimal models constituted the optimal prediction system for genome-wide selection of alfalfa yield traits. Genome-wide selection prediction analysis of alfalfa yield traits was performed using 14 different genome-wide selection statistical models and 8 different SNP genotype locus sets. The results are as follows: Figure 3, Figure 4 and Figure 5 As shown, the machine learning-based model exhibits significantly improved prediction accuracy (highest prediction accuracy r = 0.99), demonstrating a clear advantage. In the genome-wide selection prediction model, utilizing trait-related loci obtained from genome-wide association analysis (GWAS) as the genotype dataset can greatly improve the prediction accuracy of GWAS. As the marker density gradually increases from 200 to 5000, the Ridge, GBLUP, and BayesA_Ridge machine learning models show the best GS prediction model characteristics, with their prediction accuracy gradually increasing from 0.65 to 0.99. After the number of SNPs reaches 5000, the prediction accuracy tends to plateau, and the Ridge, GBLUP, and BayesA_Ridge models achieve the highest prediction accuracy, reaching 0.99. In conclusion, the Ridge, GBLUP, and BayesA_Ridge models, along with a genotype locus set of 5000 SNPs, are selected as the optimal prediction system for genome-wide selection of alfalfa yield traits.

[0034] It should be noted that the above embodiments are intended to more clearly illustrate the technical solution of the present invention, rather than to limit its scope of protection. Equivalent modifications or variations made by those skilled in the art without inventive effort should be considered to fall within the scope of protection of the present invention.

Claims

1. An optimal prediction system for alfalfa yield traits, characterized in that, The genome-wide selection statistical model is any one of the Ridge model, GBLUP model, or BayesA_Ridge model; the SNP genotype locus set is a set of 5000 SNP genotype loci.

2. The method for constructing an optimal prediction system for alfalfa yield traits as described in claim 1, characterized in that: It includes the following steps: (1) The genomes of 300 alfalfa plants were resequencing and genotyping, and 3,425,009 high-quality SNP loci were obtained; (2) Based on the 3,425,009 high-quality SNP loci obtained, combined with alfalfa yield phenotypic data, genome-wide association analysis was performed. The top 200, 300, 500, 1,000, 2,000, 3,000, 4,000 and 5,000 SNPs were extracted according to the P value from small to large, forming 8 SNP genotype locus sets; (3) The alfalfa population was divided into a training set and a test set by using the 5-fold cross-validation method. In the training set, multiple alfalfa yield phenotypic data, 14 whole-genome selection statistical models and 8 different SNP genotype locus sets were used to establish alfalfa yield phenotypic whole-genome selection prediction models. (4) Using the average Pearson correlation coefficient between the breeding value and the measured value of the test population as the evaluation index, the optimal whole genome selection statistical model and the optimal SNP genotype locus set were determined after 100 repeated iterations, thereby constructing the optimal prediction system for alfalfa yield traits.

3. The construction method according to claim 2, characterized in that: The specific method of step (1) includes: using WGS technology, performing paired-end PE150 sequencing on 300 alfalfa plants using the Illumina HiSeq6000 high-throughput sequencing platform; using the BWA tool to align the sequencing data to the reference genome figshare_version_3 of Medicago sativa, obtaining alignment results in BAM format; to improve the accuracy of subsequent variant detection, the alignment results are preprocessed, including removing PCR repetitive sequences, quality control, local re-alignment, and base quality value correction; variant detection is performed using the HaplotypeCaller tool in GATK, followed by detection of single nucleotide variants and insertions / deletions using PLINK and VCFtools software; the variant results are first initially filtered based on quality and depth indicators using the VariantFiltration tool in GATK to remove false positives and pseudo-variants; then, the genotypes are strictly filtered using PLINK and VCFtools software, where the filtering criteria include integrity greater than 0.8, minimum allele frequency not less than 0.05, deletion rate less than 20%, and conformity to the Hardy-Weinberg equilibrium law, LD 50kb 0.2; These SNP sites were annotated and functionally predicted using SnpEff software; a total of 3,425,009 high-quality SNP sites were obtained.

4. The construction method according to claim 2, characterized in that: The specific method of step (2) is as follows: genomic selection analysis is performed based on the 3,425,009 high-quality SNP loci and alfalfa yield phenotypic data obtained in step (1); in the data preprocessing stage, samples with missing phenotypes are first removed, and then outliers are identified and removed using the interquartile range (IQR) method. The SNP effect is estimated based on the mixed linear model mixed.solve in the rrBLUP package, and key SNPs that significantly contribute to yield traits are screened through a multi-iteration weighted strategy; in each iteration, 80% of the samples are randomly selected as the training set, and the remaining 20% ​​are used as the test set; genotype missing values ​​are filled using rrBLUP::A.mat, and the SNP effect value is estimated using the mixed.solve() function; a mixed linear model is constructed on the training set data, and the absolute effect value of each SNP is extracted as its contribution. SNPs were ranked in descending order based on their absolute effect values. The following sets of SNP genotype loci were selected from the top 200, top 300, top 500, top 1000, top 2000, top 3000, top 4000, and top 5000 SNP genotype loci ranked by P-value from smallest to largest in the genome-wide association analysis (GWAS) analysis: [List of sets would be inserted here].

5. The construction method according to claim 2, characterized in that: The 14 genome-wide selection statistical models are: Bayesian models: BayesA_Ridge, BayesB_Lasso, BayesC_ElasticNet, BayesLasso_Lasso; Linear regression models: Ridge, LinearLasso, ElasticNet, LinearRegression; Machine learning models: KernelRidge, PLSRegression, RandomForest, SVRlinear, SVRpoly; and the genome selection-specific model GBLUP.

6. The construction method according to claim 2, characterized in that: The alfalfa yield phenotypic data are fresh weight, dry weight, and greening time.

7. The construction method according to claim 2, characterized in that: The optimal genome-wide selection statistical model is Ridge, GBLUP, or BayesA_Ridge.

8. The construction method according to claim 2, characterized in that: The optimal SNP genotype locus set is a set of 5000 SNP genotype loci.

9. The application of the optimal prediction system for alfalfa yield traits as described in claim 1 in the process of breeding superior alfalfa germplasm.