A genomic phenotypic value prediction method for pig litter traits based on machine learning

Through the genomic phenotype value prediction method of pig birth traits based on machine learning, the problem of low accuracy of early genetic evaluation caused by the small local pig breeding population is solved, and the prediction accuracy of breeding traits is improved.

CN119049549BActive Publication Date: 2025-06-24CHONGQING ACAD OF ANIMAL SCI +1
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202411116021.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-14
Publication Date
2025-06-24
Estimated Expiration
2044-08-14

AI Technical Summary

Technical Problem

The early estimated breeding values ​​and later actual phenotypes are different from those of the late stage due to the small size of local pig breeding population and the low accuracy of early genetic assessment.

Method used

The genomic phenotype value prediction method for pig birth traits based on machine learning was used to predict the genomic phenotype value of traits by collecting sows’ reproductive trait data, performing genomic DNA extraction and SNP genotyping, genome-wide association analysis and machine learning model construction.

Benefits of technology

The prediction accuracy of local pig breeding traits is improved, and the difference between early estimates and later actual phenotypes is reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119049549B_ABST
    Figure CN119049549B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for predicting genomic phenotypic values of pig litter traits based on machine learning, belonging to the field of predicting production traits of local pig breeds. The method includes determining local pig breeds, collecting reproductive traits of sows to obtain an original phenotypic data set; correcting the original phenotypic data set and removing abnormal phenotypes to obtain a target phenotypic data set; extracting genomic DNA of sows corresponding to each target phenotypic data in the target phenotypic data set, performing GWAS (Genome-Wide Association Study) genome-wide association analysis to obtain an SNP training data set; obtaining a phenotypic sample set based on the SNP training data set and the target phenotypic data set, constructing a reproductive trait prediction model, and using the reproductive trait prediction model according to the genotype data of the sows to be tested to obtain the genomic phenotypic prediction values of litter traits. The present invention solves the problem of large differences between the early estimated breeding values and the later actual phenotypes caused by the small scale of the local pig breeding population and the low accuracy of early genetic evaluation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of predicting production traits of local breeding pigs, and particularly relates to a method for predicting genomic phenotypic values of pig litter traits based on machine learning. Background Art

[0002] With the development of the livestock industry, the technology for predicting pig production traits has increasingly become an important means to improve breeding efficiency and optimize breed selection. These technologies not only rely on traditional genetic knowledge but also incorporate the latest achievements in modern biotechnology, big data analysis, artificial intelligence, and other fields.

[0003] Predicting the reproductive traits of local sows is a research field that intersects multiple disciplines such as genetics, nutrition, environmental science, and information technology. With the rapid development of modern biotechnology, especially the rise of molecular genetics, genomics, and bioinformatics, it provides strong technical support for the accurate prediction of local sow reproductive traits.

[0004] Genetic Basis:

[0005] Gene Markers and Linkage Analysis: Using molecular marker technologies (such as SNPs, SSRs, etc.) to locate and perform linkage analysis on the genetic genes of sows, revealing the association between reproductive traits and genetic genes.

[0006] Genome-Wide Association Study (GWAS): Through large-scale genome sequencing, genetic variant sites significantly associated with reproductive traits are discovered, providing a genetic basis for trait prediction.

[0007] Nutrition and Environmental Science:

[0008] Nutritional Regulation: Studying the effects of different nutritional components on the reproductive performance of sows, and improving the reproductive efficiency and production performance of sows by optimizing feed formulations and feeding management.

[0009] Environmental Stress Management: Analyzing the effects of environmental factors (such as temperature, humidity, light, etc.) on the reproductive performance of sows, proposing corresponding environmental control measures, reducing stress responses, and improving reproductive success rates.

[0010] Information Technology:

[0011] Big Data Analysis: Collecting multi-source information such as the reproductive performance data, genetic information, nutrition, and environmental parameters of sows, and using big data technology for in-depth mining and analysis to establish a reproductive trait prediction model.

[0012] Artificial Intelligence and Machine Learning: Using artificial intelligence algorithms (such as neural networks, decision trees, etc.) to predict reproductive traits, improving the accuracy and efficiency of prediction.

[0013] The application prospect of the prediction method for the reproductive traits of local sows is broad. On the one hand, it can provide a scientific basis for breeding farms to select and breed pigs, improving the production performance and reproductive efficiency of breeding pigs; on the other hand, it can provide personalized feeding management plans for pig farms, optimizing feed formulations and feeding environments, reducing production costs, and increasing economic benefits.

[0014] With the continuous development of biotechnology and information technology, the prediction method for the reproductive traits of local sows will be more accurate and efficient. In the future, through the cross-integration and collaborative innovation of multiple disciplines, it is expected to promote the local sow breeding industry to a higher level. Summary of the Invention

[0015] Aiming at the above deficiencies in the prior art, the genomic phenotypic value prediction method for pig litter traits based on machine learning provided by the present invention solves the problem of large differences between the early estimated breeding values and the later actual phenotypes caused by the small scale of the local pig breeding population and the low accuracy of early genetic evaluation.

[0016] In order to achieve the above invention purpose, the technical solution adopted by the present invention is: a genomic phenotypic value prediction method for pig litter traits based on machine learning, including the following steps:

[0017] S1. Determine the local pig breed, and based on the local pig breed, collect the reproductive traits of a number of sows to obtain an original phenotypic data set; the reproductive traits include litter weight at birth, litter size, and number of live piglets;

[0018] S2. Use a linear model to correct the original phenotypic data set and remove abnormal phenotypes to obtain a target phenotypic data set;

[0019] S3. Extract the genomic DNA of the sows corresponding to each target phenotypic data in the target phenotypic data set;

[0020] S4. Perform SNP genotyping on each genomic DNA to obtain a genotyping data set;

[0021] S5. Perform genotype imputation on the genotyping data set to obtain an SNP data set;

[0022] S6. Perform GWAS genome-wide association analysis on the SNP data set to obtain an SNP training data set;

[0023] S7. Based on the SNP training data set and the target phenotypic data set, obtain a phenotypic sample set, and construct a reproductive trait prediction model according to machine learning, determine the sows to be tested, and use the reproductive trait prediction model according to the genotype data of the sows to be tested to obtain the genomic phenotypic prediction value of the litter traits.

[0024] Further, the specific step S2 is as follows:

[0025] S201. Calibrate the original phenotypic data in the original phenotypic dataset using a linear model; the expression of the linear model is:

[0026] y = xb + e

[0027] where y is the original phenotypic data; x is the association matrix of b; b is the fixed effect vector; e is the residual vector;

[0028] S202. Normalize the residual vector, and based on the normalized residual vector, eliminate the abnormal phenotypic data according to the three - times standard deviation principle to obtain the target phenotypic dataset.

[0029] Further, the specific steps of step S4 are to perform SNP genotyping on each genomic DNA respectively to obtain an initial genotyping dataset, and use Plink software to perform quality control on the initial genotyping dataset to obtain a genotyping dataset; the specific steps of performing quality control are: eliminating the loci without chromosome position information, the loci on sex chromosomes, the loci with a minor allele frequency less than 0.01, the loci with an SNP detection rate of less than 90% for individuals, and the loci with a Hardy - Weinberg equilibrium test P < 10 -6 of the loci.

[0030] Further, the specific steps of step S5 are:

[0031] S501. Obtain the set of local pigs to be sequenced;

[0032] S502. Perform whole - genome re - sequencing on the local pigs to be sequenced, and obtain several groups of re - sequencing data with MAF ≥ 0.05, missing rate < 10 -6 , sequencing depth ≥ 6, and SNPs located on autosomes;

[0033] S503. Divide each re - sequencing data into a reference genotype dataset and an alignment genotype dataset;

[0034] S504. Using the re - sequencing data in the reference genotype dataset as a reference template, perform genotype imputation on the genotyping data in the genotyping dataset to obtain an initial imputation result;

[0035] S505. Use the genotype data in the alignment genotype dataset as a reference to compare with the loci of the initial imputation result, and calculate the imputation accuracy of each SNP:

[0036]

[0037] where, Q c is the imputation accuracy of the SNP; R SNP is the number of correctly imputed individuals; N SNP is the total number of individuals at the locus;

[0038] S506. Retain SNPs with a filling accuracy of 1 and a MAF ≥ 0.01 to obtain a SNP data set.

[0039] Furthermore, the step S6 is specifically as follows:

[0040] S601. Based on the SNP dataset and the target phenotype dataset, a number of principal components are obtained by principal component analysis, each principal component is used as a covariate for GWAS genome-wide association analysis, and GWAS genome-wide association analysis is performed based on the SNP dataset and the target phenotype dataset to obtain the GWAS genome-wide association analysis results;

[0041] S602. Based on the results of GWAS genome-wide association analysis, select -5 SNPs are used as the first weighted locus set;

[0042] S603, selecting loci related to target phenotypic data according to the pig QTL database as a second weighted locus set;

[0043] S604, using GEMMA software to calculate the SNP explanation percentage of each site in the first weighted site set and the second weighted site set, and determining the weight of each weighted site according to the explanation percentage of each weighted site;

[0044] S605 . Obtain a SNP training data set according to the first weighted site set, the second weighted site set and the weight of each weighted site.

[0045] Furthermore, the step S7 is specifically as follows:

[0046] S701, making one-to-one correspondence between the SNP training data set and the target phenotype data set, eliminating the target phenotype data without corresponding relationship, and obtaining a phenotype sample set;

[0047] S702. Based on the phenotypic sample set, a reproductive trait prediction model is constructed using machine learning based on the three reproductive traits; specifically, the reproductive trait prediction model includes a birth litter weight prediction model, a litter size prediction model, and a live piglet number prediction model;

[0048] S703, obtaining the first weighted locus set, the second weighted locus set and the weight of each weighted locus of the sow to be tested, and obtaining the genomic phenotypic prediction value of the litter weight prediction model, the litter size prediction model and the live piglet number prediction model.

[0049] Further, the step S702 is specifically as follows:

[0050] S7021. Use litter weight, litter size, and number of live piglets as labels for the phenotypic sample set based on three reproductive traits respectively to form the first training sample set, the second training sample set, and the third training sample set.

[0051] S7022. Obtain the current training sample set. Use the reproductive traits in the current training sample set as the prediction targets, and the values of the reproductive traits in the current training sample set as the actual target values.

[0052] S7023. Construct a gradient boosting decision tree based on the current training sample set to obtain a reproductive trait prediction model.

[0053] S7024. Execute steps S7022 - S7023 on the first training sample set, the second training sample set, and the third training sample set respectively to obtain a birth litter weight prediction model, a litter size prediction model, and a number of live piglets prediction model.

[0054] Further, the expression of the reproductive trait prediction model in step S7023 is:

[0055]

[0056] Among them, F(x, w) is the reproductive trait prediction value; x is the input sample, w is the model parameter; M is the total number of decision trees; α m is the weight of the m - th decision tree; h m is the m - th decision tree; w m is the parameter of the m - th decision tree; f m (x, w m ) is the result of the m - th decision tree.

[0057] The beneficial effects of the present invention are as follows: The present invention solves the problem of large differences between the early estimated breeding value and the later actual phenotype caused by the small scale of the local pig breeding population and the low accuracy of early genetic evaluation, and improves the prediction accuracy of the reproductive traits of local pigs. BRIEF DESCRIPTION OF THE DRAWINGS

[0058] Figure 1 is the flowchart of the method of the present invention.

[0059] Figure 2 is the schematic diagram of phenotypic prediction accuracy under different SNP densities in the embodiment of the present invention.

[0060] Figure 3 is the correlation test chart between the phenotypic prediction result of the GBDT_PP method in Rongchang pigs and the true phenotypic value in the embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0061] The specific embodiments of the present invention will be described below to facilitate those skilled in the art of this technology to understand the present invention. However, it should be clear that the present invention is not limited to the scope of the specific embodiments. For those of ordinary skill in the art of this technology, as long as various changes are within the spirit and scope of the present invention defined and determined by the appended claims, these changes are obvious, and all inventions created using the concept of the present invention are within the scope of protection.

[0062] As Figure 1 shown, in an embodiment of the present invention, a method for predicting genomic phenotypic values of pig litter traits based on machine learning includes the following steps:

[0063] S1. Determine local pig breeds, and based on the local pig breeds, collect the reproductive traits of several sows to obtain an original phenotypic data set; the reproductive traits include litter weight at birth, litter size, and number of live piglets;

[0064] S2. Use a linear model to correct the original phenotypic data set and remove abnormal phenotypes to obtain a target phenotypic data set;

[0065] S3. Extract the genomic DNA of the sows corresponding to each target phenotypic data in the target phenotypic data set;

[0066] S4. Perform SNP genotyping on each genomic DNA to obtain a genotyping data set;

[0067] S5. Perform genotype imputation on the genotyping data set to obtain an SNP data set;

[0068] S6. Perform GWAS (Genome-Wide Association Study) on the SNP data set to obtain an SNP training data set;

[0069] S7. Based on the SNP training data set and the target phenotypic data set, obtain a phenotypic sample set, and construct a reproductive trait prediction model according to machine learning. Determine the sows to be tested, and use the reproductive trait prediction model based on the genotype data of the sows to be tested to obtain the genomic phenotypic prediction values of the litter traits.

[0070] In this embodiment, aiming at the problems of small scale of the local pig breeding population and large differences between the early estimated breeding values and the later actual phenotypes caused by low accuracy of early genetic evaluation, taking Rongchang pigs as the research object, a machine learning model was used to carry out early prediction research on the reproductive traits of Rongchang pigs based on whole-genome resequencing, explore the effects of different SNP densities and different prediction models on genetic evaluation, and construct an optimal model for early prediction of the reproductive traits of Rongchang pigs based on machine learning, providing data support for the genomic collaborative selection technology with estimated breeding values and phenotypic prediction values as the breeding selection criteria. For small-population local pigs, an optimal model for phenotypic prediction of pig reproductive traits based on whole-genome information was constructed - the genomic phenotypic prediction method for reproductive traits based on gradient boosting decision tree (Gradient Boosting Decision Tree for Phenotypic Prediction, GBDT_PP), which has the advantages of high prediction accuracy, fast calculation speed, etc., and has excellent prediction effects in small-population local pigs.

[0071] The specific steps of step S2 are as follows:

[0072] S201. Use a linear model to correct the original phenotypic data in the original phenotypic dataset; the expression of the linear model is:

[0073] y = xb + e

[0074] where y is the original phenotypic data; x is the correlation matrix of b; b is the fixed effect vector; e is the residual vector;

[0075] S202. Normalize the residual vector, and based on the normalized residual vector, eliminate the abnormal phenotypic data according to the three-standard-deviation principle to obtain the target phenotypic dataset.

[0076] In this embodiment, the core breeding population of Rongchang pigs in the national core breeding farm of Rongchang pigs was selected in this experiment. The reproductive traits (litter weight at birth, number of piglets born, number of live piglets) of 550 sows were collected, and ear tissue samples were collected and frozen in 75% alcohol for preservation. To avoid the influence of other factors, a linear model was used to correct the recorded phenotypes. The correction model is as follows:

[0077] y = xb + e

[0078] where b is the fixed effect vector, including factors to be eliminated such as production year, month, parity, etc., and e is the residual vector. After calculating e and normalizing it, 35 abnormal phenotypic data outside three standard deviations were eliminated according to the three-standard-deviation principle, and 515 corrected phenotypic data of Rongchang pigs were retained for use as label data for subsequent machine learning model training.

[0079] Step S4 specifically involves performing SNP genotyping on each genomic DNA to obtain an initial genotyping dataset, and using Plink software to perform quality control on the initial genotyping dataset to obtain a genotyping dataset; the quality control specifically includes: removing loci without chromosomal position information, loci on sex chromosomes, loci with a minor allele frequency less than 0.01, loci with an SNP detection rate of less than 90% in an individual, and loci with a Hardy-Weinberg equilibrium test P < 10 -6 loci.

[0080] In this embodiment, genomic DNA of pig ear tissue samples is extracted using the chloroform extraction method, and DNA quality is detected by an ultraviolet spectrophotometer (Nanodrop2000) and 1% agarose gel electrophoresis. When the concentration > 100 ng·μL -1 , 1.8 ≤ OD 260nm / OD 280nm ≤ 2.0, the band is clear without trailing degradation, then the DNA quality inspection is qualified; otherwise, DNA is prepared again until the quality inspection is passed.

[0081] The qualified genomic DNA samples are subjected to SNP genotyping using the "Zhongxin No. 1" chip, and the obtained genotyping data is subjected to quality control using Plink (V1.90) software. Loci without chromosomal position information and loci on the Y chromosome, loci with a minor allele frequency (MAF) less than 0.01, loci with an SNP detection rate of less than 90% in an individual, and loci with a Hardy-Weinberg equilibrium test P < 10 -6 loci are removed, and finally 31,756 SNPs are retained.

[0082] Step S5 specifically includes:

[0083] S501. Obtain a set of local pigs to be sequenced;

[0084] S502. Perform whole-genome resequencing on the local pigs to be sequenced to obtain several groups of resequencing data with MAF ≥ 0.05, a deletion rate < 10 -6 , a sequencing depth ≥ 6, and SNPs located on autosomes;

[0085] S503. Divide each resequencing data into a reference genotype dataset and a comparison genotype dataset;

[0086] S504. Using the resequencing data in the reference genotype dataset as a reference template, perform genotype filling on the genotyping data in the genotyping dataset to obtain an initial filling result;

[0087] S505. Compare the genotypes in the comparison genotype dataset with the loci of the initial imputation results as a reference, and calculate the imputation accuracy of each SNP:

[0088]

[0089] Among them, Q c is the imputation accuracy of the SNP; R SNP is the number of correctly imputed individuals; N SNP is the total number of individuals at the locus;

[0090] S506. Retain SNPs with an imputation accuracy of 1 and MAF≥0.01 to obtain an SNP dataset.

[0091] In this example, 150bp paired-end reads on the BGI Wuhan platform were used to perform 30× whole-genome sequencing on 120 Rongchang pigs. After sequencing, the FastQC software was used to evaluate the quality of the original reads by a minimum Phred score of 20 to filter out adapter-contaminated reads and reads containing excessive "N" characters (where "N" accounts for more than 10% of a single read) to obtain Clean reads. Then, the BWA software was used to align the clean reads with the pig reference genome (Sscrofa11.1) with the parameters of "mem -t 10 -k32 -M". The generated SAM file was converted to a BAM file using SAMtools. The Picard 1.119 version's markduplduplicate (mark duplicate reads elimination) program was used to eliminate potential PCR (polymerase chain reaction) duplicates. Subsequently, the GATK software using the multi-sample method was used to detect SNPs (single nucleotide polymorphisms) using the BAM file. The original SNPs generated by GATK were quality controlled according to the following criteria: QualByDepth<2.0, FisherStrand<60.0, RMSMappingQuality<40.0, MappingQualityRankSumTest<-12.5, ReadPosRankSumTest<-8.0. Finally, a total of 21,104,245 SNPs were retained. In subsequent analyses, resequencing data with MAF≥0.05, a missing rate<10 -6 , a sequencing depth≥6, and the retained SNPs located on autosomes were used as reference gene data for subsequent genotype imputation.

[0092] Based on the obtained whole-genome sequencing data of 120 individuals, the whole-genome sequencing data of 90 Rongchang pigs were randomly selected as the reference template, and the genotype imputation was performed using the genotype dataset obtained by Beagle. To evaluate the accuracy of imputation, the sequencing data of another 30 Rongchang pigs were used as the reference to compare with the imputed loci, the genotypes of each locus were compared, and any incorrect loci were correctly imputed through this comparison process. Subsequently, the imputation accuracy of the locus was calculated as the ratio of the number of correctly imputed individuals to the total number of individuals at that locus, and SNPs with an imputation accuracy of 1 and MAF≥0.01 were retained. Therefore, a total of 8,823,367 SNPs were retained, which were used for subsequent GWAS analysis on the one hand, that is, to construct a prior information dataset; on the other hand, they were used for machine learning model training to construct the optimal machine learning model for predicting litter traits.

[0093] The specific steps of step S6 are as follows:

[0094] S601. Based on the SNP dataset and the target phenotype dataset, several principal components are obtained through principal component analysis, and each principal component is used as a covariate for GWAS (genome-wide association study) genome-wide association analysis. And according to the SNP dataset and the target phenotype dataset, GWAS genome-wide association analysis is performed to obtain the GWAS genome-wide association analysis results;

[0095] S602. According to the GWAS genome-wide association analysis results, SNPs with a P-value not less than 10 -5 are selected as the first weighted locus set;

[0096] S603. According to the pig QTL database, loci related to the target phenotype data are selected as the second weighted locus set;

[0097] S604. Use the GEMMA software to calculate the SNP explanation percentage of each locus in the first weighted locus set and the second weighted locus set, and determine the weight of each weighted locus according to the explanation percentage occupied by each weighted locus;

[0098] S605. According to the first weighted locus set, the second weighted locus set and the weights of each weighted locus, an SNP training dataset is obtained.

[0099] In this embodiment, plink1.9 was used to perform genome-wide association analysis (GWAS) on 515 Rongchang pigs, 30-population principal component (PCA) analysis was performed using the GCTA software, the P-values of each PCA were calculated using the EIGENSTRAT software, and the PCAs with P-values less than 0.01 were taken as covariates.

[0100] In addition, according to the GWAS results, the loci significantly associated with the target trait are screened out as weighted locus 1; according to the pig QTL database, the loci related to the target trait are screened out as weighted locus 2; the GEMMA software is used to calculate the SNP explanation percentage of each locus, and the weight is determined according to the explanation percentage occupied by the weighted locus.

[0101] The specific steps of step S7 are as follows:

[0102] S701. One-to-one correspondence is made between the SNP training data set and the target phenotype data set, and the target phenotype data without corresponding relationship is removed to obtain a phenotype sample set;

[0103] S702. Based on the phenotype sample set, machine learning is used to construct a reproductive trait prediction model for each of the three reproductive traits; specifically, the reproductive trait prediction model includes a prediction model for litter weight at birth, a prediction model for litter size, and a prediction model for the number of live piglets;

[0104] S703. Obtain the first weighted locus set, the second weighted locus set of the sows to be tested, and the weights of each weighted locus, and based on the prediction models for litter weight at birth, litter size, and the number of live piglets, obtain the genomic phenotype prediction values of the litter traits.

[0105] The specific steps of step S702 are as follows:

[0106] S7021. Based on the three reproductive traits of the phenotype sample set, the litter weight, litter size, and the number of live piglets are used as labels respectively to form the first training sample set, the second training sample set, and the third training sample set;

[0107] S7022. Obtain the current training sample set, with the reproductive trait in the current training sample set as the prediction target and the value of the reproductive trait in the current training sample set as the actual target value;

[0108] S7023. According to the current training sample set, construct a gradient boosting decision tree to obtain a reproductive trait prediction model;

[0109] S7024. Steps S7022 - S7023 are respectively executed on the first training sample set, the second training sample set, and the third training sample set to correspondingly obtain a prediction model for litter weight at birth, a prediction model for litter size, and a prediction model for the number of live piglets.

[0110] The expression of the reproductive trait prediction model in step S7023 is as follows:

[0111]

[0112] Among them, F(x, w) is the reproductive trait prediction value; x is the input sample, w is the model parameter; M is the total number of decision trees; α m is the weight of the m-th decision tree; hm is the m-th decision tree; w m is the parameter of the m-th decision tree; f m (x, w m ) is the result of the m-th decision tree.

[0113] In this embodiment, in order to avoid reusing population genotype information, additional 50K chip data (genotyping dataset) of 500 Rongchang pigs were collected, and the chip data were imputed for genotypes. Randomly select 400 of them for machine learning model training (the remaining 100 are used for subsequent independent verification), and calculate the correlation coefficient by means of 5-fold cross-validation, and take the average value as an evaluation of the accuracy of the model.

[0114] In addition, select the data of the remaining 100 Rongchang pigs for independent verification, that is, simulate the phenotypic estimation of the candidate population using the trained model in actual production. The results of the two kinds of tests are synthesized to evaluate the differences between different models.

[0115] The results are as follows:

[0116] The present invention uses 50K SNP chips (genotyping dataset) and whole-genome sequencing data of Rongchang pigs for genotype imputation. The results show that: after genotype imputation, a total of 21,104,245 SNPs are available for analysis, and the average accuracy is 0.94. The imputation accuracy was evaluated by comparing the sites between the whole-genome resequencing and imputed data of 30 Rongchang pigs. At three thresholds of 0.3, 0.6, and 0.9, the imputation accuracy ranges are 0.930 - 0.953, 0.934 - 0.955, and 0.967 - 0.977 respectively. Subsequently, after screening 100% imputation accuracy and excluding sites with MAF < 0.01, a total of 8,823,367 SNP sites were obtained for subsequent analysis.

[0117] For the three reproductive traits of Rongchang pigs, the significant sites related to LW (litter weight at birth) are mainly located on SSC12, while the sites related to TNB (total number of born) and NBA (number of live born) are mainly located on SSC14 and SSC15. Select the sites with a significance level of P < 10 -5 as the first weighted site set, retaining 639, 250, and 266 sites respectively.

[0118] In addition, through the Pig QTL database, 75,933 SNPs significantly associated with total number of born and number of live born, and 69,591 SNPs significantly associated with litter weight at birth were screened out, and these sites were used as the second weighted site set.

[0119] Taking the uniform distribution across the whole genome as the standard, 10 gene datasets with 1 million, 2 million, 3 million, 4 million, 5 million, 6 million, 7 million, 8 million, 9 million, and 10 million SNPs were constructed based on the obtained imputed data. Five-fold cross-validation with 10 repetitions was used to evaluate the accuracy of total number born, number born alive, and litter weight at birth.

[0120] As Figure 2 shown, at different SNP densities, the phenotypic prediction accuracy first increased and then decreased with the increase in the number of loci, suggesting that the higher the density of genomic data is not necessarily better. The peaks of the phenotypic prediction accuracy for the three traits all appeared in the range of 800,000 - 900,000 SNPs, and the prediction accuracies were 0.581 ± 0.085, 0.642 ± 0.044, and 0.669 ± 0.029 respectively, indicating that the prediction effect of reproductive traits based on the GBDT model is good. Figure 2 In [reference], Weighted indicates whether weighting is applied.

[0121] Two types of training sample sets were constructed respectively. One assigned weights to the loci in the first weighted locus set and the second weighted locus set, and the other did not assign weights. The results of the training sets with different weights showed that the weighting of prior information data on reproductive traits had different effects on machine learning models in different situations, and there were significant differences. However, when the number of loci was large, the phenotypic prediction accuracy fluctuated less after weighting with prior information, indicating that locus weighting was beneficial to improving the stability of the model in the case of a very large number of loci.

[0122] Based on machine learning methods, the prediction model was constructed and optimized using the best-performing hyperparameters in cross-validation for evaluation. In the independent validation results, the prediction results of litter weight at birth and number born traits were basically consistent with the cross-validation results, with the maximum correlation coefficient being 0.998 and the minimum being 0.935, indicating that the trained machine learning model had good generality.

[0123] The CPU nodes in the Slurm high-performance cluster were used for calculation, with an operating frequency of 2.60 GHz and a memory of 192 G. In the stage of training the machine learning model, a large amount of effort and time were required, and the input cost was proportional to the SNP density. After the model training was completed, the GBDT_PP model could quickly predict the individual phenotype based on the input data, with the advantages of high throughput and fast speed.

[0124] Table 1 Time consumption of each link in the training of the machine learning model for pig litter weight trait at different SNP densities

[0125]

[0126] Table 2 Optimal model hyperparameters for genomic phenotypic prediction of pig litter size traits

[0127]

[0128] As Figure 3 shown, after the GBDT_PP method (the present invention) was optimized and improved, 30 pig samples were additionally collected from the pig farm to evaluate the phenotypic prediction effect. The results showed that among the three reproductive traits of litter weight, total number born (TNB), and number born alive (NBA), the prediction results of the GBDT_PP model had no significant difference from the actual phenotypes of individuals (P>0.05), and the prediction accuracies were 0.46, 0.65, and 0.51, respectively. Generally speaking, GBDT_PP had a good phenotypic prediction effect on reproductive traits in small populations of indigenous pigs and could be applied to actual production breeding. Figure 3 In which, Litter Weight is litter weight, TNB is total number born, and NBA is number born alive.

Claims

1. A method for predicting genomic phenotypic values ​​of pig litter traits based on machine learning, characterized in that: The following steps are involved: S1. Determine local pig breeds, and based on the local pig breeds, collect the reproductive traits of several sows to obtain an original phenotypic data set; the reproductive traits include birth litter weight, litter size, and live piglets; S2, using a linear model to correct the original phenotypic data set and remove abnormal phenotypes to obtain a target phenotypic data set; the step S2 is specifically as follows: S201, correcting the original phenotypic data in the original phenotypic data set using a linear model; the expression of the linear model is: y=xb+e Among them, y is the original phenotypic data; x is the association matrix of b; b is the fixed effect vector; e is the residual vector; S202, normalizing the residual vector, and according to the normalized residual vector, eliminating abnormal phenotypic data based on the principle of three times the standard deviation, to obtain a target phenotypic data set; S3, extracting the genomic DNA of the sow corresponding to each target phenotype data in the target phenotype data set; S4, performing SNP genotyping on each genomic DNA to obtain a genotyping data set; S5, performing genotype filling on the genotyping data set to obtain a SNP data set; S6, performing GWAS genome-wide association analysis on the SNP data set to obtain a SNP training data set; the step S6 is specifically as follows: S601. Based on the SNP dataset and the target phenotype dataset, a number of principal components are obtained by principal component analysis, each principal component is used as a covariate for GWAS genome-wide association analysis, and GWAS genome-wide association analysis is performed based on the SNP dataset and the target phenotype dataset to obtain the GWAS genome-wide association analysis results; S602. Based on the results of GWAS genome-wide association analysis, select -5 SNPs are used as the first weighted locus set; S603, selecting loci related to target phenotypic data according to the pig QTL database as a second weighted locus set; S604, using GEMMA software to calculate the SNP explanation percentage of each site in the first weighted site set and the second weighted site set, and determining the weight of each weighted site according to the explanation percentage of each weighted site; S605, obtaining a SNP training data set according to the first weighted site set, the second weighted site set and the weight of each weighted site; S7. Based on the SNP training data set and the target phenotype data set, a phenotypic sample set is obtained, and a reproductive trait prediction model is constructed according to machine learning to determine the sows to be tested. According to the genotype data of the sows to be tested, the reproductive trait prediction model is used to obtain the genomic phenotypic prediction value of the littering trait.

2. The method for predicting genomic phenotypic values ​​of pig litter traits based on machine learning according to claim 1, characterized in that: The step S4 specifically includes performing SNP genotyping on each genomic DNA to obtain an initial genotyping data set, and using Plink software to perform quality control on the initial genotyping data set to obtain a genotyping data set; the quality control specifically includes: eliminating sites without chromosomal location information, sites on sex chromosomes, sites with a minimum allele frequency of less than 0.01, sites with an individual SNP detection rate of less than 90%, and sites with a Hardy-Weinberg equilibrium test P<10 -6 The location of .

3. The method for predicting genomic phenotypic values ​​of pig litter traits based on machine learning according to claim 1, characterized in that: The step S5 is specifically as follows: S501, obtaining a local pig collection to be sequenced; S502. Perform whole genome resequencing on the local pigs to be sequenced, and obtain several groups with MAF ≥ 0.05 and missing rate < 10 -6 , sequencing depth ≥ 6 and resequencing data with SNPs located on autosomes; S503, dividing each resequencing data into a reference genotype data set and a comparison genotype data set; S504, using the resequencing data in the reference genotype data set as a reference template, performing genotype filling on the genotyping data in the genotyping data set to obtain an initial filling result; S505, using the genotype data in the comparison genotype data set as a reference to compare with the sites of the initial filling result, calculate the filling accuracy of each SNP: Among them, Q c is the filling accuracy of SNP; R SNP is the number of individuals filled correctly; N SNP is the total number of individuals at the locus; S506. Retain SNPs with a filling accuracy of 1 and a MAF ≥ 0.01 to obtain a SNP data set.

4. The method for predicting genomic phenotypic values ​​of pig litter traits based on machine learning according to claim 1, characterized in that: The step S7 is specifically as follows: S701, making one-to-one correspondence between the SNP training data set and the target phenotype data set, eliminating the target phenotype data without corresponding relationship, and obtaining a phenotype sample set; S702. Based on the phenotypic sample set, a reproductive trait prediction model is constructed using machine learning based on the three reproductive traits; specifically, the reproductive trait prediction model includes a birth litter weight prediction model, a litter size prediction model, and a live piglet number prediction model; S703, obtaining the first weighted locus set, the second weighted locus set and the weight of each weighted locus of the sow to be tested, and obtaining the genomic phenotypic prediction value of the litter size prediction model, the litter size prediction model and the live piglet number prediction model.

5. The method for predicting genomic phenotypic values ​​of pig litter traits based on machine learning according to claim 4, characterized in that: The step S702 is specifically as follows: S7021, using litter weight, litter size and live piglets as labels for the phenotypic sample sets as the first training sample set, the second training sample set and the third training sample set based on the three reproductive traits; S7022, obtaining a current training sample set, taking the reproductive traits in the current training sample set as prediction targets, and the values ​​of the reproductive traits in the current training sample set as actual target values; S7023. According to the current training sample set, a gradient boosting decision tree is constructed to obtain a reproductive trait prediction model; S7024, executing steps S7022 to S7023 for the first training sample set, the second training sample set, and the third training sample set, respectively, to obtain a birth litter weight prediction model, a litter size prediction model, and a live piglet number prediction model accordingly.

6. The method for predicting genomic phenotypic values ​​of pig litter traits based on machine learning according to claim 5, characterized in that: The expression of the reproductive trait prediction model in step S7023 is: Among them, F(x,w) is the predicted value of reproductive traits; x is the input sample, w is the model parameter; M is the total number of decision trees; α m is the weight of the mth decision tree; h m is the mth decision tree; w m For the m The parameters of a decision tree; f m (x,w m ) is the result of the mth decision tree.

Citation Information

Patent Citations

  • Genome breeding value estimation method based on high-throughput sequencing platform and application thereof

    CN110564832A

  • Machine learning method and system for egg laying character genome selection of laying hens

    CN116072226A

  • SNP (Single Nucleotide Polymorphism) site combination for identifying Rongchang pig variety and application of SNP site combination

    CN117431325A

  • Genomic estimated breeding value for predicting a trait and its application

    TWI786916B