Genome selective breeding method and system based on EN-CNN model

By using the EN-CNN model in genome selection breeding, using an autoencoder to extract important features of genotype data and combining convolutional neural networks for prediction, the problem of low prediction accuracy in the existing technology is solved, and more efficient breeding prediction is achieved.

CN120220818APending Publication Date: 2025-06-27WUHAN ACADEMY OF AGRI SCI

Patent Information

Application Number
CN202510321499.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-18
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The prediction accuracy of existing genome selection breeding methods is low, and their application is limited, making it difficult to effectively capture the complex interactions and nonlinear relationships between SNPs.

Method used

A genome selection breeding method based on the EN-CNN model is adopted to capture important feature information from the genotype data through an autoencoder and use a convolutional neural network model to make predictions to improve the accuracy of breeding.

Benefits of technology

It improves the accuracy and robustness of genome selection breeding, can more effectively capture the nonlinear relationship between genotype and phenotype, and improves the predictive performance of breeding.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120220818A_ABST
    Figure CN120220818A_ABST
Patent Text Reader

Abstract

The invention provides a genome selective breeding method and system based on an EN-CNN model. The method comprises the following steps: capturing important feature information from genotype data of a to-be-bred object by using an auto-encoder; predicting growth traits of the to-be-bred object according to the important feature information by using a convolutional neural network model; and according to the growth traits of the to-be-bred object, determining whether to select the to-be-bred object for breeding. According to the method, the precision and robustness of genome selective breeding are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of genetic breeding, and particularly relates to a genomic selection breeding method and system based on an EN-CNN model. Background Art

[0002] As one of the most important livestock and poultry, pigs play a crucial role in both economic and social aspects. Pork is an indispensable meat consumer product in people's daily diet, and the stability of its supply directly affects people's quality of life. Pig breeding plays a decisive role in improving pork production efficiency. Therefore, it is particularly crucial to explore and apply efficient breeding technologies. Pig breeding is a technology aimed at improving the production performance and economic benefits of pigs through selection and breeding. Through the genetic breeding of excellent pig breeds, the reproductive capacity, growth rate, and feed conversion rate of pigs can be effectively improved, thereby enhancing economic benefits. However, in actual production, traditional pig breeding methods usually require six months of feeding and data collection at intelligent measurement stations, resulting in a long time cycle, low efficiency, and high cost. Therefore, there is an urgent need to adopt more efficient breeding technologies.

[0003] In genomic research, resequencing technology and chip technology are two main methods for obtaining genomic information. Although whole genome sequencing (WGS) can provide more comprehensive genomic information, resequencing technology is more costly compared to chip technology. Genotype imputation technology can supplement missing single nucleotide polymorphism (SNP) data and integrate SNP sites of commercial chips with different specifications, thereby reducing costs and improving the accuracy of breeding.

[0004] Genomic Selection (GS) technology has shown significant advantages in the early prediction of genomic breeding values of animals and plants by analyzing the association between SNPs markers and phenotypic data of individuals. Compared with traditional phenotypic selection methods, GS technology not only improves the accuracy of prediction but also accelerates the process of genetic improvement, providing strong scientific and technological support for the sustainable development of the pig breeding industry.

[0005] Genomic Best Linear Unbiased Prediction (GBLUP) is considered to be an effective and accurate genomic selection method. It usually predicts unobserved complex traits of individuals based on genomic information (usually SNPs) obtained through sequencing under a priori assumptions. It cannot capture the nonlinear relationship between genotype and phenotype, and has low prediction accuracy and computational efficiency in large-scale high-dimensional genomic data. Machine learning is an efficient data prediction method and has been widely used in the field of animal genetics and breeding. Random forest model and support vector machine model are commonly used supervised learning models in machine learning. Random forest model is an integrated learning method that can be applied to genotype breeding prediction based on single nucleotide polymorphism (SNP). Compared with other machine learning algorithms, support vector machine model (SVM) is very powerful in identifying subtle patterns in small samples or complex data sets, and has been applied to the prediction of animal growth traits. Machine learning still has limitations in genomic prediction. Machine learning methods may not be able to fully capture the complex interactions and nonlinear relationships between SNPs, resulting in insufficient prediction ability.

[0006] Deep learning (DL), as a subset of machine learning, uses the hierarchical level of neural networks to process data and exhibits excellent automatic feature extraction capabilities. Compared with traditional methods, deep learning has obvious advantages in the field of breeding genetic estimation and can mine complex genetic patterns. The core advantage of deep learning (DL) lies in its excellent automatic feature extraction ability, which can mine complex genetic patterns from high-dimensional genomic data, which are often difficult to identify by traditional methods. In addition, deep learning models, especially deep neural networks, have powerful nonlinear mapping capabilities, which enables them to capture the complexity of genetic effects, including interactions between genes and nonlinear relationships between phenotypes and genotypes.

[0007] In the evolution of organisms, traits are mainly determined by genetic factors. However, for a particular trait, not all gene segments of an organism play a decisive role. Existing research uses feature selection methods to extract important SNPs and identifies SNPs that match the selection characteristics of the pig genome through the random forest method (RF), but the prediction results are disappointing. Feature selection methods based on machine learning are used to perform feature selection on genomic information related to the residual feed intake of pigs. However, filtering methods, embedding methods, and their combination may overly rely on univariate statistics. They do not separate the learning and feature selection parts, such as Spearman correlation, which may not be sufficient to capture the complex interactions and non-linear relationships between SNPs, resulting in insufficient prediction ability. Moreover, when the size of the SNP subset is large (such as 1,000 to 1,500 SNPs), these methods have no significant impact on the prediction quality, indicating that in a large SNP set, feature selection may not bring the expected improvement in stability. Kernel principal component analysis (KPCA) adds a kernel function to principal component analysis and can capture non-linear features. However, it cannot reconstruct data, preserve the important features of the original data, and is not suitable for small samples and ultra-high-dimensional features. XGBoost with the characteristics of supervised learning automatically performs feature selection by constructing decision trees, giving priority to features that contribute the most to the prediction results, thereby effectively reducing the complexity of the model and improving the prediction performance. Although XGBoost performs excellently in processing structured data, its dependence on labels and difficulty in effectively capturing complex non-linear relationships restrict the application of the model. Summary of the Invention

[0008] The present invention provides a genomic selection breeding method and system based on an EN-CNN model to solve the defects of low prediction accuracy and limited application in genomic selection breeding in the prior art, and to improve the accuracy of genomic selection breeding.

[0009] The present invention provides a genomic selection breeding method based on an EN-CNN model, including:

[0010] Using an autoencoder to capture important feature information from the genotype data of the object to be bred;

[0011] Using a convolutional neural network model to predict the growth traits of the object to be bred according to the important feature information;

[0012] Determining whether to select the object to be bred for breeding according to the growth traits of the object to be bred.

[0013] According to the genomic selection breeding method based on an EN-CNN model provided by the present invention, the autoencoder includes an encoder and a decoder. Before using the autoencoder to capture important feature information from the genotype data of the object to be bred, it further includes:

[0014] Using the encoder to map the original genotype data of the first breeding object sample to a latent space representation, so as to reduce the dimension of the original genotype data and extract important feature information;

[0015] Using the decoder to reconstruct the original genotype data of the first breeding object sample according to the latent space representation;

[0016] Taking the mean squared error between the original genotype data of the first breeding object sample and its reconstructed data as a loss function to train the autoencoder.

[0017] According to a genomic selection breeding method based on an EN-CNN model provided by the present invention, the encoder includes a fully connected layer, a batch normalization layer, an activation function, and a Dropout layer connected in sequence.

[0018] According to a genomic selection breeding method based on an EN-CNN model provided by the present invention, the convolutional neural network model includes an input layer, a first convolutional layer, a second convolutional layer, a first pooling layer, a third convolutional layer, a fourth convolutional layer, a second pooling layer, a first fully connected layer, a second fully connected layer, and an output layer connected in sequence.

[0019] According to a genomic selection breeding method based on an EN-CNN model provided by the present invention, before using the convolutional neural network model to predict the growth traits of the breeding object to be bred according to the important feature information, it further includes:

[0020] After performing quality control operations on the phenotypic data and genotype data of the second breeding object sample, forming phenotypic-genomic data, and taking the phenotypic data as the actual growth traits;

[0021] Using the autoencoder to extract the important feature information of the second breeding object sample from the genotype data in the phenotypic-genomic data;

[0022] Using the convolutional neural network model to obtain the predicted growth traits of the second breeding object sample according to the important feature information of the second breeding object sample;

[0023] According to the predicted growth traits and actual growth traits of the second breeding object sample, using 10-fold cross-validation to train the convolutional neural network model.

[0024] According to a genomic selection breeding method based on an EN-CNN model provided by the present invention, the phenotypic data includes feed conversion ratio and corrected age at reaching a preset weight. The feed conversion ratio refers to the amount of feed consumed by the breeding object to gain one kilogram of weight, and the corrected age at reaching a preset weight refers to the number of days required for the breeding object to reach the preset weight under corrected conditions.

[0025] A genomic selection breeding method based on an EN-CNN model provided by the present invention performs quality control operations on the phenotypic data and genotypic data of the second breeding object sample, including:

[0026] Eliminating the phenotypic data of the second breeding object sample lacking birth weight or having abnormal phenotypic data;

[0027] Using GATK software to filter SNP sites in the genotypic data with a quality score less than a first preset threshold, an FS value greater than a second preset threshold, a sequencing alignment quality less than a third preset threshold, a rank sum test value of the sequencing alignment quality less than a fourth preset threshold, a rank sum test value of the read segment position less than a fifth preset threshold, or a strand ratio greater than a sixth preset threshold;

[0028] Using the GATK software to filter INDEL sites in the genotypic data with a sequencing depth less than a seventh preset threshold, a rank sum test value of the sequencing alignment quality less than the fourth preset threshold, or a rank sum test value of the read segment position less than the fifth preset threshold;

[0029] Using bcftools software to filter out SNP sites within a preset number of base pairs near INDEL sites in the genotypic data;

[0030] Using Plink software to perform quality control operations on the genotypic data so that the missing rate of the genotypic data is less than an eighth preset threshold, the minor allele frequency is greater than a ninth preset threshold, and the sample missing rate is less than a tenth preset threshold.

[0031] A genomic selection breeding method based on an EN-CNN model provided by the present invention, before extracting important feature information of the second breeding object sample from the genotypic data in the phenotype-genome data using the autoencoder, further includes:

[0032] Calculating the degree of linkage disequilibrium between gene loci in the genomic data of the phenotype-genome data;

[0033] Removing gene loci with a linkage disequilibrium degree greater than a correlation threshold from the genomic data.

[0034] The present invention also provides a genomic selection breeding system based on an EN-CNN model, including:

[0035] An extraction module for capturing important feature information from the genotypic data of the object to be bred using an autoencoder;

[0036] A prediction module, configured to use a convolutional neural network model to predict the growth traits of the breeding object to be bred according to the important feature information;

[0037] A selection module, configured to determine whether to select the breeding object to be bred for breeding according to the growth traits of the breeding object to be bred.

[0038] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the genomic selection breeding method based on the EN-CNN model as described in any one of the above.

[0039] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the genomic selection breeding method based on the EN-CNN model as described in any one of the above.

[0040] The present invention also provides a computer program product, including a computer program. When the computer program is executed by a processor, it implements the genomic selection breeding method based on the EN-CNN model as described in any one of the above.

[0041] A genomic selection breeding method and system based on the EN-CNN model provided by the present invention first trains an auto-encoder as a feature selection module (EN), then uses the feature selection module (EN) to capture important feature information (such as SNPs and Indels) in genomic data, and finally uses the designed convolutional neural network model for prediction, thereby improving the accuracy and robustness of genomic selection breeding. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0043] Figure 1 is a schematic flowchart of the genomic selection breeding method based on the EN-CNN model provided by the present invention;

[0044] Figure 2 is a schematic diagram of the model structure of the genomic selection breeding method based on the EN-CNN model provided by the present invention;

[0045] Figure 3 is a schematic diagram of the PCC value of different gene feature numbers FCR in the genomic selection breeding method based on the EN-CNN model provided by the present invention;

[0046] Figure 4 It is a schematic diagram of the PCC value of different gene feature numbers C_Day_120kg in the genomic selection breeding method based on the EN-CNN model provided by the present invention;

[0047] Figure 5 It is a schematic diagram of the training and validation process of the EN-GS model in the genomic selection breeding method based on the EN-CNN model provided by the present invention;

[0048] Figure 6 It is a schematic diagram for comparing evaluation indexes of genomic estimated breeding values based on the XG-GS model in the genomic selection breeding method based on the EN-CNN model provided by the present invention;

[0049] Figure 7 It is a schematic diagram for comparing evaluation indexes of genomic estimated breeding values based on the EN-GS model in the genomic selection breeding method based on the EN-CNN model provided by the present invention;

[0050] Figure 8 It is a schematic diagram of the structure of the genomic selection breeding system based on the EN-CNN model provided by the present invention. Detailed implementation manners

[0051] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below with reference to the accompanying drawings in the present invention. Obviously, the described embodiments are some but not all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without making creative efforts shall fall within the protection scope of the present invention.

[0052] Data in the field of genomics has shown explosive growth, which has led to unique challenges in the analysis process, especially when faced with a relatively limited number of samples in the face of high-dimensional datasets, namely the so-called "large P, small N" dilemma. This data characteristic, combined with the high-dimensional and complex gene structure, has greatly hindered the effective analysis of large genomic datasets. In the field of animal and plant breeding, the application of feature selection technology is crucial, which enables researchers to identify genetic markers related to specific traits from high-dimensional genomic datasets and design genomic information accordingly to more efficiently evaluate breeding candidate individuals.

[0053] However, despite the increasing widespread application of feature selection techniques in genomics, research on using neural networks for feature dimensionality reduction with whole-genome information is relatively scarce. This study provides a new perspective and tools for in-depth analysis of genomic data by comparing two dimensionality reduction methods. The objectives of this study are: (1) to compare the prediction performance of genomic selection models after feature dimensionality reduction of genomic data using traditional machine learning algorithms and neural network structures; (2) to improve the accuracy of genomic prediction using a specific combination of neural network models.

[0054] XGBoost is a supervised machine learning algorithm based on Gradient Boosting. It achieves high-performance prediction by constructing multiple weak learners (usually decision trees) and gradually improving their prediction capabilities. The basic idea is for a given dataset by training a set of trees F = {f1(x), f2(x), …, f K (x)}, each tree makes a prediction on the input sample feature x i and the final prediction result is the sum of all prediction results. The specific model expression is as follows:

[0055]

[0056] where F is the space of all regression trees, is the prediction result of the sample feature x i and f k (x i ) represents the prediction score of the k-th tree on the sample feature x i .

[0057] The objective function can be defined as:

[0058]

[0059] where θ is the parameter of the model, and the objective function consists of the error function and the regularization term .

[0060] When XGBoost performs feature selection, it trains a regression model to calculate the importance score of each feature. These scores are based on the information gain brought by each feature at the split nodes of all decision trees. The information gain is measured by calculating the change in the loss function before and after splitting. The loss function commonly used is the squared error, and the information gain is:

[0061]

[0062] where G L and G R are the sums of the first-order derivatives of the left and right child nodes respectively, and HL and H R are the sums of the second-order derivatives of the left and right child nodes, respectively, and λ and γ are regularization parameters. Finally, SNPs and Indels with non-zero gain values are selected as feature SNPs and Indels, and feature importance will be calculated using xgb in the Python package "xgboost".

[0063] Genomic best linear unbiased prediction (GBLUP) is a genomic prediction method based on the mixed linear model (MLM), and its structure is generally as follows:

[0064] y = Xβ + Zu + e

[0065] where y is the phenotypic record carrier of each trait (FCR and C_Day_120kg), a phenotypic vector of N×1 (N is the number of individuals); X is an N×p fixed-effect design matrix (p is the number of fixed effects), and β is a p×1 fixed-effect vector; Z is an N×N individual-effect design matrix, and u is an N×1 random-effect vector, representing the genomic breeding value (GBV) of each individual, assuming where G is the genomic relationship matrix

[25] , is the random-effect variance; e is an N×1 error vector, assuming where I is an N×N identity matrix.

[0066] When using the support vector machine regression model (SVR) for training, the Gaussian kernel function (kernel='rbf') is set to handle non-linear regression problems. In addition, the cost parameter (C) and the parameter (gamma) are used to optimize the performance of the model, where the cost parameter controls the tolerance of errors, and the gamma parameter affects the model complexity and the number of support vectors. In the specific implementation, the cost parameter C is 1.0, and gamma is set to the automatic calculation mode scale.

[0067] When using the random forest regression model for training, the model expression is as follows:

[0068]

[0069] where, is the final predicted value, N is the number of trees in the forest, set to 100; f i (X) is the prediction of the i-th tree for the input feature X.

[0070] This study used a fully connected neural network (MLP) model with four hidden layers. Batch Normalization and ReLU activation function were adopted after each hidden layer, and Dropout regularization was introduced to enhance the generalization ability of the model. Through this network structure, complex non-linear patterns in high-dimensional feature data can be effectively captured.

[0071] The input layer accepts the feature vector x ∈ R input _ shape , and the output of each layer is calculated by the following formula:

[0072] h l =ReLU(BatchNorm(Linear(h l-1 ,n l ))) for l = 1, 2, 3, 4

[0073] where n l are the number of neurons in each layer (1024, 512, 256, 128) respectively, and the initial condition is h0 = x. The output layer generates the final prediction value:

[0074]

[0075] To optimize the model, the mean squared error (MSE) was used as the loss function, and its formula:

[0076]

[0077] The following combines Figure 1 to describe a genomic selection breeding method based on the EN-CNN model of the present invention, including:

[0078] Step 101: Use an autoencoder to capture important feature information from the genotype data of the breeding object to be bred;

[0079] Step 102: Use a convolutional neural network model to predict the growth traits of the breeding object to be bred according to the important feature information;

[0080] Step 103: Determine whether to select the breeding object to be bred for breeding according to the growth traits of the breeding object to be bred.

[0081] So far, the research on deep learning (DL) methods applied to the breeding estimation of pig growth traits is still very limited, and there is no method for predicting the genotype to phenotype of pigs that combines a deep learning model as a feature selection method with another deep learning model as a prediction method.

[0082] Due to the ultra-high dimensionality of the original genomic data after sequencing and quality control of live pigs in the farm, SNPs and Indels are extracted from the original genomic data as features and transformed and encoded into a form that can be recognized by deep learning. If directly used as the input of the deep learning prediction model, it will cause the model to fail. Therefore, in this embodiment, an autoencoder is introduced as a feature reduction tool to be fused with the convolutional neural network model, and a convolutional neural network model based on feature selection (EN-CNN) is built to predict and analyze growth traits. The prediction process of the gene EN-CNN model is as Figure 2 shown.

[0083] This embodiment proposes a new CNN-based genomic selection (GS) method called EN-CNN. This method first trains an autoencoder as a feature selection module (EN), then uses the feature selection module (EN) to capture important feature information (such as SNPs and Indels) in the genomic data, and finally uses the designed convolutional neural network model for prediction, so as to improve the accuracy and robustness of genomic selection breeding.

[0084] Based on the above embodiment, the autoencoder in this embodiment includes an encoder and a decoder. Before using the autoencoder to capture important feature information from the genotype data of the object to be bred, it further includes:

[0085] Using the encoder to map the original genotype data of the first breeding object sample to a latent space representation to reduce the dimension of the original genotype data and extract important feature information;

[0086] Using the decoder to reconstruct the original genotype data of the first breeding object sample according to the latent space representation;

[0087] Taking the mean squared error between the original genotype data of the first breeding object sample and its reconstructed data as the loss function to train the autoencoder.

[0088] The autoencoder is an unsupervised learning algorithm designed to identify low-dimensional features of high-dimensional data and has become the core tool for genomic feature reduction in bioinformatics. In order to retain as much of the original information of the high-dimensional genomic data as possible, the autoencoder is applied as an effective deep learning technique to achieve dimensionality reduction and extract effective features.

[0089] The autoencoder is a neural network composed of an encoder network and a decoder network. The former compresses the input data into a low-dimensional representation, also called the latent space, and the latter reconstructs the original data from the compressed representation. By optimizing the network to reconstruct the input data from this compact representation, the autoencoder can learn to filter out noise and retain the most informative features, thus being particularly suitable for complex ultra-high-dimensional genomic data.

[0090] First, the encoder of the autoencoder maps the input \(x\in\mathbb{R}\) d to a low-dimensional latent space represented as \(z\in\mathbb{R}\) k , that is:

[0091] \(z = f(w\) e \(x + b\) e )

[0092] where \(w\) e \(\in\mathbb{R}\) k×d is the weight matrix of the encoding process, \(b\) e \(\in\mathbb{R}\) k is the bias vector of the encoding process, and \(f\) is the activation function.

[0093] Then, the decoder reconstructs the latent space representation \(z\) into the input that is:

[0094]

[0095] where \(w\) d \(\in\mathbb{R}\) d×k is the weight matrix of the decoding process, \(b\) d \(\in\mathbb{R}\) d is the bias vector of the decoding process, and \(g\) is the activation function.

[0096] Finally, its parameters are optimized by minimizing the reconstruction error. Here, the mean square error (MSE) is used as the loss function. The mean square error has good analytical properties, smooth gradients, and is easy to optimize. This makes the training process more stable and avoids the problems of gradient explosion or disappearance. The calculation formula of the mean square error is as follows:

[0097]

[0098] where \(N\) is the number of samples, \(x\) i is the original genotype data of the \(i\)-th sample, is the reconstructed data of the \(i\)-th sample.

[0099] To retain the effective original information of the high-dimensional genomic data, the encoder is applied as a deep learning technique to achieve dimensionality reduction and extract effective genomic features. The encoder compresses the input high-dimensional data into a low-dimensional latent space through a series of nonlinear transformations. The encoded low-dimensional features can be used for prediction tasks such as regression and classification in downstream models.

[0100] Based on the above embodiments, in this embodiment, the encoder includes a fully connected layer, a batch normalization layer, an activation function, and a Dropout layer connected in sequence.

[0101] Among them, Batch Normalization (BatchNorm) can stabilize and accelerate the training process. By normalizing the output values, it can make the data in different batches perform consistently in the activation function. Subsequently, a ReLU activation function is used to introduce non-linearity. At the same time, the regularization technique of Dropout (P = 0.05) is combined to improve the stability and robustness of the model.

[0102] Based on the above embodiments, in this embodiment, the convolutional neural network model includes an input layer, a first convolutional layer, a second convolutional layer, a first pooling layer, a third convolutional layer, a fourth convolutional layer, a second pooling layer, a first fully connected layer, a second fully connected layer, and an output layer, which are connected in sequence.

[0103] The convolutional neural network model is constructed using deep learning technology. It can adopt a 1D convolutional neural network (1D-CNN) structure with an architecture of 64-64-128-128-1. The model includes an input layer, four convolutional layers (with 64 and 128 neurons), two pooling layers, two fully connected layers (with 128 and 1 neuron), and an output layer. The ReLU activation function is used in the convolutional layers and the first linear layer to enhance the non-linear expression ability.

[0104] Based on the above embodiments, before using the convolutional neural network model to predict the growth traits of the breeding target to be based on the important feature information, the following steps are also included:

[0105] After performing quality control operations on the phenotypic data and genotypic data of the second breeding object sample, phenotypic-genomic data is formed, and the phenotypic data is used as the actual growth traits.

[0106] Using the autoencoder to extract the important feature information of the second breeding object sample from the genotypic data in the phenotypic-genomic data.

[0107] Using the convolutional neural network model to obtain the predicted growth traits of the second breeding object sample based on the important feature information of the second breeding object sample.

[0108] Based on the predicted growth traits and actual growth traits of the second breeding object sample, the convolutional neural network model is trained using 10-fold cross-validation.

[0109] The prediction process of the gene EN-CNN model in this embodiment includes:

[0110] Step 1, collect and organize the quality-controlled phenotypic-genomic data.

[0111] Step 2: Without feature selection on the phenotype-genome data, directly use the traditional GBLUP model and GS models (SVM (Support Vector Machine), RF (Random Forest), MLP, and CNN (Convolutional Neural Networks)) that include machine learning and neural networks for prediction, and conduct comparative experimental analysis;

[0112] Step 3: Use XGboost and autoencoder for comparative experiments to extract features from the genome;

[0113] Step 4: Use the best feature subset obtained in Step 3 to divide the validation set and the test set, select the 10-fold cross-validation method, and perform prediction and evaluation through the CNN model.

[0114] First, train autoencoders with different structures to select the best model with the minimum loss. Subsequently, use the trained encoder to reduce the dimension of the genome features, and the obtained low-dimensional features are then used as the input of the CNN to perform the prediction task. By combining the end-to-end form of the CNN, the downstream loss can be effectively minimized and the prediction can be completed.

[0115] In EN-CNN, the EN block can adopt batch normalization (BatchNorm1d, 1024) to accelerate training and improve the stability of the activation function. The parameters of the EN block and the CNN block are optimized using the ADAM optimizer. The initial learning rate is set to 0.001, and the weight decay value for the CNN block is set to 0.00001. The batch normalization sizes of the two are 64 and 32 respectively. In addition, the mean squared error (MSE) is used as the loss function, and an early stopping mechanism is implemented during training to prevent overfitting and improve the generalization ability of the model. All experiments are implemented in PyTorch and conducted on a single NVIDIA GeForce RTX 490. The EN-CNN structure is robust and stable, and it uses the above parameters for all datasets and all features.

[0116] To improve the performance of the classification model, 10-fold cross-validation is introduced to evaluate the effectiveness of the EN-CNN model. First, the individuals in the entire genome prediction dataset are randomly divided into 10 groups of equal size. The genotype and phenotype data of the individuals from 9 groups are used to train and validate the model (90% of the individuals are the training set; 10% of the individuals are used as the validation set). This process is repeated 10 times until each group is used for testing once.

[0117] The root mean square error (RMSE) and mean absolute error (MAE) were used as regression evaluation indicators to verify the performance of the EN-CNN model. Additionally, the Pearson correlation coefficient (PCC) of the model was calculated to describe the correlation between the true values and predicted values. The calculation formulas are as follows:

[0118]

[0119] Based on the above embodiments, in this embodiment, the phenotypic data includes the feed conversion ratio and the corrected age at a preset body weight. The feed conversion ratio refers to the amount of feed consumed by the breeding object to gain one kilogram of weight, and the corrected age at a preset body weight refers to the number of days required for the breeding object to reach the preset body weight under corrected conditions.

[0120] Duroc pigs in a certain pig core breeding farm can be used as the experimental object for research. The numbers, strains, farms, birth years, measurement years, genders, and birth weights of 463 individuals were measured and collected. The phenotypic traits studied included the feed conversion ratio and the age at 120 kg body weight, represented by FCR and C_Day_120kg respectively.

[0121] The feed conversion ratio refers to the amount of feed consumed by a pig to gain one kilogram of weight, and it is a key indicator for evaluating feed utilization efficiency. Reducing the feed conversion ratio means that pigs can gain weight with less feed, thereby reducing feed costs. Feed costs account for a large proportion in pig farming, so optimizing the feed conversion ratio is an important way to improve economic benefits.

[0122] The corrected age at 120 kg body weight refers to the number of days required for a pig to reach 120 kg body weight under corrected conditions. This indicator comprehensively considers the growth rate and feeding conditions. Shortening the time to reach the target body weight can improve the production turnover rate, increase the number of pigs sold per unit time, and improve the overall production efficiency. The two together can promote the breeding farm to better plan resource allocation and production cycles.

[0123] The genetic data of the experimental pig group can be measured on samples using paired-end sequencing and a sequencing platform to obtain the original genomic data.

[0124] Based on the above embodiments, in this embodiment, quality control operations were performed on the phenotypic data and genotype data of the second breeding object samples, including:

[0125] Eliminating the phenotypic data of the second breeding object samples lacking birth weight or with abnormal phenotypic data;

[0126] Filter the SNP sites in the genotype data whose quality score is less than the first preset threshold, FS value is greater than the second preset threshold, sequencing alignment quality is less than the third preset threshold, rank sum test value of sequencing alignment quality is less than the fourth preset threshold, rank sum test value of read segment position is less than the fifth preset threshold, or strand ratio is greater than the sixth preset threshold by using the GATK software;

[0127] Filter the INDEL sites in the genotype data whose sequencing depth is less than the seventh preset threshold, rank sum test value of sequencing alignment quality is less than the fourth preset threshold, or rank sum test value of read segment position is less than the fifth preset threshold by using the GATK software;

[0128] Filter out the SNP sites within a preset number of base pairs near the INDEL sites in the genotype data by using the bcftools software;

[0129] Perform quality control operations on the genotype data by using the Plink software so that the missing rate of the genotype data is less than the eighth preset threshold, the minor allele frequency is greater than the ninth preset threshold, and the sample missing rate is less than the tenth preset threshold.

[0130] The quality control of phenotypic data includes excluding individuals with abnormal birth weight, FCR, and C_Day_120kg. The number of remaining valid individuals is 412, as shown in Table 1.

[0131] Table 1 Description of phenotypic data set

[0132]

[0133] For genotype data quality control, first, use GATK and Bcftools to perform strict quality control and filtering on genomic variant data to ensure the accuracy and reliability of downstream analysis. The specific process is as follows: Use GATK to filter SNP and INDEL sites. The specific criteria are as follows: SNP sites with a quality score (QD) less than 3.0, an FS (Fisher's Strand, Fisher test value) value greater than 60.0, a sequencing alignment quality (MQ) less than 40.0, a sequencing alignment quality rank sum test value (MQRankSum) less than -12.5, a read position rank sum test value (ReadPosRankSum) less than -8.0, and a strand odds ratio (SOR) greater than 3.0 are filtered; INDEL sites with a sequencing depth (DP) less than 10, a sequencing alignment quality rank sum test value (MQRankSum) less than -12.5, and a read position rank sum test value (ReadPosRankSum) less than -8.0 are filtered. To improve the reliability of SNP data, use bcftools to filter out SNP sites within 5 base pairs of an INDEL. These steps ensure that the SNP variant data we use has high quality and reliability, enhancing the credibility of the research results.

[0134] Secondly, use Plink software to perform quality control operations on the genotype data. The quality control criteria are set as follows: the genotype missing rate is less than 10%, the minor allele frequency (MAF) is more than 5%, and the sample missing rate is less than 10%. Finally, the number of gene loci obtained is 9,618,10, including 8,014,850 SNPs and 1,603,254 Indels.

[0135] Based on the above embodiments, before using the autoencoder to extract the important feature information of the second breeding object sample from the genotype data in the phenotype-genome data in this embodiment, it further includes:

[0136] Calculate the linkage disequilibrium degree between gene loci of genomic data in the phenotype-genome data;

[0137] Remove gene loci in the genomic data with a linkage disequilibrium degree greater than the correlation threshold.

[0138] From the perspective of genetic markers, the linkage disequilibrium degree (LD) between a marker and a causal variant has an important impact on prediction accuracy. In the case of limited sample size, the use of WGS data does not improve the accuracy of pig traits. Since the increase in gene loci is random, a large number of LD noise markers without any causal variants are included in WGS data. Under this assumption, when using millions of variants in whole-genome sequencing data (WGS data), most variants with causal variants in low LD should be excluded.

[0139] To remove highly correlated loci to reduce data redundancy and computational burden, the linkage disequilibrium between gene loci was calculated. The parameters were that the window size was selected as 25 and 50, the step size was 5 and 10, and the correlation threshold was 0.3 and 0.5. Finally, according to the random combination of parameters, the optimal parameters suitable for each population were obtained. The population used 412 individuals and 461,357 independent gene loci for genomic selection analysis, and the optimal parameters were 50, 10, and 0.5.

[0140] Using R software, the gene loci obtained by linkage disequilibrium pruning and the growth traits FCR and C_Day_120kg in the phenotypic data were used to obtain phenotypic-genotypic data.

[0141] To improve the efficiency of population breeding with the characteristics of low sample size and ultra-high-dimensional genomic data, this paper introduced a feature dimensionality reduction tool to reduce the dimensionality of the original input features. During the experiment, using an autoencoder, when the dimensionality of the low-dimensional space was set to 2000, 4000, 6000, 8000, 10000, and 12000 feature dimensions, as Figure 3 and 4 shown, in the results of the two traits, when using 8000 gene features, the prediction effect of the EN-CNN model was the best, and the gene selection model achieved relatively stable performance. The important features of feed conversion ratio (FCR) and corrected age at 120 kg body weight (C_Day_120kg) are shown in Table 2 below. Therefore, in subsequent experiments, the number of gene loci of all feature selection models was selected as 8000.

[0142] Table 2 Optimal number of gene features selected for growth traits

[0143]

[0144] Figure 5 The training and validation processes of the gene selection model (GS model) are shown, where the fixed effects are the sex and birth weight of the population.

[0145] Table 3 shows the prediction accuracies of four GS models (SVM, RF, MLP, CNN) and the GBLUP model on the full genomic data of the validation population for feed conversion ratio (FCR) and corrected age at 120 kg body weight (C-Day_120kg). Our research results show that in the two breeding traits, the GS model generally has higher prediction accuracy than the traditional GBLUP model. However, the complexity of the convolution layer calculation and the number of high-dimensional features of the CNN model lead to model failure. Therefore, this problem is solved by selecting a dimensionality reduction tool.

[0146] Table 3 Prediction accuracy of the model based on all SNPs and Indels in the validation population

[0147]

[0148]

[0149] Tables 4 and 5 show the evaluation metrics predicted by the XG-GS model and the EN-GS model in feed conversion ratio (FCR) and age at 120 kg body weight corrected (C-Day_120kg). The results show that the RMSE of the EN-GS model is less than that of the XG-GS model, and the EN-GS model can achieve a smaller prediction loss. The XGBoost model with feature selection and the random forest prediction model (RF) have the best prediction accuracy in both traits; when dealing with the features screened by XGBoost, the random forest model shows good anti-noise ability and robustness, and can produce a synergistic effect, thus performing excellently. In addition, the combination of autoencoder and neural network has an overall better effect than the combination of XGBoost and neural network, and is more stable. The non-linear features extracted by the autoencoder are more suitable for the neural network, especially when dealing with complex genomic data, and can capture those features that are difficult to extract by linear feature selection methods (such as XGBOOST). Therefore, the combination of autoencoder and neural network performs significantly better than the combination of XGBOOST and neural network in growth traits.

[0150] Table 4 Comparison of evaluation metrics of the XG-GS model

[0151]

[0152] Table 5 Comparison of evaluation metrics of the EN-GS model

[0153]

[0154] For the two important breeding goals of pig breeding, FCR and C-Day_120kg, the combination of the encoder (EN) and the genomic selection model (GS) prediction model has an overall good performance in the three performance evaluation metrics, and the values of each performance evaluation metric are relatively more excellent, indicating that the EN-MLP and EN-CNN models can better capture data features and improve prediction accuracy, that is, the combination of several network models based on deep learning can effectively carry out breeding work according to population phenotypic data and high-dimensional genomic information compared with the combination of machine learning models.

[0155] The purpose of calculating genomic estimated breeding values is to select individuals with excellent genetic characteristics to promote the genetic progress of the species. In pig breeding, this can help select individuals with fast growth rates and high fertility. This study will use the EN-GS model to predict the genomic estimated breeding values (GEBVs) calculated by GBLUP, providing a new perspective and higher accuracy in pig breeding, thereby promoting the progress of breeding research.

[0156] The evaluation metrics for the prediction of genomic estimated breeding values (GEBVs) of the XG-GS model in two growth traits are as Figure 6 shown. When using the XG-RF model for prediction, the PCC values of FCR_GEBV and C_Day_120kg_GEBV are 0.747 and 0.648 respectively, ranking first among the XG-GS models. It can be seen that the XG-CNN model has relatively poor prediction performance when predicting genomic estimated breeding values. Evidently, the important gene features screened by the XGBoost model cannot be effectively recognized by the CNN model. Therefore, a feature dimensionality reduction method more suitable for the CNN model is considered.

[0157] Figure 7 The results of the evaluation metrics for the prediction of genomic estimated breeding values (GEBVs) of the EN-GS model in two growth traits are shown. The EN-CNN model performs optimally in all three evaluation metrics for the prediction of genomic estimated breeding values in the two traits. The PCC value of FCR_GEBV is 0.8890, and the RMSE and MAE are 0.016 and 0.013 respectively; the PCC value of C-Day_120kg_GEBV is 0.7790, and the RMSE and MAE are 3.7670 and 3.0150 respectively.

[0158] The contributions of the present invention are summarized as follows:

[0159] (1) A novel CNN-based genomic selection (GS) method called EN-CNN is proposed. This method first trains an auto-encoder, then uses a feature selection module (EN) to capture important feature (SNP and Indel) information in genomic data, and finally uses a designed convolutional neural network model for prediction;

[0160] (2) Comparing four EN-GS models with the GBLUP model, the experimental results in two growth trait datasets show that the GS models have better generalization ability;

[0161] (3) The superiority of EN-CNN is verified by comparison with the XGBoost-GS model;

[0162] (4) The EN-CNN is robust to hyperparameters and the prediction results for different growth traits are more stable. Therefore, it is applicable to the actual pig breeding process.

[0163] The genomic selection breeding system based on the EN-CNN model provided by the present invention will be described below. The genomic selection breeding system based on the EN-CNN model described below can be mutually referred to the genomic selection breeding method based on the EN-CNN model described above.

[0164] As Figure 8 shown, the system includes an extraction module 801, a prediction module 802, and a selection module 803, where:

[0165] The extraction module 801 is used to capture important feature information from the genotype data of the object to be bred by using an autoencoder;

[0166] The prediction module 802 is used to predict the growth traits of the object to be bred according to the important feature information by using a convolutional neural network model;

[0167] The selection module 803 is used to determine whether to select the object to be bred for breeding according to the growth traits of the object to be bred.

[0168] The present invention first trains an autoencoder (Auto-Encoder) as a feature selection module (EN), then uses the feature selection module (EN) to capture important feature information (such as SNPs and Indels) in the genomic data, and preferably uses the designed convolutional neural network model for prediction, so as to improve the accuracy and robustness of genomic selection breeding.

[0169] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A genomic selection breeding method based on the EN-CNN model, characterized in that: include: Use autoencoders to capture important feature information from the genotype data of the breeding object; Predicting the growth traits of the breeding object according to the important characteristic information using a convolutional neural network model; Determine whether to select the object to be bred for breeding based on the growth traits of the object to be bred.

2. The genomic selection breeding method based on the EN-CNN model according to claim 1, characterized in that: The autoencoder includes an encoder and a decoder. Before using the autoencoder to capture important feature information from the genotype data of the breeding object, the autoencoder further includes: Mapping the original genotype data of the first breeding object sample to a latent space representation using the encoder to reduce the dimension of the original genotype data and extract important feature information; Reconstructing original genotype data of the first breeding object sample according to the latent space representation using the decoder; The autoencoder is trained by taking the average error between the original genotype data of the first breeding object sample and its reconstructed data as a loss function.

3. The genomic selection breeding method based on the EN-CNN model according to claim 2, characterized in that: The encoder includes a fully connected layer, a batch normalization layer, an activation function and a Dropout layer which are connected in sequence.

4. The genomic selection breeding method based on the EN-CNN model according to claim 1, characterized in that: The convolutional neural network model includes an input layer, a first convolutional layer, a second convolutional layer, a first pooling layer, a third convolutional layer, a fourth convolutional layer, a second pooling layer, a first fully connected layer, a second fully connected layer and an output layer which are connected in sequence.

5. The genomic selection breeding method based on the EN-CNN model according to claim 1, characterized in that: Before using the convolutional neural network model to predict the growth traits of the breeding object according to the important characteristic information, the method further includes: After performing a quality control operation on the phenotypic data and genotypic data of the second breeding object sample, phenotypic-genomic data are formed, and the phenotypic data are used as actual growth traits; extracting important characteristic information of the second breeding object sample from the genotype data in the phenotype-genomic data using the autoencoder; Obtaining predicted growth traits of the second breeding object sample based on important characteristic information of the second breeding object sample using the convolutional neural network model; The convolutional neural network model is trained using 10-fold cross validation based on the actual growth traits of the predicted growth traits of the second breeding object sample.

6. The genomic selection breeding method based on the EN-CNN model according to claim 5, characterized in that: The phenotypic data include feed-to-meat ratio and age at which the preset weight is reached. The feed-to-meat ratio refers to the amount of feed required for each kilogram of weight gain of the breeding object, and the age at which the preset weight is reached refers to the age required for the breeding object to reach the preset weight under corrected conditions.

7. The genomic selection breeding method based on the EN-CNN model according to claim 5, characterized in that: Perform quality control operations on the phenotypic data and genotypic data of the second breeding object sample, including: Eliminating the phenotypic data of the second breeding object samples that lack birth weight or have abnormal phenotypic data; Using GATK software, filter the SNP sites in the genotype data whose quality score is less than the first preset threshold, FS value is greater than the second preset threshold, sequencing alignment quality is less than the third preset threshold, rank sum test value of sequencing alignment quality is less than the fourth preset threshold, rank sum test value of read position is less than the fifth preset threshold, or chain ratio value is greater than the sixth preset threshold; Using the GATK software, filter the INDEL sites in the genotype data whose sequencing depth is less than the seventh preset threshold, whose rank sum test value of sequencing alignment quality is less than the fourth preset threshold, or whose rank sum test value of read position is less than the fifth preset threshold; Using bcftools software to filter out SNP sites in the genotype data that are located within a preset number of base pairs near the INDEL site; The genotype data are quality controlled using Plink software, so that the missing rate of the genotype data is less than the eighth preset threshold, the minimum allele frequency is greater than the ninth preset threshold, and the sample missing rate is less than the tenth preset threshold.

8. The genomic selection breeding method based on the EN-CNN model according to claim 5, characterized in that: Before using the autoencoder to extract the important characteristic information of the second breeding object sample from the genotype data in the phenotype-genomic data, the method further includes: Calculating the degree of linkage disequilibrium between gene loci of genomic data in the phenotype-genomic data; The gene loci in the genomic data whose linkage disequilibrium degree is greater than the correlation threshold are removed.

9. A genomic selection breeding system based on the EN-CNN model, characterized in that: include: An extraction module, used to capture important feature information from the genotype data of the breeding object using an autoencoder; A prediction module, used for predicting the growth traits of the breeding object according to the important characteristic information using a convolutional neural network model; The selection module is used to determine whether to select the object to be bred for breeding according to the growth characteristics of the object to be bred.

10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the genomic selection breeding method based on the EN-CNN model as described in any one of claims 1 to 8 is implemented.

Citation Information

Patent Citations

  • Corn phenotype prediction method and system

    CN113593635A

  • Machine learning method and system for egg laying character genome selection of laying hens

    CN116072226A

  • Model training method, clustering method, equipment and medium

    CN116522143A

  • Group of genes influencing reproductive performance of sows and screening method thereof

    CN116694776A

  • Rice phenotype prediction method and system based on whole genome selection

    CN118072823A

Cited By

  • Peanut quality character selective breeding method based on whole genome SNP (Single Nucleotide Polymorphism) and application

    CN121171335A

  • Peanut quality trait selection breeding method based on whole genome snp and application

    CN121171335B