A genome selection method based on deep learning reaGP
Through the deep learning reaGP method, a convolutional neural network model is constructed and the residual module and attention mechanism is combined, which solves the shortcomings of linear models fitting nonlinear problems in biological breeding, realizes efficient prediction of genome selection, and improves breeding accuracy and efficiency.
Patent Information
- Application Number
- CN202411437824.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-15
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2044-10-15
AI Technical Summary
Existing linear models are difficult to effectively fit nonlinear problems in biological breeding, especially when dealing with the non-additive effects of complex phenotypic control, which leads to insufficient overfitting and prediction accuracy and cannot meet the needs of biological breeding.
A genome selection method based on deep learning reaGP is adopted, and a convolutional neural network model is constructed, combining genomic data and frequency information, a residual module and SE attention mechanism are added, and cross-validation is used to improve prediction accuracy.
It significantly improves the prediction accuracy of genome selection, improves 0.5% to 18%, shortens the breeding generation interval, improves breeding efficiency, and breaks through the limitations of relying solely on genomic and phenotypic data.
Smart Images

Figure CN119360956B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of genomic selection, and particularly relates to a genomic selection method based on deep learning reaGP. Background Art
[0002] As a non-linear model, deep learning can obtain better results through simple fine-tuning of parameters and has achieved good results in image processing tasks. In the primary stage of biological breeding, when using a linear model to predict phenotypes, there is no need to establish a model like a linear model and then solve gebv. When facing non-linear problems, it is impossible to better fit the equation based on a linear model. In addition, a linear model cannot better solve the problem of p>n, is prone to overfitting, and has some non-additive effects (dominance, epistasis, etc.) in complex phenotype control. A deep learning model can better capture these performances, which brings more advantages to animal and plant breeding. For some biological populations, economic traits are complex phenotypes controlled by non-additive effects, and the deep learning method may bring infinite possibilities for improving the accuracy of biological breeding value estimation. Summary of the Invention
[0003] The purpose of the present invention is to provide a genomic selection method based on deep learning reaGP to solve the problems in the above background.
[0004] The purpose of the present invention can be achieved by the following technical solutions:
[0005] A genomic selection method based on deep learning reaGP includes the following steps:
[0006] Step 1: Construct a reference population and measure the economic traits of the population;
[0007] Step 2: Collect blood samples from the reference population for preservation, extract DNA and perform genotyping using a biochip, and process and quality control the data after genotyping;
[0008] Step 3: Convert the genomic data (0, 1, 2) into 0, 0, 0, 1, 1, 1; calculate the frequency information of the three allele genomes; combine the genomic data and the frequency information into three-channel data, and the three-channel data is used to provide richer feature information for the neural network and serve as the input data of the neural network;
[0009] Step 4: Build a convolutional neural network model, and the network structure of the convolutional neural network model is two convolutional layers and two fully connected layers. Among them, the number of nodes is 64 and 1, the size of each convolutional kernel is 7*7 and 5*5 respectively, and the number of channels in the convolutional layers is 64;
[0010] Step 5: Use cross-validation to divide the training population into a training set and a test set. Use the trained model to evaluate the test set data, and the predicted value (y) in the test set data is used as the genomic estimated breeding value (GEBV) estimated by deep learning prediction. Calculate the Pearson correlation between the y value and the true phenotype, and use the correlation coefficient as the evaluation criterion for accuracy.
[0011] As a further aspect of the present invention: The economic traits of the population include three traits in the fattening traits, namely carcass weight (WWT), weaning weight (CWT), and daily weight gain (FDG).
[0012] As a further aspect of the present invention: In step 2, quality control is to eliminate unqualified individuals and SNPs.
[0013] The biochip is an Illumina BovineHD high-density chip.
[0014] As a further aspect of the present invention: The data form of the neural network input data in step 3 is: a: 0, 0, P(aa); a: 1, 0, P(Aa); a: 1, 1, P(AA).
[0015] As a further aspect of the present invention: In step 3, calculating allele frequency information is to calculate the frequency of the corresponding allele in each marker in all samples; among them, the allele frequency information is used to provide an explanatory perspective from a biological point of view for the three-channel input information of the neural network.
[0016] Specifically, it includes:
[0017] Assume that for a specific genetic marker, the genotypes AA and Aa appear D times and H times respectively; if the population size is N, the total number of alleles will be 2N, then the frequency P of allele A is:
[0018] The frequency of allele a is q = 1 - p. After calculating the allele frequencies of the genetic markers, the genotype frequencies of the genotypes AA, Aa, and aa of a specific marker are obtained from p2, 2pq, and q2.
[0019] As a further aspect of the present invention: In step 3, the three-channel data set is used to expand the ordinary one-dimensional genomic data into a feature dimension with richness, and the input format of each sample changes from (1, n) to (3, M, M).
[0020] The calculation formula is as follows: where n is the number of markers of the genomic data, and M is the side length of the three-way data. The calculation formula is as follows:
[0021] P L ,P RAs the number of symmetrically padded 0s, the calculation formula is as follows:
[0022] P R = P L + 1.
[0023] As a further solution of the present invention: in the fourth step, a residual module is added between the fully connected layer and the second convolutional layer; the residual module is used to prevent the problem of gradient disappearance or explosion during the training of the deep network.
[0024] As a further solution of the present invention: the structure of the residual module is two convolutional kernels of size 3*3, and a 1*1 convolutional layer is added to make the number of input channels correspond to the number of output channels to achieve information transmission;
[0025] The residual module has two convolutional layers with stride = 1, and on this basis, an identity mapping layer with stride = 1 is added to ensure that the input dimension and the output dimension are consistent;
[0026] An SE attention mechanism is added between the two convolutional layers and after the residual module; it is used to enhance the recognition ability of important traits related to specific genotypes.
[0027] As a further solution of the present invention: in the fifth step, according to the phenotype of each trait and the three-channel data, the predicted result y value is obtained through the network model, where y WWT is the genomic estimated breeding value of weaning weight, y FDG is the genomic estimated breeding value of daily weight gain during the fattening period, y CWT is the genomic estimated breeding value of carcass weight.
[0028] Advantages of the present invention:
[0029] By constructing a deep learning model (reaGP), compared with linear models, existing machine learning models, and deep learning models, the overall improvement in three traits is 0.5% - 18%; compared with traditional genomic selection methods: improving the prediction accuracy of complex traits;
[0030] Traditional methods have limitations in dealing with complex traits and non-linear problems, while deep learning models can significantly improve these problems and enhance the prediction accuracy and stability through the advantages of multi-layer neural networks;
[0031] By fine-tuning the model parameters, the prediction performance can be further improved to ensure good results under different traits and conditions.
[0032] This method not only utilizes genomic and phenotypic data, but also combines other biological information as auxiliary variables, significantly improving the estimation accuracy of genomic selection breeding values and breaking through the limitations of relying solely on genomic and phenotypic data.
[0033] Compared with traditional breeding methods, deep learning models can screen and predict excellent individuals more quickly, shorten the generation interval, improve the breeding efficiency, and ultimately accelerate genetic progress. Brief Description of the Drawings
[0034] The present invention will be further described below with reference to the accompanying drawings.
[0035] Figure 1 It is a schematic flowchart of a genomic selection method for deep learning reaGP of the present invention;
[0036] Figure 2 It is a schematic diagram of the three-channel form of the data of reaGP in the present invention;
[0037] Figure 3 It is a schematic diagram of the model structure of reaGP in the present invention;
[0038] Figure 4 It is a schematic diagram of the residual structure of reaGP in the present invention;
[0039] Figure 5 It is a schematic diagram comparing the genomic prediction accuracies of three traits of West China cattle in the present invention. Detailed Embodiments
[0040] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0041] Embodiment 1
[0042] A genomic selection method based on deep learning reaGP includes the following steps:
[0043] Step 1: Construct a genomic reference population and measure the economic traits of the population; the measured economic traits include three economic traits of fattening traits, specifically: daily weight gain FDG, carcass weight CWT, and weaning weight WWT;
[0044] Step 2: Calibrate the phenotypes of the economic traits of the selected population and perform quality control on the genomic data;
[0045] Step 3: Combine the genomic data with the frequency encoding to form a 3D data structure for input into the convolutional neural network;
[0046] Step 4: Build a convolutional neural network with residual units and attention mechanism to stabilize the gradient and enhance feature extraction capabilities;
[0047] Step 5: Use the above convolutional neural network to select the genomic breeding value of the population.
[0048] Embodiment 2
[0049] See also Figure 1 - Figure 4 As shown, the present invention is a genome selection method based on deep learning reaGP, comprising the following steps:
[0050] Step 1: Construct a reference population and measure the economic traits of the population, including three fattening traits, namely, keto body weight WWT, weaning weight CWT, and daily weight gain FDG;
[0051] Step 2: Collect blood from the reference group, extract DNA, perform genotyping using biochips, and process and quality control the genotyping data;
[0052] Among them, quality control was to eliminate unqualified individuals and SNPs, and the biochip was the Illumina BovineHD high-density chip;
[0053] Step 3: Convert the genomic data (0, 1, 2) to 0, 0, 0, 1, 1, 1; calculate the frequency information of the three allele groups; combine the genomic data and frequency information to convert them into three-channel data, which is used to provide richer feature information for the neural network and serve as the input of the neural network; the data format is: a: 0, 0, P(aa), a: 1, 0, P(Aa), a: 1, 1, P(AA);
[0054] Calculating the allele frequency information is to calculate the frequency of the corresponding allele in each marker in all samples; the allele frequency information is used to provide biological explanation for the three-channel input information of the neural network; the calculation formula is as follows:
[0055] We apply Hardy-Weinberg equilibrium and assume that for a particular genetic marker, genotypes AA and Aa appear D and H times respectively; if the population size is N, the total number of alleles will be 2N, then the frequency P of allele A is:
[0056] The frequency of allele a is q = 1 - p. After calculating the allele frequencies of the genetic markers, the genotype frequencies of the specific markers AA, Aa, and aa are obtained from p2, 2pq, and q2;
[0057] The three-channel dataset is used to expand the ordinary one-dimensional genomic data into the feature dimension of richness. The input format of each sample changes from (1, N) to (3, M, M);
[0058] The calculation formula is as follows: where N is the number of markers of the genomic data, and M is the side length of the three-way data. The calculation formula is as follows: That is, assuming that the dimension N of the genomic data can directly take the complete square root, then M is the side length of the three-channel data; otherwise, first use 0 for symmetric padding to make the genomic data dimension become N + P L +P R , satisfying the condition of taking the complete square root, then M is the side length of the three-channel data.
[0059] P L , P R , where P is the number of 0s for symmetric padding. The calculation formula is as follows:
[0060] P R = P L + 1, where P L is the 0 padding at the head position of the one-dimensional genomic data N, and P R is the 0 padding at the end position of the one-dimensional genomic data N.
[0061] Step 4: Build a convolutional neural network model. The network structure of the convolutional neural network model is two convolutional layers with stride = 2, and two fully connected layers. Among them, the number of nodes is 64 and 1, and the size of the convolutional kernel of each layer is 7*7 and 5*5 respectively. The number of channels of the convolutional layer is 64; among them, after each convolutional layer, a BN layer, a Droput layer, and a ReLU activation function layer are added.
[0062] To alleviate gradient explosion or gradient disappearance, a residual module is added between the fully connected layer and the second convolutional layer; the residual module is used to prevent gradient disappearance or explosion problems during the training of deep networks;
[0063] Among them, the structure of the residual module is two convolutional kernels with a size of 3*3. Among them, after each convolutional layer, a BN layer, a Droput layer, and a ReLU activation function layer are added.
[0064] And a convolutional layer of 1*1 is added to make the input channel number and the output channel number correspond. After the convolutional layer, a BN layer, a Droput layer, and a ReLU activation function layer are added to realize information transmission.
[0065] To enhance the ability to mine feature information, an SE attention mechanism is added between two convolutional layers and after the residual module; the attention mechanism is used to enhance the ability to identify important traits related to specific genotypes;
[0066] The residual module consists of two convolutional layers with stride = 1. On this basis, an identity mapping layer with stride = 1 is added to ensure that the input dimension is consistent with the output dimension;
[0067] The specific formula for the residual is as follows:
[0068] H(X) = F(X) + 0;
[0069] The output of the residual network H(X) consists of two parts, including the two residual layers F(X) built in Step 4 and the identity mapping 0. When backpropagation is performed to calculate the gradient, due to activation functions such as sigmoid or Tanh in the network layer, as the number of network layers deepens, continuous derivation may lead to the gradient of the entire network causing gradient disappearance. By introducing the identity mapping x, gradient disappearance is prevented.
[0070] The network gradient at this time is expressed as
[0071]
[0072] where X in F(X) is the output of the convolutional neural network when adding the residual layer, θ is the continuously updated weight parameter, and α is the fixed learning rate.
[0073] Step 5: Use cross - validation to divide the training population into a training set and a test set, and use the trained model to evaluate the test set data. The specific formula is y = f(x), where f(*) is the optimal weight of model training, x is the input data, and the predicted value y is obtained by accumulating x and the internal weights. The predicted value (y) in the test set data is used as the genomic estimated breeding value (GEBV) estimated by deep learning; Calculate the Pearson correlation between the y value and the true phenotype, and the correlation coefficient is used as the evaluation criterion for accuracy;
[0074] The predicted value is equivalent to the genomic estimated breeding value (GEBV) obtained by solving based on a linear model;
[0075] Based on the phenotype of each trait and the three - channel data, the predicted result y value is obtained through the network model, where y WWT is the genomic estimated breeding value for weaning weight, y FDG is the genomic estimated breeding value for daily weight gain during the fattening period, y CWT is the genomic estimated breeding value for carcass weight.
[0076] By constructing a deep learning model (reaGP), compared with linear models, existing machine learning models, and deep learning models, the overall improvement in three traits is 0.5% - 18%; compared with traditional genomic selection methods, it improves the prediction accuracy of complex traits;
[0077] Traditional methods have limitations in dealing with complex traits and non-linear problems. However, deep learning models, through the advantages of multi-layer neural networks, can significantly improve these problems and enhance the accuracy and stability of predictions;
[0078] By fine-tuning the model parameters, the prediction performance can be further improved to ensure good results under different traits and conditions.
[0079] This method not only utilizes genomic and phenotypic data but also combines other biological information as auxiliary variables, significantly improving the estimation accuracy of genomic selection breeding values and breaking through the limitations of relying solely on genomic and phenotypic data.
[0080] Compared with traditional breeding methods, deep learning models can screen and predict excellent individuals more quickly, shorten the generation interval, improve the breeding efficiency, and ultimately accelerate genetic progress.
[0081] Example 3
[0082] Refer to Figure 5 As shown, based on the above technical solution, in this example, genomic prediction of three traits is carried out in the Huaxi cattle population. Specifically:
[0083] Step 1: Measure the traits of Huaxi cattle, including three fattening traits, namely daily weight gain (FDG), carcass weight (CWT), and weaning weight (WWT);
[0084] Huaxi cattle are weaned at 6 months old. After weaning, they are fattened under consistent feeding and management conditions, and the growth traits of the animals are measured every six or twelve months until slaughter; the live weight is recorded after a 24-hour fasting period; Huaxi cattle are slaughtered at 18 to 24 months old;
[0085] Step 2: Collect blood samples from each cow in the reference population for preservation, extract DNA, perform genotyping using the Illumina BovineHD (770k) high-density chip, and use the plinkV1.90 software to process and quality control the genotyped data. The number of three types of quality-controlled genomic data is 1299, 1362, and 1303 for 550K;
[0086] Step 3: Convert the genomic data and frequency information into a three-channel data format. Specifically, first calculate the allele frequencies of AA, Aa, and aa in 550K. Secondly, encode the three alleles: AA: 1, 1, P(AA); Aa: 1, 0, P(Aa); aa: 0, 0, P(aa). Secondly, determine whether 550K can be square-rooted. 574547≈757. (It cannot be completely square-rooted). At this time, the length of the one-dimensional genomic data is (757 + 1)×(757 + 1) = 574564, and the side length M of the three channels is 757. The dimension is 17 more than before. At this time, symmetric 0-padding is performed on the original one-dimensional genome. Zero-padding P L is rounded down to 8 for the left zero-padding at the beginning of the genome; zero-padding P R is 17 - 8 = 9 for the right zero-padding at the end of the genome; finally, a three-channel data with the shape of (3, M, M) is obtained as the input X of the neural network. Build a neural network based on the residual and channel attention mechanisms. And use the three-channel data and the phenotype y as the inputs of the deep learning model for model training;
[0087] Step 4: Use ten-fold cross-validation to divide the training dataset into a training set and a test set, and use the trained model to predict the test set data; calculate the predicted value GEBV and the true phenotype y * of the Pearson correlation coefficient, specifically:
[0088]
[0089] Step 5: Repeat Step 4 ten times, and take the average of the results;
[0090] Step 6: Use the genomic data and phenotype data as the input data for the other five models, namely DNNGP, SVR, RKHS, BayesB, and GBLUP; repeat the steps of Step 4 and Step 5 to calculate the Pearson correlation coefficients of the five methods respectively.
[0091] Figure 4 is a comparison of the Pearson correlation coefficients of the six methods under ten-fold cross-validation. It can be seen from the box plot that reaGP has the highest average value and lower standard deviation among the three traits, and the model is more stable.
[0092] An efficient and high-quality breeding method for Huaxi cattle based on deep learning genomic selection technology. This method has improved in prediction accuracy compared with other linear models. It provides a new way for Huaxi cattle in genomic selection prediction accuracy;
[0093] Through the application of this embodiment, the breeding efficiency and accuracy of Huaxi cattle can be significantly improved, and the improvement and optimization of Huaxi cattle population can be promoted, which has broad application prospects and economic value. This method can quickly estimate the genome breeding value of Huaxi cattle in my country, which has great application value and promotion prospects for beef cattle breeding in my country.
[0094] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modification, equivalent substitution, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention. Any reference numeral in a claim shall not be regarded as limiting the claim involved.
[0095] In addition, it should be understood that although the present specification is described according to implementation modes, not every implementation mode contains only one independent technical solution. This description of the specification is only for the sake of clarity. Those skilled in the art should regard the specification as a whole. The technical solutions in each embodiment may also be appropriately combined to form other implementation modes that can be understood by those skilled in the art.
[0096] The above is a detailed description of an embodiment of the present invention, but the content is only a preferred embodiment of the present invention and cannot be considered to limit the scope of implementation of the present invention. All equivalent changes and improvements made within the scope of the present invention should still fall within the scope of the patent coverage of the present invention.
Claims
1. A genomic selection method based on deep learning reaGP, characterized in that The following steps are involved: Step 1: Construct a reference population and measure the economic traits of the population; Step 2: Collect blood from the reference group, extract DNA, perform genotyping using biochips, and process and quality control the genotyping data; Step 3: Convert the genome data (0, 1, 2) to 0, 0, 0, 1, 1, 1; calculate the frequency information of the three allele groups; combine the genome data and the frequency information to convert them into three-channel data, which is used to provide richer feature information for the neural network and serve as the input data of the neural network; The data format of the neural network input data in step 3 is: a: 0, 0, P (aa), a: 1, 0, P (Aa), a: 1, 1, P (AA); In step 3, the three-channel data set is used to expand the common one-dimensional genome data into the feature dimension of richness, and the input format of each sample is changed from (1, n) to (3, M, M); The calculation formula is as follows: where n is the number of markers of genomic data and M is the side length of the three-way data, the calculation formula is as follows: P L , P R As the number of symmetrically filled 0s, the calculation formula is as follows: P R = P L + 1; Step 4: Build a convolutional neural network model. The network structure of the convolutional neural network model is two layers of convolution and two fully connected layers. The number of nodes is 64 and 1. The size of each convolution kernel is 7*7 and 5*5 respectively, and the number of channels in the convolution layer is 64. Step 5: Use cross-validation to divide the training population into training set and test set, use the trained model to evaluate the test set data, and use the predicted value y in the test set data as the genomic breeding value estimated by deep learning GEBV; use the y value and the true phenotype for Pearson correlation, and use the correlation coefficient as the accuracy evaluation criterion.
2. The genomic selection method based on deep learning reaGP according to claim 1, wherein The economic traits of the group include: three types of fattening traits, namely, keto body weight WWT, weaning weight CWT, and daily weight gain FDG.
3. A genomic selection method based on deep learning reaGP according to claim 1, characterized in that, In step 2, quality control is to eliminate unqualified individuals and SNPs; The biochip is an Illumina BovineHD high-density chip.
4. A genomic selection method based on deep learning reaGP according to claim 1, characterized in that, In the step 3, calculating the allele frequency information is to calculate the frequency of the allele corresponding to each marker in all samples; wherein the allele frequency information is used to provide biological explanation for the three-channel input information of the neural network; Specifically include: Suppose for a particular genetic marker, the genotypes AA and Aa occur D and H times respectively; if the population size is N, the total number of alleles will be 2N, then the frequency P of allele A is: The frequency of allele a is q=1-p. After calculating the allele frequencies of the genetic markers, the genotype frequencies of genotypes AA, Aa and aa of the specific marker are obtained from p2, 2pq and q2.
5. A genomic selection method based on deep learning reaGP according to claim 1, characterized in that, In the step 4, a residual module is added between the fully connected layer and the second convolutional layer; the residual module is used to prevent the gradient from disappearing or exploding during the deep network training process.
6. The genomic selection method based on deep learning reaGP according to claim 5, wherein, The structure of the residual module is two layers of 3*3 convolution kernels, and a 1*1 convolution layer is added to make the number of input channels correspond to the number of output channels to achieve information transmission; The residual module has two stride=1 convolutional layers, on top of which an identity mapping layer with stride=1 is added to ensure that the input dimension and output dimension are consistent; The SE attention mechanism is added between the two convolutional layers and after the residual module; Used to enhance the identification of important traits associated with specific genotypes.
7. A genomic selection method based on deep learning reaGP according to claim 1, characterized in that, In the fifth step, the predicted result y value is obtained through the network model based on the phenotypes of each trait and the three-channel data, where y WWT is the genomic estimated breeding value of weaning weight, and y FDG is the genomic estimated breeding value of daily weight gain during the fattening period, and y CWT is the genomic estimated breeding value of carcass weight.
Citation Information
Patent Citations
Western China cattle genome selection method
CN111243667A
Deep learning algorithm of whole genome prediction model
CN117877587A