Whole genome selection method based on deep learning

By introducing the EBMGP method in genome-wide prediction, combined with Elastic Net feature selection, BERT embedding and multi-head attention pooling, the problems of restricted development of GP models and neglected SNP relationships in the existing technology are solved, and efficient feature streamlining and complex correlation capture are achieved, improving the prediction effect.

CN120015110APending Publication Date: 2025-05-16HUNAN AGRI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510062885.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-15
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The existing technology has the problem that the number of features far exceeds the number of individuals in the whole genome prediction, which leads to the limited development of GP deep learning models; at the same time, the existing model uses one-hot to represent SNPs, ignores the relationship between SNPs, and adopts the traditional maximum pooling and average pooling methods, which may lead to information loss.

Method used

A genome-wide selection method based on deep learning is proposed, combined with Elastic Net for feature selection, conceptualize SNP into a natural language form using BERT embedding, and capture the interaction between SNP and LD block levels through multi-headed attention pooling.

Benefits of technology

EBMGP effectively streamlines feature space, can capture complex associations at the SNP and LD block levels, performs well on small datasets, and is comparable to top pooling methods on large datasets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120015110A_ABST
    Figure CN120015110A_ABST
Patent Text Reader

Abstract

The invention relates to the field of deep learning and animal and plant molecular breeding, and discloses a whole genome selection method EBMGP based on deep learning. According to the method, a feature set is optimized by combining a feature selection scheme based on an elastic network (EN), so that the calculation burden is reduced, and the prediction accuracy is improved. According to the present invention, the SNP is simulated into the structure similar to the human language by using the bidirectional encoder representation (BERT) embedding, such that the complex genetic interaction of the single SNP and the linkage imbalance block level can be effectively captured, and the single SNP and the complex genetic interaction of the linkage imbalance block level can be effectively captured. In addition, the invention further provides a multi-head attention pooling technology, when the features from multiple subspaces are learned, weights can be intelligently distributed to the features, and deep understanding of high-level semantic information is achieved. Compared with seven reference models, the method shows excellent average prediction precision in five different tasks. Therefore, the EBMGP is a promising whole genome selection model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of deep learning and animal and plant molecular breeding, and more specifically, to a whole genome selection method based on deep learning. Background Art

[0002] Genomic Prediction (GP) was first proposed by Meuwissen et al. (2001). It predicts the breeding value of breeding / prediction populations through genome-wide single nucleotide polymorphism (SNP) genotype markers, thereby accelerating the identification of excellent genotypes and promoting the breeding process (Li et al., 2023; Meuwissen et al., 2001). With the sharp decline in the cost of SNP genotyping, GP has been widely used in the past decade and has made significant progress in various plant and animal breeding projects. GP research mainly focuses on optimizing marker density, training population size, kinship, and the selection of GP models. Genomic Best Linear Unbiased Prediction (GBLUP) is a common GP model that predicts the kinship matrix constructed based on marker genotypes. In contrast, the Bayesian model incorporates the prior distribution, and different models are required for different traits. For example, Bayes B uses a Gaussian mixture model, assuming that not all markers contribute to genetic variance (Pérez and de los Campos, 2014). Bayesian Lasso (BL) uses a double exponential prior for continuous shrinkage and variable selection, and uses a long-tailed Student-t distribution to describe marker effects (Li et al., 2010). However, in practice, the exact effect of a single SNP is still difficult to determine, and it does not necessarily conform to a specific distribution. In addition, these parameterized models often fail to capture the complex interactions between SNPs, especially in complex traits caused by gene interactions.

[0003] As a branch of machine learning, deep learning uses complex neural networks with multiple layers of nonlinear transformations, making it very suitable for GP challenges. The DNNGP model integrates three convolutional neural network (CNN) layers, a batch normalization (BN) layer to prevent overfitting, and two dropout layers (Wang et al., 2023). The model efficiently processes complex omics data and outperforms common GP methods such as GBLUP, LightGBM, SVR, DeepGS, and DLGWAS (Wang et al., 2023). SoyDNGP is a deep network with 12 convolutional blocks and one fully connected layer. The coordinate attention (CA) mechanism is introduced after the first and last convolutional layers to enhance spatial information extraction (Gao et al., 2023). In classification tasks, SoyDNGP outperforms AdaBoost, decision trees, naive Bayes, and random forests; in regression tasks, it outperforms DeepGS and DNNGP, demonstrating its versatility and advantages in GP (Gao et al., 2023).

[0004] Although deep learning has made significant progress in GP, ​​there is still a broad space for exploration. First, in the p>>n problem, the number of features (p) far exceeds the number of individuals (n), which has become one of the limiting factors for the development of GP deep learning models. Second, most existing GP deep learning models use one-hot to represent SNPs, treat each SNP independently, and ignore the relationships between them. In addition, this poses a great challenge to the model in identifying the functional (semantic) differences between SNPs with the same genotype. Third, most GP models use traditional maximum pooling and average pooling methods, which may lead to information loss and cannot dynamically optimize features. Summary of the invention

[0005] In order to overcome the above-mentioned defects existing in the prior art, the present invention provides a whole genome selection method EBMGP (Joint Elastic Net feature selection, Bidirectional Encoder Representations from Transformers embedding and Multi-head attention pooling for Genomic Prediction) based on deep learning, which uses Elastic Net to pre-select important SNPs. By using BERT embedding, SNPs are conceptualized as a form similar to human natural language, so that interactions can be dynamically detected at the SNP and linkage disequilibrium (LD) block level. Multi-head attention pooling is proposed, which assigns adaptive weights to features and captures features of different subspaces through multiple heads.

[0006] The above technical objectives of the present invention are achieved through the following technical solutions: A whole genome selection method based on deep learning, comprising the following steps:

[0007] S1. Feature selection:

[0008] Before training, Elastic Net is used to pre-select the top n features in terms of importance to reduce noise and computational cost.

[0009] S2. BERT embedding layer construction:

[0010] After feature selection, BERT embedding is used to represent SNPs. Each SNP is represented by two letters. The first letter corresponds to the genotype of the SNP, where the major allele is marked as "H", the heterozygous state is "M", and the minor allele is "L". The second letter represents the linkage disequilibrium coefficient R between adjacent SNPs. 2 , when the LD value is 0.8 or higher, the second letter is marked as "Y"; conversely, if the LD value is lower than 0.8, it is marked as "J";

[0011] S3, convolution layer construction:

[0012] The convolutional layer consists of 5 Conv-MAP modules, each of which includes batch normalization, convolutional blocks, MAP modules, and Dropout. Three of the five Conv-MAP modules use larger convolution kernels (30), and the other two use smaller convolution kernels (3), which are strategically cross-stacked to effectively capture fine-grained local changes and broader high-level conceptual patterns.

[0013] S4, full connection layer construction;

[0014] S5. Model training:

[0015] EBMGP uses a five-fold crossover method to divide the training set (training population) and the test set (breeding / testing population). The model training hyperparameters are set as follows:

[0016] Batch_size 32 Learning_rate 0.0005 epoch 100 Optimizer AdamW(weight_decay=0.00002) scheduler CosineAnnealingLR(max_epoch=train_epoch)

[0017] Furthermore, the Elastic Net described in step S1 combines L1 and L2 penalty terms, and controls the ratio of the two through parameter adjustment to ensure that the number of non-zero coefficient markers exceeds 500, 3000, 6000, 9000 and 12000 respectively.

[0018] Furthermore, the feature selection described in step S1 is performed only on the training set to avoid artificially improving the prediction accuracy.

[0019] Furthermore, the pooling strategy multi-head attentionpooling described in step S3 specifically includes the following steps:

[0020] (1) Use the expansion function to split the feature into multiple sub-features;

[0021] (2) Each sub-feature passes through multiple 1D convolutions to obtain multi-spatial features, and all multi-spatial features are concatenated to capture a wider range of latent semantic associations; this step is performed only when multi-head is set to 1 or higher; if it is set to 0, this step is skipped;

[0022] These sub-features are weighted and summed through softmax, and the model architecture can be flexibly controlled by modifying multi-head hyperparameters.

[0023] In summary, the present invention has the following beneficial effects:

[0024] (1) This application introduces a deep learning model for whole genome prediction, EBMGP; this model combines EN feature selection, a novel SNP representation method, and a multi-head attention pooling mechanism; through EN feature selection, EBMGP effectively simplifies the feature space.

[0025] (2) This application uses BERT embedding to treat SNPs as natural language, and EBMGP captures complex associations at the SNP and LD block levels. This SNP representation method has demonstrated its effectiveness in various applications of EBMGP and SoyDNGP.

[0026] (3) The multi-head attention pooling method in this application performs well on small datasets and is comparable to the top pooling methods on large datasets.

[0027] (4) The effectiveness of EBMGP in this application has been verified in five prediction tasks on a rice dataset, and we believe it has great potential in promoting data-driven decision-making in animal and plant breeding programs. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] Figure 1 is a flow chart of EBMGP analysis in an embodiment of the present invention;

[0029] Figure 2 It is a diagram of the prediction accuracy of EBMGP and the reference model in five rice traits in the embodiment of the present invention. DETAILED DESCRIPTION

[0030] The following is combined with Figure 1-2 The present invention is described in further detail.

[0031] Example: The rice dataset contains 413 different rice inbred lines from 82 countries. Each plant was genotyped using a 44K chip (44,100 SNPs in total). SNPs with a minor allele frequency (MAF) below 5% were removed. After quality control, 36,901 SNPs were retained. Five traits were evaluated: grain width (SW), flag leaf width (FLW), plant height (PH), amylose content (AC), and seeds per panicle (SNPP).

[0032] In this implementation case, the whole genome prediction is performed using the method of the present invention, and the specific steps are as follows:

[0033] 1) Use EN for feature selection and select the top 6000 SNPs by adjusting the penalty parameter.

[0034] 2) After feature selection, each SNP is represented by two letters. The first letter corresponds to the genotype of the SNP, where the major allele is marked as "H", the heterozygous state is "M", and the minor allele is "L". The second letter represents the linkage disequilibrium (LD) coefficient R between adjacent SNPs. 2 . When the LD value is 0.8 or higher, the second letter is marked as "Y"; conversely, if the LD value is lower than 0.8, it is marked as "J". Therefore, the model can distinguish those SNPs that are adjacent and closely related, namely LD blocks. For the last SNP on each chromosome, the second letter is uniformly "N". The encoded SNPs are used as model input.

[0035] 3) Model training, hyperparameter selection batch_size is 32, learning_rate is 0.0005, epoch50, optimizer is AdamW (weight_decay = 0.00002), and cosine annealing is used to dynamically adjust the learning rate (CosineAnnealingLR (max_epoch = train_epoch).

[0036] 4) Reference model.

[0037] To verify the effectiveness of EBMGP, it is compared with seven frontier models, such as Figure 1As shown, including GBLUP, RKHS, Bayes B, Bayes LASSO, DLGWAS, SoyDNGP and DNNGP. GBLUP was implemented using the R package “rrBLUP” (Endelman 2011). Reproducing kernel Hilbert space (RKHS) was implemented using the R package “BGLR”, which is a semi-parametric method using a Gaussian kernel function (Friedman et al., 2010). The Bayes B model was also implemented using the R package “BGLR”, based on a Monte Carlo–Markov chain (MCMC) strategy, with 12,000 iterations and a burn-in period of 2,000 (Endelman 2011). In addition, Bayes Lasso (BL) was implemented using the R package “glmnet”, using the same MCMC settings (Endelman 2011). These models are able to accurately predict breeding values ​​based on additive effects. DLGWAS and SoyDNGP were implemented in PyTorch 11.6 based on the reference code (Gao et al., 2023; Liu et al., 2019). The DNNGP model was downloaded from https: / / github.com / AIBreeding / DNNGP / releases / download / v1.0.0 / DNNG P-v1.0.0.zip. All models were cross-validated with 5-fold cross validation, and the coefficient of determination (R 2 ) average value is used as the evaluation criterion for the model.

[0038] This specific embodiment is merely an explanation of the present invention and is not a limitation of the present invention. After reading this specification, those skilled in the art may make non-creative modifications to the present embodiment as needed. However, as long as they are within the scope of the claims of the present invention, they are protected by the patent law.

Claims

1. A whole genome selection method based on deep learning, characterized in that: The following steps are involved: S1. Feature selection: Before training, Elastic Net is used to pre-select the top n features in terms of importance to reduce noise and computational cost. S2. BERT embedding layer construction: After feature selection, BERT embedding is used to represent SNPs. Each SNP is represented by two letters. The first letter corresponds to the genotype of the SNP, where the major allele is marked as "H", the heterozygous state is "M", and the minor allele is "L". The second letter represents the linkage disequilibrium coefficient R between adjacent SNPs. 2 , when the LD value is 0.8 or higher, the second letter is marked as "Y"; conversely, if the LD value is lower than 0.8, it is marked as "J"; S3, convolution layer construction: The convolutional layer consists of 5 Conv-MAP modules, all of which use ConvBlock. The 5 Conv-MAP modules are strategically cross-stacked to effectively capture fine-grained local changes and a wide range of high-level features. The pooling strategy multi-head attentionpooling is used to meet the needs of high-level semantic understanding. S4, full connection layer construction; S5. Model training: EBMGP uses the five-fold crossover method to divide the training set and the test set. The model training hyperparameters are set as follows:

2. A whole genome selection method based on deep learning according to claim 1, characterized in that: The Elastic Net described in step S1 combines L1 and L2 penalty terms, and controls the ratio of the two through parameter adjustment to ensure that the number of non-zero coefficient markers exceeds 500, 3000, 6000, 9000 and 12000 respectively.

3. The whole genome selection method based on deep learning according to claim 1, characterized in that: The feature selection described in step S1 is performed only on the training set to avoid artificially improving the prediction accuracy.

4. The whole genome selection method based on deep learning according to claim 1, characterized in that: The pooling strategy multi-head attention pooling described in step S3 specifically includes the following steps: (1) Use the expansion function to split the feature into multiple sub-features; (2) Each sub-feature passes through multiple 1D convolutions to obtain multi-spatial features, and all multi-spatial features are concatenated to capture a wider range of latent semantic associations; this step is performed only when multi-head is set to 1 or higher; if it is set to 0, this step is skipped; (3) These sub-features are weighted and summed through softmax, and the model architecture can be flexibly controlled by modifying multi-head hyperparameters.