Transformer-based crop phenotype prediction method
Patent Information
- Application Number
- PCT/CN2025/087692
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-03-26
- Filing Date
- 2025-04-08
- Publication Date
- 2026-10-01
Smart Images

Figure PCTCN2025087692-FTAPPB-I100001 
Figure PCTCN2025087692-FTAPPB-I100002 
Figure PCTCN2025087692-FTAPPB-I100003
Abstract
Description
A Transformer-based crop phenotypic prediction method
[0001] Citation of relevant applications
[0002] This application claims priority to Chinese Patent Application No. 2025103671561, filed on March 26, 2025, entitled "A Crop Phenotype Prediction Method Based on Transformer", the entire contents of which are incorporated herein by reference. Technical Field
[0003] This invention belongs to the field of agricultural intelligent analysis technology, specifically relating to a crop phenotypic prediction method based on Transformer. Background Technology
[0004] With the rapid development of genomics technologies, single nucleotide polymorphism (SNP) data has provided a new tool for crop phenotypic prediction. SNP data can reflect genetic variation in the crop genome, providing important information for elucidating the genetic basis of complex traits. However, traditional linear models have limitations when processing high-dimensional SNP data, failing to effectively capture complex nonlinear relationships, resulting in insufficient prediction accuracy. Furthermore, crop quality and yield-related traits are often jointly regulated by multiple genes, and these genes have complex interactions, further increasing the difficulty of prediction.
[0005] In recent years, the application of deep learning and artificial intelligence technologies in agriculture has gradually increased, especially the successful application of the Transformer architecture in natural language processing and image recognition, which has provided new ideas for crop phenotypic prediction. Transformer models, through their self-attention mechanism, can effectively capture long-range dependencies in data and are suitable for processing high-dimensional and complex genomic data. Compared with traditional linear models, deep learning-based models can better handle nonlinear relationships and improve prediction accuracy. Therefore, it is necessary to develop a method that can accurately predict complex traits controlled by multiple genes. Summary of the Invention
[0006] This invention aims to solve the problem that traditional models cannot accurately predict complex traits controlled by multiple genes in crops.
[0007] To address the above problems, this invention provides a crop phenotypic prediction method based on Transformer, the method comprising:
[0008] a) Construct a crop phenotypic prediction model; the model consists of a pre-convolutional block, N layers of Transformer blocks, and terminal convolutional blocks; the input of the model is the SNP data of crop samples, and the output is the predicted phenotypic trait value of crop samples; the SNP data includes the SNP genotype matrix and SNP weight data;
[0009] b) The model is trained using SNPs data from multiple crop samples and their corresponding phenotypic trait measurements as a training set to obtain the trained model;
[0010] c) Use the trained model to predict the phenotype of the crop sample based on the SNPs data of the crop sample.
[0011] Furthermore, the SNP genotype matrix is a two-dimensional matrix A obtained by arranging the genotype values of each SNP locus in the crop sample with the crop sample and the SNP locus as rows and columns. t×s The SNP weight data is a one-dimensional vector W corresponding to the weight score of each SNP locus in the SNP genotype matrix. s Where t represents the number of crop samples and s represents the number of SNP loci.
[0012] Furthermore, the genomic data of each crop sample is compared with the reference genome, and for each SNP locus, the genotype value of each SNP locus in each sample is obtained based on the consistency between the genomic data and the reference genome.
[0013] Furthermore, genome-wide association analysis was performed on the phenotypic traits of the multiple crop samples. Based on the results of the genome-wide association analysis, the weight scores of each SNP locus were extracted and organized into a genotype matrix A corresponding to the SNPs. t×s The one-dimensional vector W of each SNP site s .
[0014] Furthermore, in the pre-convolutional block of the model, A t×s and W s The SNPs are transformed into a four-dimensional tensor and a two-dimensional matrix, respectively. The four-dimensional tensor and the two-dimensional matrix are then subjected to Hadamard multiplication to obtain SNPs data containing weight information. The SNPs data containing weight information are subjected to a two-dimensional convolution and regularized using BatchNorm2d to obtain L1. L1 is then subjected to m layers of two-dimensional convolution, Sigmoid activation function and BatchNorm2d regularization to obtain L2, where m is an integer greater than or equal to 1.
[0015] Further, after L2 enters the N-layer Transformer block of the model, in the first layer, L2 is subjected to layer regularization to obtain T1; after T1 is processed by the multi-head attention mechanism, it is joined with T1 by an additive residual connection to obtain T2; on the main branch, T2 is subjected to LayerNorm, linear transformation, Sigmoid mapping, and another linear transformation, and on the secondary branch, T2 is subjected to linear transformation and Sigmoid mapping to obtain the residual connection T3; in each subsequent layer, the same operation as the first layer is performed, and the input of each layer is provided by the output of the previous layer to obtain the output T of the N-layer Transformer block. 3N , where N is an integer greater than or equal to 1.
[0016] Furthermore, the T 3N After entering the terminal convolutional block of the model, for T 3N LayerNorm processing is applied, followed by feature extraction using a single-layer neural network with Sigmoid activation, yielding E1. E1 is then processed with k layers of one-dimensional convolution and Sigmoid activation to output the predicted phenotypic value P for the sample. t , where k is an integer greater than or equal to 1.
[0017] Further, step b) includes: calculating the loss function results of the phenotypic trait measurements and phenotypic trait predictions of the plurality of crop samples, and training the model until the loss function reaches a stable state, thereby obtaining the trained model.
[0018] Furthermore, the crops include soybeans and corn; the phenotypic traits include at least one of protein content, oil content, pollen shedding period, ear height, ear length, and number of ear rows.
[0019] The present invention also provides a system for predicting crop phenotypes, the system comprising a phenotype prediction model construction module, an SNPs data acquisition module, a phenotypic trait measurement value acquisition module, and a model training module;
[0020] The phenotypic prediction model construction module is used to construct the crop phenotypic prediction model;
[0021] The SNPs data acquisition module is used to acquire SNPs data from multiple crop samples;
[0022] The phenotypic trait measurement module is used to acquire the phenotypic trait measurement values of the multiple crop samples;
[0023] The model training module is used to train the crop phenotypic prediction model using the SNPs data of the multiple crop samples and their corresponding phenotypic trait measurements as a training set, so as to obtain the trained model.
[0024] This invention employs an improved Transformer architecture and feature engineering techniques, integrating SNP site weight information to enable the constructed model to maintain stable and high-precision prediction capabilities for crop phenotypic traits under different environments. Compared with the traditional linear model rrBLUP, the model of this invention not only improves the prediction accuracy of complex traits controlled by multiple genes, but also possesses good generalization ability and has high practical application value. Detailed Implementation
[0025] The present invention will be further described in detail below with reference to embodiments. It should be understood that the following embodiments are for explanation and illustration only and do not limit the scope of the invention in any way. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.
[0026] This invention provides a crop phenotypic prediction method based on Transformer, comprising: a) constructing a crop phenotypic prediction model; the model consists of a pre-convolutional block, N layers of Transformer blocks, and terminal convolutional blocks; the input of the model is the SNPs data of crop samples, and the output is the predicted phenotypic trait value of the crop samples; the SNPs data includes an SNPs genotype matrix and SNPs weight data; b) training the model using the SNPs data of multiple crop samples and their corresponding phenotypic trait measurements as a training set to obtain a trained model; c) using the trained model to predict the phenotype of the crop sample to be tested based on the SNPs data of the crop sample to be tested.
[0027] In some embodiments, the SNP genotype matrix is a two-dimensional matrix A obtained by arranging the genotype values of each SNP locus in the crop sample in rows and columns, with the crop sample and the SNP locus as the rows and columns. t×s The SNP weight data is a one-dimensional vector W corresponding to the weight score of each SNP locus in the SNP genotype matrix. s Where t represents the number of crop samples and s represents the number of SNP loci.
[0028] In some embodiments, the genomic data of each crop sample is compared with a reference genome. For each SNP locus, a homozygous genotype consistent with the reference genome is assigned a value of 1, a homozygous genotype different from the reference genome is assigned a value of -1, and a heterozygous genotype is assigned a value of 0, thereby obtaining the genotype assignment for each SNP locus of each crop sample.
[0029] In some embodiments, genome-wide association analysis (GWAS) is performed on the phenotypic traits of the multiple crop samples. Based on the GWAS results, the weight scores of each SNP locus are extracted and organized into a corresponding SNP genotype matrix A. t×sThe one-dimensional vector W of each SNP site s .
[0030] In some embodiments, A is placed in the first convolutional block of the model. t×s and W s Transform them into four-dimensional tensors A′ respectively t×1×r×u and two-dimensional matrix W r×u , where s = r × u; A' t×1×r×u and W r×u Perform Hadamard multiplication to obtain SNP data W' containing weight information. t×1×r×u , i.e. W' t×1×r×u =A' t×1×r×u ⊙W r×u ; To W' t×1×r×u Perform a 2D convolution and regularize it using BatchNorm2d to obtain L1, where L1 = BatchNorm2d(Conv2d(W' t×1×r×u L1 is processed with m layers of 2D convolution, Sigmoid activation function, and BatchNorm2d regularization to obtain L2. Where m is an integer greater than or equal to 1.
[0031] In some embodiments, after L2 enters the N-layer Transformer block of the model, in the first layer, L2 is subjected to layer regularization to obtain T1, T1 = LayerNorm(L2); after T1 is processed by the multi-head attention mechanism, it is added to T1 and residually connected to obtain T2, T2 = T1 + MultiHeadAttention(T1); on the main branch, T2 is subjected to LayerNorm, linear transformation, Sigmoid mapping and another linear transformation, and on the secondary branch, T2 is subjected to linear transformation and Sigmoid mapping to obtain residual connection T3, T3 = Linear(Sigmoid(Linear(LayerNorm(T2)))) + Sigmoid(Linear(T2)); in each subsequent layer, the same operation as the first layer is performed, and the input of each layer is provided by the output of the previous layer to obtain the output T of the N-layer Transformer block. 3N , where N is an integer greater than or equal to 1.
[0032] In some embodiments, the T 3N After entering the terminal convolutional block of the model, for T 3N Perform LayerNorm processing and feature extraction using a neural network with a Sigmoid activation function to obtain E1, where E1 = Sigmoid(Linear(LayerNorm(T)).3N ))); Process E1 with k layers of one-dimensional convolution and the Sigmoid activation function to output the phenotypic prediction value P of the sample. t , Where k is an integer greater than or equal to 1.
[0033] In some embodiments, step b) includes: calculating the loss function results of the phenotypic trait measurements and phenotypic trait predictions of the plurality of crop samples, and training the model until the loss function reaches a stable state, thereby obtaining the trained model.
[0034] In some embodiments, the training process of the model is accelerated by using a backpropagation algorithm and a method for optimizing model learning parameters.
[0035] In some embodiments, SNPs data and their corresponding phenotypic trait measurements from multiple crop samples outside the training set are used as a test set to test the performance of the trained model.
[0036] In some embodiments, the crop is soybean or corn.
[0037] In some embodiments, the phenotypic traits include at least one of protein content, oil content, pollen shedding period, ear height, ear length, and number of ear rows.
[0038] The present invention also provides a system for predicting crop phenotypes, which includes a phenotype prediction model construction module, an SNPs data acquisition module, a phenotypic trait measurement value acquisition module, and a model training module;
[0039] The phenotypic prediction model construction module is used to construct the crop phenotypic prediction model;
[0040] The SNPs data acquisition module is used to acquire SNPs data from multiple crop samples;
[0041] The phenotypic trait measurement module is used to acquire the phenotypic trait measurement values of the multiple crop samples;
[0042] The model training module is used to train the crop phenotypic prediction model using the SNPs data of the multiple crop samples and their corresponding phenotypic trait measurements as a training set, so as to obtain the trained model.
[0043] To verify the effectiveness of the crop phenotypic prediction method of the present invention, the present invention also provides the following examples.
[0044] Example 1
[0045] This example uses genomic and phenotypic data from 1553 soybean samples. The genomic data of the soybean samples were obtained from Illumina Hiseq X high-throughput sequencing of soybean samples (Li YH, Qin C, Wang L, et al. Genome-wide signatures of the geographic expansion and breeding of soybean[J]. Science China Life Sciences, 2023, 66(2):350-365.). The genomic data of each soybean sample was compared with the soybean Williams 82 reference genome (Williams 82 V2.1, https: / / phytozome-next.jgi.doe.gov / info / Gmax_Wm82_a2_v1). To conduct further research, genotype data were quality controlled according to the following criteria: removal of gene missing data (gene status . / .) greater than 80%, heterozygosity 0.348, deletion rate 0.2, and minimum gene frequency 0.01. A total of 826,020 high-quality SNPs were obtained (Li et al., 2023). The oil and protein content of soybean seeds were determined using a Fourier transform near-infrared spectrometer (MATRIX-1, Bruker, Germany) and its built-in multi-seed near-infrared model. Soybean seeds were evenly filled to 2 / 3 of the sample chamber to prevent near-infrared light leakage. All harvested materials were analyzed using the soybean multi-seed near-infrared model, with each material measured three times. The final average value was used as the phenotypic data for the soybean samples.
[0046] 1. Obtaining the SNP genotype matrix
[0047] Each SNP locus is assigned a genotype as follows: homozygous (0 / 0 or 1 / 1), heterozygous (0 / 1), or unknown (. / .). Homozygous genotypes consistent with the reference genome (1 / 1) are assigned a value of 1, homozygous genotypes different from the reference genome (0 / 0) are assigned a value of -1, and heterozygous genotypes (0 / 1) are assigned a value of 0. This yields the SNP genotype values for each soybean sample. These SNP genotype values are stored in a text file, where rows represent soybean samples and columns represent SNP loci, resulting in an SNP genotype matrix.
[0048] 2. Obtaining SNP weight data
[0049] Genome-wide association analysis (GWAS) was performed on the protein and oil content of soybean seeds obtained under different environments (soybean seeds harvested in Beijing, Anhui, and Hainan in 2017, and soybean seeds harvested in Beijing and Hainan in 2018) using EMMAX software. Principal component analysis (PCA) was used to extract the first three principal components controlling population structure, and the Bonferroni correction method was used to set a significance threshold, obtaining one-dimensional significant loci results. Based on the GWAS results, SNP weights corresponding to each SNP locus in the SNP genotype matrix were extracted and organized into a one-dimensional vector W with the same length as the SNP loci in the soybean samples. 826020 That is, SNP weight data.
[0050] 3. Processing of soybean sample phenotypic data
[0051] Phenotypic data includes the protein and oil content of soybean seeds obtained under different environments (soybean seeds harvested in Beijing, Anhui, and Hainan in 2017, and soybean seeds harvested in Beijing and Hainan in 2018). Outliers were filtered out from the phenotypic data. The measured values of protein and oil content of soybean samples were stored in a text file, where rows represent soybean samples and the position of each soybean sample in the row is the same as in the SNP genotype matrix above, and columns represent traits (protein content, oil content), resulting in a column vector P. s .
[0052] 4. Training and testing of crop phenotypic prediction models
[0053] SNP data (SNP genotype matrix and SNP weight data) and their corresponding phenotypic data (column vector P) from 1024 soybean samples were analyzed. s The training set is used as the training set. The training set is input into the constructed crop phenotypic prediction model in batches, with each batch containing a genotype matrix of 32 SNPs (batch_size=32).
[0054] In the first convolutional block of the model, a genotype matrix A is constructed by assigning genotype values to 826,020 SNP loci from 32 soybean samples. 32×826020 The genotype tensor A'32×1×180×4589 was restructured into a (32,1,180,4589) tensor to accommodate subsequent convolution operations. Simultaneously, the weight scores W of the 826020 SNPs were... 826020 Reshaped into a matrix W of (180, 4589) 180×4589 The genotype tensor A'32×1×180×4589 is coupled with the matrix W 180×4589Perform Hadamard multiplication to generate a tensor W'32×1×180×4589 with weight information and shape (32,1,180,4589). A two-dimensional convolution is performed on W'32×1×180×4589, with a kernel shape of kernel = (90, 101) and a stride of (1, 4). Then, BatchNorm2d is used for regularization, resulting in L1 = BatchNorm2d(Conv2d(W'32×1×180×4589)). Finally, three two-dimensional convolutions are performed on L1, with kernel shapes of (60, 101), (1, 101), and (1, 101), and strides of (1, 2), (1, 1), and (1, 1). After each convolution, the result is activated by the Sigmoid function, resulting in an output tensor shape of (32, 1, 32, 312). The output data is then subjected to BatchNorm2d operation to obtain... These convolutional layers fuse row and column information, deeply extract SNP features, and reduce redundant information. The Sigmoid activation function increases the non-linearity of the model, while BatchNorm improves model stability and accelerates training.
[0055] In the N-layer Transformer block of the model, N=3 is set, and the embedding size and number of heads in each multi-head attention mechanism are as follows: First layer: {"embed_size1":312,"embed_size2":260,"num_heads":12}; Second layer: {"embed_size1":260,"embed_size2":100,"num_heads":10}; Third layer: {"embed_size1":100,"embed_size2":20,"num_heads":5}. In the first layer, layer regularization is first applied to L2 to obtain T1 = LayerNorm(L2). After processing by the multi-head attention mechanism, T1 is added to T2 with residual connection to obtain T2 = T1 + MultiHeadAttention(T1). On the main branch, LayerNorm, linear transformation, Sigmoid mapping, and another linear transformation operation are performed on T2. On the secondary branch, linear transformation and Sigmoid mapping are performed on T2 to obtain residual connection T3 = Linear(Sigmoid(Linear(LayerNorm(T2)))) + Sigmoid(Linear(T2)). In the second and third layers, the same operation as the first layer is performed. The input of each layer is provided by the output of the previous layer. Finally, the output T of the three Transformer blocks is obtained. 3NIts data shape is (32,32,20).
[0056] In the terminal convolutional block of the model, the above T is first applied. 3N LayerNorm processing is performed, followed by feature extraction using a neural network with Sigmoid activation function, resulting in E1 = Sigmoid(Linear(LayerNorm(T)). 3N Then, E1 is processed with two layers of one-dimensional convolution and a sigmoid activation function to output the predicted phenotypic traits (oil content, protein content) P of the sample. t , The kernel shapes of the one-dimensional convolutions are (20,) and (7,), and the sliding strides are 2 and 1, respectively.
[0057] The model training employs the backpropagation algorithm with the mean squared error function as the loss function. The Adam method is used to optimize the training process, with a learning rate set to 0.001. The Pearson correlation coefficient is used to evaluate the model's predictive accuracy. On a GeForce RTX 3080, prediction models for each phenotypic trait are built and trained using PyTorch (version 1.13.0). The training runs for 200 epochs with a batch size of 32. Training stops when the loss function stabilizes, yielding the trained model.
[0058] The SNP data (SNP genotype matrix and SNP weight data) and their corresponding phenotypic data (column vector P) of the remaining 529 soybean samples were analyzed. s The test set was used. The trained model and the traditional linear model rrBLUP (ridge regression best linear unbiased prediction) were used to predict soybean phenotypes on the test set. The Pearson correlation coefficients between the predicted and actual oil and protein contents were calculated to assess the model's prediction accuracy. The traditional linear model rrBLUP used was referenced from "Endelman J B. Ridge regression and other kernels for genomic selection with R package rrBLUP[J].The plant genome,2011,4(3)".
[0059] As shown in Table 1, compared with the rrBLUP model, the model of this invention significantly improves prediction accuracy under different environments and for different traits. The rrBLUP model predicts soybean oil content with an accuracy of 0.50-0.76, while the model of this invention predicts it with an accuracy of 0.76-0.87. In Beijing in 2018, the Pearson correlation coefficient between the predicted and actual soybean oil content values of this invention and the rrBLUP model increased by 65.38%. The rrBLUP model predicts protein content with an accuracy of 0.50-0.76, while the model of this invention predicts it with an accuracy of 0.68-0.85. In Beijing in 2018, the Pearson correlation coefficient between the predicted and actual soybean protein content values of this invention and the rrBLUP model increased by 48.08%. Therefore, the model of this invention possesses good generalization ability, strong robustness, and practical application potential.
[0060] Table 1
[0061] Pearson correlation coefficient between predicted and actual values of traits in a soybean population
[0062] Note: In the table, B17, A17, H17, B18, and H18 represent soybean seeds harvested in Beijing in 2017, Anhui in 2017, Hainan in 2017, Beijing in 2018, and Hainan in 2018, respectively; P represents protein content, and O represents oil content. The "increase" in the table indicates the increase in the Pearson correlation coefficient between the predicted and actual values compared to the rrBLUP model.
[0063] Example 2
[0064] This example uses genomic and phenotypic data from 244 maize natural population inbred lines. The genomic data comes from "Li Yuan, Fan Kaijian, An Tai, et al. A preliminary study on multi-environmental whole-genome prediction of agronomic traits in maize natural population inbred lines [J]. Acta Botanica Sinica, 2024, 59(6):1041.", of which 308,136 high-quality SNPs were used for predictive analysis. In 2022, the maize was planted at the Shunyi Base of the Institute of Crop Science, Chinese Academy of Agricultural Sciences, Beijing. Four agronomic traits were investigated, including days to anthesis (DTA), ear height (EH), ear length (EL), and kernel number per row (KNR).
[0065] 1. Obtaining the SNP genotype matrix
[0066] Each SNP locus is assigned a genotype as follows: homozygous (0 / 0 or 1 / 1), heterozygous (0 / 1), or unknown (. / ). Using maize B73V5 as the reference genome, homozygous genotypes (1 / 1) consistent with the reference genome are assigned a value of 1, homozygous genotypes (0 / 0) different from the reference genome are assigned a value of -1, and heterozygous genotypes (0 / 1) are assigned a value of 0. This yields the SNP genotype values for each maize sample. The SNP genotype values are stored in a text file, where rows represent maize samples and columns represent SNP loci, resulting in an SNP genotype matrix.
[0067] 2. Obtaining SNP weight data
[0068] Genome-wide association analysis (GWAS) was performed on pollen shedding date, ear height, ear length, and ear row number of maize seeds obtained in Shunyi, Beijing in 2022 using EMMAX software. Principal component analysis (PCA) was used to extract the first three principal components controlling population structure. The Bonferroni correction method was used to set a significance threshold, obtaining one-dimensional significant loci. Based on the GWAS results, SNP weights corresponding to each SNP locus in the SNP genotype matrix were extracted and organized into a one-dimensional vector W with the same length as the SNP loci in the maize samples. 308136 That is, SNP weight data.
[0069] 3. Processing of maize sample phenotypic data
[0070] Phenotypic data includes pollen shedding date, ear height, ear length, and number of rows per ear of maize seeds obtained in Shunyi, Beijing in 2022. Outliers were filtered out from the phenotypic data. The measurements of pollen shedding date, ear height, ear length, and number of rows per ear for maize samples were stored in a text file, where rows represent maize samples and the position of each sample in the row corresponds to the SNP genotype matrix described above. Columns represent traits (pollen shedding date, ear height, ear length, and number of rows per ear), resulting in a column vector P. s .
[0071] 4. Training and testing of crop phenotypic prediction models
[0072] SNP data (SNP genotype matrix and SNP weight data) and their corresponding phenotypic data (column vector P) from 195 maize samples were analyzed. s The training set is used as the training set. The training set is input into the constructed crop phenotypic prediction model in batches, with each batch containing a genotype matrix of 32 SNPs (batch_size=32).
[0073] In the first convolutional block of the model, a genotype matrix A is constructed by assigning genotype values to 308,136 SNP loci from 32 maize samples. 32×308136The genotype tensor A'32×1×111×2776 with the structure (32,1,111,2776) is reconstructed to fit subsequent convolution operations. Simultaneously, the weight scores W of the 308136 SNPs are... 308136 Reshaped into a matrix W of (111, 2776) 111×2776 The genotype tensor A'32×1×111×2776 is compared with the matrix W. 111×2776 Perform Hadamard multiplication to generate a tensor W'32×1×111×2776 containing weight information and with shape (32,1,111,2776). A two-dimensional convolution is performed on W'32×1×111×2776, with a kernel shape of kernel = (40, 100) and a stride of (1, 2). Then, BatchNorm2d is used for regularization, resulting in L1 = BatchNorm2d(Conv2d(W'32×1×111×2776)). Finally, three two-dimensional convolutions are performed on L1, with kernel shapes of (40, 101), (1, 101), and (1, 101), and strides of (1, 2), (1, 1), and (1, 1). After each convolution, the result is activated by the Sigmoid function, resulting in an output tensor shape of (32, 1, 33, 420). The output data is then subjected to BatchNorm2d operation to obtain... These convolutional layers fuse row and column information, deeply extract SNP features, and reduce redundant information. The Sigmoid activation function increases the non-linearity of the model, while BatchNorm improves model stability and accelerates training.
[0074] In the N-layer Transformer block of the model, N=3 is set, and the embedding size and number of heads in each multi-head attention mechanism are as follows: First layer: {"embed_size1":420,"embed_size2":260,"num_heads":21}; Second layer: {"embed_size1":260,"embed_size2":100,"num_heads":10}; Third layer: {"embed_size1":100,"embed_size2":20,"num_heads":5}. In the first layer, layer regularization is first applied to L2 to obtain T1 = LayerNorm(L2). After processing by the multi-head attention mechanism, T1 is added to T2 with residual connection to obtain T2 = T1 + MultiHeadAttention(T1). On the main branch, LayerNorm, linear transformation, Sigmoid mapping, and another linear transformation operation are performed on T2. On the secondary branch, linear transformation and Sigmoid mapping are performed on T2 to obtain residual connection T3 = Linear(Sigmoid(Linear(LayerNorm(T2)))) + Sigmoid(Linear(T2)). In the second and third layers, the same operation as the first layer is performed. The input of each layer is provided by the output of the previous layer. Finally, the output T of the three Transformer blocks is obtained. 3N Its data shape is (32,33,20).
[0075] In the terminal convolutional block of the model, the above T is first applied. 3N LayerNorm processing is performed, followed by feature extraction using a neural network with Sigmoid activation function, resulting in E1 = Sigmoid(Linear(LayerNorm(T)). 3N Then, E1 is processed with a two-layer one-dimensional convolution and a sigmoid activation function to output the predicted values P of the phenotypic traits (pollen shedding period, ear height, ear length, and number of ear rows) of the sample. t , The kernel shapes of the one-dimensional convolutions are (21,) and (7,), and the sliding strides are 2 and 1, respectively.
[0076] The model training employs the backpropagation algorithm with the mean squared error function as the loss function. The Adam method is used to optimize the training process, with a learning rate set to 0.001. The Pearson correlation coefficient is used to evaluate the model's predictive accuracy. On a GeForce RTX 3080, prediction models for each phenotypic trait are built and trained using PyTorch (version 1.13.0). The training runs for 200 epochs with a batch size of 6 (drop last set to true). Training stops when the loss function stabilizes, resulting in the trained model.
[0077] The SNP data (SNP genotype matrix and SNP weight data) and their corresponding phenotypic data (column vector P) of the remaining 49 maize samples were analyzed. s The test set was used. The trained model and the traditional linear model rrBLUP (ridge regression best linear unbiased prediction) were used to predict maize phenotypes on the test set. The Pearson correlation coefficients between the predicted and actual values of pollen shedding time, ear height, ear length, and ear row number were calculated to test the model's prediction accuracy. The traditional linear model rrBLUP used was referenced from "Endelman J B. Ridge regression and other kernels for genomic selection with R package rrBLUP[J].The plant genome,2011,4(3)".
[0078] As shown in Table 2, compared with the traditional linear model rrBLUP, the model of this invention significantly improves the prediction accuracy of maize agronomic traits. The rrBLUP model's prediction accuracy for pollen shedding date, ear height, ear length, and number of ear rows is 0.62, 0.49, 0.26, and 0.39, respectively. The model of this invention's prediction accuracy for pollen shedding date, ear height, ear length, and number of ear rows is 0.79, 0.72, 0.78, and 0.75, respectively. Compared with the rrBLUP model, the model of this invention improves the prediction accuracy for pollen shedding date, ear height, ear length, and number of ear rows by 27.42%, 46.94%, 200.00%, and 92.31%, respectively.
[0079] Table 2
[0080] Pearson correlation coefficient between predicted and actual values of traits in a maize population
[0081] Note: The “increase” in the table represents the increase in the Pearson correlation coefficient between the predicted and actual values of the model of this invention compared with the rrBLUP model.
Claims
1. A crop phenotypic prediction method based on Transformer, characterized in that, include: a) Construct a crop phenotypic prediction model; the model consists of a pre-convolutional block, N layers of Transformer blocks, and terminal convolutional blocks; the input of the model is the SNP data of crop samples, and the output is the predicted phenotypic trait value of crop samples; the SNP data includes the SNP genotype matrix and SNP weight data; b) The model is trained using SNPs data from multiple crop samples and their corresponding phenotypic trait measurements as a training set to obtain the trained model; c) Use the trained model to predict the phenotype of the crop sample based on the SNPs data of the crop sample.
2. The crop phenotypic prediction method according to claim 1, characterized in that, The SNP genotype matrix is a two-dimensional matrix A obtained by arranging the genotype values of each SNP locus in the crop sample in rows and columns, with the crop sample and SNP loci as the rows and columns. t×s The SNP weight data is a one-dimensional vector W corresponding to the weight score of each SNP locus in the SNP genotype matrix. s Where t represents the number of crop samples and s represents the number of SNP loci.
3. The crop phenotypic prediction method according to claim 2, characterized in that, The genomic data of each crop sample is compared with the reference genome. For each SNP locus, the genotype value of each SNP locus in each sample is obtained based on the consistency between the genomic data and the reference genome.
4. The crop phenotypic prediction method according to claim 3, characterized in that, Genome-wide association analysis (GWAS) was performed on the phenotypic traits of the multiple crop samples. Based on the GWAS results, the weight scores of each SNP locus were extracted and organized into a genotype matrix A corresponding to the SNPs. t×s The one-dimensional vector W of each SNP site s .
5. The crop phenotypic prediction method according to claim 4, characterized in that, In the first convolutional block of the model, A t×s and W s The SNPs are transformed into a four-dimensional tensor and a two-dimensional matrix, respectively. The four-dimensional tensor and the two-dimensional matrix are then subjected to Hadamard multiplication to obtain SNPs data containing weight information. The SNPs data containing weight information are subjected to a two-dimensional convolution and regularized using BatchNorm2d to obtain L1. L1 is then subjected to m layers of two-dimensional convolution, Sigmoid activation function and BatchNorm2d regularization to obtain L2, where m is an integer greater than or equal to 1.
6. The crop phenotypic prediction method according to claim 5, characterized in that, After L2 enters the N-layer Transformer block of the model, in the first layer, layer regularization is performed on L2 to obtain T1. After T1 is processed by the multi-head attention mechanism, an additive residual connection is formed with T1 to obtain T2. On the main branch, LayerNorm, linear transformation, Sigmoid mapping, and another linear transformation are performed on T2, and on the secondary branch, linear transformation and Sigmoid mapping are performed on T2 to obtain the residual connection T3. In each subsequent layer, the same operation as the first layer is performed, and the input of each layer is provided by the output of the previous layer to obtain the output T of the N-layer Transformer block. 3N , where N is an integer greater than or equal to 1.
7. The crop phenotypic prediction method according to claim 6, characterized in that, The T 3N After entering the terminal convolutional block of the model, for T 3N LayerNorm processing is applied, followed by feature extraction using a single-layer neural network with Sigmoid activation, yielding E1. E1 is then processed with k layers of one-dimensional convolution and Sigmoid activation to output the predicted phenotypic value P for the sample. t , where k is an integer greater than or equal to 1.
8. The crop phenotypic prediction method according to claim 7, characterized in that, Step b) includes: calculating the loss function results of the phenotypic trait measurements and phenotypic trait predictions of the multiple crop samples, and training the model until the loss function reaches a stable state, thereby obtaining the trained model.
9. The crop phenotypic prediction method according to any one of claims 1-8, characterized in that, The crops include soybeans and corn; the phenotypic traits include at least one of protein content, oil content, pollen shedding period, ear height, ear length, and number of ear rows.
10. A system for predicting crop phenotypes, characterized in that, It includes a phenotypic prediction model building module, an SNPs data acquisition module, a phenotypic trait measurement value acquisition module, and a model training module; The phenotypic prediction model building module is used to build the crop phenotypic prediction model as described in claims 1-9; The SNPs data acquisition module is used to acquire SNPs data from multiple crop samples; The phenotypic trait measurement module is used to acquire the phenotypic trait measurement values of the multiple crop samples; The model training module is used to train the crop phenotypic prediction model using the SNPs data of the multiple crop samples and their corresponding phenotypic trait measurements as a training set, so as to obtain the trained model.