Whole genome prediction method and system based on coordinate attention deep learning model
Through a genome-wide prediction method based on coordinate attention deep learning model, the problem of insufficient genotype and phenotype relationship capture in the prior art is solved, and higher phenotype prediction accuracy is achieved.
Patent Information
- Application Number
- CN202510093581.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-21
- Publication Date
- 2025-05-16
AI Technical Summary
The prior art has limitations in capturing complex nonlinear relationships between genotypes and phenotypes, resulting in insufficient accuracy in phenotype prediction.
A whole genome prediction method based on the coordinate attention deep learning model is adopted. By obtaining the data set of genotype data and phenotype data, the data is preprocessed, a deep learning model is established, and features are extracted using convolutional neural network and coordinate attention modules are used to perform full connection operations to predict phenotype values.
Improves the accuracy of phenotype prediction, can better capture the complex relationship between genotype and phenotype, with greater potential and advantages in genomic selection.
Smart Images

Figure CN120015130A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of bioinformatics, and specifically relates to a whole genome prediction method and system based on a coordinate attention deep learning model. Background Art
[0002] As the global population continues to increase, the problem of food security will become more serious, especially adverse factors such as extreme weather, pests and diseases, and scarce land resources will further threaten crop yields. Researching crop varieties with excellent traits such as high yield and resistance to pests and diseases has become a key task in agricultural breeding. Whole genome prediction is a method of predicting crop phenotypes through whole genome genetic markers, which has been applied to crop breeding in recent years.
[0003] There is often a high correlation between genotype data and phenotypic data, and it is crucial to design a model that effectively learns the complex relationship between the two. Although traditional breeding methods based on statistics or machine learning can perform genomic selection, they have limitations in capturing the complex nonlinear relationship between genotype and phenotype. In contrast, deep learning methods, due to their powerful nonlinear modeling capabilities, can more accurately reveal the deep and complex relationship between genotype and phenotype, thus having greater potential and advantages in genomic selection. Summary of the invention
[0004] The technical problem to be solved by the present invention is to provide a whole genome prediction method and system based on a coordinate attention deep learning model to improve the accuracy of phenotype prediction.
[0005] The technical solution adopted by the present invention to solve the above technical problems is: a whole genome prediction method based on a coordinate attention deep learning model, comprising the following steps: S1: Obtain a dataset including genotype data and phenotype data; S2: Preprocess the data set to obtain a two-dimensional feature matrix; S3: Establish a deep learning model to obtain the predicted phenotypic value through a two-dimensional feature matrix; S4: Divide the dataset into training set and test set, train and test the deep learning model until the final whole genome prediction model is obtained.
[0006] According to the above scheme, in step S2, the specific steps are: S21: Encode genotype data and normalize phenotypic data; S22: Convert the encoded genotype data into a two-dimensional feature matrix.
[0007] Furthermore, in step S21, the specific steps for encoding the genotype data are: representing the genotype data as a matrix, each row represents the genotype data of a sample, and each column represents the genotype data of the same SNP site in different samples, that is, the number of rows represents the number of samples, and the number of columns represents the number of SNP sites, and the genotype data are encoded by row to obtain the encoded genotype data.
[0008] According to the above scheme, in step S3, the specific steps are: S31: Input the two-dimensional feature matrix into the convolutional neural network model to extract features; S32: extracting weighted output features using a coordinate attention operation on the features extracted in step S31; S33: Activate the weighted output features and map them through a fully connected operation to obtain the predicted phenotypic value.
[0009] Furthermore, in the step S31, the specific steps are: S311: Extract features from the two-dimensional feature matrix through convolution operation, and the convolution kernel size is , the step length is , and set the bias term; S312: Use the LeakyReLU function to perform activation operation and convert the convolution feature matrix into a nonlinear relationship; S313: Use size The window with a step size of , perform the maximum pooling operation on the activated feature matrix to obtain the feature matrix after maximum pooling; S314: Perform dropout processing on the feature matrix after the maximum pooling operation to randomly discard a part of neurons to reduce overfitting, and obtain the feature matrix after convolution and dropout processing.
[0010] Furthermore, in step S32, the specific steps are: S321: Perform global average pooling on the feature matrix after convolution and dropout processing in the horizontal and vertical directions respectively; S322: concatenate the pooled horizontal feature matrix and the vertical feature matrix in the channel dimension, apply a convolution operation, compress the channel dimension by 16 times, and obtain a compressed feature matrix; S323: performing a nonlinear transformation on the compressed feature matrix through a ReLU activation function, and then performing a batch normalization operation to obtain a batch normalized feature matrix; S324: Segment the batch-normalized feature matrix, and then use the Sigmoid function to generate horizontal and vertical attention weights; S325: Perform a dot multiplication operation on the horizontal and vertical attention weights and the feature matrix after convolution and dropout processing to obtain a weighted output feature matrix.
[0011] Furthermore, in step S33, the specific steps are: S331: Flatten the weighted output feature matrix into a one-dimensional feature vector; S332: Perform dropout processing on the one-dimensional feature vector and use the LeakReLU activation function for nonlinear activation; S333: Mapping the activated feature vector into a 32-dimensional feature vector through the first full connection operation; S334: Perform dropout processing on the 32-dimensional feature vector and use the LeakReLU activation function for nonlinear activation; S335: Through the second full connection operation, the 32-dimensional feature vector is mapped to a single value as the phenotypic prediction value.
[0012] According to the above scheme, in step S4, the specific steps are: S41: Divide the data set into a training set and a test set according to a certain ratio; S42: Set the loss function of the deep learning model to , including Smooth L1 function and L2 regularization term; S43: Loss function of deep learning model using Adam optimization method Optimize and adjust the model parameters to obtain the final whole genome prediction model.
[0013] Furthermore, in step S43, the specific steps are: The Adam optimizer was used to train the convolutional neural network model. Five-fold cross validation was used for model training. The Pearson correlation coefficient and mean square error between the predicted value and the true value were used as evaluation indicators of the prediction performance, and the average of the five results was used as the final result of the evaluation indicator.
[0014] A whole genome prediction system based on coordinate attention deep learning model, A data acquisition submodule, used to acquire a data set including genotype data and phenotype data; The preprocessing submodule is used to preprocess the data set to obtain a two-dimensional feature matrix; The deep learning submodule is used to build a deep learning model to predict phenotypic values through a two-dimensional feature matrix; The training prediction submodule is used to divide the data set into a training set and a test set, and to train and test the deep learning model until the final whole genome prediction model is obtained.
[0015] The beneficial effects of the present invention are: 1. The present invention provides a whole genome prediction method and system based on a coordinate attention deep learning model, which obtains a data set including genotype data and phenotypic data, preprocesses the data set to obtain a two-dimensional feature matrix; establishes a deep learning model, and predicts phenotypic values through the two-dimensional feature matrix; divides the data set into a training set and a test set, trains and tests the deep learning model until the final whole genome prediction model is obtained, thereby achieving the function of improving the accuracy of phenotypic prediction.
[0016] 2. This paper proposes a new whole-genome prediction method, which uses a coordinate attention module to better capture the complex relationship between genotype and phenotype and improve the accuracy of whole-genome prediction.
[0017] Of course, any product implementing the present invention does not necessarily need to achieve all of the advantages described above at the same time. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0019] Figure 1 is a flow chart of an embodiment of the present invention.
[0020] Figure 2 It is a schematic diagram of a flow chart of an embodiment of the present invention.
[0021] Figure 3 This is a comparison chart of the Pearson correlation coefficient of whole genome prediction on a soybean dataset using the deep learning model of an embodiment of the present invention and other prediction methods.
[0022] Figure 4 This is a comparison chart of the mean square error of whole genome prediction of the deep learning model of an embodiment of the present invention and other prediction methods on a soybean dataset. DETAILED DESCRIPTION
[0023] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0024] Example 1 See also Figure 1,The specific steps of a whole genome prediction method based on the coordinate attention deep learning model are as follows: S1: Obtain a dataset including genotype data and phenotype data; S2: Preprocess the data set to obtain a two-dimensional feature matrix; S3: Establish a deep learning model to obtain the predicted phenotypic value through a two-dimensional feature matrix; S4: Divide the dataset into training set and test set, train and test the deep learning model until the final whole genome prediction model is obtained.
[0025] Furthermore, in step S2, the specific steps are: S21: Encode genotype data and normalize phenotypic data; S22: Convert the encoded genotype data into a two-dimensional feature matrix.
[0026] Furthermore, in step S21, the specific steps for encoding the genotype data are: representing the genotype data as a matrix, each row represents the genotype data of a sample, and each column represents the genotype data of the same SNP site in different samples, that is, the number of rows represents the number of samples, and the number of columns represents the number of SNP sites, and the genotype data are encoded by row to obtain the encoded genotype data.
[0027] In step S3, the specific steps are: S31: Input the two-dimensional feature matrix into the convolutional neural network model to extract features; S32: extracting weighted output features using a coordinate attention operation on the features extracted in step S31; S33: Activate the weighted output features and map them through a fully connected operation to obtain the predicted phenotypic value.
[0028] Furthermore, in step S31, the specific steps are: S311: Extract features from the two-dimensional feature matrix through convolution operation, and the convolution kernel size is , the step length is , and set the bias term; S312: Use the LeakyReLU function to perform activation operation and convert the convolution feature matrix into a nonlinear relationship; S313: Use size The window with a step size of , perform the maximum pooling operation on the activated feature matrix to obtain the feature matrix after maximum pooling; S314: Perform dropout processing on the feature matrix after the maximum pooling operation to randomly discard a part of neurons to reduce overfitting, and obtain the feature matrix after convolution and dropout processing.
[0029] Further, in step S32, the specific steps are: S321: Perform global average pooling on the feature matrix after convolution and dropout processing in the horizontal and vertical directions respectively; S322: concatenate the pooled horizontal feature matrix and the vertical feature matrix in the channel dimension, apply a convolution operation, compress the channel dimension by 16 times, and obtain a compressed feature matrix; S323: performing a nonlinear transformation on the compressed feature matrix through a ReLU activation function, and then performing a batch normalization operation to obtain a batch normalized feature matrix; S324: Segment the batch-normalized feature matrix, and then use the Sigmoid function to generate horizontal and vertical attention weights; S325: Perform a dot multiplication operation on the horizontal and vertical attention weights and the feature matrix after convolution and dropout processing to obtain a weighted output feature matrix.
[0030] Furthermore, in step S33, the specific steps are: S331: Flatten the weighted output feature matrix into a one-dimensional feature vector; S332: Perform dropout processing on the one-dimensional feature vector and use the LeakReLU activation function for nonlinear activation; S333: Mapping the activated feature vector into a 32-dimensional feature vector through the first full connection operation; S334: Perform dropout processing on the 32-dimensional feature vector and use the LeakReLU activation function for nonlinear activation; S335: Through the second full connection operation, the 32-dimensional feature vector is mapped to a single value as the phenotypic prediction value.
[0031] In step S4, the specific steps are: S41: Divide the data set into a training set and a test set according to a certain ratio; S42: Set the loss function of the deep learning model to , including Smooth L1 function and L2 regularization term; S43: Loss function of deep learning model using Adam optimization method Optimize and adjust the model parameters to obtain the final whole genome prediction model.
[0032] Furthermore, in step S43, the specific steps are: The Adam optimizer was used to train the convolutional neural network model. Five-fold cross validation was used for model training. The Pearson correlation coefficient and mean square error between the predicted value and the true value were used as evaluation indicators of the prediction performance, and the average of the five results was used as the final result of the evaluation indicator.
[0033] This embodiment obtains a data set including genotype data and phenotypic data, pre-processes the data set to obtain a two-dimensional feature matrix; establishes a deep learning model to predict phenotypic values through the two-dimensional feature matrix; divides the data set into a training set and a test set, trains and tests the deep learning model until a final whole genome prediction model is obtained, thereby achieving the function of improving the accuracy of phenotypic prediction.
[0034] Example 2 The steps of this embodiment are the same as those of Example 1, except that each step is applied to a specific example. Taking the YHSBLP Soybean dataset as an example, single nucleotide polymorphisms (SNPs) obtained by RAD-seq simplified genome sequencing technology from the Jianghuai soybean breeding germplasm population (Soybean573) established by the National Soybean Improvement Center are used. The dataset includes genotyping of 573 samples, with a total of 61,166 SNPs, including phenotypic data of the average 100-grain weight of soybeans from the YHSBLP population obtained in 2013, 2014, 2017, and 2018. (Yield_13: average 100-grain weight in 2013; Yield_14: average 100-grain weight in 2014; Yield_17: average 100-grain weight in 2017; Yield_18: average 100-grain weight in 2018).
[0035] This example compares the performance of a genome-wide prediction method based on a deep learning model of coordinate attention (DLCA) with machine learning methods (RF, SVM) and deep learning methods (DeepGS, DNNGP, SoyDNGP) on four trait predictions of the YHSBLP Soybean dataset, and uses the Pearson correlation coefficient (PCC) and mean square error (MSE) as evaluation indicators of prediction performance.
[0036] See also Figure 2 , this embodiment specifically includes the following steps: S1: Get a dataset including genotypes and phenotypes; S2: Preprocess the data set to obtain a two-dimensional feature matrix; the specific steps are: S21: Encode genotype data and normalize phenotypic data; Represent the genotype data as , each row represents the genotype data of a sample, and each column represents the genotype data of the same SNP site in different samples, that is, represents the number of samples, represents the number of SNP sites, where ; Encode the genotype data by row, The code is 0, Coded as 1, Coded as 2, The code is 3, and the encoded genotype data is .
[0037] S22: Convert the encoded genotype data into a two-dimensional feature matrix; the specific steps are: The genotype data after encoding is constructed with a size of The two-dimensional feature matrix of ,in Indicates rounding up, and the missing genotype data are filled with the mean value, where .
[0038] S3: Establish a deep learning model and obtain the predicted phenotypic value through a two-dimensional feature matrix; the specific steps are: S31: Input the two-dimensional feature matrix into the convolutional neural network model to extract features; the specific steps are: S311: The two-dimensional feature matrix is transformed into The convolution kernel of is In this embodiment, the convolution kernel size is 3×3 and the step length is 1; and a bias item is set to extract feature information; S312: Use the LeakyReLU function to convert the convolutional feature matrix into a nonlinear relationship to enhance the expressiveness of the model. The calculation formula is as follows:
[0039] in, Represents a small constant with a value of 0.01.
[0040] S313: Use another size The window with a step size of In this embodiment, the window size is 4×4 and the step length is 4; a maximum pooling operation is performed on the activated feature matrix to obtain a feature matrix after maximum pooling; S314: Apply the dropout layer to the feature matrix after the maximum pooling operation to randomly discard some neurons to reduce overfitting, and finally obtain the feature matrix after convolution and dropout processing .
[0041] S32: using the coordinate attention module to enhance the feature extraction capability of the features extracted in step S31, and extracting weighted output features; the specific steps are: S321: The feature matrix Global average pooling is performed in the horizontal and vertical directions respectively to obtain ; S322: Pooling the horizontal feature matrix and the vertical feature matrix Splicing is performed on the channel dimension, and a convolution operation is applied to reduce the channel dimension by 16 times to obtain the compressed feature matrix ; S323: The compressed feature matrix The nonlinear transformation is performed through the ReLU activation function, and the calculation formula is as follows:
[0042] Then use batch normalization to get the batch normalized feature matrix ,The normalized calculation formula for each feature is as follows:
[0043] in represents the features after nonlinear transformation, represents the mean of the feature, represents the variance of the feature, Represents a small constant.
[0044] S324: Batch normalized feature matrix Segmentation is performed, and then the Sigmoid function is used to generate horizontal and vertical attention weights ; S325: Attention weight With the input genotype feature matrix Perform a dot multiplication operation to obtain the final weighted output feature matrix .
[0045] S33: Activate the weighted output features and map them through two fully connected layers to obtain the predicted phenotypic values; the specific steps are: S331: The final weighted output feature matrix Flattened to a one-dimensional feature vector; S332: Apply a dropout layer to the one-dimensional feature vector and use the LeakReLU activation function for nonlinear activation; S333: Through the first fully connected layer, the activated feature vector is mapped into a 32-dimensional feature vector; S334: The 32-dimensional feature vector is passed through a dropout layer and nonlinearly activated using the LeakReLU activation function; S335: Then, through the second fully connected layer, the 32-dimensional feature vector is mapped to a single value as the phenotypic prediction value.
[0046] S4: Divide the data set into a training set and a test set, train and test the established deep learning model until the final whole genome prediction model is obtained. The specific steps are: S41: Divide the data set into a training set and a test set in a ratio of 8:2; S42: The loss function of the overall deep learning model is , including the Smooth L1 function and the L2 regularization term, the calculation formula is as follows:
[0047] in, represents the predicted value, represents the true value, represents the Smooth L1 function, Indicates the model weight parameters, Indicates the total number of parameters. is the regularization coefficient.
[0048] S43: Loss function of deep learning model using Adam optimization method Optimize and adjust the corresponding parameters in the model to obtain the final whole genome prediction model.
[0049] The Adam optimizer is used to train the convolutional neural network model. The parameters related to back propagation selected in this example are as follows: the regularization coefficient is 1e-4, the learning rate is 0.00009, the number of iterations is 150 epochs, the batch size is set to 16, and the model training uses five-fold cross validation. The Pearson correlation coefficient (PCC) and mean square error (MSE) between the predicted value and the true value are used as evaluation indicators of the prediction performance, and the average of the five results is used as the final result of the evaluation indicator.
[0050] The performance of the whole genome prediction method based on the coordinate attention deep learning model was compared with machine learning methods (RF, SVM) and deep learning methods (DeepGS, DNNGP, SoyDNGP) on the prediction of four traits of the YHSBLP Soybean dataset, as follows: Figure 3 This is a comparison of the Pearson correlation coefficient of the whole genome prediction of the deep learning model of the embodiment of the present invention and other prediction methods on the soybean data set. In general, the prediction accuracy of this embodiment is higher than that of the other five existing methods. Figure 4 This is a comparison of the whole genome prediction mean square error of the deep learning model of the embodiment of the present invention and other prediction methods on the soybean dataset. The mean square error value of this embodiment is lower than that of other methods.
[0051] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.
[0052] Example 3 This embodiment is used to implement the principle construction system of the above method embodiment, including a data acquisition submodule, a preprocessing submodule, a deep learning submodule and a training prediction submodule.
[0053] A data acquisition submodule, used to acquire a data set including genotype data and phenotype data; The preprocessing submodule is used to preprocess the data set to obtain a two-dimensional feature matrix; The deep learning submodule is used to build a deep learning model to predict phenotypic values through a two-dimensional feature matrix; The training prediction submodule is used to divide the data set into a training set and a test set, and to train and test the deep learning model until the final whole genome prediction model is obtained.
[0054] Each sub-module is mainly used to implement each step of the method embodiment, which will not be described in detail here.
[0055] It should be pointed out that, according to the needs of implementation, the various steps / components described in this application can be split into more steps / components, and two or more steps / components or partial operations of steps / components can be combined into new steps / components to achieve the purpose of the present invention.
[0056] This embodiment also includes a processor, a communication interface, a memory and a communication bus; wherein the processor, the communication interface and the memory communicate with each other through the communication bus; a computer program is stored in the memory, and when the program is executed by the processor, the processor executes the steps of a whole genome prediction method based on a coordinate attention deep learning model.
[0057] This embodiment also provides a computer-readable storage medium having executable instructions stored thereon, which, when executed by a processor, enables the processor to implement a whole genome prediction method based on a coordinate attention deep learning model.
[0058] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware.
[0059] Moreover, the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes.
[0060] The present application is described with reference to the flowchart of the method and computer program product according to Embodiment 1 of the present application and the block diagram of the device (system) according to Embodiment 3. It should be understood that each process or box in the flowchart or block diagram, and the combination of processes or boxes in the flowchart or block diagram can be implemented by computer program instructions.
[0061] These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the process Figure 1 A process or multiple processes or boxes Figure 1 A genome-wide prediction system based on a coordinate attention deep learning model for functions specified in one or more boxes.
[0062] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes or boxes Figure 1 A function specified in one or more boxes.
[0063] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process in the computer or other programmable device. Figure 1 A process or multiple processes or boxes Figure 1 The steps of a genome-wide prediction method based on a coordinate attention deep learning model are specified in one or more boxes.
[0064] The above embodiments are only used to illustrate the design ideas and features of the present invention, and their purpose is to enable those skilled in the art to understand the content of the present invention and implement it accordingly. The protection scope of the present invention is not limited to the above embodiments. Therefore, any equivalent changes or modifications made based on the principles and design ideas disclosed by the present invention are within the protection scope of the present invention.
Claims
1. A whole genome prediction method based on a coordinate attention deep learning model, characterized in that: The following steps are involved: S1: Obtain a dataset including genotype data and phenotype data; S2: Preprocess the data set to obtain a two-dimensional feature matrix; S3: Establish a deep learning model to obtain the predicted phenotypic value through a two-dimensional feature matrix; S4: Divide the dataset into training set and test set, train and test the deep learning model until the final whole genome prediction model is obtained.
2. The whole genome prediction method based on the coordinate attention deep learning model according to claim 1, characterized in that: In the step S2, the specific steps are: S21: Encode genotype data and normalize phenotypic data; S22: Convert the encoded genotype data into a two-dimensional feature matrix.
3. The whole genome prediction method based on the coordinate attention deep learning model according to claim 2, characterized in that: In step S21, the specific steps of encoding the genotype data are: representing the genotype data as a matrix, each row represents the genotype data of a sample, and each column represents the genotype data of the same SNP site in different samples, that is, the number of rows represents the number of samples, and the number of columns represents the number of SNP sites. The genotype data is encoded by row to obtain the encoded genotype data.
4. The whole genome prediction method based on the coordinate attention deep learning model according to claim 1, characterized in that: In the step S3, the specific steps are: S31: Input the two-dimensional feature matrix into the convolutional neural network model to extract features; S32: extracting weighted output features using a coordinate attention operation on the features extracted in step S31; S33: Activate the weighted output features and map them through a fully connected operation to obtain the predicted phenotypic value.
5. The whole genome prediction method based on the coordinate attention deep learning model according to claim 4, characterized in that: In the step S31, the specific steps are: S311: Extract features from the two-dimensional feature matrix through convolution operation, and the convolution kernel size is , the step length is , and set the bias term; S312: Use the LeakyReLU function to perform activation operation and convert the convolution feature matrix into a nonlinear relationship; S313: Use size The window with a step size of , perform the maximum pooling operation on the activated feature matrix to obtain the feature matrix after maximum pooling; S314: Perform dropout processing on the feature matrix after the maximum pooling operation to randomly discard a part of neurons to reduce overfitting, and obtain the feature matrix after convolution and dropout processing.
6. The whole genome prediction method based on the coordinate attention deep learning model according to claim 5, characterized in that: In the step S32, the specific steps are: S321: Perform global average pooling on the feature matrix after convolution and dropout processing in the horizontal and vertical directions respectively; S322: concatenate the pooled horizontal feature matrix and the vertical feature matrix in the channel dimension, apply a convolution operation, compress the channel dimension by 16 times, and obtain a compressed feature matrix; S323: performing a nonlinear transformation on the compressed feature matrix through a ReLU activation function, and then performing a batch normalization operation to obtain a batch normalized feature matrix; S324: Segment the batch-normalized feature matrix, and then use the Sigmoid function to generate horizontal and vertical attention weights; S325: Perform a dot multiplication operation on the horizontal and vertical attention weights and the feature matrix after convolution and dropout processing to obtain a weighted output feature matrix.
7. The whole genome prediction method based on the coordinate attention deep learning model according to claim 6, characterized in that: In the step S33, the specific steps are: S331: Flatten the weighted output feature matrix into a one-dimensional feature vector; S332: Perform dropout processing on the one-dimensional feature vector and use the LeakReLU activation function for nonlinear activation; S333: Mapping the activated feature vector into a 32-dimensional feature vector through the first full connection operation; S334: Perform dropout processing on the 32-dimensional feature vector and use the LeakReLU activation function for nonlinear activation; S335: Through the second full connection operation, the 32-dimensional feature vector is mapped to a single value as the phenotypic prediction value.
8. The whole genome prediction method based on the coordinate attention deep learning model according to claim 1, characterized in that: In the step S4, the specific steps are: S41: Divide the data set into a training set and a test set according to a certain ratio; S42: Set the loss function of the deep learning model to , including Smooth L1 function and L2 regularization term; S43: Loss function of deep learning model using Adam optimization method Optimize and adjust the model parameters to obtain the final whole genome prediction model.
9. The whole genome prediction method based on the coordinate attention deep learning model according to claim 8, characterized in that: In the step S43, the specific steps are: The Adam optimizer was used to train the convolutional neural network model. Five-fold cross validation was used for model training. The Pearson correlation coefficient and mean square error between the predicted value and the true value were used as evaluation indicators of the prediction performance, and the average of the five results was used as the final result of the evaluation indicator.
10. A whole genome prediction system based on a coordinate attention deep learning model, characterized in that: A data acquisition submodule, used to acquire a data set including genotype data and phenotype data; The preprocessing submodule is used to preprocess the data set to obtain a two-dimensional feature matrix; The deep learning submodule is used to build a deep learning model to predict phenotypic values through a two-dimensional feature matrix; The training prediction submodule is used to divide the data set into a training set and a test set, and to train and test the deep learning model until the final whole genome prediction model is obtained.
Citation Information
Cited By
Trait heritable variation site prediction method fusing transfer learning and interpretable mechanism
CN120895090A
Whole genome prediction method and system based on Kolmogov-Arnod network
CN121260235A
Intelligent system for correlation analysis of whole genome of pyrus ussuriensis on basis of deep learning
CN121641206A