Plant phenotype prediction

Through the combination of graph convolutional neural network and Transformer, multiomic data are integrated to predict plant phenotypes, solving the problem of low singleomic prediction effect and achieving more accurate phenotype prediction effect.

WO2025065954A9PCT designated stage expired Publication Date: 2025-05-22ZHEJIANG LAB +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2023/143520
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-09-28
Filing Date
2023-12-29
Publication Date
2025-05-22

AI Technical Summary

Technical Problem

In the prior art, the span from monomicrologic gene to phenotypic prediction is too large, resulting in low phenotypic prediction effect and ineffective integration of multiomics data for accurate phenotypic prediction.

Method used

Using a combination of graph convolutional neural network and Transformer, we construct multi-omic graph structure data, integrate multi-omic data such as genome, transcriptome and metabolic groups, and use graph neural network and Transformer network to mine the correlation relationship between multi-omics to achieve prediction of multi-omic phenotypes.

Benefits of technology

It effectively solved the problem of low accuracy of single-omic phenotype prediction, and improved the effect of phenotype prediction by integrating multi-omic data, achieving more accurate plant phenotype prediction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2023143520_22052025_PF_FP_ABST
    Figure CN2023143520_22052025_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed are a plant phenotype prediction method and apparatus. According to one example of the method, graph structure data for each plant species is constructed by using multiomics data, such as a genome, transcriptome, and metabolome, as graph nodes and by using degrees of association between different omics as graph edges. Then, the constructed graph structure data is inputted into a graph convolutional neural network to extract node features, the node features are updated by means of a Transformer network, the node features are spliced, and the spliced node features are inputted into a fully connected layer, so as to output a phenotype prediction value. Thus, a graph convolutional neural network is combined with a Transformer network to achieve gene-to-phenotype prediction, and a multiomics constructed graph structure is used to fuse multiomics data to achieve accurate phenotype prediction, effectively improving a phenotype prediction effect.
Need to check novelty before this filing date? Find Prior Art

Description

Plant phenotype prediction Technical Field

[0001] The present invention relates to the field of deep learning technology and bioinformatics, and particularly to plant phenotype prediction. Background Art

[0002] Gene-to-phenotype prediction plays a crucial role in disease prediction, pathogenic gene discovery, and crop yield prediction, and is a core problem in genetic biology. Currently, phenotypic prediction is based on the genome. However, the expression structure of the phenotype is not solely related to the genome but also involves multiple omics, such as the transcriptome and metabolome. Direct gene-to-phenotype prediction involves a large gap, resulting in low phenotypic prediction effectiveness. Therefore, it is necessary to link multiple omics for accurate phenotypic prediction. Furthermore, graph structures can link nodes of multiple types. By leveraging graph neural networks to mine the relationships between these nodes, multiple omics can be integrated for phenotypic prediction.

[0003] In recent years, the Transformer has been widely used in computer vision. Its attention mechanism can learn the correlation between different local features. In particular, there has been relevant work applying the Transformer to protein coding region prediction. However, research on applying the Transformer to phenotype prediction is relatively rare.

[0004] Summary of the Invention

[0005] To this end, the present disclosure proposes a phenotypic prediction method that combines a graph convolutional neural network and a transformer. This method can achieve phenotypic prediction by integrating multi-omics data, which can effectively solve the problem of low phenotypic prediction effect of single-omics. In particular, the present disclosure provides a plant phenotypic prediction method and device. Among them, by constructing graph structure data from multi-omics and using a combination of a graph convolutional neural network and a transformer network to mine the correlation between multi-omics, the prediction of multi-omics phenotypes is achieved.

[0006] According to a first aspect of an embodiment of the present invention, a method for predicting plant phenotypes is provided, comprising the following steps:

[0007] (1) obtaining multi-omics data and phenotypic data of the target species, performing quality control and dimensionality reduction on the multi-omics data to obtain multiple omics features, and processing the phenotypic data for outliers and missing values ​​to obtain true phenotypic data;

[0008] (2) normalizing the multiple omics features obtained in step (1), calculating a first similarity between the multiple omics features of each sample of the target species, and constructing graph structure data of each sample using the first similarity;

[0009] (3) For each sample, the graph structure data of each sample obtained in step (2) is input into a three-layer graph neural network to extract the features of each node of each sample;

[0010] (4) For each sample, the features of each node of each sample extracted in step (3) are spliced ​​in the order of multi-omics and input into two first fully connected layers to obtain a first phenotype prediction value, and a first loss is calculated based on the first phenotype prediction value and the corresponding true phenotype data;

[0011] (5) For each sample, the features of each node of each sample obtained in step (3) are input into the Transformer network as local identification embedding features and global identification embedding features, and the updated global identification embedding features and the local identification embedding features of each node are obtained. The second similarity between the updated global identification embedding features and the local identification embedding features of each node is calculated, and the local identification embedding features and the global identification embedding features of the two nodes with the largest second similarity are selected and spliced ​​into two second fully connected layers to obtain a second phenotype prediction value, and the second loss is calculated according to the second phenotype prediction value and the corresponding true phenotype data;

[0012] (6) Repeat steps (3) to (5) to perform training, and according to the total loss of the sum of the first loss and the second loss, use the stochastic gradient descent method to back-propagate and adjust the network parameters of the graph neural network, the first fully connected layer, the Transformer network, and the second fully connected layer to obtain the trained graph neural network, the first fully connected layer, the Transformer network, and the second fully connected layer;

[0013] (7) For the material to be tested, a first phenotypic prediction value is obtained through the trained graph neural network and the first fully connected layer, and a second phenotypic prediction value is obtained through the trained graph neural network, the Transformer network and the second fully connected layer. The average of the first phenotypic prediction value and the second phenotypic prediction value is taken as the final predicted phenotypic result of the material to be tested.

[0014] Furthermore, the multi-omics data includes genomic data, transcriptomic data and metabolomic data; the graph neural network includes a graph convolution layer and a ReLU activation layer.

[0015] Furthermore, in step (1), the multi-omics data are subjected to quality control and dimensionality reduction processing to obtain multiple omics features, and the phenotypic data are subjected to outlier and missing value processing to obtain true phenotypic data, which specifically includes the following sub-steps:

[0016] (1.1) Remove missing values ​​and abnormal outliers from the phenotypic data of each sample of the target species;

[0017] (1.2) Delete positions with a missing rate greater than 90% in the multi-omics data of each sample of the target species, and delete positions with a minimum allele frequency less than 5% in the genomic data to complete the quality control filtering of the multi-omics data;

[0018] (1.3) encoding the multi-omics data after quality control filtering in step (1.2), wherein the genomic data are encoded as -1, 0, and 1 according to the homozygous state, heterozygous state, and homozygous variant state, respectively, and other omics data are not encoded according to the measured values;

[0019] (1.4) For the data at each position of each omics, the Pearson correlation coefficient between the coding features of all samples and the corresponding phenotypic data is calculated to obtain the degree of correlation between each position of each omics data and the phenotypic data. For each omics data, the L positions with the largest correlation degree are selected as omics features to obtain multiple omics features.

[0020] Furthermore, the calculation formula of the Pearson correlation coefficient is:

[0021] Among them, ρ XY is the Pearson correlation coefficient, X and Y represent the coding features and the corresponding phenotypic data, respectively, cov(*) represents the covariance, and σ represents the standard deviation.

[0022] Furthermore, the step (2) includes the following sub-steps:

[0023] (2.1) Each omics feature obtained in step (1.4) is normalized as follows: Assume that there are M samples, each sample has N omics, and the jth omics feature after dimensionality reduction of the i-th sample is denoted as f ij , i = 1…M, j = 1…N, and normalized according to the following formula:

[0024] Among them, f′ ij represents the jth omics feature after normalization of the i-th sample, u j is the mean of all samples of the j-th omics, var j is the standard deviation of all samples of the j-th omics;

[0025] (2.2) For each sample, the cosine similarity between the normalized omics features calculated according to step (2.1) is used as the first similarity, and the calculation formula is as follows:

[0026] Where similarity is the cosine similarity, A and B represent the normalized omics features, A·B is the inner product, and ‖A‖ represents the norm corresponding to the normalized omics feature A.

[0027] (2.3) For each sample, the graph structure data of each sample is constructed by taking the different omics features of the sample as graph nodes and taking the first similarity between the different omics features obtained in step (2.2) as the edge weight of the graph.

[0028] Furthermore, the first loss includes prediction regression loss and Pearson correlation coefficient loss, and its expression is:

[0029] Among them, Loss1 represents the first loss, M is the number of samples, and Y i is the true phenotypic data of node sample i, Y′ i is the first phenotype prediction value of node sample i, is the true phenotypic data mean of all samples, is the mean of the first phenotype prediction values ​​of all samples, and α is the adjustment coefficient between the prediction regression loss and the Pearson correlation coefficient loss.

[0030] Furthermore, the step (5) includes the following sub-steps:

[0031] (5.1) Generate a global identity embedding feature using random normal distribution parameters initialization. The value of each dimension of the global identity embedding feature comes from a random value generated by the normal distribution, and its feature length is the same as the feature length of each node extracted in step (3);

[0032] (5.2) For each sample, calculate the first attention matrix of the sample based on the features of all nodes extracted in step (3);

[0033] (5.3) For each sample, the features of each node extracted in step (3) are used as local identification embedding features, and the global identification embedding features and the local identification embedding features of each node obtained in step (5.1) are input into a Transformer network, the Transformer network including a self-attention module and a feedforward neural network module, the second attention matrix corresponding to the sample is calculated by the self-attention module, the second attention matrix is ​​combined with the first attention matrix obtained in step (5.2) to update the global identification embedding features and the local identification embedding features of each node, and the updated global identification embedding features and the local identification embedding features of each node are obtained through the feedforward neural network module;

[0034] (5.4) Calculate the cosine similarity between the updated global identifier embedding feature obtained in step (5.3) and the local identifier embedding feature of each node as the second similarity, select the local identifier embedding features and the global identifier embedding features of the two nodes with the largest second similarity, and concatenate them to obtain a second concatenated feature, input the second concatenated feature into the two second fully connected layers to obtain a second phenotype prediction value, and calculate the second loss based on the second phenotype prediction value and the corresponding true phenotype data.

[0035] Furthermore, the second loss is the prediction regression loss, which is expressed as:

[0036] Among them, Loss2 is the second loss, M is the number of samples, Y i is the true phenotypic data of node sample i, Yi ″ is the second phenotype prediction value of node sample i.

[0037] According to a second aspect of an embodiment of the present invention, a plant phenotype prediction device is provided, comprising one or more processors and a memory. The memory is coupled to the processor and configured to store program data; the processor is configured to execute the program data to implement the plant phenotype prediction method described above.

[0038] According to a third aspect of an embodiment of the present invention, a computer-readable storage medium is provided, on which a program is stored, wherein when the program is executed by a processor, it is used to implement the above-mentioned plant phenotype prediction method.

[0039] By using multiple omics as nodes, constructing a graph structure based on the correlation between multiple omics, and using graph neural networks to implement the regression task of the entire graph, the graph neural network and Transformer are combined to explore the association between different omics, integrating the attention of the graph nodes into the attention update of the Transformer, and dynamically selecting the omics features that contribute more to phenotypic prediction based on the degree of correlation with the global features. This can effectively solve the problem of low accuracy of phenotypic prediction of a single omics. In other words, the present invention creatively uses the method of combining graph neural networks with Transformers to achieve phenotypic prediction of multiple omics, which is conducive to improving the effect of phenotypic prediction. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] FIG1 is a detailed diagram of the network structure of multi-omics phenotype prediction in the present invention.

[0041] FIG2 is a flow chart of the plant phenotype prediction method of the present invention.

[0042] FIG3 is a test flow chart of the plant phenotype prediction method of the present invention.

[0043] FIG4 is a comparison diagram of the experimental results of the present invention.

[0044] FIG5 is a schematic structural diagram of a plant phenotype prediction device according to the present invention. DETAILED DESCRIPTION

[0045] Exemplary embodiments will be described in detail herein, examples of which are illustrated in the accompanying drawings. In the following description, when referring to the drawings, like numbers in different figures represent like or similar elements unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all possible embodiments consistent with the present invention. Rather, they are merely examples of apparatus and methods consistent with certain aspects of the present invention, as detailed in the appended claims.

[0046] The terms used in this invention are for the purpose of describing specific embodiments only and are not intended to limit the invention. The singular forms "a," "the," and "the" used in this invention and the appended claims are also intended to include plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" as used herein refers to and includes any or all possible combinations of one or more of the associated listed items.

[0047] It should be understood that although the terms "first," "second," "third," etc. may be used in the present invention to describe various information, such information should not be limited to these terms. These terms are merely used to distinguish information of the same type from one another. For example, first information may also be referred to as second information, and similarly, second information may also be referred to as first information, without departing from the scope of the present invention. Depending on the context, the term "if" as used herein may be interpreted as "when," "when," or "in response to determining."

[0048] The present invention will be described in detail below with reference to the accompanying drawings. Unless there is any conflict, the features of the following embodiments and implementations may be combined with each other.

[0049] 2 , the plant phenotype prediction method of the present invention specifically comprises the following steps:

[0050] (1) Obtain multi-omics data and phenotypic data of the target species, perform quality control and dimensionality reduction on the multi-omics data to obtain multiple omics features, and process outliers and missing values ​​on the phenotypic data to obtain true phenotypic data.

[0051] Furthermore, multi-omics data includes but is not limited to genomic data, transcriptomic data, metabolomic data, etc., that is, multi-omics includes but is not limited to genomics, transcriptomic, metabolomic, etc.

[0052] In this example, an open-source tomato dataset was used, from which multi-omics and phenotypic data of the target species can be obtained. The open-source tomato dataset includes 332 tomato accessions, each containing 6,971,059 SNPs (single nucleotide polymorphisms), 657,549 indels (base insertions and deletions), 54,838 SVs (structural variations), and approximately 170,000 RAN-Seg gene expression data.

[0053] Taking the acquisition of multi-omics data and phenotypic data of the target species from the open source tomato dataset as an example, this paper introduces in detail how to perform quality control and dimensionality reduction on the multi-omics data, and how to handle outliers and missing values ​​on the phenotypic data. In particular, step (1) can specifically include the following sub-steps:

[0054] (1.1) Remove missing values ​​and abnormal outliers from the phenotypic data of each sample of the target species in the tomato dataset;

[0055] (1.2) Delete the multi-omics data of each sample of the target species in the tomato dataset at each position where the missing rate is greater than 90%, and delete the sites where the minimum allele frequency in the genomic data is less than 5% to complete the quality control filtering of the multi-omics data;

[0056] (1.3) Encoding the multi-omics data after quality control filtering in step (1.2), wherein genomic data are coded as -1, 0, and 1 according to the homozygous state, heterozygous state, and homozygous variant state, respectively, and other omics data are not coded according to the measured values;

[0057] (1.4) For the data at each position of each omics data, calculate the Pearson correlation coefficient between the coding features of all samples and the corresponding phenotypic data to obtain the degree of correlation between each position of each omics data and the phenotypic data. Sorting the data from large to small according to the degree of correlation, select the first L positions of each omics data as the omics features corresponding to the omics data to obtain multiple omics features.

[0058] It should be understood that selecting the top L positions with the highest degree of correlation for each omics data set as the corresponding omics features for that omics data achieves dimensionality reduction for the omics data. In the tomato data of this example, there are four omics, so each plant sample will obtain four omics features, which can be summarized and expressed as a 4*L omics feature. The value of L can be selected based on the actual needs of the actual application scenario. For example, in this example, the top L = 1000 positions of each omics data set are selected as the omics features after dimensionality reduction.

[0059] Furthermore, the calculation formula of the Pearson correlation coefficient is:

[0060] Among them, ρ XYis the Pearson correlation coefficient, X and Y represent the coding characteristics and the corresponding phenotypic data, cov(*) represents the covariance, and σ represents the standard deviation. ρ XY is the Pearson correlation coefficient calculated by the above formula. The value range of this coefficient is -1 to 1. A Pearson correlation coefficient less than 0 indicates a negative correlation, a Pearson correlation coefficient greater than 0 indicates a positive correlation, and a Pearson correlation coefficient equal to 0 indicates no correlation.

[0061] It should be understood that the larger the absolute value of the Pearson correlation coefficient, the closer the correlation between the two variables, that is, the greater the degree of association. In this way, the degree of association between each position of each omics data and the phenotypic data can be obtained based on the Pearson correlation coefficient.

[0062] (2) Normalizing the omics features obtained in step (1), and then calculating the first similarity between the multiple omics features of each sample of the target species, and using the first similarity to construct graph structure data for each sample. The nodes in the graph represent different omics, and the edges in the graph represent the first similarities between different omics.

[0063] (2.1) Each omics feature obtained in step (1.3) is normalized as follows. Assume there are M samples, each sample has N omics, and the jth omics feature after dimensionality reduction of the i-th sample is denoted as fi j , i = 1…M, j = 1…N, and normalized according to the following formula:

[0064] Among them, f′ ij represents the jth omics feature after normalization of the i-th sample, u j is the mean of all samples of the j-th omics, var j is the standard deviation of all samples of the j-th omics.

[0065] It should be understood that each sample has N types of omics, that is, each sample, or each plant, has N types of omics data. For example, in this embodiment, M=332 and N=4.

[0066] (2.2) For each sample, the cosine similarity between the normalized omics features calculated in step (2.1) is used as the first similarity. The calculation formula is as follows:

[0067] Among them, similarity is the cosine similarity, A and B represent the normalized omics features, A·B is the inner product, and ‖A‖ represents the norm corresponding to the normalized omics feature A.

[0068] It should be understood that the value range of similarity is -1 to 1. The larger the value of similarity, the higher the feature similarity and the higher the degree of association.

[0069] (2.3) For each sample, the different omics features of the sample are used as graph nodes, and the first similarities between the different omics features obtained in step (2.2) are used as the edge weights of the graph to construct the graph structure data of each sample.

[0070] It should be understood that the weight of the edge reflects the degree of association between different omics, that is, the greater the first similarity, the higher the degree of association between different omics.

[0071] (3) For each sample, the graph structure data obtained in step (2) is input into a three-layer graph neural network to extract the features of each node, where the graph neural network includes a graph convolution layer and a ReLU activation layer.

[0072] Specifically, in a three-layer graph neural network, each graph convolution layer is followed by a ReLU activation layer, and the input and output feature dimensions of the three graph convolution layers are set to, for example, 1000. The graph structure data of each sample is input into the three-layer graph neural network to extract the features of each node, as shown in Figure 1.

[0073] (4) For each sample, the features of each node of the sample extracted in step (3) are spliced ​​in the order of multi-omics to obtain a first spliced ​​feature; the first spliced ​​feature is input into the two first fully connected layers to obtain a first phenotype prediction value; and the first loss is calculated based on the first phenotype prediction value and the corresponding true phenotypic data.

[0074] It should be understood that the true phenotypic data is the phenotypic data after outlier and missing value processing in step (1).

[0075] Specifically, taking the open source tomato dataset in this embodiment as an example, for each sample, the features of each node extracted in step (3) are spliced ​​in the order of the four omics of SNPs, InDels, SVs, and RAN-Seg to obtain a 4000-dimensional first spliced ​​feature. Then, the 4000-dimensional first spliced ​​feature is input into two first fully connected layers, and the first phenotype prediction value is output, as shown in Figure 1. Among them, the input dimension of the first first fully connected layer is 4000 and the output dimension is 32; the input dimension of the second first fully connected layer is 32 and the output dimension is 1. Finally, the first loss is calculated based on the first phenotype prediction value and the corresponding true phenotypic data.

[0076] In particular, the first loss can include two parts, namely, prediction regression loss and Pearson correlation coefficient loss, and its expression can be specifically as follows:

[0077] Wherein, Loss1 represents the first loss; M is the number of samples, which is 332 in this embodiment; Y i is the true phenotypic data of node sample i; Y′ i is the first phenotype prediction value of node sample i; is the true phenotypic data mean of all samples; is the mean of the first phenotype prediction values ​​of all samples; α is the adjustment coefficient between the prediction regression loss and the Pearson correlation coefficient loss, ranging from 0 to 1. In this embodiment, it is 0.5. Of course, other values ​​can also be taken according to actual needs.

[0078] (5) For each sample, the global identity embedding feature and the feature of each node of the sample extracted in step (3) are input into the Transformer network as the local identity embedding feature to obtain the updated global identity embedding feature and the local identity embedding feature of each node, and the second similarity between the global identity embedding feature and the local identity embedding feature of each node is calculated. The local identity embedding features and the global identity embedding features of the two nodes with the largest second similarity are selected and spliced ​​and then input into the two second fully connected layers to obtain the second phenotype prediction value, and the second loss is calculated based on the second phenotype prediction value and the corresponding true phenotype data.

[0079] (5.1) The global identifier embedding feature can be generated by initializing the parameters of the random normal distribution. Specifically, the value of each dimension of the global identifier embedding feature comes from a random value generated by the normal distribution, and the feature length of the global identifier feature is the same as the feature length of each node extracted in step (3).

[0080] Specifically, the xavier_normal function of the Pytorch framework can be used to obtain a random normal distribution, and then the global identifier embedding feature can be generated by initializing the random normal distribution parameters. The feature length of the global identifier embedding feature is the same as the feature length of each node extracted in step (3), that is, the feature length of the global identifier embedding feature is 1000.

[0081] (5.2) For each sample, the first attention matrix of the sample is calculated based on the features of all nodes of the sample extracted in step (3). The calculation formula is: Attention(F i ,F i )=softmax(F i F i T )

[0082] Among them, Attention(Fi ,F i ) represents the first attention matrix of the i-th sample; softmax(·) represents the softmax activation function; F i represents the feature matrix of the sample composed of the features of all nodes of the sample obtained in step (3), where the feature vector of each node is F ij , in this embodiment, j=1, 2, 3, 4.

[0083] (5.3) For each sample, the features F of each node of the sample extracted in step (3) are ij As the local identity embedding feature, the global identity embedding feature G obtained in step (5.1) and the local identity embedding feature of each node are input into the Transformer network. The Transformer network includes a self-attention module and a feedforward neural network module. In this way, the second attention matrix corresponding to the sample can be calculated by the self-attention module, and the feature is updated by combining the second attention matrix with the first attention matrix obtained in step (5.2). The updated global identity embedding feature and the local identity embedding feature of each node are obtained through the feedforward neural network module.

[0084] The calculation formula of the second attention matrix is: Q=[G,F i1 ,F i2 ,F i3 ,F i4 ] K=[G,F i1 ,F i2 ,F i3 ,F i4 ]

[0085] Among them, Attention(Q,K) represents the second attention matrix of the sample, Q and K are the feature matrices composed of the global identification embedding features and the local identification features of each node, respectively. k It is the characteristic dimension of the feature matrix Q,K.

[0086] The updated global identity embedding features and local identity embedding features of each node can be expressed as: V′=(Attention(F,F)+Attention(Q,K))*VV=[G,F i1 ,F i2 ,F i3 ,F i4 ]

[0087] Among them, V′ represents the updated global logo embedding feature and the local logo embedding feature of each node, and V represents the global logo embedding feature G and the local logo embedding feature F of each node before the update. ij .

[0088] (5.4) For each sample, calculate the cosine similarity between the updated global identity embedding feature obtained in step (5.3) and the local identity embedding feature of each node as the second similarity, select the local identity embedding features and the global identity embedding features of the two nodes with the largest second similarity, and concatenate them to obtain the second concatenated feature. Input the second concatenated feature into the two second fully connected layers to obtain the second phenotype prediction value; and calculate the second loss based on the second phenotype prediction value and the corresponding true phenotype data.

[0089] Specifically, the second similarity between each omics feature and the global feature is calculated, that is, the cosine similarity between the updated global identity embedding feature obtained in step (5.3) and the local identity embedding feature of each node is calculated. The cosine similarity is the second similarity, and the local identity embedding features and the global identity embedding features of the two nodes with the largest second similarity are selected for splicing to obtain a 3000-dimensional second splicing feature. The splicing method is the same as the splicing method in step (4). Then, the second splicing feature is input into two second fully connected layers to obtain a second phenotype prediction value, as shown in Figure 1. The input dimension of the first second fully connected layer is 3000 and the output dimension is 32; the input dimension of the second second fully connected layer is 32 and the output dimension is 1. Finally, the second loss is calculated based on the second phenotype prediction value and the corresponding true phenotypic data. In particular, the second loss can be a prediction regression loss, and its expression can be:

[0090] Among them, Loss2 is the second loss; M is the number of samples, which is 332 in this embodiment; Y i is the true phenotypic data of node sample i; Y i ″ is the second phenotype prediction value of node sample i, which is predicted by the Transformer network.

[0091] (6) Repeat steps (3) to (5) for iterative training, and use the stochastic gradient descent method to adjust the network parameters of the three-layer graph neural network in step (3), the two first fully connected layers in step (4), and the Transformer network and the two second fully connected layers in step (5) based on the total loss of the sum of the first loss and the second loss, so as to obtain the trained three-layer graph neural network and the two first fully connected layers, the Transformer network and the two second fully connected layers.

[0092] It should be understood that the iterative training is performed by repeating steps (3) to (5), and the training is stopped after reaching the set number of iterations to obtain the final network parameters. For example, the number of iterations in this embodiment can be set to 100.

[0093] (7) For the material to be tested, the first phenotypic prediction value is obtained through the trained three-layer graph neural network and two first fully connected layers, and the second phenotypic prediction value is obtained through the trained three-layer graph neural network, Transformer network and two second fully connected layers. The average of the first phenotypic prediction value and the second phenotypic prediction value is taken as the final predicted phenotypic result of the material to be tested.

[0094] In this embodiment, the process of using the trained three-layer graph neural network with two first fully connected layers, the Transformer network, and two second fully connected layers to predict phenotypes may include the following steps, as shown in FIG3 :

[0095] (7.1) Collect multi-omics data of the material to be tested: SNPs (single nucleotide polymorphisms), InDels (base insertions and deletions), SVs (structural variations), and RAN-Seg data.

[0096] (7.2) Delete the positions with a missing rate greater than 90% at each position in the multi-omics data of each sample, and delete the sites with a minimum allele frequency less than 5% in the genomic data to complete the quality control filtering of the multi-omics data.

[0097] (7.3) Encode the multi-omics data after quality control filtering, where the genomic data are encoded as -1, 0, and 1 according to the homozygous state, heterozygous state, and homozygous variant state, respectively, and other omics data are not encoded according to the measured values; then, each omics data is dimensionality reduced according to the L positions selected for each omics data in step (1.3), and the graph structure data is constructed according to the method of step (2).

[0098] (7.4) Input the graph structure data into the process of predicting phenotypes using a trained three-layer graph neural network and two first fully connected layers, a Transformer network, and two second fully connected layers, wherein the network parameters are the network parameters trained in step (6), and the predicted first phenotype prediction value and the second phenotype prediction value are obtained respectively, and the average of the first phenotype prediction value and the second phenotype prediction value is taken as the final predicted phenotypic result of the material to be tested.

[0099] The experimental results of this example on the tomato dataset are shown in Table 1, where SNPs (single nucleotide polymorphisms), InDels (base insertions and deletions), SVs (structural variations), and RAN-Seg gene expression are the results of the traditional RRBLUP method. In this example, the Pearson correlation coefficient results for each are shown in Figure 4. The method described in this example achieved improvements of 11.94% to 84.02% compared to the traditional method.

[0100] Table 1 Experimental results

[0101] Referring to Figure 5, an embodiment of the present invention provides a plant phenotype prediction device, which includes one or more processors and a memory, and the memory is coupled to the processor; wherein the memory is used to store program data, and the processor is used to execute the program data to implement the plant phenotype prediction method in the above embodiment.

[0102] The embodiment of the plant phenotype prediction device of the present invention can be applied to any device with data processing capabilities, and the arbitrary device with data processing capabilities can be a device or apparatus such as a computer. The device embodiment can be implemented by software, or by hardware or a combination of software and hardware. Taking software implementation as an example, as a device in a logical sense, it is formed by the processor of any device with data processing capabilities in which it is located to read the corresponding computer program instructions in the non-volatile memory into the internal memory and run them. From the hardware level, as shown in Figure 5, it is a hardware structure diagram of any device with data processing capabilities in which the plant phenotype prediction device of the present invention is located. In addition to the processor, memory, network interface, and non-volatile memory shown in Figure 5, the arbitrary device with data processing capabilities in which the device in the embodiment is located can also include other hardware according to the actual function of the arbitrary device with data processing capabilities, which will not be described in detail.

[0103] The implementation process of the functions and effects of each unit in the above-mentioned device is specifically described in the implementation process of the corresponding steps in the above-mentioned method, and will not be repeated here.

[0104] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to the partial description of the method embodiments. The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the present invention. A person of ordinary skill in the art can understand and implement the present invention without inventive work.

[0105] An embodiment of the present invention further provides a computer-readable storage medium having a program stored thereon. When the program is executed by a processor, the plant phenotype prediction method in the above embodiment is implemented.

[0106] The computer-readable storage medium may be an internal storage unit of any device with data processing capabilities described in any of the aforementioned embodiments, such as a hard disk or memory. The computer-readable storage medium may also be any device with data processing capabilities, such as a plug-in hard disk, a smart media card (SMC), an SD card, a flash card, etc. equipped on the device. Furthermore, the computer-readable storage medium may also include both an internal storage unit of any device with data processing capabilities and an external storage device. The computer-readable storage medium is used to store the computer program and other programs and data required by any device with data processing capabilities, and may also be used to temporarily store data that has been output or is to be output.

[0107] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A method for predicting plant phenotypes, The following steps are involved: (1) obtaining multi-omics data and phenotypic data of the target species, performing quality control and dimensionality reduction processing on the multi-omics data to obtain multiple omics features, and performing outlier and missing value processing on the phenotypic data to obtain true phenotypic data; (2) normalizing the omics features obtained in step (1), calculating a first similarity between multiple omics features of each sample of the target species, and constructing graph structure data of each sample using the first similarity; (3) For each sample, the graph structure data obtained in step (2) is input into a three-layer graph neural network to extract the features of each node; (4) for each sample, the features of each node of the sample extracted in step (3) are spliced ​​in a multi-omics order and input into two first fully connected layers to obtain a first phenotype prediction value, and a first loss is calculated based on the first phenotype prediction value and the corresponding true phenotype data; (5) For each sample, the features of each node of the sample obtained in step (3) are input into the Transformer network as local identifier embedding features and global identifier embedding features, and the updated global identifier embedding features and the local identifier embedding features of each node are obtained, and the second similarity between the updated global identifier embedding features and the local identifier embedding features of each node is calculated, and the local identifier embedding features and the global identifier embedding features of the two nodes with the largest second similarity are selected and spliced ​​to input into two second fully connected layers to obtain a second phenotype prediction value, and a second loss is calculated according to the second phenotype prediction value and the corresponding true phenotype data; (6) Repeat steps (3) to (5) to perform training, and adjust the network parameters of the graph neural network, the first fully connected layer, the Transformer network, and the second fully connected layer by back-propagation using the stochastic gradient descent method according to the total loss of the sum of the first loss and the second loss, so as to obtain the trained graph neural network, the first fully connected layer, the Transformer network, and the second fully connected layer; (7) For the material to be tested, a first phenotypic prediction value is obtained through the trained graph neural network and the first fully connected layer, and a second phenotypic prediction value is obtained through the trained graph neural network, the Transformer network and the second fully connected layer. The first phenotypic prediction value and the second phenotypic prediction value are averaged as the final predicted phenotypic result of the material to be tested.

2. The plant phenotype prediction method according to claim 1, It is characterized in that The multi-omics data includes genomic data, transcriptomic data and metabolomic data; The graph neural network includes a graph convolution layer and a ReLU activation layer.

3. The plant phenotype prediction method according to claim 1, It is characterized in that In the step (1), quality control and dimensionality reduction processing are performed on the multi-omics data to obtain multiple omics features, and outlier and missing value processing is performed on the phenotypic data to obtain true phenotypic data, which specifically includes the following sub-steps: (1.1) Remove missing values ​​and abnormal outliers from the phenotypic data of each sample of the target species; (1.2) Deleting positions with a missing rate greater than 90% in the multi-omics data of each sample of the target species, and deleting sites with a minimum allele frequency less than 5% in the genomic data, to complete quality control filtering of the multi-omics data; (1.3) encoding the multi-omics data after quality control filtration in step (1.2), wherein the genomic data is encoded as -1, 0, and 1 according to the homozygous state, heterozygous state, and homozygous variant state, respectively, and other omics data are not encoded according to the measured values; (1.4) For the data at each position of each omics data, the Pearson correlation coefficient between the coding features of all samples and the corresponding phenotypic data is calculated to obtain the degree of correlation between each position of each omics data and the phenotypic data. For each omics data, the L positions with the largest degree of correlation are selected as omics features to obtain multiple omics features.

4. The plant phenotype prediction method according to claim 3, It is characterized in that The calculation formula of the Pearson correlation coefficient is: Among them, ρ XY is the Pearson correlation coefficient, X and Y represent the coding features and the corresponding phenotypic data, respectively. cov(*) represents covariance, σ represents standard deviation.

5. The plant phenotype prediction method according to claim 1, It is characterized in that The step (2) comprises the following sub-steps: (2.1) Each omics feature obtained in step (1.3) is normalized according to the following formula: Among them, f ij represents the j-th omics feature of the i-th sample, i = 1…M, M represents the number of samples, j = 1…N, N represents the number of omics; f′ ij represents the j-th omics feature of the i-th sample after normalization. u j is the mean of all samples of the j-th omics, var j is the standard deviation of all samples of the j-th omics; (2.2) For each sample, the cosine similarity between the normalized omics features calculated in step (2.1) is used as the first similarity, and the calculation formula is as follows: Among them, similarity is the cosine similarity, A and B represent the omics features after normalization. A·B is the inner product, ‖A‖ represents the norm corresponding to the omics feature A after normalization; (2.3) For each sample, the graph structure data of each sample is constructed by taking the different omics of the sample as graph nodes and taking the first similarity between the different omics features obtained in step (2.2) as the edge weight of the graph.

6. The plant phenotype prediction method according to claim 1, It is characterized in that The first loss includes prediction regression loss and Pearson correlation coefficient loss, and its expression is: Among them, Loss1 represents the first loss, M is the sample size, Y i is the true phenotypic data of node sample i, Y′ i is the first phenotype prediction value of node sample i, is the true phenotypic data mean of all samples, is the mean of the first phenotype prediction values ​​of all samples, α is the adjustment coefficient between the prediction regression loss and the Pearson correlation coefficient loss.

7. The plant phenotype prediction method according to claim 1, It is characterized in that The step (5) comprises the following sub-steps: (5.1) Generate a global identity embedding feature using random normal distribution parameter initialization. The value of each dimension of the global identity embedding feature comes from a random value generated by a normal distribution, and its feature length is the same as the feature length of each node extracted in step (3); (5.2) For each sample, calculate the first attention matrix of the sample according to the features of all nodes of the sample extracted in step (3); (5.3) For each sample, the features of each node of the sample extracted in step (3) are used as local identification embedding features, and the global identification embedding features and the local identification embedding features of each node obtained in step (5.1) are input into a Transformer network, the Transformer network includes a self-attention module and a feedforward neural network module, the second attention matrix corresponding to the sample is calculated by the self-attention module, the second attention matrix is ​​combined with the first attention matrix obtained in step (5.2) to update the global identification embedding features and the local identification embedding features of each node, and the updated global identification embedding features and the local identification embedding features of each node are obtained by the feedforward neural network module; (5.4) For each sample, the cosine similarity between the updated global identifier embedding feature obtained in step (5.3) and the local identifier embedding feature of each node is calculated as the second similarity, the local identifier embedding features and the global identifier embedding features of the two nodes with the largest second similarity are selected and concatenated to obtain a second concatenated feature, the second concatenated feature is input into two second fully connected layers, a second phenotype prediction value is obtained, and a second loss is calculated based on the second phenotype prediction value and the corresponding true phenotype data.

8. The plant phenotype prediction method according to claim 1 or 7, It is characterized in that The second loss is the prediction regression loss, which is expressed as: Among them, Loss2 is the second loss, M is the sample size, Y i is the true phenotypic data of node sample i, Y″ i is the second phenotype prediction value of node sample i.

9. A plant phenotype prediction device, comprising one or more processors and a memory coupled to the processors, It is characterized in that The memory is used to store program data. The processor is used to execute the program data to implement the plant phenotype prediction method according to any one of claims 1 to 8.

10. A computer-readable storage medium having a program stored thereon, wherein when the program is executed by a processor, the program is used to implement the plant phenotype prediction method according to any one of claims 1 to 8.