Field crop gene data prediction method, device, electronic device and medium
Through the prediction neural network model combining the two-dimensional coding matrix and the embedding layer, the problems of high coding dimension of long DNA sequences and cumbersome training of multi-environment models are solved, and efficient genetic data prediction is achieved.
Patent Information
- Application Number
- CN202410929565.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-11
- Publication Date
- 2025-09-16
- Estimated Expiration
- 2044-07-11
AI Technical Summary
The encoding method used by existing technologies when processing long DNA or protein sequences results in excessively high data input dimensions, and existing prediction models need to be trained separately for different environmental and climatic conditions, resulting in inefficient methods for predicting field crop genetic data.
A two-dimensional coding matrix based on the frequency and order relationship of gene words is used, and combined with the embedding layer of environmental and climatic factors, a predictive neural network model is established. The two-dimensional coding matrix is input into the predictive neural network, and only one model is needed to adapt to different environmental and climatic conditions.
It effectively reduces the input dimension of the encoded data, improves the model training and processing speed, and can accurately predict the relationship between field crop genetic data and phenotypic traits under different environmental and climatic conditions, reducing the tediousness of the prediction work.
Smart Images

Figure CN118800321B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of crop gene technology, and in particular to a method, device, electronic equipment and medium for predicting field crop gene data. Background Art
[0002] Current methods for exploring the relationship between genetic data and phenotypic traits in field crops often suffer from the following two shortcomings:
[0003] On the one hand, when processing DNA sequences or protein sequences, the currently more common encoding methods can complete the encoding process more simply by directly assigning values or taking 1 in the corresponding dimension. However, the input length of the data after these two encodings is equal to the length of the gene sequence. When the gene sequence is too long, the data input dimension will be too high, which will make the common machine learning processing time too long or even fail.
[0004] On the other hand, existing prediction models only consider the relationship between genetic data and phenotypic traits of field crops under a single environmental and climatic condition. Training separate machine learning models for different environmental and climatic conditions is cumbersome. Therefore, an efficient prediction method for field crop genetic data is urgently needed. Summary of the Invention
[0005] In view of this, it is necessary to provide a method, device, electronic device and medium for predicting field crop genetic data to solve the problem that the methods for predicting field crop genetic data in the prior art are relatively inefficient.
[0006] In order to achieve the above technical objectives, the present invention adopts the following technical solutions:
[0007] In a first aspect, the present invention provides a method for predicting field crop gene data, comprising:
[0008] Obtain the variant position and specific variant information of each sample from the VCF variant information file, then combine the reference genome information on both sides of this variant information. Each SNP site forms a 5nt variant unit. These 5nt variant units are connected in the order of the size of the variant sites in the VCF to form the original gene sequence;
[0009] According to the original gene sequence and a preset word length, dividing the original gene sequence into a plurality of gene words based on the preset word length to obtain a gene word sequence;
[0010] Establishing a two-dimensional coding matrix based on the frequency and order relationship of each gene word in the gene word sequence;
[0011] Encoding environmental climate factors and establishing an embedding layer, and establishing a prediction neural network model based on the preset word length and the embedding layer;
[0012] The two-dimensional coding matrix is input into the prediction neural network model to obtain a prediction result.
[0013] Furthermore, the two-dimensional coding matrix is established according to the frequency and order relationship of each gene word in the gene word sequence, including:
[0014] According to the preset word length, all types of gene words are obtained based on the types of gene bases;
[0015] Establish a mapping relationship between each gene word and each row in the two-dimensional encoding matrix, and establish a mapping relationship between each gene word and each column in the two-dimensional encoding matrix;
[0016] Based on the mapping relationship, the values of the matrix elements are obtained according to the gene word sequence, and the two-dimensional coding matrix M is established;
[0017] Among them, the matrix element M in the two-dimensional coding matrix ij Used to represent: in the gene word sequence, the frequency of occurrence of the gene word corresponding to the jth column of the two-dimensional encoding matrix after the gene word corresponding to the i-th row of the two-dimensional encoding matrix appears for the first time, where i and j are both positive integers.
[0018] Furthermore, the establishing of a mapping relationship between each gene word and each row in the two-dimensional coding matrix, and the establishing of a mapping relationship between each gene word and each column in the two-dimensional coding matrix, include:
[0019] All types of gene words are sequentially coded, with each code number corresponding to a gene word, and the code numbers are all natural numbers;
[0020] Establish a mapping relationship between the gene word with code number a and the a+1th row in the two-dimensional coding matrix;
[0021] Establish a mapping relationship between the gene word with coding number a and the a+1th column in the two-dimensional coding matrix.
[0022] Furthermore, the environmental climate factors are encoded and an embedding layer is established, and a prediction neural network model is established based on the preset word length and the embedding layer, including:
[0023] Encoding environmental climate factors to obtain a multidimensional vector, and establishing the embedding layer according to the multidimensional vector;
[0024] According to the preset word length, a plurality of residual blocks are established and an arrangement order among the plurality of residual blocks is obtained;
[0025] The prediction neural network model is established according to the embedding layer, the plurality of residual blocks and the arrangement order between the residual blocks.
[0026] Furthermore, encoding the environmental climate factors to obtain a multidimensional vector, and establishing the embedding layer according to the multidimensional vector, including:
[0027] Sequentially encode various environmental and climatic factors to obtain coded values;
[0028] Performing embedding encoding on the encoded value to obtain a multi-dimensional vector;
[0029] The embedding layer is established with the multi-dimensional vector as an output terminal.
[0030] Furthermore, the prediction neural network model includes two first convolutional layers, a second convolutional layer, a maximum pooling layer, the embedding layer, multiple residual blocks, an average pooling layer and a fully connected layer; the input end of the first convolutional layer is used to input the two-dimensional coding matrix, and the output end of the first convolutional layer is connected to the input end of the second convolutional layer; the output end of the second convolutional layer is connected to the input end of the maximum pooling layer, and the output end of the second convolutional layer and the output end of the embedding layer are cross-connected to form a cross product output end; multiple residual blocks are connected in sequence according to the arrangement order between the multiple residual blocks, the input end of the residual block at the head end is connected to the cross product output end, and the output end of the residual block at the end is connected to the input end of the average pooling layer; the output end of the average pooling layer is connected to the input end of the fully connected layer, and the output end of the fully connected layer is used to output the prediction result.
[0031] Furthermore, the multiple residual blocks include a first residual block and a second residual block; the first residual block includes a third convolutional layer and a fourth convolutional layer, the input end of the third convolutional layer is the input end of the first residual block, the output end of the third convolutional layer is connected to the input end of the fourth convolutional layer, the input end of the third convolutional layer and the output end of the fourth convolutional layer are added and connected to form the output end of the first residual block; the second residual block includes a fifth convolutional layer, a sixth convolutional layer and a seventh convolutional layer, the input end of the fifth convolutional layer is the input end of the second residual block, the output end of the fifth convolutional layer is connected to the input end of the sixth convolutional layer, and the input end of the fifth convolutional layer is also connected to the input end of the seventh convolutional layer; the output end of the sixth convolutional layer and the output end of the seventh convolutional layer are added and connected to form the output end of the second residual block.
[0032] In a second aspect, the present invention further provides a field crop gene data prediction device, comprising:
[0033] a word division module, configured to obtain an original gene sequence and a preset word length, and divide the original gene sequence into a plurality of gene words based on the preset word length to obtain a gene word sequence;
[0034] A gene encoding module, configured to establish a two-dimensional encoding matrix based on the frequency and order of occurrence of each gene word in the gene word sequence;
[0035] A model building module is used to encode environmental climate factors and establish an embedding layer, and to establish a prediction neural network model based on the preset word length and the embedding layer;
[0036] The gene prediction module is used to input the two-dimensional coding matrix into the prediction neural network model to obtain a prediction result.
[0037] In a third aspect, the present invention further provides an electronic device comprising a memory and a processor, wherein:
[0038] Memory, used to store programs;
[0039] The processor is coupled to the memory and is used to execute the program stored in the memory to implement the steps in the field crop gene data prediction method in any of the above implementations.
[0040] In a fourth aspect, the present invention also provides a computer-readable storage medium for storing computer-readable programs or instructions. When the programs or instructions are executed by a processor, the steps in the field crop genetic data prediction method in any of the above-mentioned implementation methods can be implemented.
[0041] The present invention provides a method, device, electronic device, and medium for predicting gene data for field crops. The method first extracts 2 bp of base information before and after the corresponding position of the reference genome based on the SNP site in the VCF file, and then assembles the base information into a sequence in order of position. The method then divides the original gene sequence into multiple gene words based on the preset word length according to the original gene sequence and a preset word length, obtaining a gene word sequence. A two-dimensional encoding matrix is then established based on the frequency and order relationship of each gene word in the gene word sequence. Environmental and climatic factors are then encoded and an embedding layer is established. A prediction neural network model is established based on the preset word length and the embedding layer. Finally, the two-dimensional encoding matrix is input into the prediction neural network model to obtain a prediction result. Compared to the prior art, the present invention adopts a two-dimensional matrix encoding method that combines the frequency and order relationship of gene words at the encoding level, so that the encoding retains the information on the mutual relationship between gene words. At the same time, the encoding can effectively reduce the input dimension of the encoded data, allowing subsequent model training to be performed more quickly, reducing the training and processing time of the prediction neural network model. In addition, the present invention also establishes an embedding layer based on environmental and climatic factors, and then establishes a predictive neural network model in combination with the embedding layer, so that the entire process only requires training one model to obtain the relationship between field crop genetic data and phenotypic traits under different environmental and climatic conditions. Combined with the encoding method in the present invention, the problem of relatively low efficiency of the existing method for predicting field crop genetic data is solved. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Figure 1 A method flow chart of an embodiment of a method for predicting field crop gene data provided by the present invention;
[0043] Figure 2 for Figure 1 A method flow chart of an embodiment of step S103;
[0044] Figure 3 A schematic diagram of the gene sequence encoding process in one embodiment of the field crop gene data prediction method provided by the present invention;
[0045] Figure 4 for Figure 1 A method flow chart of an embodiment of step S104;
[0046] Figure 5 A schematic diagram of the structure of a prediction neural network in an embodiment of the field crop gene data prediction method provided by the present invention;
[0047] Figure 6 A schematic diagram of the structure of a prediction neural network in another embodiment of the field crop gene data prediction method provided by the present invention;
[0048] Figure 7This is a schematic structural diagram of an embodiment of a field crop gene data prediction device provided by the present invention;
[0049] Figure 8 This is a structural diagram of an embodiment of an electronic device provided by the present invention. DETAILED DESCRIPTION
[0050] The preferred embodiments of the present invention will be described in detail below in conjunction with the accompanying drawings, wherein the accompanying drawings constitute a part of this application and are used together with the embodiments of the present invention to illustrate the principles of the present invention, and are not used to limit the scope of the present invention.
[0051] Before describing the specific embodiments, this article first explains some technical terms in the article:
[0052] K-mer encoding: A K-mer is a DNA fragment of length k, obtained by shearing a portion of a gene sequence. k is an integer, and the number of k equals the number of mer.
[0053] ResNet model: ResNet neural network, also known as residual neural network, adds the idea of residual learning to the traditional convolutional neural network, solving the problems of gradient diffusion and precision degradation (training set) in deep networks, allowing the network to become deeper and deeper, ensuring both precision and speed.
[0054] Embedding layer: The embedding layer is a layer in the neural network structure, consisting of several neurons.
[0055] It is understandable that other technical terms, English abbreviations, etc. appearing in the following text are all prior art, and those skilled in the art can understand their meanings based on the context. Due to space constraints, they are not explained in detail in this article.
[0056] In the description of the present application, “plurality” means two or more, unless otherwise clearly defined.
[0057] References herein to "embodiments" mean that a particular feature, structure, or characteristic described in connection with the embodiments may be included in at least one embodiment of the present invention. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor does it constitute a separate or alternative embodiment that is mutually exclusive of other embodiments. It is understood, both explicitly and implicitly, by those skilled in the art that the embodiments described herein may be combined with other embodiments.
[0058] On the one hand, the present invention combines the frequency and sequence relationship of gene words at the coding level and encodes them in the form of a two-dimensional matrix, reducing the dimension of the input data after encoding while retaining the sequence relationship between genes. On the other hand, at the level of establishing the prediction model, the present invention combines environmental and climatic factors for establishment. Only one model is needed to adapt to different environmental predictions, greatly reducing the tediousness of the prediction work and improving efficiency.
[0059] The present invention provides a method, device, electronic device and storage medium for predicting gene data of field crops, which are described below respectively.
[0060] Combine Figure 1 As shown, a specific embodiment of the present invention discloses a method for predicting field crop gene data, comprising:
[0061] S101. Obtain the mutation position and specific mutation information of each sample from the VCF mutation information file, then combine the reference genome information on both sides of the mutation information. Each SNP site forms a 5nt mutation unit. These 5nt mutation units are connected in the order of the size of the mutation sites in the VCF to form the original gene sequence;
[0062] S102, according to the original gene sequence and a preset word length, dividing the original gene sequence into a plurality of gene words based on the preset word length to obtain a gene word sequence;
[0063] S103, establishing a two-dimensional coding matrix according to the frequency and order of occurrence of each gene word in the gene word sequence;
[0064] S104, encoding the environmental climate factors and establishing an embedding layer, and establishing a prediction neural network model based on the preset word length and the embedding layer;
[0065] S105: Input the two-dimensional coding matrix into the prediction neural network model to obtain a prediction result.
[0066] Compared to existing technologies, the present invention employs a two-dimensional matrix encoding method that combines the frequency and order of gene words at the encoding level. This method preserves the interrelationships between gene words and effectively reduces the input dimensionality of the encoded data, enabling faster model training and reducing the training and processing time of the predictive neural network model. Furthermore, the present invention establishes an embedding layer based on environmental and climatic factors, and then uses this embedding layer to establish a predictive neural network model. This allows the entire process to train only a single model to capture the relationship between field crop genetic data and phenotypic traits under different environmental and climatic conditions. Combined with the encoding method of the present invention, this method addresses the relatively inefficient methods of predicting field crop genetic data in the prior art.
[0067] In the above step S102, the gene word is a short sequence composed of bases of a preset word length (four types: A, G, C, and T), and the gene word sequence is a sequence composed of multiple gene words obtained by dividing the original gene sequence according to the preset word length.
[0068] Combine Figure 2 As shown, in a preferred embodiment, the above step S103, based on the frequency and order of occurrence of each gene word in the gene word sequence, establishes a two-dimensional coding matrix, specifically including:
[0069] S201, obtaining all types of gene words based on the types of gene bases according to the preset word length;
[0070] S202, establishing a mapping relationship between each gene word and each row in the two-dimensional coding matrix, and establishing a mapping relationship between each gene word and each column in the two-dimensional coding matrix;
[0071] S203, based on the mapping relationship, obtain the value of the matrix element according to the gene word sequence, and establish the two-dimensional coding matrix M;
[0072] Among them, the matrix element M in the two-dimensional coding matrix ij Used to represent: in the gene word sequence, the frequency of occurrence of the gene word corresponding to the jth column of the two-dimensional encoding matrix after the gene word corresponding to the i-th row of the two-dimensional encoding matrix appears for the first time, where i and j are both positive integers.
[0073] Furthermore, in a preferred embodiment, the above step S202, establishing a mapping relationship between each gene word and each row in the two-dimensional encoding matrix, and establishing a mapping relationship between each gene word and each column in the two-dimensional encoding matrix, includes:
[0074] All types of gene words are sequentially coded, with each code number corresponding to a gene word, and the code numbers are all natural numbers;
[0075] Establish a mapping relationship between the gene word with code number a and the a+1th row in the two-dimensional coding matrix;
[0076] Establish a mapping relationship between the gene word with coding number a and the a+1th column in the two-dimensional coding matrix.
[0077] It will be appreciated that, in this embodiment, each row and column of the two-dimensional coding matrix corresponds to each gene word in the same manner. In practice, different mapping methods can be used for the rows and columns of the two-dimensional coding matrix, depending on specific needs, to represent different relationships between gene words. In this embodiment, each gene word is first sequentially encoded so that it can be directly mapped to the row and column numbers in the matrix, which is easier for people to understand. Furthermore, the code numbers are natural numbers, and starting from zero also facilitates code writing and computer processing.
[0078] Combine Figure 3 As shown, the present invention also provides a more detailed embodiment to more clearly illustrate the above steps S201-S202:
[0079] First, the original gene sequence is divided into a set of "words" (short for gene words) of length k (preset word length) (i.e. gene word sequence) by using k-mer counting. The number of word types divided out is at most 4 k kind.
[0080] Then arrange all possible words in a certain order to create 4 k *4 k A two-dimensional matrix M, where M ij It represents the frequency of the jth word appearing after the ith word appears in the original gene sequence.
[0081] For example, Figure 3 During the processing, set k=4, then the word set divided out will have at most 4 4 = 256 possible words, which can then be encoded sequentially into 0~255 (i.e. code numbers).
[0082] Then create a 256*256 two-dimensional matrix M. The meaning of the value a after encoding the 256 words can be understood as the row name of the a+1th row or the column name of the a+1th column. ij It represents the frequency of the jth word appearing after the ith word appears in the original gene sequence.
[0083] Through the above method, the original gene sequence can be uniformly converted into a 256*256 two-dimensional coding matrix, in which the matrix has the following characteristics:
[0084] 1. The sum of the data in the i-th row is equal to the sum of the data in the i-th column;
[0085] 2. If the sample sequences have the same length, the sum of all the data in the two-dimensional matrix M encoded by each sequence is the same value.
[0086] This greatly improves the efficiency of encoding and processing. It is also worth noting that, compared with the existing k-mer encoding, the encoding method in this embodiment can not only effectively reduce the input dimension of the encoded data, but also preserve the mutual information between "gene words".
[0087] Further, combined Figure 4 As shown, in a preferred embodiment, the above step S104, encoding the environmental climate factors and establishing an embedding layer, and establishing a prediction neural network model based on the preset word length in combination with the embedding layer, specifically includes:
[0088] S401, encoding environmental climate factors to obtain a multidimensional vector, and establishing the embedding layer according to the multidimensional vector;
[0089] S402: establishing a plurality of residual blocks according to the preset word length and obtaining an arrangement order among the plurality of residual blocks;
[0090] S403: Establish the prediction neural network model according to the embedding layer, the plurality of residual blocks, and the arrangement order between the residual blocks.
[0091] The embedding layer enables the prediction neural network to make predictions in combination with environmental and climatic factors. At the same time, the residual block part in the prediction neural network can be flexibly adjusted according to the preset word length.
[0092] Specifically, in a preferred embodiment, the above step S401, encoding the environmental climate factors to obtain a multidimensional vector, and establishing the embedding layer according to the multidimensional vector, specifically includes:
[0093] Sequentially encode various environmental and climatic factors to obtain coded values;
[0094] Performing embedding encoding on the encoded value to obtain a multi-dimensional vector;
[0095] The embedding layer is established with the multi-dimensional vector as an output terminal.
[0096] Furthermore, in a preferred embodiment, the prediction neural network model includes two first convolutional layers, a second convolutional layer, a maximum pooling layer, the embedding layer, multiple residual blocks, an average pooling layer and a fully connected layer; the input end of the first convolutional layer is used to input the two-dimensional coding matrix, and the output end of the first convolutional layer is connected to the input end of the second convolutional layer; the output end of the second convolutional layer is connected to the input end of the maximum pooling layer, and the output end of the second convolutional layer and the output end of the embedding layer are cross-connected to form a cross product output end; multiple residual blocks are connected in sequence according to the arrangement order between the multiple residual blocks, the input end of the residual block at the head end is connected to the cross product output end, and the output end of the residual block at the end is connected to the input end of the average pooling layer; the output end of the average pooling layer is connected to the input end of the fully connected layer, and the output end of the fully connected layer is used to output the prediction result.
[0097] Specifically, in a preferred embodiment, the multiple residual blocks include a first residual block and a second residual block; the first residual block includes a third convolutional layer and a fourth convolutional layer, the input end of the third convolutional layer is the input end of the first residual block, the output end of the third convolutional layer is connected to the input end of the fourth convolutional layer, the input end of the third convolutional layer and the output end of the fourth convolutional layer are added together to form the output end of the first residual block; the second residual block includes a fifth convolutional layer, a sixth convolutional layer and a seventh convolutional layer, the input end of the fifth convolutional layer is the input end of the second residual block, the output end of the fifth convolutional layer is connected to the input end of the sixth convolutional layer, and the input end of the fifth convolutional layer is also connected to the input end of the seventh convolutional layer; the output end of the sixth convolutional layer and the output end of the seventh convolutional layer are added together to form the output end of the second residual block.
[0098] The present invention also provides two more detailed embodiments to more clearly illustrate the above steps S104-S105:
[0099] At the model prediction level, to address the need to train separate machine learning models for different environmental and climatic conditions, this example designs a deep learning algorithm that embeds environmental and climatic factors. This allows a single model to be trained to capture the relationship between genetic data and phenotypic traits of field crops under diverse environmental and climatic conditions. Because this model learns from data under diverse conditions, it can achieve faster and better predictions than models trained under a single condition.
[0100] This example uses the ResNet model as its core. An embedding layer is added to the model to incorporate environmental and climate factors. For ease of understanding, this example introduces two different prediction neural network models, one for a preset word length of three and one for a preset word length of four.
[0101] In a preferred embodiment, when k=4 is set in the new encoding method (i.e., the preset word length is 4), the prediction neural network model is as follows: Figure 5 As shown, the processing process is:
[0102] At input, add a three-dimensional tensor of (1, 256, 256) channel dimensions to the two-dimensional encoding matrix.
[0103] First, the (1,256,256) tensor is convolved by the first convolutional layer (kernel_size=7,stride=1,padding=3) to obtain a (12,256,256) tensor, and then convolved by the second convolutional layer (kernel_size=7,stride=2,padding=3) to obtain a (64,128,128) tensor, and then passed through the maximum pooling layer (kernel_size=3,stride=2,padding=1) after maximum pooling to obtain a (64,64,64) tensor.
[0104] Afterwards, the sequentially encoded values of the regional information (a representation of environmental and climatic factors) are re-embedded, with each region encoded as a 64-dimensional vector (the multidimensional vector). At this point, the 64-dimensional tensor A output by the embedding layer is multiplied by the values in each dimension of tensor A, according to the channel dimensions of tensor B output by the max pooling layer, with the two-dimensional tensors corresponding to the corresponding channels of tensor B, resulting in a (64, 64, 64) tensor C (the output of the cross product).
[0105] After that, the (64,64,64) tensor C is first convolved by the third convolutional layer (kernel_size=3,stride=1,padding=1) to obtain a (64,64,64) tensor, and then convolved by the fourth convolutional layer (kernel_size=3,stride=1,padding=1) to obtain a (64,64,64) tensor C1. At this time, the (64,64,64) tensor C and the (64,64,64) tensor C1 are added according to the corresponding positions to obtain the (64,64,64) tensor D (that is, the output of the first residual block).
[0106] After that, the (64,64,64) tensor D is first convolved with (kernel_size=3,stride=1,padding=1) to obtain a (64,64,64) tensor, and then convolved with (kernel_size=3,stride=1,padding=1) to obtain a (64,64,64) tensor D1. At this time, the (64,64,64) tensor D and the (64,64,64) tensor D1 are added according to the corresponding positions to obtain a (64,64,64) tensor E (that is, the output of the second first residual block).
[0107] After that, the (64,64,64) tensor E is first convolved by the fifth convolutional layer (kernel_size=3,stride=2,padding=1) to obtain a (128,32,32) tensor, and then convolved by the sixth convolutional layer (kernel_size=3,stride=1,padding=1) to obtain a (128,32,32) tensor E1, and then the (64,64,64) tensor E is convolved by the seventh convolutional layer (kernel_size=1,stride=2,padding=0) to obtain a (128,32,32) tensor E2. At this time, the (128,32,32) tensor E1 and the (128,32,32) tensor E2 are added according to the corresponding positions to obtain a (128,32,32) tensor F (that is, the output of the first and second residual blocks).
[0108] After that, the (128, 32, 32) tensor F is first convolved with (kernel_size=3, stride=1, padding=1) to obtain a (128, 32, 32) tensor, and then convolved with (kernel_size=3, stride=1, padding=1) to obtain a (128, 32, 32) tensor F1. At this time, the (128, 32, 32) tensor F and the (128, 32, 32) tensor F1 are added according to the corresponding positions to obtain a (128, 32, 32) tensor G (that is, the output of the third first residual block).
[0109] After that, the (128, 32, 32) tensor G is first convolved with (kernel_size=3, stride=2, padding=1) to obtain a (256, 16, 16) tensor, and then convolved with (kernel_size=3, stride=1, padding=1) to obtain a (256, 16, 16) tensor G1. Then, the (128, 32, 32) tensor G is convolved with (kernel_size=1, stride=2, padding=0) to obtain a (256, 16, 16) tensor G2. At this time, the (256, 16, 16) tensor G1 and the (256, 16, 16) tensor G2 are added according to the corresponding positions to obtain the (256, 16, 16) tensor H (that is, the output of the second residual block).
[0110] After that, the (256, 16, 16) tensor H is first convolved with (kernel_size=3, stride=1, padding=1) to obtain a (256, 16, 16) tensor, and then convolved with (kernel_size=3, stride=1, padding=1) to obtain a (256, 16, 16) tensor H1. At this time, the (256, 16, 16) tensor H and the (256, 16, 16) tensor H1 are added according to the corresponding positions to obtain a (256, 16, 16) tensor I (that is, the output of the fourth first residual block).
[0111] After that, the (256,16,16) tensor I is first convolved with (kernel_size=3,stride=2,padding=1) to obtain a (512,8,8) tensor, and then convolved with (kernel_size=3,stride=1,padding=1) to obtain a (512,8,8) tensor I1. Then, the (256,16,16) tensor G is convolved with (kernel_size=1,stride=2,padding=0) to obtain a (512,8,8) tensor I2. At this time, the (512,8,8) tensor I1 and the (512,8,8) tensor I2 are added according to the corresponding positions to obtain the (512,8,8) tensor J (that is, the output of the third second residual block).
[0112] After that, the (512, 8, 8) tensor J is first convolved with (kernel_size=3, stride=1, padding=1) to obtain a (512, 8, 8) tensor, and then convolved with (kernel_size=3, stride=1, padding=1) to obtain a (512, 8, 8) tensor J1. At this time, the (512, 8, 8) tensor J and the (512, 8, 8) tensor J1 are added according to the corresponding positions to obtain the (512, 8, 8) tensor K (that is, the output of the fifth first residual block).
[0113] After that, the (512,8,8) tensor K is average pooled by the average pooling layer (kernel_size=8) to obtain a 512-dimensional tensor, and finally the fully connected output of the fully connected layer (512,1) is obtained to obtain the final prediction value.
[0114] In another preferred embodiment, when k=3 (i.e., the preset word length is 3) is set in the new encoding method, the prediction neural network model is as follows: Figure 6 As shown, the processing process is:
[0115] As input, add a three-dimensional tensor of (1, 64, 64) channel dimensions to the two-dimensional encoding matrix.
[0116] First, the (1, 64, 64) tensor is convolved by the first convolutional layer (kernel_size=7, stride=1, padding=3) to obtain a (12, 64, 64) tensor, and then convolved by the second convolutional layer (kernel_size=7, stride=2, padding=3) to obtain a (64, 32, 32) tensor, and then passed through the maximum pooling layer (kernel_size=3, stride=2, padding=1) after maximum pooling to obtain a (64, 16, 16) tensor.
[0117] Afterwards, the sequentially encoded values of the regional information (a representation of environmental and climatic factors) are re-embedded, with each region encoded as a 64-dimensional vector (the aforementioned multidimensional vector). At this point, the 64-dimensional tensor A output by the embedding layer is multiplied by the values in each dimension of tensor A, according to the channel dimensions of tensor B output by the max pooling layer, with the two-dimensional tensors corresponding to the corresponding channels of tensor B, resulting in a (64, 16, 16) tensor C (the output of the cross product).
[0118] After that, the (64, 16, 16) tensor C is first convolved by the third convolutional layer (kernel_size=3, stride=1, padding=1) to obtain a (64, 16, 16) tensor, and then convolved by the fourth convolutional layer (kernel_size=3, stride=1, padding=1) to obtain a (64, 16, 16) tensor C1. At this time, the (64, 16, 16) tensor C and the (64, 16, 16) tensor C1 are added according to the corresponding positions to obtain the (64, 16, 16) tensor D (that is, the output of the first residual block).
[0119] After that, the (64, 16, 16) tensor D is first convolved with (kernel_size=3, stride=1, padding=1) to obtain a (64, 16, 16) tensor, and then convolved with (kernel_size=3, stride=1, padding=1) to obtain a (64, 16, 16) tensor D1. At this time, the (64, 16, 16) tensor D and the (64, 16, 16) tensor D1 are added according to the corresponding positions to obtain a (64, 16, 16) tensor E (that is, the output of the second first residual block).
[0120] After that, the (64,16,16) tensor E is first convolved through the fifth convolutional layer (kernel_size=3,stride=2,padding=1) to obtain a (128,8,8) tensor, and then convolved through the sixth convolutional layer (kernel_size=3,stride=1,padding=1) to obtain a (128,8,8) tensor E1. Then, the (64,16,16) tensor E is convolved through the seventh convolutional layer (kernel_size=1,stride=2,padding=0) to obtain a (128,8,8) tensor E2. At this time, the (128,8,8) tensor E1 and the (128,8,8) tensor E2 are added according to the corresponding positions to obtain the (128,8,8) tensor F (that is, the output of the first and second residual blocks).
[0121] After that, the (128,8,8) tensor F is first convolved with (kernel_size=3,stride=1,padding=1) to obtain a (128,8,8) tensor, and then convolved with (kernel_size=3,stride=1,padding=1) to obtain a (128,8,8) tensor F1. At this time, the (128,8,8) tensor F and the (128,8,8) tensor F1 are added according to the corresponding positions to obtain a (128,8,8) tensor G (that is, the output of the third first residual block).
[0122] After that, the (128, 32, 32) tensor G is average pooled by the average pooling layer (kernel_size=8) to obtain a 128-dimensional tensor, and finally the fully connected output of the fully connected layer (128, 1) is obtained to obtain the final prediction value.
[0123] Experiments have shown that, first, data encoded using the encoding method of this embodiment achieves superior prediction results for long gene sequences, and prediction time is significantly reduced compared to sequential and one-hot encoding methods. Second, adding an embedding layer to incorporate environmental and climatic factors into the model can improve overall prediction performance. As shown in the table below, the prediction performance of the predictive neural network model in this embodiment for five provinces (i.e., the prediction results after adding location information in the table) is superior to the prediction results of each province trained and predicted separately.
[0124] Table 1 Comparison of prediction effects between the model in the prior art and the model in this embodiment
[0125]
[0126] In order to better implement the field crop gene data prediction method in the embodiment of the present invention, based on the field crop gene data prediction method, correspondingly, please refer to Figure 7 , Figure 7 This is a schematic diagram of the structure of an embodiment of a field crop gene data prediction device provided by the present invention. The field crop gene data prediction device 700 provided by the embodiment of the present invention includes:
[0127] A word segmentation module 710 is configured to obtain an original gene sequence and a preset word length, and to segment the original gene sequence into a plurality of gene words based on the preset word length to obtain a gene word sequence;
[0128] Gene encoding module 720, for establishing a two-dimensional encoding matrix based on the frequency and order of occurrence of each gene word in the gene word sequence;
[0129] A model building module 730 is used to encode environmental climate factors and establish an embedding layer, and to establish a prediction neural network model based on the preset word length and the embedding layer;
[0130] The gene prediction module 740 is used to input the two-dimensional coding matrix into the prediction neural network model to obtain a prediction result.
[0131] It should be noted here that the corresponding device 700 provided in the above embodiment can implement the technical solutions described in the above method embodiments. The specific implementation principles of the above modules or units can be found in the corresponding contents in the above method embodiments, which will not be repeated here.
[0132] See also Figure 8 , Figure 8 This is a schematic diagram of the structure of an electronic device provided by an embodiment of the present invention. Based on the aforementioned field crop genetic data prediction method, the present invention also provides a field crop genetic data prediction device 800, i.e., the aforementioned electronic device. Field crop genetic data prediction device 800 can be a computing device such as a mobile terminal, desktop computer, notebook, PDA, or server. Field crop genetic data prediction device 800 includes a processor 810, memory 820, and a display 830. Figure 8 Only some components of the field crop gene data prediction device are shown, but it should be understood that it is not required to implement all of the shown components, and more or fewer components may be implemented instead.
[0133] In some embodiments, the memory 820 can be an internal storage unit of the field crop genetic data prediction device 800, such as the hard drive or memory of the field crop genetic data prediction device 800. In other embodiments, the memory 820 can also be an external storage device of the field crop genetic data prediction device 800, such as a plug-in hard drive, a Smart Media Card (SMC), a Secure Digital (SD) card, a flash memory card, etc. equipped with the field crop genetic data prediction device 800. Furthermore, the memory 820 can include both the internal storage unit of the field crop genetic data prediction device 800 and an external storage device. The memory 820 is used to store application software installed in the field crop genetic data prediction device 800 and various data, such as the program code for installing the field crop genetic data prediction device 800. The memory 820 can also be used to temporarily store data that has been output or is about to be output. In one embodiment, the memory 820 stores a field crop gene data prediction program 840 , which can be executed by the processor 810 , thereby implementing the field crop gene data prediction method of each embodiment of the present application.
[0134] In some embodiments, the processor 810 may be a central processing unit (CPU), a microprocessor, or other data processing chip, configured to execute program codes or process data stored in the memory 820 , such as executing a field crop genetic data prediction method.
[0135] In some embodiments, display 830 can be an LED display, a liquid crystal display, a touch-sensitive liquid crystal display, or an OLED (Organic Light-Emitting Diode) touchscreen. Display 830 is used to display information on field crop genetic data prediction device 800 and to display a visual user interface. Components 810-830 of field crop genetic data prediction device 800 communicate with each other via a system bus.
[0136] In one embodiment, when the processor 810 executes the field crop gene data prediction program 840 in the memory 820 , the steps in the above field crop gene data prediction method are implemented.
[0137] This embodiment further provides a computer-readable storage medium on which a field crop gene data prediction program is stored. When the field crop gene data prediction program is executed by a processor, the steps in the above embodiment can be implemented.
[0138] The present invention provides a method, device, electronic device, and medium for predicting gene data for field crops. The method first extracts 2 bp of base information before and after the corresponding position of the reference genome based on the SNP site in the VCF file, and then assembles the base information into a sequence in order of position. The method then divides the original gene sequence into multiple gene words based on the preset word length according to the original gene sequence and a preset word length, obtaining a gene word sequence. A two-dimensional encoding matrix is then established based on the frequency and order relationship of each gene word in the gene word sequence. Environmental and climatic factors are then encoded and an embedding layer is established. A prediction neural network model is established based on the preset word length and the embedding layer. Finally, the two-dimensional encoding matrix is input into the prediction neural network model to obtain a prediction result. Compared to the prior art, the present invention adopts a two-dimensional matrix encoding method that combines the frequency and order relationship of gene words at the encoding level, so that the encoding retains the information on the mutual relationship between gene words. At the same time, the encoding can effectively reduce the input dimension of the encoded data, allowing subsequent model training to be performed more quickly, reducing the training and processing time of the prediction neural network model. In addition, the present invention also establishes an embedding layer based on environmental and climatic factors, and then establishes a predictive neural network model in combination with the embedding layer, so that the entire process only requires training one model to obtain the relationship between field crop genetic data and phenotypic traits under different environmental and climatic conditions. Combined with the encoding method in the present invention, the problem of relatively low efficiency of the existing method for predicting field crop genetic data is solved.
[0139] The above description is only a preferred specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by any technician familiar with this technical field within the technical scope disclosed by the present invention should be covered by the scope of protection of the present invention.
Claims
1. A method for predicting field crop gene data, characterized in that: include: Obtain the variant position and specific variant information of each sample from the VCF variant information file, then combine the reference genome information on both sides of this variant information. Each SNP site forms a 5nt variant unit. These 5nt variant units are connected in the order of the size of the variant sites in the VCF to form the original gene sequence; According to the original gene sequence and a preset word length, dividing the original gene sequence into a plurality of gene words based on the preset word length to obtain a gene word sequence; Establishing a two-dimensional coding matrix based on the frequency and order relationship of each gene word in the gene word sequence; Encoding environmental climate factors and establishing an embedding layer, and establishing a prediction neural network model based on the preset word length and the embedding layer; The two-dimensional coding matrix is input into the prediction neural network model to obtain a prediction result.
2. The method for predicting field crop gene data according to claim 1, wherein: The method of establishing a two-dimensional coding matrix based on the frequency and order of occurrence of each gene word in the gene word sequence includes: According to the preset word length, all types of gene words are obtained based on the types of gene bases; Establish a mapping relationship between each gene word and each row in the two-dimensional encoding matrix, and establish a mapping relationship between each gene word and each column in the two-dimensional encoding matrix; Based on the mapping relationship, the values of the matrix elements are obtained according to the gene word sequence, and the two-dimensional coding matrix M is established; Among them, the matrix element M in the two-dimensional coding matrix ij Used to represent: in the gene word sequence, the frequency of occurrence of the gene word corresponding to the jth column of the two-dimensional encoding matrix after the gene word corresponding to the i-th row of the two-dimensional encoding matrix appears for the first time, where i and j are both positive integers.
3. The method for predicting field crop gene data according to claim 2, wherein: The step of establishing a mapping relationship between each gene word and each row in the two-dimensional coding matrix and establishing a mapping relationship between each gene word and each column in the two-dimensional coding matrix includes: All types of gene words are sequentially coded, with each code number corresponding to a gene word, and the code numbers are all natural numbers; Establish a mapping relationship between the gene word with code number a and the a+1th row in the two-dimensional coding matrix; Establish a mapping relationship between the gene word with coding number a and the a+1th column in the two-dimensional coding matrix.
4. The method for predicting field crop gene data according to claim 1, wherein: The encoding of environmental climate factors and establishing an embedding layer, and establishing a prediction neural network model based on the preset word length in combination with the embedding layer, include: Encoding environmental climate factors to obtain a multidimensional vector, and establishing the embedding layer according to the multidimensional vector; According to the preset word length, a plurality of residual blocks are established and an arrangement order among the plurality of residual blocks is obtained; The prediction neural network model is established according to the embedding layer, the plurality of residual blocks and the arrangement order between the residual blocks.
5. The method for predicting field crop gene data according to claim 4, wherein: Encoding environmental climate factors to obtain a multidimensional vector, and establishing the embedding layer according to the multidimensional vector, including: Sequentially encode various environmental and climatic factors to obtain coded values; Performing embedding encoding on the encoded value to obtain a multi-dimensional vector; The embedding layer is established with the multi-dimensional vector as an output terminal.
6. The method for predicting field crop gene data according to claim 5, wherein: The prediction neural network model includes two first convolutional layers, a second convolutional layer, a maximum pooling layer, the embedding layer, a plurality of residual blocks, an average pooling layer and a fully connected layer; the input end of the first convolutional layer is used to input the two-dimensional coding matrix, and the output end of the first convolutional layer is connected to the input end of the second convolutional layer; the output end of the second convolutional layer is connected to the input end of the maximum pooling layer, and the output end of the second convolutional layer and the output end of the embedding layer are cross-connected to form a cross-product output end; The multiple residual blocks are connected in sequence according to the arrangement order between the multiple residual blocks, the input end of the residual block connection at the head end is connected to the cross product output end, and the output end of the residual block at the end end is connected to the input end of the average pooling layer; the output end of the average pooling layer is connected to the input end of the fully connected layer, and the output end of the fully connected layer is used to output the prediction result.
7. The method for predicting field crop gene data according to claim 6, wherein: The multiple residual blocks include a first residual block and a second residual block; the first residual block includes a third convolutional layer and a fourth convolutional layer, the input end of the third convolutional layer is the input end of the first residual block, the output end of the third convolutional layer is connected to the input end of the fourth convolutional layer, the input end of the third convolutional layer and the output end of the fourth convolutional layer are added together to form the output end of the first residual block; the second residual block includes a fifth convolutional layer, a sixth convolutional layer and a seventh convolutional layer, the input end of the fifth convolutional layer is the input end of the second residual block, the output end of the fifth convolutional layer is connected to the input end of the sixth convolutional layer, and the input end of the fifth convolutional layer is also connected to the input end of the seventh convolutional layer; the output end of the sixth convolutional layer and the output end of the seventh convolutional layer are added together to form the output end of the second residual block.
8. A field crop gene data prediction device, characterized in that: include: a word division module, configured to obtain an original gene sequence and a preset word length, and divide the original gene sequence into a plurality of gene words based on the preset word length to obtain a gene word sequence; A gene encoding module, configured to establish a two-dimensional encoding matrix based on the frequency and order of occurrence of each gene word in the gene word sequence; A model building module is used to encode environmental climate factors and establish an embedding layer, and to establish a prediction neural network model based on the preset word length and the embedding layer; The gene prediction module is used to input the two-dimensional coding matrix into the prediction neural network model to obtain a prediction result.
9. An electronic device, characterized in that: comprising a memory and a processor, wherein, The memory is used to store programs; The processor is coupled to the memory and is used to execute the program stored in the memory to implement the steps in the field crop gene data prediction method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that Used to store computer-readable programs or instructions, which, when executed by a processor, can implement the steps of the field crop genetic data prediction method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Aberrant splicing detection using convolutional neural networks (CNNS)
CN110870020A
Analysis method suitable for gene-environment interaction of complex characters and storage medium
CN114898809A