Crop whole genome prediction method and system based on machine learning

By acquiring gene expression data from multiple growth cycles for feature extraction and fusion prediction, whole-genome prediction results are generated, which solves the problem of the inability to accurately predict crop trait expression trends in existing technologies, realizes intelligent and precise crop breeding, and improves the adaptability and yield stability of crops in different environments.

CN120164521BActive Publication Date: 2025-09-09INSTITUTE OF CROP SCIENCE CHINESE ACADEMY OF AGRICULTURAL SCIENCES +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510643432.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-19
Publication Date
2025-09-09
Estimated Expiration
2045-05-19

AI Technical Summary

Technical Problem

Existing technologies lack the comprehensive utilization of crop whole genome information and are unable to accurately predict the trait expression trends of crops under different environmental conditions, making it difficult to formulate precise adaptive optimization strategies, affecting the efficiency and quality of crop breeding.

Method used

By acquiring gene expression data from multiple growth cycles, feature extraction and fusion prediction are performed to generate whole-genome prediction results. Based on the machine learning model, an adaptive optimization strategy is generated and fed back to the crop breeding system to dynamically adjust the breeding parameters.

Benefits of technology

It has achieved accurate prediction of the trait expression trends of target crops under different environmental conditions, improved the accuracy and comprehensiveness of the prediction, enhanced the efficiency and quality of crop cultivation, and enhanced the adaptability and yield stability of crops in different environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120164521B_ABST
    Figure CN120164521B_ABST
Patent Text Reader

Abstract

The present invention provides a crop whole-genome prediction method and system based on machine learning. First, a gene expression data set within multiple growth cycles of a target crop is obtained. The gene expression data set contains multiple gene sequence samples composed of genetic markers and phenotypic trait data. Next, feature extraction is performed on the gene expression data set to obtain gene association features and growth trait features. Then, a preset machine learning model is used to perform a fusion prediction of the gene association features and growth trait features to generate a fusion prediction feature. A whole-genome prediction result is determined based on the fusion prediction feature. The whole-genome prediction result can indicate the trait expression trend of the crop under different environments. Finally, an adaptive optimization strategy is generated based on the whole-genome prediction result and fed back to the crop cultivation system to trigger the adjustment of cultivation parameters, thereby realizing precise and intelligent crop cultivation and improving cultivation efficiency and crop adaptability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of machine learning technology, and in particular to a crop whole genome prediction method and system based on machine learning. Background Art

[0002] In the agricultural sector, crop breeding has always been a crucial step, directly impacting crop yield, quality, and adaptability to different environments. Traditional crop breeding methods rely primarily on experience and simple genetic theories. In the early days, breeders selected and bred crops based on observations of phenotypic characteristics, such as plant height and fruit size, relying on their own experience. This approach was highly subjective, inefficient, and difficult to accurately predict crop performance in different environments.

[0003] With the development of genetics, researchers have begun using gene marker technology to assist crop breeding. By detecting genetic markers at specific gene loci in crops, crops with specific genetic traits can be screened to a certain extent. However, this approach focuses on only a few known gene loci, ignoring the complex relationships between genes and the impact of environmental factors on gene expression. Furthermore, it primarily analyzes data from a single growth cycle and cannot fully reflect changes in gene expression throughout the crop's growth process.

[0004] In addition, some statistically based prediction models have also been applied to crop trait prediction. These models typically utilize large amounts of historical data and establish relationships between traits and environmental factors through statistical analysis. However, due to the limitations of statistical models, they struggle to handle the complex nonlinear relationships in genetic data, resulting in low accuracy and reliability in prediction results.

[0005] In actual crop breeding, existing technologies lack the comprehensive utilization of crop genome information, making it impossible to accurately predict trait expression trends under different environmental conditions. This makes it difficult to develop precise adaptive optimization strategies to guide crop breeding. Therefore, in the face of a changing environment and growing food demand, a more efficient and accurate crop genome prediction method is urgently needed to advance crop breeding technology. Summary of the Invention

[0006] In view of the above-mentioned problems, in combination with the first aspect of the present invention, an embodiment of the present invention provides a crop whole genome prediction method based on machine learning, the method comprising:

[0007] Acquire a gene expression data set of a target crop over multiple growth cycles, wherein the gene expression data set includes multiple gene sequence samples, each gene sequence sample consisting of a genetic marker of at least one gene locus and corresponding phenotypic trait data;

[0008] Performing feature extraction on the gene expression data set to obtain gene association features and growth trait features of each gene sequence sample;

[0009] Performing a fusion prediction on the gene association feature and the growth trait feature based on a preset machine learning model to generate a fusion prediction feature of the gene sequence sample;

[0010] Determining a whole-genome prediction result of the target crop based on the fused prediction features, wherein the whole-genome prediction result is used to indicate a trait expression trend of the target crop under different environmental conditions;

[0011] An adaptive optimization strategy is generated based on the whole genome prediction result, and the adaptive optimization strategy is fed back to the crop cultivation system to trigger a cultivation parameter adjustment operation.

[0012] On the other hand, an embodiment of the present invention also provides a crop whole genome prediction system based on machine learning, including a processor and a machine-readable storage medium, wherein the machine-readable storage medium is connected to the processor, the machine-readable storage medium is used to store programs, instructions or codes, and the processor is used to execute the programs, instructions or codes in the machine-readable storage medium to implement the above method.

[0013] Based on the above aspects, the embodiment of the present invention integrates the gene expression data set of multiple growth cycles to perform feature extraction, fusion prediction, determine the whole genome prediction results and generate an adaptive optimization strategy to feed back to the cultivation system, thereby achieving accurate prediction of the trait expression trend of the target crop under different environmental conditions and dynamic adjustment of the cultivation parameters. Compared with traditional methods, the comprehensive consideration of gene association characteristics and growth trait characteristics effectively improves the accuracy and comprehensiveness of the prediction. At the same time, based on the prediction results, an adaptive optimization strategy is generated and fed back to the cultivation system, which can respond to crop growth needs and environmental changes in real time, providing an intelligent and precise solution for crop cultivation, significantly improving the efficiency and quality of crop cultivation, and enhancing the adaptability and yield stability of crops in different environments. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] Figure 1 1 is a schematic diagram of the execution flow of the crop whole genome prediction method based on machine learning provided by an embodiment of the present invention.

[0015] Figure 2 Schematic diagram of exemplary hardware and software components of a crop whole genome prediction system based on machine learning provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0016] The present invention will be described in detail below with reference to the accompanying drawings. Figure 1This is a flow chart of a crop whole genome prediction method based on machine learning provided by an embodiment of the present invention. The crop whole genome prediction method based on machine learning is introduced in detail below.

[0017] Step S110: obtaining a gene expression data set of a target crop in multiple growth cycles, wherein the gene expression data set includes multiple gene sequence samples, each gene sequence sample consisting of a genetic marker of at least one gene locus and corresponding phenotypic trait data.

[0018] This embodiment aims to perform genome-wide predictions on target crops. First, it is necessary to obtain an accurate and comprehensive set of gene expression data. Specifically, the target crop undergoes multiple growth cycles during its growth process, and the crop's gene expression exhibits different characteristics during different growth cycles. Taking a hypothetical crop C as an example, its growth cycle may include the germination period, seedling period, growth period, and maturity period. Data collection on the crop's gene expression is required during each growth cycle.

[0019] A gene expression dataset consists of multiple gene sequence samples. Each gene sequence sample contains genetic markers for at least one locus and corresponding phenotypic trait data. Genetic markers, such as single nucleotide polymorphisms (SNPs) and insertion / deletions (InDels), are specific identifiers that reflect genetic characteristics and are represented by the symbol M. Phenotypic trait data refers to the external characteristics of crops, such as plant height, leaf length, and fruit size, and is represented by the symbol P. Assuming the gene expression dataset D is D, it can be represented as D = {S1, S2, …, Sn}, where Si (i = 1, 2, …, n) represents the i-th gene sequence sample. Each gene sequence sample Si can be further represented as Si = {Mi1, Mi2, …, Mij; Pi1, Pi2, …, Pik}, where Mij represents the genetic marker at the j-th gene locus in the i-th sample, and Pik represents the k-th phenotypic trait data in the i-th sample. For example, the first gene sequence sample S1 may contain genetic markers M11, M12, and M13 of three gene loci, as well as two phenotypic trait data P11 (plant height) and P12 (leaf length).

[0020] Step S120: performing feature extraction on the gene expression data set to obtain gene association features and growth trait features of each gene sequence sample.

[0021] After acquiring the gene expression data set, feature extraction is required to mine key information related to crop growth and traits. The extraction of gene association features and growth trait features will provide a basis for subsequent predictions.

[0022] Step S121: traverse each gene sequence sample in the gene expression data set and extract the linkage disequilibrium coefficient between the genetic markers of the gene loci.

[0023] To extract gene association features, we first need to traverse each gene sequence sample in the gene expression dataset. For each gene sequence sample Si, we need to extract the linkage disequilibrium coefficient between the genetic markers at its loci. Linkage disequilibrium refers to the degree of non-random association between alleles at different loci. Suppose there are m gene loci in the gene sequence sample Si, and their genetic markers are Mi1, Mi2, …, Mim.

[0024] For any two genetic markers Mia and Mib (a ≠ b) at a locus, the linkage disequilibrium coefficient between them can be calculated as follows. First, the frequency of the two genetic markers appearing together in all samples is counted, denoted as f(Mia, Mib). Then, the frequency of the genetic markers Mia and Mib appearing separately is counted, denoted as f(Mia) and f(Mib). The linkage disequilibrium coefficient D can be expressed as D = f(Mia, Mib) - f(Mia) * f(Mib).

[0025] In order to eliminate the influence of frequency differences, it is necessary to further calculate the standardized linkage disequilibrium coefficient D'. The calculation of D' needs to take into account the positive and negative conditions of D. When D>0, Dmax=min(f(Mia)f(Mib), (1-f(Mia))(1-f(Mib))); when D<0, Dmax=min(f(Mia)(1-f(Mib)), (1-f(Mia))f(Mib)). The standardized linkage disequilibrium coefficient D'=D / Dmax. By performing such calculations for all gene loci pairs, the linkage disequilibrium coefficient matrix Ci between the genetic markers of the gene loci in the gene sequence sample Si can be obtained. Ci is an m×m matrix, in which the elements in the ath row and bth column represent the linkage disequilibrium coefficients between the genetic markers Mia and Mib.

[0026] Step S122: calling a pre-trained gene feature encoder to perform multi-level encoding processing on each linkage disequilibrium coefficient to generate a gene interaction feature vector of the gene sequence sample.

[0027] In this embodiment, unsupervised pre-training (e.g., an autoencoder) or transfer learning (e.g., pre-training on similar crop data) can be used to initialize the genetic signature encoder. If using an autoencoder for unsupervised pre-training, a large amount of genetic data is first prepared as pre-training data. The autoencoder consists of two parts: an encoder and a decoder. The encoder compresses the input genetic data into a low-dimensional representation, and the decoder reconstructs the low-dimensional representation into the original genetic data. The autoencoder is trained by minimizing the reconstruction error. After training, the encoder is used as the genetic signature encoder.

[0028] If transfer learning is used, pre-training is performed using genetic data from similar crops to generate a pre-trained genetic feature encoder model. The parameters of this model are then transferred to the current genetic feature encoder and fine-tuned on the current genetic data. After pre-training, the genetic feature encoder can better perform multi-level encoding processing on the input linkage disequilibrium coefficients to generate gene interaction feature vectors for the genetic sequence samples. The specific multi-level encoding process steps are described above in steps S1221-S1224.

[0029] Step S1221: inputting each linkage disequilibrium coefficient into the first encoding layer of the gene feature encoder to generate initial interaction weights between gene loci.

[0030] The first encoding layer of the genetic signature encoder receives the linkage disequilibrium matrix Ci as input. This first encoding layer can be viewed as a linear transformation layer, which linearly combines each element in the linkage disequilibrium matrix Ci to generate initial interaction weights between loci. Let the weight matrix of the first encoding layer be W1, with dimensions m×m (the same as Ci), and the bias vector be b1, with dimensions m. For each row vector Ci_row (of dimension m) in the matrix Ci, a linear transformation is performed to obtain the corresponding initial interaction weight vector Wi1_row: Wi1_row = W1*Ci_row+b1. The multiplication here is matrix multiplication, and the addition is vector addition. By performing this calculation on each row of the matrix Ci, a new matrix Ii1 is obtained, representing the initial interaction weight matrix between loci. Ii1 has dimensions m×m.

[0031] Step S1222: In the second encoding layer of the gene feature encoder, local aggregation processing is performed on the genetic markers of adjacent gene sites based on the initial interaction weights to obtain local aggregation features of the gene site group.

[0032] The main function of the second encoding layer is to perform local aggregation of genetic markers at adjacent loci. Assuming that the loci are arranged in a predetermined order, for each locus, several adjacent loci are considered. Let the range of adjacent loci be r. That is, for locus j, the r loci before and after it are considered. For a gene sequence sample Si, based on the initial interaction weight matrix Ii1, a weighted sum of the genetic markers at adjacent loci is performed. Let the set of adjacent loci for locus j be Nj = {jr, …, j-1, j, j+1, …, j+r}. For each locus k in Nj, the corresponding initial interaction weight is Ii1jk. The genetic markers Mik (k∈Nj) of loci in Nj are weighted and summed according to Ii1jk to obtain the local aggregated feature value Li2j for locus j. The specific calculation method is Li2j = Σ(Ii1jk*Mik) (k∈Nj). By performing such calculations on all gene loci, we can obtain the local aggregate feature vector Li2 of the gene sequence sample Si, whose dimension is m.

[0033] Step S1223: In the third encoding layer of the gene feature encoder, after the initial interaction weight is adjusted to the same dimension as the local aggregate feature through the fully connected layer, the initial interaction weight and the local aggregate feature are batch normalized respectively, and the normalized local aggregate feature and the initial interaction weight are channel-attention weighted splicing processed according to the preset cross-layer connection rules to generate cross-layer fusion features.

[0034] The third encoding layer first resizes the initial interaction weight matrix Ii1 through a fully connected layer to the same dimension as the local aggregate feature vector Li2. Let the fully connected layer weight matrix be W2, with dimensions m×m, and the bias vector be b2, with dimensions m. Each row vector Ii1_row (of dimension m) of Ii1 is converted to a vector Ii3_row of the same dimension as Li2 through a linear transformation: Ii3_row = W2 * Ii1_row + b2. This yields the resized initial interaction weight matrix Ii3.

[0035] Next, batch normalization is performed on the adjusted initial interaction weight matrix Ii3 and the local aggregate feature vector Li2. The goal of batch normalization is to ensure that the input data has the same mean and variance within each batch, thereby accelerating model training. Let the batch normalization parameters be γ and β. For the vectors Ii3_row and Li2, after batch normalization, the vectors Ii3_norm and Li2_norm are obtained respectively.

[0036] Then, according to the preset cross-layer connection rules, the normalized local aggregate feature vector Li2_norm is concatenated with the initial interaction weight matrix Ii3_norm using channel-wise attention weighting. The channel-wise attention mechanism adaptively adjusts the weights of different channels (i.e., different gene loci). Let the channel-wise attention weight vector be α, with dimension m. The cross-layer fused feature vector Fi3 is obtained by weighted concatenation of Li2_norm and Ii3_norm according to α. Specifically, the corresponding elements of Li2_norm and Ii3_norm are weightedly combined according to the weights in α, i.e., Fi3j = αj * Li2_normj + (1-αj) * Ii3_normj (j = 1, 2, …, m).

[0037] Step S1224: performing global average pooling processing on the cross-layer fusion features to obtain a gene interaction feature vector of the gene sequence sample.

[0038] After obtaining the cross-layer fusion feature vector Fi3, global max pooling is performed on it. Global max pooling selects the maximum value from all elements in vector Fi3 to obtain an m-dimensional vector. For vector Fi3, whose elements are Fi31, Fi32, ​​…, Fi3m, global max pooling will produce a gene interaction feature vector Gi, where the elements of Gi are the values ​​after global max pooling. This results in an m-dimensional gene interaction feature vector Gi.

[0039] Step S123: performing environmental adaptability analysis on the phenotypic trait data to obtain trait stability scores of the gene sequence samples under different growth environments.

[0040] In addition to gene association features, phenotypic trait data must be analyzed to derive trait stability scores for gene sequence samples under different growth environments. Assume that the phenotypic trait data for gene sequence sample Si is Pi = {Pi1, Pi2, …, Pik}, and different growth environments can be represented by E = {E1, E2, …, El}.

[0041] Each phenotypic trait data point, Pik, will have different values ​​under different growth environments, Ej (j = 1, 2, …, 1). First, calculate the mean μik and standard deviation σik for each phenotypic trait data point under different growth environments. The trait stability score can be defined as a score based on the coefficient of variation between environments: stability = μik / σik. This stability score is then normalized using the Z-score to obtain the trait stability score for the gene sequence sample under different growth environments.

[0042] Step S124: Perform feature alignment processing on the gene interaction feature vector and the trait stability score to obtain the gene association feature of the gene sequence sample.

[0043] After obtaining the gene interaction feature vector Gi and the trait stability score vector Si_stable, it is necessary to perform feature alignment processing on them to obtain the gene association feature vector Ai of the gene sequence sample Si.

[0044] Step S125: Perform trait association mapping processing on the genetic markers of the gene sequence sample to obtain the growth trait feature of the gene sequence sample.

[0045] Step S1251: Obtain the phenotypic trait data of the gene sequence sample in different growth periods, and extract the dynamic change gradient of the phenotypic trait data.

[0046] To obtain the growth trait feature, first, it is necessary to obtain the phenotypic trait data of the gene sequence sample in different growth periods. Let the phenotypic trait data of the gene sequence sample Si in different growth periods T = {T1, T2,..., Tq} be Pi_T = {Pi_T1, Pi_T2,..., Pi_Tq}, where Pi_Tt (t = 1, 2,..., q) represents the phenotypic trait data vector under the growth period Tt.

[0047] For each phenotypic trait data element, calculate its dynamic change gradient between adjacent growth periods. Taking the phenotypic trait data element Pi_Ttk (the k-th phenotypic trait data under the t-th growth period) as an example, its dynamic change gradient ΔPi_Ttk can be calculated by Pi_T(t + 1)k - Pi_Ttk (assuming t < q). By performing such calculations on all phenotypic trait data elements, the dynamic change gradient matrix ΔPi_T of the phenotypic trait data can be obtained.

[0048] Step S1252: Construct a trait association mapping matrix according to the Spearman rank correlation coefficient between the dynamic change gradient and the linkage disequilibrium coefficient of the genetic marker.

[0049] Next, a trait association mapping matrix is ​​constructed based on the Spearman rank correlation coefficient between the dynamic gradients of phenotypic trait data and the linkage disequilibrium coefficients of genetic markers. The dynamic gradient matrix ΔPi_T is a k×q matrix (k traits, q periods), and the linkage disequilibrium coefficient matrix Ci is an m×m matrix. Therefore, a time-slicing aggregation method can be used to segment each row (m-dimensional) of Ci according to growth period q. For the j-th row vector Ci_j of Ci, it is divided according to growth period q, with each segment corresponding to one growth period. Statistics, such as the mean, are calculated for the elements within each segment to transform Ci_j into a q-dimensional vector Ci_j_transformed.

[0050] The Spearman rank correlation coefficient is a statistic that measures the monotonic relationship between two variables. For each trait k's dynamic gradient vector ΔPi_Tk (of dimension q) and each locus j's q-dimensional vector Ci_j_transformed, the Spearman rank correlation coefficient ρkj is calculated. By performing this calculation for all trait and locus pairs, a k×m matrix R is obtained, where the element in the kth row and jth column represents the Spearman rank correlation coefficient between the dynamic gradient of the kth trait and the linkage disequilibrium value of the jth locus.

[0051] Step S1253: performing singular value decomposition processing on the trait association mapping matrix to obtain the principal component characteristics of the phenotypic trait data.

[0052] To extract the principal component characteristics of the phenotypic trait data, the trait association mapping matrix R is subjected to singular value decomposition. Singular value decomposition decomposes a matrix into the product of three matrices: R = U * Σ * Vᵀ, where U is the left singular matrix, Σ is a diagonal matrix with singular values ​​on the diagonal, and Vᵀ is the transpose of the right singular matrix. By selecting the first p singular values ​​with the largest singular values ​​and their corresponding left and right singular vectors, the principal component characteristic matrix PCi of the phenotypic trait data is obtained. PCi has a dimension of p × m, where p is the number of principal components selected and m is the number of loci.

[0053] Step S1254: Standardize the principal component features and normalize the gene interaction feature vectors, align the feature dimensions of the standardized principal component features and the normalized gene interaction feature vectors, perform nonlinear spatial projection processing through a multi-layer perceptron, and obtain the growth trait characteristics of the gene sequence sample.

[0054] First, normalize the principal component feature matrix PCi. The goal of normalization is to ensure that each column of the principal component feature matrix has zero mean and unit variance. Let the jth column vector of PCi be PCi_j, with mean μj and standard deviation σj. Then, the normalized vector PCi_j_std can be expressed as PCi_j_std = (PCi_j - μj) / σj. By performing this normalization on all column vectors, we obtain the normalized principal component feature matrix PCi_std (of dimension p × m).

[0055] At the same time, the gene interaction feature vector Gi is normalized. Normalization maps the elements of the vector to the interval [0, 1]. Let the elements of Gi be Gi1, Gi2, …, Gim, with their minimum value min(Gi) and maximum value max(Gi). The normalized vector Gi_norm can be expressed as Gi_norm = (Gi - min(Gi)) / (max(Gi) - min(Gi)) (with dimension m).

[0056] Because the dimensions of the standardized principal component feature matrix PCi_std and the normalized gene interaction feature vector Gi_norm do not match, PCi_std undergoes principal component dimensionality reduction. A fully connected layer is used to compress PCi_std to m dimensions. Let the weight matrix of the fully connected layer be W_pca, with dimensions p × m, and the bias vector be b_pca, with dimensions m. After a linear transformation, PCi_std is compressed into the m-dimensional matrix PCi_std_compressed, where PCi_std_compressed = W_pca * PCi_std + b_pca.

[0057] Next, concatenate PCi_std_compressed with the normalized gene interaction feature vector Gi_norm to produce a 2m-dimensional vector, Si_growth_concat. Si_growth_concat is then projected to the appropriate dimension through a fully connected layer. Let the weight matrix of this fully connected layer be W_final, with dimensions of 2m × r (r is the dimension of the final growth trait feature vector), and the bias vector be b_final, with dimensions of r. After a linear transformation, the growth trait feature vector Si_growth for the gene sequence sample is obtained: Si_growth = W_final * Si_growth_concat + b_final.

[0058] Step S130: performing fusion prediction on the gene association feature and the growth trait feature based on a preset machine learning model to generate a fusion prediction feature of the gene sequence sample.

[0059] Step S131: performing standardization processing on the gene association features to obtain standardized gene association features.

[0060] Before performing fusion prediction, the gene association feature vector Ai needs to be normalized. The purpose of normalization is to ensure that each element of the gene association feature vector has zero mean and unit variance. Assume that the elements of the gene association feature vector Ai are Ai1, Ai2, …, Aiik, with a mean of μA and a standard deviation of σA. The normalized gene association feature vector Ai_std can be expressed as Ai_std = (Ai - μA) / σA, where the subtraction and division are performed element by element.

[0061] Step S132: normalizing the growth trait characteristics to obtain normalized growth trait characteristics.

[0062] At the same time, the growth trait characteristic vector Si_growth is normalized. Normalization can map the elements of the vector to the interval [0, 1]. Assume that the elements of the growth trait characteristic vector Si_growth are Si_growth1, Si_growth2, ..., Si_growthr, with its minimum value being min(Si_growth) and its maximum value being max(Si_growth). Then the normalized growth trait characteristic vector Si_growth_norm can be expressed as Si_growth_norm=(Si_growth-min(Si_growth)) / (max(Si_growth)-min(Si_growth)). Similarly, subtraction and division are performed element by element.

[0063] Step S133: calling the dynamic weight allocation module in the machine learning model to generate feature fusion weights according to the correlation scores between the standardized gene association features and the normalized growth trait features.

[0064] Step S1331: inputting the standardized gene association features into the first fully connected layer of the dynamic weight allocation module to generate a gene association weight vector.

[0065] The normalized gene association feature vector Ai_std is input into the first fully connected layer of the dynamic weight assignment module. The first fully connected layer can be considered a linear transformation layer, which linearly combines the input vectors to generate the gene association weight vector. Let the weight matrix of the first fully connected layer be W3, with dimensions k × s (k is the dimension of the normalized gene association feature vector, s is the dimension of the gene association weight vector), and the bias vector be b3, with dimensions s. Through linear transformation, the gene association weight vector Wi_gene can be expressed as Wi_gene = W3 * Ai_std + b3, where the multiplication is matrix multiplication and the addition is vector addition.

[0066] Step S1332: inputting the normalized growth trait features into the second fully connected layer of the dynamic weight allocation module to generate a trait association weight vector.

[0067] The normalized growth trait feature vector Si_growth_norm is input into the second fully connected layer of the dynamic weight allocation module. This second fully connected layer is also a linear transformation layer. Its weight matrix is ​​W4, with dimensions r × s (r is the dimension of the normalized growth trait feature vector, s is the dimension of the trait-associated weight vector), and its bias vector is b4, with dimensions s. After the linear transformation, the trait-associated weight vector Wi_trait can be expressed as Wi_trait = W4 * Si_growth_norm + b4.

[0068] Step S1333: Calculate the cosine similarity between the gene association weight vector and the trait association weight vector, and generate an initial weight score through a nonlinear activation function.

[0069] To avoid dimensionality mismatch when generating feature fusion weights, first ensure that the dimensions of the gene-association weight vector Wi_gene and the trait-association weight vector Wi_trait are consistent. When inputting the standardized gene-association feature vector Ai_std into the first fully connected layer of the dynamic weight assignment module and the normalized growth trait feature vector Si_growth_norm into the second fully connected layer of the dynamic weight assignment module, set the output dimensions of these two fully connected layers to be the same, both t-dimensional.

[0070] Compute the multi-channel cosine similarity between the gene association weight vector Wi_gene and the trait association weight vector Wi_trait. For each channel (element) in Wi_gene and Wi_trait, calculate the cosine similarity by dividing the dot product by the product of their moduli. Let the elements of Wi_gene be Wi_gene1, Wi_gene2, …, Wi_genet, and the elements of Wi_trait be Wi_trait1, Wi_trait2, …, Wi_traitt. For the i-th channel, their dot product is dot_product_i = Wi_genei * Wi_traiti. The modulus of the i-th channel element of Wi_gene is norm_genei = the absolute value of Wi_genei, and the modulus of the i-th channel element of Wi_trait is norm_traiti = the absolute value of Wi_traiti. Then, the cosine similarity of the i-th channel is cos_sim_i = dot_product_i / (norm_genei * norm_traiti).

[0071] The cosine similarity of each channel is input into a nonlinear activation function, such as the Sigmoid function. The Sigmoid function maps the input value to the range (0, 1). Let the Sigmoid function be S(x) = 1 / (1 + exp(-x)), and substitute the cosine similarity cos_sim_i of each channel into the function to obtain the initial weight score vector score_init, whose elements are score_initi = S(cos_sim_i). This obtains a t-dimensional weight vector that matches the element-by-element weighting of the subsequent t-dimensional unified features.

[0072] Step S1334: performing nonlinear activation processing on the initial weight score to generate a feature fusion weight between the standardized gene association feature and the normalized growth trait feature.

[0073] The initial weight score score_init is reactivated nonlinearly. Another nonlinear activation function can be used, such as the ReLU (rectified linear unit) function, where ReLU(x) = max(0, x). Substitute the initial weight score score_init into the ReLU function to obtain the feature fusion weight vector weight_fusion. If score_init is greater than 0, the corresponding element in weight_fusion is equal to score_init; if score_init is less than or equal to 0, the corresponding element in weight_fusion is 0.

[0074] Step S134: After unifying the standardized gene association features and the normalized growth trait features to the same dimension through the third fully connected layer of the machine learning model, weighted splicing processing is performed based on the feature fusion weight to obtain the fusion prediction features of the gene sequence sample.

[0075] The standardized gene-association feature vector Ai_std and the normalized growth trait feature vector Si_growth_norm are fed into the third fully connected layer of the machine learning model. The third fully connected layer unifies these two vectors to the same dimension. For the standardized gene-association feature vector, the weight matrix is ​​W5, with dimensions k × t (k is the dimension of the standardized gene-association feature vector, t is the dimension after unification), and the bias vector is b5, with dimensions t. For the normalized growth trait feature vector, the weight matrix is ​​W6, with dimensions r × t (r is the dimension of the normalized growth trait feature vector, t is the dimension after unification), and the bias vector is b6, with dimensions t.

[0076] After linear transformation, the transformed vector of the standardized gene association feature vector is Ai_std_transformed=W5*Ai_std+b5, and the transformed vector of the normalized growth trait feature vector is Si_growth_norm_transformed=W6*Si_growth_norm+b6.

[0077] Based on the feature fusion weight vector weight_fusion, the two transformed vectors are weightedly concatenated. Let the elements of weight_fusion be weight_fusion1, weight_fusion2, …, weight_fusions. Ai_std_transformed and Si_growth_norm_transformed are weightedly combined according to weight_fusion. For each element position i (i ranges from 1 to t), the i-th element of the fused prediction feature vector Fi_fusion can be expressed as Fi_fusioni = weight_fusioni * Ai_std_transformedi + (1-weight_fusioni) * Si_growth_norm_transformedi. In this way, the fused prediction feature vector Fi_fusion for the gene sequence sample is obtained.

[0078] Step S140: determining a whole-genome prediction result of the target crop based on the fused prediction features, wherein the whole-genome prediction result is used to indicate a trait expression trend of the target crop under different environmental conditions.

[0079] Step S141: calling a pre-trained genome prediction model to perform cross-environment generalization processing on the fused prediction features to obtain the trait expression probability distribution of the target crop under preset environmental variables.

[0080] Step S1411: Input the fused prediction feature into the cross-attention module, calculate the interaction weight between the fused prediction feature and the preset environmental variable, and dynamically weighted fuse the preset environmental variable according to the interaction weight to generate an environmental extension feature.

[0081] The fused prediction feature vector Fi_fusion is input into the cross-attention module. The preset environment variables can be represented by the vector E={E1, E2, …, Eu}. The cross-attention module first calculates the interaction weight between the fused prediction feature vector Fi_fusion and the preset environment variable vector E. Specifically, for each element Fi_fusionj in Fi_fusion (j ranges from 1 to t) and each element Ek in E (k ranges from 1 to u), the correlation score between them is calculated. This can be calculated through a linear transformation. Let the correlation score matrix be S, and its element Sjk in the jth row and kth column can be expressed as Sjk=W7*Fi_fusionj*Ek+b7, where W7 is a scalar weight and b7 is a scalar bias.

[0082] Then the correlation score matrix S is normalized, for example, using the Softmax function, Softmax(xi)=exp(xi) / Σ(exp(xj)) (j ranges from 1 to u), to obtain the interaction weight matrix W_interaction, where the element in the j-th row and k-th column represents the interaction weight between the j-th element of Fi_fusion and the k-th element of the preset environment variable.

[0083] Dynamically weight the preset environmental variable vector E according to the interaction weight matrix W_interaction. Let the elements of the extended environmental feature vector Ei_extended be Ei_extendedl (l ranges from 1 to t), then Ei_extendedl = Σ(W_interactionlj * Ek) (k ranges from 1 to u). This calculation yields the extended environmental feature vector Ei_extended.

[0084] Step S1412: Inputting the environmental extension features into the multi-task learning layer of the genome prediction model to generate multi-task prediction results for the target crop under different environmental combinations.

[0085] The extended environmental feature vector Ei_extended is input into the multi-task learning layer of the genomic prediction model. The multi-task learning layer can simultaneously process multiple related prediction tasks, each corresponding to a different environmental combination. Assume that the multi-task learning layer has v tasks, each corresponding to a different environmental combination. The multi-task learning layer consists of multiple hidden layers and an output layer. Its input is the extended environmental feature vector Ei_extended. After a series of nonlinear and linear transformations, it outputs a multi-task prediction result vector Pi_multi for the target crop under different environmental combinations. Pi_multi has a dimension of v.

[0086] Step S1413: performing probability normalization processing on the multi-task prediction results to obtain the trait expression probability distribution of the target crop under preset environmental variables.

[0087] The multi-task prediction result vector Pi_multi is probabilistically normalized to ensure that its elements sum to 1 and that each element lies within the interval [0, 1]. A softmax function can be used to process Pi_multi. Let the processed vector be Pi_prob, whose elements Pi_probi = exp(Pi_multi_i) / Σ(exp(Pi_multi_j)) (j ranges from 1 to v). Here, Pi_prob is the probability distribution vector of the target crop's trait expression under the preset environmental variables.

[0088] Step S142: extracting gene locus identifiers whose variance in the trait expression probability distribution exceeds a preset threshold, and generating a set of candidate key loci.

[0089] For the trait expression probability distribution vector Pi_prob, calculate the variance of each element of the corresponding locus. Suppose there are m loci, each with a sequence of probability values ​​under different environmental combinations. Calculate the variance of these probability value sequences. Let the probability value sequence for the nth locus be Pi_prob_n = {Pi_prob_n1, Pi_prob_n2, …, Pi_prob_nv}, and its variance be variance_n = Σ((Pi_prob_ni - mean(Pi_prob_n))²) / v (where i ranges from 1 to v), where mean(Pi_prob_n) is the mean of the probability value sequence.

[0090] The identifiers of the loci whose variance exceeds the preset threshold are extracted to form the candidate key loci set C_candidate. If variance_n>threshold, the identifier of locus n is added to C_candidate.

[0091] Step S143: Calculate the characteristic activation intensity of each locus in the candidate key locus set in the environmental extension feature, and select loci with an activation intensity mean higher than the genetic significance level to form a key locus set.

[0092] For each locus in the candidate key locus set C_candidate, calculate its feature activation strength in the extended environment feature vector Ei_extended. Let the locus identifiers in the candidate key locus set C_candidate be c1, c2, …, cw, and for each locus ci, its corresponding element in the extended environment feature vector Ei_extended is Ei_extended_ci.

[0093] Calculate the mean of the feature activation strength of each locus. Let the mean of the feature activation strength of locus ci be mean_activation_ci=Σ(Ei_extended_ci) / t (where t is the dimension of the extended feature vector of the environment).

[0094] Select loci whose mean activation intensity is higher than the genetic significance level significance_level to form the key loci set C_key. If mean_activation_ci>significance_level, locus ci is added to C_key.

[0095] Step S144: performing feature association binding processing on the trait expression probability distribution and the key gene site set to generate the genome-wide prediction result.

[0096] The trait expression probability distribution vector Pi_prob is associated with the key loci set C_key through feature binding. Specifically, each locus in the key locus set C_key is associated with the corresponding element in the trait expression probability distribution vector Pi_prob. This can be accomplished by establishing a mapping relationship, such as a dictionary, where the key is the locus identifier and the value is the probability value for the corresponding position in the trait expression probability distribution vector. This binding process yields genome-wide prediction results that include information on trait expression trends in the target crop under different environmental conditions.

[0097] Step S150: generating an adaptive optimization strategy based on the whole genome prediction result, and feeding the adaptive optimization strategy back to the crop cultivation system to trigger a cultivation parameter adjustment operation.

[0098] Step S151: Determine the contribution score of the key gene locus set to the expression of the target trait based on the whole genome prediction result.

[0099] For example, step S1511: extracting the site identifiers of the trait expression probability distribution and the key gene site set bound in the whole genome prediction result.

[0100] The bound trait expression probability distribution vector Pi_prob and the locus identifiers of the key gene locus set C_key are extracted from the whole genome prediction results. These locus identifiers can uniquely identify each key gene locus, while the trait expression probability distribution vector contains the trait expression probability information of the target crop under different environmental conditions.

[0101] Step S1512: extracting the probability value vector of the corresponding gene locus from the trait expression probability distribution according to the locus identifier, and performing maximum and minimum value normalization processing on the probability value vector to generate a standardized probability vector in the interval [0, 1].

[0102] Based on the locus identifiers of the key locus set C_key, extract the probability value vector of the corresponding locus from the trait expression probability distribution vector Pi_prob. Suppose there are w loci in the key locus set C_key, whose locus identifiers are c1, c2, ..., cw, and the corresponding probability value vector is Pi_prob_key={Pi_prob_c1, Pi_prob_c2, ..., Pi_prob_cw}.

[0103] Normalize the probability value vector Pi_prob_key to its maximum and minimum values. Assume that the minimum value in Pi_prob_key is min(Pi_prob_key) and the maximum value is max(Pi_prob_key). Then the element Pi_prob_key_normi of the standardized probability vector Pi_prob_key_norm is Pi_prob_key_normi=(Pi_prob_keyi-min(Pi_prob_key)) / (max(Pi_prob_key)-min(Pi_prob_key)) (i ranges from 1 to w). This will map the probability value vector to the interval [0, 1].

[0104] Step S1513: Obtain the genetic marker data of the key gene locus set, and calculate the genetic association strength vector of each gene locus based on the linkage disequilibrium coefficient of the genetic marker data.

[0105] Obtain the genetic marker data of the key gene locus set C_key. For each key gene locus, its genetic marker data can be expressed as a series of eigenvalues. Based on the linkage disequilibrium coefficient of these genetic marker data, calculate the genetic association strength of each gene locus. Let the genetic marker data of the gene locus ci in the key gene locus set C_key be Mi_key, and its linkage disequilibrium coefficient matrix be Ci_key. By analyzing Ci_key, for example, calculating the value of its determinant or certain eigenvalues, the genetic association strength value of the gene locus ci can be obtained. The genetic association strength values ​​of all key gene loci are combined into a genetic association strength vector Gi_association={Gi_association1, Gi_association2,…, Gi_associationw}.

[0106] Step S1514: performing Z-score normalization processing on the genetic association strength vector to generate a normalized association vector with zero mean and unit variance.

[0107] Perform Z-score normalization on the genetic association strength vector Gi_association. The purpose of Z-score normalization is to ensure that the elements of the vector have zero mean and unit variance. Assume that the mean of the genetic association strength vector Gi_association is μG and the standard deviation is σG. Then, the elements of the standardized association vector Gi_association_std are Gi_association_stdi = (Gi_associationi - μG) / σG (where i ranges from 1 to w).

[0108] Step S1515: performing element-wise multiplication of the standardized probability vector and the standardized association vector to generate a comprehensive weight vector corresponding one-to-one to each key locus in the key locus set.

[0109] Multiply the standardized probability vector Pi_prob_key_norm by the standardized association vector Gi_association_std element by element. Let Wi_composite be the element of the comprehensive weight vector Wi_composite, then Wi_compositei = Pi_prob_key_norm i * Gi_association_std (where i ranges from 1 to w). This yields a comprehensive weight vector corresponding to each key locus in the key locus set.

[0110] Step S1516: performing linear mapping processing on the comprehensive weight vector according to the dimensional range of the historical phenotypic data of the target trait, and generating a contribution score sequence consistent with the expression dimension of the target trait.

[0111] Since the dimensions of different traits are different, in order to avoid the problem of inconsistent dimensions when calculating the contribution score, the method of independent mapping of traits is adopted.

[0112] For each trait in the target trait set, determine the dimensional range of its historical phenotypic data. Suppose there are k target traits, and the dimensional range of the historical phenotypic data for the i-th trait is [min_phenotype_i, max_phenotype_i].

[0113] Process the composite weight vector Wi_composite (dimension w) and partition it by trait, assuming that each trait corresponds to a certain number of loci. For the composite weight subvector Wi_composite_i corresponding to the i-th trait, first perform Z-score normalization on it. Assume that the mean of Wi_composite_i is μ_i and the standard deviation is σ_i. Then, the element Wi_composite_i_stdj of the standardized subvector Wi_composite_i_std is Wi_composite_i_stdj = (Wi_composite_ij - μ_i) / σ_i (where j represents the element index in the subvector).

[0114] Then, the standardized sub-vector Wi_composite_i_std is linearly mapped to obtain the contribution score sub-sequence Si_contribution_i of the i-th trait, whose element Si_contribution_ij=min_phenotype_i+(max_phenotype_i-min_phenotype_i)*(Wi_composite_i_stdj-min(Wi_composite_i_std)) / (max(Wi_composite_i_std)-min(Wi_composite_i_std)).

[0115] The contribution score subsequences of all traits are combined to obtain the contribution score sequence Si_contribution that is consistent with the expression dimension of the target trait.

[0116] Step S152: generating a site optimization priority sequence according to the difference between the contribution score and a preset genetic gain threshold.

[0117] Calculate the difference between each element in the contribution score sequence Si_contribution and the preset genetic gain threshold threshold_gain. Let the difference sequence be Si_diff, where element Si_diffi = Si_contributioni - threshold_gain (i ranges from 1 to w).

[0118] Sort the key loci based on the difference sequence Si_diff to generate a locus optimization priority sequence. Sorting can be done from largest to smallest difference, with loci with larger differences receiving higher priorities. The sorted sequence of key loci identifiers is the locus optimization priority sequence.

[0119] Step S153: selecting the first N key gene loci as optimization seed loci according to the locus optimization priority sequence, where N is a preset positive integer.

[0120] Select the top N key gene loci from the site optimization priority sequence as optimization seed loci. Suppose the site optimization priority sequence is C_priority = {c_priority1, c_priority2, ..., c_priorityw}, then the optimization seed loci set C_seed = {c_priority1, c_priority2, ..., c_priorityN}.

[0121] Step S154: performing crossover recombination simulation processing on the genetic markers of the optimized seed site to generate a candidate gene combination set.

[0122] Perform a crossover and recombination simulation on the genetic markers at each locus in the optimized seed locus set C_seed. Assuming each locus has different alleles, simulate the crossover and recombination process to generate different gene combinations. For example, for the locus c_seedi, its alleles are a1, a2, …, am. By combining these alleles, different gene combinations can be obtained. Perform this combination operation on all loci in the optimized seed locus set C_seed to generate the candidate gene combination set C_candidate_combination.

[0123] Step S155: Calling the genome prediction model that has undergone gene recombination data enhancement training and adversarial regularization processing to perform virtual phenotype prediction processing on the candidate gene combination set, and introducing a Monte Carlo dropout layer in the prediction process to simulate data distribution deviation to obtain the predicted trait expression level of each candidate gene combination.

[0124] A genomic prediction model, trained with gene recombination data augmentation and adversarial regularization, was used to predict virtual phenotypes for the candidate gene combination set C_candidate_combination. A Monte Carlo dropout layer was introduced during the prediction process. To verify the impact of the Monte Carlo dropout layer on the stability of the prediction results, a multiple sampling ensemble method was employed.

[0125] First, we experimentally determined the optimal deactivation ratio. We tested the model at different deactivation ratios (e.g., 10%, 20%, and 30%), observing the stability and accuracy of the prediction results. We then selected the deactivation ratio that yielded the best prediction results. We assumed that the optimal deactivation ratio was α.

[0126] For each candidate gene combination, input it into a genomic prediction model with a Monte Carlo dropout layer for multiple (e.g., N) predictions. During each prediction, the Monte Carlo dropout layer randomly sets elements of the input features to 0 with a probability of α. The results of these N predictions are averaged to obtain the final predicted trait expression level for that candidate gene combination. Assume that there are q candidate gene combinations in the candidate gene combination set C_candidate_combination, and their predicted trait expression level vector is Pi_prediction = {Pi_prediction1, Pi_prediction2, …, Pi_predictionq}.

[0127] Step S156: Screening out an optimized gene combination scheme from the candidate gene combination set according to the matching degree between the predicted trait expression level and the target optimized trait.

[0128] Calculate the degree of match between each element in the predicted trait expression level vector Pi_prediction and the target optimized trait. A distance metric, such as Euclidean distance, can be used to calculate the match. Let the target optimized trait be T_optimize, the j-th element in the predicted trait expression level vector Pi_prediction be Pi_predictionj, and the Euclidean distance between them be distance_j = √(Σ((Pi_predtionj_i - T_optimize_i)²) (i represents the feature dimension). Calculate the distance between all candidate gene combinations and the target optimized trait and sort them in ascending order of distance. The smaller the distance, the higher the match between the candidate gene combination and the target optimized trait. Select the top candidate gene combinations with the smallest distances, and these combinations constitute the optimized gene combination solution set C_optimized_combination.

[0129] Step S157: associating the optimized gene combination scheme with the adaptive optimization strategy and storing them in the crop cultivation system.

[0130] The optimized gene combination set C_optimized_combination is associated with an adaptive optimization strategy. The adaptive optimization strategy can include information such as the direction and magnitude of cultivation parameter adjustments recommended for each optimized gene combination. For example, for a specific optimized gene combination, the adaptive optimization strategy may recommend adjusting cultivation parameters such as fertilizer application rate, irrigation frequency, and light duration. This associated information is stored in the crop cultivation system so that during the actual crop cultivation process, corresponding cultivation parameter adjustments can be automatically triggered based on the currently used gene combination.

[0131] For example, the method may further include:

[0132] Step S210: obtaining a training sample set, wherein the training sample set includes sample fusion prediction features of multiple historical gene sequence samples and corresponding actual phenotypic trait data, wherein each sample fusion prediction feature is associated with at least one sample preset environmental variable.

[0133] To train a genomic prediction model, a training sample set is required. The training sample set consists of multiple historical gene sequence samples. Each historical gene sequence sample contains a sample fusion prediction feature and corresponding actual phenotypic trait data. The sample fusion prediction feature is obtained by performing the previously described series of feature extraction and fusion operations on the historical gene sequence samples and is represented by F_sample. The actual phenotypic trait data is observed and measured in the actual growth environment and is represented by P_actual. Each sample fusion prediction feature is associated with at least one sample preset environmental variable, represented by E_sample. Assuming the training sample set is S_train, it can be represented as S_train = {S_sample1, S_sample2, …, S_samplen}, where S_samplei (i = 1, 2, …, n) represents the i-th historical gene sequence sample. Each historical gene sequence sample S_samplei can be represented as S_samplei = {F_samplei; P_actuali; E_samplei}.

[0134] Step S220: constructing an initial network architecture of the genome prediction model, wherein the initial network architecture includes a cross-attention module and a multi-task learning layer.

[0135] Construct the initial network architecture of the genomic prediction model. This initial network architecture primarily consists of a cross-attention module and a multi-task learning layer. The cross-attention module calculates the interaction weights between the sample fusion prediction features and the sample's preset environmental variables and dynamically weights the sample's preset environmental variables. The multi-task learning layer simultaneously processes multiple related prediction tasks, each corresponding to a different environmental combination. The cross-attention module can include multiple fully connected layers and nonlinear activation functions to transform and process the input features. The multi-task learning layer also includes multiple hidden layers and an output layer to predict trait expression under different environmental combinations. Let the parameters of the cross-attention module be W_attention and b_attention, and the parameters of the multi-task learning layer be W_multi and b_multi. Constructing the initial network architecture involves initializing these parameters. For example, random initialization can be used to initialize the parameters to random values.

[0136] Step S230: Input the sample fusion prediction feature into the cross-attention module, calculate the sample dynamic interaction weight of the sample fusion prediction feature and the sample preset environment variable, and perform weighted fusion on the sample preset environment variable according to the sample interaction weight to generate a sample environment extension feature.

[0137] The sample fusion prediction features F_samplei in the training sample set are input into the cross-attention module. For each sample fusion prediction feature F_samplei, the cross-attention module first calculates the sample dynamic interaction weight between it and the sample preset environment variable E_samplei. The specific calculation process is similar to the process of calculating the interaction weight in the previous step S1411. Let the sample dynamic interaction weight matrix be W_sample_interaction. Its calculation method is the same as the previous interaction weight matrix calculation method, obtained through a series of linear transformations and normalization processes. Then, the sample preset environment variable E_samplei is weightedly fused according to the sample dynamic interaction weight matrix W_sample_interaction. Let the sample environment extended feature vector be E_sample_extendedi, and its calculation method is E_sample_extendedi=Σ(W_sample_interactionij*E_sampleik) (j represents the dimension of the sample fusion prediction feature, and k represents the dimension of the sample preset environment variable). Through this calculation, the sample environment extended feature vector of each historical gene sequence sample is obtained.

[0138] Step S240: Input the sample environment extended features into the multi-task learning layer, and output the predicted trait expression probability distribution of the historical gene sequence samples under multiple environment combinations.

[0139] The sample-environment extended feature vector E_sample_extendedi is input to the multi-task learning layer. The multi-task learning layer performs a series of processing on the input sample-environment extended feature vector, including nonlinear and linear transformations. After processing by the multi-task learning layer, it outputs the predicted trait expression probability distribution vector P_predictedi for the historical gene sequence samples under multiple environmental combinations. The processing of the multi-task learning layer can be viewed as a mapping process, mapping the sample-environment extended feature vector to a vector space representing the trait expression probability under different environmental combinations.

[0140] Step S250: Calculating the parameter gradient of the genomic prediction model based on the cross entropy loss between the predicted trait expression probability distribution and the actual phenotypic trait data.

[0141] Calculate the cross-entropy loss between the predicted trait expression probability distribution vector P_predictedi and the actual phenotypic trait data P_actuali. Cross-entropy loss is a commonly used loss function that measures the difference between two probability distributions. Let the cross-entropy loss function be L_cross_entropy. Its calculation involves performing a series of logarithmic operations and summation operations on the predicted trait expression probability distribution vector and the actual phenotypic trait data. Specifically, for each historical gene sequence sample, calculate its cross-entropy loss value L_cross_entropy_i. Then, sum the cross-entropy loss values ​​for all historical gene sequence samples to obtain the total cross-entropy loss L_total = Σ(L_cross_entropy_i) (i ranges from 1 to n).

[0142] Based on the total cross-entropy loss L_total, the backpropagation algorithm is used to calculate the parameter gradients of the genomic prediction model. Backpropagation is a method for calculating parameter gradients using the chain rule. For the parameters W_attention and b_attention of the criss-cross attention module, and the parameters W_multi and b_multi of the multi-task learning layer, their gradients with respect to the total cross-entropy loss L_total are calculated. Let ∇W_attention and ∇b_attention be the gradients of the criss-cross attention module parameters, and ∇W_multi and ∇b_multi be the gradients of the multi-task learning layer parameters.

[0143] Step S260: After performing gradient clipping on the parameter gradient, an adaptive moment estimation optimizer is used to adjust the network weights of the genome prediction model until the cross entropy loss converges to a preset threshold.

[0144] Gradient clipping is performed on the calculated parameter gradients. The purpose of gradient clipping is to prevent exploding or vanishing gradients. By setting a gradient clipping threshold, threshold_clip, the absolute value of the parameter gradient is limited to the threshold range. Gradient clipping is performed on the gradients ∇W_attention and ∇b_attention of the cross-attention module parameters, and the gradients ∇W_multi and ∇b_multi of the multi-task learning layer parameters. For example, if the absolute value of an element in ∇W_attention is greater than threshold_clip, the value of the element is set to threshold_clip or -threshold_clip.

[0145] An adaptive moment estimation optimizer (such as the Adam optimizer) is used to adjust the network weights of the genomic prediction model. The adaptive moment estimation optimizer adaptively adjusts the learning rate based on the first-order moment estimates (mean) and second-order moment estimates (variance) of the parameter gradients. Let the learning rate be η. In each iteration, the parameters W_attention and b_attention of the cross-attention module, as well as the parameters W_multi and b_multi of the multi-task learning layer, are updated based on the parameter gradients after gradient clipping and the update rule of the adaptive moment estimation optimizer. This process is repeated until the total cross-entropy loss L_total converges to the preset threshold threshold_loss. At each iteration, a new cross-entropy loss value is calculated and the convergence condition is determined.

[0146] Step S270: During each iterative training process, the sample environment extended features are subjected to adversarial perturbation injection processing to generate sample perturbation enhanced features, and the sample perturbation enhanced features are input into the multi-task learning layer for robust regularization training.

[0147] During each training iteration, adversarial perturbations are injected into the extended feature vector E_sample_extendedi (of dimension d). Adversarial samples are generated using the Fast Gradient Signed Method (FGSM). First, the gradient ∇L(E_sample_extendedi) of the extended feature vector E_sample_extendedi with respect to the cross-entropy loss is calculated, ensuring that the dimension of this gradient is consistent with E_sample_extendedi, i.e., d.

[0148] Then, the perturbation vector δi is generated based on the sign of the gradient, δi = ε * sign (∇L (E_sample_extendedi)), where ε is the perturbation strength, which needs to be limited to avoid destroying the validity of the feature. The sample perturbation enhanced feature vector E_sample_perturbedi can be expressed as E_sample_perturbedi = E_sample_extendedi + δi. Since δi and E_sample_extendedi have the same dimension, there is no dimension conflict when adding them.

[0149] The perturbation-enhanced feature vector E_sample_perturbedi is fed into the multi-task learning layer for robust regularization training. During training, the model learns to extract useful information from perturbed samples, thereby improving its resilience to noise and uncertainty. In each iteration, the predicted trait expression probability distribution vector P_predicted_perturbedi is calculated after the perturbation-enhanced feature vector passes through the multi-task learning layer. The cross-entropy loss is then calculated between this vector and the actual phenotypic trait data P_actuali. This cross-entropy loss is added to the overall cross-entropy loss and used together for parameter updates.

[0150] Step S280: performing dimension compression verification on the trained genome prediction model to ensure that the output dimension of the multi-task learning layer is consistent with the probabilistic encoding dimension of the actual phenotypic trait data.

[0151] In this embodiment, the actual phenotypic data is binned or softmax normalized to achieve probabilistic encoding. If binning is used, the range of the actual phenotypic data is divided into several intervals (bins). The frequency of each sample's actual phenotypic data falling within each bin is then counted. These frequencies are used as the probabilistic encoding results to obtain the dimension d_prob of the probabilistically encoded actual phenotypic data.

[0152] During the model building phase, a pre-dimensionality check is performed. Based on the actual data binning number or the probabilistic encoding dimension d_prob, the number of neurons in the multi-task learning layer's output layer is dynamically adjusted to ensure that the output dimension d_output of the multi-task learning layer is consistent with d_prob. After training is complete, d_output and d_prob are rechecked for equality to ensure that the output of the genomic prediction model accurately reflects the probability distribution of the actual phenotypic trait data, thereby improving the prediction accuracy of the genomic prediction model.

[0153] Figure 2A schematic diagram illustrates exemplary hardware and software components of a machine learning-based crop genome-wide prediction system 100 that can implement the concepts of the present application, as provided in some embodiments of the present application. For example, a processor 120 can be used in the machine learning-based crop genome-wide prediction system 100 to perform the functions described in the present application.

[0154] The machine learning-based crop genome-wide prediction system 100 can be a general-purpose server or a special-purpose server, both of which can be used to implement the machine learning-based crop genome-wide prediction method of this application. Although only one server is shown in this application, for convenience, the functions described in this application can be implemented in a distributed manner on multiple similar platforms to balance the processing load.

[0155] For example, the crop genome-wide prediction system 100 based on machine learning may include a network port 110 connected to a network, one or more processors 120 for executing program instructions, a communication bus 130, and storage media 140 in different forms, such as a disk, ROM, or RAM, or any combination thereof. Exemplarily, the crop genome-wide prediction system 100 based on machine learning may also include program instructions stored in ROM, RAM, or other types of non-transitory storage media, or any combination thereof. The method of the present application may be implemented according to these program instructions. The crop genome-wide prediction system 100 based on machine learning also includes an I / O interface 150 between the computer and other input and output devices.

[0156] For ease of explanation, only one processor is described in the machine learning-based crop whole genome prediction system 100. However, it should be noted that the machine learning-based crop whole genome prediction system 100 in this application may also include multiple processors, so the steps performed by one processor described in this application may also be performed jointly or individually by multiple processors. For example, if the processor of the machine learning-based crop whole genome prediction system 100 executes step A and step B, it should be understood that step A and step B may also be performed jointly by two different processors or individually in one processor. For example, the first processor executes step A, the second processor executes step B, or the first processor and the second processor execute steps A and B together.

[0157] In addition, an embodiment of the present invention further provides a readable storage medium, in which computer-executable instructions are preset. When a processor executes the computer-executable instructions, the above-mentioned crop whole genome prediction method based on machine learning is implemented.

[0158] It should be noted that in order to simplify the description of the present invention and thus help understand one or more embodiments of the invention, in the foregoing description of the embodiments of the present invention, multiple features are sometimes combined into one embodiment, figure or description thereof.

Claims

1. A crop genome-wide prediction method based on machine learning, characterized in that: The method comprises: Acquire a gene expression data set of a target crop over multiple growth cycles, wherein the gene expression data set includes multiple gene sequence samples, each gene sequence sample consisting of a genetic marker of at least one gene locus and corresponding phenotypic trait data; Performing feature extraction on the gene expression data set to obtain gene association features and growth trait features of each gene sequence sample; Performing a fusion prediction on the gene association feature and the growth trait feature based on a preset machine learning model to generate a fusion prediction feature of the gene sequence sample; Determining a whole-genome prediction result of the target crop based on the fused prediction features, wherein the whole-genome prediction result is used to indicate a trait expression trend of the target crop under different environmental conditions; generating an adaptive optimization strategy based on the whole genome prediction result, and feeding the adaptive optimization strategy back to the crop cultivation system to trigger a cultivation parameter adjustment operation; The feature extraction of the gene expression data set to obtain gene association features and growth trait features of each gene sequence sample includes: Traversing each gene sequence sample in the gene expression data set, and extracting linkage disequilibrium coefficients between genetic markers at the gene loci; Calling a pre-trained gene feature encoder to perform multi-level encoding processing on each linkage disequilibrium coefficient to generate a gene interaction feature vector of the gene sequence sample; Performing environmental adaptability analysis on the phenotypic trait data to obtain trait stability scores of the gene sequence samples under different growth environments; Performing feature alignment processing on the gene interaction feature vector and the trait stability score to obtain the gene association feature of the gene sequence sample; performing trait association mapping processing on the genetic markers of the gene sequence sample to obtain growth trait characteristics of the gene sequence sample; The performing trait association mapping processing on the genetic markers of the gene sequence sample to obtain the growth trait characteristics of the gene sequence sample includes: Obtaining phenotypic trait data of the gene sequence sample in different growth cycles, and extracting the dynamic change gradient of the phenotypic trait data; constructing a trait association mapping matrix based on the Spearman rank correlation coefficient between the dynamic change gradient and the linkage disequilibrium coefficient of the genetic marker; Performing singular value decomposition on the trait association mapping matrix to obtain principal component characteristics of the phenotypic trait data; Standardizing the principal component features and normalizing the gene interaction feature vectors, aligning the feature dimensions of the standardized principal component features and the normalized gene interaction feature vectors, and performing nonlinear spatial projection processing through a multi-layer perceptron to obtain growth trait characteristics of the gene sequence samples; The fusion prediction of the gene association feature and the growth trait feature based on a preset machine learning model to generate a fusion prediction feature of the gene sequence sample includes: performing a standardization process on the gene association feature to obtain a standardized gene association feature; performing normalization processing on the growth trait characteristics to obtain normalized growth trait characteristics; Invoking a dynamic weight allocation module in the machine learning model to generate a feature fusion weight according to a correlation score between the standardized gene association feature and the normalized growth trait feature; After unifying the standardized gene association features and the normalized growth trait features to the same dimension through the third fully connected layer of the machine learning model, weighted splicing processing is performed based on the feature fusion weight to obtain the fusion prediction features of the gene sequence sample; Determining the whole genome prediction result of the target crop according to the fused prediction features includes: Calling a pre-trained genomic prediction model to perform cross-environment generalization processing on the fused prediction features to obtain the trait expression probability distribution of the target crop under preset environmental variables; Extracting gene site identifiers whose variance in the trait expression probability distribution exceeds a preset threshold value to generate a set of candidate key sites; Calculating the characteristic activation intensity of each gene locus in the candidate key locus set in the environmental extension feature, and screening the gene loci whose activation intensity mean is higher than the genetic significance level to form a key locus set; The trait expression probability distribution is subjected to feature association binding processing with the key gene site set to generate the genome-wide prediction result.

2. The crop genome-wide prediction method based on machine learning according to claim 1, characterized in that: The calling of the pre-trained gene feature encoder to perform multi-level encoding processing on each linkage disequilibrium coefficient to generate a gene interaction feature vector of the gene sequence sample includes: Inputting each linkage disequilibrium coefficient into the first encoding layer of the gene feature encoder to generate initial interaction weights between gene loci; In the second encoding layer of the gene feature encoder, local aggregation processing is performed on the genetic markers of adjacent gene sites based on the initial interaction weights to obtain local aggregation features of the gene site group; In the third encoding layer of the gene feature encoder, after adjusting the initial interaction weight to the same dimension as the local aggregate feature through a fully connected layer, the initial interaction weight and the local aggregate feature are batch normalized respectively, and the normalized local aggregate feature and the initial interaction weight are subjected to channel attention weighted concatenation according to a preset cross-layer connection rule to generate a cross-layer fusion feature; Performing global average pooling processing on the cross-layer fusion features to obtain a gene interaction feature vector of the gene sequence sample.

3. The crop genome-wide prediction method based on machine learning according to claim 1, characterized in that: The calling of the dynamic weight allocation module in the machine learning model to generate a feature fusion weight according to the correlation score between the standardized gene association feature and the normalized growth trait feature includes: Inputting the standardized gene association features into the first fully connected layer of the dynamic weight allocation module to generate a gene association weight vector; Inputting the normalized growth trait features into the second fully connected layer of the dynamic weight allocation module to generate a trait association weight vector; Calculating the cosine similarity between the gene association weight vector and the trait association weight vector, and generating an initial weight score through a nonlinear activation function; Nonlinear activation processing is performed on the initial weight score to generate a feature fusion weight between the standardized gene association feature and the normalized growth trait feature.

4. The crop genome-wide prediction method based on machine learning according to claim 3, characterized in that: The calling of the pre-trained genome prediction model to perform cross-environment generalization processing on the fused prediction features to obtain the trait expression probability distribution of the target crop under preset environmental variables includes: Inputting the fused prediction features into a cross attention module, calculating the interaction weights between the fused prediction features and the preset environmental variables, and dynamically weighted fusion of the preset environmental variables according to the interaction weights to generate environmental extension features; Inputting the environmental extension features into the multi-task learning layer of the genomic prediction model to generate multi-task prediction results of the target crop under different environmental combinations; The multi-task prediction results are subjected to probability normalization processing to obtain the trait expression probability distribution of the target crop under preset environmental variables.

5. The crop genome-wide prediction method based on machine learning according to claim 1, characterized in that: Generating an adaptive optimization strategy based on the whole genome prediction result, and feeding the adaptive optimization strategy back to the crop cultivation system to trigger a cultivation parameter adjustment operation, includes: Determining a contribution score of the key gene locus set to the expression of the target trait based on the whole genome prediction result; generating a site optimization priority sequence according to the difference between the contribution score and a preset genetic gain threshold; Selecting the top N key gene loci as optimization seed loci according to the locus optimization priority sequence, where N is a preset positive integer; Performing crossover recombination simulation processing on the genetic markers of the optimized seed site to generate a candidate gene combination set; Calling a genomic prediction model that has undergone gene recombination data augmentation training and adversarial regularization processing to perform virtual phenotype prediction processing on the candidate gene combination set, and introducing a Monte Carlo dropout layer in the prediction process to simulate data distribution deviation to obtain the predicted trait expression level of each candidate gene combination; Screening out an optimized gene combination scheme from the candidate gene combination set according to the matching degree between the predicted trait expression level and the target optimized trait; The optimized gene combination scheme is associated with the adaptive optimization strategy and stored in the crop breeding system.

6. A crop genome-wide prediction system based on machine learning, characterized in that: The method comprises a processor and a memory, wherein the memory is connected to the processor, the memory is used to store programs, instructions or codes, and the processor is used to execute the programs, instructions or codes in the memory to implement the crop whole genome prediction method based on machine learning as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Method for predicting whole-growth-period image based on crop genotypes of artificial intelligence

    CN119359688A

  • Genome prediction method and model for corn genotype and environment cross-modal feature fusion

    CN119560010A