Tree phenotype deep learning modeling method for multi-character collaborative prediction

By using multimodal data fusion and genotype-environment interaction algorithms, a multi-trait collaborative prediction model was constructed, which solved the problems of low data acquisition efficiency and poor model generalization ability in traditional tree breeding, and achieved efficient and accurate tree phenotypic prediction and breeding.

CN121658995APending Publication Date: 2026-03-13EXPERIMENTAL CENT OF SUBTROPICAL FORESTRY CHINESE ACAD OF FORESTRY
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202511702081.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-19
Publication Date
2026-03-13

AI Technical Summary

Technical Problem

Traditional tree breeding techniques suffer from problems such as low data acquisition efficiency, insufficient consideration of environmental factors, weak ability to predict multiple traits in a coordinated manner, and poor model generalization ability, making it difficult to meet the needs of efficient and accurate tree breeding.

Method used

Multimodal data were acquired using a high-throughput phenotyping platform, whole-genome sequencing technology, and soil nutrient detection equipment. Principal component analysis, attention mechanism, and convolutional neural network were combined for deep data fusion. Genotype-environment interaction algorithm and multi-task learning framework were introduced, and Bayesian optimization algorithm was used to adjust model hyperparameters to construct a multi-trait collaborative prediction model.

Benefits of technology

It improves the accuracy and stability of tree phenotypic prediction, enables stable prediction under different environments, and provides a scientific basis to accelerate the breeding process of superior varieties.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121658995A_ABST
    Figure CN121658995A_ABST
Patent Text Reader

Abstract

The invention discloses a tree phenotype deep learning modeling method for multi-character collaborative prediction, and relates to the crossing field of tree breeding and artificial intelligence technologies. The method comprises the following components: S1, a multi-modal data acquisition step, S2, a multi-modal data preprocessing and fusion step, S3, a genotype-environment interaction algorithm development step, S4, a multi-character collaborative prediction model construction and optimization step and S5, a model verification and application step. According to the method, the contribution degree of key factors to phenotypes is quantified through an interpretability analysis method, core factors for regulating and controlling phenotype formation are mined in combination with gene function annotation, the process is beneficial to deep understanding of genetic and environmental bases of tree growth, a scientific basis is provided for rapid screening of good varieties, and particularly, the method has the advantages of being high in practicability and convenient to popularize and use. By analyzing a genotype-environment interaction effect, a genotype-environment combination having significant influence on phenotypes is identified, and then a core genotype-core environment factor-key phenotype regulation network is constructed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the intersection of tree breeding and artificial intelligence technology, specifically to a deep learning modeling method for tree phenotypic prediction of multiple traits. Background Technology

[0002] With the increasing severity of global climate change and resource constraints, forestry production has an increasingly urgent need for efficient and precise tree breeding technologies. Tree phenotype, as the result of the combined effects of genetic characteristics and environmental factors, directly determines timber yield, quality, and ecosystem service functions.

[0003] Traditional tree phenotypic prediction methods mainly rely on manual observation and statistical analysis, which have several significant drawbacks: First, data acquisition efficiency is low; manual observation is not only time-consuming and labor-intensive, but also makes it difficult to obtain large-scale, high-precision phenotypic data. Second, environmental factors are not adequately considered; traditional methods often ignore the influence of environmental factors on phenotypes or only consider a single environmental factor, making it difficult to comprehensively reflect the complex interaction between genotype and environment. Third, the ability to predict multiple traits synergistically is weak; tree phenotypes are influenced by multiple traits, and traditional methods cannot predict multiple traits simultaneously, leading to low breeding selection efficiency. Finally, the model generalization ability is poor; prediction methods based on limited data and simple statistical models often perform poorly in new environments or new varieties, failing to meet actual breeding needs. In addition, traditional methods lack in-depth analysis of the genotype-environment interaction mechanism, limiting the rapid screening of superior varieties and accelerating the breeding process.

[0004] To address the problems of low data acquisition efficiency, insufficient consideration of environmental factors, weak multi-trait collaborative prediction ability, and poor model generalization ability in traditional tree breeding techniques, it is particularly important to propose a deep learning modeling method for tree phenotypes for multi-trait collaborative prediction. Summary of the Invention

[0005] The purpose of this invention is to overcome the shortcomings of existing technologies and provide a deep learning modeling method for tree phenotypic prediction based on multi-trait collaborative prediction. This method integrates multimodal data obtained from high-throughput phenotyping platforms, whole-genome sequencing technology, and soil nutrient detection equipment. It utilizes principal component analysis, attention mechanisms, and convolutional neural networks to achieve deep data fusion, effectively removing data redundancy and noise. Furthermore, it introduces genotype-environment interaction algorithms and a multi-task learning framework, combined with Bayesian optimization algorithms to adjust model hyperparameters, ensuring stable predictive performance under different environments. This improves the accuracy and stability of tree phenotypic prediction, providing a scientific basis and technical support for tree breeding and accelerating the cultivation of superior varieties.

[0006] To solve the above-mentioned technical problems, this invention provides the following technical solution: a deep learning modeling method for tree phenotypic prediction based on multi-trait collaborative prediction, the specific steps of which are as follows:

[0007] S1. Multimodal data acquisition steps: High-throughput phenotyping platform, whole genome sequencing technology and soil nutrient detection equipment are used to acquire tree genotype data, environmental data and phenotypic data respectively; the genotype data is SNP locus information within the whole genome, the environmental data is soil physicochemical properties and climatic factor indicators, and the phenotypic data includes tree height, diameter at breast height, crown width, biomass and wood properties parameters;

[0008] S2. Multimodal data preprocessing and fusion steps: Quality control, outlier removal and feature optimization are performed on each modality of data. Principal component analysis is used to reduce the dimensionality of high-dimensional data and remove redundancy. Attention mechanism and convolutional neural network are combined to complete the deep fusion of multimodal data and generate a standardized fusion dataset.

[0009] S3. Genotype-environment interaction algorithm development steps: Based on the standardized fusion dataset, construct a deep neural network with a multilayer perceptron as the basic architecture, extract the nonlinear interaction features between genotype and environmental factors through an improved attention mechanism, and construct a genotype-environment interaction algorithm with quantification interaction effect function.

[0010] S4. Steps for constructing and optimizing a multi-trait collaborative prediction model: Introduce a contrastive learning strategy to optimize the feature representation space, adopt a multi-task learning framework and adaptive weight allocation mechanism, combine Bayesian optimization algorithm to adjust the model hyperparameters, use interpretability analysis method to quantify the contribution of key factors to the phenotype, and establish a multi-trait collaborative prediction model.

[0011] S5. Model Validation and Application Steps: Validate the model's prediction accuracy and stability through biological replication data and field measurement data. Combine gene function annotation to identify the core factors regulating phenotype formation and apply the model to tree phenotype prediction and superior variety selection.

[0012] Furthermore, in S1, genotypic data collection employed the Illumina NovaSeq 6000 high-throughput sequencing platform for whole-genome resequencing. For each plot, 15-20 trees were randomly selected using a gradient selection method, representing individuals with different growth vigors. Each sample was configured with three biological replicates, meaning DNA was extracted from different healthy leaves of the same plant for independent sequencing. SNP selection criteria were: minimum allele frequency (MAF) ≥ 0.05, deletion rate ≤ 10%, and Hardy-Weinberg equilibrium test p-value ≥ 0.01. Environmental data collection utilized a SU-LFPro intelligent soil analyzer and a small automatic weather station. Soil indicators included pH, water content, organic matter content, available nitrogen content, available phosphorus content, and available potassium content. Three biological replicates were set for each plot. Twenty sampling points were set up, and soil samples were collected from two layers, 0-20cm and 20-40cm, at each sampling point. Each layer of sample was measured three times. Climatic factors, including the annual average temperature, annual precipitation, and sunshine duration during the growing season, were obtained through continuous monitoring of automatic weather stations for 12 months. Phenotypic data were collected using a DJI Matrice 350RTK drone and a Trimble X7 ground-based lidar system. The drone was set to a flight altitude of 50m, with a forward overlap of 80% and a lateral overlap of 60%. The ground-based lidar used a grid-based scanning method with key area densification, and a scanning resolution of 0.05m. Point cloud denoising, individual tree segmentation, and parameter extraction were performed using LiDAR 360 6.0 software. Finally, 12 phenotypic parameters, including tree height, diameter at breast height (DBH), crown width, crown height, branch angle, bark thickness, and wood density, were obtained.

[0013] Furthermore, the phenotypic data preprocessing in S2 includes outlier removal and accuracy calibration: outliers are identified using box plots, and values ​​exceeding 1.5 times the interquartile range are marked as outliers and removed after confirmation by field surveys; accuracy calibration uses ground-measured data as a benchmark, constructing a calibration model based on random forest regression. The model input is the raw parameters extracted by lidar, and the output is the ground-measured values. During model training, 5-fold cross-validation is used to optimize parameters, ensuring that the relative error of the phenotypic data after calibration is ≤5%; environmental data preprocessing uses a combination of variance inflation factor analysis and principal component analysis: first, the VIF value of each environmental indicator is calculated, high-collinearity indicators with VIF>5 are removed, and the remaining indicators are standardized using the following formula: ,in For the first The first sample The original values ​​of each indicator For the first The average of the indicators, For the first The standard deviation of each indicator was used to extract principal components using PCA. Principal components with eigenvalues ​​>1 were selected as comprehensive environmental factors, requiring a cumulative variance explanation rate of ≥85%. Genotype data preprocessing included quality control and population structure correction: SNP site filtering was performed using GATK4 software to remove low-quality sites, and genotype imputation was performed using PLINK software, with a missing site rate of ≤3% after imputation. Population structure was analyzed using ADMIXTURE software, and the population structure coefficient was included as a covariate in subsequent models to avoid false positive associations caused by population stratification. Multimodal data fusion adopted a convolutional fusion network based on an attention mechanism. The network input was the preprocessed genotype feature matrix, with dimensions of 1. , For the sample size, For SNP loci, environmental feature matrix, and dimensions , To integrate the number of environmental factors, phenotypic feature matrix, and dimensions , For phenotypic parameters, 1x1 convolutional kernels are used to map data from different modalities to the same feature space, with dimensions... , To fuse feature dimensions, the weights of each modality feature are calculated using a self-attention mechanism, as shown in the formula: ,in For the first Feature matrix of each modality The average matrix of the three modal features. The cosine similarity function is used to finally fuse the features. .

[0014] Furthermore, the genotype-environment interaction algorithm in S3 is constructed based on an improved Transformer-multilayer perceptron hybrid network. The network structure includes an input layer, an interaction feature extraction layer, an interaction effect quantization layer, and an output layer. The input layer receives a standardized fusion feature matrix F and performs positional encoding on the genotype features and environment features respectively. Genotype positional encoding: , ,in This refers to the position number of the SNP site in the genome. Indexed by feature dimensions, Genotype feature dimension; Environmental location encoding: , ,in This refers to the monitoring time / spatial sequence number of the environmental factor. For the environmental feature dimension; the interaction feature extraction layer adopts an improved multi-head self-attention mechanism, introducing a genotype-environment interaction attention head, and the attention weight calculation formula is as follows: ,in For the query matrix, the feature matrix F is fused. The genotype feature matrix contains positional encoding. The environmental feature matrix contains location encoding. for Feature dimensions, for and Interaction feature dimension This is the interaction attention weight coefficient, ranging from 0.1 to 0.5, optimized through 5-fold cross-validation, with an initial value of 0.3. The normalization coefficient is the reciprocal of the number of attention heads. This algorithm uses 8 attention heads. , The normalization function is used; the interaction effect quantification layer calculates the interaction effect value between each SNP locus and the environmental factor through a fully connected network, using the following formula: ,in For the first The SNP locus and the first The interaction effect value of each environmental factor, ranging from 0 to 1, with larger values ​​indicating stronger interactions. For the first The SNP locus at the _ in the _ ... Characteristic values ​​under each environmental factor For the first The environmental factor in the first Feature values ​​for each SNP site , This is the weight matrix. , The bias term, weight matrix and bias term are obtained through training with Adam optimizer, learning rate is set to 0.001, training 500 epochs, learning rate decays by 10% every 100 epochs; output layer outputs genotype-environment interaction effect matrix of each sample, providing a basis for subsequent phenotypic prediction.

[0015] Furthermore, the contrastive learning strategy in S4 combines positive sample mining based on phenotypic similarity with hard negative sample construction: first, the phenotypic similarity between samples is calculated using the formula: ,in For the sample No. Each phenotypic parameter value, , The first The maximum and minimum values ​​of each phenotypic parameter, where l is the total number of phenotypic parameters. To determine the total number of phenotypic parameters, sample pairs with a similarity ≥ 0.8 are considered positive sample pairs. Hard negative samples are constructed using a genotype similarity-environmental difference or environment similarity-genotype difference principle, i.e., sample pairs with genotype similarity ≥ 0.7 and environment similarity ≤ 0.3, or sample pairs with environment similarity ≥ 0.7 and genotype similarity ≤ 0.3, are selected as hard negative sample pairs. The contrastive loss function is: ,in For the sample size, For the sample The fusion characteristics For the sample The interaction effect matrix, , For the sample The positive samples correspond to the features. For the sample The negative sample set, The temperature parameter is set to 0.1 and optimized using a grid search. This represents the loss weight for the interaction effect features, with a value of 0.5, balancing the contributions of the fusion features and the interaction effect features. The cosine similarity function is used. The multi-task learning framework adopts a shared-private feature structure. The shared feature layer consists of the interaction features output by the interaction algorithm, while the private feature layer is a fully connected subnetwork designed for each phenotypic trait. The adaptive weight allocation mechanism dynamically adjusts the weights based on the real-time loss of each task. The weight calculation formula is as follows: ,in For the first Round Weights for each phenotypic prediction task For the first Round The loss value of each task, For the first The average loss across all tasks in the round. The total number of phenotypic prediction tasks; the total loss function of the model is: ,in For the first The mean squared error loss or cross-entropy loss for each task.

[0016] Furthermore, the Bayesian optimization algorithm in S4 employs a tree-based Parzen estimator as the probabilistic model, and the optimized hyperparameters include: the number of hidden layer nodes in the deep neural network, the learning rate, the batch size, the number of attention heads, and the contrastive learning temperature parameter. The objective function is the mean absolute error of the model, and the formula is: ,in For the sample No. Measured values ​​for each phenotype The predicted value is used for the optimization process, which is divided into an exploration phase and a utilization phase. The first 20 iterations are the exploration phase, where random sampling is used to select hyperparameter combinations. The next 80 iterations are the utilization phase, where the optimal combination is selected based on the probability distribution of hyperparameter performance predicted by the TPE model. The interpretability analysis uses an improved SHAP value calculation method, introducing the spatiotemporal weights of environmental factors. The SHAP value calculation formula is: ,in For the sample The Middle The SNP locus and the first The combined SHAP value of the environmental factors, It is the set of all genotype-environment factor combinations. for a subset of For subset The predicted contribution value is calculated using the model's marginal effects. For the first The spatiotemporal weights of each environmental factor For the first The genetic weight of each SNP locus will Genotype-environment combinations with an absolute value ≥ 0.01 are considered key combinations that significantly contribute to the phenotype and are used for subsequent core regulatory factor discovery.

[0017] Furthermore, in S4, a transfer learning framework is introduced for small sample data scenarios, employing a pre-training-fine-tuning mode: In the pre-training phase, a multimodal dataset of closely related tree species is used as training data to train the interaction algorithm and multi-task prediction network. The pre-training cycle is 800 epochs, preserving the weights of the shared feature layer. In the fine-tuning phase, a dataset of the target tree species is used as training data. The first 70% of the weights of the shared feature layer are frozen, and only the last 30% of the weights and the private feature layer weights are trained. The fine-tuning learning rate is set to 1 / 10 of the pre-training learning rate, with 300 training epochs. Model performance is evaluated every 50 epochs. If two consecutive epochs pass... Training is stopped if there is no improvement in performance. To further improve the model's generalization ability under small sample sizes, data augmentation strategies are adopted: genotype data augmentation is achieved through site resampling and slight mutation, that is, 5% of SNP sites are randomly selected for resampling and 1% of sites are artificially mutated; environmental data augmentation is achieved by adding Gaussian noise; phenotypic data augmentation adopts a synthesis method based on generative adversarial networks to construct a phenotypic generation network. The input is genotype and environmental features, and the output is synthetic phenotypic data. The GAN training loss adopts the least squares loss to ensure that the distribution difference between synthetic data and real data is ≤10%.

[0018] Furthermore, the model validation in S5 employs a three-layer validation system: the first layer is cross-validation, using 5-fold cross-validation, dividing the dataset into training and validation sets in a 7:3 ratio, repeating the validation 10 times, calculating the coefficient of determination R², root mean square error, and mean absolute error for each validation, and taking the average of the 10 validations as the model's basic performance index; the second layer is independent sample validation, selecting sample plots with a geographical distance ≥100km from the training set as the independent validation set, accounting for 20% of the total sample size, to verify the model's predictive ability in unfamiliar environments, requiring that the R² of the independent validation set decrease by ≤15% compared to the cross-validation R²; the third layer is time series validation, collecting phenotypic data from the target sample plots for three consecutive years, using the data from the first and second years as the training set and the data from the third year as the validation set, to verify... To verify the time stability of the model, the RMSE of time-series validation must increase by ≤20% compared to the RMSE of validation in the same year. Core regulatory factor mining combines gene functional annotation and environmental effect analysis: First, based on the genomic location of significant SNP sites, candidate genes within a 10kb range upstream and downstream of these sites are obtained from a reference genome database. These are then aligned to the functional databases of Arabidopsis thaliana and Populus tomentosa model plants using BLAST to obtain annotation information for these candidate genes. Next, the influence patterns of core environmental factors on phenotypes are analyzed. The relationship curve between environmental factors and phenotypic values ​​is fitted using a locally weighted regression scatter smoothing method to identify the optimal range of environmental factors. Finally, a regulatory network of core genotype-core environmental factors-key phenotypes is formed, providing specific guidance for parent selection and cultivation environment regulation.

[0019] Compared with existing technologies, this deep learning modeling method for tree phenotypic prediction oriented towards multi-trait collaborative prediction has the following beneficial effects:

[0020] I. This method quantifies the contribution of key factors to phenotype through interpretability analysis and identifies core factors regulating phenotype formation by combining gene function annotation. This process helps to gain a deeper understanding of the genetic and environmental basis of tree growth and provides a scientific basis for the rapid screening of superior varieties. Specifically, by analyzing genotype-environment interaction effects, it identifies genotype-environment combinations that have a significant impact on phenotype, and then constructs a regulatory network of core genotype-core environmental factors-key phenotypes. This network not only reveals the complex mechanism of phenotype formation, but also provides specific guidance for the selection of breeding parents and the regulation of the cultivation environment, accelerating the breeding process of superior varieties.

[0021] II. This method significantly improves the accuracy and stability of tree phenotypic prediction by employing multimodal data acquisition and deep fusion techniques, combined with genotype-environment interaction algorithms and multi-trait collaborative prediction models. Specifically, it utilizes high-throughput phenotyping platforms, whole-genome sequencing technology, and soil nutrient detection equipment to acquire comprehensive data. Principal component analysis, attention mechanisms, and convolutional neural networks are used to achieve deep data fusion, effectively removing data redundancy and enhancing the model's generalization ability. Simultaneously, contrastive learning strategies and multi-task learning frameworks are introduced, along with Bayesian optimization algorithms to adjust model hyperparameters, ensuring stable predictive performance under different environments. These measures collectively improve the model's prediction accuracy and stability, providing strong support for accurate tree phenotypic prediction.

[0022] Other advantages, objectives and features of the invention will be set forth in part in the description which follows, and in part will be apparent to those skilled in the art from the following examination or study, or may be learned from the practice of the invention. Attached Figure Description

[0023] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are merely some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without any creative effort.

[0024] Figure 1 Flowchart of a deep learning modeling method for tree phenotypic prediction for multi-trait collaborative prediction;

[0025] Figure 2 This is a schematic diagram of the deep learning modeling method for tree phenotypic prediction based on multi-trait collaborative prediction. Detailed Implementation

[0026] To further illustrate the technical means and effects of the present invention in achieving its intended purpose, the following detailed description of the specific implementation methods, structures, features, and effects of the present invention, in conjunction with the accompanying drawings and preferred embodiments, is provided below.

[0027] Example 1

[0028] S1, Multimodal Data Acquisition

[0029] Genotypic data were collected using the Illumina NovaSeq 6000 high-throughput sequencing platform for whole-genome resequencing. In three different geographical regions, 18 trees were randomly selected from each pine planting plot according to the principle of random sampling and gradient selection, covering individuals with three growth vigors: fast growth, medium growth, and slow growth. Each sample was set up with three biological replicates, and DNA was extracted from different healthy leaves of the same plant for independent sequencing. SNP site screening followed the criteria of minimum allele frequency (MAF) ≥ 0.05, site deletion rate ≤ 10%, and Hardy-Weinberg equilibrium test p-value ≥ 0.01.

[0030] Environmental data collection was conducted using a SU-LFPro intelligent soil tester and a small automatic weather station. Twenty sampling points were set up in each plot, and two soil samples were collected from each sampling point: 0-20cm and 20-40cm. Each sample layer was measured three times to detect indicators such as soil pH, water content, organic matter content, available nitrogen content, available phosphorus content, and available potassium content. Climate factors, including the plot's annual average temperature, annual precipitation, and sunshine duration during the growing season, were obtained through continuous monitoring for 12 months using the automatic weather station.

[0031] Phenotypic data were collected using a DJI Matrice 350RTK drone and a Trimble X7 ground-based lidar system, specifically including wood property parameters such as tree height, diameter at breast height (DBH), crown width, biomass, wood density, and fiber length.

[0032] S2, Multimodal Data Preprocessing and Fusion

[0033] The phenotypic data preprocessing uses box plots to identify outliers. Values ​​exceeding 1.5 times the interquartile range are marked as outliers and removed after confirmation by field surveys. A calibration model based on random forest regression is constructed using ground-measured data as a benchmark. The model input is the raw parameters extracted by lidar, and the output is the ground-measured values. During model training, 5-fold cross-validation is used to optimize the parameters to ensure that the relative error of the phenotypic data after calibration is ≤5%.

[0034] Environmental data preprocessing first calculates the variance inflation factor (VIF) for each environmental indicator, removes high-collinearity indicators with a VIF > 5, and then standardizes the remaining indicators using the following formula: ,in For the first The first sample The original values ​​of each indicator For the first The average of the indicators, For the first The standard deviation of each indicator is used to extract principal components through principal component analysis. Principal components with eigenvalues ​​> 1 are selected as comprehensive environmental factors, requiring a cumulative variance explanation rate of ≥ 85%.

[0035] Genotype data preprocessing used GATK4 software to filter SNP sites and remove low-quality sites. PLINK software was used for genotype imputation, and the missing site rate after imputation was ≤3%. ADMIXTURE software was used to analyze the population structure, and the population structure coefficient was included as a covariate in the subsequent model.

[0036] Multimodal data fusion employs a convolutional fusion network based on an attention mechanism. The network input consists of preprocessed genotype feature matrix, environmental feature matrix, and phenotypic feature matrix. A 1x1 convolutional kernel maps each modality's data to the same feature space, and then a self-attention mechanism calculates the weights of each modality's features. The formula is as follows: ,in For the first Feature matrix of each modality The average matrix of the three modal features. Using the cosine similarity function, the final standardized fused dataset is obtained. .

[0037] S3, Genotype-Environment Interaction Algorithm Development

[0038] Based on a standardized fused dataset, a deep neural network is constructed using an improved Transformer-multilayer perceptron hybrid network. The network input layer receives the standardized fused feature matrix and performs positional encoding on genotype features and environmental features respectively. Genotype positional encoding: , ,in This refers to the position number of the SNP site in the genome. Indexed by feature dimensions, Genotype feature dimension; Environmental location encoding: , ,in This refers to the monitoring time / spatial sequence number of the environmental factor. This refers to the environmental characteristics dimension.

[0039] The interaction feature extraction layer employs an improved multi-head self-attention mechanism, introducing a genotype-environment interaction attention head. The attention weight calculation formula is as follows: ,in For querying the matrix, This is a genotype feature matrix. This is the environmental feature matrix. for Feature dimensions, for and Interaction feature dimension For interactive attention weight coefficients, The normalization coefficient is... As a normalization function, it integrates information from the query matrix, genotype feature matrix, and environmental feature matrix through an attention weight calculation method.

[0040] The interaction effect quantification layer calculates the interaction effect value between each SNP locus and the environmental factor using a fully connected network, and outputs the quantification result using a combination of the sigmoid and tanh functions. The formula is as follows: ,in For the first The SNP locus and the first The interaction effect values ​​of the environmental factors For the first The SNP locus at the _ in the _ ... Characteristic values ​​under each environmental factor For the first The environmental factor in the first Feature values ​​for each SNP site , This is the weight matrix. , Genetic-environment interaction algorithm with quantification of interaction effects is formed for the bias term.

[0041] S4. Construction and Optimization of Multi-Trait Collaborative Prediction Model

[0042] Introducing a contrastive learning strategy to optimize the feature representation space, we first calculate the phenotypic similarity between samples, using the following formula: ,in For the sample No. Each phenotypic parameter value, , The first The maximum and minimum values ​​of each phenotypic parameter, where l is the total number of phenotypic parameters. Given the total number of phenotypic parameters, sample pairs with similarity ≥ 0.8 are used as positive sample pairs. Hard negative samples are constructed based on the principle of genotype similarity-environmental difference or environment similarity-genotype difference, and the corresponding contrastive loss function is used for training.

[0043] A multi-task learning framework employs a shared-private feature structure. The shared feature layer consists of interactive features output by the interaction algorithm, while the private feature layer comprises fully connected subnetworks designed for each phenotypic trait, such as tree height and diameter at breast height. Based on the real-time loss of each task, weights are dynamically adjusted through an adaptive weight allocation mechanism, and overall optimization is performed in conjunction with the model's total loss function. The weight calculation formula is as follows: ,in For the first Round Weights for each phenotypic prediction task For the first Round The loss value of each task, For the first The average loss across all tasks in the round. Given the total number of phenotypic prediction tasks, the model's total loss function is: ,in For the first The mean squared error loss or cross-entropy loss for each task.

[0044] The model hyperparameters were adjusted using a Bayesian optimization algorithm. The optimized hyperparameters included the number of hidden layer nodes, learning rate, batch size, number of attention heads, and contrastive learning temperature parameter of the deep neural network. The mean absolute error of the model was used as the optimization objective function. The first 20 iterations were the exploration phase, which used random sampling. The next 80 iterations were the utilization phase, which selected the optimal combination of hyperparameters based on the tree-structured Parzen estimator model.

[0045] An improved SHAP value calculation method was used for interpretability analysis, which introduced the spatiotemporal weights of environmental factors and the genetic weights of SNP loci. Genotype-environment combinations with an absolute value of ≥0.01 of the combined SHAP value were regarded as key combinations that made a significant contribution to the phenotype.

[0046] Due to the limited amount of Masson pine sample data, a transfer learning framework was introduced, employing a pre-training-fine-tuning mode. In the pre-training stage, a multimodal dataset of loblolly pine was used as the training data, while in the fine-tuning stage, the Masson pine dataset was used as the training data. The fine-tuning learning rate was set to 1 / 10 of the pre-training learning rate, with 300 training epochs. The model performance was evaluated every 50 epochs, and training was stopped if there was no performance improvement for two consecutive epochs. Simultaneously, data augmentation strategies were adopted. Genotype data was obtained through site resampling and slight mutations, environmental data was obtained by adding Gaussian noise, and phenotypic data was obtained using a synthesis method based on generative adversarial networks. The least squares loss was used for GAN training.

[0047] S5, Model Validation and Application

[0048] A three-layer validation system was used to verify the model performance: The first layer was 5-fold cross-validation, in which the dataset was divided into training and validation sets in a 7:3 ratio, and the validation was repeated 10 times. The coefficient of determination (R²), root mean square error (RMSE), and mean absolute error (MAE) were calculated, and the average of the 10 validations was taken as the basic performance index. The second layer was independent sample validation, in which sample plots with a geographical distance of ≥100km from the training set were selected as independent validation sets, accounting for 20% of the total sample size. The R² of the independent validation set was required to decrease by ≤15% compared to the R² of the cross-validation. The third layer was time series validation, in which phenotypic data of the target sample plots for three consecutive years were collected. The data of the first and second years were used as the training set, and the data of the third year were used as the validation set. The RMSE of the time series validation was required to increase by ≤20% compared to the RMSE of the validation in the same year.

[0049] By combining gene function annotation and environmental effect analysis, core regulatory factors are identified. Based on the genomic location of significant SNP sites, candidate genes within a 10kb range upstream and downstream of them are obtained from the reference genome database. The relationship curve between core environmental factors and phenotypic values ​​is fitted by the local weighted regression scatter smoothing method to identify the optimal range of environmental factors and form a regulatory network of core genotype-core environmental factors-key phenotype. The validated model is applied to predict the phenotype of Masson pine to screen out varieties that grow rapidly in tree height and diameter at breast height and have excellent wood properties under the target growth environment.

[0050] Example 2

[0051] S1, Multimodal Data Acquisition

[0052] Genotypic data were collected using the Illumina NovaSeq 6000 high-throughput sequencing platform for whole-genome resequencing. In fast-growing and high-yield forest plots in five major poplar producing areas, 20 trees were randomly selected from each plot as samples according to the principle of random sampling + gradient selection, covering individuals with different growth years and health conditions. Each sample was set up with 3 biological replicates, and DNA was extracted from different healthy leaves of the same plant for independent sequencing. SNP site screening followed the criteria of minimum allele frequency (MAF) ≥ 0.05, site deletion rate ≤ 10%, and Hardy-Weinberg equilibrium test P value ≥ 0.01.

[0053] Environmental data collection was conducted using a SU-LFPro intelligent soil tester and a small automatic weather station. Twenty sampling points were set up in each plot, and two soil samples were collected from each sampling point: 0-20cm and 20-40cm. Each sample layer was measured three times to detect indicators such as soil pH, water content, organic matter content, available nitrogen content, available phosphorus content, and available potassium content. Climate factors, including the plot's annual average temperature, annual precipitation, and sunshine duration during the growing season, were obtained through continuous monitoring for 12 months using the automatic weather station.

[0054] Phenotypic data were collected using a DJI Matrice 350RTK drone and a Trimble X7 ground-based lidar system, specifically including tree height, diameter at breast height (DBH), crown width, biomass, and wood property parameters such as wood fiber width and cellulose content.

[0055] S2, Multimodal Data Preprocessing and Fusion

[0056] The phenotypic data preprocessing uses box plots to identify outliers. Values ​​exceeding 1.5 times the interquartile range are marked as outliers and removed after confirmation by field surveys. A calibration model based on random forest regression is constructed using ground-measured data as a benchmark. The model input is the raw parameters extracted by lidar, and the output is the ground-measured values. During model training, 5-fold cross-validation is used to optimize the parameters to ensure that the relative error of the phenotypic data after calibration is ≤5%.

[0057] Environmental data preprocessing first calculates the variance inflation factor of each environmental indicator, removes high collinearity indicators with VIF>5, standardizes the remaining indicators, extracts principal components through principal component analysis, selects principal components with eigenvalues>1 as comprehensive environmental factors, and requires a cumulative variance explanation rate of ≥85%.

[0058] Genotype data preprocessing used GATK4 software to filter SNP sites and remove low-quality sites. PLINK software was used for genotype imputation, and the missing site rate after imputation was ≤3%. ADMIXTURE software was used to analyze the population structure, and the population structure coefficient was included as a covariate in the subsequent model.

[0059] Multimodal data fusion employs a convolutional fusion network based on an attention mechanism. The network input consists of preprocessed genotype feature matrix, environmental feature matrix, and phenotypic feature matrix. Each modality data is mapped to the same feature space through a 1x1 convolutional kernel. The weights of each modality feature are then calculated through a self-attention mechanism, ultimately yielding a standardized fusion dataset.

[0060] S3, Genotype-Environment Interaction Algorithm Development

[0061] Based on a standardized fusion dataset, a deep neural network is constructed with an improved Transformer-multilayer perceptron hybrid network. The network input layer receives the standardized fusion feature matrix and performs positional encoding on genotype features and environmental features respectively.

[0062] The interactive feature extraction layer adopts an improved multi-head self-attention mechanism, introducing a genotype-environment interactive attention head, and integrating the information of the query matrix, genotype feature matrix and environment feature matrix through an attention weight calculation method.

[0063] The interaction effect quantification layer calculates the interaction effect value between each SNP locus and environmental factor through a fully connected network, and outputs the quantification result by combining the sigmoid function and the tanh function, forming a genotype-environment interaction algorithm with the function of quantifying interaction effects.

[0064] S4. Construction and Optimization of Multi-Trait Collaborative Prediction Model

[0065] A contrastive learning strategy is introduced to optimize the feature representation space. First, the phenotypic similarity between samples is calculated. Sample pairs with similarity ≥ 0.8 are used as positive sample pairs. Hard negative samples are constructed according to the principle of genotype similarity-environmental difference or environment similarity-genotype difference. The corresponding contrastive loss function is used for training.

[0066] A multi-task learning framework with a shared-private feature structure is adopted. The shared feature layer consists of interactive features output by the interaction algorithm, while the private feature layer consists of fully connected subnetworks designed for each phenotypic trait such as tree height and diameter at breast height. Based on the real-time loss of each task, the weights are dynamically adjusted through an adaptive weight allocation mechanism, and the overall model is optimized by combining the total loss function.

[0067] The model hyperparameters were adjusted using a Bayesian optimization algorithm. The optimized hyperparameters included the number of hidden layer nodes, learning rate, batch size, number of attention heads, and contrastive learning temperature parameter of the deep neural network. The mean absolute error of the model was used as the optimization objective function. The first 20 iterations were the exploration phase, which used random sampling. The next 80 iterations were the utilization phase, which selected the optimal combination of hyperparameters based on the tree-structured Parzen estimator model.

[0068] An improved SHAP value calculation method was used for interpretability analysis, which introduced the spatiotemporal weights of environmental factors and the genetic weights of SNP loci. Genotype-environment combinations with an absolute value of ≥0.01 of the combined SHAP value were regarded as key combinations that made a significant contribution to the phenotype.

[0069] To address the limited sample size of some newly introduced poplar germplasm resources, a transfer learning framework was introduced, employing a pre-training-fine-tuning model. During the pre-training phase, a multimodal dataset of local dominant poplar varieties was used as training data, while the fine-tuning phase used a dataset of newly introduced poplar varieties. The fine-tuning learning rate was set to 1 / 10 of the pre-training learning rate, with 300 training epochs. Model performance was evaluated every 50 epochs, and training was stopped if there was no performance improvement for two consecutive epochs. Simultaneously, data augmentation strategies were employed: genotype data was obtained through site resampling and minor mutations, environmental data was obtained by adding Gaussian noise, and phenotypic data was synthesized using a generative adversarial network (GAN)-based method. The least squares loss was used for GAN training.

[0070] S5, Model Validation and Application

[0071] A three-layer validation system was used to verify the model performance: The first layer was 5-fold cross-validation, in which the dataset was divided into training and validation sets in a 7:3 ratio, and the validation was repeated 10 times. The coefficient of determination (R²), root mean square error (RMSE), and mean absolute error (MAE) were calculated, and the average of the 10 validations was taken as the basic performance index. The second layer was independent sample validation, in which sample plots with a geographical distance of ≥100km from the training set were selected as independent validation sets, accounting for 20% of the total sample size. The R² of the independent validation set was required to decrease by ≤15% compared to the R² of the cross-validation. The third layer was time series validation, in which phenotypic data of the target sample plots for three consecutive years were collected. The data of the first and second years were used as the training set, and the data of the third year were used as the validation set. The RMSE of the time series validation was required to increase by ≤20% compared to the RMSE of the validation in the same year.

[0072] By combining gene function annotation and environmental effect analysis, core regulatory factors are identified. Based on the genomic location of significant SNP sites, candidate genes within a 10kb range upstream and downstream of these sites are obtained from a reference genome database. The relationship curve between core environmental factors and phenotypic values ​​is fitted using a locally weighted regression scatter smoothing method to identify the optimal range of environmental factors and form a regulatory network of core genotype-core environmental factors-key phenotype. The validated model is then applied to the cultivation of fast-growing and high-yield poplar forests to predict the growth performance of different genotypes of poplar under different soil and climatic conditions. This provides a scientific basis for forest site selection and variety matching, and allows for the screening of fast-growing poplar varieties with excellent timber quality for widespread planting.

[0073] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention in any way. Although the present invention has been disclosed above with reference to preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make some modifications or alterations to the above-disclosed technical content to create equivalent embodiments without departing from the scope of the present invention. Any simple modifications, equivalent changes and alterations made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the scope of the present invention.

Claims

1. A deep learning modeling method for tree phenotypic prediction based on multi-trait collaborative prediction, characterized in that, The specific steps of this method are as follows: S1. Multimodal data acquisition steps: High-throughput phenotyping platform, whole-genome sequencing technology and soil nutrient detection equipment are used to acquire tree genotype data, environmental data and phenotypic data respectively; the genotype data is SNP (Single Nucleotide Polymorphism) site information across the whole genome, the environmental data is soil physicochemical properties and climatic factor indicators, and the phenotypic data includes tree height, diameter at breast height, crown width, biomass and wood properties parameters; S2. Multimodal data preprocessing and fusion steps: Quality control, outlier removal and feature optimization are performed on each modality of data. Principal component analysis is used to reduce the dimensionality of high-dimensional data and remove redundancy. Attention mechanism and convolutional neural network are combined to complete the deep fusion of multimodal data and generate a standardized fusion dataset. S3. Genotype-environment interaction algorithm development steps: Based on the standardized fusion dataset, construct a deep neural network with a multilayer perceptron as the basic architecture, extract the nonlinear interaction features between genotype and environmental factors through an improved attention mechanism, and construct a genotype-environment interaction algorithm with quantification interaction effect function. S4. Steps for constructing and optimizing a multi-trait collaborative prediction model: Introduce a contrastive learning strategy to optimize the feature representation space, adopt a multi-task learning framework and adaptive weight allocation mechanism, combine Bayesian optimization algorithm to adjust the model hyperparameters, use interpretability analysis method to quantify the contribution of key factors to the phenotype, and establish a multi-trait collaborative prediction model. S5. Model Validation and Application Steps: Validate the model's prediction accuracy and stability through biological replication data and field measurement data. Combine gene function annotation to identify the core factors regulating phenotype formation and apply the model to tree phenotype prediction and superior variety selection.

2. The deep learning modeling method for tree phenotypic prediction based on multi-trait collaborative prediction according to claim 1, characterized in that, Genotypic data collection in S1 was performed using the Illumina NovaSeq 6000 high-throughput sequencing platform for whole-genome resequencing. For each plot, 15-20 trees were randomly selected using a gradient selection method, representing individuals with different growth vigors. Each sample was configured with three biological replicates, meaning DNA was extracted from different healthy leaves of the same plant and sequenced independently. SNP selection criteria were: minimum allele frequency (MAF) ≥ 0.05, deletion rate ≤ 10%, and Hardy-Weinberg equilibrium test p-value ≥ 0.

01. Environmental data collection was performed using SU... - The LFPro intelligent soil tester and a small automatic weather station were used to measure soil indicators including pH, moisture content, organic matter content, available nitrogen content, available phosphorus content, and available potassium content. Twenty sampling points were set up in each plot, and two soil samples were collected from each sampling point: 0-20cm and 20-40cm. Each sample layer was measured three times. Climate factors included the plot's annual average temperature, annual precipitation, and sunshine duration during the growing season, which were obtained through continuous monitoring for 12 months using the automatic weather station. Phenotypic data were collected using a DJI Matrice 350RTK drone and a Trimble X7 ground lidar system.

3. The deep learning modeling method for tree phenotypic prediction based on multi-trait collaborative prediction according to claim 1, characterized in that, The phenotypic data preprocessing in S2 includes outlier removal and accuracy calibration: outliers are identified using box plots, and values ​​exceeding 1.5 times the interquartile range are marked as outliers and removed after confirmation by field surveys; accuracy calibration uses ground-measured data as a benchmark, constructing a calibration model based on random forest regression. The model input is the raw parameters extracted by lidar, and the output is the ground-measured values. During model training, 5-fold cross-validation is used to optimize parameters, ensuring that the relative error of the phenotypic data after calibration is ≤5%; environmental data preprocessing uses a combination of variance inflation factor analysis and principal component analysis: first, the VIF value of each environmental indicator is calculated, high-collinearity indicators with VIF>5 are removed, and the remaining indicators are standardized using the following formula: ,in For the first The first sample The original values ​​of each indicator For the first The average of the indicators, For the first The standard deviation of each indicator was used to extract principal components using PCA. Principal components with eigenvalues ​​>1 were selected as comprehensive environmental factors, requiring a cumulative variance explanation rate of ≥85%. Genotype data preprocessing included quality control and population structure correction: SNP site filtering was performed using GATK4 software to remove low-quality sites, and genotype imputation was performed using PLINK software, with a missing site rate of ≤3% after imputation. Population structure was analyzed using ADMIXTURE software, and the population structure coefficient was included as a covariate in the subsequent model to avoid false positive associations caused by population stratification. Multimodal data fusion adopted a convolutional fusion network based on an attention mechanism. The network input consisted of the preprocessed genotype feature matrix, environmental feature matrix, and phenotypic feature matrix. A 1x1 convolutional kernel was used to map each modality of data to the same feature space, and then the weights of each modality feature were calculated using a self-attention mechanism. The formula is as follows: ,in For the first Feature matrix of each modality The average matrix of the three modal features. The cosine similarity function is used to finally fuse the features. .

4. The deep learning modeling method for tree phenotypic prediction based on multi-trait collaborative prediction according to claim 1, characterized in that, The genotype-environment interaction algorithm in S3 is based on an improved Transformer-multilayer perceptron hybrid network. The network structure includes an input layer, an interaction feature extraction layer, an interaction effect quantization layer, and an output layer. The input layer receives a standardized fusion feature matrix F and performs positional encoding on genotype and environment features respectively. Genotype positional encoding: , ,in This refers to the position number of the SNP site in the genome. Indexed by feature dimensions, Genotype characteristics dimension; Environmental location coding: , ,in This refers to the monitoring time / spatial sequence number of the environmental factor. For environmental characteristics; The interaction feature extraction layer employs an improved multi-head self-attention mechanism, introducing a genotype-environment interaction attention head. The attention weight calculation formula is as follows: ,in For querying the matrix, This is a genotype feature matrix. This is the environmental feature matrix. for Feature dimensions, for and Interaction feature dimension For interactive attention weight coefficients, The normalization coefficient is... The normalization function is used; the interaction effect quantification layer calculates the interaction effect value between each SNP locus and the environmental factor through a fully connected network, using the following formula: ,in For the first The SNP locus and the first The interaction effect values ​​of the environmental factors For the first The SNP locus at the _ in the _ ... Characteristic values ​​under each environmental factor For the first The environmental factor in the first Feature values ​​for each SNP site , This is the weight matrix. , This is a bias term.

5. The deep learning modeling method for tree phenotypic prediction based on multi-trait collaborative prediction according to claim 1, characterized in that, The contrastive learning strategy in S4 combines positive sample mining based on phenotypic similarity with hard negative sample construction: First, the phenotypic similarity between samples is calculated using the following formula: ,in For the sample No. Each phenotypic parameter value, , The first The maximum and minimum values ​​of each phenotypic parameter, where l is the total number of phenotypic parameters. For the total number of phenotypic parameters, sample pairs with a similarity ≥ 0.8 are considered positive sample pairs; hard negative samples are constructed by screening based on the principle of genotype similarity-environmental difference or environment similarity-genotype difference; the contrastive loss function is: ,in For the sample size, For the sample The fusion characteristics For the sample The interaction effect matrix, , For the sample The positive samples correspond to the features. For the sample The negative sample set, For temperature parameters, The loss weights are characteristics of the interaction effect. The cosine similarity function is used; the multi-task learning framework adopts a shared-private feature structure, where the shared feature layer consists of the interactive features output by the interaction algorithm, and the private feature layer consists of a fully connected sub-network designed for each phenotypic trait. The adaptive weight allocation mechanism dynamically adjusts the weights based on the real-time loss of each task. The weight calculation formula is as follows: ,in For the first Round Weights for each phenotypic prediction task For the first Round The loss value of each task, For the first The average loss across all tasks in the round. The total number of phenotypic prediction tasks; the total loss function of the model is: ,in For the first The mean squared error loss or cross-entropy loss for each task.

6. The deep learning modeling method for tree phenotypic prediction based on multi-trait collaborative prediction according to claim 1, characterized in that, The Bayesian optimization algorithm in S4 uses a tree-based Parzen estimator as the probabilistic model. The optimized hyperparameters include: the number of hidden layer nodes in the deep neural network, the learning rate, the batch size, the number of attention heads, and the contrastive learning temperature parameter. The objective function is the mean absolute error of the model, and the formula is: ,in For the sample No. Measured values ​​for each phenotype The predicted value is used for the optimization process, which is divided into an exploration phase and a utilization phase. The first 20 iterations are the exploration phase, where random sampling is used to select hyperparameter combinations. The next 80 iterations are the utilization phase, where the optimal combination is selected based on the probability distribution of hyperparameter performance predicted by the TPE model. The interpretability analysis uses an improved SHAP value calculation method, introducing the spatiotemporal weights of environmental factors. The SHAP value calculation formula is: ,in For the sample The Middle The SNP locus and the first The combined SHAP value of the environmental factors, It is the set of all genotype-environment factor combinations. for a subset of For subset The predicted contribution value is calculated using the model's marginal effects. For the first The spatiotemporal weights of environmental factors For the first The genetic weight of each SNP locus will Genotype-environment combinations with an absolute value ≥ 0.01 are considered key combinations that significantly contribute to the phenotype and are used for subsequent core regulatory factor discovery.

7. The deep learning modeling method for tree phenotypic prediction based on multi-trait collaborative prediction according to claim 1, characterized in that, In section S4, a transfer learning framework is introduced for scenarios with small sample data, employing a pre-training-fine-tuning mode: During the pre-training phase, a multimodal dataset of closely related tree species is used as training data to train the interaction algorithm and the multi-task prediction network; during the fine-tuning phase, a dataset of the target tree species is used as training data, with the fine-tuning learning rate set to 1 / 10 of the pre-training learning rate, and 300 training rounds are conducted. Model performance is evaluated every 50 rounds, and training stops if there is no performance improvement for two consecutive rounds. A data augmentation strategy is employed: genotype data augmentation is achieved through site resampling and minor mutations, i.e., 5% of SNP sites are randomly selected for resampling, and 1% of sites are artificially mutated. Environmental data augmentation is achieved by adding Gaussian noise; phenotypic data augmentation adopts a synthesis method based on generative adversarial networks to construct a phenotypic generation network. The input is genotype and environmental features, and the output is synthesized phenotypic data. The least squares loss is used for GAN training loss.

8. The deep learning modeling method for tree phenotypic prediction based on multi-trait collaborative prediction according to claim 1, characterized in that, The model validation in S5 adopts a three-layer validation system: the first layer is cross-validation, which uses 5-fold cross-validation. The dataset is divided into training set and validation set in a 7:3 ratio. The validation is repeated 10 times. The coefficient of determination R², root mean square error and mean absolute error are calculated for each validation. The average value of the 10 validations is taken as the basic performance index of the model. The second layer is independent sample validation, selecting plots geographically ≥100km from the training set as the independent validation set, accounting for 20% of the total sample size. This validates the model's predictive ability in unfamiliar environments, requiring the R² of the independent validation set to decrease by ≤15% compared to the R² of the cross-validation. The third layer is time-series validation, collecting phenotypic data from the target plots for three consecutive years, using data from the first and second years as the training set and data from the third year as the validation set. This validates the model's temporal stability, requiring the RMSE of the time-series validation to increase by ≤20% compared to the RMSE of the validation in the same year. Core regulatory factor mining combined with gene function annotation and environmental effect analysis: First, based on the genomic location of significant SNP sites, candidate genes within a 10kb range upstream and downstream of them are obtained from the reference genome database. Then, the influence pattern of core environmental factors on phenotypes is analyzed, and the relationship curve between environmental factors and phenotypic values ​​is fitted using a locally weighted regression scatter smoothing method to identify the optimal range of environmental factors. Finally, a regulatory network of core genotype-core environmental factors-key phenotypes is formed.

Citation Information

Patent Citations

  • Gating and linear attention mechanism G*E interaction fused genome prediction method

    CN118866092A

  • Mulberry SNP (Single Nucleotide Polymorphism) marker mining and stress resistance character prediction method based on Transform architecture

    CN120636548A

  • Methods and systems to enhance a plant breeding pipeline

    WO2023250482A1