Intelligent forest tree breeding method based on multi-modal data and large model

By using a multimodal data and large-scale model-based intelligent breeding method for forest trees, the coupling mechanism between forest tree genotypes and dynamic environment is systematically analyzed, and a dynamic whole-genome selection model is constructed. This achieves precision and systematization in forest tree breeding, solves the problems of long breeding cycles and low trait selection efficiency in forest tree breeding, and improves breeding efficiency and multi-trait synergistic genetic gains.

CN121963850APending Publication Date: 2026-05-01RES INST OF FOREST RESOURCE INFORMATION TECHN CHINESE ACADEMY OF FORESTRY
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
RES INST OF FOREST RESOURCE INFORMATION TECHN CHINESE ACADEMY OF FORESTRY
Filing Date
2026-01-19
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Forest tree breeding suffers from problems such as long breeding cycles, low efficiency of trait selection, and slow improvement of genetic gain. Existing technologies are unable to systematically reveal the coupling mechanism of multi-level biomolecular networks such as genome, transcriptome, and metabolome in dynamic environments, resulting in a high false positive rate in the discovery of key regulatory genes, insufficient generalization ability of prediction models across environments and growth stages, and a lack of multi-trait synergistic optimization in breeding decisions.

Method used

We will construct an intelligent breeding method for forest trees based on multimodal data and large models. By integrating genotype, phenotype, multi-omics and environmental time-series data, we will use genotype-environment interaction algorithms and multi-omics coupling analysis to mine key genes or gene modules, construct a dynamic whole-genome selection model, realize intelligent selection and optimization decision-making for multiple traits, and establish a multi-level breeding decision support system for parents, combinations and offspring.

Benefits of technology

It has enabled a systematic and dynamic analysis of the formation mechanism of complex traits in forest trees, significantly improved the accuracy of key gene mining and prediction, enhanced the intelligence and automation of the breeding process, achieved efficient optimization of synergistic genetic gains of multiple traits, and reduced reliance on human experience and overall costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121963850A_ABST
    Figure CN121963850A_ABST
Patent Text Reader

Abstract

The invention provides an intelligent forest tree breeding method based on multi-modal data and a large model, and relates to the technical field of forest tree genetic breeding. The method comprises the steps of obtaining and integrating genotype data, phenotype data, multi-omics data and environment time sequence data of a target tree species, and performing preprocessing and feature fusion to form a multi-modal data set for analysis; on the basis of the multi-modal data set, key genes or gene modules for regulating and controlling target characters are mined by adopting a genotype-environment interaction algorithm and multi-omics coupling analysis; constructing a dynamic model fusing genotypes, environmental factors and time dimensions, and predicting breeding values of individuals or combinations on a whole genome level based on mined key genes or gene module information; and based on a prediction result, performing multi-objective optimization and intelligent screening at three levels of parents, combinations and offspring, and outputting a breeding decision scheme. According to the method, a complete, efficient and accurate forest tree intelligent breeding solution is formed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of forest tree genetics and breeding technology, and in particular to a smart forest tree breeding method based on multimodal data and large models. Background Technology

[0002] As a strategic foundation for national forest resources and timber security, the efficiency and precision of genetic improvement of trees are of paramount importance. However, the long growth cycle of trees, their highly heterogeneous genetic background, and the strong influence of genotype, environment, and their interactions on trait expression lead to inherent challenges in traditional breeding, such as long cycles, low efficiency of trait selection, and slow improvement of genetic gain.

[0003] With the widespread adoption of high-throughput sequencing technology, forestry research has entered the genomics era. Genome-wide association studies (GWAS) and quantitative trait locus (QTL) mapping have been widely applied to identify genetic markers associated with traits such as growth and stress resistance. Furthermore, multi-omics technologies such as transcriptomics and metabolomics provide more dimensions for understanding the molecular mechanisms of trait formation. In breeding prediction, genome-wide selection technology, by utilizing genome-wide molecular markers to predict the breeding value of individuals, has gradually become a core tool in forestry molecular breeding, and its predictive accuracy has been validated in various tree species.

[0004] Despite continuous technological advancements, current research on intelligent breeding of forest trees still faces a series of severe challenges and unresolved bottlenecks. At the level of genetic mechanism analysis, existing studies largely focus on the static association between genotype and phenotype, or simply correct for environmental effects as fixed effects. They fail to systematically reveal how multi-level biomolecular networks, including the genome, transcriptome, and metabolome, respond to and couple with environmental factors under dynamically changing environmental stresses, thereby jointly driving the final phenotype formation. This lack of a dynamic coupling mechanism between genotype, environment, and multiple omics means that the discovery of key regulatory genes still relies on statistical association, resulting in a high false positive rate and difficulty in elucidating the underlying biological causal pathways.

[0005] At the level of predictive model construction, currently widely used genomic selection models, such as Genomic Best Linear Unbiased Prediction (GBLUP) and various Bayesian models, are mostly based on the assumption of static matching between genotype and phenotypic data. However, as perennial plants, forest trees exhibit trait expression as a strongly temporal process, profoundly influenced by dynamic environmental factors such as seasonal changes and interannual climate fluctuations. Existing static models cannot effectively capture and model the patterns of genetic effects changing with environmental time series, resulting in insufficient predictive generalization ability across environments and growth stages, thus limiting their application accuracy in diverse real-world breeding scenarios.

[0006] At the level of breeding decision optimization, forest tree breeding goals typically involve the synergistic improvement of multiple traits, such as simultaneously enhancing growth rate, timber quality, and stress resistance. These traits often involve complex genetic correlations and phenotypic trade-offs. Currently, selection decisions in breeding practice largely rely on independent assessments of single traits or simple index selection, lacking an intelligent and systematic decision-making framework capable of integrating multi-source heterogeneous data (including genomics, phenotyping, and environmental data) and simultaneously considering multiple target traits and their interrelationships. This results in inefficient allocation of breeding resources and makes it difficult to achieve simultaneous optimal genetic progress for multiple traits.

[0007] In conclusion, the current field of forest tree breeding urgently needs a comprehensive technical system that can connect the analysis of microscopic molecular mechanisms to the optimization of macroscopic breeding decisions, in order to overcome the limitations of existing technologies that are scattered, static, and single-objective, and thus achieve precise, efficient, and systematic forest tree genetic improvement. Summary of the Invention

[0008] This application provides a smart forest tree breeding method based on multimodal data and large-scale models. The primary objective is to construct an analytical framework and intelligent method capable of systematically analyzing the coupling mechanisms between forest tree genotypes, dynamic environments, and multi-omics data, thereby more accurately and reliably identifying core genes and pathways regulating the formation of key traits. Another core objective is to develop a dynamic genome-wide selection model that effectively integrates gene, environmental, and temporal dimensions to significantly improve the prediction accuracy and stability of complex traits under different spatiotemporal environments. A further objective is to establish a multi-trait intelligent selection and optimization decision-making method for multiple breeding levels, including parents, hybrid combinations, and offspring individuals, forming a closed-loop breeding decision support system. Ultimately, this aims to maximize the synergistic genetic gains of multiple forest tree traits, fundamentally transforming forest tree breeding from a traditional experience-based model to a data-driven and model-guided intelligent model.

[0009] This application provides a method for intelligent forest tree breeding based on multimodal data and large models, including: Acquire and integrate genotype data, phenotypic data, multi-omics data, and environmental time-series data of the target tree species, perform preprocessing and feature fusion, and form a multimodal dataset for analysis; Based on the aforementioned multimodal dataset, a genotype-environment interaction algorithm and multi-omics coupling analysis are used to identify key genes or gene modules that regulate target traits. A dynamic model integrating genotype, environmental factors, and time dimension is constructed to predict the breeding value of individuals or combinations at the whole genome level based on the information of key genes or gene modules mined. Based on the prediction results, multi-objective optimization and intelligent screening are carried out at three levels: parents, combinations, and offspring, and breeding decision schemes are output.

[0010] As a preferred technical solution, dimensionality reduction and feature engineering of environmental data includes: performing variance inflation factor diagnosis on climate indicators to remove highly collinear variables, and performing principal component analysis on the remaining factors to extract comprehensive environmental principal components with a cumulative variance contribution rate of ≥80%.

[0011] As a preferred technical solution, after obtaining genotype data, the method further includes quality control of the genotype data, including filtering low-quality sites, identifying population structure using principal component analysis, and incorporating principal components as covariates into subsequent models.

[0012] As a preferred technical solution, based on the aforementioned multimodal dataset, a genotype-environment interaction algorithm and multi-omics coupling analysis are used to identify key genes or gene modules that regulate target traits, including: A mixed linear model incorporating the main effects of SNPs, the main effects of environmental factors, and their interaction terms was used for G×E interaction analysis. (1) In the formula, Y Phenotypic observations represent the numerical values ​​of target traits (such as yield, quality indicators, or other quantitative traits) measured in the research subjects under specific environmental conditions. μ The population mean (fixed effect) represents the average level of this trait without considering the effects of genotype, environment, and other covariates. G The main effect (fixed effect) of SNPs represents the direct genetic effect of a specific single nucleotide polymorphism (SNP) on phenotypic traits. E The main environmental effect (fixed effect) represents the influence of different environmental conditions (such as location, year, treatment method, etc.) on phenotypic traits. G×E is the genotype × environment interaction effect (fixed effect), which represents the degree of difference in the expression of different genotypes under different environmental conditions. It is the core term for characterizing gene-environment interaction (G×E). PC 1 and PC 2 represents the population structure covariate (fixed effect), typically the first two principal components obtained from principal component analysis (PCA) of genotype data. This is used to correct for differences in population structure or genetic background, reducing the interference of population differentiation on association analysis results. S Random effects (or random effects) are used to characterize random variation among plots, repeated experimental units, or individuals, reflecting differences caused by unsystematic factors. The random error term represents the unexplained residual variation in the model, typically assumed to have a mean of 0 and a variance of 0. It follows a normal distribution.

[0013] As a preferred technical solution, after the mixed linear model analysis, an attention mechanism is further introduced, and the calculation of its attention weights incorporates the GWAS significance p-value of the SNP sites: (2) In the formula, Q This represents the current genotype feature vector. K This is the characteristic matrix of environmental factors. V Genotype-environment interaction mapping; p j For the first j The SNP loci were obtained through FarmCPU. p The reciprocal of the value is used to assign significance weight to the SNP; exp (·) Transform saliency information into non-negative attention weights. V j It is a value vector that carries environmental response information.

[0014] As a preferred technical solution, based on the aforementioned multimodal dataset, a genotype-environment interaction algorithm and multi-omics coupling analysis are used to identify key genes or gene modules that regulate target traits. This also includes using the SHAP algorithm to analyze key loci and constructing an environmental contribution index. Quantifying the impact of environmental factors: (3) In the formula, ECI represents the Environmental Contribution Index, which measures the overall impact of key environmental factors on the target trait; y represents the target trait value predicted by the model. This represents the target trait y in relation to the t-th environmental factor e. t The partial derivatives are used to characterize the sensitivity of the trait prediction results to changes in the environmental factor; This represents the target trait y in relation to the t-th environmental factor. e t The partial derivative of the environmental factor is used to characterize the sensitivity of the trait prediction results to changes in the environmental factor; T is the number of key environmental factors, and t is the index of the key environmental factors.

[0015] As a preferred technical solution, constructing a dynamic model that integrates genotype, environmental factors, and time dimensions, and predicting the breeding value of an individual or combination at the whole genome level based on the mined key gene or gene module information, includes the following methods: A genome-wide selection model with a dynamic three-way interaction between genes, environment, and time is constructed. A time-aware Transformer encoder captures the nonlinear interactions between genetics and environment under different seasons, and a multi-task deep neural network is used for joint prediction of multiple traits, with dynamic loss weights. Calculated using the following formula: (4) In the formula, For the first k Dynamic weights for each phenotypic task For the first k The weighted score of each task. For the first k The loss function for each task (quantifying the deviation between the predicted and actual values), where K is the total number of phenotypic tasks, exp is the exponential function, and LSTM is the Long Short-Term Memory network.

[0016] As a preferred technical solution, methods for multi-objective optimization and intelligent screening at the parental level include: Parental selection was performed based on whole-genome data using a genome association analysis model, which is: (5) In the formula, These are the phenotypic traits of the parents. It is the label effect vector. It is the residual term. It is a parental genotype matrix; Constructing a multivariate predictive model to assess the breeding value and environmental adaptability of parental lines: (6) In the formula, For the predicted phenotypic values, For multivariate prediction models, For environment variables; Predicting parental breeding values ​​using a regression model: (7) In the formula, For the first Predicted breeding value of individual parents; For the first A genetic marker or omics characteristic; Characteristic effect; For the first The residuals of each parent; The number of genetic markers or omics features.

[0017] As a preferred technical solution, the methods for multi-objective optimization and intelligent screening at the combinatorial level include: Based on the combination force analysis, a mixed linear model was used to evaluate the performance of the combination: (8) In the formula, Let be the sub-representation vector, and b represent the fixed effect. For general fit force GCA, For special cohesive force SCA, For random error, , , It is a design matrix; Maximize the heterosis index by optimizing the objective function: (9) In the formula, and These represent weighting coefficients, indicating the degree of importance attached to GCA and SCA, respectively. Analysis of G×E interactions in multiple environments: (10) In the formula, For the first A combination of genetic effects, For the first An environmental effect, For the first i The genotype in the first j Phenotypic observations under various environmental conditions e j No. j The environmental effects of an environment (i.e., the main environmental effects). The population mean This is the random error term, representing the variation that was not captured by the model.

[0018] As a preferred technical solution, methods for multi-objective optimization and intelligent selection at the offspring level include: Construct a multi-trait genomic selection model and / or a deep learning model for prediction; wherein, the multi-trait genomic selection model is represented as: (11) In the formula, This is a multi-trait phenotypic matrix for offspring. For the fixed effects matrix, This is a genome effects matrix. The error matrix; The deep learning model is represented as follows: (12) In the formula, This is a high-throughput phenotypic characteristic. For environment variables, For the estimated value of the target trait, A deep learning model that integrates genomic, phenotypic, and environmental information.

[0019] Furthermore, the fineness modulus of the manufactured sand is in the range of 2.1 to 3.8, and meets the standard gradation of Zone I or Zone II.

[0020] This application provides a smart forest tree breeding method based on multimodal data and large models, which has at least the following beneficial effects: (1) This application achieves a systematic and dynamic analysis of the formation mechanism of complex traits in forest trees, significantly improving the accuracy and reliability of key gene discovery. Breaking through the limitations of traditional single-omics static association analysis, this application constructs a multidimensional coupling analysis framework of genotype-environmental factors-multi-omics time-series data, and integrates various algorithms such as hybrid linear models, attention mechanisms, and causal inference. This systematically reveals the multi-level regulatory network and causal pathways from genetic variation to phenotypic output under dynamic environmental perturbations. This method not only significantly reduces the false positive rate of candidate gene screening but also helps to elucidate the function of key genes and their regulatory logic under specific environments, providing a solid and interpretable theoretical foundation and precise gene targets for molecular design breeding.

[0021] (2) A predictive model based on the dynamic interaction of genes, environment, and time was constructed, which greatly improved the prediction accuracy and generalization ability of complex traits under different spatiotemporal environments. This application creatively incorporates the time dimension (such as growing season and interannual variation) as a core variable into the whole genome selection model, and uses architectures with time-series modeling capabilities such as Transformer to automatically learn and quantify the nonlinear laws of genetic effects changing with environmental time series. Compared with the traditional static GS model, the dynamic model of this application can more accurately predict the performance of varieties in future environments or different geographical regions that they have not experienced before, significantly enhancing the environmental adaptability and predictive stability of the model, and providing a more reliable decision-making tool for cross-regional breeding and climate change response.

[0022] (3) A multi-level intelligent decision-making system covering the entire process of parents, combinations, and offspring has been established, achieving efficient optimization of multi-trait collaborative breeding. This application deconstructs the breeding process into three key levels: parent selection, combination configuration, and offspring selection, and integrates dedicated prediction and optimization algorithms for each level. By introducing multi-objective optimization algorithms, the system balances the genetic trade-offs among multiple traits such as growth, quality, and resistance, and finally outputs a series of decision-making schemes from optimal parent pairing to high-potential offspring individuals. This systematic approach changes the fragmented state of traditional breeding decision-making, realizes the optimal allocation of breeding resources, and can maximize the comprehensive genetic gain of multiple traits while shortening the breeding cycle.

[0023] (4) It improves the intelligence and automation level of the entire breeding process, effectively reducing reliance on human experience and overall costs. This application constructs a standardized data preprocessing, model training, and decision output process, and integrates large language models and knowledge graph technology into the knowledge management and reasoning stages, forming a closed-loop intelligent system from data to decision. This greatly reduces the repetitive labor intensity of breeding experts in data processing, model building, and complex decision-making, reduces decision bias caused by differences in subjective experience, and makes large-scale, precise intelligent breeding possible, which is conducive to accelerating the breeding process and promoting the large-scale application of scientific and technological achievements.

[0024] In summary, this application has developed a complete, efficient, and precise intelligent forest tree breeding solution through technological innovations in multiple aspects, including multimodal data fusion, dynamic mechanism analysis, intelligent predictive modeling, and system decision optimization. This solution has significant practical application value for breaking through the current core bottlenecks in forest tree breeding and promoting the advancement of seed industry technology. Attached Figure Description

[0025] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0026] Figure 1 A technical roadmap for a forest tree intelligent breeding method based on multimodal data and large models is provided for embodiments of this application; Figure 2 A flowchart illustrating an intelligent forest tree breeding method based on multimodal data and a large model, provided for embodiments of this application; Figure 3 This is a schematic diagram of PubMed literature database retrieval provided in an embodiment of this application; Figure 4 A structural diagram of the whole-genome selection and intelligent prediction model provided in the embodiments of this application; Figure 5 A flowchart illustrating the construction process of a parent-combination-offspring multi-level dataset provided in this application embodiment; Figure 6 A flowchart illustrating the multi-level intelligent selection process of parent-combination-offspring for embodiments of this application; Figure 7 A flowchart illustrating the large language model-assisted genomic selection breeding process provided in this application embodiment; Figure 8 This is a diagram of the WISEED smart breeding platform interface provided in an embodiment of this application. Figure 9 A schematic diagram of the tree species library provided in the embodiments of this application; Figure 10 This is a diagram of the Lin Longda model application interface provided in an embodiment of this application.

[0027] The accompanying drawings illustrate specific embodiments of this application, which will be described in more detail below. These drawings and descriptions are not intended to limit the scope of the concept in any way, but rather to illustrate the concept of this application to those skilled in the art through reference to particular embodiments. Detailed Implementation

[0028] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.

[0029] The collection, storage, use, processing, transmission, provision, and disclosure of relevant data and information in the technical solution of this application all comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0030] It should be noted that in the embodiments of this application, certain software, components, models and other existing solutions in the industry may be mentioned. These should be regarded as exemplary and are only intended to illustrate the feasibility of implementing the technical solution of this application. However, it does not mean that the applicant has used or necessarily used the solution.

[0031] The technical solution of this application and how the technical solution of this application solves the above-mentioned technical problems are described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments. The embodiments of this application will now be described with reference to the accompanying drawings.

[0032] This application provides an intelligent forest tree breeding method based on multimodal data and large models, such as... Figure 1The diagram shown is a technical roadmap for the intelligent forest tree breeding method based on multimodal data and large models provided in this application. Figure 1 As shown, this method revolves around the core objective of "the coupling mechanism of forest tree genotype-environment type-multi-omics and the construction of intelligent breeding models," and comprehensively deploys three research directions: systematically analyzing the multi-omics regulatory network of complex traits in forest trees, constructing a dynamic whole-genome selection algorithm and improving phenotypic prediction capabilities, and developing a large-scale modular genome selection breeding model with autonomous reasoning and intelligent optimization capabilities. Taking poplar, pine, and other major tree species as research objects, the project adopts a four-stage progressive technical route of multimodal data acquisition, molecular mechanism analysis, intelligent modeling and prediction, and decision-making optimization breeding.

[0033] Specifically, this embodiment integrates cross-environmental genotype, phenotype, environmental factor, and multi-omics data based on the PubMed literature database, the National Forest Tree Germplasm Resource Bank, and the Huazhi Biotechnology Breeding Big Data Platform. After standardization and time-series normalization, a high-quality multimodal database is constructed. Combining GWAS, eQTL, and WGCNA co-expression networks with differential expression analysis, deep learning technology is used to analyze the dynamic interaction network of genotype-multi-omics-phenotype under environmental gradients, and a key gene mining algorithm for forest tree genotype-environment type-multi-omics coupling is developed. Furthermore, it integrates genome-wide SNPs, long-term environmental monitoring, and seasonal data. Based on seasonal phenotypic data, we developed a dynamic GS algorithm with a hybrid architecture of MLP-Transformer-multi-task learning. We introduced transfer learning and Bayesian optimization to overcome the bottleneck of spatiotemporal modeling of genetic effects, and constructed a genome-wide selection intelligent algorithm and intelligent prediction model with dynamic interaction of genes, environment and time. We also constructed a breeding agent driven by a large language model and knowledge graph, integrating gene mining, phenotypic prediction, combinatorial optimization and candidate recommendation modules to form a multi-trait collaborative breeding decision system. We developed a large-scale intelligent selection breeding model for important traits of parents, combinations and offspring, and realized closed-loop precision breeding design.

[0034] like Figure 2 As shown, the intelligent forest tree breeding method based on multimodal data and large models includes the following steps S10-S40.

[0035] S10: Acquire and integrate genotype data, phenotypic data, multi-omics data, and environmental time-series data of the target tree species, perform preprocessing and feature fusion, and form a multimodal dataset for analysis.

[0036] In this embodiment, the genotype, phenotype, and multi-omics data are acquired as follows: Based on the PubMed literature database, the National Forest Germplasm Resource Bank, and the Huazhi Bio-breeding Big Data Platform, genotype, phenotype, and multi-omics data of major tree species such as poplar are mined and acquired, and a dataset is constructed. Based on a structured search strategy, core literature from the past decade is screened in the PubMed literature database using keywords, and the full text and supplementary materials are deeply analyzed. Public database access numbers (NCBISRA, GEO) are extracted, and raw sequencing data (FASTQ), genotype files (VCF), and standardized phenotype matrices are downloaded in batches. Figure 3 ).

[0037] The environmental data was acquired as follows: Based on the latitude and longitude of the sample plots, climate data was obtained using a Python web crawler and the ClimateAP software (access address: https: / / web.climateap.net). Based on the geographic coordinates of the plants in the sample plots, location-based spatial interpolation analysis was performed to obtain data on various climate factors, including annual mean temperature, extreme temperatures, annual precipitation, accumulated temperature (≥5℃), and relative humidity.

[0038] Based on the data obtained above, preprocessing and feature fusion are performed through the following steps S101-S103 to form a multimodal dataset for analysis.

[0039] S101: Phenotypic data calibration and optimization.

[0040] A multiple regression model was constructed based on phenotypic data such as tree height, diameter at breast height (DBH), crown width, and biomass. Measured phenotypic parameters were calibrated to eliminate sensor errors (point cloud segmentation bias was corrected using a random forest model). Random forest was used to correct point cloud segmentation errors, and semi-automatic clustering with DBSCAN in CloudCompare, supplemented by manual labeling of outlier clusters, refined individual tree boundaries. Subsequently, phenotypic data was spatially mapped to plot environmental information based on GPS coordinates, achieving seamless matching of genotype-environment-phenotypic three-dimensional data.

[0041] S102: Environmental Data Dimensionality Reduction and Feature Engineering.

[0042] Variance inflation factor (VIF) diagnostics were performed on climate indicators such as annual mean temperature, extreme temperature, annual precipitation, accumulated temperature, and relative humidity to eliminate highly collinear variables with a VIF > 5. Principal component analysis (PCA) was then performed on the remaining climate factors to extract 2–3 comprehensive environmental principal components (such as temperature stress index and water availability index), with a cumulative variance contribution rate ≥ 80%, significantly compressing the input dimensions and providing concise yet information-rich climate characteristics for subsequent models.

[0043] S103: Quality control and standardization of genotype data.

[0044] Low-quality loci (MAF < 0.05 or deletion rate > 10%) were filtered out, triplet sequencing data were merged, and a consistency threshold of Kappa > 0.8 was used to ensure the reliability of the genotype matrix. Subsequently, PCA / STRUCTURE was used to identify population structure caused by geographical isolation, and the first three principal components were included as covariates in the model to suppress false positives. Furthermore, SNPs were matched with environmental principal components to construct an SNP × environment interaction matrix, achieving precise coupling between genotype and environmental dimensions.

[0045] S20: Based on multimodal datasets, using genotype-environment interaction algorithms and multi-omics coupling analysis, we can discover key genes or gene modules that regulate target traits.

[0046] In this embodiment, based on multimodal data and the mixed linear model (MLM) framework, phenotype is used as the response variable, and the fixed effects include SNP main effects, PCA environmental factors and their interaction (SNP×environmental factors). Spatial heterogeneity and population structure principal components are incorporated into the random effects. The multilevel model is shown in formula (1).

[0047] (1) In the formula, Y Phenotypic observations represent the numerical values ​​of target traits (such as yield, quality indicators, or other quantitative traits) measured in the research subjects under specific environmental conditions. μ The population mean (fixed effect) represents the average level of this trait without considering the effects of genotype, environment, and other covariates. G The main effect (fixed effect) of SNPs represents the direct genetic effect of a specific single nucleotide polymorphism (SNP) on phenotypic traits. E The main environmental effect (fixed effect) represents the influence of different environmental conditions (such as location, year, treatment method, etc.) on phenotypic traits. G×E is the genotype × environment interaction effect (fixed effect), which represents the degree of difference in the expression of different genotypes under different environmental conditions. It is the core term for characterizing gene-environment interaction (G×E). PC 1 and PC 2 represents the population structure covariate (fixed effect), typically the first two principal components obtained from principal component analysis (PCA) of genotype data. This is used to correct for differences in population structure or genetic background, reducing the interference of population differentiation on association analysis results. S Random effects (or random effects) are used to characterize random variation among plots, repeated experimental units, or individuals, reflecting differences caused by unsystematic factors. The random error term represents the unexplained residual variation in the model, typically assumed to have a mean of 0 and a variance of 0. It follows a normal distribution.

[0048] A hybrid linear model was used to screen significant SNPs and environmental factors (p < 0.01), and the model structure was simplified according to the Akaike Information Criterion (AIC). Subsequently, a Transformer-MLP hybrid architecture was introduced, in which the SNP embedding stream is used to fuse the GWAS significance p-value with positional encoding (Equation 2). At the same time, LSTM was used to capture temporal fluctuations in soil moisture and a threshold triggering submodule was connected to achieve dynamic coupling between genotype and environment.

[0049] (2) In the formula, Q This represents the current genotype feature vector. K This is the characteristic matrix of environmental factors. V Genotype-environment interaction mapping; p j For the first j The SNP loci were obtained through FarmCPU. p The reciprocal of the value is used to assign significance weight to the SNP; exp (·) Transform saliency information into non-negative attention weights. V j It is a value vector that carries environmental response information.

[0050] The DeepExplainer algorithm based on SHAP was used to analyze key sites (SHAP≥0.01), and the Environmental Contribution Index (ECI) was constructed, as shown in formula (3). (3) In the formula, ECI represents the Environmental Contribution Index, which measures the overall impact of key environmental factors on the target trait; y represents the target trait value predicted by the model. This represents the target trait y in relation to the t-th environmental factor e. t The partial derivatives are used to characterize the sensitivity of the trait prediction results to changes in the environmental factor; This represents the target trait y in relation to the t-th environmental factor. e t The partial derivative is used to characterize the sensitivity of the trait prediction results to changes in the environmental factor. T represents the number of key environmental factors, and t represents the index of the key environmental factors.

[0051] S30: Construct a dynamic model that integrates genotype, environmental factors, and time dimensions, and predict the breeding value of an individual or combination at the whole genome level based on the information of key genes or gene modules mined.

[0052] In this embodiment, the genotype-environment interaction analysis employed a modified MTG2 model, introducing an SNP × environmental factor interaction term into the basic linear model (Y = μ + G + E + ε) to detect environment-specific QTLs (p < 1e-6). The coloc R package was used for eQTL-mQTL colocalization analysis to identify hotspot regions that simultaneously regulate gene expression and metabolite abundance. Subsequently, a time-series co-expression regulatory network (TO-GCN) and a time-weighted gene co-expression network (TWGCNA) were constructed to screen core modules significantly regulated by environmental factors. Finally, structural equation modeling (SEM) was used to analyze the multi-level regulatory pathway of SNP → chromatin opening → gene expression → metabolic flux → phenotype, quantifying the effect values ​​(β coefficients) of each stage and identifying key candidate gene modules regulating forest tree growth traits.

[0053] By analyzing the association between SNPs and target traits, potential functional sites can be effectively identified and used for prediction. Figure 4 The Transformer encoder is used to further identify the complex interactions between SNPs and environmental factors in different seasons, automatically learning which genetic information plays an important role in specific seasons. This algorithm integrates whole-genome SNP data, long-term environmental monitoring factors (such as temperature, light, and moisture), and cross-seasonal phenotypic observation data to achieve dynamic and accurate modeling of phenotypic fluctuations caused by seasonal changes in genetic effects, providing a scientific basis and technical support for intelligent breeding of varieties under different spatiotemporal environments. Specific implementation steps include: (1) Constructing a seasonal dynamic GEBV calculation model: By introducing a Transformer encoder with time-aware capabilities, the interaction between SNP sites and environmental factors at different times is modeled, capturing their temporal dependence and nonlinear response characteristics in the growing seasons of spring, summer, autumn and winter, and then calculating the genome breeding value (GEBV) with seasonal resolution, providing precise support for variety selection and evaluation in stages and regions.

[0054] (2) Design of a dynamic coupling module: A dual-pathway subnetwork of genetics and environment is constructed in the model structure. The genotype pathway embeds seasonally specific SNP sites selected through GWAS analysis, emphasizing genetic variations that have significant regulatory effects at different growth stages. The environmental pathway introduces key seasonal environmental factors (such as temperature, light, and precipitation) identified by the random forest algorithm, highlighting their driving effect on trait expression. The two pathways are coupled through a dynamic weighting mechanism, automatically adjusting the weights of genetic and environmental contributions to phenotype under different seasons, to achieve high-precision modeling of the nonlinear interaction effect of "phenotype-gene-environment".

[0055] In this embodiment, a Long Short-Term Memory (LSTM) network is used to learn the loss weights of each phenotypic task in real time, and a multi-task deep neural network (MT-DNN) is constructed. The genotype-environment interaction feature layer is shared, and shape-specific prediction heads such as tree height and diameter at breast height are learned independently, as shown in formula (4). An adaptive weight allocator is designed to dynamically balance the contribution of the loss function of different phenotypic tasks.

[0056] (4) In the formula, For the first k Dynamic weights for each phenotypic task (higher weights indicate that the model focuses more on predicting that trait). For the first k The weight score of each task (generated by the LSTM network based on the loss value of each task, capturing the dependencies between tasks). Let K be the loss function for the k-th task (quantifying the deviation between the predicted and actual values), where K is the total number of phenotypic tasks (corresponding to the 3 phenotypic indicators of interest). exp The function is an exponential function (mapping the weight scores to positive numbers), and LSTM is a long short-term memory network (dynamically learning the trend of loss changes for each task and outputting weight scores).

[0057] The SHAP (SHapley Additive exPlanations) algorithm is introduced to identify key SNP sites and environmental factor combinations that significantly influence prediction results, and to quantify the relative contribution of each variable to the prediction of the target trait. This helps breeders identify key decision factors and improves the understandability, transparency, and practicality of the model output. A gene co-expression network is constructed to systematically mine gene modules co-expressed under different growing seasons, identify regulatory pathways highly correlated with key complex traits, and further analyze their potential molecular mechanisms, providing systematic support for functional gene screening and precise location of candidate sites. Combining Bayesian methods, and addressing the sample size limitation (6 plots × 10 samples), a Markov Chain Monte Carlo (MCMC) sampling algorithm is introduced to construct a Bayesian hierarchical model. The strength of the G×E effect is estimated through posterior probability, improving the statistical power for small samples. A mixed linear model is integrated with nonlinear algorithms such as random forest and XGBoost. A stacking strategy is used to validate the G×E effect across models, screening interaction terms stable in ≥80% of models to reduce the risk of false positives. Leave-one-out cross-validation (LOOCV) was used to evaluate the generalization ability of the model, with a focus on the predictive stability of the interaction term in independent plots (R²>0.6 is the acceptable threshold). The confidence interval of parameter estimation was optimized by Bootstrap resampling (1000 iterations) (Table 1).

[0058] Table 1 Expected Technical Indicators

[0059] S40: Based on the prediction results, multi-objective optimization and intelligent screening are carried out at three levels: parents, combinations, and offspring, and breeding decision schemes are output.

[0060] Using keywords such as forest tree breeding, tree species, region, environment, phenotype, genome, multi-omics, and G×E interaction, web crawling was used to automatically extract multimodal data, including text, tables, images, remote sensing data, and time series data, from target websites, databases, and online platforms (such as the National Germplasm Resource Bank, Remote Sensing Image Center, and meteorological and ecological monitoring stations). Subsequently, natural language processing (NLP), entity recognition, and image pattern recognition technologies were used to accurately extract core elements such as phylogenetic relationships, trait indicators, and environmental factors from unstructured data. Combined with machine learning algorithms such as association rule mining, cluster analysis, and classification prediction, the potential structure and evolutionary patterns among key variables were uncovered. Finally, through data cleaning, standardization, and semantic fusion techniques, the mining results were deeply integrated with measured data and local databases to construct a unified, multi-level, and multi-dimensional intelligent breeding dataset for parents, combinations, and offspring. Specific information for the three levels of the dataset (parents, combinations, and offspring) is shown below. Figure 5 ): (1) Parent layer dataset This study focuses on the systematic characterization of high-quality genetic resources, primarily encompassing genome-wide genotyping data (such as SNPs, InDel, and CNVs), phenotypic traits (such as growth, stress resistance, and timber properties), pedigree information, and ecological adaptation characteristics. High-throughput sequencing results and phenotypic data are acquired and integrated through web mining of the National Forest Germplasm Bank, literature databases, and laboratory records. Simultaneously, by combining remote sensing products (such as MODIS and Sentinel-2 NDVI), meteorological station data, and topographic information, multimodal data fusion and feature extraction methods are employed to assess the response mechanisms and adaptability of parents under various ecological environments. Multiple regression, principal component analysis (PCA), and cluster analysis are used to form a data layer that supports parent selection and kinship construction.

[0061] (2) Combined layer dataset This study focuses on hybridization design and combination effect assessment, primarily including mating pedigrees, combination design methods, breeding geographical environment background, and preliminary screening data of combination traits. By automatically retrieving data from the breeding registration system, seedling database, and combination experiment records, combined with remote sensing imagery (soil moisture, light, land use) and site factor data, a ternary relationship between combination design, environment, and initial phenotype is constructed. Using mixed linear modeling (MLM) and combining ability analysis (GCA / SCA) methods, a database for predicting combination performance can be established. Simultaneously, the study explores heterosis, environmental adaptability, and stability characteristics among combinations, providing scientific support for optimizing parental pairing strategies and constructing new combinations.

[0062] (3) Subgeneration layer dataset Centered on continuous monitoring data from multiple locations, generations, and modalities, this study focuses on the interaction between offspring performance, genotype, and environment. Through web crawling and API interfaces, it acquires real-time data from high-throughput phenotyping platforms (such as UAV remote sensing, multi-angle imaging, and TLS point clouds) and environmental monitoring systems (such as ground weather stations and pest and disease monitoring platforms). By integrating manual observation records, omics sequencing data (such as GBS and transcriptomics), and remote sensing products, and employing feature selection, classification models (such as SVM and RF), and deep learning algorithms (such as CNN and LSTM), it can construct an intelligent model system for offspring performance-gene-environment. This dataset supports complex quantitative trait prediction, early screening, and individual selection decisions, forming a closed-loop intelligent breeding decision support system.

[0063] A multi-level intelligent selection and prediction method for parents, combinations, and offspring is based on a comprehensive multi-level dataset to achieve precise decision-making throughout the breeding process. This method integrates high-throughput genomic data, multi-environmental phenotypic observations, ecological and environmental variables, and remote sensing monitoring information. Utilizing data mining, machine learning, and deep learning techniques, it achieves precise screening of parental resources, optimal configuration of combinations, and early and accurate prediction of offspring performance. Figure 6 Through multi-level, multi-data source fusion and model iterative optimization, genetic potential and environmental adaptability are systematically identified, enabling intelligent selection and efficient decision-making from germplasm resources to breeding populations and then to superior individuals, significantly improving genetic gain and breeding efficiency.

[0064] (1) Parental hierarchy: Intelligent screening and potential prediction of high-quality core parents At the parental level, combining whole-genome resequencing, transcriptomics, metabolomics, and multi-year, multi-environmental stable phenotypic data, techniques such as cluster analysis, feature selection (LASSO, Boruta), and genome association analysis (GWAS) were used to precisely screen parents with broad genetic diversity and abundant functional genes. Let the parental genotype matrix be... ,in For parental numbers, Number of genotype markers. Utilizing the GWAS model (Formula 5): (5) In the formula, These are the phenotypic traits of the parents. It is the label effect vector. It is the residual term. It is the parental genotype matrix.

[0065] Based on salience markers and environmental variables Construct a multivariate prediction model, as shown in formula (6): (6) In the formula, For the predicted phenotypic values, For multivariate prediction models, For environment variables.

[0066] Machine learning models (such as random forests and support vector machines) can be used to predict the breeding value of parents, assess their genetic potential and environmental adaptability, provide a basis for subsequent combination design, avoid the risk of inbreeding, and maintain genetic diversity.

[0067] The regression model framework for predicting parental potential is shown in formula (7): (7) In the formula, For the first Predicted breeding value of individual parents; For the first A genetic marker or omics characteristic; Characteristic effect; For the first The residuals of each parent; The number of genetic markers or omics features.

[0068] (2) Combination level: Optimal configuration and performance prediction of genetic complementarity and heterosis Based on parental genotypes and offspring phenotypic data from multiple experimental sites, the combination layer uses mixed linear models (such as GBLUP), combining ability analysis (GCA, SCA), and graphical model optimization methods to intelligently screen parental pairing schemes and uncover potential heterosis and genetic complementarity.

[0069] Let the sub-representation be a vector. The model is in the form of (Formula 8): (8) In the formula, Let b be a sub-representative vector, where b represents fixed effects (environmental, operational measures, etc.). For general fit force GCA, For special cohesive force SCA, For random error, , , It is a design matrix.

[0070] By optimizing the objective function, see formula (9): (9) In the formula, and These represent weighting coefficients, indicating the degree of importance attached to GCA and SCA, respectively.

[0071] Combining multiple environments Interactive analysis, see formula (10): (10) In the formula, For the first A combination of genetic effects, For the first An environmental effect, Let i be the phenotypic observation value of the i-th genotype in the j-th environment. e j For the first j The environmental effects of an environment (i.e., the main environmental effects). The population mean This is the random error term, representing the variation that was not captured by the model.

[0072] Finally, the model integrates remote sensing imagery, soil and climate data to construct a multidimensional ecological model based on G×E interaction, which can further accurately predict the performance stability and adaptability of different combinations in multiple environments, realize targeted combination recommendations, and optimize the structure and performance of breeding populations.

[0073] (3) Offspring level: Accurate phenotypic prediction and early screening of superior individuals through multi-source data fusion The progeny layer uses high-density genotype and phenotypic data from multiple environments and time points as its core, and integrates UAV remote sensing, multi-angle imaging and microenvironment monitoring data to construct a multi-trait, multi-environment genomic selection (GS) and deep learning models (such as CNN and Transformer) for joint prediction of complex quantitative traits.

[0074] Let the offspring multi-trait phenotypic matrix be... ,in The number of offspring individuals. The number of traits is used, and the model is in the form of multi-trait genomic selection (MTGS), as shown in formula (11): (11) In the formula, This is a multi-trait phenotypic matrix for offspring. For the fixed effects matrix, This is a genome effects matrix. This is the error matrix.

[0075] Combining deep learning models (such as CNN and Transformer) to perform feature extraction and complex quantitative trait prediction on genotype and environmental time series data, see formula (12): (12) In the formula, This is a high-throughput phenotypic characteristic. For environment variables, For the estimated value of the target trait, A deep learning model that integrates genomic, phenotypic, and environmental information.

[0076] This model balances genetic effects and environmental responses, improving prediction accuracy and stability, and enabling precise early screening of offspring. Through intelligent sorting and selection, it quickly identifies candidate individuals with superior target traits and strong environmental adaptability, significantly shortening the breeding cycle and increasing genetic gain.

[0077] In practical implementation, an intelligent selection program can be constructed based on the above three levels of prediction methods.

[0078] Specifically, the complex genetic background of forest tree breeding, the difficulty of synergistic improvement of multiple traits, and the lengthy breeding cycle severely restrict breeding efficiency and the speed of results transformation. To address these challenges, it is necessary to construct an integrated, multi-level, progressive intelligent selection system that deeply integrates predictive modeling, genetic diversity assessment, and multi-objective optimization decision-making to achieve automated decision support throughout the entire process, from identifying superior parents and optimizing combinations to recommending offspring. Based on high-throughput genomic and phenotypic data, this system constructs precise and interpretable selection paths, providing theoretical support and technical pathways for improving complex traits in forest trees.

[0079] The system's functional modules are designed around a multi-level intelligent selection process of "parents – combinations – offspring," and consist of six core functional units (Table 2): data preprocessing, genetic evaluation, multi-trait prediction model, combination optimization, offspring selection, and decision output. The data preprocessing module integrates heterogeneous data from multiple sources, including parent genotypes, phenotypes, and environmental factors, and inputs it into the system after standardization. The genetic evaluation module provides basic support for parent selection based on combining ability analysis and genetic diversity calculation. The multi-trait prediction model module constructs algorithm models such as GS-BLUP, GBDT, or multi-task deep learning to achieve high-precision prediction of offspring trait performance. The combination optimization module uses the non-dominated sorting genetic algorithm (NSGA-II) for multi-objective decision-making to select the Pareto optimal combination set. The offspring selection module accurately selects offspring individuals with excellent multi-trait performance and genetic representativeness based on trait thresholds, breeding values, and population structure control strategies. Finally, the decision output module outputs parent combination recommendations, offspring selection lists, and genetic diversity assessment results in a visualized report format, providing a scientific basis for breeding practice.

[0080] Table 2. Design of Multi-Level Intelligent Selection Functional Module for Parent-Combination-Offspring

[0081] The key modules and implementation details are as follows: ① Parental screening module This module aims to construct a parental resource system with a broad genetic base and excellent combining ability. First, a genetic distance matrix (such as Jaccard coefficient and IBS similarity) is constructed based on high-throughput SNP genotyping data to systematically characterize the genetic differences between parents. UPGMA clustering or spectral clustering methods are then used to identify genetic groups and reveal population structure characteristics. Based on this, a linear mixture model (LMM) or genome-wide BLUP (GS-BLUP) method is used to estimate the general combining ability (GCA) of parental individuals, serving as a genetic parameter to measure the average offspring's phenotypic ability. Finally, a dual screening strategy is constructed, combining the "top 30% GCA within a group" and "maximizing genetic distance between groups," to select candidate parents with both genetic diversity and combining potential, laying the genetic foundation for efficient combinatorial design.

[0082] ② Combined performance prediction module This module aims to predict the multi-trait breeding potential of offspring by simulating the performance of possible parental combinations. Using genotypic data from screened parents, nonlinear models such as multi-task deep learning (MT-DL) or gradient boosting tree (GBDT) are constructed to accurately predict the offspring's performance in multiple key traits, including tree height, timber properties, and stress resistance. Simultaneously, specific combining ability (SCA) is modeled based on combination bias and GCA differences to further evaluate the specific advantages of combinations. The final output structure includes: predicted multi-trait values ​​for each combination, SCA indices, and the degree of genetic difference between combinations, used for subsequent optimization decisions and combination selection.

[0083] ③ Multi-objective combined optimization module This module constructs a multi-objective optimization system focusing on the multidimensionality of breeding objectives and the efficiency of genetic resource utilization. It sets the maximum multi-trait genetic gain, the maximum SCA value, and the minimum genetic distance redundancy as objective functions. The NSGA-II multi-objective evolutionary algorithm is used to construct the Pareto optimal combination set, systematically exploring the trade-offs among multiple breeding objectives. By visualizing the Pareto front, decision-makers can dynamically adjust weights and select optimal combinations based on actual breeding needs, effectively improving the relevance and scientific rigor of combination design.

[0084] ④ Offspring Selection Module Building upon the combined assessment, this module further enables precise selection of offspring individuals. A first round of screening is conducted based on predefined phenotypic thresholds (such as tree height, diameter at breast height, growth, and disease resistance score). Subsequently, the genomic breeding value (GEBV) of offspring individuals is calculated using the GS model to quantify their genetic potential. Based on this, a multi-trait comprehensive index (such as linear weighted index or TOPSIS) is constructed, and principal component analysis (PCA) or population differentiation coefficient (Fst) is used to control population structure, achieving intelligent selection of high-potential offspring and ensuring coordinated development of multiple traits and the rational maintenance of genetic structure.

[0085] In terms of platform deployment and expansion, the system adopts a development architecture based on the JavaScript ecosystem. The backend uses Node.js to build the service framework, and the frontend uses Vue or React to build a responsive interface. R language is used for genetic data statistics and visual analysis, forming a highly efficient system structure with separation of frontend and backend and decoupling of computation and display. The system supports high-throughput training and prediction of complex forest genetic models and can be flexibly deployed on local servers or cloud platforms.

[0086] The data management layer supports mainstream databases such as MongoDB and PostgreSQL, balancing the document-oriented flexibility of genetic data with the consistency requirements of relational data. It integrates a genetic diversity assessment module and model caching strategies to ensure efficient data access and performance in high-concurrency environments. For visualization, the platform integrates interactive visualization tools such as Plotly.js, ECharts, Deck.gl, and WebGIS, enabling parental combination recommendations, offspring selection path backtracking, graphical display of model results, and spatial distribution visualization, significantly improving the system's interpretability and user experience.

[0087] The system's overall architecture adopts a microservice and RESTful API design, supporting seamless integration with the smart breeding cloud platform. It can also interface with inoculum resource databases, phenotypic observation systems, and breeding process management platforms, demonstrating excellent cross-platform adaptability and modular scalability. To meet future intelligent breeding needs, the system reserves access interfaces for Large Language Models (LLMs) and deep learning inference frameworks (such as HuggingFace Transformers, ONNX Runtime, TensorFlow.js, etc.), supporting functions such as semantic question answering, knowledge graph inference, text generation, and automatic report writing. Combined with cloud-based GPU / TPU acceleration, it enables large-scale genetic combination inference and dynamic model optimization, creating a comprehensive forestry breeding support platform with continuous learning and intelligent decision-making capabilities.

[0088] In some embodiments, the construction and validation of a large-scale intelligent selection model for important traits of parents, combinations, and offspring is carried out, which specifically includes the following four steps.

[0089] Step 1: Building a breeding knowledge graph.

[0090] This study systematically extracts core information related to forest tree breeding from relevant domestic and international literature and materials, forming a multi-dimensional basic dataset covering the association between genes and traits, hybridization patterns between parents and combinations, and the genetic relationship between offspring and traits. For the diverse data in the literature, a multi-path knowledge extraction strategy is adopted: structured data is cleaned and field-mapped to extract gene-trait, parent-combination, and combination-offspring relationships; unstructured text is combined with a large language model to extract trait regulation-related triples, which are then manually verified and improved. Based on knowledge extraction, knowledge fusion is achieved through entity mapping and attribute matching, unifying knowledge from different literature sources under a standardized semantic framework to form a breeding knowledge dataset. Finally, based on the Neo4j graph database, genes, traits, parents, combinations, and offspring are defined as nodes, and regulation, hybridization, and heredity are defined as relationships, constructing a visualized and queryable breeding knowledge graph. On this basis, multimodal content such as text descriptions, tabular data, and image information from the literature is integrated to form a multimodal dataset. By establishing standardized relationships and precise mappings to multimodal content, nodes in the knowledge graph can be directly linked to corresponding original text paragraphs, phenotypic data tables, and related image materials. Ultimately, this enables a coherent application of relational querying, multimodal content tracing, and in-depth interpretation, providing multidimensional evidence for knowledge graph reasoning and rich multimodal knowledge support for breeding research.

[0091] Step 2, design multi-layered prompts.

[0092] This study focuses on the core tasks of genome selection and breeding decision-making, constructing a cue word and instruction set system adapted to the complex genetic logic of forest tree breeding. Through task-layered design, it precisely guides the reasoning process of large-scale models, ensuring efficient handling of the entire process from basic analysis to decision optimization. For basic analysis tasks, universally applicable instruction templates are designed to support core aspects such as genome-level association analysis and phenotypic prediction. These instructions guide the large-scale model to access genetic marker information and gene-trait association patterns in the knowledge graph, combining them with adapted algorithm models to complete basic analysis. The output includes analytical results encompassing association effects, predictions, and the underlying genetic logic, providing data support and theoretical basis for subsequent decision-making. For decision optimization tasks, the instruction design focuses on key aspects such as parental selection, combination configuration, and progeny screening, guiding the large-scale model to integrate multi-dimensional genetic data and breeding objectives, and using scientific decision-making algorithms for comprehensive evaluation. Through instruction-guided modeling, core indicators such as genetic diversity and combining ability are considered during the screening process, constructing a multi-objective optimization decision logic, and ultimately outputting optimized solutions and corresponding analysis and justifications that meet breeding requirements.

[0093] Step 3: Modular large model construction and verification.

[0094] Constructing an intelligent workflow of analysis, decision-making, and feedback, and organically integrating algorithms such as genetic analysis, predictive reasoning, and decision optimization through a unified integration framework, achieves efficient data sharing and dynamic optimization, and promotes collaborative completion of the entire breeding process across all stages. Figure 7 The system was developed and its functional modules were validated based on experimental data. The genetic analysis module integrates algorithms such as genome-wide association analysis and genetic effect analysis. Relying on entity data such as genes, traits, and parents in the knowledge graph, it accurately outputs genetic law results such as the strength of gene-trait associations and allele transmission probabilities through in-depth mining of multi-omics data, providing underlying genetic logic support for subsequent prediction and decision-making. The prediction and reasoning module integrates algorithms such as phenotypic prediction and genome-wide selection. Utilizing relationship information such as gene interactions and parental genetic distance in the knowledge graph, it achieves accurate prediction of multiple traits in offspring phenotypes. The decision optimization module embeds multi-objective optimization algorithms and population genetic structure analysis tools. Combining decision rules such as parental combining ability thresholds and combination gain laws in the knowledge graph, it first clarifies the genetic background differences of candidate parents through population genetic structure analysis, and then uses multi-objective optimization algorithms to balance indicators such as trait gain and genetic diversity, completing multi-level intelligent decision-making from parental selection to combination optimization and offspring selection. Simultaneously, data interfaces were designed, and standardized data interaction formats were defined to achieve dynamic feedback between modules, ensuring the efficient and stable operation of the modular large model.

[0095] Step 4, large model inference and prediction.

[0096] The large-scale model adapts to different breeding scenarios by endowing it with dynamic learning capabilities. It associates the genetic logic of corresponding traits in the knowledge graph with breeding goals such as fast-growing forest breeding and high stress resistance breeding. It designs strategy switching instructions to realize scenario-based switching of reasoning strategies. On this basis, it connects the knowledge base, reasoning strategies and modular workflows, calls relevant knowledge and algorithms from the knowledge graph, and outputs a list of high-quality parental combinations and analysis reports, Pareto optimal combination sets and visualization analysis, and offspring individuals that meet the conditions and related reports. At the same time, the large-scale model can be connected to breeding test data. By comparing the differences between the predicted results and the actual phenotypes and growth performance, it can be verified, providing a basis for model parameter optimization and reasoning logic calibration. It completes the intelligent process from parents to offspring, improving the efficiency and accuracy of breeding complex traits of forest trees.

[0097] This embodiment also provides a feasibility analysis of the method of this application, which specifically includes the following three aspects.

[0098] First, the platform's guarantee is feasible. For example... Figure 8 As shown, the applicant has independently developed a multi-species, large-scale, full-process intelligent breeding platform (WISEED), which provides a system platform guarantee for the project from data analysis to breeding decision-making.

[0099] This application utilizes a core technological approach combining biotechnology (BT), data technology (DT), and large-scale modeling (LM). It focuses on the integration and application of genotype, phenotype, and environmental data, and has developed a one-stop, end-to-end bio-breeding system solution encompassing genotyping, multi-omics analysis, single-cell spatiotemporal omics, precise identification of germplasm resources, gene editing, digital breeding, and smart farm construction. The project team's core products, including mGPS / cGPS liquid phase chips, the WISEED smart breeding platform, and Breeding Butler 4.0, have been applied on a large scale across multiple species, including forest trees. These products have served over 1,300 industry-academia-research institutions and over 4,000 research teams, and have been implemented in 58 digital farms, covering an area of ​​over 2 million mu (approximately 133,333 hectares). This has formed a comprehensive service system that is scalable, replicable, and capable of supporting complex breeding scenarios. Leveraging the advantages of the WISEED intelligent breeding platform, its mature technology system, and extensive practical experience, this project will receive strong, platform-based, and feasible support across the entire chain for functional gene mining, multi-omics intelligent analysis, and application of research results in forest species.

[0100] Second, the data foundation is feasible. The applicant has integrated data such as whole-genome resequencing of forest trees, constructing a dataset of over 100 GB, which provides a rich data foundation for the implementation of the methodology proposed in this application.

[0101] This embodiment has successfully constructed a high-density SNP molecular marker system and a whole-genome liquid chip covering multiple important forest tree species, possessing a solid foundation in molecular marker research and development and mature productization capabilities, providing strong support for the implementation of this project. Figure 9 In the field of poplar, a high-density SNP molecular marker combinatorial system covering five major poplar lineages (white poplar, black poplar, green poplar, large-leaved poplar, and desert poplar) (totaling 60,944 loci) has been developed, and a whole-genome liquid-phase chip has been prepared (patent application number: CN202410920746.8). This chip has been applied to poplar germplasm resource identification and molecular-assisted breeding, demonstrating good universality and interspecific differentiation capabilities. The related results have completed intellectual property rights planning and are ready for widespread application. In the field of Masson pine, a Masson pine SNP molecular marker combinatorial system containing 113,709 loci and its whole-genome chip have been established (patent application number: CN202510071504.0). This chip, based on the Pinus tabuliformis reference genome, possesses good universality and stability and has been widely used in Masson pine germplasm resource background analysis, QTL mapping, GWAS analysis, and whole-genome selection research, demonstrating a high level of technological maturity. This dataset and technology will provide a solid guarantee for the molecular data acquisition, genotypic analysis of breeding materials, and association analysis research of this application.

[0102] Third, the algorithm model is feasible. This embodiment systematically proposes a genotype-environment interaction algorithm and constructs an industry-wide large model (Linlong Large Model) suitable for smart forestry applications, providing an algorithmic and large model foundation for the development of intelligent forest breeding.

[0103] In phenomics research, a platform for 3D reconstruction and structural analysis of single-tree point clouds was constructed, and algorithmic models were developed to assist in the extraction of key growth trait features. This embodiment uses the DeepSeek large-scale model as a foundation to develop the first large-scale model for the forestry and grassland industry—the Linlong large-scale model (…). Figure 10 Based on the characteristics of forestry and grassland industry data and operations, a spatiotemporal large model of forestry and grassland multimodal data is constructed, breaking through the limitations of general models in spatiotemporal data understanding, analysis, and reasoning capabilities, thereby improving the computing and processing capabilities of forestry and grassland operations by more than 50%.

[0104] In summary, the outstanding progressiveness of the method in this application lies in: 1) Intelligent mining scheme for key genes coupled with genotype-environment-multi-omics.

[0105] This application breaks through the limitations of single-omics or static associations, and for the first time systematically and dynamically analyzes the G×E interaction and spatiotemporal dynamic coupling mechanism of forest tree growth traits from multiple levels (genotype-environment-phenotype-multi-omics). It integrates multi-omics time-series data (genome, transcriptome, metabolome, high-throughput phenotype) under different environmental gradients, and combines multiple model methods such as GWAS, temporal gene networks (To-GCN), WGCNA, and causal inference (e.g., dynamic Bayesian networks) to reveal the coupling mechanism of forest tree genotype-environment-multi-omics. Under environmental perturbation, it reveals the dynamic interaction network and causal pathway from gene activation, regulatory cascade to metabolic effects, mines and identifies key genes, develops an intelligent algorithm for key gene mining in genotype-environment-multi-omics coupling, elucidates the mechanism of multi-level synergistic driving trait formation, and lays the core theoretical foundation for constructing an environment-adaptive genomic selection model for forest trees.

[0106] 2) A genome-wide selection intelligent scheme based on the triple dynamic interaction of genes, environment, and time.

[0107] This project is the first to integrate cutting-edge technologies from multiple disciplines, including plant biology, environmental science, and computer and information science, particularly in innovative applications of large-scale artificial intelligence models and multimodal big data analysis. Addressing three major bottlenecks—difficulty in trait genetic analysis, low efficiency in selection modeling, and weak collaborative decision-making capabilities—it constructs an innovative system that deeply integrates genetics, environmental science, and artificial intelligence. At the genetic level, it is the first to integrate GWAS, WGCNA, and multi-omics (transcriptomics + metabolomics + methylome) into a four-dimensional coupling framework to systematically analyze the regulatory network of growth traits under environmental stress, accurately locate key candidate genes, and reveal the molecular mechanisms of genotype-environment-multi-omics coupling. At the environmental science level, it breaks through the limitations of static environmental variables by embedding dynamic factors of "time-microclimate-stress event" to construct a high-resolution environmental time-series database, realizing the nonlinear mapping of environmental stress on genetic effects. At the level of artificial intelligence, we have proposed an original whole-genome selection intelligent algorithm that integrates the gene-environment-time triple dynamic interaction of Transformer-MLP-multi-task learning. By capturing the gene-environment-time triple interaction through the attention mechanism, it significantly improves the accuracy of cross-environment phenotypic prediction and model interpretability, providing a theoretical paradigm and technical support for the precision and efficient breeding of forest trees and the construction of large models.

[0108] 3) Establishment of a large-scale model for multi-level intelligent selection of important traits in parents, combinations, and offspring.

[0109] This application proposes and constructs for the first time a large-scale intelligent selection model for important traits of parents, combinations, and offspring in the field of forest tree genetics and breeding. It constructs a multi-modal dataset of phenotypic and genotype of parents, combinations, and offspring, and adopts multi-objective optimization and intelligent decision-making methods to realize intelligent decision-making throughout the entire process of parent selection, combination optimization, and offspring prediction. First, this model identifies and integrates key gene loci affecting target traits under different developmental stages and environmental conditions by embedding a key gene intelligent mining algorithm that couples genotype, environment, and multiple omics, and a genome-wide selection intelligent algorithm that dynamically interacts with genes, environment, and time. This provides a genetic basis for parental combination optimization. Second, the model addresses the interaction effects between complex traits by using multi-phenotypic and genotypic datasets of parents, combinations, and offspring, along with multi-level intelligent selection algorithms for important traits. This enables parallel optimization of multiple target traits and overcomes the limitations of single-trait optimization in traditional breeding. Finally, relying on a breeding knowledge graph, the model forms a gene selection breeding agent with autonomous logical reasoning, prediction, and decision-making capabilities. It can predict and optimize complex traits across regions and generations, demonstrating broad applicability in various environments and for multiple breeding objectives.

[0110] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.

Claims

1. A method for intelligent forest tree breeding based on multimodal data and large models, characterized in that, include: Acquire and integrate genotype data, phenotypic data, multi-omics data, and environmental time-series data of the target tree species, perform preprocessing and feature fusion, and form a multimodal dataset for analysis; Based on the aforementioned multimodal dataset, a genotype-environment interaction algorithm and multi-omics coupling analysis are used to identify key genes or gene modules that regulate target traits. A dynamic model integrating genotype, environmental factors, and time dimension is constructed to predict the breeding value of individuals or combinations at the whole genome level based on the information of key genes or gene modules mined. Based on the prediction results, multi-objective optimization and intelligent screening are carried out at three levels: parents, combinations, and offspring, and breeding decision schemes are output.

2. The intelligent forest tree breeding method based on multimodal data and large models according to claim 1, characterized in that, Dimensionality reduction and feature engineering of environmental data include: diagnosing variance inflation factors for climate indicators to remove highly collinear variables, and performing principal component analysis on the remaining factors to extract comprehensive environmental principal components with a cumulative variance contribution rate of ≥80%.

3. The intelligent forest tree breeding method based on multimodal data and large models according to claim 1, characterized in that, After acquiring genotype data, the method further includes quality control of the genotype data, including filtering low-quality loci, and using principal component analysis to identify population structure, incorporating principal components as covariates into subsequent models.

4. The intelligent forest tree breeding method based on multimodal data and large models according to claim 1, characterized in that, Based on the aforementioned multimodal dataset, a genotype-environment interaction algorithm and multi-omics coupling analysis are used to identify key genes or gene modules that regulate target traits, including: A mixed linear model incorporating the main effects of SNPs, the main effects of environmental factors, and their interaction terms was used for G×E interaction analysis. (1) In the formula, Y These are phenotypic observations, representing the numerical values ​​of the target traits measured in the research subjects under specific environmental conditions. μ The population mean represents the average level of this trait without considering the influence of genotype, environment, and other covariates. G This represents the main effect of SNPs, indicating the direct genetic effect of a specific single nucleotide polymorphism on phenotypic traits. E The environmental main effect represents the influence of different environmental conditions on phenotypic traits. G×E represents the genotype×environment interaction effect, which represents the degree of difference in the expression of different genotypes under different environmental conditions. PC 1 and PC 2 is a covariate for group structure. S For random effects of sample plots, This is the random error term.

5. The intelligent forest tree breeding method based on multimodal data and large models according to claim 4, characterized in that, Following the mixed linear model analysis, an attention mechanism is further introduced, whose attention weight calculation incorporates the GWAS significance p-value of the SNP locus: (2) In the formula, Q This represents the current genotype feature vector. K This is the characteristic matrix of environmental factors. V Genotype-environment interaction mapping; p j For the first j The SNP loci were obtained through FarmCPU. p The reciprocal of the value is used to assign significance weight to the SNP; exp (·) is an exponential function with the natural constant as its base. V j It is a value vector that carries environmental response information.

6. The intelligent forest tree breeding method based on multimodal data and large models according to claim 4, characterized in that, Based on the aforementioned multimodal dataset, a genotype-environment interaction algorithm and multi-omics coupling analysis are employed to identify key genes or gene modules regulating target traits. This also includes using the SHAP algorithm to analyze key loci and constructing an environmental contribution index. Quantifying the impact of environmental factors: (3) In the formula, The Environmental Contribution Index measures the overall impact of key environmental factors on target traits. r This represents the target trait value predicted by the model; Indicates target traits r For the first t Environmental factors The partial derivatives are used to characterize the sensitivity of the trait prediction results to changes in the environmental factor; Indicates target traits r For the first t Environmental factors e t The partial derivatives are used to characterize the sensitivity of the trait prediction results to changes in the environmental factor; T The number of key environmental factors, t An index of key environmental factors.

7. The intelligent forest tree breeding method based on multimodal data and large models according to claim 1, characterized in that, Methods for constructing dynamic models that integrate genotype, environmental factors, and time dimensions to predict the breeding value of individuals or combinations at the whole genome level based on the information of key genes or gene modules discovered include: A genome-wide selection model with a dynamic three-way interaction between genes, environment, and time is constructed. A time-aware Transformer encoder captures the nonlinear interactions between genetics and environment under different seasons, and a multi-task deep neural network is used for joint prediction of multiple traits, with dynamic loss weights. Calculated using the following formula: (4) In the formula, For the first k Dynamic weights for each phenotypic task For the first k The weighted score of each task. For the first k The loss function for each task, where K is the total number of phenotypic tasks, exp is the exponential function, and LSTM is the Long Short-Term Memory network.

8. The intelligent forest tree breeding method based on multimodal data and large models according to claim 1, characterized in that, Methods for multi-objective optimization and intelligent selection at the parent level include: Parental selection was performed based on whole-genome data using a genome association analysis model, which is: (5) In the formula, These are the phenotypic traits of the parents. It is the label effect vector. It is the residual term. It is a parental genotype matrix; Constructing a multivariate predictive model to assess the breeding value and environmental adaptability of parental lines: (6) In the formula, For the predicted phenotypic values, For multivariate prediction models, For environment variables; Predicting parental breeding values ​​using a regression model: (7) In the formula, For the first Predicted breeding value of individual parents; For the first A genetic marker or omics characteristic; Characteristic effect; For the first The residuals of each parent; The number of genetic markers or omics features.

9. The intelligent forest tree breeding method based on multimodal data and large models according to claim 1, characterized in that, Methods for multi-objective optimization and intelligent selection at the combinatorial level include: Based on the combination force analysis, a mixed linear model was used to evaluate the performance of the combination: (8) In the formula, Let be the sub-representation vector, and b represent the fixed effect. For general fit force GCA, For special cohesive force SCA, For random error, , , It is a design matrix; Maximize the heterosis index by optimizing the objective function: (9) In the formula, and These represent weighting coefficients, indicating the degree of importance attached to GCA and SCA, respectively. Analysis of G×E interactions in multiple environments: (10) In the formula, For the first A combination of genetic effects, For the first An environmental effect, Let i be the phenotypic observation value of the i-th genotype in the j-th environment. e j For the first j The environmental effects of an environment The population mean This is the random error term, representing the variation that was not captured by the model.

10. The intelligent forest tree breeding method based on multimodal data and large models according to claim 1, characterized in that, Methods for multi-objective optimization and intelligent selection at the offspring level include: Construct a multi-trait genomic selection model and / or a deep learning model for prediction; wherein, the multi-trait genomic selection model is represented as: (11) In the formula, This is a multi-trait phenotypic matrix for offspring. For the fixed effects matrix, This is a genome effects matrix. The error matrix; The deep learning model is represented as follows: (12) In the formula, This is a high-throughput phenotypic characteristic. For environment variables, For the estimated value of the target trait, A deep learning model that integrates genomic, phenotypic, and environmental information.

Citation Information

Patent Citations

  • A SNP molecular marker combination of poplar, a whole-genome liquid chip prepared therefrom, and applications thereof

    CN118895377B

  • Pinus massoniana snp molecular marker combination and application thereof

    CN119979753B