A method and apparatus for constructing predictive models for predicting plant phenotypes.

The prediction model built through deep learning, utilizing chromosome perception and comparative learning functions, solves the problems of traditional breeding methods relying on experience and models not considering non-additive effects, and achieves efficient and accurate plant phenotypic prediction.

CN120048340BActive Publication Date: 2026-04-03INST OF GENETICS & DEVELOPMENTAL BIOLOGY CHINESE ACAD OF SCI
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-26
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

Existing breeding methods rely on experience and are inefficient. Furthermore, traditional genome prediction models fail to effectively consider non-additive effects, making it difficult to improve prediction accuracy and stability in the context of synergistic improvement of multiple traits and high-dimensional genome data.

Method used

A prediction model is constructed using deep learning methods. Local and global SNP sites are obtained through the feature selection module of chromosome perception function. Combined with the output network module of contrastive learning function, a multi-task model is constructed to consider additive and non-additive genetic effects and optimize feature selection.

Benefits of technology

It improves the accuracy and stability of plant phenotypic prediction, effectively captures complex nonlinear relationships in high-dimensional genomic data, and enhances breeding efficiency and precision.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120048340B_ABST
    Figure CN120048340B_ABST
Patent Text Reader

Abstract

This disclosure presents a method and apparatus for constructing a predictive model for predicting plant phenotypes. The method includes: acquiring whole-genome SNP loci and phenotypic values ​​of each first plant sample and pairing them; inputting the paired whole-genome SNP loci and paired phenotypic values ​​into a feature selection module with chromosome sensing function to obtain target whole-genome SNP loci and maintain the pairing between the target whole-genome SNP loci and the paired phenotypic values; inputting different second plant samples into an output network module with contrastive learning function to generate feature output and low-dimensional feature output, wherein the second plant samples have paired target whole-genome SNP loci and paired phenotypic values; calculating the total loss based on the feature output, low-dimensional feature output, and paired phenotypic values; and constructing a predictive model by minimizing the total loss. The established chromosome sensing and contrastive learning functions enable the constructed model to achieve higher accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention generally relates to the technical field of constructing predictive models. More specifically, this invention relates to a method for constructing predictive models for predicting plant phenotypes. Background Technology

[0002] Traditional crop breeding is based on phenotypic selection, choosing superior offspring by observing variations in crop phenotypes. While breeders can design and select field trials using general principles of biological inheritance, comprehensive selection indices, comparisons of contemporaneous populations, and statistical analysis of field trials, the success of such breeding relies heavily on the breeder's experience, resulting in low efficiency. Currently, the application of molecular marker-assisted selection (MMR) breeding technology is becoming increasingly mature, but it is only applicable to traits determined by a small number of major QTLs (quantitative trait loci). Actual crop breeding work requires the synergistic improvement of multiple traits. Breeding projects offer breeders hundreds or even thousands of materials, with even more combinations available. However, due to limitations in trial scale, many important materials remain untested, hindering the breeding process. Both of these methods have significant limitations.

[0003] Genomic selection (GS) utilizes genome-wide molecular markers and phenotypic data from samples to build predictive models for early genetic assessment of individuals. GS breeding technology eliminates the need to identify SNPs (Single Nucleotide Polymorphisms) significantly associated with the target trait. Even the genetic effects of a single locus with a small effect on phenotypic variation can be captured by high-density genetic markers, and the breeding value can be evaluated through individual genotype. This significantly improves breeding precision, shortens the breeding cycle, and increases breeding efficiency, achieving a leap from empirical breeding to precision breeding, and has become a cutting-edge technology in plant and animal breeding.

[0004] Genetic effects in genomics (GS) primarily include additive and non-additive effects. Additive effects can be interpreted as linear relationships between loci, while non-additive effects include epistatic and dominant effects. Epistatic effects refer to non-linear relationships between loci, while dominant effects are the interactions between alleles at the same locus. Traditional genome prediction models mostly consider only additive effects when estimating genetic effects, lacking design considerations for non-additive effects. Therefore, constructing genome prediction models should take non-additive effects into account.

[0005] Deep learning is a crucial approach for assessing non-additive genetic effects. Deep learning models, through multi-layered neural network structures, effectively capture complex nonlinear relationships in genomic data and automatically learn the characteristics between genotypes and phenotypes. Compared to traditional statistical models, deep learning does not require pre-defined distributional assumptions, making it particularly suitable for predicting complex crop traits. Furthermore, deep learning offers flexible modeling methods, adaptable to various application scenarios, including single-task, multi-task, multimodal, and transfer learning. Multi-task learning, in particular, simultaneously addresses multiple related tasks, enhancing the model's generalization ability through shared representation learning, thereby improving the accuracy of the constructed model. Multi-task models, by simultaneously optimizing multiple related tasks within the same model, effectively utilize shared information, improving predictive robustness and generalization ability, especially with limited data. However, existing deep learning genome prediction models lack multi-task framework designs.

[0006] Therefore, there is an urgent need to provide a method for constructing predictive models for plant phenotypes. Predictive models constructed using this method should be able to design for both additive and non-additive effects, thereby improving the accuracy of the prediction model. Furthermore, the model should also possess multi-task capabilities to further enhance its accuracy. In addition, considering that the feature dimensionality of genomic data is much higher than the sample size, an optimized feature selection method is needed that ensures comprehensive feature selection while avoiding the selection of redundant or irrelevant genetic information, thereby further improving the model's predictive power and stability. Summary of the Invention

[0007] In order to at least address the technical problems mentioned above, this disclosure proposes a technical solution for determining the authenticity of data collection behavior of power metering boxes in several aspects.

[0008] In a first aspect, this disclosure provides a method for constructing a predictive model for predicting plant phenotypes, the method comprising: acquiring whole-genome SNP loci and phenotypic values ​​of each first plant sample, and pairing them; inputting the paired whole-genome SNP loci and paired phenotypic values ​​into a feature selection module with chromosome sensing function to obtain target whole-genome SNP loci, and maintaining the pairing of the target whole-genome SNP loci with the paired phenotypic values; inputting different second plant samples into an output network module with contrastive learning function to generate feature output and low-dimensional feature output, wherein the second plant samples have paired target whole-genome SNP loci and paired phenotypic values; calculating a total loss based on the feature output, the low-dimensional feature output, and the paired phenotypic values; and constructing a predictive model by minimizing the total loss.

[0009] In a second aspect, this disclosure provides an apparatus for constructing a predictive model for predicting plant phenotypes, comprising: a processor configured to execute program instructions; and a memory configured to store the program instructions; wherein, when the program instructions are loaded and executed by the processor, the apparatus causes the apparatus to perform any of the methods described in this embodiment.

[0010] The embodiments disclosed herein, by setting a chromosome sensing function in the feature selection module, can acquire local SNP sites and global SNP sites, thereby capturing both local and global genetic effects simultaneously. Therefore, they take into account the additive, epistatic, and dominant effects of genome-wide selection genetic effects, avoiding the inaccuracies of the model caused by a single type of SNP site. By setting a contrastive learning function in the output network module, a multi-task model is constructed, which can simultaneously input two different plant samples for loss calculation and add contrastive loss, making the constructed model more accurate. Attached Figure Description

[0011] The above and other objects, features, and advantages of exemplary embodiments of this disclosure will become readily apparent upon reading the following detailed description with reference to the accompanying drawings. In the drawings, several embodiments of this disclosure are illustrated by way of example and not limitation, and like or corresponding reference numerals denote like or corresponding parts, wherein:

[0012] Figure 1 A schematic flowchart illustrating a method for constructing a predictive model for predicting plant phenotypes according to an embodiment of this disclosure is shown.

[0013] Figure 2 An overall block diagram of a method for constructing a predictive model for predicting plant phenotypes, according to an embodiment of this disclosure, is shown.

[0014] Figure 3 A schematic flowchart illustrating a method for obtaining target whole-genome SNP sites according to an embodiment of this disclosure is shown;

[0015] Figure 4 A schematic flowchart of a method for obtaining local SNP sites according to an embodiment of this disclosure is shown;

[0016] Figure 5 A schematic flowchart illustrating a method for inputting different second plant samples to an output network module according to an embodiment of this disclosure is shown.

[0017] Figure 6 A structural block diagram of an apparatus for constructing a predictive model for predicting plant phenotypes, according to an embodiment of this disclosure, is shown. Detailed Implementation

[0018] The technical solutions in the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this disclosure, not all of them. Based on the embodiments in this disclosure, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this disclosure.

[0019] It should be understood that the terms “comprising” and “including” used in this disclosure and claims indicate the presence of the described features, integrals, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or collections thereof.

[0020] It should also be understood that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the scope of this disclosure. As used in this disclosure and claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used in this disclosure and claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes such combinations.

[0021] As used in this specification and claims, the term "if" may be interpreted, depending on the context, as "when," "once," "in response to determination," or "in response to detection." Similarly, the phrase "if determined" or "if [described condition or event] is detected" may be interpreted, depending on the context, as "once determined," "in response to determination," "once [described condition or event] is detected," or "in response to detection of [described condition or event]."

[0022] The specific embodiments disclosed herein will now be described in detail with reference to the accompanying drawings.

[0023] Genome-wide selection (GS) is a technique that uses whole-genome molecular markers and phenotypic data to construct genetic models, thereby predicting the breeding value and phenotype of candidate populations. Traditional hybridization breeding is time-consuming, labor-intensive, and heavily reliant on experience. Later-developed molecular breeding can select molecular markers associated with the target trait, suitable for traits determined by a few major markers, thus limiting its application. Genome-wide selection, however, does not require marker selection, is applicable to traits determined by numerous molecular markers, and is widely used. This approach can establish genotype-phenotype association models, enabling early screening of offspring and tracking generations to update the model, achieving effects that the previous two approaches could not. This approach advances the breeding process with efficient and precise selection, its core being the construction of a reliable genomic prediction model (GP).

[0024] Genome prediction faces the "large p, small n" problem, where the number of molecular markers far exceeds the sample size. This situation easily leads to multicollinearity and overparameterization. To address these issues, many genome prediction models have been developed, mainly divided into parametric and nonparametric algorithms. The difference between the two lies in whether or not a manually pre-defined data distribution is used. Parametric algorithms require a hypothetical distribution, while nonparametric algorithms learn updated parameters through data learning, without the need for manual pre-determining. Therefore, they have a wider range of applications and can build better models.

[0025] This disclosure discloses a method for constructing a predictive model for predicting plant phenotypes. This method can construct a deep learning-based GP model that does not require manual parameter pre-determination for predicting plant phenotypes. Figure 1 A schematic flowchart illustrating a method for constructing a predictive model for predicting plant phenotypes according to an embodiment of this disclosure is shown.

[0026] like Figure 1 As shown, a method for constructing a predictive model for predicting plant phenotypes includes: obtaining whole-genome SNP loci and phenotypic values ​​of each first plant sample and pairing them; inputting the paired whole-genome SNP loci and paired phenotypic values ​​into a feature selection module with chromosome sensing function to obtain target whole-genome SNP loci and maintain the pairing between target whole-genome SNP loci and paired phenotypic values; inputting different second plant samples into an output network module with contrastive learning function to generate feature output and low-dimensional feature output, wherein the second plant samples have paired target whole-genome SNP loci and paired phenotypic values; calculating the total loss based on the feature output, low-dimensional feature output, and paired phenotypic values; and constructing a predictive model by minimizing the total loss.

[0027] In summary, the method for constructing a predictive model for predicting plant phenotypes according to the present disclosure includes five main steps: data acquisition and pairing, target feature extraction, feature generation and output, total loss calculation, and model optimization and construction.

[0028] Data acquisition and pairing: The whole genome SNP loci and corresponding phenotypic values ​​of each first plant sample were obtained from the breeding population, and the whole genome SNP loci and phenotypic values ​​were paired.

[0029] Target Feature Extraction: Paired whole-genome SNP loci and phenotypic values ​​are input into a feature selection module with chromosome sensing capabilities to screen and extract target whole-genome SNP loci. After extraction, the pairing relationship between target whole-genome SNP loci and phenotypic values ​​is maintained.

[0030] Feature generation and output: Different second plant samples containing target whole-genome SNP loci and phenotypic values ​​are input into an output network module with contrast learning function to generate feature output and low-dimensional feature output.

[0031] Total loss calculation: The total loss is calculated based on the feature output, low-dimensional feature output, and paired phenotypic values. The total loss includes prediction error (i.e., the loss in the traditional model) and contrastive learning loss (i.e., the loss calculated through the contrastive learning function of the output network module in this disclosed embodiment).

[0032] Model optimization and construction: The model parameters are optimized by minimizing the total loss, and a deep learning prediction model for plant phenotypic prediction is finally constructed.

[0033] Specifically, the method 100 for constructing a predictive model for predicting plant phenotypes includes the following steps:

[0034] Step S110: Obtain the whole genome SNP loci and phenotypic values ​​of each first plant sample and perform pairing.

[0035] The first plant sample is a plant individual sample within the breeding population, containing a complete genome information of that individual plant and the corresponding phenotypic values. The breeding population can be understood as a cluster of individuals in a specific environment, sharing common kinship and varietal characteristics. Through in-depth research and breeding work within the breeding population, improved varieties with specific agricultural value can be developed. The complete genome information of the first plant sample can be used to extract whole-genome SNP loci according to specified partitioning rules, representing genetic variation information at specific locations in the whole genome. Furthermore, each first plant sample also possesses a specific phenotype, which can be quantitatively represented using phenotypic values. Therefore, the complete genome SNP loci and phenotypic values ​​of each first plant sample within the breeding population can be obtained. These two are then paired for input in step S120.

[0036] Step S120: Input the paired whole-genome SNP loci and paired phenotypic values ​​into the feature selection module with chromosome sensing function to obtain the target whole-genome SNP loci and maintain the pairing between the target whole-genome SNP loci and the paired phenotypic values.

[0037] Figure 2 An overall block diagram of a method for constructing a predictive model for predicting plant phenotypes, according to an embodiment of this disclosure, is shown. Figure 2 As shown, the overall structure includes a feature selection module (i.e. Figure 2 The part indicated by label 1) and the output network module (i.e. Figure 2 The part indicated by label 2 in the diagram has a sequential relationship. The paired whole-genome SNP loci and paired phenotypic values ​​are input into the feature selection module as the initial input. Here, "paired phenotypic value" represents the phenotypic value that has been paired. In step S120, because the feature selection module has chromosome sensing capabilities, the whole-genome SNP loci can undergo certain changes and screenings, such as grouping SNP loci by chromosomal intervals and then screening them to become target whole-genome SNP loci. At this point, because the whole-genome SNP loci and phenotypic values ​​have been paired, after being transformed into target whole-genome SNP loci, the target whole-genome SNP loci can continue to maintain their pairing with the phenotypic value. The plant sample, i.e., the second plant sample, can be determined using the target whole-genome SNP loci and paired phenotypic values.

[0038] Step S130: Input different second plant samples into the output network module with contrast learning function to generate feature output and low-dimensional feature output.

[0039] Step S120 yields a second plant sample containing paired target whole-genome SNP loci and paired phenotypic values. At this stage, the second plant sample is not directly input into the output network module. Because the output network module has a contrastive learning function, it can simultaneously receive multiple different second plant samples as input for a single round. The output network module can include multiple branches to output various types of results, such as results with different dimensions; in this embodiment, this includes feature output and low-dimensional feature output. During the construction of the prediction model, the output of the output network module can be used to optimize the prediction model.

[0040] The feature output includes the predicted phenotypic value of the second plant sample, which is the phenotypic value calculated by the model. In contrast to the predicted phenotypic value are the paired phenotypic values. These paired phenotypic values ​​are obtained from the plant sample and can be paired with SNP loci.

[0041] The low-dimensional feature output includes a low-dimensional representation of the target whole-genome SNP sites from the second plant sample.

[0042] Step S140: Calculate the total loss based on the feature output, low-dimensional feature output, and paired phenotypic values.

[0043] By obtaining the feature output and low-dimensional feature output, as well as the paired phenotypic values ​​of multiple different second plant samples through the above steps, the total loss can be calculated using a loss function. This total loss is the difference between the predicted value and the true value; the smaller the difference, the more accurate the prediction model.

[0044] Step S150: Construct a prediction model by minimizing the total loss.

[0045] The model's parameters may change after each round. By comparing the total loss under different parameters, the parameter with the minimum total loss is found, and the model at this point is the constructed prediction model.

[0046] Figure 3 A schematic flowchart illustrating a method for obtaining target whole-genome SNP sites according to an embodiment of this disclosure is shown.

[0047] like Figure 3 As shown, in some embodiments, paired whole-genome SNP sites and paired phenotypic values ​​are input to a feature selection module with chromosome sensing function to obtain target whole-genome SNP sites, including: obtaining target whole-genome SNP sites that interpret local information, i.e., local SNP sites; obtaining target whole-genome SNP sites that interpret global information, i.e., global SNP sites; and sequentially splicing local SNP sites and global SNP sites to generate a representation of target whole-genome SNP sites.

[0048] Specifically, method 200 for obtaining target whole-genome SNP sites includes the following steps:

[0049] Step S210: Obtain the target whole-genome SNP sites for interpreting local information.

[0050] The target genome-wide SNP loci consist of two parts: local SNP loci and global SNP loci. In this step, the genome-wide SNP loci and paired phenotypic values ​​of the input first plant sample are processed to obtain the target genome-wide SNP loci that interpret local information, i.e., the local SNP loci.

[0051] Local SNP loci are used to capture key genetic variation information in specific regions of the genome, focusing on describing the direct effects of SNP loci within these regions on phenotypic changes. The acquisition of local SNP loci is typically based on chromosomal regions or gene functional regions. The entire genome is divided into several groups according to chromosomes or functional units, and SNP loci in each group are screened using feature selection methods to retain SNP loci that contribute significantly to phenotypic changes. This approach reduces interference from redundant features, thereby improving the accuracy and computational efficiency of phenotypic prediction.

[0052] The role of local SNP sites is mainly reflected in the following aspects:

[0053] (i) Capturing local genetic effects: By identifying SNP sites that are directly related to the phenotype, local SNP sites reflect the mechanism of action of local genetic variations.

[0054] (ii) Improve the generalization ability of the model: noise and irrelevant SNP sites are removed, making the model more focused on the core genetic information related to the phenotype.

[0055] (iii) Reduce computational complexity: By filtering and reducing the dimensionality of features, the dimensionality of the model input data is significantly reduced, thereby improving computational efficiency.

[0056] In this step, local SNP sites and global SNP sites combine to form target genome-wide SNP sites, which are then used together for subsequent model training to maximize the use of local and global genetic variation information and improve phenotypic prediction performance.

[0057] Step S220: Obtain the target whole-genome SNP sites for interpreting global information.

[0058] In this step, the whole genome SNP sites and paired phenotypic values ​​of the first plant sample are processed to obtain the target whole genome SNP sites that interpret global information, namely global SNP sites.

[0059] Global SNPs represent SNPs with strong explanatory power for phenotypic changes using all genome-wide SNPs. They represent the direct interaction between all genome-wide SNPs in the first plant sample and paired phenotypic values. Global SNPs are selected from the entire genome to identify SNPs that significantly explain phenotypic changes, reflecting the overall association between genetic variation and phenotype across the entire genome. By assessing the contribution of genome-wide SNPs to paired phenotypic values, global SNPs retain key SNPs that significantly influence phenotypic changes at the entire genome scale, while removing noisy or irrelevant SNPs.

[0060] The role of global SNP sites is reflected in the following aspects:

[0061] (i) Capturing global genetic effects: Global SNP sites cover the entire genome, integrating information from different chromosomes and gene functional regions, which helps to reflect the global characteristics of phenotypic changes.

[0062] (ii) Enhancing the comprehensiveness of the model: By integrating whole-genome information, global SNP sites provide the model with broader genetic background support, avoiding over-reliance on local effects of certain specific regions.

[0063] (III) Optimize model performance: Screening out key SNP sites across the entire genome reduces noise from redundant features while preserving the comprehensive contribution of genome-wide information to phenotypic prediction, thereby improving the model's generalization ability and prediction accuracy.

[0064] Step S230: Sequentially splice the local SNP sites and global SNP sites to generate SNP sites representing the target whole genome.

[0065] After obtaining local target genome-wide SNP sites and global SNP sites, these SNP sites need to be spliced ​​into a plant sample. This step is also called SNP splicing. After splicing, a second plant sample containing the target genome-wide SNP sites is generated. Through the above operations, SNP splicing effectively connects local and global information, providing an efficient and high-quality feature foundation for phenotypic prediction. This process is not only a core step in data preprocessing, but also an important guarantee for the efficient modeling of deep learning models. The purpose of this step is to construct an optimized data input format for downstream models to perform efficient learning and prediction.

[0066] Main functions and roles of SNP splicing:

[0067] (i) Feature integration: Local SNP loci capture local genetic effects, while global SNP loci capture overall genetic background information. By splicing, the two are combined to form the target genome-wide SNP loci, giving the data stronger interpretability at both the local and global levels.

[0068] (ii) Improving data quality: By selecting and retaining key SNPs with high explanatory power and removing noisy or redundant SNP sites, the representativeness of the samples is enhanced. This improvement not only optimizes the quality of the sample data but also reduces the errors caused by irrelevant input features in the model.

[0069] (III) Reducing Data Dimensionality: Traditional genomic data has extremely high dimensionality, and directly inputting it into the model may lead to excessive computational complexity and the risk of overfitting. By selecting and splicing SNP sites, the data dimensionality is significantly reduced, thereby improving computational efficiency and reducing the burden on the model.

[0070] (iv) Maintaining phenotypic consistency: Although the SNP data were filtered and spliced, the phenotypic value (target variable) of each sample remained unchanged. This consistency ensures that the biological significance and statistical consistency of the samples are not compromised while the data features are optimized.

[0071] (v) Providing optimized features for model input: The second plant sample generated by splicing not only has highly representative target genome-wide SNP sites, but also maintains the accuracy of phenotype. This sample provides optimized feature input for subsequent deep learning models, which helps to improve the accuracy and reliability of phenotype prediction.

[0072] Figure 4 A schematic flowchart of a method for obtaining local SNP sites according to an embodiment of this disclosure is shown.

[0073] like Figure 4 As shown, in some embodiments, obtaining the target whole-genome SNP site for interpreting local information, i.e., the local SNP site, includes: grouping whole-genome SNP sites by chromosomal intervals; constructing a feature selection model and performing feature selection on the paired phenotypic values ​​of the grouped whole-genome SNP sites and paired with the whole-genome SNP sites respectively, to obtain local SNP site groups; and splicing the local SNP sites of each local SNP site group to obtain the local SNP site.

[0074] Obtaining local SNP sites mainly involves three steps:

[0075] (i) The input whole genome SNP sites are divided into several chromosomal intervals according to the chromosome structure or specified rules to form multiple groups, each group containing SNP sites belonging to the corresponding chromosomal interval.

[0076] (ii) For each chromosomal region after grouping, a feature selection model is constructed based on the paired phenotypic values ​​of the whole genome SNP loci. This model is used to screen SNP loci within the group to obtain local SNP loci groups with strong phenotypic explanatory power.

[0077] (iii) The selected local SNP loci groups within each chromosome interval are spliced ​​together to form a complete set of local SNP loci, and the local SNP loci are obtained.

[0078] Specifically, the method 300 for obtaining local SNP sites includes the following steps:

[0079] Step S310: Group the whole genome SNP sites by chromosomal region.

[0080] The whole genome SNP loci of the first plant sample are grouped according to chromosomal intervals. For example, the whole genome can be divided into several groups according to chromosomes or functional units based on chromosomal regions or gene functional regions, with each group of SNP loci having a specific chromosomal interval.

[0081] Step S320: The grouped whole-genome SNP loci are paired with the paired phenotypic values ​​of the whole-genome SNP loci, a feature selection model is constructed and feature selection is performed to obtain local SNP locus groups.

[0082] Since a plant sample has a single phenotypic value, and whole-genome SNP loci are paired with phenotypic values, the whole-genome SNP loci of the same first plant sample after grouping also have the same phenotypic values. Using the grouped whole-genome SNP loci and paired phenotypic values, a feature selection model can be constructed and features can be selected. The aim is to obtain the importance values ​​of the SNP nodes in each group to the paired phenotypic values, and to select certain groups with importance values ​​greater than a certain threshold, i.e., local SNP locus groups.

[0083] In this embodiment, feature selection for the feature model can be performed using the LightGBM algorithm, a tree-based method that calculates the non-linear relationships between SNP sites, aiding in the identification of key sites. This algorithm obtains the information gain (IG) of the SNP site and its paired phenotypic value (equivalent to the importance value mentioned earlier). In this embodiment, SNP sites with positive information gain (IG > 0) are retained, while those with IG = 0 are discarded; that is, SNP sites with a positive correlation to the paired phenotypic value are retained.

[0084] Step S330: The local SNP sites of each local SNP site group are spliced ​​together to obtain local SNP sites.

[0085] For the preserved local SNP locus groups, the local SNP loci can be spliced ​​together to obtain local SNP loci.

[0086] In some embodiments, obtaining the target genome-wide SNP site for interpreting global information, i.e., the global SNP site, includes: constructing a feature selection model and performing feature selection by pairing the genome-wide SNP site with the paired phenotypic values, thereby obtaining the global SNP site.

[0087] Specifically, the whole genome SNP sites and paired phenotypic values ​​of the first plant sample can be used to construct a feature selection model and perform feature selection. The purpose is to obtain the importance values ​​of whole genome SNP nodes to paired phenotypic values ​​and select certain groups with importance values ​​greater than a certain threshold, i.e., global SNP sites.

[0088] Similar to the previous example, in this embodiment, feature selection of the feature model can be performed using the LightGBM algorithm, retaining SNP sites with positive information gain (IG > 0) and discarding SNP sites with IG = 0, that is, retaining SNP sites that have a positive correlation with the paired phenotypic value.

[0089] like Figure 2 As shown, in some embodiments, the output network module includes a first branch and a second branch, which are merged by addition.

[0090] Specifically, the output network module includes two branches: a first branch and a second branch. The input to both branches is the same: the second plant sample. In this embodiment, the output dimensions of both the first and second branches are 1024-dimensional; therefore, the first and second branches can be merged by addition.

[0091] like Figure 2 As shown, in some embodiments, the first branch includes three nonlinear fully connected layers, and the second branch includes one linear fully connected layer.

[0092] Specifically, the first branch may contain three nonlinear fully connected layers. In this embodiment, the activation functions of these three nonlinear fully connected layers are all ReLU activation functions, and the dimensions of these three nonlinear fully connected layers decrease proportionally, to 4096, 2048, and 1024 dimensions, respectively. The second main branch may contain one linear fully connected layer. In this embodiment, the dimension of this linear fully connected layer is 1024.

[0093] The genetic effects of GS mainly include additive effects, epistatic effects, and dominant effects. Additive effects can be interpreted as linear relationships between loci, epistatic effects refer to non-linear relationships between loci, and dominant effects are the influence of different alleles at the same locus; the latter two are called non-additive effects. Traditional genome prediction models mostly consider only additive effects when estimating genetic effects, but these are very important for traits closely related to fitness and traits with low heritability. Therefore, non-additive effects should be considered when constructing genome prediction models. Through the two branches mentioned above, the additive, epistatic, and dominant effects of GS genetic effects can be reflected simultaneously. Compared with other models, this model can reflect non-additive effects due to the effect of the first branch.

[0094] Figure 5 A schematic flowchart illustrating a method for inputting different second plant samples to an output network module according to an embodiment of this disclosure is shown.

[0095] like Figure 5 As shown, in some embodiments, different second plant samples are input into the output network module with contrast learning function, including: when sample A, which belongs to the second plant sample, is input into the output network module, another sample B is randomly matched from the remaining second plant samples; sample A and sample B, the target whole genome SNP sites of the two second plant samples and the paired phenotypic values ​​are jointly input into the output network module.

[0096] Specifically, the output network module in this disclosure has a comparative learning function, that is, by simultaneously inputting multiple plant samples, the constructed model has better predictive ability. In this embodiment, the number of plant samples input simultaneously is two.

[0097] This step ensures that in each training epoch, the input sample A is not the same sample as sample B, but rather a different sample randomly selected from the remaining samples. This random selection helps enhance the model's generalization ability, avoids overfitting, and promotes the learning of differences between samples.

[0098] The method 400 for inputting different second plant samples into the output network module includes the following steps:

[0099] Step S410: When a sample A belonging to the second plant sample is input into the output network module, another sample B is randomly matched from the remaining second plant samples.

[0100] When inputting a second plant sample into the output network module, several second plant samples can be calculated first. Then, one of these second plant samples, namely sample A, is selected as one of the input plant samples. Next, from the remaining calculated second plant samples, another sample B is randomly matched, and sample A and sample B are not identical.

[0101] Step S420: Input the target whole genome SNP sites and paired phenotypic values ​​of the two second plant samples, A and B, into the output network module.

[0102] After obtaining samples A and B, the target whole genome SNP sites and paired phenotypic values ​​of these two plant samples are combined and input into the output network module.

[0103] like Figure 2 As shown, in some embodiments, the output network module further includes a third branch and a fourth branch, wherein the third branch is used to calculate the predicted phenotypic value of the total loss, and the fourth branch is used to calculate the low-dimensional representation of the target whole-genome SNP site of the total loss.

[0104] Specifically, the output network module further includes a third branch and a fourth branch. These two branches are located after the sum of the first and second branches. The third branch may include a Flatten layer, two non-linear fully connected layers, and an Output layer, used to calculate the predicted phenotypic value of the total loss, with dimensions of 4096, 2048, and 1 dimension for each layer, respectively; while the fourth branch may include a Flatten layer, two non-linear fully connected layers, and an Enbedding layer, used to calculate the low-dimensional representation of the target genome-wide SNP sites for the total loss, with dimensions of 4096, 2048, and 1024 dimensions for each layer, respectively.

[0105] In some embodiments, the method for calculating the total loss includes: calculating a first principal loss using the predicted phenotypic value of sample A and the paired phenotypic value of sample A; calculating a second principal loss using the predicted phenotypic value of sample B and the paired phenotypic value of sample B; calculating a contrastive loss using the low-dimensional representation of the target genome-wide SNP site of sample A, the low-dimensional representation of the target genome-wide SNP site of sample B, and the difference between the paired phenotypic value of sample A and the paired phenotypic value of sample B; and summing the first principal loss, the second principal loss, and the contrastive loss to obtain the total loss.

[0106] Specifically, the loss function used for optimization in this disclosure embodiment includes the main loss and the contrast loss. The main loss is a common algorithm that uses the predicted phenotypic value and the paired phenotypic value of the plant sample for calculation. That is, for a certain plant sample, the main loss is the mean squared error (MSE) of the predicted phenotypic value and the paired phenotypic value.

[0107] The main loss function is shown below:

[0108]

[0109] Where LossMSE represents the main loss for the plant samples, N represents the number of plant samples, and y i This represents the paired phenotypic value of the i-th plant sample. This represents the predicted phenotypic value of the i-th plant sample.

[0110] In this disclosed embodiment, since the output network module has a contrastive learning function, an additional contrastive loss function is added when calculating the total loss. The low-dimensional feature output includes the low-dimensional representation of the target whole-genome SNP loci of the second plant sample. Therefore, the low-dimensional representations of the target whole-genome SNP loci of the two plant samples represent the corresponding low-dimensional feature outputs. The low-dimensional representations of the target whole-genome SNP loci of the two plant samples, and the difference between the paired phenotypic values ​​of the target whole-genome SNP loci of the two plant samples are calculated to obtain the contrastive loss function for the output network module with contrastive learning function in this disclosed embodiment.

[0111] The contrast loss function is shown below:

[0112]

[0113] Among them, Loss contrastive This represents the contrast loss between a pair of plant samples, where N is the number of plant samples. The diff represents the Euclidean distance between the i-th pair of embeddings. i This represents the phenotypic difference in the i-th pair of plant samples.

[0114] Therefore, the total loss can be simply expressed as:

[0115] Total loss = First principal loss + Second principal loss + Comparative loss

[0116] First primary loss = Loss MSE (Predicted phenotypic value A, paired phenotypic value A)

[0117] Second primary loss = Loss MSE (Predicted phenotypic value B, paired phenotypic value B)

[0118] Comparative loss = Loss contrastive (Low-dimensional representation A, low-dimensional representation B, paired phenotypic difference)

[0119] Wherein, the paired phenotypic value of sample A is the paired phenotypic value A, the predicted phenotypic value is the predicted phenotypic value A, and the low-dimensional representation is the low-dimensional representation A; the paired phenotypic value of sample B is the paired phenotypic value B, the predicted phenotypic value is the predicted phenotypic value B, and the low-dimensional representation is the low-dimensional representation B; the paired phenotypic value difference is the difference between the paired phenotypic value of sample A and the paired phenotypic value of sample B.

[0120] Predicted phenotypic values ​​are the phenotypic values ​​calculated by the model. During model building, predicted phenotypic values ​​represent the phenotypic values ​​calculated by the model under the current parameters. In contrast, paired phenotypic values ​​are the paired phenotypic values. These phenotypic values ​​can be directly obtained from real plant samples and can be paired with SNP loci. The smaller the difference between predicted and paired phenotypic values, the closer the model is to the observed situation, and the higher the accuracy of the constructed model.

[0121] After calculating the total loss, the model parameters are optimized by minimizing the total loss, and finally a deep learning prediction model for plant phenotypic prediction is constructed.

[0122] Figure 6 A structural block diagram of an apparatus for constructing a predictive model for predicting plant phenotypes, according to an embodiment of this disclosure, is shown.

[0123] like Figure 6 As shown, an apparatus for constructing a predictive model for predicting plant phenotypes includes: a processor configured to execute program instructions; and a memory configured to store the program instructions; wherein, when the program instructions are loaded and executed by the processor, the apparatus causes the apparatus to perform any of the methods described in this embodiment.

[0124] Specifically, the device 500 can be implemented as various types of devices, including but not limited to mainframes, personal computers (PCs), and mobile devices. The device 500 includes a processor 501 and a memory 502. The processor 501 controls the operation of the device 500 by executing programs stored in the memory 502. The processor 501 can be implemented using, but is not limited to, a central processing unit (CPU), a graphics processing unit (GPU), an application processor (AP), and an artificial intelligence processor chip (IPU). The memory 502 can be used to store various data and instructions processed in the device 500, including the method for constructing a predictive model for predicting plant phenotypes in this disclosure embodiment. The memory 502 can include at least one of volatile memory and non-volatile memory. Non-volatile memory can include read-only memory (ROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), etc. Volatile memory can include dynamic RAM (DRAM), static RAM (SRAM), synchronous DRAM (SDRAM), etc. In addition, the memory 502 may include at least one of a hard disk drive (HDD), a solid-state drive (SSD), a high-density flash memory (CF), a secure digital card (SD), a micro-secure digital card (Micro-SD), a mini-secure digital card (Mini-SD), an extreme digital card (xD), a cache, or a memory stick.

[0125] While numerous embodiments of this disclosure have been shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. Many modifications, alterations, and alternatives will occur to those skilled in the art without departing from the spirit and intent of this disclosure. It should be understood that various alternatives to the embodiments of this disclosure described herein may be employed in the practice of this disclosure. The appended claims are intended to define the scope of this disclosure and therefore cover equivalents or alternatives within the scope of these claims.

Claims

1. A method for constructing a predictive model for predicting plant phenotypes, characterized in that, The method includes: Obtain the whole genome SNP loci and phenotypic values ​​of each first plant sample, and then pair them. Input the paired whole-genome SNP loci and paired phenotypic values ​​into the feature selection module with chromosome sensing function to obtain the target whole-genome SNP loci and maintain the pairing of the target whole-genome SNP loci with the paired phenotypic values; Different second plant samples are input into an output network module with contrastive learning capabilities to generate feature output and low-dimensional feature output. The second plant sample has paired target whole-genome SNP loci and paired phenotypic values. The feature output includes the predicted phenotypic values ​​of the second plant sample. The low-dimensional feature output includes a low-dimensional representation of the target whole-genome SNP loci of the second plant sample. The total loss is calculated based on the feature output, the low-dimensional feature output, and the paired phenotypic values. A prediction model is constructed by minimizing the total loss. Input the paired whole-genome SNP loci and the paired phenotypic values ​​into the feature selection module with chromosome sensing function to obtain the target whole-genome SNP loci, including: The target genome-wide SNP sites for interpreting local information are obtained, i.e., local SNP sites; The target genome-wide SNP sites for interpreting global information are obtained, i.e., global SNP sites; The local SNP sites and the global SNP sites are sequentially spliced ​​together to generate SNP sites representing the target whole genome. The target genome-wide SNP sites for obtaining interpretation of local information, i.e., local SNP sites, include: The SNP sites across the entire genome were grouped by chromosomal region; The grouped whole-genome SNP loci are paired with the paired phenotypic values ​​of the whole-genome SNP loci, and a feature selection model is constructed and feature selection is performed to obtain local SNP locus groups. The local SNP sites of each of the local SNP site groups are spliced ​​together to obtain the local SNP sites; The target genome-wide SNP sites for obtaining global information, i.e., global SNP sites, include: The whole genome SNP sites are paired with the paired phenotypic values ​​to construct a feature selection model and perform feature selection to obtain the global SNP sites; The output network module includes a first branch and a second branch, which are merged by addition. The first branch contains three nonlinear fully connected layers, and the second branch contains one linear fully connected layer.

2. The method according to claim 1, characterized in that, Different second plant samples are input into the output network module with contrast learning capabilities, including: When a sample A belonging to the second plant sample is input into the output network module, another sample B is randomly matched from the remaining second plant samples; The target whole-genome SNP loci and paired phenotypic values ​​of the two second plant samples, namely, sample A and sample B, are input together into the output network module.

3. The method according to claim 2, characterized in that, The output network module further includes a third branch and a fourth branch, wherein the third branch is used to calculate the predicted phenotypic value of the total loss, and the fourth branch is used to calculate the low-dimensional representation of the target whole-genome SNP site of the total loss.

4. The method according to claim 3, characterized in that, The method for calculating the total loss includes: The first principal loss is calculated using the predicted phenotypic value of sample A and the paired phenotypic value of sample A. The second principal loss is calculated using the predicted phenotypic value of sample B and the paired phenotypic value of sample B. The contrast loss is calculated using the low-dimensional representation of the target genome-wide SNP locus in sample A, the low-dimensional representation of the target genome-wide SNP locus in sample B, and the difference between the paired phenotypic values ​​of sample A and sample B. The first principal loss, the second principal loss, and the comparison loss are summed to obtain the total loss.

5. An apparatus for constructing a predictive model for predicting plant phenotypes, comprising: A processor, configured to execute program instructions; as well as A memory configured to store the program instructions; The characteristic is that, when the program instructions are loaded and executed by the processor, the device causes the device to perform the method according to any one of claims 1-4.