Method and device for constructing prediction model for predicting plant phenotype
By setting up chromosome perception and contrast learning functions in the feature selection module and the output network module, a multi-task genome prediction model that can capture additive and non-additive genetic effects was built, solving the problem of insufficient model accuracy and generalization capabilities in the prior art.
Patent Information
- Application Number
- CN202411943779.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-26
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2044-12-26
AI Technical Summary
The prior art lacks the design of non-additive effects when constructing genomic prediction models and lacks a multi-task framework, resulting in insufficient accuracy and generalization capabilities of the prediction models.
By setting chromosome perception functions in the feature selection module, local and global SNP sites are obtained to capture additive and non-additive genetic effects; setting up comparison learning functions in the output network module to build a multi-task model to enhance the robustness and generalization capabilities of the model.
It improves the accuracy and stability of the prediction model, can capture both additive and non-additive genetic effects at the same time, and enhances the multi-task ability and data adaptability of the model.
Smart Images

Figure CN120048340A_ABST
Abstract
Description
Technical Field
[0001] The present invention generally relates to the technical field of constructing prediction models. More specifically, the present invention relates to a method for constructing a prediction model for predicting plant phenotypes. Background Art
[0002] Traditional crop breeding is based on phenotypic selection, and excellent offspring are selected by observing the variation of crop phenotypes. Although breeders can use means such as the general laws of biological inheritance, comprehensive selection indices, comparison of contemporary populations, and field trial statistics for field trial design and selection, the success of their breeding depends on the experience of breeders and the efficiency is low. At present, the application of molecular marker-assisted selection breeding technology has become increasingly mature, but it is only applicable to traits determined by a small number of major QTLs (Quantitative Trait Loci). The actual crop breeding work requires the coordinated improvement of multiple traits. There are hundreds or thousands of materials available for breeders in breeding projects, and there are even more mating combinations. However, due to the limitation of the test scale, a large number of important materials have not been tested, which restricts the breeding process. Both of the above methods have great limitations.
[0003] The Genomic Selection (GS) method uses molecular markers covering the entire genome and phenotypic data of samples to establish a prediction model to achieve early genetic evaluation of individuals. The GS breeding technology does not require the identification of SNP loci (Single Nucleotide Polymorphism) significantly related to the target trait. Even the genetic effects of single loci with small effects causing phenotypic variation can be captured by high-density genetic markers, and the breeding value of an individual can be evaluated through its genotype, greatly improving breeding accuracy, shortening the breeding cycle, increasing breeding efficiency, and achieving a leap from empirical breeding to precision breeding. It has become a cutting-edge technology in animal and plant breeding.
[0004] The genetic effects of GS mainly include additive effects and non-additive effects. Among them, the additive effect can be interpreted as the linear relationship between loci, and the non-additive effects include epistatic effects and dominance effects. The epistatic effect refers to the non-linear relationship between loci, and the dominance effect is the interaction between alleles at the same locus. Most traditional genomic prediction models only consider additive effects when estimating genetic effects and lack the design of non-additive effects. Therefore, the factors of non-additive effects should be considered when constructing a genomic prediction model.
[0005] Deep learning methods are an important approach to achieving non-additive genetic effect evaluation. Deep learning models can effectively capture complex non-linear relationships in genomic data through a multi-layer neural network structure and can automatically learn the features between genotypes and phenotypes. Compared with traditional statistical models, deep learning does not require pre-set distribution assumptions and is particularly suitable for predicting complex crop traits. In addition, deep learning has flexible modeling methods and can adapt to various application scenarios, including single-task, multi-task, multi-modal, and transfer learning. Among them, multi-task learning simultaneously solves multiple related tasks and improves the generalization ability of the model through shared representation learning, thereby improving the accuracy of the component model. By optimizing multiple related tasks simultaneously in the same model, multi-task models can effectively utilize shared information, thereby improving the robustness and generalization ability of prediction, especially having obvious advantages in the case of limited data volume. Existing deep learning genomic prediction models lack the design of a multi-task framework.
[0006] In view of this, there is an urgent need to provide a method for constructing a prediction model for predicting plant phenotypes. The prediction model for predicting plant phenotypes constructed by this method can not only design additive effects but also non-additive effects, thereby improving the accuracy of the prediction model. At the same time, the model also needs to have the ability of multi-tasking to further improve the accuracy of the prediction model. In addition, in view of the characteristic that the genomic data feature dimension is much higher than the number of samples, a method for optimizing feature selection is also required, which can not only ensure the comprehensiveness of feature selection but also avoid selecting redundant or irrelevant genetic information, thereby further improving the prediction ability and stability of the model. Summary of the Invention
[0007] In order to at least solve the above-mentioned technical problems, the present disclosure proposes a technical solution for a method for determining the authenticity of data collection behavior of an electric power metering box in multiple aspects.
[0008] In a first aspect, the present disclosure provides a method for constructing a prediction model for predicting plant phenotypes, the method comprising: obtaining the whole-genome SNP loci of each first plant sample and the phenotypic value of each first plant sample and pairing them; inputting the paired whole-genome SNP loci and paired phenotypic values into a feature selection module with chromosome perception function to obtain target whole-genome SNP loci and maintaining the pairing of the target whole-genome SNP loci and the paired phenotypic values; inputting different second plant samples into an output network module with contrast learning function for generating a feature output and a low-dimensional feature output, the second plant samples having the paired target whole-genome SNP loci and the paired phenotypic values; calculating a total loss based on the feature output, the low-dimensional feature output, and the paired phenotypic values; and constructing a prediction model by minimizing the total loss.
[0009] In a second aspect, the present disclosure provides an apparatus for constructing a prediction model for predicting plant phenotypes, including: a processor configured to execute program instructions; and a memory configured to store the program instructions; wherein, when the program instructions are loaded and executed by the processor, the apparatus is caused to execute the method according to any one of the embodiments in the present embodiment.
[0010] In the embodiments of the present disclosure, by setting a chromosome perception function in the feature selection module, local SNP sites and global SNP sites can be obtained, and local genetic effects and global genetic effects can be captured simultaneously. Therefore, the additive effects, epistatic effects, and dominant effects of the genome-wide selection genetic effects are considered, and the inaccuracy of the model caused by a single type of SNP site is avoided; by setting a contrastive learning function in the output network module, a multi-task model is constructed, so that two different plant samples can be input simultaneously for loss calculation, and the contrastive loss is increased, making the constructed model have higher accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] By reading the following detailed description with reference to the accompanying drawings, the above and other objects, features, and advantages of the exemplary embodiments of the present disclosure will become readily understood. In the drawings, several embodiments of the present disclosure are shown in an exemplary and non-limiting manner, and the same or corresponding reference numerals represent the same or corresponding parts, wherein:
[0012] Figure 1 A schematic flowchart of a method for constructing a prediction model for predicting plant phenotypes according to an embodiment of the present disclosure is shown;
[0013] Figure 2 An overall block diagram of a method for constructing a prediction model for predicting plant phenotypes according to an embodiment of the present disclosure is shown;
[0014] Figure 3 A schematic flowchart of a method for obtaining target genome-wide SNP sites according to an embodiment of the present disclosure is shown;
[0015] Figure 4 A schematic flowchart of a method for obtaining local SNP sites according to an embodiment of the present disclosure is shown;
[0016] Figure 5 A schematic flowchart of a method for inputting different second plant samples into the output network module according to an embodiment of the present disclosure is shown;
[0017] Figure 6 A structural block diagram of an apparatus for constructing a prediction model for predicting plant phenotypes according to an embodiment of the present disclosure is shown. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0018] The technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present disclosure. Obviously, the described embodiments are some, but not all, of the embodiments of the present disclosure. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present disclosure without creative efforts shall fall within the scope of protection of the present disclosure.
[0019] It should be understood that the terms "comprising" and "including" used in the specification and claims of the present disclosure indicate the presence of the described features, wholes, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.
[0020] It should also be understood that the terms used in the specification of the present disclosure are only for the purpose of describing specific embodiments and are not intended to limit the present disclosure. As used in the specification and claims of the present disclosure, unless the context clearly indicates otherwise, the singular forms "a", "an", and "the" are intended to include the plural forms. It should be further understood that the term "and / or" used in the specification and claims of the present disclosure refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0021] As used in this specification and the claims, the term "if" can be interpreted as "when", "once", "in response to determining", or "in response to detecting" depending on the context. Similarly, the phrase "if determined" or "if [the described condition or event] is detected" can be interpreted as meaning "once determined", "in response to determining", "once [the described condition or event] is detected", or "in response to detecting [the described condition or event]" depending on the context.
[0022] The specific implementation manners of the present disclosure will be described in detail below with reference to the accompanying drawings.
[0023] Genomic selection (GS) is a technology that uses whole-genome molecular markers and phenotypic data to construct a genetic model for predicting the breeding values and phenotypes of candidate populations. Traditional crossbreeding is time-consuming and laborious and highly dependent on experience. The later-developed molecular breeding can select molecular markers related to target traits and is applicable to traits determined by major and a small number of molecular markers. Therefore, this solution has limitations. Genomic selection, without the need to select molecular markers, is applicable to traits determined by numerous molecular markers and is widely used. This solution can establish a genotype-phenotype association model, enabling early screening of offspring and tracking generations to update the model, achieving effects that the first two solutions cannot achieve. This solution promotes the breeding process with efficient and accurate selection, and its core lies in constructing a reliable genomic prediction model (Genomic Prediction, i.e., GP).
[0024] Genomic prediction faces the "large p, small n" problem, that is, the number of molecular markers is much larger than the sample size. This situation easily leads to multicollinearity and overparameterization. To solve these problems, many genomic prediction models have been developed, mainly divided into parametric and non-parametric algorithms. The difference between the two lies in whether there is an artificially preset data distribution. Parametric algorithms require giving a hypothesized distribution, while non-parametric algorithms obtain updated parameters through data learning without artificial predetermination. Therefore, their application scenarios are more extensive and the constructed models are better.
[0025] This disclosure discloses a method for constructing a prediction model for predicting plant phenotypes. This method can construct a GP model based on deep learning that does not require artificially preset parameters for predicting plant phenotypes. Figure 1 A schematic flowchart of the method for constructing a prediction model for predicting plant phenotypes according to the embodiments of this disclosure is shown.
[0026] As Figure 1 shown, a method for constructing a prediction model for predicting plant phenotypes includes: obtaining the whole-genome SNP loci of each first plant sample and the phenotypic value of each first plant sample and pairing them; inputting the paired whole-genome SNP loci and paired phenotypic values into a feature selection module with chromosome perception function to obtain target whole-genome SNP loci and maintaining the pairing of the target whole-genome SNP loci and the paired phenotypic values; inputting different second plant samples into an output network module with contrast learning function to generate a feature output and a low-dimensional feature output, where the second plant samples have paired target whole-genome SNP loci and paired phenotypic values; calculating a total loss based on the feature output, low-dimensional feature output, and paired phenotypic values; and constructing a prediction model by minimizing the total loss.
[0027] Generally speaking, the method for constructing a prediction model for predicting plant phenotypes according to the embodiments of the present disclosure includes five main steps: data acquisition and pairing, target feature extraction, feature generation and output, total loss calculation, and model optimization and construction.
[0028] Data acquisition and pairing: Obtain the whole-genome SNP loci and corresponding phenotypic values of each first plant sample from the breeding population, and pair the whole-genome SNP loci with the phenotypic values.
[0029] Target feature extraction: Input the paired whole-genome SNP loci and phenotypic values into a feature selection module with chromosome perception function to screen and extract the target whole-genome SNP loci. After the extraction is completed, maintain the paired relationship between the target whole-genome SNP loci and the phenotypic values.
[0030] Feature generation and output: Input different second plant samples containing the target whole-genome SNP loci and phenotypic values into an output network module with contrast learning function to generate feature output and low-dimensional feature output.
[0031] Total loss calculation: Calculate the total loss based on the feature output, low-dimensional feature output, and paired phenotypic values. The total loss includes prediction error (i.e., the loss in the traditional model) and contrast learning loss (i.e., the loss calculated through the contrast learning function of the output network module in the embodiments of the present disclosure).
[0032] Model optimization and construction: Optimize the model parameters by minimizing the total loss, and finally construct a deep learning prediction model for plant phenotype prediction.
[0033] Specifically, the method 100 for constructing a prediction model for predicting plant phenotypes includes the following steps:
[0034] Step S110, obtain the whole-genome SNP loci of each first plant sample and the phenotypic value of each first plant sample, and pair them.
[0035] The first plant sample is a sample of a plant individual in a breeding population, containing a set of whole-genome information of the plant individual and the phenotypic value corresponding to the first plant sample. Among them, the breeding population can be understood as a cluster of individuals in a specific environment, and these individuals have common genetic relationships and varietal characteristics. Through in-depth research and breeding work on the breeding population, improved varieties with specific agricultural values can be cultivated. The whole-genome information of the first plant sample can extract whole-genome SNP loci according to the specified partitioning rules to represent the genetic variation information at specific positions in the whole genome. In addition, each first plant sample has a specific phenotype, and the phenotype can be quantitatively represented by a phenotypic value. Therefore, through each first plant sample in the breeding population, the whole-genome SNP loci of the plant sample and the phenotypic value of the plant sample can be obtained. Pairing the two is used as the input for step S120.
[0036] Step S120: Input the paired whole-genome SNP loci and paired phenotypic values into the feature selection module with chromosome perception function to obtain the target whole-genome SNP loci and maintain the pairing of the target whole-genome SNP loci and the paired phenotypic values.
[0037] Figure 2 The overall block diagram of the method for constructing a prediction model for predicting plant phenotypes according to an embodiment of the present disclosure is shown. As Figure 2 shown, the overall structure includes a feature selection module (i.e., the part indicated by reference numeral 1 in Figure 2 ) and an output network module (i.e., the part indicated by reference numeral 2 in Figure 2 ), and there is a sequential relationship between the two. The paired whole-genome SNP loci and paired phenotypic values are used as the most initial input and input into the feature selection module. Here, the "paired phenotypic value" is used to represent the completed paired phenotypic value. In step S120, due to the chromosome perception function of the feature selection module, the whole-genome SNP loci can undergo certain changes and screening, such as grouping the SNP loci by chromosome intervals and then screening them to become the target whole-genome SNP loci. At this time, since the whole-genome SNP loci and the phenotypic values are paired, after changing to the target whole-genome SNP loci, the target whole-genome SNP loci can continue to maintain the pairing with the phenotypic value. Through the target whole-genome SNP loci and the paired phenotypic values, the plant sample, that is, the second plant sample, can be determined.
[0038] Step S130: Input different second plant samples into the output network module with contrast learning function to generate feature outputs and low-dimensional feature outputs.
[0039] Through step S120, a second plant sample can be obtained. The second plant sample has paired target whole-genome SNP loci and paired phenotypic values. At this time, the second plant sample is not directly input into the output network module. Since the output network module has a contrastive learning function, the output network module can receive multiple different second plant samples as inputs in one round. The output network module can include multiple branches for outputting various types of results, such as results with different dimensions. In this embodiment, they are feature output and low-dimensional feature output. During the process of constructing the prediction model, the output of the output network module can be used to optimize the prediction model.
[0040] The feature output includes the predicted phenotypic values of the second plant sample, that is, the phenotypic values calculated by the model. Opposite to the predicted phenotypic values are the paired phenotypic values. The paired phenotypic values are the phenotypic values after pairing. Here, the phenotypic values can be obtained from the plant sample and can be paired with the SNP loci.
[0041] The low-dimensional feature output includes the low-dimensional representation of the target whole-genome SNP loci of the second plant sample.
[0042] Step S140, calculate the total loss based on the feature output, low-dimensional feature output, and paired phenotypic values.
[0043] By obtaining the feature output, low-dimensional feature output, and paired phenotypic values of multiple different second plant samples through the above steps, the total loss can be calculated using a loss function. This total loss is the gap between the predicted value and the true value. The smaller the gap, the more accurate the prediction model.
[0044] Step S150, construct a prediction model by minimizing the total loss.
[0045] After each round, the parameters of the model may change. By comparing the total losses under different parameters, the parameters with the smallest total loss are found. At this time, the model is the constructed prediction model.
[0046] Figure 3 Shows a schematic flowchart of the method for obtaining target whole-genome SNP loci according to an embodiment of the present disclosure.
[0047] As Figure 3 shown, in some embodiments, inputting paired whole-genome SNP loci and paired phenotypic values into a feature selection module with chromosome perception function to obtain target whole-genome SNP loci includes: obtaining target whole-genome SNP loci that explain local information, that is, local SNP loci; obtaining target whole-genome SNP loci that explain global information, that is, global SNP loci; concatenating the local SNP loci and global SNP loci in order to generate the representation of the target whole-genome SNP loci.
[0048] Specifically, the method 200 for obtaining target genome-wide SNP loci includes the following steps:
[0049] Step S210, obtain the target genome-wide SNP loci that explain local information.
[0050] The target genome-wide SNP loci include two parts, local SNP loci and global SNP loci. In this step, the genome-wide SNP loci and paired phenotypic values of the input first plant sample are processed to obtain the target genome-wide SNP loci that explain local information, that is, local SNP loci.
[0051] Local SNP loci are used to capture key genetic variation information in certain specific regions of the genome, and mainly describe the direct relationship between SNP loci in these regions and phenotypic changes. The acquisition of local SNP loci is usually based on chromosomal regions or gene functional regions. The whole genome is divided into several groups according to chromosomes or functional units, and each group of SNP loci is screened by feature selection methods to retain SNP loci that contribute more to phenotypic changes. Such a processing method can reduce the interference of redundant features, thereby improving the accuracy and computational efficiency of phenotypic prediction.
[0052] The functions of local SNP loci are mainly reflected in the following aspects:
[0053] (1) Capturing local genetic effects: By identifying SNP loci directly related to phenotypes, local SNP loci reflect the action mechanism of local genetic variation.
[0054] (2) Improving the generalization ability of the model: Removing noise and irrelevant SNP loci enables the model to focus more on the core genetic information related to phenotypes.
[0055] (3) Reducing computational complexity: By screening and dimensionality reduction of features, the dimensionality of the model input data is significantly reduced, thereby improving computational efficiency.
[0056] In this step, local SNP loci and global SNP loci are combined to form target genome-wide SNP loci, which are jointly used for subsequent model training to maximize the utilization of local and global genetic variation information and improve phenotypic prediction performance.
[0057] Step S220, obtain the target genome-wide SNP loci that explain global information.
[0058] In this step, the genome-wide SNP loci and paired phenotypic values of the input first plant sample are processed to obtain the target genome-wide SNP loci that explain global information, that is, global SNP loci.
[0059] The global SNP loci represent the SNP loci with strong explanatory power in explaining phenotypic changes using all genome-wide SNP loci, that is, the direct interaction relationship between all genomic SNP loci in the whole genome of the first plant sample and the paired phenotypic values. The global SNP loci are obtained by selecting SNP loci with significant explanatory power for phenotypic changes from the whole genome, reflecting the overall correlation between genetic variation and phenotype at the whole genome level. By evaluating the contribution of genome-wide SNP loci to paired phenotypic values, the global SNP loci retain the key SNPs that significantly affect phenotypic changes at the whole genome scale, while excluding noisy or irrelevant SNP loci.
[0060] The roles of the global SNP loci are reflected in the following aspects:
[0061] (I) Capturing global genetic effects: The global SNP loci cover the entire genome range, integrating information from different chromosomal and gene functional regions, and helping to reflect the global characteristics of phenotypic changes.
[0062] (II) Enhancing the comprehensiveness of the model: By integrating genome-wide information, the global SNP loci provide broader genetic background support for the model, avoiding over-reliance on local effects in certain specific regions.
[0063] (III) Optimizing model performance: Screening out key SNP loci across the whole genome not only reduces the noise brought by redundant features but also retains the comprehensive contribution of genome-wide information to phenotypic prediction, enhancing the generalization ability and prediction accuracy of the model.
[0064] Step S230: Sequentially splice the local SNP loci and the global SNP loci to generate the target genome-wide SNP loci.
[0065] After obtaining the local target genome-wide SNP loci and the global SNP loci, these SNP loci need to be spliced into a plant sample. This step is also called SNP splicing. After splicing, a second plant sample containing the target genome-wide SNP loci is generated. Through the above operations, SNP splicing effectively connects local and global information, providing an efficient and high-quality feature basis for phenotypic prediction. This process is not only the core link of data preprocessing but also an important guarantee for the efficient modeling of deep learning models. The purpose of this step is to construct an optimized data input format for downstream models to perform efficient learning and prediction.
[0066] The main functions and roles of SNP splicing:
[0067] (1) Feature integration: Local SNP sites capture local genetic effects, and global SNP sites capture overall genetic background information. Through splicing, the two are combined to form target genome-wide SNP sites, enabling the data to have stronger explanatory power at both the local and global levels.
[0068] (2) Improving data quality: By screening and retaining key SNPs with high explanatory power and removing noisy or redundant SNP sites, the representativeness of the samples is enhanced. This improvement not only optimizes the quality of the sample data but also reduces the errors caused by the model's input of irrelevant features.
[0069] (3) Reducing data dimensionality: Traditional genomic data has extremely high dimensionality. Directly inputting it into the model may lead to overly high computational complexity and the risk of overfitting. The SNP sites obtained through screening and splicing significantly reduce the data dimensionality, thereby improving computational efficiency and reducing the model's burden.
[0070] (4) Maintaining phenotypic consistency: Although operations such as screening and splicing are performed on the SNP data, the phenotypic values (target variables) of each sample remain unchanged. This consistency ensures that while the data features are optimized, the biological significance and statistical consistency of the samples are not disrupted.
[0071] (5) Providing optimized features for model input: The second plant sample generated by splicing has both representative target genome-wide SNP sites and maintains the accuracy of the phenotype. This sample provides optimized feature input for subsequent deep learning models, helping to improve the accuracy and reliability of phenotype prediction.
[0072] Figure 4 The schematic flowchart of the method for obtaining local SNP sites according to the embodiments of the present disclosure is shown.
[0073] As Figure 4 shown, in some embodiments, obtaining the target genome-wide SNP sites that explain local information, i.e., local SNP sites, includes: grouping the genome-wide SNP sites by chromosomal intervals; constructing a feature selection model and performing feature selection on the grouped genome-wide SNP sites and the paired phenotypic values paired with the genome-wide SNP sites respectively to obtain local SNP site groups; splicing the local SNP sites of each local SNP site group to obtain local SNP sites.
[0074] Obtaining local SNP sites mainly includes three steps:
[0075] (1) Dividing the input genome-wide SNP sites into several chromosomal intervals according to the chromosomal structure or specified rules to form multiple groups, and each group contains SNP sites belonging to the corresponding chromosomal interval.
[0076] (2) For each chromosomal interval after grouping of the whole-genome SNP loci, construct a feature selection model with the paired phenotypic values. Screen the SNP loci within the group through this model to obtain a local SNP locus group with strong phenotypic interpretation ability.
[0077] (3) Stitch together the local SNP locus groups screened within each chromosomal interval to form a complete set of local SNP loci and obtain the local SNP loci.
[0078] Specifically, the method 300 for obtaining local SNP loci includes the following steps:
[0079] Step S310, group the whole-genome SNP loci by chromosomal interval.
[0080] Group the whole-genome SNP loci of the first plant sample by chromosomal interval. For example, based on chromosomal regions or gene functional regions, the whole genome can be divided into several groups according to chromosomes or functional units, and each group of SNP loci has a specific chromosomal interval.
[0081] Step S320, for the grouped whole-genome SNP loci, respectively construct a feature selection model with the paired phenotypic values of the whole-genome SNP loci and perform feature selection to obtain a local SNP locus group.
[0082] Since a plant sample has one phenotypic value and the whole-genome SNP loci are paired with the phenotypic values, the grouped whole-genome SNP loci of the same first plant sample also have the same phenotypic value. Through the grouped whole-genome SNP loci and the paired phenotypic values, a feature selection model can be constructed and feature selection can be performed. The purpose is to obtain the importance values of the SNP nodes in each group for the paired phenotypic values and select some groups with importance values greater than a certain threshold, that is, the local SNP locus group.
[0083] In this embodiment, the feature selection of the feature model can be performed using the LightGBM algorithm, which is a tree-based method that can calculate the non-linear relationship between SNP loci and helps to identify key loci. This algorithm can obtain the information gain IG (equivalent to the importance value in the previous text) between the SNP loci and the paired phenotypic values. In this embodiment, the SNP loci with positive information gain (IG > 0) are retained, and the SNP loci with IG = 0 are discarded, that is, the SNP loci with positive correlation to the paired phenotypic values in the SNP loci are retained.
[0084] Step S330, stitch together the local SNP loci of each local SNP locus group to obtain the local SNP loci.
[0085] For the remaining local SNP locus groups, the local SNP loci of the local SNP locus groups can be spliced to obtain local SNP loci.
[0086] In some embodiments, obtaining the target whole-genome SNP loci that explain global information, i.e., global SNP loci, includes: constructing a feature selection model with the whole-genome SNP loci and paired paired phenotypic values and performing feature selection to obtain global SNP loci.
[0087] Specifically, the whole-genome SNP loci and paired paired phenotypic values of the first plant sample can be used to construct a feature selection model and perform feature selection, aiming to obtain the importance values of the whole-genome SNP nodes for the paired phenotypic values and select certain groups with importance values greater than a certain threshold, i.e., global SNP loci.
[0088] Similar to the previous text, in this embodiment, the feature selection of the feature model can be performed using the LightGBM algorithm, retaining the SNP loci with positive information gain (IG>0) and discarding the SNP loci with IG = 0, that is, retaining the SNP loci with positive correlation to the paired phenotypic values among the SNP loci.
[0089] As Figure 2 shown, in some embodiments, the output network module includes a first branch and a second branch, and the first branch and the second branch are combined by addition.
[0090] Specifically, the output network module includes two branches, the first branch and the second branch. The inputs of these two branches are the same, both being the second plant sample. In this embodiment, the output dimensions of the first branch and the second branch are both 1024 dimensions. Therefore, the first branch and the second branch can be combined by addition.
[0091] As Figure 2 shown, in some embodiments, the first branch includes three layers of non-linear fully connected layers, and the second branch includes one layer of linear fully connected layer.
[0092] Specifically, the first branch can include three layers of non-linear fully connected layers. In this embodiment, the activation functions of these three layers of non-linear fully connected layers are all ReLU activation functions, and the dimensions of these three layers of non-linear fully connected layers decrease in geometric progression, being 4096 dimensions, 2048 dimensions, and 1024 dimensions respectively. The second main branch can include one layer of linear fully connected layer. In this embodiment, the dimension of this layer of linear fully connected layer is 1024 dimensions.
[0093] The genetic effects of GS mainly include additive effect, epistatic effect, and dominance effect. The additive effect can be interpreted as the linear relationship between loci. The epistatic effect refers to the non-linear relationship between loci. The dominance effect is the influence of different alleles at the same locus. The latter two are called non-additive effects. Most traditional genomic prediction models only consider the additive effect when estimating genetic effects, but they are very important for traits closely related to adaptability and traits with low heritability. Therefore, factors of non-additive effects should be considered when constructing genomic prediction models. Through the above two branches, the additive effect, epistatic effect, and dominance effect of the genetic effects of GS can be simultaneously reflected. Compared with other models, due to the role of the first branch, this model can reflect non-additive effects.
[0094] Figure 5 A schematic flowchart showing a method of inputting different second plant samples into an output network module according to an embodiment of the present disclosure
[0095] As Figure 5 shown, in some embodiments, inputting different second plant samples into an output network module with contrast learning function includes: when inputting sample A belonging to the second plant sample into the output network module, randomly matching another sample B from the remaining second plant samples; jointly inputting the target whole-genome SNP loci and paired phenotypic values of sample A and sample B, the two second plant samples, into the output network module.
[0096] Specifically, the output network module in the present disclosure has a contrast learning function, that is, by inputting multiple plant samples simultaneously, the constructed model has better prediction ability. In this embodiment, the number of plant samples input simultaneously is 2.
[0097] This step ensures that in each training round, the input sample A and sample B are not the same sample, but another sample randomly selected from the remaining samples. This random selection helps to enhance the generalization ability of the model, avoid overfitting, and at the same time promote the learning of differences between samples.
[0098] The method 400 of inputting different second plant samples into an output network module includes the following steps:
[0099] Step S410, when inputting sample A belonging to the second plant sample into the output network module, randomly matching another sample B from the remaining second plant samples.
[0100] When inputting the second plant sample into the output network module, several second plant samples can be first calculated and obtained, and then one of the second plant samples, that is, sample A, is selected as one of the input plant samples. Then, randomly match another sample B from the remaining calculated second plant samples, and sample A and sample B are not the same.
[0101] Step S420: Input the target whole-genome SNP loci and paired phenotypic values of the two second plant samples, namely sample A and sample B, into the output network module together.
[0102] After obtaining sample A and sample B, combine the target whole-genome SNP loci and paired phenotypic values of these two plant samples, namely sample A and sample B, and input them into the output network module together.
[0103] As Figure 2 shown, in some embodiments, the output network module further includes a third branch and a fourth branch. Among them, the third branch is used to calculate the predicted phenotypic value of the total loss, and the fourth branch is used to calculate the low-dimensional representation of the target whole-genome SNP loci of the total loss.
[0104] Specifically, the output network module further includes two branches, namely the third branch and the fourth branch. These two branches are located after the addition of the first branch and the second branch. The third branch may include a Flatten layer, two non-linear fully connected layers, and an Output layer, and is used to calculate the predicted phenotypic value of the total loss. The dimensions of each layer are 4096 dimensions, 2048 dimensions, and 1 dimension respectively; while the fourth branch may include a Flatten layer, two non-linear fully connected layers, and an Embedding layer, and is used to calculate the low-dimensional representation of the target whole-genome SNP loci of the total loss. The dimensions of each layer are 4096 dimensions, 2048 dimensions, and 1024 dimensions respectively.
[0105] In some embodiments, the method for calculating the total loss includes: calculating the first main loss through the predicted phenotypic value of sample A and the paired phenotypic value of sample A; calculating the second main loss through the predicted phenotypic value of sample B and the paired phenotypic value of sample B; calculating the contrast loss through the difference between the low-dimensional representations of the target whole-genome SNP loci of sample A, the low-dimensional representations of the target whole-genome SNP loci of sample B, the paired phenotypic value of sample A, and the paired phenotypic value of sample B; adding the first main loss, the second main loss, and the contrast loss to obtain the total loss.
[0106] Specifically, the loss function used for optimization in the embodiments of the present disclosure includes a main loss and a contrast loss. The main loss is a relatively common algorithm and is calculated using the predicted phenotypic value and the paired phenotypic value of the plant sample. That is, for a certain plant sample, its main loss is the mean squared error (MSE) between the predicted phenotypic value and the paired phenotypic value.
[0107] The main loss function is as follows:
[0108]
[0109] Among them, LossMSE represents the main loss of plant samples, N represents the number of plant samples, and y i represents the paired phenotypic value of the i-th plant sample, represents the predicted phenotypic value of the i-th plant sample.
[0110] In the disclosed embodiment, since the output network module has a comparative learning function, a comparative loss function is additionally added when calculating the total loss. The low-dimensional feature output includes a low-dimensional representation of the target whole-genome SNP site of the second plant sample. Therefore, the low-dimensional representation of the target whole-genome SNP site of the two plant samples represents the corresponding low-dimensional feature output. The low-dimensional representation of the target whole-genome SNP site of the two plant samples and the difference in the paired phenotypic values of the target whole-genome SNP site of the two plant samples are calculated to obtain a comparative loss function for the output network module with a comparative learning function in the disclosed embodiment.
[0111] The contrastive loss function is as follows:
[0112]
[0113] Among them, Loss contrastive represents the contrast loss of a pair of plant samples, N is the number of plant samples, represents the Euclidean distance between the i-th pair of embeddings, diff i Represents the phenotypic difference between the i-th pair of plant samples.
[0114] Therefore, the total loss can be simply expressed as:
[0115] Total loss = first main loss + second main loss + contrast loss
[0116] First main loss = Loss MSE (Predicted phenotypic value A, paired phenotypic value A)
[0117] Second main loss = Loss MSE (Predicted phenotypic value B, paired phenotypic value B)
[0118] Contrast loss = Loss contrastive (low dimension represents A, low dimension represents B, paired phenotype value difference)
[0119] Among them, the paired phenotype value of sample A is paired phenotype value A, the predicted phenotype value is predicted phenotype value A, and the low-dimensional representation is low-dimensional representation A; the paired phenotype value of sample B is paired phenotype value B, the predicted phenotype value is predicted phenotype value B, and the low-dimensional representation is low-dimensional representation B; the paired phenotype value difference is the difference between the paired phenotype value of sample A and the paired phenotype value of sample B.
[0120] The predicted phenotypic value, which is the phenotypic value calculated by the model. During the model construction process, the predicted phenotypic value can represent the phenotypic value calculated by the model under the current parameters. In contrast to the predicted phenotypic value is the paired phenotypic value. The paired phenotypic value is the phenotypic value after pairing. Here, the phenotypic value can be directly obtained from real plant samples and can be paired with SNP loci. The smaller the gap between the predicted phenotypic value and the paired phenotypic value, the closer the model is to the actual situation and the higher the accuracy of the constructed model.
[0121] After calculating the total loss, optimize the model parameters by minimizing the total loss, and finally construct a deep learning prediction model for plant phenotype prediction.
[0122] Figure 6 The structural block diagram of the device for constructing a prediction model for predicting plant phenotypes according to an embodiment of the present disclosure is shown.
[0123] As Figure 6 shown, a device for constructing a prediction model for predicting plant phenotypes includes: a processor configured to execute program instructions; and a memory configured to store the program instructions; wherein, when the program instructions are loaded and executed by the processor, the device is caused to execute any one of the methods in this embodiment.
[0124] Specifically, the device 500 can be implemented as various types of devices, including but not limited to mainframes, personal computers (PCs), mobile devices, etc. The device 500 includes a processor 501 and a memory 502. The processor 501 controls the operation of the device 500 by executing the programs stored in the memory 502. The processor 501 can be implemented using, including but not limited to, a central processing unit (CPU), a graphics processing unit (GPU), an application processor (AP), an artificial intelligence processor chip (IPU), etc. The memory 502 can be used to store various data and instructions processed in the device 500, including the method for constructing a prediction model for predicting plant phenotypes according to an embodiment of the present disclosure. The memory 502 can include at least one of volatile memory or non-volatile memory. The non-volatile memory can include read-only memory (ROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), etc. The volatile memory can include dynamic RAM (DRAM), static RAM (SRAM), synchronous DRAM (SDRAM), etc. In addition, the memory 502 can also include at least one of a hard disk drive (HDD), a solid state drive (SSD), a high density flash (CF), a secure digital (SD) card, a micro secure digital (Micro-SD) card, a mini secure digital (Mini-SD) card, an extreme digital (xD) card, a cache, or a memory stick.
[0125] While numerous embodiments of the present disclosure have been shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. Many variations, changes, and alternative approaches may be contemplated by those skilled in the art without departing from the spirit and scope of the present disclosure. It should be understood that various alternatives to the embodiments of the present disclosure described herein may be employed in practicing the present disclosure. The appended claims are intended to define the scope of the present disclosure and thus cover equivalents or alternatives within the scope of these claims.
Claims
1. A method for constructing a prediction model for predicting plant phenotypes, characterized in that: The method comprises: Obtaining the whole genome SNP loci of each first plant sample and the phenotypic value of each first plant sample, and performing pairing; Inputting the paired whole-genome SNP site and the paired phenotypic value into a feature selection module with chromosome perception function to obtain a target whole-genome SNP site, and maintaining the pairing of the target whole-genome SNP site and the paired phenotypic value; Inputting a different second plant sample into an output network module with a contrastive learning function to generate a feature output and a low-dimensional feature output, wherein the second plant sample has a paired target whole-genome SNP site and a paired phenotypic value; Calculate the total loss based on the feature output, the low-dimensional feature output, and the paired phenotype value; By minimizing the total loss, a prediction model is constructed.
2. The method according to claim 1, characterized in that Inputting the paired whole-genome SNP sites and the paired phenotype values into a feature selection module with chromosome perception function to obtain a target whole-genome SNP site, including: Obtain the target genome-wide SNP site for explaining local information, i.e., the local SNP site; Obtain the target genome-wide SNP site for explaining global information, i.e., the global SNP site; The local SNP sites and the global SNP sites are sequentially spliced to generate a whole genome SNP site representing the target.
3. The method according to claim 2, characterized in that Obtaining the target genome-wide SNP site for explaining local information, i.e., the local SNP site, comprises: Grouping the whole genome SNP sites by chromosome interval; The grouped whole genome SNP sites are respectively paired with the paired phenotypic values of the whole genome SNP sites, and a feature selection model is constructed and feature selection is performed to obtain a local SNP site group; The local SNP sites of each of the local SNP site groups are spliced to obtain the local SNP sites.
4. The method according to claim 2, characterized in that: Obtaining the target genome-wide SNP site for explaining global information, i.e., the global SNP site, includes: The whole genome SNP sites are paired with the paired phenotype values to construct a feature selection model and perform feature selection to obtain the global SNP sites.
5. The method according to claim 1, characterized in that The output network module includes a first branch and a second branch, and the first branch and the second branch are combined by adding.
6. The method according to claim 5, characterized in that The first branch includes three nonlinear fully connected layers, and the second branch includes one linear fully connected layer.
7. The method according to claim 1, characterized in that Different second plant samples are input to the output network module with contrastive learning function, including: When a sample A belonging to the second plant sample is input to the output network module, another sample B is randomly matched from the remaining second plant samples; The sample A and the sample B, the target whole genome SNP sites of the two second plant samples and the paired phenotype values are input to the output network module together.
8. The method according to claim 1, characterized in that The output network module also includes a third branch and a fourth branch, wherein the third branch is used to calculate the predicted phenotypic value of the total loss, and the fourth branch is used to calculate the low-dimensional representation of the target whole-genome SNP site of the total loss.
9. The method according to claim 8, characterized in that Methods for calculating the total loss include: Calculating a first principal loss by using the predicted phenotypic value of the sample A and the paired phenotypic value of the sample A; Calculating a second principal loss by using the predicted phenotypic value of sample B and the paired phenotypic value of sample B; Calculate the comparison loss by using the low-dimensional representation of the target whole-genome SNP site of the sample A, the low-dimensional representation of the target whole-genome SNP site of the sample B, and the difference between the paired phenotypic value of the sample A and the paired phenotypic value of the sample B; The first main loss, the second main loss and the comparative loss are added together to obtain the total loss.
10. A device for constructing a prediction model for predicting plant phenotypes, comprising: a processor configured to execute program instructions; as well as a memory configured to store the program instructions; It is characterized in that when the program instructions are loaded and executed by the processor, the device executes the method according to any one of claims 1-9.
Citation Information
Patent Citations
Whole genome association analysis algorithm based on parent genotypes and progeny phenotypes
CN113793637A
Plant phenotype prediction method and device based on multiple omics
CN116992919A
Complex character efficient genome prediction method based on deep learning
CN117711484A
Rice phenotype prediction method and system based on whole genome selection
CN118072823A
Phenotype assisted lemon breeding method based on deep learning
CN118216422A