A multi-trait collaborative screening method and system fusing known gene information

By integrating known gene information into a multi-trait collaborative screening method, the BLUP model and TOPv2 algorithm are used to screen candidate materials with superior genotypes, solving the problem that traditional methods fail to fully utilize known functional gene information and achieving efficient breeding in complex environments.

CN121075416BActive Publication Date: 2026-03-31HUAZHONG AGRI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-20
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

Traditional TOP methods fail to fully incorporate known functional gene information, limiting the ability of multi-trait collaborative screening models to identify key functional sites and the stability and interpretability of material selection in complex environments, making it difficult to meet the needs of multiple challenges in crop breeding.

Method used

A multi-trait collaborative screening method integrating known gene information was adopted. Genotype prediction values ​​were obtained through the BLUP model, and superior genotypes were screened by combining t-test. The TOPv2 model was constructed and a weighted optimization algorithm was used to screen out candidate materials that are closest to the target material under specific conditions.

Benefits of technology

This technology enables the precise screening of candidate materials with superior genotypes within a multi-trait synergistic framework, improving the efficiency and scientific rigor of the breeding process and enhancing the accuracy and stability of material selection in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121075416B_ABST
    Figure CN121075416B_ABST
Patent Text Reader

Abstract

The application provides a multi-traits collaborative screening method and system fusing known gene information, which is based on a training population, applies a TOP algorithm to estimate environment-specific trait weights and gene-specific weights, combines known excellent gene information in a test population, and comprehensively screens candidate materials similar to a target reference individual in multiple traits and better in key traits by combining the similarity of phenotypes and genotypes, constructs a material screening system with multiple traits, better key traits and enrichment of excellent genes, and realizes the function of collaborative integration of multi-trait information and genotype information. The application helps breeders to accurately screen breeding materials with target trait advantages and enrichment of excellent alleles in a specific ecological environment, improves the efficiency and scientificity of material selection, provides a decision basis for multi-trait collaborative improvement and genotype-driven intelligent breeding, and has significant application and promotion value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of bioinformatics technology, specifically relating to a method and system for collaborative screening of multiple traits that integrates known gene information. Background Technology

[0002] With the dual pressures of global population growth and climate change, crop breeding faces multiple challenges in improving yield, quality, and environmental adaptability. In modern crop genetic improvement, accurately and efficiently screening candidate materials with excellent comprehensive traits is a crucial step in breeding work. Utilizing high-throughput sequencing technology and multi-environment, multi-phenotypic acquisition platforms, researchers can obtain large-scale genomic and multi-dimensional phenotypic data, providing a data foundation and technical support for precise material screening methods based on the synergistic integration of genotype and multiple traits.

[0003] Crop breeding objectives often involve the synergistic optimization of multiple key traits (such as yield, stress resistance, and quality), and screening strategies based on a single trait are insufficient to meet practical breeding needs. The TOP (Target-oriented prioritization) algorithm is a comprehensive prediction framework that integrates phenotypic and genotypic data, and it has been applied to generalized modeling of trait performance and candidate material screening in crops such as maize.

[0004] Some important functional genes play a core regulatory role in the coordination of multiple traits. Incorporating this prior information into models can improve the accuracy and biological relevance of breeding screening. However, traditional TOP methods mainly rely on statistical models to model the relationship between traits and genotypes, failing to fully incorporate prior biological information such as known functional genes. This limits the model's ability to identify key functional sites and reduces the stability and interpretability of material selection in complex environments. Therefore, a candidate material screening method is needed that can incorporate known functional gene information within a multi-trait synergistic framework and achieve multi-trait synergistic optimization. This would improve the ability to analyze complex traits and the efficiency of material selection during the breeding process, facilitating the implementation of molecular design breeding. Summary of the Invention

[0005] The technical problem to be solved by the present invention is to provide a method and system for multi-trait collaborative screening that integrates known gene information, for the collaborative integration of multi-trait information and genotype information.

[0006] The technical solution adopted by this invention to solve the above-mentioned technical problems is as follows: a multi-trait collaborative screening method integrating known gene information, comprising the following steps:

[0007] S0: Obtain biological materials and divide them into training and testing sets according to a certain ratio;

[0008] S1: Based on the multi-trait phenotypic data, genotype information and kinship information of the training set, the genotype prediction value of each material at the known gene loci is obtained through the BLUP model, and the phenotypic prediction value is obtained through ten-fold cross-validation of the training set.

[0009] S2: Use t-tests to screen known genotypes that have significant phenotypic effects in the training set and identify superior genotypes with breeding significance;

[0010] S3: Select individuals that exhibit excellent multi-trait performance and carry superior genotypes as target materials, or select widely cultivated commercial varieties as target materials based on prior knowledge.

[0011] S4: Construct a TOPv2 model by combining the genotype predictions, phenotypic predictions, and information from the target material in the training set; solve for the weight parameters of the multi-trait phenotypes and known genes using a weighted optimization algorithm;

[0012] S5: In the test set, based on the multi-trait phenotype-genotype similarity between candidate materials and target materials, candidate materials with the phenotypes that are closest to the target materials under specific environments are selected.

[0013] According to the above scheme, the specific steps in step S1 are as follows:

[0014] S11: Construct a BLUP model based on genotype information and multi-trait phenotypic data from the training set;

[0015] S12: A single-view kinship matrix construction method using the MVBLUP approach, which constructs a kinship matrix based on individual multi-trait phenotypic data to measure phenotypic similarity between different materials;

[0016] S13: Use the BLUP model to obtain the predicted genotype values ​​of the material at known gene loci.

[0017] Furthermore, in step S11, it is assumed that... For the first Genotype values ​​of known genes in an individual vector, for matrix, for The fixed effects vector, for The random effects vector, For residual error vector, and These are the variances associated with random effects and residuals, respectively. Let be the identity matrix; the expression for the BLUP model is:

[0018] (1)

[0019] In step S12, the MVBLUP single-view kinship matrix method is used to obtain a kinship matrix constructed based on individual multi-trait phenotypic data. The expression for the random effects vector is:

[0020] (2)

[0021] in, , .

[0022] According to the above scheme, the specific steps in step S2 are as follows:

[0023] S21: Based on the multi-trait phenotypic data of the training set, t-tests were performed on different haplotype groups of each gene locus to screen gene loci with significant differences in phenotype.

[0024] S22: Based on the multi-phenotypic data of the training set, determine the direction of allele effect of haplotypes in gene loci with significant differences in phenotype, and define the alleles that perform well in the target trait as superior genotypes.

[0025] According to the above scheme, the specific steps in step S4 are as follows:

[0026] S41: Construct pseudo-observation vectors and predicted vectors for the phenotype and genotype of the target individuals in the training set;

[0027] S42: Use the TOP method to establish a collaborative optimization objective function, and obtain the optimal weights of multi-trait phenotypes and known genotypes by maximizing the objective function.

[0028] Furthermore, in step S41, the training set is used in the first... Pseudo-observation vectors under various environments and predicted value vector and genotype vectors of target materials carrying superior genotypes. and the corresponding predicted value vector Build the TOPv2 model;

[0029] In step S42, the weight vector of the multi-trait phenotype is integrated. Weight vector of known genes :

[0030] (4)

[0031] The optimal weight vector is derived using the TOP method:

[0032] (3).

[0033] According to the above scheme, the specific steps in step S5 are as follows:

[0034] S51: Construct a multi-phenotypic phenotype-genotype similarity function using the optimal weights obtained in step S4;

[0035] S52: Calculate the overall similarity between candidate materials and target materials in all test sets;

[0036] S53: Sort the candidate materials according to the similarity score, and select a number of candidate materials that are similar to the target material and have better key traits than the target material, as the preferred material set with breeding potential in the target environment.

[0037] Furthermore, in step S51, it is assumed that... and These are the phenotypic value vectors of the target material and the first value in the test set, respectively. A vector of phenotypic values ​​for each individual; and These are known genes and the first gene in the target material, respectively. The corresponding genotype of the test material; calculate the genotype of the first test material; In the first environment The phenotypic-genotypic similarity between the test material and the target material is as follows:

[0038] (5).

[0039] A multi-trait collaborative screening system that integrates known genetic information.

[0040] The materials acquisition submodule is used to acquire biological materials and divide them into training and testing sets according to a certain ratio.

[0041] The prediction submodule is used to obtain the genotype prediction value of each material at known gene loci through the BLUP model based on the multi-trait phenotypic data, genotype information and kinship information of the training set, and obtain the phenotypic prediction value through ten-fold cross-validation of the training set.

[0042] The gene screening submodule is used to screen known genotypes that have significant phenotypic effects in the training set using t-tests, and to identify superior genotypes with breeding significance.

[0043] The target submodule is used to select individuals that exhibit excellent multi-trait performance and carry superior genotypes as target materials, or to select widely cultivated commercial varieties as target materials based on prior knowledge.

[0044] The model weight submodule is used to construct the TOPv2 model by combining the genotype predictions, phenotypic predictions, and information from the target material in the training set; and to solve the weight parameters of the multi-trait phenotypes and known genes through a weighted optimization algorithm.

[0045] The phenotypic screening submodule is used to screen candidate materials in the test set based on the multi-trait phenotypic-genotypic similarity between candidate materials and target materials, and to select candidate materials whose phenotypes are closest to those of target materials under specific conditions.

[0046] A computer memory storing a computer program executable by a computer processor, the computer program performing a multi-trait collaborative screening method that incorporates known genetic information.

[0047] The beneficial effects of this invention are as follows:

[0048] 1. The present invention provides a multi-trait collaborative screening method and system that integrates known gene information. Based on a training population, the TOP algorithm is applied to estimate the weights of environment-specific traits and gene-specific traits. In the test population, combined with known superior gene information and the similarity between phenotype and genotype, candidate materials that are similar to the target reference individuals in multiple traits but have better performance in key traits are screened out. A material screening system of "similar multiple traits, better key traits, and enrichment of superior genes" is constructed, realizing the function of collaboratively integrating multi-trait information and genotype information.

[0049] 2. This invention can assist breeders in accurately screening breeding materials with target trait advantages and rich in superior alleles in specific ecological environments, improving the efficiency and scientific nature of material selection and providing a basis for decision-making for multi-trait synergistic improvement and genotype-driven intelligent breeding.

[0050] 3. Compared with traditional methods, this invention can help breeders more accurately identify materials with real breeding value under complex environmental conditions, improve the efficiency and reliability of material selection in the target environment, and has significant application and promotion value.

[0051] Of course, any product implementing this invention does not necessarily need to achieve all of the advantages described above at the same time. Attached Figure Description

[0052] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0053] Figure 1 This is a flowchart of an embodiment of the present invention.

[0054] Figure 2 This is a flowchart illustrating an embodiment of the present invention.

[0055] Figure 3 This is a diagram illustrating how integrating multi-trait phenotypes with known gene information improves material recognition rate in an embodiment of the present invention.

[0056] Figure 4 This is a distribution chart showing the yield of materials screened by two methods under different environments and their similarity to the target variety, according to an embodiment of the present invention.

[0057] Figure 5 This is a similarity distribution diagram of materials and target varieties screened by two methods under the HeB environment in an embodiment of the present invention.

[0058] Figure 6 This is a yield distribution diagram of materials and target varieties screened by two methods under the HeB environment in an embodiment of the present invention.

[0059] Figure 7 This is a comparison chart of the multi-phenotypic distribution of the top-ranked material selected by two methods under the HeB environment in an embodiment of the present invention.

[0060] Figure 8 This is a distribution diagram of the quantity of materials selected in different environments according to embodiments of the present invention.

[0061] Figure 9 This is a phenotypic distribution diagram of flowering period, plant height, and yield traits of materials screened in all environments according to embodiments of the present invention.

[0062] Figure 10 This is a heatmap showing the contribution weights of known functional genes in different environments in this invention embodiment.

[0063] Figure 11 This is a graph showing the number of gene sets that contribute to the model under different environments in embodiments of the present invention. Detailed Implementation

[0064] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.

[0065] Example 1

[0066] See Figure 1 The specific steps of a multi-trait collaborative screening method that integrates known gene information are as follows:

[0067] S0: Obtain biological materials and divide them into training and testing sets according to a certain ratio;

[0068] S1: Based on the multi-trait phenotypic data, genotype information and kinship information of the training set, the genotype prediction value of each material at the known gene loci is obtained through the BLUP model, and the phenotypic prediction value is obtained through ten-fold cross-validation of the training set.

[0069] S2: Use t-tests to screen known genotypes that have significant phenotypic effects in the training set and identify superior genotypes with breeding significance;

[0070] S3: Select individuals that exhibit excellent multi-trait performance and carry superior genotypes as target materials, or select widely cultivated commercial varieties as target materials based on prior knowledge.

[0071] S4: Construct a TOPv2 model by combining the genotype predictions, phenotypic predictions, and information from the target material in the training set; solve for the weight parameters of the multi-trait phenotypes and known genes using a weighted optimization algorithm;

[0072] S5: In the test set, based on the multi-trait phenotype-genotype similarity between candidate materials and target materials, candidate materials with the phenotypes that are closest to the target materials under specific environments are selected.

[0073] According to the above scheme, the specific steps in step S1 are as follows:

[0074] S11: Construct a BLUP model based on genotype information and multi-trait phenotypic data from the training set;

[0075] S12: A single-view kinship matrix construction method using the MVBLUP approach, which constructs a kinship matrix based on individual multi-trait phenotypic data to measure phenotypic similarity between different materials;

[0076] S13: Use the BLUP model to obtain the predicted genotype values ​​of the material at known gene loci.

[0077] Furthermore, in step S11, it is assumed that... For the first Genotype values ​​of known genes in an individual vector, for matrix, for The fixed effects vector, for The random effects vector, For residual error vector, and These are the variances associated with random effects and residuals, respectively. Let be the identity matrix; the expression for the BLUP model is:

[0078] (1)

[0079] In step S12, the MVBLUP single-view kinship matrix method is used to obtain a kinship matrix constructed based on individual multi-trait phenotypic data. The expression for the random effects vector is:

[0080] (2)

[0081] in, , .

[0082] According to the above scheme, the specific steps in step S2 are as follows:

[0083] S21: Based on the multi-trait phenotypic data of the training set, t-tests were performed on different haplotype groups of each gene locus to screen gene loci with significant differences in phenotype.

[0084] S22: Based on the multi-phenotypic data of the training set, determine the direction of allele effect of haplotypes in gene loci with significant differences in phenotype, and define the alleles that perform well in the target trait as superior genotypes.

[0085] According to the above scheme, the specific steps in step S4 are as follows:

[0086] S41: Construct pseudo-observation vectors and predicted vectors for the phenotype and genotype of the target individuals in the training set;

[0087] S42: Use the TOP method to establish a collaborative optimization objective function, and obtain the optimal weights of multi-trait phenotypes and known genotypes by maximizing the objective function.

[0088] Furthermore, in step S41, the training set is used in the first... Pseudo-observation vectors under various environments and predicted value vector and genotype vectors of target materials carrying superior genotypes. and the corresponding predicted value vector Build the TOPv2 model;

[0089] In step S42, the weight vector of the multi-trait phenotype is integrated. Weight vector of known genes :

[0090] (4)

[0091] The optimal weight vector is derived using the TOP method:

[0092] (3).

[0093] According to the above scheme, the specific steps in step S5 are as follows:

[0094] S51: Construct a multi-phenotypic phenotype-genotype similarity function using the optimal weights obtained in step S4;

[0095] S52: Calculate the overall similarity between candidate materials and target materials in all test sets;

[0096] S53: Sort the candidate materials according to the similarity score, and select a number of candidate materials that are similar to the target material and have better key traits than the target material, as the preferred material set with breeding potential in the target environment.

[0097] Furthermore, in step S51, it is assumed that... and These are the phenotypic value vectors of the target material and the first value in the test set, respectively. A vector of phenotypic values ​​for each individual; and These are known genes and the first gene in the target material, respectively. The corresponding genotype of the test material; calculate the genotype of the first test material; In the first environment The phenotypic-genotypic similarity between the test material and the target material is as follows:

[0098] (5).

[0099] This embodiment uses the TOP algorithm to estimate the weights of environment-specific traits and gene-specific traits based on the training population. In the test population, it combines known superior gene information with the similarity of phenotype and genotype to screen candidate materials that are similar to the target reference individuals in multiple traits and have better performance in key traits. It constructs a material screening system that is "similar in multiple traits, better in key traits, and rich in superior genes" and realizes the function of synergistically integrating multiple trait information and genotype information.

[0100] Example 2

[0101] The steps in this embodiment are the same as in Embodiment 1, the difference being that each step is applied to a specific instance. See also Figure 2 Specifically, it includes the following steps:

[0102] S0: Obtain biological materials and divide them into training and testing sets according to a certain ratio;

[0103] S1: Using phenotypic data in the training set, predict the genotypes of known functional genes in the training set, and obtain phenotypic prediction values ​​through 10-fold cross-validation of the training set.

[0104] The specific expression for the BLUP model is as follows:

[0105] (1)

[0106] in For the first Known genotype values ​​in individuals vector, for matrix, for The fixed effects vector, for The random effects vector, For residual error vector, and These are the variances associated with random effects and residuals, respectively. It is an identity matrix.

[0107] Using the MVBLUP single-view kinship matrix method, a kinship matrix constructed based on individual multi-trait phenotypic data is obtained. The specific expression for the random effects vector is:

[0108] (2)

[0109] in, , .

[0110] S2: Perform a t-test on different haplotype groups at each locus to determine whether they are significant for the target phenotype (P<0.05). If significant, determine the directionality based on the haplotype phenotype mean and retain the directional information of the superior genotype.

[0111] S3: Identify the target individuals for population screening. This can be done by selecting representative target individuals based on their high performance in the training set across multiple traits; or by selecting widely cultivated commercial varieties as target individuals based on prior knowledge.

[0112] S4: Construct the TOPv2 model and solve for the optimal weights.

[0113] (1) TOPv2 uses the training population in the first quarter. Pseudo-observation vectors under various environments and predicted value vector And the genotype vector of the target hybrid carrying genotypes related to superior yield. and the prediction vectors of these genes Build the model.

[0114] (2) The optimal weight vector is derived using the TOP method.

[0115] (3)

[0116] (4)

[0117] in, The weight vector representing the multi-trait phenotype. This represents the weight vector of known genes.

[0118] S5: Define the phenotypic-genotypic probability similarity function between candidate materials and target materials, and filter the test set based on the similarity with the target materials.

[0119] No. In the first environment Phenotypic-genotypic similarity between individual test materials and the target ideal type

[0120] (5)

[0121] in, and Represent the phenotypic vector of the target ideal type and the phenotypic vector of the test group, respectively. A vector of phenotypic values ​​for each individual. Similarly, and Representing the known yield-related genes and the first [gene] in the target [are] respectively. The corresponding genotypes of the test hybrids.

[0122] This embodiment specifically designed the following experiment for verification:

[0123] Experiment 1: Predict the genotypes of known functional genes and assess their phenotypic effects

[0124] Genotypes of materials at known functional loci are predicted using multi-trait phenotypic data from the training set, and the phenotypic effects of these loci in the population are evaluated. A BLUP model is constructed based on the multi-trait phenotypic data of the training population, using the phenotypic kinship matrix within the MVBLUP framework as the random effects component to infer the genotype values ​​of each individual at all known functional genes. After prediction, genotype estimates at key gene loci are obtained for each material.

[0125] Experiment 2: Evaluate the phenotypic effects of known functional genes and determine the direction of superior alleles.

[0126] Based on the phenotypic data of the training population, it is determined whether the known functional genes at each locus have a significant effect on the target trait, and the direction of action of superior and inferior alleles is identified. A t-test is performed on different haplotype groups at each locus to assess whether the locus has a significant effect on the target phenotype (e.g., ear weight, plant height, flowering time). If the test result for a locus is significant, the phenotypic mean of different haplotype groups is further calculated to determine the direction of superiority or inferiority, and haplotypes superior to the population mean are defined as "superior genotypes," providing a basis for subsequent construction of target individuals and selection criteria.

[0127] Experiment 3: Constructing a multi-trait-genotype collaborative model and solving for the optimal weights

[0128] Based on multi-trait phenotypic data and known functional gene information in the training set, a multi-trait-genotype collaborative model is constructed, and quantitative criteria for candidate material screening are obtained through weight optimization. First, the phenotypic vector and superior genotype vector of the target material are used as ideal type inputs, serving as comparison standards in the screening process. Next, in each training environment, the predicted phenotypic values ​​of all materials in the training population and the predicted values ​​of the known genotype of the target material are constructed as pseudo-observation vectors and predicted value vectors, respectively. These vectors together form the input of the TOPv2 model, used to quantitatively model the differences between the target material and other materials in the training population in multi-trait performance and key functional gene composition. Then, a collaborative objective function is constructed using the TOP algorithm, setting weights for the phenotypic and genotype components to evaluate the importance of various information in material screening. The differences between the training set materials and the target material in both phenotypic and genotype dimensions are calculated, and the objective function is iteratively solved using a gradient optimization algorithm with the goal of minimizing these differences, ultimately obtaining a set of optimal collaborative weights for multi-trait phenotypes and known functional genes. This set of weights reflects the relative contribution of multiple trait expressions and key functional genes to material performance under the current training set and specific environment. It is used to calculate the similarity between candidate materials and target materials in the test set, serving as a basis for subsequent screening. For example... Figure 3 As shown, under specific conditions, with the increase in the number of yield-related genes, the recognition rate of TOPv2 significantly improved compared to the hybrid library without added known functional gene information under different hybrid library sizes, eventually reaching a plateau. In the HeB environment, when the hybrid library size was 100, the recognition accuracy without added known functional genes was 0.46; after adding 50 genes, the recognition accuracy of TOPv2 reached 0.91; after adding 100 genes, the accuracy increased to 0.97; and after adding 180 genes, the accuracy reached 0.98. The same trend was observed in the other four test environments and at hybrid library sizes of 400 and 1000.

[0129] Experiment 4: Target material-oriented screening in a specific environment of the test set

[0130] Under specific conditions of the test set, candidate materials are screened based on a pre-constructed multi-trait-genotype synergistic model, using the target material as a reference, to improve breeding efficiency under specific ecological conditions. First, the predicted values ​​of multiple traits and the target genotype are obtained for each candidate material, and these are used as inputs, substituted into the obtained TOPv2 optimal weighting parameters for comprehensive evaluation. Then, by calculating the phenotypic and genotypic differences between each test material and the target material, and combining the weighted information of multi-trait phenotypes and functional genes, the overall similarity between the test material and the target material is evaluated. Finally, all candidate materials are ranked according to their similarity scores with the target material, and several candidate materials with high overall similarity and superior performance in core breeding traits (such as yield) are selected to form a set of superior materials under specific conditions. Figure 4 As shown in the five test environments, the yields of materials identified by TOP and TOPv2 were significantly better than those of unidentified materials, and the materials identified by TOPv2 exhibited higher yield levels and stability under all environmental conditions. In the HeB environment, such as... Figure 5 As shown, compared to unidentified materials, the hybrid combinations identified by TOP and TOPv2 closely resembled the target varieties in overall phenotype and genotype; for example... Figure 6 As shown, TOPv2 achieves higher yield while maintaining a high similarity between the identified material and the target material; Figure 7 The distribution of the top-ranked materials selected by the two methods across 17 traits is shown, with the materials selected by TOPv2 showing superior performance in yield and yield-related traits (ERN, KNPR, KNPE, and KWPE). Figure 8 Most of the materials shown exhibit strong environment specificity in their environmental adaptability; however, 13 materials were consistently selected in all five environments, demonstrating good broad-spectrum adaptability. For example... Figure 9 The example shown is of one of the selected materials. Compared with the target material, this material exhibits a stable flowering period, shorter plant height, and stable yield in multiple different environments, demonstrating broad environmental adaptability while maintaining stable yield.

[0131] The contribution weights of 109 known genes included in the model were analyzed in various test environments. The results showed that 74 genes contributed to the model output in all environments. For example... Figure 10 As shown, after cluster analysis based on the contribution weight of each gene in the environment, these genes can be divided into two categories: "stable genes" (42 genes) and "environmentally malleable genes" (67 genes). Figure 11 As shown, although the distribution of the two types of genes differs in different environments, the overall difference is small.

[0132] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0133] Example 3

[0134] This embodiment is used to implement the principle of the above method embodiment to construct a multi-trait collaborative screening system that integrates known gene information, including a material acquisition submodule, a prediction submodule, a gene screening submodule, a target submodule, a model weight submodule, and a phenotypic screening submodule;

[0135] The materials acquisition submodule is used to acquire biological materials and divide them into training and testing sets according to a certain ratio.

[0136] The prediction submodule is used to obtain the genotype prediction value of each material at known gene loci through the BLUP model based on the multi-trait phenotypic data, genotype information and kinship information of the training set, and obtain the phenotypic prediction value through ten-fold cross-validation of the training set.

[0137] The gene screening submodule is used to screen known genotypes that have significant phenotypic effects in the training set using t-tests, and to identify superior genotypes with breeding significance.

[0138] The target submodule is used to select individuals that exhibit excellent multi-trait performance and carry superior genotypes as target materials, or to select widely cultivated commercial varieties as target materials based on prior knowledge.

[0139] The model weight submodule is used to construct the TOPv2 model by combining the genotype predictions, phenotypic predictions, and information from the target material in the training set; and to solve the weight parameters of the multi-trait phenotypes and known genes through a weighted optimization algorithm.

[0140] The phenotypic screening submodule is used to screen candidate materials in the test set based on the multi-trait phenotypic-genotypic similarity between candidate materials and target materials, and to select candidate materials whose phenotypes are closest to those of target materials under specific conditions.

[0141] Each submodule is mainly used to implement the various steps of the method embodiment, which will not be elaborated here.

[0142] It should be noted that, depending on the implementation needs, the various steps / components described in this application can be broken down into more steps / components, or two or more steps / components or parts of the operation of steps / components can be combined into new steps / components to achieve the purpose of this invention.

[0143] This embodiment also includes a processor, a communication interface, a memory, and a communication bus; wherein the processor, the communication interface, and the memory communicate with each other through the communication bus; the memory stores a computer program, and when the program is executed by the processor, the processor performs the steps of a multi-trait collaborative screening method that integrates known gene information.

[0144] This embodiment also provides a computer-readable storage medium storing executable instructions that, when executed by a processor, enable the processor to implement a multi-trait collaborative screening method that integrates known genetic information.

[0145] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects.

[0146] Furthermore, this application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0147] This application is described with reference to the flowchart of the method and computer program product according to Embodiment 1 and the block diagram of the device (system) according to Embodiment 3. It should be understood that each step or block in the flowchart or block diagram, as well as combinations of steps or blocks in the flowchart or block diagram, can be implemented by computer program instructions.

[0148] These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing device to produce a machine, such that the instructions, which are executable by the processor of the computer or other programmable data processing device, produce instructions for implementing the process. Figure 1 One or more processes or boxes Figure 1 A multi-trait collaborative screening system that integrates known genetic information and specifies functions within one or more boxes.

[0149] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes or boxes Figure 1 The function specified in one or more boxes.

[0150] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes or boxes Figure 1 The steps of a multi-trait collaborative screening method that integrates known genetic information are specified in one or more boxes.

[0151] The above embodiments are only used to illustrate the design concept and features of the present invention, and their purpose is to enable those skilled in the art to understand the content of the present invention and implement it accordingly. The protection scope of the present invention is not limited to the above embodiments. Therefore, all equivalent changes or modifications made based on the principles and design ideas disclosed in the present invention are within the protection scope of the present invention.

Claims

1. A method of multi-trait co-selection integrating known genetic information, characterized by: The method comprises the following steps: S0: obtaining biological materials and dividing them into a training set and a test set in proportion; S1: obtaining genotype prediction values of each material at known genetic loci based on multi-trait phenotype data, genotype information and kinship information of the training set by a BLUP model, and obtaining phenotype prediction values by ten-fold cross-validation of the training set; S2: screening known genotypes with significant phenotype effects in the training set by t-test to determine elite genotypes with breeding significance; S3: selecting individuals with excellent multi-trait performance and carrying elite genotypes as target materials, or selecting widely planted commercial varieties as target materials according to prior knowledge; S4: constructing a TOPv2 model combining genotype prediction values, phenotype prediction values of the training set and information of the target materials; solving weight parameters of multi-trait phenotypes and known genotypes by a weighted optimization algorithm; the specific steps are as follows: S41: constructing a pseudo-observation vector and a prediction vector of the target individual phenotype and genotype in the training set; The pseudo observation vector and the predicted value vector under the first environment are obtained by using the training set The genotype vector and the corresponding predicted value vector of the target material carrying the excellent genotype are obtained by using the test set The TOPv2 model is constructed​​​ In step S42, the weight vector of the integrated multi-trait phenotype is determined with the weight vector of the known gene : (4) deriving an optimal weight vector by using the TOP method: (3); S42: establishing a collaborative optimization objective function by using the TOP method, and obtaining optimal weights of multi-trait phenotypes and known genotypes by maximizing the objective function; S5: in the test set, screening candidate materials closest to the target materials in phenotype under a specific environment according to multi-trait phenotype-genotype similarity between the candidate materials and the target materials; the specific steps are as follows: S51: constructing a multi-trait phenotype-genotype similarity function by using the optimal weights obtained in step S4; Let and be the phenotype vectors of the target material and the th individual in the test set, respectively; and be the genotypes of the known genes in the target material and the th test material, respectively; the multi-trait phenotype-genotype similarity between the th test material and the target material in the th environment is calculated as: (5); S52: calculating overall similarity between all candidate materials in the test set and the target materials; S53: ranking the candidate materials according to the similarity scores, and screening several candidate materials similar to the target materials and superior to the target materials in key traits as an optimal material set with breeding potential under the target environment.

2. The method of claim 1, wherein the known gene information is fused to the multi-trait coordinated screening method. In step S1, the specific steps are as follows: S11: constructing a BLUP model based on genotype information and multi-trait phenotype data of the training set; S12: constructing a kinship matrix based on individual multi-trait phenotype data by using a single-view kinship matrix construction method of the MVBLUP method to measure phenotype similarity between different materials; S13: obtaining genotype prediction values of materials at known genetic loci by using the BLUP model; S14: obtaining phenotype prediction values by using the BLUP model to perform ten-fold cross-validation on the training set.

3. The method of claim 2, wherein the known gene information is fused to the multi-trait collaborative screening method. The step S11 is provided with the genotype value of the known gene in the individual vector, the genotype value of the known gene in the individual matrix, the genotype value of the known gene in the individual fixed effect vector, the genotype value of the known gene in the individual random effect vector, the genotype value of the known gene in the individual residual error vector, and and are the variances related to the random effect and the residual error respectively, is the unit matrix; the expression of the BLUP model is: (1) In the step S12, the single view kinship matrix method of the MVBLUP is adopted to obtain the kinship matrix constructed based on the individual multi-trait phenotype data The expression of the random effect vector is: (2) wherein , .

4. The method of claim 1, wherein the known gene information is fused to the multi-trait collaborative screening method. In step S2, the specific steps are as follows: S21: performing t-test on different haplotype groups of each genetic locus based on multi-trait phenotype data of the training set to screen genetic loci with significant differences in phenotype; S22: judging the allelic effect direction of the haplotype in the genetic locus with significant differences in phenotype based on multi-trait phenotype data of the training set, and defining an excellent allele showing excellent performance in the target trait as an elite genotype.

5. A multi-trait collaborative screening system based on the multi-trait collaborative screening method of fusing known genetic information according to any one of claims 1 to 4. A material obtaining submodule is configured to obtain biological materials and proportionally divide the biological materials into a training set and a test set; A prediction submodule is configured to obtain a genotype prediction value of each material at a known genetic locus based on multi-trait phenotype data, genotype information and kinship information of the training set by using a BLUP model, and obtain a phenotype prediction value by ten-fold cross-validation of the training set; A gene screening submodule is configured to screen known genotypes with significant phenotype effects in the training set by t-test, and determine excellent genotypes with breeding significance; A target submodule is configured to select an individual with excellent multi-trait performance and carrying excellent genotypes as a target material, or select a widely planted commercial variety as the target material according to prior knowledge; A model weight submodule is configured to construct a TOPv2 model by combining the genotype prediction value, the phenotype prediction value of the training set and the information of the target material, and solve weight parameters of multi-trait phenotypes and known genes by using a weighted optimization algorithm; A phenotype screening submodule is configured to screen a candidate material closest to the target material in a specific environment in terms of multi-trait phenotype-genotype similarity between the candidate material and the target material in the test set.

6. A computer memory, characterized by: The computer program is stored in the memory and can be executed by the computer processor, and the computer program performs the multi-trait collaborative screening method fusing known gene information according to any one of claims 1-4.

Citation Information

Patent Citations

  • Specialized whole-genome SNP chip for comprehensive breeding of intramuscular fat and abdominal fat of chicken and application of specialized whole-genome SNP chip

    CN118147326A

  • Method for predicting phenotype of new environmental material by fusing multiple environmental factors

    CN118629491A