Method for identifying harmful mutation sites and total value of harmful mutations in a target plant and application

CN117423383BActive Publication Date: 2026-09-29AGRICULTURAL GENOMICS INSTITUTE AT SHENZHEN CHINESE ACADEMY OF AGRICULTURAL SCIENCES (SHENZHEN BRANCH GUANGDONG LABORATORY FOR LINGNAN MODERN AGRICULTURE) +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202310387389.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-03
Publication Date
2026-09-29
Estimated Expiration
2043-04-03

AI Technical Summary

Technical Problem

然而,目前针对二倍体马铃薯还没有全基因组预测模型的报道

Benefits of technology

[0106](1)本申请可全面鉴定全基因组有害突变(包含同义突变及非编码区域)及其效应大小,通过计算每个材料的有害突变位点及其效应大小来估算纯合有害突变总值和杂合有害突变总值,并准确估算其传给后代的有害突变总值,从而指导自交系构建的起始材料的选择及自交系的构建,了解有害突变对马铃薯等植物物种育种的影响,进一步开发新的全基因组预测模型预测其表型指导马铃薯等植物物种的选育。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117423383B_ABST
    Figure CN117423383B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of biological information, in particular to a method for identifying harmful mutation sites and total harmful mutation values of target plants and application. The method comprises the following steps: identifying evolutionarily conserved sites in the genome of the target plant and their conservation degree values; taking the target plant as a material, identifying SNP mutation sites of the target plant; according to the information of the evolutionarily conserved sites and the SNP mutation sites, identifying harmful mutation sites of the target plant and their harmful mutation effect degree values; and calculating total harmful mutation values of the target plant. The method can comprehensively identify harmful mutations and their effect sizes in the whole genome, and accurately estimate the total harmful mutation values passed to offspring, thereby guiding the selection of starting materials for constructing selfing lines and the construction of selfing lines. According to the harmful mutation information, a new whole-genome prediction model is further developed to predict phenotypes and guide the breeding of plants, and the plant breeding cycle is shortened.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of bioinformatics, specifically to a method and application for identifying harmful mutation sites and the total value of harmful mutations in a target plant. Background Technology

[0002] Potatoes are the most important tuber crop and a major source of food for people. However, potato breeding is slow. The Russet Burbank variety, developed 121 years ago, remains the most important processing potato variety despite its susceptibility to viruses, which reduces yield. Complex tetraploid genetics and asexual reproduction are the main obstacles to potato breeding. Diploid hybrid potato breeding based on inbred lines, abandoning the asexual tetraploid breeding method and instead improving seed-based sexual reproduction through diploid hybridization, will greatly accelerate the potato breeding process.

[0003] The first and most crucial step in diploid hybrid potato breeding is constructing highly homozygous inbred lines. However, potatoes accumulate a large number of heterozygous recessive or partially recessive harmful mutations during asexual reproduction. These harmful mutations, originally hidden at heterozygous sites, become homozygous during inbred line construction, exposing their harmful effects and reducing potato adaptability and fertility, thus causing inbreeding depression and making it extremely difficult to obtain highly homozygous inbred lines. Although harmful mutations with large effects can be screened out and neutral alleles retained during inbred line selection, since approximately 50% or more of the heterozygous harmful mutations are mosaicked in both genomes, extensive recombination is required to screen out harmful alleles, making it difficult to construct highly homozygous inbred lines with current technology. For example, 'Solyntus', after nine generations of continuous inbreeding, still has 20% heterozygous regions. Therefore, there is an urgent need for an effective method to predict harmful mutations and their effects, thereby guiding the construction of highly homozygous inbred lines, enabling precise selection of breeding materials, accelerating the potato breeding process, and improving breeding efficiency. Since the breeding of crops such as sweet potatoes and cassava, which have historically relied on asexual reproduction, also faces the same breeding challenges as potatoes, this method has universal applicability and the same application value in the breeding of crops such as sweet potatoes and cassava, which have historically relied on asexual reproduction or have accumulated a large number of harmful mutations.

[0004] Furthermore, potato breeding is a slow process, and breeding decisions can only be made based on phenotypes at the end of the growth stage. Therefore, developing whole-genome prediction models and then conducting whole-genome selection breeding at an early stage is an effective breeding method that can shorten the breeding cycle. However, there are currently no reports on whole-genome prediction models for diploid potatoes. Summary of the Invention

[0005] In view of this, the present invention provides a method and application for identifying harmful mutation sites and the total value of harmful mutations in target plants. This method can comprehensively identify harmful mutations across the entire genome and accurately estimate the total value of harmful mutations passed on to offspring, thereby guiding the selection of starting materials and the construction of inbred lines. Using the information on the total value of harmful mutations, new genome prediction models can be further developed to predict phenotypes and guide the breeding of materials such as potatoes.

[0006] To achieve the above-mentioned objectives, the present invention provides the following technical solution:

[0007] In a first aspect, the present invention provides a method for identifying harmful mutation sites and the total value of harmful mutations in a target plant, comprising the following steps:

[0008] Step (11) The genomes of species in the target plant's internal and external taxa are compared with the target plant's reference genome; based on the phylogenetic trees of the internal and external taxa and the results of the whole genome comparison, the evolutionary conserved sites and their degree of conservation in the target plant's genome are identified.

[0009] Step (12) Using the target plant as material, identify the SNP mutation sites of the target plant;

[0010] Step (13) Based on evolutionarily conserved sites and their degree of conservation and SNP mutation site information, identify the harmful mutation sites of the target plant and their degree of harmful mutation effects; the harmful mutation sites include homozygous harmful mutation sites and / or heterozygous harmful mutation sites.

[0011] Step (14) Calculate the total value of harmful mutations of the target plant using formulas (1) and (2) based on the genotype, harmful mutation sites, and degree of harmful mutation effect of the target plant. The total value of harmful mutations includes the total value of homozygous harmful mutations and / or the total value of heterozygous harmful mutations.

[0012]

[0013]

[0014] Among them, b hom The total number of homozygous harmful mutations is represented by i, where i represents the homozygous harmful mutation site and L(hom) represents the number of homozygous harmful mutation sites. Effect represents the magnitude of the effect of the harmful mutation site, i.e., the evolutionary conservation value of that site.

[0015] b het denoted as the total number of heterozygous harmful mutations, j represents the heterozygous harmful mutation site, and L(het) represents the number of heterozygous harmful mutation sites.

[0016] The inventors of this application discovered through research that:

[0017] (1) Materials with good growth and excellent phenotype (such as the parents of 'Solyntus') often have a lower total homozygous harmful mutation value, but a higher total heterozygous harmful mutation value and a higher total harmful mutation genetic value, which will pass on more harmful mutations to offspring. Highly homozygous offspring cannot survive, while surviving offspring often have a high degree of heterozygosity, thus making it impossible to construct highly homozygous inbred lines;

[0018] (2) Considering only the number of heterozygous harmful mutations that are not synonymous can lead to bias in the assessment and make it impossible to accurately predict the magnitude of the harmful mutation effect. For example, the number of heterozygous harmful mutations in potatoes calculated by SIFT can only predict harmful mutations that affect amino acid sequence changes, i.e., non-synonymous harmful mutations, and cannot be used to predict synonymous mutations and harmful mutations in non-coding regions. Using diploid potatoes with few harmful mutations that could be self-crossed according to SIFT calculations as starting materials, inbred lines were constructed through continuous self-crossing. Two of the materials were successfully constructed (homozygosity of 98.16% and 98.54%). However, the self-crossed progeny of the other two materials were already very weak before reaching high homozygosity and could not continue to self-cross, so highly homozygous inbred lines could not be constructed.

[0019] To address the two issues mentioned above, the inventors of this application have developed the aforementioned technical solution. To better identify harmful mutations, in a specific embodiment of this application, 38 Solanaceae materials were sequenced, and the genomes of 100 materials (92 species) were analyzed to identify evolutionarily conserved sites and their effect sizes across the entire genome. Combined with potato SNP information and the genotypes of 190 local potato varieties, harmful mutations and their effect sizes were identified. Comprehensive identification of harmful mutations across the entire genome, and estimation of the total number of harmful mutations passed on to offspring based on the genotypes and effect sizes of these mutations, not only further clarifies the impact of harmful mutations on the breeding of target plants such as potatoes, but also guides the selection of initial materials for inbred lines and the construction of inbred lines, as well as the development of new genome prediction models to predict phenotypes and guide the breeding of target plants such as potatoes.

[0020] As a preferred option, the total branch length of the species evolutionary tree is ≥2, and the differentiation time of the intragroup is ≥60 million years.

[0021] Preferably, the total branch length of the species evolutionary tree is ≥3, and the differentiation time of Solanaceae is ≥70 million years.

[0022] In the embodiments provided by the present invention, the total branch length of the species evolutionary tree is ≥4, and the differentiation time of Solanaceae is ≥80 million years.

[0023] In a specific embodiment provided by the present invention, the total branch length of the species evolution tree is 4.05, and the differentiation time of Solanaceae is 80 million years.

[0024] Preferably, in step (11), the software used to identify the evolutionarily conserved sites and their degree of conservation in the genome of the target plant includes, but is not limited to, GERP software or LIST software, as long as the software can calculate the size of harmful mutations and their effects.

[0025] As a preferred option, step (11) includes: calculating the conservation value of conserved sites based on the species evolutionary tree and whole genome alignment results of the in-group and out-group, and identifying the evolutionary conserved sites and their conservation values ​​in the genome of the target plant based on the screening threshold of the conservation value of the evolutionary conserved sites.

[0026] Preferably, the top 2.4% of conserved sites in the whole genome are selected as the screening threshold for the degree of conservation.

[0027] In a specific embodiment of the present invention, the software used to identify evolutionarily conserved sites and their degree of conservation in the genome of the target plant is GERP software, and steps (11) and (13) include:

[0028] Step (11): Based on the species evolutionary tree and whole genome alignment results of the in-group and out-group, calculate the GERP value of the conserved sites. With GERP ≥ 2.75 as the threshold, which is about 2.4% of the whole genome, the evolutionary conserved sites and their GERP values ​​in the target plant genome are identified.

[0029] Step (13): Based on evolutionarily conserved sites and their GERP values ​​and SNP mutation site information, identify the harmful mutation sites and their GERP values ​​of the target plant; the harmful mutation sites include homozygous harmful mutation sites and / or heterozygous harmful mutation sites.

[0030] In this embodiment of the invention, the target plant needs to be bred using highly homozygous inbred lines, including but not limited to self-incompatible plants, asexually reproducing plants, or cross-pollinated plants. The target plant has accumulated a large number of harmful mutations.

[0031] As a preferred choice, self-incompatible plants include, but are not limited to, at least one of sweet potato, rye, rapeseed, sunflower, beet, cabbage or kale;

[0032] As a preferred choice, asexually reproducing plants include, but are not limited to, at least one of the following: potato, cassava, arrowroot, sugarcane, castor bean, bamboo, lily, water lily, rose, and rose.

[0033] As a preferred option, cross-pollinating plants include, but are not limited to, at least one of tobacco, hemp, alfalfa, clover, and sweet clover.

[0034] In this embodiment of the invention, the number of species in the ingroup and the outgroup is ≥10; for example, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, etc. The more species, the higher the accuracy.

[0035] Preferably, the number of species in the ingroup and the outgroup is ≥50.

[0036] More preferably, the number of species in the ingroup and the number of species in the outgroup are ≥90.

[0037] In a specific embodiment provided by the present invention, the number of species in the ingroup and the outgroup is 92.

[0038] As a preferred option, the number of strains of both inogroup and outogroup species is ≥10.

[0039] Preferably, the number of strains of in-group and out-group species is ≥50.

[0040] In a specific embodiment provided by the present invention, the number of strains of ingroup species and outgroup species is 100.

[0041] In a specific embodiment of the present invention, the target plant is potato, the inner group species are Solanaceae species other than potato, and the outer group species are Convolvulaceae species.

[0042] In this embodiment of the invention, the number of target plant strains in step (12) is not limited. One or more strains can be used to identify the SNP mutation sites of the target plant.

[0043] Preferably, in step (12), the software used to identify the SNP mutation sites of the target plant includes at least one of GATK, samtools, bedtools, or plink.

[0044] In this embodiment of the invention, the software used to identify the SNP mutation sites of the target plant, such as potato, is GATK, and the following conditions are used for filtering: GATK "QD<2.0||FS>60.0||MQ<40.0||SOR>3.0||MQRankSum<-12.5||ReadPosRankSum<-8.0"; VCFtools (v0.1.16): "-minDP 4--maxDP 100--minGQ 10--minQ 30--max-missing 0.5--min-alleles 2--max-alleles 2".

[0045] In this embodiment of the invention, step (13) includes: obtaining the overlapping sites of evolutionarily conserved sites and their degree of conservation and SNP mutation site information, identifying the overlapping sites as harmful mutation sites of the target plant, and the degree of evolutionary conservation of the site is the magnitude of the harmful mutation effect.

[0046] In this embodiment of the invention, step (11) includes: performing whole-genome alignment of the genomes of the species in the inner and outer taxa of the target plant with the reference genome of the target plant, extracting the 4D loci of the reference genome of the target plant, and constructing the species evolutionary tree of the inner and outer taxa; and identifying the evolutionary conserved loci and their degree of conservation in the genome of the target plant based on the species evolutionary tree of the inner and outer taxa and the whole-genome alignment results.

[0047] Secondly, the present invention provides a method for selecting starting materials for a target plant inbred line, comprising the following steps:

[0048] Step (21) The genomes of species in the target plant's internal and external taxa are compared with the target plant's reference genome; based on the phylogenetic trees of the internal and external taxa and the results of the whole genome comparison, the evolutionary conserved sites and their degree of conservation in the target plant's genome are identified.

[0049] Step (22) Using the target plant as material, the SNP mutation sites of the target plant were identified;

[0050] Step (23) Based on evolutionarily conserved sites and their degree of conservation and SNP mutation site information, the harmful mutation sites of the target plant and their degree of harmful mutation effects are identified; the harmful mutation sites include homozygous harmful mutation sites and / or heterozygous harmful mutation sites.

[0051] Step (24) Calculate the total value of harmful mutations of the target plant using formulas (1) and (2) based on the genotype, harmful mutation sites, and degree of harmful mutation effect of the target plant. The total value of harmful mutations includes the total value of homozygous harmful mutations and / or the total value of heterozygous harmful mutations.

[0052] Step (25) Calculate the total genetic value of harmful mutations in the target plant material to be selected using formula (3);

[0053]

[0054]

[0055] b genetic = b hom + b het ×0.5 Formula (3)

[0056] Among them, bhom The total number of homozygous harmful mutations is represented by i, where i represents the homozygous harmful mutation site and L(hom) represents the number of homozygous harmful mutation sites. Effect represents the magnitude of the effect of the harmful mutation site, which is the value of evolutionary conservation here.

[0057] b het denoted as the total value of heterozygous harmful mutations, j represents the heterozygous harmful mutation sites, and L(het) represents the number of heterozygous harmful mutation sites;

[0058] b genetic This represents the total ancestry value of harmful mutations;

[0059] Step (26) Based on the total genetic value of harmful mutations in the target plant material to be selected, determine whether to use the target plant material to be selected as the starting material for the construction of inbred lines.

[0060] Preferably, the number of species of target plant materials to be selected is ≥10. For example, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, etc.

[0061] In one embodiment of the present invention, step (26), specifically determining whether to use the target plant material to be selected as the starting material for the construction of inbred lines, involves:

[0062] The selected target plant materials are ranked according to the total genetic value of harmful mutations.

[0063] When the percentage of the target plant material to be selected is within 40% (ranked from smallest to largest), it is used as the starting material for constructing inbred lines; for example, 1%, 5%, 10%, 15%, 20%, 25%, 30%, 35%, 38%, 40%, etc. The higher the ranking, the easier it is to successfully construct highly homozygous inbred lines.

[0064] If the percentage of the selected target plant material in the ranking is higher than 90%, the material will fail to be used as the starting material for constructing an inbred line.

[0065] And / or, when the percentage of the target plant material ranked from largest to smallest is above 60%, the target plant material is used as the starting material for constructing inbred lines. Examples include 60%, 62%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 100%, etc. The lower the ranking, the easier it is to successfully construct highly homozygous inbred lines.

[0066] In another embodiment of the present invention, step (26), specifically determining whether to use the target plant material to be selected as the starting material for the construction of inbred lines, involves:

[0067] When the total genetic value of harmful mutations in the target plant material to be selected is ≤57000, the target plant material to be selected will be used as the starting material.

[0068] As a preferred option, when the total genetic value of harmful mutations in the target plant material to be selected is ≤55000, the target plant material to be selected is used as the starting material.

[0069] Preferably, when the total genetic value of harmful mutations in the target plant material to be selected is ≤53000, the target plant material to be selected is used as the starting material.

[0070] More preferably, when the total genetic value of harmful mutations in the target plant material to be selected is ≤52000, the target plant material to be selected is used as the starting material. When the total genetic value of harmful mutations in the target plant material to be selected is >57000, the target plant material to be selected is not used as the starting material.

[0071] Thirdly, the present invention provides a method for constructing a target plant inbred line, comprising the following steps:

[0072] Step (31) Select starting materials using the above method, and perform self-crossing on the starting materials to obtain offspring;

[0073] Step (32) Using the offspring as the target plant material to be selected, repeat step (31) to obtain the target plant inbred line.

[0074] Preferably, the number of times step (31) is repeated is ≥2.

[0075] Preferably, step (31) is repeated 3 to 8 times.

[0076] More preferably, the number of times step (31) is repeated is 4 to 6.

[0077] Preferably, the homozygosity of the inbred line is not less than 95%, such as 95%, 96%, 97%, 98%, 99%, or 100%.

[0078] Fourthly, this invention provides a phenotypic prediction method based on the whole genome of a target plant, comprising the following steps:

[0079] Step (41) The genomes of species in the target plant and species in the out-of-group are compared with the reference genome of the target plant. Based on the phylogenetic tree of the in-group and out-of-group and the results of the whole genome comparison, the evolutionary conserved sites and their degree of conservation in the genome of the target plant are identified.

[0080] Step (42) Using the target plant as material, identify the SNP mutation sites of the target plant;

[0081] Step (43) Based on evolutionarily conserved sites and their degree of conservation and SNP mutation site information, identify the harmful mutation sites of the target plant and their degree of harmful mutation effect; the harmful mutation sites include homozygous harmful mutation sites and / or heterozygous harmful mutation sites;

[0082] Step (44) Obtain a genome-wide phenotype prediction model through training:

[0083] Obtain a similar population of the target plant material to be tested as a training population, and use the genotype and phenotype of the training population as the training set (preferably, the number of materials in the training population is greater than or equal to 100, and the larger the number, the more accurate the model training).

[0084] Based on the genotype of the target plant, the harmful mutation site, and the degree of its harmful mutation effect, the total harmful mutation value of each material in the training population is calculated using formulas (1) and (2); the total harmful mutation value includes the total homozygous harmful mutation value and / or the total heterozygous harmful mutation value.

[0085] The kinship matrix G of the training set is calculated based on the genotypes of the training population using the grm function of the qgg package;

[0086] Based on the mixed linear model using the greml function in the qgg package of R programming, the total homozygous and heterozygous harmful mutation values ​​of the training population are used as fixed effect factors, the kinship matrix is ​​used as random effect factors, and the observed phenotype of the training population is used as y. The model is then trained to obtain μ and α corresponding to formula (4). hom α het Specific numerical values ​​(μ, α) for different species, population types, phenotypes, etc. hom α het The specific values ​​will vary and need to be trained based on the actual population and phenotype to obtain a genome-wide phenotypic prediction model for the target plant.

[0087]

[0088]

[0089] y = μ + α hom b hom +α het b het Formula (4) +u+e

[0090] Among them, b homThe total number of homozygous harmful mutations is represented by i, where i represents the homozygous harmful mutation site, and L(hom) represents the number of homozygous mutant sites. Effect represents the magnitude of the effect of the harmful mutation site, which is the value of evolutionary conservation here.

[0091] b het denoted as the total value of heterozygous harmful mutations, j represents the heterozygous harmful mutation sites, and L(het) represents the number of heterozygous harmful mutation sites;

[0092] y represents the phenotypic value; μ represents the intercept, which is the average phenotypic value when there is no harmful mutation; α hom The coefficient representing the total homozygous harmful mutation value is a fixed effect of the total homozygous harmful mutation value; α het represents the coefficient of total heterozygous harmful mutation, which is the fixed effect of the total heterozygous harmful mutation; u represents the additive genetic breeding value, where The random effects are calculated based on the additive kinship matrix G; G is the additive kinship matrix between materials, calculated from the genotypes using the qgg package, and e represents the random residuals;

[0093] Step (45) Obtain the phenotypic values ​​of the test material based on the whole-genome phenotypic prediction model and the genotype of the test material:

[0094] The total number of harmful mutations is calculated according to formulas (1) and (2) and the genotype of the material to be tested; the total number of harmful mutations includes the total number of homozygous harmful mutations and / or the total number of heterozygous harmful mutations.

[0095] Based on the grm function of the qgg package and the genotype of the test material, calculate the kinship matrix between the test material and the training population;

[0096] Based on the total harmful mutation value of the species to be tested and its phylogenetic relationship matrix with the training population, the phenotypic value of the material to be tested is obtained by formula (4).

[0097] Preferably, the group includes F n At least one of a segregated population, an F1 population, or a natural population, wherein n ≥ 2. For example, F1... n Separating populations include F2, F3, F4, F5, etc.

[0098] In this embodiment of the invention, the group includes only F. n In a group, F n The population has no fewer than 500 genotypes.

[0099] In embodiments of the present invention, when the population includes only the F1 population or a natural population, the population includes multiple different genotypes, with ≥2 genotype types, preferably ≥10, and more preferably ≥50.

[0100] In embodiments of the present invention, the phenotype includes at least one of yield, plant height, tuber size, number of tubers, flowering time, and disease resistance.

[0101] In a specific embodiment of the present invention, the phenotype includes at least one of yield, plant height, and tuber size.

[0102] In this embodiment of the invention, the formula for calculating the kinship matrix G is as follows:

[0103]

[0104] Where X is the genotype matrix of the population, that is, the genotype matrix of the number of materials multiplied by the number of mutation sites; X T It is the transpose of the X matrix; n is the number of mutation sites; p g It represents the gene frequency of mutation site g in the population.

[0105] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0106] (1) This application can comprehensively identify whole-genome harmful mutations (including synonymous mutations and non-coding regions) and their effect sizes. By calculating the harmful mutation sites and their effect sizes of each material, the total value of homozygous harmful mutations and the total value of heterozygous harmful mutations can be estimated, and the total value of harmful mutations passed on to offspring can be accurately estimated. This can guide the selection of starting materials for inbred line construction and the construction of inbred lines, understand the impact of harmful mutations on the breeding of plant species such as potatoes, and further develop new whole-genome prediction models to predict their phenotypes and guide the breeding of plant species such as potatoes.

[0107] (2) This application calculates and obtains genome-wide harmful mutations and estimates the total number of harmful mutations passed from the starting material to the offspring. This allows for the effective selection of starting materials for inbred line construction, ensuring that highly homozygous inbred lines can be obtained after continuous self-pollination of these materials, thus greatly accelerating the construction of highly homozygous inbred lines for plant species such as potatoes. For example, during the inbred line construction process, the total number of harmful mutations passed from each F2 plant to F3, as well as the total number of homozygous and heterozygous harmful mutations, can be calculated. Materials with a smaller total number of harmful mutations passed to the offspring are more likely to successfully construct inbred lines after continuous self-pollination. Then, F2 plants with a small total number of harmful mutations and a low heterozygosity rate can be selected for further self-pollination, and so on for F3 and F4, accelerating the inbred line construction process. This application avoids situations like the 'Solyntus' case, where nine generations of continuous self-pollination failed to yield highly homozygous inbred lines, resulting in wasted breeding time, land, management resources, and increased breeding costs. This application can improve the success rate of constructing highly homozygous inbred lines, effectively solve the problems in the key step (constructing highly homozygous inbred lines) in the breeding of hybrid potatoes and other plant species, and provide an effective technical solution for diploid hybrid potatoes and other plant species.

[0108] (3) The harmful mutations identified in this application can guide the selection of parent plants such as hybrid potatoes, so that hybrids are exposed to fewer harmful mutations and have better growth potential.

[0109] (4) This application utilizes harmful mutation information to develop a novel genome prediction model. This application uses the total homozygous and heterozygous harmful mutation values ​​of plant species such as potato as factors in the whole-genome prediction model to train and optimize the model for predicting yield, plant height, and tuber number. Utilizing harmful mutation information and genomic SNP information from plant species such as potato, as long as seedling materials are sequenced, the model can predict complex traits such as yield, plant height, and tuber number, enabling early plant selection and breeding decisions, significantly shortening the breeding cycle of plant species. Compared to traditional mixed linear models, the model in this application significantly improves the accuracy of the whole-genome prediction model. This model can be used to guide breeding decisions through genome selection, improve the efficiency of early selection of plant species, shorten the breeding cycle, save costs, and accelerate the breeding process of hybrid plant species. Attached Figure Description

[0110] Figure 1 This application illustrates the technical approach used in this application;

[0111] Figure 2 This application presents the phylogenetic tree of species constructed in this application;

[0112] Figure 3A The phenotypic distribution of individual plants in the population is shown.

[0113] Figure 3B The correlation between the total number of homozygous and heterozygous harmful mutations in individual plants of the population and their phenotypes;

[0114] Figure 3C This demonstrates the accuracy of the prediction model in this application. Detailed Implementation

[0115] This invention discloses a method and application for identifying harmful mutation sites and the total value of harmful mutations in target plants. Those skilled in the art can refer to the content of this document and appropriately modify the process parameters to achieve the desired result. It should be particularly noted that all similar substitutions and modifications are obvious to those skilled in the art and are considered to be included in this invention. The method and application of this invention have been described through preferred embodiments. Those skilled in the art can obviously modify or appropriately change and combine the methods and applications described herein without departing from the content, spirit, and scope of this invention to realize and apply the technology of this invention.

[0116] Terminology Explanation:

[0117] Outgroup species: Species that are most closely related to other taxa.

[0118] Whole genome alignment: This refers to sequence alignment performed at the whole genome level. The whole genome refers to the genetic information of all the base sequences carried by an organism.

[0119] 4D site: tetradonic degenerate site, which refers to a site where the amino acid remains the same regardless of how the nucleotide is mutated.

[0120] Evolutionary tree: In biology, it is used to represent the evolutionary relationships and scales between species. Biologists and evolutionary theorists arrange organisms on a branching tree-like diagram based on their kinship, clearly representing the evolutionary process and kinship relationships. Each leaf node in the evolutionary tree represents a species, and if each edge is assigned an appropriate weight, the shortest distance between two leaf nodes represents the degree of difference between the corresponding two species.

[0121] SNP mutation sites: Single nucleotide polymorphism sites, mainly refer to the diversity of DNA sequences caused by variations in a single nucleotide at the genomic level. They are numerous and highly polymorphic. Theoretically, each SNP site can have four different forms of variation (substitution (transition or transversion), insertion, or deletion). In other words, replacing a base at this position with another base creates a SNP mutation site.

[0122] English-Chinese bilingual edition:

[0123]

[0124] The reagents, instruments, or biological materials used in this invention can all be obtained through commercial channels.

[0125] Among them, the Solanaceae materials came from the U.S. Germplasm Bank and the Southwest China Germplasm Bank.

[0126] Potato materials such as RH, C10-20, E86-69, PG6359, and RH1015 are all publicly available and belong to the category of biomaterials that are readily available to the public.

[0127] RH and PG6359 are previously reported materials (Clot et al., 2020; Peterson et al., 2016; Zhang et al., 2019). RH1015 is a descendant of RH after four generations of self-pollination.

[0128] C10-20 and E86-69 are BC1 cells of materials CIP701165 and CIP703767 with the Sli gene introduced, respectively (Zhang et al., 2021).

[0129] The type and source of the materials used do not affect the implementation of this invention. Other materials that meet the technical requirements and are publicly available can be used to implement this invention.

[0130] The present invention will be further illustrated below with reference to the embodiments:

[0131] Example 1: Detection of harmful mutations in potatoes

[0132] 1. Experimental Method:

[0133] Table 1. Names of Potato Materials

[0134]

[0135]

[0136]

[0137] Note: " / " indicates unknown.

[0138] Thirty-eight materials (32 species) were cultivated, and leaf DNA was extracted and library constructed according to the standard procedures of Pacific Biosciences (PacBio). Sequencing was performed using the circular consensus sequencing (CCS) mode on the PacBio Sequel II platform, and high-fidelity reads were extracted for genome assembly. An average of 25X sequencing data was obtained per material. Whole genome assembly was then performed using hifiasm software (Cheng et al., 2021) to obtain high-quality genomes. Combined with 62 published reference genomes of Solanaceae and Convolvulaceae plants, a total of 100 genome sequences from Solanaceae and adjacent Convolvulaceae materials were obtained (Table 1), encompassing 92 species. Whole genome alignment was performed using cactus software (Armstrong et al., 2020). Alignment to the potato reference genome DMV6 yielded sequences ranging in size from 32 Mb to 446 Mb. Alignment files of the 4d loci of the potato reference genome were extracted from the whole genome multiple sequence alignments of Solanaceae and Convolvulaceae species. Using IQ-TREE (Nguyen et al., 2015) software, a high-density, high-evolutionary-scale phylogenetic tree of 100 materials (92 species) was constructed with 5 Convolvulaceae materials as outgroups. Figure 2The total branch length of this phylogenetic tree is 4.05, and the divergence time of Solanaceae is 80 MYA (MYA represents millions of years). Compared with the previously reported maximum branch length of 0.7 (Tang et al., 2022), the phylogenetic tree of 92 species obtained in this example has a total branch length 5.8 times longer than previously reported phylogenetic trees. This whole-genome multiple sequence alignment and phylogenetic tree provides a sufficient evolutionary scale for identifying and quantifying evolutionarily conserved sites and harmful mutation sites.

[0139] Based on the phylogenetic trees and whole-genome alignments of Solanaceae and Convolvulaceae species, GERP values ​​were calculated using GERP software (Davydov et al., 2010) to represent the degree of conservation of each locus in the potato genome. Using a GERP value ≥ 2.75 as the threshold, 17,362,955 evolutionarily conserved loci were identified, representing 2.4% of the potato genome. Of all evolutionarily conserved loci, 36% were located in non-coding regions.

[0140] To identify harmful mutations in diploid potatoes, we used a potato population consisting of 179 local varieties and 2 inbred lines (Table 2). Mutation sites were identified in this population, ultimately identifying 58,597,787 SNPs with an average SNP heterozygosity of 0.02. Mutations at evolutionarily conserved sites are considered harmful mutations; therefore, combining information from evolutionarily conserved sites and the potato population's mutation sites allows for the identification of harmful mutation sites in potatoes.

[0141] The identification method for mutation sites included: using BWA software to align the DNA-seq of 181 potato materials (including 179 local varieties) to the reference genome S. tuberosum group Phureja DMv6.1; removing duplicate reads using SAMtools software; and then using GATK to identify mutation sites in the population, filtering with the following conditions: GATK "QD<2.0||FS>60.0||MQ<40.0||SOR>3.0||MQRankSum<-12.5||ReadPosRankSum<-8.0"; VCFtools (v0.1.16): "-minDP 4--maxDP 100--minGQ 10--minQ 30--max-missing 0.5--min-alleles 2--max-alleles 2", thus obtaining 58,597,787 potato population SNP mutation sites.

[0142] 2. Experimental Results

[0143] By combining evolutionarily conserved sites and SNPs from 190 potato endemic populations, we identified 367,499 harmful mutation sites in potato endemic populations, accounting for 0.6% of all mutation sites. The study found that rare variants were more likely to be identified as harmful mutations, a result consistent with population purification selection often controlling the frequency of harmful mutations. Furthermore, we found that sites with higher degrees of harmful mutations tended to be enriched in coding regions, synonymous mutations, and other functional regions, indicating accurate identification of harmful mutation sites and their severity.

[0144] Example 2: Method for Selecting Starting Materials for the Construction of Self-crossing Lines

[0145] 1. Experimental Method:

[0146] Experimental materials: 179 local potato varieties.

[0147] Local potato varieties are an important source of starting material for inbred line construction, but this invention is not limited to them; other publicly available breeding materials may also be used. Due to recessive or partially recessive effects, heterozygous harmful mutations often do not express or only partially express their harmful effects. Homozygous harmful mutations fully express their harmful effects. During inbred line construction, through continuous self-pollination, heterozygous harmful mutations become homozygous, and the hidden harmful effects become apparent, severely reducing the potato's adaptability and rendering it non-homozygous or sterile. For example, previous studies have shown that *solyntus* is a single plant that has been self-pollinated for nine generations, but its genome still contains 20% heterozygosity. Genetic burden includes both hidden and exposed burdens, representing the burden passed on to offspring. Similarly, the total genetic value of harmful mutations passed on to offspring should include the total value of both hidden and exposed harmful mutations, representing the total value of harmful mutations passed on to offspring. Predicting the total value of harmful mutations passed on to offspring can help estimate whether the material can successfully construct an inbred line.

[0148] We first identified homozygous and heterozygous harmful mutation sites for each material. Based on the genotype of each local variety, we calculated the number and size of homozygous and heterozygous harmful mutations, thereby calculating the total value of homozygous and heterozygous harmful mutations. The total value of homozygous harmful mutations is obtained by summing the GERP scores corresponding to the homozygous harmful mutation alleles of each material (see formula (1)); the total value of heterozygous harmful mutations is obtained by summing the GERP scores corresponding to the heterozygous harmful mutation alleles (see formula (2)).

[0149] Then, we used formula (3) to calculate the total genetic value of harmful mutations that each local variety may pass on to its offspring. The smaller the total genetic value of harmful mutations, the smaller the harmful mutations passed on to the offspring. Therefore, the smaller the harmful mutations in homozygous inbred lines constructed from these materials, the easier it is to successfully construct highly homozygous inbred lines.

[0150] The inventors discovered that materials with a low total value of harmful mutations tend to have a lower heterozygous total value of harmful mutations, a higher homozygous total value of harmful mutations, and a higher exposed total value of harmful mutations. In other words, materials with a low total value of harmful mutations passed on to offspring have a higher homozygous total value of harmful mutations and are exposed to more harmful mutations, leading to lower fitness and weaker growth. Therefore, starting materials for constructing inbred lines based on genome-guided selection often exhibit weak growth, contrary to phenotypic guidance, i.e., counterintuitive.

[0151]

[0152]

[0153] b genetic = b hom + b het ×0.5 Formula (3)

[0154] Among them, b hom The total number of homozygous harmful mutations is represented by i, where i represents the homozygous harmful mutation site and L(hom) represents the number of homozygous harmful mutation sites. Effect represents the magnitude of the effect of the harmful mutation site, which is the value of evolutionary conservation here.

[0155] b het denoted as the total value of heterozygous harmful mutations, j represents the heterozygous harmful mutation sites, and L(het) represents the number of heterozygous harmful mutation sites;

[0156] b genetic This represents the total trait value of harmful mutations, and 0.5 is the probability that a heterozygous harmful mutation site will be passed on to offspring.

[0157] 2. Experimental Results

[0158] Table 2 Number and total number of harmful mutations

[0159]

[0160]

[0161]

[0162]

[0163]

[0164] RH, C10-20, E8669, and PG6359 are four diploid heterozygous potatoes. Taking these four materials as examples, the calculated total genotype of harmful mutations and the self-crossing verification results are shown below:

[0165] Table 3

[0166]

[0167] It is evident that the total genetic values ​​of harmful mutations are relatively high in RH and C10-20, at 60,254 and 57,419 respectively, both exceeding 57,000. We predict that highly homozygous inbred lines cannot be successfully constructed using RH and C10-20 as starting materials.

[0168] PG6359 and E8669 are relatively few, with 49,927 and 51,648 respectively, both below 52,000. We predict that highly homozygous inbred lines can be successfully constructed by using PG6359 and E8669 as starting materials and continuously self-pollinating for 4 to 6 generations.

[0169] Furthermore, using RH and C10-20 as starting materials, continuous self-pollination failed to produce offspring before reaching high homozygosity, thus preventing the successful construction of highly homozygous inbred lines. Using PG6359 and E8669 as starting materials, continuous self-pollination resulted in inbred lines with homozygosity of 98.16% and 98.54%, respectively.

[0170] As can be seen, the results of the self-crossing experiment are consistent with the predicted results.

[0171] Example 3: Accuracy Verification of Starting Material Selection Method

[0172] Taking RH1015 material as an example, predict whether it is possible to successfully construct an inbred line.

[0173] 1. Experimental Grouping

[0174] Experimental group: The total genetic value of harmful mutations was calculated using the method in Example 2 to predict whether inbred lines could be successfully constructed.

[0175] Control group: The technique of Zhang et al. (Zhang, C., Yang, Z., Tang, D., Zhu, Y., Wang, P., Li, D., Zhu, G., Xiong, X., Shang, Y., Li, C., and Huang, S. (2021). Genome design of hybrid potato. Cell 184, 3873–3883.) was used to predict whether inbred lines could be successfully constructed.

[0176] 2. Prediction Results

[0177] Experimental group: The results showed that the number of heterozygous harmful mutations in RH1015 was 3,913, and the total genetic value of harmful mutations was 85,640. The total genetic value of harmful mutations in RH1015 was higher than that of E8669 and PG6359 (49,927 and 51,648, respectively), and higher than 57,000. We predict that highly homozygous inbred lines cannot be successfully constructed using RH1015 as the starting material.

[0178] Comparative group: According to the technique of Zhang et al., the number of heterozygous harmful mutations in RH1015 is lower than that in E8669 and PG6359 (49,538 and 55,858, respectively), suggesting that inbred lines can be constructed.

[0179] 3. Field verification results

[0180] The study found that RH1015 cannot produce offspring through self-pollination and cannot be used to construct highly homozygous inbred lines.

[0181] 4. Conclusion

[0182] It is evident that the prediction method of this invention is more comprehensive and accurate.

[0183] Example 4: Method for Constructing Inbred Lines

[0184] According to the selection method in Example 2, the selected materials were continuously self-pollinated for 4 to 6 generations. The homozygosity and total genetic value of harmful mutations were calculated for each generation. Individual plants with low total genetic value and high homozygosity were selected for further self-pollination. Finally, highly homozygous inbred lines can be constructed.

[0185] Example 5: Whole Genome Prediction Method

[0186] 1. Experimental Method:

[0187] Experimental materials: A population consisting of 1064 F2 individual plants, which were obtained by crossing A6-26 (maternal parent) and E4-63 (paternal parent) to obtain F1, and then self-crossing the F1 to obtain the F2 population.

[0188] ① Control group method:

[0189] A conventional whole-genome prediction model was used, with the additive sex-linked kinship matrix G calculated based on genotypes and used as the random effects factor in the model. The model was trained using the greml function in the qgg package of R programming, employing the sex-linked matrix and phenotypes. Specifically, 1064 samples were divided into five sets using cross-validation: four sets for training and one set for validation. Training and validation were performed according to the whole-genome prediction model's framework, repeated 20 times, to calculate the model's prediction accuracy. Accuracy was calculated as the square of the correlation coefficient (r) between predicted and observed phenotypic values. 2 ).

[0190] y = μ + u + e

[0191] y represents the phenotypic value, μ represents the intercept, the total phenotypic mean, and u represents the additive genetic breeding value, where... The random effects are calculated based on the additive affinity matrix G; G is the additive affinity matrix between materials, calculated from genotypes using the qgg package in R programming, and e represents the random residuals.

[0192] ② Experimental group method:

[0193] Based on the parental genotype and population genotype, the total homozygous and heterozygous harmful mutation values ​​for each material are calculated according to formulas (1) and (2). The total harmful mutation value is added as a fixed-effect factor in the traditional model. Using the greml function in the qgg package of R programming, formula (4) is used as the model skeleton for training and cross-validation. Specifically, the 1064 materials are divided into 5 sets according to cross-validation, with 4 sets as the training set and 1 set as the test set. The total harmful mutation value is used as the fixed-effect factor, and the kinship matrix is ​​used as the random-effect factor. The greml function in the qgg package of R programming is used for model training and phenotypic prediction of the test set. This is repeated 20 times to calculate the model prediction accuracy. The accuracy is the square of the correlation coefficient between the predicted phenotypic value and the observed phenotypic value (r). 2 ).

[0194] The formula after incorporating harmful mutation information is as follows:

[0195]

[0196]

[0197] y = μ + α hom b hom +α het b het Formula (4) +u+e

[0198] Among them, b hom The total number of homozygous harmful mutations is represented by i, where i represents the homozygous harmful mutation site, and L(hom) represents the number of homozygous mutant sites. Effect represents the magnitude of the effect of the harmful mutation site, which is the value of evolutionary conservation here.

[0199] b het denoted as the total value of heterozygous harmful mutations, j represents the heterozygous harmful mutation sites, and L(het) represents the number of heterozygous harmful mutation sites;

[0200] y represents the phenotypic value; μ represents the intercept; α is the mean phenotypic value when there is no harmful mutation; α is the mean phenotypic value. homThe coefficient representing the total homozygous harmful mutation value is a fixed effect of the total homozygous harmful mutation value; α het represents the coefficient of total heterozygous harmful mutation, which is the fixed effect of the total heterozygous harmful mutation; u represents the additive genetic breeding value, where The random effects are calculated based on the additive kinship matrix G; G is the kinship matrix between materials, calculated from the genotypes using the qgg package, and e represents the random residual.

[0201] 2. Experimental Results

[0202] Model comparison method: The coefficient of determination between the predicted phenotype and the observed (true) phenotype is the square of the correlation coefficient (r²). 2 The accuracy of model predictions is represented by a ) to allow for comparison of models.

[0203] We found that materials with higher homozygous harmful mutation rates tended to have lower yields, plant heights, and tuber numbers, with correlation coefficients of -0.33, -0.25, and -0.22, respectively. This indicates that harmful mutations are associated with these agronomic traits, and harmful mutation factors should be considered in potato breeding. Therefore, the experimental group added homozygous harmful mutation total value and heterozygous harmful mutation total value factors to the existing genome-wide prediction model to predict potato yield, plant height, and tuber size.

[0204] Figures 3A-3C The analysis results showed that, with 20 random shufflings in the control group, the prediction accuracy (r) was improved by 5 times cross-validation. 2 The mean values ​​were 0.183, 0.212, and 0.248, respectively. The experimental group underwent 20 random shufflings, and the phenotypic prediction accuracy (r) after 5-fold cross-validation was... 2 The average values ​​were 0.264, 0.250, and 0.289, respectively, representing improvements in prediction accuracy of 44.6% (yield), 17.8% (plant height), and 16.4% (tuberous size) compared to the control group. This indicates that incorporating the total value of harmful mutations into the whole-genome prediction model can significantly improve the accuracy of phenotypic prediction.

[0205] Therefore, the prediction model of the experimental group in this application has higher accuracy in predicting the phenotype of potatoes.

[0206] In summary, the harmful mutation information detected in this application can not only guide the selection of hybrid potato materials, such as the starting materials for the construction of inbred lines, but also guide the improvement of potato agronomic traits in the early stages, greatly shortening the breeding cycle and promoting hybrid potato breeding.

[0207] The above description is only a preferred embodiment of the present invention. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.

Claims

1. A method for identifying harmful mutation sites and the total value of harmful mutations in a target plant, characterized in that, Includes the following steps: Step (11) Perform whole-genome alignment of the genomes of species in the target plant and species in the out-of-group with the reference genome of the target plant; based on the phylogenetic tree of the in-group and out-of-group and the whole-genome alignment results, identify the evolutionary conserved sites and their degree of conservation in the genome of the target plant. Step (12) Using the target plant as material, identify the SNP mutation sites of the target plant; Step (13) Based on evolutionarily conserved sites and their degree of conservation and SNP mutation site information, identify the harmful mutation sites of the target plant and their degree of harmful mutation effect; the harmful mutation sites include homozygous harmful mutation sites and / or heterozygous harmful mutation sites; Step (14) Calculate the total value of harmful mutations in the target plant using formulas (1) and (2) based on the genotype, harmful mutation sites, and degree of harmful mutation effect of the target plant. The total value of harmful mutations includes the total value of homozygous harmful mutations and / or the total value of heterozygous harmful mutations. Among them, b hom The total number of homozygous harmful mutations is represented by i, where i represents the homozygous harmful mutation site, and L(hom) represents the number of homozygous harmful mutation sites; Effect represents the degree of harmful mutation effect of the harmful mutation site. b het denoted as the total number of heterozygous harmful mutations, j represents the heterozygous harmful mutation sites, and L(het) represents the number of heterozygous harmful mutation sites; The target plant is a plant of the Solanaceae family.

2. The identification method according to claim 1, characterized in that, The total branch length of the species evolutionary tree is ≥2, and the differentiation time of the intragroups is ≥60 million years.

3. The identification method according to claim 1, characterized in that, In step (11), the software used to identify evolutionarily conserved sites and their degree of conservation in the genome of the target plant includes GERP software or LIST software.

4. The identification method according to claim 1, characterized in that, Step (11) includes: calculating the conservation value of conserved sites based on the species evolutionary tree and whole genome alignment results of the in-group and out-group, and identifying the evolutionary conserved sites and their conservation values ​​in the genome of the target plant based on the screening threshold of the conservation value of evolutionary conserved sites.

5. The identification method according to claim 4, characterized in that, The selection threshold for the conservation value of the top 2.4% of conserved sites across the entire genome was chosen.

6. The identification method according to claim 3, characterized in that, The software used to identify evolutionarily conserved sites and their degree of conservation in the genome of the target plant is GERP software. Step (11) includes: calculating the GERP values ​​of conserved loci based on the phylogenetic trees and whole-genome alignment results of the in- and out-groups, and identifying evolutionarily conserved loci and their GERP values ​​in the target plant genome using a GERP ≥ 2.75 threshold; and / or, Step (13) includes: identifying harmful mutation sites and their GERP values ​​of the target plant based on evolutionarily conserved sites and their GERP values ​​and SNP mutation site information; the harmful mutation sites include homozygous harmful mutation sites and / or heterozygous harmful mutation sites.

7. The identification method according to claim 1, characterized in that, In step (12), the software used to identify the SNP mutation sites of the target plant includes at least one of GATK, samtools, bedtools, or plink.

8. The identification method according to any one of claims 1-7, characterized in that, The target plant is potato.

9. A method for selecting starting material for a target plant inbred line, characterized in that, Includes the following steps: Step (21) Perform whole-genome alignment of the genomes of species in the target plant and species in the out-of-group with the reference genome of the target plant; based on the phylogenetic tree of the in-group and out-of-group and the whole-genome alignment results, identify the evolutionary conserved sites and their degree of conservation in the genome of the target plant. Step (22) Using the target plant as material, identify the SNP mutation sites of the target plant; Step (23) Based on evolutionarily conserved sites and their degree of conservation and SNP mutation site information, identify the harmful mutation sites of the target plant and their degree of harmful mutation effect; the harmful mutation sites include homozygous harmful mutation sites and / or heterozygous harmful mutation sites; Step (24) Calculate the total value of harmful mutations in the target plant using formulas (1) and (2) based on the genotype, harmful mutation sites, and degree of harmful mutation effect of the target plant. The total value of harmful mutations includes the total value of homozygous harmful mutations and / or the total value of heterozygous harmful mutations. Step (25) Calculate the total genetic value of harmful mutations in the target plant material to be selected using formula (3); Among them, b hom The total number of homozygous harmful mutations is represented by i, where i represents the homozygous harmful mutation site, and L(hom) represents the number of homozygous harmful mutation sites; Effect represents the degree of harmful mutation effect of the harmful mutation site. b het denoted as the total number of heterozygous harmful mutations, j represents the heterozygous harmful mutation sites, and L(het) represents the number of heterozygous harmful mutation sites; b genetic This represents the total ancestry value of harmful mutations; Step (26) Based on the total genetic value of harmful mutations in the target plant material to be selected, determine whether to use the target plant material to be selected as the starting material for the construction of inbred lines; when the total genetic value of harmful mutations in the target plant material to be selected is ≤52000, the target plant to be selected is used as the starting material; The target plant is a plant of the Solanaceae family.

10. The selection method according to claim 9, characterized in that, The target plant is potato.

11. A method for constructing an inbred line of a target plant, characterized in that, Includes the following steps: Step (31) Select starting materials using the method described in any one of claims 9-10, and perform self-crossing on the starting materials to obtain offspring; Step (32): Using the offspring as the target plant material to be selected, repeat step (31) to obtain the target plant inbred line.

12. The construction method according to claim 11, characterized in that, The number of times the repeating step (31) is performed is ≥2.

13. The construction method according to claim 11 or 12, characterized in that, The homozygosity of the inbred line is not less than 95%.

14. A phenotypic prediction method based on the whole genome of a target plant, characterized in that, Includes the following steps: Step (41) Perform whole-genome alignment of the genomes of species in the target plant and species in the out-of-group with the reference genome of the target plant; Based on the phylogenetic tree of the in-group and out-of-group and the whole-genome alignment results, identify the evolutionary conserved sites and their degree of conservation in the genome of the target plant. Step (42) Using the target plant as material, identify the SNP mutation sites of the target plant; Step (43) Based on evolutionarily conserved sites and their degree of conservation and SNP mutation site information, identify the harmful mutation sites of the target plant and their degree of harmful mutation effect; the harmful mutation sites include homozygous harmful mutation sites and / or heterozygous harmful mutation sites; Step (44) Obtain a genome-wide phenotype prediction model through training: A similar population of the target plant material to be tested is obtained as a training population, and the genotype and phenotype of the training population are used as the training set. Based on the genotype of the target plant, the harmful mutation sites, and the degree of harmful mutation effect, the total harmful mutation value of each material in the training population is calculated using formulas (1) and (2); the total harmful mutation value includes the total homozygous harmful mutation value and / or the total heterozygous harmful mutation value. The kinship matrix G of the training set is calculated from the genotypes of the training population; Based on the mixed linear model, the total homozygous and heterozygous harmful mutation values ​​of the training population are used as fixed-effect factors, their kinship matrix is ​​used as random-effect factors, and the observed phenotype of the training population is used as y. The model is then trained to obtain the formula (4). μ , α hom , α het The specific values ​​are used to obtain the prediction model formula (4) for the whole genome phenotypic value of the target plant; Among them, b hom The total number of homozygous harmful mutations is represented by i, where i represents the homozygous harmful mutation site, and L(hom) represents the number of homozygous mutant sites. Effect represents the degree of harmful mutation effect of the harmful mutation site. b het denoted as the total number of heterozygous harmful mutations, j represents the heterozygous harmful mutation sites, and L(het) represents the number of heterozygous harmful mutation sites; y represents the phenotypic value. μ The intercept represents the average phenotype without the influence of harmful mutations. α hom The coefficient representing the total value of homozygous harmful mutations is a fixed effect of the total value of homozygous harmful mutations. α het The coefficient represents the total heterozygous harmful mutation value, which is the fixed effect of the total heterozygous harmful mutation value, and u represents the genetic breeding value, where... The random effects are calculated based on the kinship matrix G; G is the kinship matrix between materials, and e represents the random residuals. Step (45) Obtain the phenotypic values ​​of the test material based on the whole-genome phenotypic prediction model and the genotype of the test material: The total number of harmful mutations is calculated according to formulas (1) and (2) and the genotype of the material to be tested; the total number of harmful mutations includes the total number of homozygous harmful mutations and / or the total number of heterozygous harmful mutations. Calculate the kinship matrix between the test material and the training population based on the genotype of the test material; Based on the total harmful mutation value of the species to be tested and its phylogenetic matrix with the training population, the phenotypic value of the material to be tested is obtained by formula (4); The target plant is a plant of the Solanaceae family.

15. The phenotypic prediction method according to claim 14, characterized in that, The population includes at least one of the following: the Fn segregating population, the F1 population, or the natural population, wherein n ≥ 2.

16. The phenotypic prediction method according to claim 14 or 15, characterized in that, The target plant is potato.

17. The phenotypic prediction method according to claim 16, characterized in that, The phenotype includes at least one of the following: yield, plant height, tuber size, number of tubers, flowering time, and disease resistance.