Method for constructing a base population library for crop domestication breeding using natural populations

By combining phylogenetic analysis and morphological feature screening with reproductive ecology characteristics, a basic population bank of fruit-type Elaeagnus genus crops was constructed, which solved the problems of high time and economic costs in conventional breeding methods and achieved rapid and efficient screening of genetic resources and improved hybridization success rate.

CN115691668BActive Publication Date: 2026-06-02江西省 中国科学院庐山植物园 +1

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
江西省 中国科学院庐山植物园
Filing Date
2022-05-18
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

In existing technologies, constructing a basic population library of fruit-type Elaeagnus genus crops from wild species using conventional breeding methods presents challenges in terms of time and economic costs, especially since interspecific hybrid offspring are difficult to obtain due to self-incompatibility.

Method used

By using phylogenetic analysis and morphological feature screening, PCA or OPLS-DA models can be constructed to expand the number and types of hybrid parents. Combined with reproductive ecology characteristics, suitable species for database construction can be selected, thus shortening the time for resource collection and selection mating experiments.

Benefits of technology

It has improved the selectivity of genetic resources and the success rate of hybridization, shortened the time and cost of building a basic population bank, expanded the range of choices, and enriched genetic resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115691668B_ABST
    Figure CN115691668B_ABST
Patent Text Reader

Abstract

The application discloses a method for constructing a basic population library of crop domestication breeding by using natural populations. The first aspect of the application provides a method for screening a library construction species of a basic population library of crop domestication breeding, which comprises the following steps: obtaining a target species of a crop; performing phylogenetic analysis on the target species according to genomic markers; clustering the target species according to a result of the phylogenetic analysis by using morphological characteristics to obtain a plurality of groups; obtaining species information of parents of a hybrid species of the crop, and screening a group containing the parents of the hybrid species from the plurality of groups; and screening the group according to reproductive ecological characteristics to obtain the library construction species. In this way, a suitable library construction species is screened, and an original population is quickly and effectively constructed, so that the time for resource collection and selection of mating experiments is greatly shortened, and the manpower, material resources and time required for construction of the basic population library are reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of crop breeding technology, and in particular to a method for constructing a basic population library for crop domestication breeding using natural populations. Background Technology

[0002] When breeding relatively mature food crops such as rice and corn, genetic resources are typically obtained from populations that have undergone artificial cultivation and selection to build a basic population bank. However, for some potential or undeveloped crops that lack sufficient domestication processes and only have wild or semi-wild types, this method is less feasible. In the 20th century, only four fruit crops were domesticated from wild species: kiwifruit, blueberry, avocado, and macadamia nut. Looking at the trajectory of wild plant domestication and crop adoption, the mainstream method for breeding new crops (such as food, oilseeds, sugar crops, spices, fruits, and medicinal crops) remains the systematic population improvement and new variety selection of superior germplasm resources from widely distributed wild species using conventional breeding methods. Taking the fruit-bearing Elaeagnus as an example, the genus Elaeagnus L. is a genus in the Elaeagnaceae family, native to temperate and subtropical regions of Asia, Australia, Southern Europe, and North America. It is cultivated as an ornamental or fruit crop due to its dense, shrub-like structure, fragrant flowers, and ripe fruits rich in lycopene. Although there are approximately 100 recognized wild species in the genus Elaeagnus, only about 15 are currently widely reported to have edible or medicinal value. In most countries and regions, most varieties of Elaeagnus are typically planted around gardens or courtyards as ornamental shrubs, with very few being developed for fruit production.

[0003] Self-incompatibility is a common characteristic of Elaeagnus genus plants, offering the possibility of creating new cultivars through interspecific hybridization. In conventional breeding methods, recurrent selection is an important method for improving crop populations, especially cross-pollinated varieties. Since large-scale geographically adaptable distribution is a necessary condition for selecting hybrid fruit crops with ideal characteristics, it is necessary to establish an effective basic population bank with more genetic resources to select and breed the next generation of fruit-type Elaeagnus genus crops. Therefore, how to quickly and effectively construct original populations using local and interregional kinship within the same genus has become an urgent issue to consider in breeding design. Currently, it is generally believed that the offspring of interspecific hybridization are difficult to obtain through broad hybridization due to strong self-incompatibility. To obtain fertile recombinant offspring, the primary condition for parental selection is to overcome this obstacle. In field breeding, experimental hybridization is usually used to overcome this obstacle, but selecting wild resources of perennial shrubs requires a significant amount of time for resource collection and selective mating experiments. Therefore, it is necessary to provide a method for rapidly constructing a basic population bank. Summary of the Invention

[0004] This application aims to address at least one of the technical problems existing in the prior art. To this end, this application proposes a method for constructing a basic population library for crop domestication and breeding using natural populations. By pre-screening the species to be used in the library, the construction time and economic cost of the basic population library are significantly reduced.

[0005] A first aspect of this application provides a method for screening species for the construction of a basic population bank of crop domestication breeding, the method comprising the following steps:

[0006] Obtain the target species of the crop;

[0007] Phylogenetic analysis of the target species based on genomic markers;

[0008] Based on the phylogenetic analysis results, the target species were clustered using morphological characteristics to obtain several groups;

[0009] Obtain species information of the parents of hybrids of the same family of crops, and screen out groups containing the parents of hybrids from several groups;

[0010] Based on the selection of groups according to reproductive ecology characteristics, species for database construction were obtained.

[0011] The method according to the embodiments of this application has at least the following beneficial effects:

[0012] The offspring of interspecific hybridization are highly self-incompatible, making them difficult to obtain through broad hybridization. Therefore, obtaining fertile recombinant offspring is a major challenge in constructing a basic population bank. In this approach, firstly, by constructing a phylogenetic tree, the selection of parents for the bank is limited to the evolutionary branches of existing hybrid parents. These existing hybrid parents clearly belong to populations directly capable of interspecific hybridization. Clustering further expands the number and groups of hybrid-compatible species, broadening the range of candidates for the bank and enriching genetic resources, thereby improving the selectivity and tolerance of the bank construction. Secondly, reproductive ecology characteristics are used to screen groups within the group, ensuring a better match between different species in terms of reproductive ecology-related plant traits and external geographical distribution and dispersal. This verifies that the species within the group have the environmental permitting to produce natural hybridization, increasing the hybridization success rate. In this way, suitable species for the bank are selected, enabling efficient and rapid construction of a basic population bank of original species. This significantly shortens the time required for resource collection and selection mating experiments, reducing the manpower, resources, and time required for basic population bank construction.

[0013] In some embodiments of this application, the genomic markers are derived from at least one of the nuclear genome and the chloroplast genome.

[0014] In some embodiments of this application, morphological characteristics include at least one of the following: flowering period, fruiting period, minimum leaf size, maximum leaf size, whether it is evergreen, plant height, presence or absence of thorns, whether it is climbing, flower color, calyx tube length, style with long soft hairs, and number of flowers.

[0015] In some embodiments of this application, the method of clustering target species using morphological characteristics based on phylogenetic analysis results to obtain several groups includes:

[0016] Based on different evolutionary branches in phylogenetic analysis, the target species are clustered using quantified morphological characteristics to construct PCA or OPLS-DA models, resulting in several groups.

[0017] In some embodiments of this application, reproductive ecology characteristics include at least one of plant type, species distribution, flowering period, flower color, and style morphology.

[0018] In some embodiments of this application, phylogenetic analysis of a target species based on genomic markers further includes: correcting the phylogenetic analysis results by incorporating taxonomic characteristics.

[0019] In some embodiments of this application, the crop is a fruit tree crop, and more specifically a berry crop.

[0020] In some embodiments of this application, the crop is a crop of the Elaeagnus genus.

[0021] In some embodiments of this application, hybrids of crops belonging to the same family can be natural hybrids and / or artificial hybrids formed by species from the same genus and / or different genera within the same family classification level of crops, with at least one parent being either the male or female parent.

[0022] A second aspect of this application provides a computer-readable storage medium storing computer-executable instructions for causing a computer to perform the aforementioned method.

[0023] A third aspect of this application provides an apparatus comprising a processor and a memory, wherein the memory stores a computer program executable on the processor, and the processor implements the aforementioned method when executing the computer program.

[0024] A fourth aspect of this application provides an apparatus for screening genera of a basic population bank for crop domestication breeding, the apparatus comprising:

[0025] The acquisition module is used to acquire the target species of crops.

[0026] The phylogenetic analysis module is used to perform phylogenetic analysis on target species based on genomic markers.

[0027] The clustering module is used to cluster target species based on morphological characteristics according to the results of phylogenetic analysis, resulting in several groups;

[0028] The screening module is used to obtain species information of the parents of hybrids of the same family of crops, screen out groups containing the parents of hybrids from several groups, and screen groups according to reproductive ecology characteristics to obtain species for database construction.

[0029] The fifth aspect of this application provides a method for constructing a basic population library for crop domestication and breeding using natural populations. The method includes the steps of obtaining genetic material from the corresponding natural population based on information about the species to be bred and constructing the basic population library.

[0030] Among them, the species used for establishing the database are obtained through the aforementioned methods, or by screening the species used for establishing the database of crop basic populations as described above.

[0031] A sixth aspect of this application provides a method for breeding crops, the method comprising the step of constructing a basic population library according to the foregoing method.

[0032] The technical highlights of the method for constructing a basic crop population library and the breeding method provided in this application are as follows:

[0033] First, it proposes using several hybrids to assess evolutionary branches and corresponding species with high hybridization affinity in natural populations, thus significantly narrowing the target scope of natural material collection based solely on subjective imagination in conventional breeding methods. Second, it proposes a new model, namely a quantitative and cluster analysis model of morphological characteristics guided by molecular phylogenetic evolutionary relationships. This model can directly and accurately guide information on species with high interspecific hybridization affinity, shortening subsequent library construction time and costs.

[0034] Additional aspects and advantages of this application will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of this application. Attached Figure Description

[0035] Figure 1 This is a flowchart illustrating the method for selecting species for the construction of a basic population library of crop domestication and breeding, as described in an embodiment of this application.

[0036] Figure 2The embodiments of this application use different genomic markers and construct phylogenetic trees in different ways. Among them, A is a phylogenetic tree constructed using the ML method based on matK sequences, B is a phylogenetic tree constructed using the ML method based on ITS sequences, and C is a phylogenetic tree constructed using the MP method based on ITS sequences.

[0037] Figure 3 This is a phylogenetic tree of the genus Elaeagnus, constructed using matK sequences (molecular markers derived from chloroplast genomes) and grouped in conjunction with some plant morphological characteristics, as described in the embodiments of this application.

[0038] Figure 4 The following are the clustering results of 52 species of Elaeagnus genus using different methods in the embodiments of this application. Among them, A is the result of principal component analysis (PCA) using morphological features; B is the result of orthogonal partial least squares discriminant analysis (OPLS-DA) using morphological features based on the phylogenetic evolutionary branch of matK; C is the variables in B sorted from largest to smallest according to the variable importance projection (VIP) value; and D is the result of evaluating the validity of different evolutionary branch boundaries using the receiver operating characteristic curve (ROC).

[0039] Figure 5 This document presents the comparative analysis results of morphological characteristics among different populations of Elaeagnus pungens G1–G3 in the embodiments of this application. Wherein, A represents flowering period (V1), B represents fruiting period (V2), C represents minimum leaf size (V3), D represents maximum leaf size (V4), E represents evergreen status (V5), F represents plant height (V6), G represents thorn presence (V7), H represents climbing ability (V8), I represents flower color (V9), J represents calyx tube length (V10), K represents style pubescence (V11), and L represents flower quantity (V12). Wherein, N>9; t-test, ****p<0.0001; ***p<0.001; **p<0.01; *p<0.05.

[0040] Figure 6 This is the result of neural network analysis of the G1 to G3 population distribution of Elaeagnus species in China in the embodiments of this application. Detailed Implementation

[0041] The following will clearly and completely describe the concept and technical effects of this application in conjunction with embodiments, so as to fully understand the purpose, features and effects of this application. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are all within the scope of protection of this application.

[0042] The embodiments of this application are described in detail below. The described embodiments are exemplary and are only used to explain this application, and should not be construed as limiting this application.

[0043] In the description of this application, "several" means one or more, "multiple" means two or more, "greater than," "less than," and "exceeding" are understood to exclude the stated number, while "above," "below," and "within" are understood to include the stated number. The use of "first" and "second" is for the purpose of distinguishing technical features only and should not be construed as indicating or implying relative importance, or implicitly indicating the number of indicated technical features, or implicitly indicating the order of the indicated technical features. In the following description, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than that shown in the flowchart.

[0044] Unless otherwise defined, all technical and scientific terms used in this application have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.

[0045] In the description of this application, the terms "one embodiment," "some embodiments," "illustrative embodiment," "example," "specific example," or "some examples," etc., refer to specific features, structures, materials, or characteristics described in connection with that embodiment or example, which are included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.

[0046] This application provides a method for screening species for the construction of a base pool for crop domestication breeding. The base pool refers to the collection of original breeding materials. Domestic and international breeding practices have shown that the beneficial genes needed are usually present in various different germplasm resources. Therefore, it is necessary to construct an original population with abundant genetic resources, especially superior genes, for improvement. Recurrent selection is an important breeding method for improving crop populations, especially cross-pollinated varieties. It refers to selecting ideal individuals from a population for cross-pollination to achieve gene and trait recombination, thereby forming a morphological population. Therefore, in some specific embodiments, the base pool refers to the base population used in domestication breeding using recurrent selection. In this application embodiment, the method for screening species for the construction of a base pool for crop domestication breeding includes the following steps:

[0047] S110, Obtain the target species of the crop.

[0048] Specifically, the target species can refer to all species of crops at the genus level, or it can be a subset of species selected at the genus level, such as species with specific geographical distributions or traits selected for pre-screening based on actual needs. Specifically, the crop can refer to any food or economic crop that meets specific crop breeding objectives and possesses specific conditions or uses, such as food crops, fruit crops, medicinal crops, fiber crops, oil crops, sugar crops, three-ingredient (beverage, spice, condiment) crops, dye crops, ornamental crops, and other crops for different purposes. Taking food as an example, it can be a fruit crop, such as berries from domesticated and improved wild plants. The plant type of a crop can be multiple, classified according to its spatial hierarchy, including at least one of trees, shrubs, and herbs. Taking berry crops as an example, they mainly consist of trees and shrubs.

[0049] S120. Perform phylogenetic analysis on the target species based on genomic markers.

[0050] In phylogenetic analysis, genomic markers refer to molecular sequences carrying genetic information upon which the phylogenetic relationships between different species are calculated. Specifically, genomic markers can originate from at least one of the nuclear genome or the chloroplast genome. Genomic markers derived from the nuclear genome can be at least one of the ribosomal transcriptional spacer regions (ITS) (including ITS1 and ITS2) or intergenic spacer sequences (IGS). These non-coding regions exhibit relatively high intraspecific homogeneity but significant interspecific differences, thus making them suitable for phylogenetic analysis at both intraspecific and interspecific levels. Genomic markers derived from the chloroplast genome can be at least one of the following chloroplast genes: matK, rbcL, rps4, psbA, psbI, psbK, trnH, trnL, trnF, rpoC1, rpoB, atpF, and atpH. It is understood that the matK (cpDNA) sequence demonstrates significant stability and reliable sequencing information in phylogenetic analysis, constructing genetic relationships that more closely approximate actual plant classification and evolutionary relationships; therefore, genomic markers containing the matK sequence are preferred. Furthermore, since the nuclear genome and chloroplast genome represent the characteristics of the parent and maternal parent of a species, respectively, the optional genomic markers include markers derived from the nuclear genome and chloroplast genome.

[0051] Phylogenetic analysis refers to inferring phylogenetic relationships between different target species, typically represented by a phylogenetic tree. Specifically, in constructing a phylogenetic tree, specific outgroups are selected to root the tree, and each species is used as a node, dividing it into several evolutionary branches containing one or more nodes. Within the same evolutionary branch, different species are closely related. In this case, since the parents of natural or artificial hybrids are confirmed to have interspecific self-compatibility, other closely related species within the same evolutionary branch have a greater chance of acting as substitutes for the parents and producing fertile offspring compared to species in other branches. In some implementations, the phylogenetic tree can be constructed using at least one of the following methods: distance method, maximum parsimony method, maximum likelihood method, Bayesian method, etc. For the phylogenetic tree, support characterizes the reliability of a specific evolutionary branch. Specific calculation methods for support include bootstrap value, UFBoot (Ultra-fast bootstrap), posterior probability, etc., with higher values ​​indicating higher reliability. In some implementations, a bootstrap value is used to characterize the credibility of an evolutionary branch. When the bootstrap value is less than 70% (percentage or decimal), the evolutionary branch it represents is deemed unreliable.

[0052] In some implementations, because the genomic markers relied upon for phylogenetic analysis originate from different countries and laboratories, there may be issues with sequencing and identification, potentially interfering with the construction of the phylogenetic tree. Therefore, all haplotypes of the same species are retained for analysis. Furthermore, the results of phylogenetic analysis can be further corrected using traditional plant taxonomy. Non-limiting implementations include introducing basic taxonomic features during phylogenetic analysis, such as deciduousness (evergreen / deciduous), flowering period, fruiting period, and plant type (shrub / tree), to construct a phylogenetic tree incorporating these features. This allows for the selection of broader groupings based on different evolutionary branches, with each group containing one or more evolutionary branches.

[0053] S130. Based on the phylogenetic analysis results, the target species are clustered using morphological characteristics to obtain several groups.

[0054] Specifically, phylogenetic analysis results refer to several evolutionary branches containing one or more species obtained through the construction of a phylogenetic tree. These different evolutionary branches are used to provide basic grouping principles for clustering, constructing a grouping model based solely on plant morphological characteristics to obtain several groups. Morphological characteristics can be any morphological trait, including, understandably, chemotypes based on chemical composition or content. In some embodiments, morphological characteristics can be obtained by referring to relevant flora descriptions. Specifically, morphological characteristics include traits related to plant life form, ecological type, pollination ecology, reproductive ecology, and chemotype, such as at least one of flowering period, fruiting period, minimum leaf size, maximum leaf size, evergreen status, plant height, thorn presence or absence, climbing ability, flower color, calyx tube length, style pubescence, and number of flowers. In some embodiments, morphological characteristics are quantified, and the further quantified morphological characteristics are standardized. Quantification and standardization can be performed using any relevant data processing methods well known in the art.

[0055] Specifically, based on different evolutionary branches in phylogenetic analysis, the target species are clustered using quantified morphological characteristics to construct a model, resulting in several groups. The target species are then clustered based on these morphological characteristics to form several different groups. Different species within the same group share one or more identical or similar morphological characteristics, while specific morphological characteristics or groups of morphological characteristics differ between different groups.

[0056] In some implementations, the clustering methods described above can include some or all species distributed at different latitudes, expanding the population of species with interspecific hybridization compatibility. Moreover, under strict molecular phylogenetic support for grouping, morphological clustering also provides a new means to infer interspecific self-compatibility.

[0057] The specific methods for constructing the model can be any one or more of the analysis methods that can be used for clustering, such as principal component analysis (PCA), partial least squares (PLS), partial least squares discriminant analysis (PLS-DA), orthogonal partial least squares (OPLS), or orthogonal partial least squares discriminant analysis (OPLS-DA) or heatmap analysis (HM).

[0058] S140. Obtain species information of the parents of hybrids of the same family as the crop, and select groups containing the parents of the hybrids from several groups.

[0059] Among them, the crop in the same family refers to the wild species or the wild species from which the domesticated species originated. The same family refers to the same genus and / or different genera under the same family classification level of the crop. The hybrid can be at least one of natural hybrids or artificial hybrids. Therefore, the hybrid of the same family of the crop can be a natural hybrid and / or artificial hybrid formed by species within the same genus and / or different genera under the same family classification level of the crop as at least one of the male or female parents.

[0060] The position of the parents of successfully hybridized hybrids in the phylogenetic tree provides a reference example for obtaining the genetic characteristics and distribution patterns of hybrid-compatible species, thus enabling the analysis of species for constructing a basic population library even in the absence of sufficient DNA sequences. This step expands the number and taxa of hybrid-compatible species by extending from the parents of the hybrid to other closely related species within the group containing the parents that share one or more key morphological features.

[0061] S150. Based on the reproductive ecology characteristics, the selected groups were used to obtain the species for the database.

[0062] It is understandable that reproductive ecological characteristics refer to any characteristics related to species reproduction, such as traits related to pollination, like flower color, style morphology, and flowering period, as well as the geographical distribution and dispersal of species related to reproduction. Geographical distribution and dispersal are used to verify the environmental permitting for natural hybridization between species. Depending on the final breeding objectives of the basic population bank, reproductive ecological characteristics can be adaptively adjusted. For example, when cultivating fruit crops, reproductive ecological characteristics mainly include fruit traits, including but not limited to peel thickness, fruit size, fruit yield, fruit stalk length, flowering period, and fruiting period.

[0063] In the embodiments of this application, the species to be bred are first investigated using genomic marker regions, phylogenetic analysis is performed, and morphological clustering is combined to assist in the construction of an effective population bank, serving as the genetic basis for superior parents for breeding practice. Specifically, firstly, the phylogenetic positions of the parents of successfully hybridized hybrids provide reference examples for obtaining the genetic characteristics and distribution patterns of hybrid-compatible species; secondly, statistical and model-based clustering of morphological traits indicates the natural dispersion of different species clusters; based on this, species are screened according to relevant reproductive ecology characteristics; finally, the screened species are used to guide the selection of original parents and the construction of the basic breeding population bank. The time required for establishing the bank, estimated at 5 years based on the number of Elaeagnus species, can be shortened to 2 years; the cost of establishing the bank, estimated based on species distribution, can be reduced from 280,000 to 130,000, significantly reducing both time and labor costs.

[0064] This application also provides a computer-readable storage medium storing computer-executable instructions for causing a computer to perform the aforementioned method.

[0065] This application also provides an apparatus including a processor and a memory, wherein the memory stores a computer program that can run on the processor, and the processor implements the aforementioned method when running the computer program.

[0066] The memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs, as described in the aforementioned methods in the embodiments of this application. The processor performs the screening of library species by running the non-transitory software programs and instructions stored in the memory.

[0067] The memory may include a program storage area and a data storage area. The program storage area may store the operating system and application programs required for at least one function; the data storage area may store programs that execute the aforementioned programs. Furthermore, the memory may include high-speed random access memory and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device.

[0068] Specifically, the memory may optionally include memory remotely located relative to the processor, which can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.

[0069] The non-transient software programs and instructions required to implement the above screening method are stored in memory. When executed by one or more processors, they perform the screening of species for the construction of the basic population library of crop domestication and breeding.

[0070] This application embodiment also provides an apparatus for screening species for the establishment of a basic population bank of crop domestication breeding, the apparatus comprising:

[0071] The acquisition module is used to acquire the target species of crops;

[0072] The phylogenetic analysis module is used to perform phylogenetic analysis on target species based on genomic markers;

[0073] The clustering module is used to cluster target species based on morphological characteristics according to the results of phylogenetic analysis, resulting in several groups;

[0074] The screening module is used to obtain species information of the parents of hybrids of the same family of crops, screen out groups containing the parents of hybrids from several groups, and screen groups according to reproductive ecology characteristics to obtain species for database construction.

[0075] The device implementation described above is merely illustrative. The modules described as separate components may or may not be physically separate; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.

[0076] It is understood that all or some of the steps disclosed above can be implemented as software, firmware, hardware, and suitable combinations thereof. Some or all of the physical components can be implemented as software executed by a processor, such as a central processing unit, digital signal processor, or microprocessor, or as hardware, or as an integrated circuit, such as an application-specific integrated circuit. Such software can be distributed on a computer-readable medium, which can include computer storage media (or non-transitory media) and communication media (or transient media). It is understood that computer storage media includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information (such as computer-readable instructions, data structures, program modules, or other data). Computer storage media includes, but is not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disc (DVD) or other optical disc storage, magnetic cartridges, magnetic tape, disk storage or other magnetic storage devices, or any other medium that can be used to store desired information and is accessible to a computer.

[0077] Furthermore, it is understood that communication media typically contain computer-readable instructions, data structures, program modules, or other data in modulated data signals such as carrier waves or other transmission mechanisms, and may include any information delivery medium.

[0078] This application also provides a method for constructing a basic population library for crop domestication and breeding using natural populations. The method includes the step of obtaining the genetic material of the corresponding natural population based on the information of the species to be constructed, and constructing the basic population library. The species to be constructed are obtained by the aforementioned method or by screening the species to be constructed in the basic population library using the aforementioned device.

[0079] In some implementations, the method includes selecting natural populations with several defined phenotypic traits from the established species based on their information, obtaining their genetic material, and using this genetic material to construct a base pool. The defined phenotypic traits can be specific advantageous traits that meet breeding objectives. Natural populations are primarily wild populations, but semi-wild populations are also included. A base pool refers to a comprehensive genetic material library constructed for a specific breeding purpose. Therefore, a base pool is essentially a new base pool constructed by selecting materials with superior traits, abundant genotypes, and high combining ability from natural populations. Selection, hybridization, and other methods are used to alter the frequency of genes and genotypes within the base pool, increasing the recombination of superior genes and improving the frequency of beneficial genes and genotypes in the population, thereby achieving the goal of population improvement. Specifically, based on the information of the species to be included in the gene bank, genetic material from corresponding natural populations is obtained. The basic population bank can be constructed by anchoring 5-10 phenotypic traits that meet the breeding objectives from the species to be included in the gene bank, collecting as many natural populations as possible that satisfy the breeding objectives, and using ex-situ cultivation techniques to construct a co-cultivation germplasm nursery, forming a gene bank with wild basic materials that meet the breeding objectives. Several individuals are selected from the dominant trait population as the "backbone parents" for breeding. Using the "backbone parents" as the maternal or paternal parents, a recurrent selection breeding method is adopted. Superior individual plants are selected from the original population for self-pollination and testcrosses. Based on the testcross results, individuals with high combining ability or superior phenotypes are selected, mixed, and cross-crossed to form the first-round improved population. The first-round improved population continues to be selected to form the second-round improved population. Through multiple rounds of selection and recombination, the frequency of beneficial genes and the proportion of superior genotypes in the population can be increased, thereby increasing the average trait value in the population while maintaining a certain degree of genetic variation. Finally, a stable, strong-dominant improved population is formed through 4-5 generations of backcrossing. Specifically, this can involve using methods such as "one father, multiple mothers" or "one mother, multiple fathers" pollination, hybridization, diallel crosses, and backcrosses. Alternatively, individuals of the species being established can be mixed with a predetermined amount of seeds in an isolated area, sown, and then freely pollinated. Based on the breeding objectives, superior individuals are selected from the original materials and propagated separately, allowing the offspring of the selected individuals to form a system (strain). Then, through repeated comparative trials, new varieties are developed. It is understandable that during selection, it is crucial to distinguish between genetic and non-genetic variations; selecting for non-genetic variations is often ineffective. Variations generally caused by environmental factors (including cultivation conditions) are considered non-genetic variations. Variations caused by hybridization (natural and artificial hybridization), gene mutations, or chromosomal aberrations are considered genetic variations.

[0080] This application also provides a crop breeding method, which includes the step of constructing a basic population library using the aforementioned construction method. Specifically, after the basic population library is constructed, breeding is carried out through methods such as systematic selection or recurrent selection. This breeding method can be specifically applied to the domestication breeding of berry crops, such as the development breeding of berries in the genus Elaeagnus.

[0081] The following explanation uses the basic population bank required for the development of berries in the genus *Gnaphalium* as an example.

[0082] Example 1

[0083] 1. Materials and Methods

[0084] 1.1 Sampling, Sequence Acquisition and Processing

[0085] DNA sequences for ITS and matK were downloaded from Genbank. ITS consisted of one sequence from an outgroup of a Rhamnaceae species (Ziziphus jujuba Mill.), 50 sequences from the genus Elaeagnus, and 5 sequences from the genus Hippophae (Elaeagnus). Each ITS sequence was at least 541 bp. matK consisted of one sequence from an outgroup of a Rhamnaceae species (Ziziphus jujuba Mill.), 53 sequences from the genus Elaeagnus, and 8 sequences from the genus Hippophae (Elaeagnus). Alignment was performed using BioEditor and Clustal W. Sequences with deletions or errors exceeding 5% of the total bases were considered low-quality and discarded. Gaps were treated as missing data. The heads and tails of aligned sequences were removed. To ensure the DNA sequence data were correctly taxonomically related to the assigned species, only haplotypes with sequence similarity exceeding 99% within the same species were retained for further phylogenetic analysis. Molecular matrices were constructed using Geneious_11 software.

[0086] 1.2 Quantitative phenotype

[0087] Referring to the *Flora of China*, *Flora of Higher Plants of China*, and *Illustrated Flora of Higher Plants of China*, all taxonomic descriptions were quantitatively summarized and analyzed. First, all recorded morphological traits were extracted. After classification, quantification, and data standardization, 12 traits related to the life form, ecological type, pollination ecology, and reproductive ecology of *Elaeagnus* species were finally identified. These 12 traits included flowering period, fruiting period, minimum leaf size, maximum leaf size, evergreen status, plant height, thorn presence or absence, climbing ability, flower color, calyx tube length, style length, and number of flowers. Flowering and fruiting periods were quantified chronologically, while other traits were quantified according to degree. A total of 55 species within the *Elaeagnus* and *Hippophae* genera were investigated using high-resolution online specimen images to extract and quantify trait information. Specimens of the same species included at least six different source regions. For a few species distributed abroad, such as *E. commutata* in the United States and *E. montana* in Japan, corresponding specimen information was also used. Data standardization was performed using SPSS 22.

[0088] 1.3 Sequence Differentiation and Phylogenetic Analysis

[0089] First, alignment was performed using MEGA-X, followed by assembly and quality assessment of the consensus sequences using SeqScape v2.5. Phylogenetic reconstruction was conducted using maximum parsimony (MP), maximum likelihood (ML), and Bayesian inference (BI) to resolve interspecific and intergeneric phylogenetic relationships. The bootstrap value was set to 1000 when constructing phylogenetic trees for MP or ML analyses. The optimal models for constructing ML trees were the Hasegawa-Kishino-Yano (HKY) + Gamma distribution (G) model or the Kimura 2-parametric (K2P) model, based on calculations in Find Best DNA / Protein Models. Phylogenetic relationships and divergence times were estimated using BEAST 2.2.

[0090] 1.4 Construction of the OPLS-DA Model

[0091] Based on the molecular evolutionary relationships of 14 species in the phylogenetic tree, an OPLS-DA model was constructed using normalized data of 12 traits. Morphological data from these 14 species were used as training data to validate the model's classification of a total of 52 species. PCA without grouping was used to demonstrate the advantages of the OPLS-DA method. The Receiver Operating Characteristic (ROC) curve was used to evaluate the effectiveness of the cluster boundaries, and the area under the curve (AUC) was used to measure the effectiveness of the clustering. Variable Importance Projection (VIP) was used to assess the importance of different traits in distinguishing different groups obtained from the clusters. Furthermore, F-tests and t-tests were used to analyze the top five contributing traits in the clustering model.

[0092] 1.5 Geographical Distribution Analysis

[0093] Referring to specimen records in the *Flora of China* and the China Digital Herbarium (https: / / www.cvh.ac.cn), the collection locations of specimens for each species were statistically analyzed, and the distribution centers of each species were determined by the frequency of specimen occurrence in the same region. The main distribution centers of different species were analyzed using ArcGIS 10.8 and separated using the OPLS-DA model. Neural network analysis based on Gephi 0.9.2 was used to display the geographical distribution of *Elaeagnus* genus within China. Nodes include the species' Latin name, section name, and geographical region (East China, Central China, South China, Southwest China, and North China).

[0094] 1.6 Statistical Analysis

[0095] Excel was used as the primary tool for recording, editing, and converting file formats. GraphpadPrism 7 software was used for data analysis. All values ​​are expressed as mean ± SD. R and open-source analysis packages were used for difference analysis, such as t-tests, f-tests, mean, and Kruskal-Wallis tests. Whether a difference was statistically significant was determined by a t-test (p < 0.05 or 0.01).

[0096] 2. Experimental Results

[0097] 2.1 Results of Phylogenetic Analysis

[0098] This study constructed matrices using 62 matK sequences from 21 species and 56 ITS sequences from 25 species for molecular phylogenetic analysis. The sequence information is shown in Table 1. 109 sympathetic features were found in the ITS matrix and 45 sympathetic features in the matK matrix. Using *Ziziphus jujuba* as the outgroup, rigorously consistent trees were constructed using the maximum parsimony (MP) method based on ITS and the maximum likelihood (ML) method based on both ITS and matK. Both phylogenetic trees showed that *Hippophae L.* was divided into a separate lineage (bootstrapping value exceeding 95%). Therefore, these two phylogenetic trees are invaluable for further analysis of genetic classification.

[0099] Table 1. Sequence information involved in phylogenetic analysis

[0100]

[0101]

[0102]

[0103]

[0104]

[0105]

[0106]

[0107] Since ITS is biparentally inherited while matK is maternally inherited, phylogenetic results based on ITS sequences can better reflect the genetic relationships between hybrids and their parents. The phylogenetic tree construction results are as follows: Figure 2 As shown in Figures A through C, analysis of the ITS trees using ML or MP methods revealed that three hybrids and one environmental variety were grouped into a single lineage with several previously reported species belonging to underutilized fruits or widely used economic crops (including *E. pungens*, *E. glabra*, *E. macrophylla*, *E. conferta*, *E. angustifolia*, *E. multiflora*, and *E. umbellata*) (bootstrapping values ​​of 73% and 70%, respectively). These varieties are considered to be underutilized fruits or widely used economic crops. Therefore, these seven species have the potential to become pioneering species for future hybridization breeding or valuable wild resources.

[0108] Grouping was performed in the phylogenetic tree constructed based on matK sequences, incorporating plant taxonomic characteristics, to obtain... Figure 3A ring-shaped phylogenetic tree, where plant taxonomic characteristics include whether it is evergreen (evergreen / semi-evergreen / deciduous), plant type (shrub / tree), flowering period, fruiting period, etc. For example... Figure 3 As shown, *E. commutata* forms a separate evolutionary branch with a bootstrap value of 98.8%; *E. angustifolia* and *E. mollis* form another branch with a bootstrap value of only 60.7%; *E. umbellata*, *E. multiflora*, and *E. montana* also form a distinct evolutionary branch with a bootstrap value of 90.5%; *E. bockii*, *E. henryi*, *E. pungens*, *E. conferta*, and *E. lanceolata* form a distinct evolutionary branch with a bootstrap value of 95.8%; the remaining species of *E. glabra*, *E. macrophylla*, and another haplotype of *E. pungens* form another branch with a bootstrap value of only 63.5%. Therefore, only evolutionary branches with bootstrap values ​​higher than 70% will be discussed subsequently. Referring to the grouping of taxonomic features on the outer side of the phylogenetic tree, species of the *Elaeagnus* genus are divided into three groups. Along the molecular phylogeny, the maternal inheritance characteristic of the *Elaeagnus* genus is the evolution of plant life forms from deciduous shrubs to erect or climbing evergreen plants. The reason for introducing the genus Hippophae in the phylogenetic tree is that when there is a conflict between traditional morphological and molecular classifications, individual species of the genus Elaeagnus may fall into the genus Hippophae.

[0109] 2.2 OPLS-DA Model Construction Results

[0110] according to Figure 3 The results showed that the molecular phylogeny of 14 Elaeagnus species exhibited clear evolutionary relationships consistent with plant taxonomy. Therefore, the evolutionary branches defined by phylogeny can be used as grouping principles and optimized training datasets to construct grouping models based solely on plant morphological data. The established morphological grouping model can then guide breeding practices to address the challenge of obtaining more large-scale species DNA sequences in the short term.

[0111] like Figure 4 As shown in A, the initially used PCA model failed to effectively group the 52 species morphologically. However, as... Figure 4 As shown in B and referring to Table 2, based on Figure 3 The OPLS-DA model of phylogenetic clades in the study of phylogenetic branches, however, exhibited clear clustering results. The results showed that the 52 species were well clustered into three groups: G1, G2, and G3. The five groups of points with different colors represent the phylogenetic branches of phylogenetic clades. Figure 3 Species within the five evolutionary branches derived from the evolutionary relationships in the model. For example... Figure 4As shown in C, the top five traits that significantly influence the clustering model in B are V9 (flower color), V1 (flowering period), V5 (whether it is evergreen), V3 (minimum leaf size), and V2 (fruiting period). Among these, flower color and flowering period play a decisive role in the morphological clustering model. Figure 4 As shown in D, the numbering is the same as the evolutionary branch in B. The dark blue 3 represents G1, and the green and light blue 1 and 2 represent G2. The AUC values ​​of both G1 and G2 exceed 95%, indicating that the OPLS-DA model has a good analytical effect on morphological grouping. Therefore, the phylogenetic-based OPLS-DA model is an effective method for clustering wild Elaeagnus angustifolia plants, and can extract key morphological features related to plant taxonomy.

[0112] Table 2. Grouping Results of OPLS-DA Model

[0113]

[0114]

[0115] The morphological characteristic data were normalized, and the results were consistent with the OPLS-DA results described above. For example... Figure 5 As shown, among the 12 morphological traits, 5 traits showed significant differences among the three groups. First, flower color differed significantly among the three groups (ANOVA, p = 0.0013), gradually changing from yellow to white from G1 to G3. In G2, the flowering and fruiting periods were relatively independent, with flowering concentrated in the fourth quarter (t-test, p < 0.0001) and fruiting in the first quarter of the following year (t-test, p < 0.05). Therefore, flowering uniformity is indeed a prerequisite for natural interspecific gene recombination, directly leading to the rich species diversity of group G2. Furthermore, the results showed that calyx tube length and style pubescence did not significantly contribute to population formation (ANOVA, p = 0.9729 and 0.4733, respectively).

[0116] In summary, the most significant trait difference between the different groups is flowering time, and the G2 group plants have relatively good natural hybridization conditions in China, mainly due to the consistency of pollination cycles and the cross-distribution of multiple species. However, the geographical distribution of Elaeagnus angustifolia in different groups remains unclear. Therefore, further analysis of geographical distribution is needed to help assess self-compatibility and identify potential geographical barriers.

[0117] 2.3 Geographical distribution of the genus Elaeagnus

[0118] like Figure 6 As shown, group G2 is in the upper left corner, containing 28 species, while G1 and G3 are in the lower right corner, each containing 24 species. (Combined) Figure 5 and Figure 6Group G3 is dominated by deciduous shrubs with narrower leaves, primarily distributed in high-altitude or high-latitude regions, including Gansu, Yunnan, and Guizhou provinces. While Group G1 also consists of some deciduous shrubs, it is mainly distributed in low-altitude areas, and its leaves are significantly larger than those of Group G3. (Reference) Figure 6 The arrows indicate the evolutionary direction of species within the genus Elaeagnus. Group G2 has reached the pinnacle of molecular systematics, including some evergreen or semi-evergreen species, which are widely distributed in areas south of the Yellow River.

[0119] The three known hybrids within the genus Elaeagnus are:

[0120] E. ×ebbingei Boom., 'Gilt Edge', parents are E. pungens and E. macrophylla;

[0121] E.reflexa &Decne., the parents are E. glabra and E. pungens;

[0122] E. × maritima Koidz, whose parents are E. glabra and E. macrophylla.

[0123] The parents of the three hybrids mentioned above all belong to group G2. Therefore, species in group G2 are more likely to produce fertile offspring when replacing the parents of the hybrids compared to species in groups G1 and G3. Therefore, fruit traits of species in group G2 were specifically examined, including pericarp thickness, fruit size, fruit yield, pedicel length, flowering period, and fruiting period. Ten widely distributed wild species in group G2 with the potential to be developed into berry crops were identified. As shown in Table 3, these include three evergreen climbing shrubs: *E. glabra*, *E. gonyanthes*, and *E. henryi*; two evergreen micro-climbing / vining shrubs: *E. macrophylla* and *E. conferta*; and five erect shrubs: *E. lanceolata*, *E. delavayi*, *E. oldhami*, *E. pungens*, and *E. bockii*. The table also summarizes some representative dominant traits of these species, preparing for the next step of collecting these wild resources to construct a targeted basic population database. For example, *E. conferta* can provide resources with large fruits for the basic population bank, while *E. gonyanthes* can provide important genetic resources with long pedicels. *E. lanceolata* and *E. delavayi* can be used to domesticate thornless hybrids, and five climbing / vining shrubs can be used to develop high-yielding commercial cultivars that are easy to automate field management. Suitable individual samples can be selected from these 10 *Elaeagnus* species to construct the basic population bank.

[0124] Table 3. Screening results of fruit-type Elaeagnus angustifolia

[0125]

[0126] Table 4 shows the estimated budget for constructing basic population banks for different genera of plants during the field species collection phase, based on the methods described above.

[0127] Table 4. Explanation of the budget for field collection of several plant genera.

[0128]

[0129] The main difference between the conventional species collection and the species collection budget in the embodiments of this application lies in the difference in the number and distribution of the species to be collected. The method in the embodiments can further narrow down the number of species to be collected and the distribution area involved from the number of target species and the distribution area of ​​target species. The budget is calculated based on the optimal sampling route kilometers, at 3 yuan per kilometer.

[0130] Based on the estimation results in the table above, and using the construction method described in the examples, the construction time for building a basic population bank of Elaeagnus pungens for the domestication and breeding of berry crops can be shortened from the estimated 5 years to 2 years based on the number of species; the cost of building the bank can be reduced from 280,000 to 130,000 based on the species distribution. Both time and manpower costs are greatly reduced.

[0131] Example 1 is further explained below. Generally, traditional breeding methods, such as genetic selection or systematic breeding, are suitable for the initial domestication of wild plant resources. Therefore, the accurate characteristics of the available genetic pool are important in breeding programs. Once the parent species are identified, assessing their taxonomic and phylogenetic positions and identifying closely related species is essential for studying their hybridization compatibility. In this application, a method for constructing a basic population pool during the domestication breeding of wild berry crops is provided. Taking the development of new berries in the genus Elaeagnus as an example, Elaeagnus has significant advantages in geographical distribution and translatitudinal adaptability, with its natural distribution extending from the North Temperate Zone to the subtropics and even the tropics. This characteristic has practical significance for new methods of fruit breeding in mid- and low-latitude regions. This construction method is based on the following concept: using natural or artificial hybrids as basic clues, the parent species are identified to determine a limited number of species with interspecific self-compatibility; the evolutionary relationship of the parent species is used as a basis to examine their evolutionary level or branch, thus making a preliminary judgment on the evolutionary branches within the genus *Hymenopterys* that may have hybridization compatibility; then, based on the evolutionary branch, an OPLS-DA grouping model of morphological characteristics is constructed; the constructed OPLS-DA grouping model is used to group the morphological characteristics of different species across latitudinal regions, marking the existing hybrids and their parent species to their respective groups; using the marked groups as research objects, their reproductive ecology-related characteristics are investigated, and the results are used to guide the construction of a breeding base population of wild fruit trees.

[0132] The results of Example 1 support the construction of an effective population bank through a combination of phylogenetic analysis and morphological clustering, and reveal the genetic basis of superior parents for breeding practices. The construction of a basic population bank mainly relies on the breeder's approach and understanding of breeding goals, as well as the genetic collection of wild relatives with superior morphological traits. However, how to analyze solely based on morphological traits in the absence of sufficient DNA sequences is a challenge. In this application, an innovative clustering model was used to preliminarily solve this problem. First, the phylogenetic positions of successful hybrids provide reference examples for obtaining the genetic characteristics and distribution patterns of hybrid-compatible species. Second, the statistical and OPLS-DA model clustering of 12 quantitative traits shows the natural distribution of different species clusters and innovatively determines the phylogenetic positions of hybrids, which has profound significance for constructing a basic population bank. Furthermore, the current inferences directly identify 10 potential wild species, consistent with breeding objectives for further improvement and domestication.

[0133] The present application has been described in detail above with reference to the embodiments. However, the present application is not limited to the above embodiments. Within the scope of knowledge possessed by those skilled in the art, various changes can be made without departing from the spirit of the present application. Furthermore, unless otherwise specified, the embodiments and features in the embodiments of the present application can be combined with each other.

Claims

1. A method for screening species for the construction of a basic population bank of crop domestication and breeding, characterized in that, Includes the following steps: Obtain the target species of the crop; Phylogenetic analysis of the target species was performed based on genomic markers, which included markers derived from the nuclear genome and markers derived from the chloroplast genome. The genomic markers derived from the nuclear genome included at least one of the ribosomal transcription spacer region (ITS) and the intergenic spacer sequence (IGS). The markers derived from the chloroplast genome included matK. Based on the phylogenetic analysis results, the target species are clustered using morphological characteristics to expand the interspecific hybridization compatible species population and obtain several groups; Obtain species information of the parents of the hybrids of the same family of the crop, and select groups containing the parents of the hybrids from the several groups. The hybrids of the same family of the crop are natural hybrids and / or artificial hybrids formed by species of the same genus and / or different genera under the same family classification level of the crop as at least one parent, either male or female. The groups were selected based on reproductive ecology characteristics to obtain the species for the database.

2. The method according to claim 1, characterized in that, The morphological characteristics include at least one of the following: flowering period, fruiting period, minimum leaf size, maximum leaf size, whether it is evergreen, plant height, presence or absence of thorns, whether it is climbing, flower color, calyx tube length, style with long soft hairs, and number of flowers.

3. The method according to claim 1 or 2, characterized in that, The method for clustering the target species using morphological characteristics based on phylogenetic analysis results to obtain several groups includes: Based on the different evolutionary branches of phylogenetic analysis, the target species are clustered using quantified morphological characteristics to construct PCA or OPLS-DA models, resulting in several groups.

4. The method according to claim 1, characterized in that, The reproductive ecology characteristics include at least one of the following: plant type, population distribution, flowering period, flower color, and style morphology.

5. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions for causing a computer to perform the method according to any one of claims 1 to 4.

6. The equipment, characterized in that, The method includes a processor and a memory, wherein the memory stores a computer program that can run on the processor, and the processor, when running the computer program, implements the method of any one of claims 1 to 4.

7. A device for screening species for the construction of a basic population bank of crop domestication and breeding, characterized in that, The device includes: The acquisition module is used to acquire the target species of the crop; A phylogenetic analysis module is used to perform phylogenetic analysis on the target species based on genomic markers, wherein the genomic markers include markers derived from the nuclear genome and the chloroplast genome; A clustering module is used to cluster the target species based on morphological characteristics according to the phylogenetic analysis results, thereby expanding the interspecific hybridization compatible species population and obtaining several groups; The screening module is used to obtain species information of the parents of hybrids of the same family as the crop, screen out groups containing the parents of the hybrids from the plurality of groups, and screen the groups according to reproductive ecology characteristics to obtain species for database construction; the hybrids of the same family as the crop are natural hybrids and / or artificial hybrids formed by species of the same genus and / or different genera under the same family classification level of the crop as at least one parent, either male or female.

8. A method for constructing a basic population bank for crop domestication and breeding using natural populations, characterized in that, include: The species for establishing a database are obtained by the method described in any one of claims 1 to 4, or by the device described in claim 7 for screening species for establishing a database of basic populations for crop domestication and breeding. Genetic materials from corresponding natural populations are obtained based on information about the species used in the database, and a basic population database is constructed.

9. A method for breeding crops, characterized in that, This includes the step of constructing a basic group library using the method described in claim 8.