Peanut core germplasm bank rapid construction method and system and medium
By grouping and clustering analysis of peanut germplasm resources, a core germplasm library was constructed, and the problem of unclear genetic diversity was solved and the utilization efficiency of germplasm resources was improved.
Patent Information
- Application Number
- CN202510354166.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-07-11
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The degree and distribution of genetic diversity in existing peanut germplasm resources are unclear, and the identification of traits is not in-depth enough, which affects the introduction and utilization of germplasm. In particular, the study of traits such as Aspergillus aflatoxin, drought, and oleic acid has not been effectively carried out.
By collecting germplasm resource data, pre-processing and grouping, clustering analysis is performed using UPGMA class averaging method or Euclidean distance method, a third group is generated, and the sampling ratio is adjusted according to the degree of genetic diversity, and a peanut core germplasm library is constructed.
The genetic diversity of the primary germplasm is achieved to maximize the preservation of the minimum sample number, and the utilization efficiency of germplasm resources is improved.
Smart Images

Figure CN120299534A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of constructing a core germplasm bank, and particularly relates to a method, a system and a medium for rapidly constructing a peanut core germplasm bank. Background Art
[0002] Peanut is an important oil and cash crop. In the past 20 years, the research on crop variety resources has received much attention. There are about more than 7,000 existing peanut resources, and systematic evaluations have been made on their main botanical traits, agronomic traits, disease and insect resistance, and seed quality traits, obtaining a large number of resources with various excellent traits. However, generally speaking, there are still some important deficiencies in the research of peanut germplasm resources, including unclear genetic diversity degree and distribution of the conserved resources, which affect the introduction and exploration of germplasm; the trait identification is not deep enough, especially the research on some important traits (such as Aspergillus flavus, drought, oleic acid, etc.) with relatively complex identification techniques has not been effectively carried out, which affects the utilization of germplasm.
[0003] To clarify the genetic diversity of germplasm resources from different sources, study the origin and evolution of peanuts, and construct a core germplasm using cultivated peanut germplasm materials and analyze its genetic diversity have important theoretical and practical significance for accelerating the development and utilization of peanut germplasm resources and improving the exploration efficiency of new peanut gene sources.
[0004] Therefore, providing a method, a system and a medium for rapidly constructing a peanut core germplasm bank to preserve the genetic diversity of the original germplasm to the greatest extent with the least number of samples is an urgent problem to be solved. Summary of the Invention
[0005] In view of the above-mentioned technical problems, the present invention provides a method, a system and a medium for rapidly constructing a peanut core germplasm bank.
[0006] In a first aspect, the present invention provides a method for rapidly constructing a peanut core germplasm bank, the method comprising the following steps:
[0007] Step 1, collecting the existing first germplasm resource data in the whole germplasm resource, preprocessing the first germplasm resource data to obtain second germplasm resource data;
[0008] Step 2, dividing the second germplasm resource data into N1 first groups according to botanical types;
[0009] Step 3, extracting any one of the first groups, dividing any one of the first groups into N2 second groups according to geographical origin, and after traversing all the first groups, obtaining N3 second groups;
[0010] Step 4, extracting any one of the second groups, performing cluster analysis on any one of the second groups to generate N4 third groups;
[0011] Step 5: After traversing all the second groups, extract the sample germplasm resource data from each third group respectively;
[0012] Step 6: Construct a peanut core germplasm bank based on all the sample germplasm resource data.
[0013] Specifically, preprocess according to whether the first germplasm resource data is complete, and delete the incomplete data. Among them, the first germplasm resource data includes basic data, evaluation data, and characteristic data. The basic data includes geographical origin, ecological and geographical conditions of the geographical origin, and breeding system. The evaluation data includes agronomic traits. The characteristic data includes morphological markers, biochemical indexes, and DNA molecular marker data.
[0014] Specifically, the botanical types include multi-seeded type, pearl bean type, runner type, common type, and intermediate type.
[0015] Specifically, the clustering method is the UPGMA average linkage method or the Euclidean distance method.
[0016] Specifically, in Step 5, the method for extracting the sample germplasm resource data is as follows:
[0017] Extract any third group, and calculate the similarity of all the second germplasm resource data in any third group;
[0018] Judge whether the similarity is greater than the first preset value. If so, sample from any third group according to the first ratio. If not, sample from any third group according to the second ratio, where the first ratio is less than the second ratio.
[0019] Specifically, in Step 4, perform cluster analysis on any second group based on agronomic traits.
[0020] Specifically, the agronomic traits include growth period, weight of 100 fruits, weight of 100 kernels, kernel percentage, plant height, plant type, flowering habit, branching type, type, crude protein content, crude fat content, oleic acid content, and linoleic acid content.
[0021] In the second aspect, the present invention also provides a system for rapidly constructing a peanut core germplasm bank, and the system includes:
[0022] A data collection module, a data grouping module, a sample extraction module, and a germplasm bank construction module;
[0023] The data collection module is used to collect the existing first germplasm resource data in the whole germplasm resources, preprocess the first germplasm resource data, and obtain the second germplasm resource data;
[0024] A data grouping module, which is used to divide the second germplasm resource data into N1 first groupings according to botanical types, then extract any one of the first groupings, divide any one of the first groupings into N2 second groupings according to geographical sources. After traversing all the first groupings, N3 second groupings are obtained. Finally, any one of the second groupings is extracted, and cluster analysis is performed on any one of the second groupings based on agronomic traits to generate N4 third groupings;
[0025] A sample extraction module, which is used to extract sample germplasm resource data from each of the third groupings respectively after traversing all the second groupings;
[0026] A germplasm bank construction module, which is used to construct a peanut core germplasm bank according to all the sample germplasm resource data.
[0027] In a third aspect, the present invention provides a computer storage medium, which stores program instructions. When the program instructions run, they control the device where the computer storage medium is located to execute the peanut core germplasm bank rapid construction method of any one of the above.
[0028] In a fourth aspect, the present invention provides a processor, which is used to run a program. When the program runs, it executes the peanut core germplasm bank rapid construction method of any one of the above.
[0029] The present invention discloses a peanut core germplasm bank rapid construction method, system and medium. First, the first germplasm resource data collected is preprocessed to obtain the second germplasm resource data to ensure the accuracy and reliability of the resource data for constructing the peanut core germplasm bank. Subsequently, the second germplasm resource data is stratified and grouped based on botanical types and geographical sources to obtain multiple second groupings, so as to ensure that the constructed core germplasm bank can contain germplasm resource data of all botanical types and all geographical sources. Cluster analysis is performed on the second groupings, and the germplasm resource data with similar genetic characteristics is divided into the same third grouping. The sampling ratio is determined based on the similarity of all the germplasm resource data in the third grouping. Further sampling is performed from the third grouping. For the grouping with a smaller degree of genetic diversity, sampling is performed according to a smaller ratio, and for the grouping with a larger degree of genetic diversity, sampling is performed according to a larger ratio, so as to preserve the genetic diversity of the original germplasm to the greatest extent with the least number of samples. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.
[0031] Figure 1Flow chart of a method for rapidly constructing a peanut core germplasm bank of the present invention;
[0032] Figure 2 Modular schematic diagram of a system for rapidly constructing a peanut core germplasm bank of the present invention. Detailed implementation manners
[0033] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Obviously, the specific embodiments described herein are only used to explain the present invention, which are a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0034] It should be noted that if there are descriptions involving "first", "second", etc. in the embodiments of the present invention, the descriptions of "first", "second", etc. are only for descriptive purposes, and cannot be understood as indicating or implying their relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of such features. In addition, the technical solutions between various embodiments may be combined with each other, but it must be based on the fact that those of ordinary skill in the art can implement them. When the combination of technical solutions is contradictory or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the protection scope required by the present invention.
[0035] Figure 1 The figure shows a flow chart of an embodiment of a method for rapidly constructing a peanut core germplasm bank provided by the present invention, and the flow chart specifically includes the following steps:
[0036] Step 1: Collect the existing first germplasm resource data in the entire germplasm resource, preprocess the first germplasm resource data, and obtain the second germplasm resource data.
[0037] Obtain the first germplasm resource data from the national germplasm resource bank, agricultural scientific research institutions, universities, etc.
[0038] Specifically, preprocess based on whether the first germplasm resource data is complete, and delete the incomplete data. Among them, the first germplasm resource data includes basic data, evaluation data, and characteristic data. The basic data includes geographical origin, ecological and geographical conditions of the geographical origin, and breeding system. The evaluation data includes agronomic traits. The characteristic data includes shape markers, biochemical indexes, and DNA molecular marker data.
[0039] For the first type of germplasm resource data collected, extract its basic information, characteristic data, and evaluation data. Such information may include variety name, place of origin, growth habit, yield performance, disease resistance, etc. Clean the extracted data to remove duplicate, incorrect, or incomplete information.
[0040] Step 2: Divide the second type of germplasm resource data into N1 first groups according to botanical types.
[0041] Specifically, botanical types include multi-seeded type, pearl bean type, runner type, common type, and intermediate type.
[0042] Step 3: Extract any one of the first groups and divide any one of the first groups into N2 second groups according to geographical origin. After traversing all the first groups, obtain N3 second groups.
[0043] Exemplarily, the geographical origin of germplasm resources can be divided into the peanut-growing area of the Yellow River Basin, the peanut-growing area of the Yangtze River Basin, the peanut-growing area of the southeast coast, the peanut-growing area of the Yunnan-Guizhou Plateau, the peanut-growing area of the Loess Plateau, the peanut-growing area of Northeast China, and the peanut-growing area of Northwest China.
[0044] As a preferred technical solution of the present invention, it is also possible to obtain peanut germplasm resource data from abroad and construct a global peanut core germplasm bank. Specifically, foreign germplasm resources can be divided into the peanut-growing area of North America, the peanut-growing area of South America, the peanut-growing area of Asia, the peanut-growing area of Africa, the peanut-growing area of ICRISAT, the peanut-growing area of Europe, and the peanut-growing area of Oceania based on geographical origin.
[0045] Step 4: Extract any one of the second groups and perform cluster analysis on any one of the second groups to generate N4 third groups.
[0046] Specifically, the clustering method is the UPGMA average linkage method or the Euclidean distance method.
[0047] According to the nature and characteristics of the data, select a suitable clustering method, such as K-means clustering, hierarchical clustering, etc. Subsequently, perform corresponding processing (such as normalization, standardization, etc.) on the data used for clustering to ensure the accuracy of the clustering results. Finally, use the selected clustering method to perform cluster analysis on the stratified second groups to form different third groups.
[0048] Specifically, in Step 4, perform cluster analysis on any one of the second groups based on agronomic traits.
[0049] Specifically, agronomic traits include growth period, weight of 100 fruits, weight of 100 kernels, kernel percentage, plant height, plant type, flowering habit, branching type, type, crude protein content, crude fat content, oleic acid content, and linoleic acid content.
[0050] Among the agronomic trait data, growth period, hundred - fruit weight, hundred - kernel weight, kernel percentage, plant height, plant type, flowering habit, branching type, type, etc. belong to phenotypic traits, and crude protein content, crude fat content, oleic acid content, linoleic acid content, etc. belong to nutritional quality traits.
[0051] Step 5: After traversing all the second groups, extract the sample germplasm resource data from each third group respectively.
[0052] Specifically, in Step 5, the method for extracting the sample germplasm resource data is as follows:
[0053] Extract any third group, and calculate the similarity of all the second germplasm resource data in any third group;
[0054] Judge whether the similarity is greater than the first preset value. If so, sample from any third group according to the first ratio; if not, sample from any third group according to the second ratio, where the first ratio is less than the second ratio.
[0055] The first preset value, the first ratio, and the second ratio are set according to the experience of those skilled in the art or according to the actual application scenario, and the embodiments of the present application do not limit this. Exemplarily, the first ratio is 5%, and the second ratio is 10%.
[0056] A larger similarity value indicates a smaller degree of genetic diversity of the germplasm resource data in the group, so sample from the group according to a smaller ratio; a smaller similarity value indicates a larger degree of genetic diversity of the germplasm resource data in the group, so sample from the group according to a larger ratio.
[0057] Preferably, several more similarity levels can also be divided, and different sampling ratios are set for each similarity level.
[0058] Step 6: Construct a peanut core germplasm bank based on all the sample germplasm resource data.
[0059] Figure 2 Shown is a schematic structural diagram of an embodiment of a rapid construction system for a peanut core germplasm bank provided by the present invention. As Figure 2 Shown, the system includes: a data collection module 10, a data grouping module 20, a sample extraction module 30, and a germplasm bank construction module 40.
[0060] The data collection module 10 is used to collect the existing first germplasm resource data in the entire germplasm resource, pre - process the first germplasm resource data, and obtain the second germplasm resource data.
[0061] The data grouping module 20 is configured to divide the second germplasm resource data into N1 first groups according to botanical types, then extract any one of the first groups, divide any one of the first groups into N2 second groups according to geographical sources. After traversing all the first groups, N3 second groups are obtained. Finally, any one of the second groups is extracted, and cluster analysis is performed on any one of the second groups based on agronomic traits to generate N4 third groups.
[0062] The sample extraction module 30 is configured to, after traversing all the second groups, extract sample germplasm resource data from each of the third groups respectively.
[0063] The germplasm bank construction module 40 is configured to construct a peanut core germplasm bank according to all the sample germplasm resource data.
[0064] According to another aspect of the embodiments of the present invention, there is provided a computer storage medium storing program instructions, wherein when the program instructions run, the device where the computer storage medium is located is controlled to execute the peanut core germplasm bank rapid construction method of any one of the above.
[0065] According to another aspect of the embodiments of the present invention, there is provided a processor for running a program, wherein when the program runs, the peanut core germplasm bank rapid construction method of any one of the above is executed.
[0066] It should be understood that although the steps in the flowcharts of the embodiments of the present invention are shown sequentially according to the arrows, these steps do not necessarily have to be executed in the order indicated by the arrows. Unless there is a clear indication in this article, there is no strict order limit for the execution of these steps, and these steps can be executed in other orders. Moreover, at least some of the steps in each embodiment may include multiple sub-steps or multiple stages. These sub-steps or stages do not necessarily have to be executed at the same time, but can be executed at different times. The execution order of these sub-steps or stages does not necessarily have to be sequential, but can be executed alternately or alternately with at least a part of other steps or sub-steps or stages of other steps.
[0067] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The above program can be stored in a non-volatile computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.
[0068] The above embodiments only express the preferred implementation modes of the present invention. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several deformations and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the invention patent of the present invention should be subject to the appended claims.
Claims
1. A rapid construction method for a peanut core germplasm bank, characterized in that, It includes the following steps: Step 1: Collect the existing data of the first germplasm resource in the whole germplasm resource, preprocess the data of the first germplasm resource, and obtain the data of the second germplasm resource; Step 2: Divide the data of the second germplasm resource into N1 first groups according to botanical types; Step 3: Extract any one of the first groups, divide any one of the first groups into N2 second groups according to geographical origin. After traversing all the first groups, obtain N3 second groups; Step 4: Extract any one of the second groups, perform cluster analysis on any one of the second groups, and generate N4 third groups; Step 5: After traversing all the second groups, extract sample germplasm resource data from each of the third groups respectively; Step 6: Construct a peanut core germplasm bank based on all the sample germplasm resource data.
2. The method according to claim 1, wherein Based on whether the data of the first germplasm resource is complete, perform the preprocessing, and delete the incomplete data. Among them, the data of the first germplasm resource includes basic data, evaluation data, and characteristic data. The basic data includes geographical origin, ecological and geographical conditions of the geographical origin, and breeding system. The evaluation data includes agronomic traits. The characteristic data includes shape markers, biochemical indexes, and DNA molecular marker data.
3. The method according to claim 1, wherein The botanical types include multi-seeded type, pearl bean type, runner type, common type, and intermediate type.
4. The method according to claim 1, characterized in that The clustering method is the UPGMA average linkage method or the Euclidean distance method.
5. The method according to claim 1, characterized in that, In the said Step 5, the method for extracting sample germplasm resource data is: Extract any one of the third groups, and calculate the similarity of all the data of the second germplasm resource in any one of the third groups; Judge whether the similarity is greater than a first preset value. If so, sample according to a first ratio from any one of the third groups. If not, sample according to a second ratio from any one of the third groups, where the first ratio is less than the second ratio.
6. The method according to claim 1, characterized in that, In the said Step 4, perform cluster analysis on any one of the second groups based on agronomic traits.
7. The method according to claim 2 or 6, characterized in that, The agronomic traits include growth period, weight of 100 fruits, weight of 100 kernels, kernel percentage, plant height, plant type, flowering habit, branching type, type, crude protein content, crude fat content, oleic acid content, and linoleic acid content.
8. A rapid construction system for a peanut core germplasm bank, which is used to implement the method described in any one of claims 1 to 7, and is characterized in that, It includes: A data collection module, a data grouping module, a sample extraction module, and a germplasm bank construction module; The data collection module is used to collect the existing data of the first germplasm resource in the whole germplasm resource, preprocess the data of the first germplasm resource, and obtain the data of the second germplasm resource; The data grouping module is used to divide the data of the second germplasm resource into N1 first groups according to botanical types, then extract any one of the first groups, divide any one of the first groups into N2 second groups according to geographical origin. After traversing all the first groups, obtain N3 second groups, and finally extract any one of the second groups, and perform cluster analysis on any one of the second groups based on agronomic traits to generate N4 third groups; The sample extraction module is used to extract sample germplasm resource data from each of the third groups respectively after traversing all the second groups; The germplasm bank construction module is used to construct a peanut core germplasm bank according to all the sample germplasm resource data.
9. A computer storage medium, characterized in that, The computer storage medium stores program instructions, wherein when the program instructions are running, the device where the computer storage medium is located is controlled to execute the method described in any one of claims 1 to 7.
10. A processor, characterized in that, The processor is used to run a program, wherein when the program is running, the method described in any one of claims 1 to 7 is executed.
Citation Information
Patent Citations
Sampling method for improving representativeness of core germplasm of plant genetic resources
CN102890755A
Rice core collection intelligent management system and method
CN105740320A
SNP chip for peanut variety identification and preparation method and application of SNP chip
CN110894540A
Method for constructing basic population library for crop domestication breeding by using natural population
CN115691668A
Screening method of selenium-rich high-yield peanut core germplasm
CN117280909A