A method and system for constructing a core germplasm of millet

By establishing a comprehensive evaluation model and hierarchical cluster analysis, the problems of insufficient comprehensiveness and accuracy in the construction of core millet germplasm in existing technologies have been solved, enabling efficient screening and dynamic management of millet germplasm resources, and improving the utilization efficiency and genetic diversity of germplasm resources.

CN120748476BActive Publication Date: 2026-04-17ANYANG ACAD OF AGRI SCI +2
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
ANYANG ACAD OF AGRI SCI
Filing Date
2025-04-27
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

Existing technologies for constructing core germplasm of millet suffer from insufficient comprehensiveness and precision, making it difficult to efficiently explore and utilize the genetic diversity of millet.

Method used

By collecting and preprocessing multidimensional data of millet germplasm resources, a comprehensive evaluation model was established. The weight values ​​were determined by combining expert experience and the analytic hierarchy process. Hierarchical clustering analysis was used for grouping and classification to screen out core germplasm resources with excellent traits and breeding potential, and their genetic diversity changes were monitored.

Benefits of technology

This approach enables a comprehensive and accurate evaluation of millet germplasm resources, improves the representativeness and genetic diversity of core germplasm, reduces computational complexity, and ensures the dynamic updating and utilization efficiency of germplasm resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120748476B_ABST
    Figure CN120748476B_ABST
Patent Text Reader

Abstract

This application relates to the field of data processing technology and discloses a method and system for constructing core germplasm of millet. The method includes: collecting a large number of millet germplasm resources, recording relevant data of the germplasm resources, and preprocessing the relevant data; determining multiple relevant evaluation indicators, determining the weight values ​​of each relevant evaluation indicator, and establishing a comprehensive evaluation model for millet germplasm resources; grouping the collected millet germplasm resources; classifying the grouped germplasm resources, determining the sampling quantity for each category of germplasm resources classified into the same category, and randomly sampling germplasm resources from each category; evaluating the randomly sampled germplasm resources to obtain a corresponding comprehensive score, and selecting core germplasm resources based on the comprehensive score. This invention can accurately select core germplasm resources.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing technology, and in particular to a method and system for constructing core germplasm of millet. Background Technology

[0002] Millet, as an important food crop, possesses rich genetic diversity. However, the vast germplasm resource population results in low research and utilization efficiency, making it difficult to efficiently explore and utilize its genetic diversity. Therefore, constructing a representative core millet germplasm bank is of great significance for improving germplasm resource utilization efficiency and promoting the development of millet breeding.

[0003] A similar prior art patent application, CN107577920A, provides a method and apparatus for constructing core germplasm of *Cynodon dactylon* using morphological markers. The method includes: using original *Cynodon dactylon* germplasm as experimental material, measuring 11 morphological traits through morphological markers, and screening the optimal strategy for constructing core germplasm of *Cynodon dactylon* from five levels: 7 overall sampling ratios, 3 sampling methods, 5 clustering methods, 3 genetic distances, and 3 intragroup sampling ratios, thereby constructing core germplasm.

[0004] Similar prior art includes Chinese patent application CN118568477A, ​​which provides a method and system for constructing core germplasm of moso bamboo based on morphological traits. The method includes: obtaining moso bamboo germplasm samples, measuring 15 morphological traits of moso bamboo germplasm, and screening target strategies for constructing core germplasm of moso bamboo from four levels: two genetic distances, three sampling methods, six sampling ratios, and four clustering methods, thereby constructing core germplasm.

[0005] However, the core germplasm constructed by the above two methods, which only focus on morphological markers, is relatively weak in terms of comprehensiveness and accuracy. Therefore, this invention provides a method and system for constructing millet core germplasm. Summary of the Invention

[0006] This application provides a method and system for constructing core germplasm of millet, which can be used to construct core germplasm of millet more accurately.

[0007] In a first aspect, this application provides a method for constructing core germplasm of millet, the method comprising:

[0008] Step S1: Collect a large amount of millet germplasm resources, record relevant data of germplasm resources, including phenotypic traits, molecular markers, geographical origin, environmental factors and breeding history information, and preprocess the relevant data according to preset standards;

[0009] Step S2: Determine multiple relevant evaluation indicators, combine multiple methods to determine the weight values ​​of each relevant evaluation indicator, and establish a comprehensive evaluation model for millet germplasm resources based on the weight values;

[0010] Step S3: Based on geographical origin, the collected millet germplasm resources are first grouped to obtain several different first groups. Each first group is then second grouped based on phenotypic morphology to obtain several different second groups.

[0011] Step S4: Perform a first classification on the germplasm resources of each second group, and perform a second classification based on the results of the first classification to obtain multiple different categories. For germplasm resources classified into the same category, determine the sampling quantity for each category by calculation, and randomly sample the germplasm resources of each category based on the sampling quantity.

[0012] Step S5: Evaluate the randomly sampled germplasm resources based on the comprehensive evaluation model to obtain the corresponding comprehensive score, screen core germplasm resources based on the comprehensive score, monitor the genetic diversity of core germplasm resources, and dynamically update the core germplasm resources based on the monitoring results.

[0013] In conjunction with the first aspect, in the first implementation of the first aspect of this application, the weight values ​​of each relevant evaluation indicator are determined, including:

[0014] Obtain the first weight values ​​assigned by the relevant expert team to each relevant evaluation indicator;

[0015] The second weight value of each relevant evaluation indicator was obtained using the analytic hierarchy process.

[0016] Calculate the information entropy corresponding to each relevant evaluation indicator, calculate the corresponding difference coefficient based on the information entropy, add the difference coefficients corresponding to all relevant evaluation indicators to obtain the first result value, and divide the difference coefficient by the first result value as the third weight value of the corresponding relevant evaluation indicator.

[0017] The final weight value is obtained by combining the first weight value, the second weight value, and the third weight value.

[0018] In conjunction with the first aspect, in the second implementation of the first aspect of this application, the germplasm resources of each second group are first classified, including:

[0019] For each second group, relevant data of all germplasm resources are obtained. Each relevant data point includes multiple data parameters. The first analysis is performed on all relevant data in the second group to obtain the first key data parameters. A clustering algorithm is used to classify all germplasm resources based on the first key data parameters to obtain multiple first categories.

[0020] In conjunction with the first aspect, in the third implementation of the first aspect of this application, a second classification is performed based on the first classification result, including:

[0021] The first classification result includes multiple different first categories. For each first category, a first analysis is performed to obtain the corresponding first key data parameters. Based on the first key data parameters and a clustering algorithm, the first category is further divided. This step is repeated until the termination condition is met.

[0022] In conjunction with the first aspect, in the fourth implementation of the first aspect of this application, a first analysis is performed to obtain the first key data parameters, including:

[0023] Calculate the standard deviation of each data parameter, and select data parameters whose standard deviation is greater than the second threshold as candidate data parameters. Calculate the first feature value of the candidate data parameters, and select candidate data parameters whose absolute value of the first feature value is less than the third threshold as the first key data parameters.

[0024] In conjunction with the first aspect, in the fifth implementation of the first aspect of this application, a clustering algorithm is used to classify all germplasm resources based on the first key data parameter, including:

[0025] If the number of germplasm resources in the target data exceeds a preset threshold, the first clustering algorithm is used to classify the target data; otherwise, the second clustering algorithm is used.

[0026] In conjunction with the first aspect, in the sixth implementation of the first aspect of this application, calculating the first feature value of the candidate data parameter includes:

[0027] The first eigenvalue of the candidate parameters is calculated based on the first formula, which is: Where Q is the total number of germplasm resources in the second group, x k These are the values ​​of the candidate parameters for the k-th germplasm resource. s refers to the average value of the candidate parameters, s refers to the standard deviation of the candidate parameters, and P refers to the first eigenvalue of the candidate parameters.

[0028] In conjunction with the first aspect, in the seventh implementation of the first aspect of this application, the sampling amount for each category is determined, including:

[0029] The sample size for each category is calculated based on the second formula, which is: Where N is the total number of germplasm resources, n i It is the total number of germplasm resources within the i-th category, n j N is the total number of germplasm resources within the j-th category, m is the total number of categories, and N is the total number of germplasm resources within the j-th category. i It is the sample size of the i-th category.

[0030] In conjunction with the first aspect, in the eighth implementation of the first aspect of this application, core germplasm resources are selected based on a comprehensive score, including:

[0031] Core germplasm resources with a comprehensive score greater than the preset fourth threshold are used as initial core germplasm resources. Multiple evaluation indicators of the initial core germplasm resources are calculated. If multiple evaluation indicators meet the preset conditions, the initial core germplasm resources are used as the final core germplasm resources. Otherwise, the selected relevant evaluation indicators are adjusted and the comprehensive evaluation model is updated.

[0032] Secondly, this application provides a millet core germplasm construction system, the system comprising:

[0033] The collection module is used to collect a large amount of millet germplasm resources, record relevant data of germplasm resources, including phenotypic traits, molecular markers, geographical origin, environmental factors and breeding history information, and preprocess the relevant data according to preset standards;

[0034] The evaluation module is used to determine multiple relevant evaluation indicators, combine various methods to determine the weight values ​​of each relevant evaluation indicator, and establish a comprehensive evaluation model for millet germplasm resources based on the weight values.

[0035] The grouping module is used to first group the collected millet germplasm resources based on their geographical origin to obtain several different first groups, and then to second group each first group based on its phenotypic shape to obtain several different second groups.

[0036] The clustering module is used to perform the first classification of germplasm resources in each second group, and to perform the second classification based on the results of the first classification, thereby obtaining multiple different categories. For germplasm resources that are classified into the same category, the sampling quantity of each category is determined by calculation, and the germplasm resources of each category are randomly sampled based on the sampling quantity.

[0037] The module is used to evaluate randomly sampled germplasm resources based on a comprehensive evaluation model to obtain corresponding comprehensive scores, screen core germplasm resources based on comprehensive scores, monitor the genetic diversity of core germplasm resources, and dynamically update core germplasm resources based on monitoring results.

[0038] Compared with the prior art, the beneficial effects of the present invention are at least as follows:

[0039] The technical solution provided in this application collects a large amount of multi-dimensional data on millet germplasm resources, including phenotypic traits, molecular markers, geographical origin, environmental factors, and breeding history information, and incorporates this data into a comprehensive evaluation model to achieve a comprehensive and accurate evaluation of germplasm resources. Through comprehensive analysis of different types of data, germplasm resources with superior traits and breeding potential can be identified more accurately, improving the representativeness of core germplasm. The weight values ​​of each evaluation indicator are determined by combining expert experience, the analytic hierarchy process (AHP), and information entropy theory, avoiding the limitations of a single method and making the weight allocation more scientific and reasonable. A hierarchical clustering analysis strategy is adopted, first grouping based on geographical origin and phenotypic shape, then further refining the classification of each group, ultimately classifying similar germplasm resources into the same category. This effectively reduces computational complexity, improves classification efficiency, and ensures the genetic diversity and representativeness of core germplasm. After screening core germplasm, its genetic diversity is monitored using indicators such as the Shannon-Weaver diversity index or genetic distance, and compared with the original germplasm resources. This allows for timely detection of trends in the genetic diversity of core germplasm, providing a basis for the dynamic updating of germplasm resources. Attached Figure Description

[0040] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0041] Figure 1 This is a schematic diagram of an embodiment of a method for constructing core germplasm of millet according to the present application.

[0042] Figure 2 This is a schematic diagram illustrating the determination of the weight values ​​of each relevant evaluation indicator in the embodiments of this application;

[0043] Figure 3 This is a schematic diagram of an embodiment of the second classification based on the first classification result in this application;

[0044] Figure 4 This is a schematic diagram of one embodiment of a millet core germplasm construction system in this application. Detailed Implementation

[0045] This application provides a method and system for constructing millet core germplasm. The terms "first," "second," "third," "fourth," etc. (if present) in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments described herein can be implemented in a sequence other than that illustrated or described herein. Furthermore, the terms "comprising" or "having" and any variations thereof are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0046] For ease of understanding, the specific process of the embodiments of this application is described below. Please refer to [link / reference]. Figure 1 One embodiment of a method for constructing core germplasm of millet in this application includes:

[0047] Step S1: Collect a large amount of millet germplasm resources, record relevant data of germplasm resources, including phenotypic traits, molecular markers, geographical origin, environmental factors and breeding history information, and preprocess the relevant data according to preset standards.

[0048] Specifically, to collect a wider range of millet germplasm resources and increase genetic diversity, a large number of millet germplasm resources are collected from different geographical regions, ecological types, and breeding backgrounds. Relevant data for each germplasm resource is recorded. For example, technologies such as drones, lidar, and image recognition are used to collect large-scale, high-precision phenotypic data of millet germplasm resources. Automated phenotypic measurement tools based on machine vision can also be used to measure traits such as ear length, ear weight, and grain size, collecting more comprehensive and accurate millet phenotypic data. Phenotypic identification is also conducted under different ecological conditions to assess the environmental adaptability of germplasm resources. For example, phenotypic identification is conducted under stress conditions such as drought, salinity, and pests and diseases to obtain the performance of germplasm resources under different environmental conditions. Molecular marker data refers to the genomic information of millet germplasm resources. By collecting molecular marker data, more direct genetic information is provided for the assessment of the genetic diversity of germplasm resources. After collecting relevant data, the data is preprocessed according to preset standards. Preprocessing includes normalizing numerical data and converting non-numerical data to numerical data before normalization.

[0049] Step S2: Determine multiple relevant evaluation indicators, combine multiple methods to determine the weight values ​​of each relevant evaluation indicator, and establish a comprehensive evaluation model for millet germplasm resources based on the weight values.

[0050] Specifically, to more comprehensively evaluate the quality of millet germplasm resources and provide a scientific basis for the selection of core germplasm resource sites, several relevant evaluation indicators were determined. These indicators include various collected data, phenotypic traits (including agronomic traits such as plant height, ear length, ear weight, and growth duration), stress resistance (cold resistance, salt and alkali resistance, disease resistance, and insect resistance), quality traits (nutrient content, grain appearance quality, and processing quality), molecular marker indicators (including genomic diversity such as SNP marker diversity, SSR marker diversity, and genetic distance), and environmental factor indicators (including ecological adaptability and climate adaptability). In addition to other relevant indicators such as nutritional value and functional components, a combination of methods is used to determine the weight values ​​of each relevant evaluation indicator in order to improve the reliability of the evaluation results. The specific methods for determining the weight values ​​will be explained in detail later. After calculating the weight values, a comprehensive evaluation model is established based on the weight values. Specifically, the weight values ​​are multiplied by the data of each relevant evaluation indicator, and then a weighted sum is performed. The result of the weighted sum is used as the comprehensive score of the germplasm resource. Subsequently, the comprehensive score of each germplasm resource can be calculated based on the comprehensive evaluation model, and the germplasm resources can be evaluated based on the comprehensive score to determine their quality level.

[0051] Step S3: Based on geographical origin, the collected millet germplasm resources are first grouped to obtain several different first groups. Each first group is then second grouped based on phenotypic morphology to obtain several different second groups.

[0052] Specifically, to ensure that different germplasm resources are appropriately representative when constructing core germplasm and to avoid over- or under-representation of certain types of germplasm resources, the germplasm resources are first grouped based on their geographical origin to obtain multiple different first groups. For example, germplasm resources can be divided into introduced varieties, local varieties, wild varieties, and cultivated varieties based on their geographical origin. Then, each first group is further grouped based on phenotypic traits to obtain several different second groups. For example, germplasm resources can be divided into multiple different second groups based on the shape of the millet ear, such as conical, cylindrical, and spindle-shaped. After that, cluster analysis is performed on each second group to reduce computational complexity.

[0053] Step S4: Perform a first classification on the germplasm resources of each second group, and perform a second classification based on the results of the first classification to obtain multiple different categories. For germplasm resources classified into the same category, calculate and determine the sampling quantity for each category, and randomly sample the germplasm resources of each category based on the sampling quantity.

[0054] Specifically, in order to improve the representativeness of core germplasm, germplasm resources are classified into a first and second classification based on grouping, further refining the classification of germplasm resources. More similar germplasm resource groups are grouped into the same category. For germplasm resources in the same category, the sampling quantity of each category is determined by calculation. The method for calculating the sampling quantity will be explained in detail later. Based on the sampling quantity, germplasm resources in each category are randomly sampled.

[0055] Step S5: Evaluate the randomly sampled germplasm resources based on the comprehensive evaluation model to obtain the corresponding comprehensive score, screen core germplasm resources based on the comprehensive score, monitor the genetic diversity of core germplasm resources, and dynamically update the core germplasm resources based on the monitoring results.

[0056] Specifically, to accurately screen core germplasm resources, a comprehensive score is calculated based on a comprehensive evaluation model for germplasm resources obtained through random sampling. Core germplasm resources are then screened based on this comprehensive score. The specific screening method will be explained in detail later. After screening core germplasm resources, to promptly detect changes in the genetic diversity of core germplasm resources, monitoring the genetic diversity of core germplasm resources involves using indicators such as the Shannon-Weaver diversity index or genetic distance to assess the genetic diversity of core germplasm resources and comparing it with the original germplasm resources. Core germplasm resources with a declining trend in genetic diversity are identified. If the detection results indicate that the genetic diversity of core germplasm resources has indeed declined, new core germplasm resources are selected to replace them, maintaining the representativeness of the core germplasm resources and ensuring their utilization value.

[0057] In one specific embodiment, determining the weight values ​​of each relevant evaluation indicator includes the following steps:

[0058] Obtain the first weight values ​​assigned by the relevant expert team to each relevant evaluation indicator;

[0059] The second weight value of each relevant evaluation indicator was obtained using the analytic hierarchy process.

[0060] Calculate the information entropy corresponding to each relevant evaluation indicator, calculate the corresponding difference coefficient based on the information entropy, add the difference coefficients corresponding to all relevant evaluation indicators to obtain the first result value, and divide the difference coefficient by the first result value as the third weight value of the corresponding relevant evaluation indicator.

[0061] The final weight value is obtained by combining the first weight value, the second weight value, and the third weight value.

[0062] Specifically, to improve the accuracy and reliability of weight allocation, methods such as... Figure 2The diagram illustrates how to determine the weight values ​​of each relevant evaluation indicator. First, the first weight value of each relevant evaluation indicator is initially determined using expert knowledge and experience. Then, the second weight value is determined using the Analytic Hierarchy Process (AHP), an existing technique that will not be explained further. Next, the information entropy of each relevant evaluation indicator is calculated (the calculation method for information entropy is also an existing technique and will not be explained further). Information entropy measures the dispersion of data; the higher the dispersion, the greater the information content, and the higher the weight of the corresponding relevant evaluation indicator should be. Therefore, the information entropy is subtracted from the data to obtain the corresponding difference coefficient. The sum of the difference coefficients of all relevant evaluation indicators is calculated as the first result value. The difference coefficient is divided by the first result value to obtain the third weight value for each corresponding relevant evaluation indicator. The final weight value is obtained by combining these three weight coefficients. For example, the average of the three weight coefficients can be used as the final weight value, or the weight coefficients can be determined based on the actual situation. For instance, if expert opinions are given more weight, the weight coefficient of the first weight value can be set to 0.4, and the weight coefficients of the second and third weight values ​​can be set to 0.3. Finally, a weighted average is calculated based on the weight value and its corresponding weight coefficient as the final weight value. The above method takes into account the knowledge and experience of experts and utilizes the data characteristics of the relevant evaluation indicators themselves, avoiding the limitations of a single method. It can scientifically determine the weights of relevant evaluation indicators for millet germplasm resources, providing a reliable basis for subsequent work.

[0063] In one specific embodiment, the germplasm resources of each second group are classified, specifically including the following steps:

[0064] For each second group, relevant data of all germplasm resources are obtained. Each relevant data point includes multiple data parameters. The first analysis is performed on all relevant data in the second group to obtain the first key data parameters. A clustering algorithm is used to classify all germplasm resources based on the first key data parameters to obtain multiple first categories.

[0065] Specifically, each second group includes multiple germplasm resources. The relevant data of each germplasm resource includes multiple data parameters, such as plant height, ear length, and ear weight. The first analysis is performed on all relevant data in terms of data parameters to obtain the first key data parameter. The first key data parameter refers to the data parameter with obvious dispersion and uniform distribution. For example, in a certain second group, the ear weight has obvious dispersion and uniform distribution, the ear length has obvious dispersion but uneven distribution, and the plant height has little separation and most of them are similar. Therefore, ear weight is taken as the first key data parameter to facilitate subsequent classification based on key data.

[0066] In one specific embodiment, the second classification based on the first classification result further includes the following steps:

[0067] The first classification result includes multiple different first categories. For each first category, a first analysis is performed to obtain the corresponding first key data parameters. Based on the first key data parameters and a clustering algorithm, the first category is further divided. This step is repeated until the termination condition is met.

[0068] Specifically, in order to identify hierarchical results in germplasm resources, further refined classification is performed on the multiple different first categories obtained after the first classification, using methods such as... Figure 3 The process shown is to perform a second classification based on the results of the first classification. First, a first analysis is performed on each first category to obtain the first key data parameters of the first category. Then, based on the first key data parameters and a clustering algorithm, the first category is further divided. At this point, the categories may not be refined enough. The categories obtained from the second division are taken as the first category. This step is repeated to further divide each first category until the termination condition is met. The termination condition is that the number of iterations is greater than the preset maximum value or the number of germplasm resources contained in the final category is less than the preset value.

[0069] It is important to note that the first key data parameter calculated through analysis will be different depending on the target data for classification. The first key data parameter corresponding to the first category is different from the first key data parameter corresponding to the second group.

[0070] Through the above steps, germplasm resources can be efficiently and accurately classified, so that similar germplasm resources are classified into the same category.

[0071] In one specific embodiment, performing a first analysis to obtain a first key data parameter includes the following steps:

[0072] Calculate the standard deviation of each data parameter, and select data parameters whose standard deviation is greater than the second threshold as candidate data parameters. Calculate the first feature value of the candidate data parameters, and select candidate data parameters whose absolute value of the first feature value is less than the third threshold as the first key data parameters.

[0073] Specifically, to improve the accuracy and efficiency of classification, the second group needs to be classified for the first time based on the first key data parameter. The first key data parameter refers to the data parameter with obvious and uniform data dispersion. Germplasm resource data includes many data parameters, some of which are useful for classification, while others are not. Classifying based on all data parameters would reduce efficiency. Therefore, during classification, it is necessary to first select data parameters with greater data dispersion. First, the standard deviation of the data parameters is calculated, and data parameters with a standard deviation greater than a second threshold are used as candidate data parameters to ensure that the candidate data parameters are differentiated. If the standard deviation is less than or equal to the second threshold, the corresponding data parameters may be more concentrated. Classifying based on this data parameter will not only reduce the accuracy of classification but also reduce its efficiency. To improve efficiency, data parameters with a standard deviation less than the second threshold are ignored during clustering. Then, the first feature value of the candidate data parameters is calculated. This first feature value measures the uniformity of the data parameter distribution. If the data parameter distribution is dispersed but uneven, such as the distribution of plant height being dispersed but uneven (e.g., a certain germplasm resource having a much larger plant height than others, resulting in a calculated standard deviation greater than the second threshold), classification based on plant height would lead to inaccurate classification. Therefore, based on the first feature value and the third threshold, unevenly dispersed data parameters are identified and ignored during classification to improve efficiency.

[0074] It is important to note that the second and third thresholds mentioned above change depending on the target data. The target data refers to the data currently being classified. For example, when classifying all germplasm resources within the second group, the relevant data corresponding to all germplasm resources within the second group is obtained. When performing the first analysis on this relevant data, the target data is all the relevant data corresponding to all germplasm resources included in the second group. If classifying all germplasm resources within the first category, the relevant data corresponding to all germplasm resources within the first category needs to be obtained. When performing the first analysis on this relevant data, the target data is all the relevant data corresponding to all germplasm resources included in the first category.

[0075] Obtaining the second threshold includes: obtaining the median based on all calculated standard deviations, and using the median as the corresponding second threshold. For example, when performing the first classification of the second group, all relevant data corresponding to all germplasm resources within all second groups are obtained, and the corresponding standard deviation is calculated for each data parameter in each relevant data. At this time, multiple standard deviation data can be obtained.

[0076] Obtaining the third threshold includes: obtaining all first feature values ​​and obtaining the median among them, and using the median as the corresponding third threshold.

[0077] In one specific embodiment, a clustering algorithm is used to classify all germplasm resources based on a first key data parameter, including:

[0078] If the number of germplasm resources in the target data exceeds a preset threshold, the first clustering algorithm is used to classify the target data; otherwise, the second clustering algorithm is used.

[0079] Specifically, since different clustering methods have different application scenarios, when the number of germplasm resources in the target data exceeds a preset threshold, indicating a large data volume, the first clustering algorithm, namely the Louvain algorithm, is used to classify the target data. The Louvain algorithm is a hierarchical clustering algorithm based on modularity, which can efficiently handle large-scale data. As the number of classifications increases, the data volume of the target data will gradually decrease. If the Louvain algorithm is used to classify small data volumes, it will reduce the classification efficiency. Therefore, when the number of germplasm resources is less than the preset threshold, indicating a relatively small data volume, the second clustering method, such as the K-means method, which is suitable for classifying small data volumes, is selected for classification.

[0080] In one specific embodiment, calculating the first feature value of the candidate data parameter specifically includes the following steps:

[0081] The first eigenvalue of the candidate parameters is calculated based on the first formula, which is: Where Q is the total number of germplasm resources in the target data, and x k These are the values ​​of the candidate parameters for the k-th germplasm resource. s refers to the average value of the candidate parameters, s refers to the standard deviation of the candidate parameters, and P refers to the first eigenvalue of the candidate parameters.

[0082] Specifically, in order to improve the accuracy and effectiveness of classification, it is necessary to calculate the first feature value of the candidate data parameters. The first feature value is used to measure the uniformity of the distribution of data parameters. The first feature value of the candidate parameters is calculated based on the first formula mentioned above, where Q is the total number of germplasm resources in the target data. For example, if the target data is the second group, then Q is the total number of germplasm resources contained in the second group.

[0083] In one specific embodiment, determining the sample size for each category includes the following steps:

[0084] The sample size for each category is calculated based on the second formula, which is: Where N is the total number of germplasm resources, ni is the total number of germplasm resources in the i-th category, nj is the total number of germplasm resources in the j-th category, m is the total number of categories, and Ni is the sampling quantity of the i-th category.

[0085] Specifically, to ensure sample representativeness while reducing sample size and improving germplasm resource utilization efficiency, the sampling size for each category is calculated based on the second formula mentioned above. Here, N is the total number of germplasm resources (e.g., if 1000 germplasm resources are collected, N is 1000), ni is the total number of germplasm resources in the i-th category (e.g., if the first category contains 25 germplasm resources, n1 is 15), nj is the total number of germplasm resources in the j-th category, m is the total number of categories (e.g., if there are 30 categories, m is 30), and Ni is the sampling size for the i-th category. This formula ensures that the sampling size for each category is proportional to the square root of the germplasm resources within that category, while also considering the proportional relationships between different categories, thus guaranteeing the overall sampling balance and representativeness.

[0086] In one specific embodiment, the selection of core germplasm resources based on comprehensive scoring includes the following steps:

[0087] Core germplasm resources with a comprehensive score greater than the preset fourth threshold are used as initial core germplasm resources. Multiple evaluation indicators of the initial core germplasm resources are calculated. If multiple evaluation indicators meet the preset conditions, the initial core germplasm resources are used as the final core germplasm resources. Otherwise, the selected relevant evaluation indicators are adjusted and the comprehensive evaluation model is updated.

[0088] Specifically, to accurately select core germplasm resources, core germplasm resources with a comprehensive score greater than a preset fourth threshold are first selected as initial core germplasm resources. Then, multiple evaluation indicators for the initial core germplasm resources are calculated. These indicators include genetic diversity index, allele retention ratio, and genetic distance. If multiple evaluation indicators meet preset conditions, such as the difference between the evaluation indicators of the original germplasm resources and the preset corresponding thresholds, it indicates that the selected initial core germplasm resources are relatively representative and are selected as the final core germplasm resources. Otherwise, it indicates that the initial core germplasm resources are not very representative, possibly because some of the selected relevant evaluation indicators are not highly relevant to the breeding objectives. The selected relevant evaluation indicators can be adjusted, such as deleting the relevant evaluation indicators or adjusting their corresponding weights, and updating the corresponding comprehensive evaluation model. New initial core germplasm resources are then selected based on the updated comprehensive evaluation model.

[0089] The above describes a method for constructing millet core germplasm in embodiments of this application. The following describes a system for constructing millet core germplasm in embodiments of this application. Please refer to [link to relevant documentation]. Figure 4One embodiment of a millet core germplasm construction system in this application includes:

[0090] The collection module is used to collect a large amount of millet germplasm resources, record relevant data of germplasm resources, including phenotypic traits, molecular markers, geographical origin, environmental factors and breeding history information, and preprocess the relevant data according to preset standards;

[0091] The evaluation module is used to determine multiple relevant evaluation indicators, combine various methods to determine the weight values ​​of each relevant evaluation indicator, and establish a comprehensive evaluation model for millet germplasm resources based on the weight values.

[0092] The grouping module is used to first group the collected millet germplasm resources based on their geographical origin to obtain several different first groups, and then to second group each first group based on its phenotypic shape to obtain several different second groups.

[0093] The clustering module is used to perform the first classification of germplasm resources in each second group, and to perform the second classification based on the results of the first classification, thereby obtaining multiple different categories. For germplasm resources that are classified into the same category, the sampling quantity of each category is determined by calculation, and the germplasm resources of each category are randomly sampled based on the sampling quantity.

[0094] The module is used to evaluate randomly sampled germplasm resources based on a comprehensive evaluation model to obtain corresponding comprehensive scores, screen core germplasm resources based on comprehensive scores, monitor the genetic diversity of core germplasm resources, and dynamically update core germplasm resources based on monitoring results.

[0095] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0096] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0097] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.

Claims

1. A method for constructing core germplasm of millet, characterized in that, The method includes: Step S1: Collect a large amount of millet germplasm resources and record relevant data of the germplasm resources, including phenotypic traits, molecular markers, geographical origin, environmental factors and breeding history information. Preprocess the relevant data according to preset standards. Among them, a large amount of millet germplasm resources are collected from different geographical regions, ecological types and breeding backgrounds. Step S2: Determine multiple relevant evaluation indicators, combine various methods to determine the weight values ​​of each relevant evaluation indicator, and establish a comprehensive evaluation model for millet germplasm resources based on the weight values; determining the weight values ​​of each relevant evaluation indicator includes: obtaining the first weight values ​​assigned by relevant expert teams to each relevant evaluation indicator; obtaining the second weight values ​​of each relevant evaluation indicator using the analytic hierarchy process (AHP); calculating the information entropy corresponding to each relevant evaluation indicator, calculating the corresponding difference coefficient based on the information entropy, adding the difference coefficients corresponding to all relevant evaluation indicators to obtain the first result value, and using the difference coefficient divided by the first result value as the third weight value of the corresponding relevant evaluation indicator; and integrating the first weight value, the second weight value, and the third weight value to obtain the final weight value. Step S3: Based on geographical origin, the collected millet germplasm resources are first grouped to obtain several different first groups. Each first group is then second grouped based on phenotypic morphology to obtain several different second groups. Step S4: Perform a first classification on the germplasm resources of each second group, and perform a second classification based on the results of the first classification to obtain multiple different categories. For germplasm resources classified into the same category, calculate and determine the sampling quantity for each category, and randomly sample the germplasm resources of each category based on the sampling quantity. The first classification of germplasm resources in each second group includes: obtaining relevant data of all germplasm resources contained in each second group, with each relevant data point including multiple data parameters; performing a first analysis on all relevant data in the second group to obtain first key data parameters; and using a clustering algorithm to classify all germplasm resources based on the first key data parameters to obtain multiple first categories. Step S5: Evaluate the randomly sampled germplasm resources based on the comprehensive evaluation model to obtain the corresponding comprehensive score, screen core germplasm resources based on the comprehensive score, monitor the genetic diversity of core germplasm resources, and dynamically update the core germplasm resources based on the monitoring results. The first analysis was conducted to obtain the first key data parameters, including: Calculate the standard deviation of each data parameter, and select data parameters whose standard deviation is greater than the second threshold as candidate data parameters. Calculate the first feature value of the candidate data parameters, and select candidate data parameters whose absolute value of the first feature value is less than the third threshold as the first key data parameters. Calculate the first eigenvalue of the candidate data parameters, including: The first eigenvalue of the candidate parameters is calculated based on the first formula, which is: Where Q is the total number of germplasm resources in the second group, and x k These are the values ​​of the candidate parameters for the k-th germplasm resource. s refers to the average value of the candidate parameters, s refers to the standard deviation of the candidate parameters, and P refers to the first eigenvalue of the candidate parameters.

2. The method for constructing core germplasm of millet according to claim 1, characterized in that, The second classification, based on the results of the first classification, also includes: The first classification result includes multiple different first categories. For each first category, a first analysis is performed to obtain the corresponding first key data parameters. Based on the first key data parameters and a clustering algorithm, the first category is further divided. This step is repeated until the termination condition is met.

3. The method for constructing core germplasm of millet according to claim 1, characterized in that, All germplasm resources were classified using a clustering algorithm based on the first key data parameter, including: If the number of germplasm resources in the target data exceeds a preset threshold, the first clustering algorithm is used to classify the target data; otherwise, the second clustering algorithm is used.

4. The method for constructing core germplasm of millet according to claim 1, characterized in that, Determine the sample size for each category, including: calculating the sample size for each category based on the second formula, which is: Where N is the total number of germplasm resources, n i It is the total number of germplasm resources within the i-th category, n j N is the total number of germplasm resources within the j-th category, m is the total number of categories, and N is the total number of germplasm resources within the j-th category. i It is the sample size of the i-th category.

5. The method for constructing core germplasm of millet according to claim 1, characterized in that, Core germplasm resources were selected based on comprehensive scoring, including: Core germplasm resources with a comprehensive score greater than the preset fourth threshold are used as initial core germplasm resources. Multiple evaluation indicators of the initial core germplasm resources are calculated. If multiple evaluation indicators meet the preset conditions, the initial core germplasm resources are used as the final core germplasm resources. Otherwise, the selected relevant evaluation indicators are adjusted and the comprehensive evaluation model is updated.

6. A millet core germplasm construction system, used to implement the millet core germplasm construction method as described in any one of claims 1-5, characterized in that, The system includes: The collection module is used to collect a large amount of millet germplasm resources, record relevant data of germplasm resources, including phenotypic traits, molecular markers, geographical origin, environmental factors and breeding history information, and preprocess the relevant data according to preset standards; The evaluation module is used to determine multiple relevant evaluation indicators, combine various methods to determine the weight values ​​of each relevant evaluation indicator, and establish a comprehensive evaluation model for millet germplasm resources based on the weight values. The grouping module is used to first group the collected millet germplasm resources based on their geographical origin to obtain several different first groups, and then to second group each first group based on its phenotypic shape to obtain several different second groups. The clustering module is used to perform the first classification of germplasm resources in each second group, and to perform the second classification based on the results of the first classification, thereby obtaining multiple different categories. For germplasm resources that are classified into the same category, the sampling quantity of each category is determined by calculation, and the germplasm resources of each category are randomly sampled based on the sampling quantity. The module is used to evaluate randomly sampled germplasm resources based on a comprehensive evaluation model to obtain corresponding comprehensive scores, screen core germplasm resources based on comprehensive scores, monitor the genetic diversity of core germplasm resources, and dynamically update core germplasm resources based on monitoring results.

Citation Information

Patent Citations

  • Method and device for constructing Cynodon dactylon core germplasms through morphological marking

    CN107577920A

  • Method and system for constructing phyllostachys pubescens core germplasm based on morphological characters

    CN118568477A

  • Plant breeding materials screening method and system

    CN103793850A

  • Sweet potato germplasm resource evaluation method based on multi-criterion decision

    CN113780845A