Efficient breeding decision support system and method based on Chinese trumpet creeper
By normalizing and screening the multidimensional phenotypic data in vetch breeding, and combining adaptive weighting algorithms and genetic algorithms, a phenotypic feature matrix and a breeding decision model are constructed. This solves the problem of incomplete data in traditional breeding and improves breeding efficiency and accuracy.
Patent Information
- Application Number
- CN202511508069.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-22
- Publication Date
- 2025-11-21
AI Technical Summary
Traditional vetch breeding suffers from incomplete phenotypic data collection, resulting in a lack of complete information to support the analysis. Differences in dimensionality affect data comparison, stability assessment is insufficient, the association analysis between genotype and phenotype is not in-depth, the breeding decision model lacks reliability, and efficiency is low.
By acquiring multi-dimensional phenotypic data, performing normalization processing and coefficient of variation screening, constructing a phenotypic feature matrix, using an adaptive weighted algorithm for dimensionality reduction, and combining it with a genetic algorithm optimization module, calculating the contribution of gene loci, constructing a vetch breeding decision model, and outputting the optimal breeding combination scheme.
It enables the scientific processing of multi-dimensional phenotypic data, eliminates unstable indicators, improves data accuracy and usability, and conducts in-depth analysis of the association between genotype and phenotype, ensuring the relevance and reliability of the breeding decision model and improving breeding efficiency and accuracy.
Smart Images

Figure CN120994983A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of Vicia seed breeding, in particular to a Vicia seed efficient breeding decision support system and method. BACKGROUND
[0002] As a crop with multiple values such as forage and green manure in agricultural production, the improvement and optimization of Vicia seed varieties have always been the focus of the agricultural technology field. In the current Vicia seed breeding practice, the traditional breeding mode faces many problems to be solved and cannot meet the demand of modern agriculture for breeding efficiency and precision. In the traditional Vicia seed breeding process, there are obvious limitations in the collection and utilization of phenotype data. Breeders often only focus on a few intuitive phenotype indicators, such as single yield data or plant height information, without fully covering multi-dimensional phenotype data such as plant height, leaf area, flowering period and yield indicators. This one-sided data collection method leads to a lack of complete information support for breeding analysis, making it difficult to grasp the correlation between Vicia seed growth and yield formation from a holistic perspective, and thus affecting the comprehensive evaluation of breeding materials. Even if part of the breeding work tries to collect multi-dimensional phenotype data, there are significant defects in the data processing link. Due to the difference in the dimension of different phenotype indicators, such as plant height in centimeters, leaf area in square centimeters, and yield in kilograms, directly using these data for analysis will seriously interfere with the comparison and correlation calculation between data, making the analysis results deviate from the actual situation and unable to accurately reflect the real impact of each phenotype indicator on Vicia seed growth. At the same time, the traditional method does not effectively evaluate the stability of the phenotype indicators, and only determines the analysis object according to single or a few test data, ignoring the fluctuation of the phenotype indicators under different environmental conditions, resulting in a large uncertainty in the analysis foundation constructed subsequently. In terms of phenotype feature selection and processing, traditional breeding lacks scientific and effective technical means. In the face of multi-dimensional phenotype data collected, breeders rely on experience to select features, which is difficult to eliminate redundant and unstable indicators, and thus cannot construct a high-quality phenotype feature matrix. This experience-dependent selection method not only has low efficiency, but also may miss phenotype features that have a key impact on breeding results, posing a hidden danger to subsequent breeding analysis. When traditional Vicia seed breeding combines with genotype data for research, it lacks systematic algorithm support. Breeders cannot accurately analyze the key information in the genotype data, and cannot calculate the contribution of different gene loci to the phenotype characteristics, resulting in insufficient correlation analysis between genotype and phenotype. This situation makes the breeding decision model constructed lack reliable genetic support, and the accuracy and pertinence of the model are insufficient, making it difficult to output the optimal breeding combination scheme according to the actual breeding demand, ultimately leading to a prolonged breeding cycle, low efficiency, and the inability to cultivate excellent varieties to meet the needs of agricultural production in a timely manner. SUMMARY
[0003] The present application aims to provide a high-efficiency breeding decision support system and method based on Vicia, to solve the problems raised in the background art.
[0004] To achieve the above-mentioned purpose, the present application provides a high-efficiency breeding decision support system based on Vicia, which comprises: Obtaining multi-dimensional phenotype data in Vicia breeding test, the multi-dimensional phenotype data including plant height, leaf area, flowering period and yield index; Normalizing the multi-dimensional phenotype data to eliminate the influence of different dimensions, and calculating the coefficient of variation of each phenotype index; Screening phenotype indexes with higher stability than a preset threshold based on the coefficient of variation, and constructing a Vicia phenotype feature matrix; Using an adaptive weighted algorithm to reduce the dimensionality of the Vicia phenotype feature matrix, and extracting key phenotype features; Inputting the key phenotype features into a genetic algorithm optimization module, and calculating the contribution of each genetic locus in combination with the Vicia genotype data; Constructing a Vicia breeding decision model according to the contribution of the genetic locus, and outputting the optimal breeding combination scheme.
[0005] Preferably, the normalizing the multi-dimensional phenotype data comprises: Calculating the maximum and minimum values of each phenotype index, and mapping each phenotype index to the interval of zero to one using a linear normalization method; Calculating the coefficient of variation of the normalized phenotype index, the coefficient of variation being the ratio of the standard deviation to the mean; Eliminating phenotype indexes with a coefficient of variation exceeding a preset fluctuation range, and retaining phenotype indexes with required stability.
[0006] Preferably, the screening of phenotype indexes with higher stability than a preset threshold based on the coefficient of variation comprises: Setting a coefficient of variation threshold, and screening phenotype indexes with a coefficient of variation lower than the threshold; Calculating the Pearson correlation coefficient between the screened phenotype indexes, and removing redundant indexes with a correlation higher than a redundancy threshold; Ordering the remaining phenotype indexes by weight, and constructing a Vicia phenotype feature matrix.
[0007] Preferably, the reducing the dimensionality of the Vicia phenotype feature matrix using an adaptive weighted algorithm comprises: Calculating the initial feature importance score according to the weight of the phenotype index; Adjusting the feature importance score using an iterative optimization method until the score converges; Select the phenotype indicators with importance score higher than the dimensionality reduction threshold as key phenotype features.
[0008] Preferably, the inputting the key phenotype features into the genetic algorithm optimization module comprises: Resolving the medicago polymorpha genotype data to extract single nucleotide polymorphism sites; Calculating the association strength of each single nucleotide polymorphism site and the key phenotype features; Based on the association strength, an initial population is constructed, and a genetic algorithm is used to iteratively optimize the gene site combination.
[0009] Preferably, the calculating the contribution degree of each gene site based on the medicago polymorpha genotype data comprises: Statistically analyzing the allele frequency of each single nucleotide polymorphism site in the breeding population; Calculating the regression coefficient of the allele frequency and the key phenotype features; According to the regression coefficient and the allele frequency, the genetic contribution degree of each gene site is calculated.
[0010] Preferably, the constructing the medicago polymorpha breeding decision model according to the contribution degree of the gene site comprises: Screening the gene sites with genetic contribution degree higher than the significance threshold; Based on the screened gene sites, a medicago polymorpha breeding decision tree model is constructed; The parameters of the decision tree model are optimized by using the cross-validation method.
[0011] Preferably, the outputting the optimal breeding combination scheme comprises: According to the decision tree model, the phenotype performance of different genotype combinations is predicted; The comprehensive breeding index of each genotype combination is calculated; The genotype combination with the highest comprehensive breeding index is selected as the optimal breeding scheme.
[0012] Preferably, the calculating the comprehensive breeding index of each genotype combination comprises: Setting the weight coefficient of each phenotype indicator, calculating the weighted score of the genotype combination on each phenotype indicator, and generating the comprehensive breeding index by summarizing the weighted score.
[0013] Preferably, the system further comprises a medicago polymorpha high-efficiency breeding decision support method, which comprises all the modules and method processes of the above-mentioned medicago polymorpha high-efficiency breeding decision support system.
[0014] Compared with the prior art, the beneficial effects of the present application are: In the data processing link, the system normalizes the multi-dimensional phenotype data, can eliminate the influence brought by different dimensions, makes the data of different phenotype indexes in the same dimension which can be directly compared and analyzed, calculates the variation coefficient of each phenotype index, provides a key reference for subsequent phenotype index screening, makes the processing of phenotype data more scientific and standardized, and improves the usability and accuracy of the data itself. Screening the phenotype indexes with higher stability than the preset threshold based on the variation coefficient and constructing the phenotype feature matrix of the vetch can accurately eliminate the phenotype indexes with insufficient stability, ensure that the phenotype indexes contained in the constructed feature matrix have high reliability and representativeness, avoid the problem that the subsequent analysis results are deviated due to the interference of unstable indexes, and lay a good foundation for subsequent feature dimension reduction and key feature extraction. Adopting the adaptive weighted algorithm to reduce the dimension of the vetch phenotype feature matrix can effectively extract key phenotype features, remove redundant information and non-key features in the matrix, reduce the complexity of subsequent data processing, reduce unnecessary calculation amount, and make the extracted key features more reflect the core needs of vetch breeding, thereby improving the pertinence and efficiency of feature analysis. The key phenotype features are input into the genetic algorithm optimization module, and the contribution of each gene locus is calculated in combination with the vetch genotype data, and with the advantage of genetic algorithm, the association between the key phenotype features and the genotype data can be accurately analyzed, and the role of each gene locus in the breeding process is clearly defined, so that the analysis of the gene locus is more in-depth and accurate, and key gene-level information is provided for subsequent construction of a breeding decision model. According to the contribution of the gene locus, a vetch breeding decision model is constructed and an optimal breeding combination scheme is output, so that the constructed model can fully combine the key information of the gene locus, ensure that the model has strong pertinence and reliability, and the output optimal breeding combination scheme can closely meet the actual needs of vetch breeding, provide clear and feasible direction for breeding personnel to carry out specific breeding work, and help to promote the orderly and efficient development of vetch breeding work, speed up the cultivation process of excellent vetch varieties, and better meet the demand of agricultural production for high-quality vetch varieties. BRIEF DESCRIPTION OF DRAWINGS
[0015] Figure 1 The timing diagram of the efficient vetch breeding decision support system based on the vetch is described in the present application; Figure 2 The working principle diagram of the multi-dimensional phenotype data normalization processing is described in the present application; Figure 3 The working principle diagram of the adaptive weighted algorithm dimension reduction is described in the present application; Figure 4 The vetch breeding decision optimization analysis diagram is described in the present application. DETAILED DESCRIPTION
[0016] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by a person of ordinary skill in the art without creative effort belong to the scope of protection of the present application.
[0017] Please refer to Figure 1 The present application provides a high-efficiency breeding decision support system and method based on Vicia, which integrates multi-dimensional phenotype data and genotype data, uses data preprocessing, feature screening and optimization algorithm to realize intelligent decision of Vicia breeding process. The core of the system is to analyze the stability of the phenotype indexes in Vicia breeding test and reduce the dimension, combined with genetic algorithm to mine key gene sites, and finally output the optimal breeding combination scheme. The overall implementation scheme is as follows: first, obtain the multi-dimensional phenotype data in Vicia breeding test, including plant height, leaf area, flowering period and yield index and other specific parameters, which come from field test or laboratory measurement, to ensure the coverage of phenotype variation in the whole growth cycle of Vicia; then normalize the multi-dimensional phenotype data to eliminate the influence of different dimensions on analysis, and calculate the coefficient of variation of each phenotype index, which is a key parameter for measuring the stability of the index, and the phenotype index with smaller fluctuation is screened out through the preset threshold to construct the Vicia phenotype feature matrix; then the adaptive weighted algorithm is used to reduce the dimension of the phenotype feature matrix, which dynamically adjusts the feature importance according to the index weight, and extracts the phenotype features that have key influence on breeding decision; input the key phenotype features into the genetic algorithm optimization module, and combine the Vicia genotype data to analyze single nucleotide polymorphism sites, calculate the correlation strength of each gene site and phenotype feature, and then evaluate the contribution degree of the gene site; based on the contribution degree result, a Vicia breeding decision model is constructed, which predicts the phenotype performance by optimizing the genotype combination, and finally outputs the optimal breeding combination scheme with the highest comprehensive breeding index. The system realizes the full-process automation from data collection to decision output, and improves the efficiency and accuracy of Vicia breeding.
[0018] Embodiment 1: Please refer to Figure 2After acquiring the multi-dimensional phenotypic data in the field experiment of Vicia sativa breeding, the data processing procedure is started. The multi-dimensional phenotypic data includes plant height, leaf area, flowering period and various yield indicators, which are derived from the regular observation records in the field experiment and the precise measurements in the laboratory. Different indicators in the original data set often have different dimensions and numerical ranges, for example, plant height is measured in centimeters and leaf area is measured in square centimeters. Such differences in dimensions will directly affect the accuracy of subsequent comparative analysis and model construction, so they must be preprocessed to eliminate their effects. Normalization is the core step in this preprocessing stage, and the linear normalization method is used to project the numerical values of each phenotypic indicator into a unified, dimensionless numerical range.
[0019] The specific operation of linear normalization is to process each phenotypic indicator one by one. First, scan all the observation values under this indicator and identify the maximum and minimum values, which define the original distribution range of the indicator values. Then apply the linear transformation formula to convert each original observation value into a new value between zero and one. The conversion process is linear and can maintain the relative distance relationship of the numerical points in the original data unchanged. After this processing, all phenotypic indicators are constrained in the same numerical range, and different indicators such as plant height and leaf area can be directly compared and operated on an equal basis, creating fair conditions for subsequent stability evaluation. After completing the normalization conversion of the data, the analysis focuses on evaluating the stability of each phenotypic indicator, and the coefficient of variation is selected as the key statistical quantity to measure stability. The calculation of the coefficient of variation requires first obtaining the standard deviation and arithmetic mean of the normalized data sequence of each indicator. The standard deviation represents the dispersion degree of the data around the mean, while the mean represents the central tendency of the data. The coefficient of variation is the ratio of the standard deviation to the mean, which is a relative indicator. Its advantage is that it can eliminate the interference of the mean size itself on the judgment of dispersion, and is particularly suitable for comparing the fluctuation of data sets with different units or large differences in mean values. The smaller the calculated coefficient of variation value, the more stable the performance of the phenotypic indicator in the breeding population, and the less affected by environmental or random errors.
[0020] The phenotypic indicators are screened according to the coefficient of variation, which sets a clear preset fluctuation range as the screening threshold. The threshold is a value determined in advance, which may be based on the analysis and summary of a large amount of historical breeding data, or may incorporate the prior knowledge of experts in the field, with the purpose of defining an acceptable stability boundary. The system compares the coefficient of variation of each phenotypic indicator with the preset threshold, and those indicators with a coefficient of variation exceeding the threshold, i.e. those with excessive fluctuation, are considered to have too much noise or insufficient reliability and are removed from the feature set. Only those phenotypic indicators with a coefficient of variation below the threshold, i.e. those with stability meeting the preset requirements, are retained. This screening mechanism effectively filters out unstable observation information, improving the overall quality of the data set. After initially screening out a set of indicators with higher stability, the set needs to be further optimized to eliminate information redundancy. Calculating the Pearson correlation coefficient is an important means to achieve this purpose. The Pearson correlation coefficient is used to quantify the linear correlation between any two retained indicators, with a value ranging from -1 to 1. The closer the absolute value is to 1, the stronger the linear correlation. The system calculates the correlation coefficient for all pairs of indicators and forms a correlation coefficient matrix. By examining this matrix, it can be identified which pairs of indicators have a high degree of correlation, which means that they may reflect the same biological characteristics of the plant from different sides, resulting in redundant expression of information.
[0021] To refine the feature set, a redundancy threshold is set to determine highly correlated pairs of indicators. When the absolute value of the correlation coefficient between two indicators exceeds this redundancy threshold, it is considered that there is redundancy between them. The strategy for handling redundancy is to retain one of the two highly correlated indicators, and the usual basis for selection is the weight of the indicator itself or its biological significance. For example, the indicator that is more directly related to the breeding goal or has a higher weight score is retained. By systematically checking and removing these redundant information, the problem of multicollinearity in the feature matrix can be avoided, which helps to improve the efficiency of subsequent modeling and the interpretability of the model. After the above two steps of stability screening and redundancy removal, a refined and representative subset of phenotypic indicators is obtained. The final step is to sort these indicators by weight and construct the final feature matrix of the phenotypic characteristics of the plant. The allocation of weights can have different bases, for example, the reciprocal of the coefficient of variation can be directly used as the weight. Because the indicator with a small coefficient of variation has high stability, its reciprocal weight also increases accordingly, emphasizing its importance. Or the weight can be directly specified by the breeding expert according to the importance of the actual breeding goal. The indicators are sorted by weight from high to low, so that the key features are in a prominent position in the matrix. The final constructed phenotypic feature matrix is a two-dimensional data table, whose rows represent different individuals or lines of the plant, and the columns represent the phenotypic indicators retained after multiple rounds of screening, sorted by weight. The matrix structure is clear and the data quality is high, providing ideal input data for subsequent feature dimensionality reduction and genetic analysis.
[0022] Example 2: see Figure 3 After constructing the phenotype feature matrix of the plant, the system enters the feature dimensionality reduction phase to extract the most representative key phenotypic features. The adaptive weighting algorithm is the core of this process, which achieves the dimensionality reduction goal by dynamically evaluating and adjusting the importance of each feature. The algorithm starts with reading and using the initial weights of each indicator in the phenotype feature matrix. These weights may come from the reciprocal of the coefficient of variation calculated in the previous processing steps, or they may be fixed values pre-set by breeding experts based on professional experience. The algorithm calculates the initial importance score of each feature based on these weights. The calculation of the initial score is not simply copying the weight, but a comprehensive evaluation of the distribution of the feature in the matrix and its preliminary relationship with other features, so as to provide a reasonable starting point for subsequent iterative optimization. After obtaining the initial importance score, the algorithm enters the core iterative optimization phase, which aims to continuously correct the bias of the initial score through loop calculation, so that it can better reflect the actual contribution of each phenotypic feature to breeding decision-making. The iterative process usually uses gradient descent or similar stochastic optimization strategies. In each iteration, the algorithm re-evaluates the performance of the entire feature set based on the importance score of the current round, and calculates a target function value. The target function is often related to the discriminability, independence of the features, or their relevance to the pre-set breeding goal.
[0023] The algorithm analyzes the trend of the objective function and fine-tunes the importance score of each feature in the direction that optimizes the function, with the magnitude of adjustment controlled by a parameter called learning rate. Setting the learning rate too large can cause the optimization process to oscillate around the optimal solution, while setting it too small can make the convergence too slow. The iterative optimization continues for multiple rounds until a preset convergence condition is met, which is usually set as the sum of the absolute value of the change in importance score of all features in consecutive iterations being less than a small tolerance threshold. Once the convergence condition is met, the algorithm has found a relatively stable feature importance evaluation result, and the scores no longer change significantly, indicating that the model has reached an optimal or near-optimal state. Throughout the iteration process, the algorithm not only considers the independent value of a single feature, but also implicitly considers the interaction effects and potential redundancy between features. For example, even if two features are important individually, if the information they carry is highly overlapping, the algorithm may lower the score of one of them to avoid redundant calculations during iteration.
[0024] When the iteration process is complete and the final importance scores are obtained, the last step of dimensionality reduction is to filter the features based on these scores, which involves a parameter called the dimensionality reduction threshold. The dimensionality reduction threshold is a manually set critical value, and its specific value can be adjusted according to the data characteristics and analysis goals of this breeding trial. For example, it can be set to the upper quartile or a percentage of the final score range. The system will compare the final importance score of each phenotypic feature with this dimensionality reduction threshold one by one, and only those features with scores higher than the threshold will be retained and recognized as key phenotypic features in this round of analysis. The initial feature importance score is calculated based on the weight of the phenotypic indicator, for example, the weight can be multiplied by the standard deviation of the indicator data to consider both prior importance and data dispersion. The feature importance score is adjusted using an iterative optimization method, which can use a stochastic gradient descent method with momentum, with a momentum coefficient of 0.9. The learning rate is gradually reduced from 0.01 through a linear decay strategy until the sum of the squares of the score changes is less than the preset tolerance threshold (e.g., 1e-6), indicating convergence.
[0025] Selecting the phenotypic indicators with importance scores higher than a dimensionality reduction threshold, which can be set as the median of all feature importance scores after convergence, thus dynamically retaining half of the relatively important features. These selected key features, such as plant height, leaf area, etc., constitute a refined feature subset that maximally retains the most relevant information to breeding goals in the original data while significantly reducing the dimensionality of the data. After dimensionality reduction is completed through the adaptive weighting algorithm, the resulting reduced feature set will serve as a bridge and be directly input into the subsequent genetic algorithm optimization module, which will correlate these phenotypic features with genotypic data. This dimensionality reduction greatly reduces the computational burden of subsequent genetic analysis, and since unimportant redundant features are eliminated, it is possible to improve the accuracy and interpretability of genotype-phenotype correlation analysis. The entire adaptive weighting dimensionality reduction process embodies the idea of data-driven decision-making, automatically completing feature evaluation and selection through algorithms, reducing the interference of subjective judgment.
[0026] Suppose a breeding team is working on a sorghum breeding project in the hilly areas of the south, and has collected field observation data from two hundred different strains, which are planted in three test sites with subtle soil and climate differences. After preliminary sorting, the phenotypic feature matrix contains thirteen observation indicators, in addition to core indicators such as plant height, leaf area, main stem flowering period, and single plant pod weight, it also includes more detailed traits such as branch number, internode length, leaf chlorophyll content, etc. Based on experience and preliminary judgment of local planting goals, breeding experts assign initial importance weights to the thirteen indicators, for example, considering the need for drought resistance in hilly areas, they set the weights of leaf area and leaf thickness, which may be related to water use efficiency, slightly higher. After the system reads the matrix containing thirteen features and their initial weights, it begins the adaptive weighting analysis process. It first calculates the initial importance score of each feature based on the combined data of two hundred strains and three test sites, this score not only considers the expert weight, but also introduces a stability measure of each indicator between different test sites. For example, if the numerical value of a certain indicator fluctuates little between three test sites, its stability bonus is high, if a indicator shows huge differences in performance at different locations, its score will be suppressed.
[0027] After the first round of computation, a few indicators such as branch number, plant height, and pod weight per plant, which have higher expert weights and good stability, have obtained leading initial scores. The system enters a stage of repeated fine-tuning, which simulates the process of repeated deliberation and adjustment of the breeder to the data. The system program will virtually try to slightly raise or lower the score of a certain feature based on the current score situation, and then observe whether this adjustment can make the entire feature set perform better in distinguishing different line populations. For example, it may tentatively increase the score of the "first flowering node" feature, and then check whether those lines known to have excellent overall performance can be more clearly clustered together in the reduced dimension feature space under the new score distribution. Each tentative adjustment is accompanied by a virtual "evaluation", if the adjustment brings better discrimination effect, the system will keep the direction of the score adjustment; if the effect is worse, it will be returned or try another direction. This process is silent and fast, and has experienced hundreds of such fine-tuning attempts in the background.
[0028] After multiple rounds of internal exploration and adjustment, the system finds that when the scores of "branch number", "pod weight per plant", "leaf thickness" and "flowering period duration" are significantly higher than those of other features, the virtual line discrimination effect is best. Further adjustment of the score cannot bring obvious improvement, the system determines that the analysis has stabilized and reached a state of convergence. It determines the above four indicators as the key phenotypic features of this round of analysis according to an internal filtering rule, such as only retaining the top thirty percent of features by score. This result is not entirely consistent with the initial expectations of the breeding experts, for example, the "leaf area" that the experts value is devalued because it fluctuates greatly at different test points, while the "flowering period duration" that was not previously particularly valued stands out because of its excellent stability and discrimination ability. The system marks these four key features and prepares to pass them to the next analysis module. For the breeding team, this result means that in subsequent genotype association analysis, they can focus on exploring genetic markers related to these four key traits, rather than dispersing resources in the massive data of all thirteen traits.
[0029] Example 3: After obtaining the key phenotypic features extracted via the adaptive weighting algorithm, these features are systematically inputted into the genetic algorithm optimization module, whose core task is to establish the deep connection between the key phenotypic features and the genotypic data of the foxtail millet, and to explore the optimal combination of genetic loci. The processing of genotypic data is the first step, and the original genotypic data is usually derived from high-throughput sequencing platforms or gene chip technologies, presenting as massive nucleotide sequence variation information. The analysis of these raw data aims to extract high-quality single nucleotide polymorphism loci, which relies on bioinformatics tool chains, including sequence alignment, genotyping calling, and strict quality control filtering. Quality control standards may include minimum allele frequency, locus detection rate, and p-value of Hardy-Weinberg equilibrium test parameters. Only reliable SNP loci that pass all these filtering conditions are retained and formatted into a standard allele matrix, where the rows represent different foxtail millet individuals or lines, the columns represent different SNP loci, and each cell records the genotype of the individual at that locus.
[0030] The association strength between each remaining single nucleotide polymorphism locus and each key phenotypic feature needs to be quantified, which is a statistical quantity that measures the size of the genetic locus's ability to explain the variation of the phenotype. The calculation of association strength usually adopts statistical methods based on linear mixed models or similar structures, which take the phenotype value as the dependent variable and the genotype as the fixed effect independent variable, while introducing the individual relationship matrix as a random effect to control the false positive associations caused by population structure. For each combination of SNP loci and phenotypic features, an independent association analysis is performed, and a significance statistic is obtained, such as the regression coefficient estimate and its standard error, or the hypothesis test p-value. These statistics, after appropriate conversion (such as taking the negative logarithm of the p-value), can be integrated into a value representing the association strength, with a higher value indicating a stronger potential genetic connection between the locus and the phenotypic feature. After obtaining the association strength of all candidate SNP loci and key phenotypic features, the genetic algorithm optimization module uses this information to construct an initial population for searching the optimal solution set. The initial population consists of a certain number of individuals, each representing a possible breeding scheme, specifically a combination of screened SNP loci. The generation of the initial population is not completely random, but is guided by the association strength information, that is, the probability of SNP loci with higher association strength with phenotypic features being randomly selected and included in the initial individuals also increases accordingly. This biased initialization strategy helps to set the search starting point in the more promising areas of the solution space, thereby accelerating the subsequent convergence process.
[0031] The main body of the genetic algorithm is an iterative loop that simulates the natural evolution process, constantly updating and optimizing the population through basic operation operators such as selection, crossover, and mutation. Selection operation evaluates the pros and cons of each individual in the population based on the fitness function, and the design of the fitness function is crucial. It usually integrates the overall association strength between the SNP combination represented by the individual and the target phenotype characteristics, and its function form can be defined as:
[0032] wherein: represents the fitness value of an individual; represents the set of SNP sites contained in the individual; is the total number of key phenotype characteristics; is the weight given to the th phenotype characteristic, reflecting the relative importance of the phenotype in breeding goals; is the pre-calculated association strength value between the th SNP site and the th phenotype characteristic. The higher the fitness value , the stronger the genetic interpretation ability of the SNP combination to important phenotype characteristics, and the greater the probability of the individual being retained to the next generation in the selection process.
[0033] Crossover operation simulates gene recombination in sexual reproduction, randomly selecting two parent individuals in the population, and exchanging part of the SNP sites they contain with a certain probability, thereby generating new offspring individuals with characteristics of both parents. Mutation operation randomly changes a specific SNP site contained in an individual with a lower probability, i.e., randomly replaces it with another site not selected by the individual, which injects new genetic diversity into the population, helping to jump out of local optimal solutions and explore wider areas in the solution space. The selection, crossover, and mutation processes are repeated, and the average fitness of individuals in each generation usually shows a gradual upward trend. The iteration process continues until the maximum number of generations is reached, or the fitness of the optimal individual in the population no longer improves significantly over multiple generations. The algorithm outputs the individual with the highest fitness in the final population, and the SNP site combination it represents is considered an optimized solution for the current key phenotype characteristics, providing a clear candidate site set for subsequent calculation of the specific contribution of each genetic site. The entire process combines principles of quantitative genetics and evolutionary computation strategies to achieve the goal of intelligently screening potential functional sites from large-scale genotype data.
[0034] For reference Figure 4, the left panel shows the results of the coefficient of variation analysis for four key phenotypic indicators. The coefficient of variation is a key parameter for measuring the stability of indicators, reflecting the fluctuation of each phenotypic characteristic in the breeding population. The bar chart in the figure represents the coefficient of variation values of plant height, leaf area, flowering period, and yield indicators, respectively. The red dotted line represents the pre-set stability threshold. From the data analysis, it can be seen that different phenotypic indicators exhibit different stability characteristics: plant height and leaf area generally have relatively low coefficients of variation, indicating that these traits are relatively stable in the population; while flowering period and yield indicators may exhibit higher variability, reflecting the influence of environmental factors and genetic background on these traits. Through this stability analysis, the system can screen phenotypic indicators with smaller fluctuations, providing a reliable data basis for subsequent feature matrix construction; the right panel shows the fitness trend of genetic algorithm in optimizing SNP site combination. The figure contains two curves: the solid line represents the evolutionary trajectory of the population average fitness, and the dotted line represents the fitness change of the best individual in each generation. The fitness function integrates the overall correlation strength between SNP combination and target phenotypic characteristics, reflecting the strength of the genetic explanation ability of the gene site combination to important phenotypic characteristics. From the optimization process, it can be seen that genetic algorithm gradually improves the fitness level of the population through selection, crossover, and mutation operators. In the initial stage, the fitness value grows rapidly, indicating that the algorithm is rapidly exploring the promising areas in the solution space; as the evolution generation increases, the fitness growth gradually tends to be stable, indicating that the algorithm is converging to a better solution. This optimization process simulates the natural evolution mechanism, and can intelligently screen potential functional sites with important value for breeding decisions from large-scale genotype data.
[0035] After completing the genetic algorithm optimization and obtaining a set of candidate single nucleotide polymorphism sites with high correlation strength with key phenotypic characteristics, the implementation process enters the quantitative calculation phase of the contribution of gene sites. The calculation of contribution needs to rely on the genotype data of the breeding population and the corresponding phenotypic observation values. Taking a hypothetical Chinese milk vetch breeding population as an example, the population contains hundreds of strains, whose genotypes have been identified by sequencing technology, and the phenotypic data comes from multi-year and multi-point field tests. The calculation process begins with the statistics of the allele frequency of each SNP site. Allele frequency is a basic parameter in population genetics, reflecting the distribution of a specific allele in the population. For each candidate SNP site, the system counts the number of times of its two alleles (e.g. allele A and allele G) appearing in all breeding strains, and then divides the count by the total number of alleles (i.e. twice the number of strains), to obtain the frequency of each allele, which is between 0 and 1. The allele frequency of a site can help understand its genetic diversity, for example, when the frequency of a certain allele is close to 0.5, it indicates that the site has high polymorphism in the population.
[0036] On the basis of allele frequency, the next step is to quantify the effect size of each SNP locus on the phenotype, which is usually achieved through regression analysis. For each candidate SNP locus and each key phenotypic trait (e.g. "single pod weight"), a simple linear regression model is constructed. In this model, the phenotypic value is treated as the dependent variable, while the genotype is treated as the independent variable; for convenience of calculation, the genotype is usually coded as the count of alleles, for example, genotype AA is coded as 0, AG as 1, and GG as 2 (assuming G is the effect allele). Regression analysis will output a regression coefficient, which represents the average change in the phenotype caused by an additional copy of the effect allele; for example, a positive regression coefficient means that the presence of the effect allele G is associated with an increase in the phenotypic value. This regression coefficient, together with its standard error and significance p-value, collectively describes the quantitative relationship between the SNP locus and the phenotype.
[0037] The calculation of genetic contribution is the result of integrating allele frequency and regression coefficient. The contribution is not a single statistic, but an index that takes into account both the effect size of the locus and its commonality in the population. One common approach is to combine the absolute value of the regression coefficient (indicating the strength of the effect size) with the frequency of the effect allele (indicating the universality of the impact), and a simple multiplication is an intuitive way. More sophisticated calculations may consider the distribution characteristics of the frequency, for example, by multiplying (where p is the frequency of the effect allele) to introduce a weight proportional to the population variance, because the locus with the highest heterozygosity can theoretically explain the most genetic variance. The contribution value calculated in this way can distinguish between loci with strong effects but extremely rare and loci with moderate effects but very common, providing more comprehensive genetic information for breeding selection. See Table 1 for a virtual contribution calculation of five SNP loci: Table 1: Allele frequency, regression coefficient and genetic contribution of candidate SNP loci
[0038] Based on the calculated genetic contributions, the next step is to select the significant loci based on a pre-set significance threshold. This threshold can be determined by statistical significance (e.g., p-value less than 0.05 or a more stringent criterion after multiple testing correction) or by setting an absolute minimum value of contribution. Assuming the contribution threshold is set to 0.15, from the above table, SNP-001, SNP-002, and SNP-005 would be selected as the core explanatory variables for building the breeding decision model. With the selected set of significant genetic loci, the next step is to build the decision tree model for the breeding of the plant. The decision tree is an intuitive and easy-to-interpret machine learning model that classifies or predicts samples through a series of "yes / no" rules. The construction of the decision tree starts with the selection of the root node, which is usually the genetic locus that can best separate the breeding materials into two initial subsets based on the phenotype (e.g., "high yield" vs. "non-high yield"). For example, the SNP-001 locus with the highest contribution can be selected as the root node, and the entire breeding population is divided into two initial subsets based on its genotype (e.g., GG vs. non-GG). Then, the algorithm recursively finds the next best splitting locus within each subset to continue dividing the samples until the stopping condition is met, such as the number of samples in the node being too small or the purity being high enough. The final decision tree model has internal nodes representing the splitting rules of genetic loci, and leaf nodes corresponding to the predicted phenotype categories or specific phenotype expectations.
[0039] After the decision tree model is constructed, its parameters must be optimized to prevent overfitting, which is the phenomenon that the model performs well on training data but has poor predictive ability on new data. Cross-validation is a standard method for optimizing parameters and evaluating the generalization ability of the model. Specifically, the entire breeding population data is randomly divided into k mutually exclusive subsets of similar size (e.g., k = 5 or 10), and each subset is sequentially used as the test set while the remaining k-1 subsets are used as the training set. The decision tree model is built using the training set data, and then its prediction accuracy is evaluated on the test set. This process is repeated k times, with each subset having a chance to be the test set, and the average of the k accuracy rates is finally obtained as an estimate of the model's generalization ability. During cross-validation, different combinations of parameters such as maximum tree depth and minimum sample size per node can be tried, and the combination that performs best in terms of average accuracy and does not have an overly complex model structure is chosen as the final parameters of the decision tree model. The cross-validated and optimized decision tree model has enhanced stability and reliability, providing a practical and genotype-based decision support tool for breeders.
[0040] Example 5: After the model is built and optimized, the model outputs the optimal breeding combination scheme. The core of this process is to evaluate the potential value of different genotype combinations and make a selection. The decision tree model itself is a rule set composed of multiple judgment nodes, and each node corresponds to the genotype condition of a significant genetic locus. When a user inputs a genotype combination to be evaluated, the combination is considered a virtual new strain, and the system substitutes its genotype data into the decision tree model for "traversal". Traversal starts from the root node of the decision tree, and according to the allele composition of the genotype combination at the root node locus (e.g. SNP-001), it determines whether it should enter the left or right subtree, then moves along the branch to the next node, and again makes a judgment based on the genotype of the corresponding genetic locus of the new node. This process is recursive until the genotype combination reaches a certain leaf node of the decision tree. Each leaf node is associated with a prediction result, which can be a specific phenotype value, such as a predicted single pod weight of 25 grams, or a phenotype category, such as "high yield type". The system can quickly predict the phenotype performance of a large number of virtual or real genotype combinations, thereby evaluating their potential before implementing hybrid selection.
[0041] After predicting the performance of each genotype combination on multiple key phenotype indicators, a unified and quantifiable standard is needed to compare the overall pros and cons of different combinations, and this standard is the comprehensive breeding index. The comprehensive breeding index is a composite indicator that integrates the predicted values of multiple breeding objectives (such as high yield, early maturity, stress resistance) into a single value, allowing for intuitive ranking between different genotype combinations. The first step in constructing the index is to set the weight coefficients of each phenotype indicator, which can be determined using the analytic hierarchy process: invite multiple breeding experts to compare each trait pairwise, construct a judgment matrix, and through calculating the characteristic vector of the matrix and passing the consistency test, obtain the weight coefficients of each phenotype indicator. Calculate the weighted score of the genotype combination on each phenotype indicator. Before weighted summation, the phenotype prediction values of different dimensions need to be normalized to convert them into dimensionless indices, such as mapping to the [0, 1] interval through the extreme value method, and then multiplied by the weight coefficients. The allocation of weight coefficients reflects the breeder's emphasis on different traits in a specific breeding project, which is a decision-making process that needs to consider breeding strategies, market demand, and expert experience. For example, in a breeding program aimed at improving yield, the weights of "single pod weight" and "grain number per pod" will be assigned higher values, while the weight of "flowering date" may be relatively low; conversely, if the goal is to breed early-maturing varieties to adapt to specific cultivation systems, the weight of "flowering date" will be significantly increased. Weight coefficients are usually expressed as percentages or decimals, and the sum of the weights of all phenotype indicators under consideration should be 100% or 1 to ensure the fairness of weighted summation.
[0042] After the weight of each phenotypic index is determined, the calculation of the comprehensive breeding index of each genotype combination can begin. The calculation process is systematic, and first requires obtaining the predicted values of the combination on all target phenotypic indexes via the decision tree model. These predicted values form the basis of the evaluation data. Each phenotypic index is processed individually, and the predicted value of the index is multiplied by its pre-set weight coefficient to obtain the weighted score of the genotype combination on that particular index. For example, assume that the predicted "single pod weight" of genotype combination A is 28 grams, and the weight coefficient of this trait is set to 0.4. Then the weighted score of this genotype combination on this trait is 28 multiplied by 0.4, i.e. 11.2 points. The weighted score of this genotype combination on all other indexes such as "flowering period" and "leaf area" is calculated in the same way. The sum of the weighted scores of the genotype combination on all phenotypic indexes is the comprehensive breeding index of the genotype combination. The higher the index value, the better the overall performance of the genotype combination is expected to be after considering all important breeding goals with weights.
[0043] The system will iterate through all genotype combinations to be evaluated, and repeat the above prediction and calculation process for each combination, thereby obtaining a series of corresponding comprehensive breeding index values. The system will sort these genotype combinations according to their comprehensive breeding indexes from high to low. After sorting, the genotype combination at the top of the list, i.e. with the highest comprehensive breeding index, will be recommended by the system as the current optimal breeding scheme. The optimal scheme is usually output in the form of a report, which not only indicates the optimal genotype configuration (i.e. the specific combination of alleles of significant loci), but also lists its predicted phenotypic performance and the final calculated comprehensive breeding index value. Breeders can refer to the report to select parents carrying the target genotype for crossing or to perform early genotype screening on the offspring population in actual breeding work, thereby greatly improving breeding efficiency and accuracy.
[0044] It should be noted that the relational terms such as first and second and the like are used merely to distinguish one entity or action from another, without necessarily requiring or implying any such actual relationship or order between or among the entities or actions. Also, the terms "comprises", "comprising", or any other variations thereof are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but can include other elements not expressly listed or inherent to such process, method, article, or apparatus.
[0045] While embodiments of the application have been shown and described, it is to be understood that the embodiments described are merely exemplary of the principles and application of the present application. Numerous modifications and adaptions can be effected without departing from the spirit and scope of the present application, which is not limited to the exact construction and arrangement described. It is intended, therefore, to cover all modifications and adaptions that fall within the scope of the claims and their equivalents.
Claims
1. A Vicia based efficient breeding decision support system, characterized in that, The method comprises the following steps: Obtain multi-dimensional phenotype data in a Vicia sativa breeding test, wherein the multi-dimensional phenotype data comprises plant height, leaf area, flowering period and yield indicators; Normalize the multi-dimensional phenotype data to eliminate the influence of different dimensions, and calculate the coefficient of variation of each phenotype indicator; Screen out phenotype indicators with stability higher than a preset threshold based on the coefficient of variation, and construct a Vicia sativa phenotype feature matrix; Use an adaptive weighting algorithm to reduce the dimensionality of the Vicia sativa phenotype feature matrix, and extract key phenotype features; Input the key phenotype features into a genetic algorithm optimization module, and calculate the contribution of each genetic locus in combination with Vicia sativa genotype data; Construct a Vicia sativa breeding decision model according to the contribution of the genetic loci, and output an optimal breeding combination scheme.
2. The Vicia-based efficient breeding decision support system according to claim 1, characterized in that, The normalization of the multi-dimensional phenotype data comprises the following steps: Calculate the maximum and minimum values of each phenotype indicator, and use a linear normalization method to map each phenotype indicator to the interval of zero to one; Calculate the coefficient of variation of the normalized phenotype indicators, wherein the coefficient of variation is the ratio of the standard deviation to the mean value; Remove the phenotype indicators with a coefficient of variation exceeding a preset fluctuation range, and retain the phenotype indicators with stability meeting the requirements.
3. The Vicia-based efficient breeding decision support system according to claim 2, characterized in that, The screening of the phenotype indicators with stability higher than the preset threshold based on the coefficient of variation comprises the following steps: Set a coefficient of variation threshold, and screen out the phenotype indicators with a coefficient of variation lower than the threshold; Calculate the Pearson correlation coefficients between the screened phenotype indicators, and remove the redundant indicators with a correlation higher than a redundancy threshold; Sort the remaining phenotype indicators according to the weights, and construct a Vicia sativa phenotype feature matrix.
4. The Vicia-based efficient breeding decision support system according to claim 3, characterized in that, The reduction of the dimensionality of the Vicia sativa phenotype feature matrix by using the adaptive weighting algorithm comprises the following steps: Calculate the initial feature importance score according to the weights of the phenotype indicators; Adjust the feature importance score by using an iterative optimization method until the score converges; Select the phenotype indicators with an importance score higher than a dimensionality reduction threshold as the key phenotype features.
5. The Vicia-based efficient breeding decision support system according to claim 4, characterized in that, The input of the key phenotype features into the genetic algorithm optimization module comprises the following steps: Analyze the Vicia sativa genotype data, and extract single nucleotide polymorphism loci; Calculate the association strength of each single nucleotide polymorphism locus with the key phenotype features; Construct an initial population based on the association strength, and iteratively optimize the genetic locus combination by using a genetic algorithm.
6. The Vicia-based efficient breeding decision support system according to claim 5, characterized in that, The calculation of the contribution of each genetic locus in combination with the Vicia sativa genotype data comprises the following steps: Statistically analyze the allele frequency of each single nucleotide polymorphism locus in the breeding population; Calculate the regression coefficient of the allele frequency and the key phenotype features; Calculate the genetic contribution of each genetic locus according to the regression coefficient and the allele frequency.
7. The Vicia-based efficient breeding decision support system according to claim 6, characterized in that, The construction of the Vicia sativa breeding decision model according to the contribution of the genetic loci comprises the following steps: Screen out the genetic loci with a genetic contribution higher than a significance threshold; Construct a Vicia sativa breeding decision tree model based on the screened genetic loci; Optimize the parameters of the decision tree model by using a cross-validation method.
8. The Vicia-based efficient breeding decision support system according to claim 7, characterized in that, The output of the optimal breeding combination scheme comprises the following steps: Predict the phenotype performance of different genotype combinations according to the decision tree model; Calculate the comprehensive breeding index of each genotype combination; Select the genotype combination with the highest comprehensive breeding index as the optimal breeding scheme.
9. The Vicia-based efficient breeding decision support system according to claim 8, characterized in that, The calculation of the comprehensive breeding index of each genotype combination comprises the following steps: The weight coefficients of each phenotype index are set, the weighted scores of the genotype combination on each phenotype index are calculated, and the comprehensive breeding index is generated by summarizing the weighted scores.
10. A Vicia based efficient breeding decision support method, characterized by, All modules and method processes based on the effective breeding decision support system of the vetch in any one of claims 1 to 9 are included.
Citation Information
Patent Citations
Breeding method and system for live pig hybrid commercial line
CN118331965A
Tea tree precise breeding phenotype prediction method based on GWAS and eQTL analysis
CN120220800A
Inplanatable machine learning genome prediction method and device
CN120564828A
Whole genome based genetic evaluation and selection process
US20080163824A1
Cited By
Seed quality screening method
CN121479153A
A method of germplasm screening
CN121479153B
Crop whole-growth-period data processing method and system oriented to mulching film recovery test
CN121767126A