A method for rapid breeding of wheat
By optimizing wheat gene combinations using a data-driven approach, combined with machine learning and genetic algorithms, the problems of long breeding time, resource waste, and instability in traditional wheat breeding have been solved, enabling rapid breeding and efficient resource utilization, and improving the stability and performance of new varieties.
Patent Information
- Application Number
- CN202311436332.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-01
- Publication Date
- 2025-12-05
- Estimated Expiration
- 2043-11-01
AI Technical Summary
Traditional wheat breeding methods are time-consuming, wasteful of resources, unstable, and difficult to accurately analyze and optimize gene combinations, making it impossible to quickly promote new varieties.
Through data collection and processing, population optimization, genome selection, data analysis and experimental verification, combined with machine learning and genetic algorithms, wheat gene combinations are optimized, predictive models are established and iteratively improved until a satisfactory wheat variety is obtained.
Significantly shorten the breeding cycle, improve resource utilization efficiency, enhance the stability and performance of new varieties, and achieve data-driven breeding innovation.
Smart Images

Figure CN117373543B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of agricultural technology, and in particular to a rapid wheat breeding method. Background Technology
[0002] Traditional wheat breeding methods typically rely on experience, experimentation, and long breeding cycles, which have several significant drawbacks:
[0003] 1. Huge time consumption: Traditional breeding methods require a breeding cycle of many years, which limits the rapid promotion and application of new varieties.
[0004] 2. Waste of resources: Traditional methods require a large amount of land, human and material resources, which is costly.
[0005] 3. Instability: The performance of wheat is affected by environmental factors, and traditional methods have failed to effectively cope with different environmental conditions.
[0006] 4. Genome complexity: The wheat genome is complex, and traditional methods are difficult to accurately analyze and optimize gene combinations. Summary of the Invention
[0007] The purpose of this invention is to provide a rapid wheat breeding method.
[0008] To achieve the above objectives, the present invention is implemented according to the following technical solution:
[0009] This invention includes the following steps:
[0010] S1: Data Collection and Processing: Collect a large amount of wheat genome data, phenotypic data, and environmental data, including growth conditions, soil type, and weather. Clean, standardize, and normalize the data to ensure data quality and consistency.
[0011] S2: Population Optimization: The population selection method is used to optimize the genetic combination of wheat. Candidate gene combinations are encoded into individuals of the genetic algorithm, and a fitness function is defined to evaluate the performance of each individual.
[0012] S3: Genome selection: Based on known gene-phenotype associations, genetic algorithms are used to screen and optimize wheat gene combinations, and machine learning models are used to predict the phenotype of candidate gene combinations.
[0013] S4: Data Analysis: Use data mining and statistical analysis techniques to discover potential gene-phenotype associations and use machine learning algorithms to build predictive models;
[0014] S5: Data optimization: Use gradient descent or Bayesian optimization to adjust the model's hyperparameters to improve model performance, iterate the training of the model, and continuously improve prediction performance based on new data;
[0015] S6: Experimental Validation: The selected gene combination was planted in the experimental field, and the actual performance data was recorded; the data was compared with the model prediction to verify the effectiveness of the model.
[0016] S7: Iterative optimization: Based on experimental results and newly collected data, continuously iterate and improve the model and gene combination, repeating the above steps until a satisfactory wheat variety is achieved.
[0017] The data cleaning, standardization, and normalization in step S1 are as follows:
[0018] Data cleaning:
[0019] Outlier detection: Use Z-scores or box plots to detect outliers and mark them as data points that require further processing;
[0020] Missing value handling: For missing data, use the mean, median, or regression to fill in the missing values;
[0021] Data consistency check: Check whether the date format of the data is correct and whether the text data is consistent;
[0022] Deduplication: Find and delete duplicate data records;
[0023] Data standardization:
[0024] Z-score standardization: Transforms data into a distribution with a mean of 0 and a standard deviation of 1, as shown below:
[0025] Z=(X-μ) / σ
[0026] Where X is the original data, μ is the mean, and σ is the standard deviation;
[0027] Min-Max standardization: Scaling data to a specified interval [0,1]; as shown below:
[0028] X'=(XX min ) / (X max -X min )
[0029] Where X' is the standardized data, X min and X max These are the minimum and maximum values, respectively.
[0030] Robust standardization: using the median and interquartile range to reduce the impact of outliers;
[0031] Data normalization:
[0032] Mean normalization: Subtract the value of each data point from the mean of the feature, as shown in the following formula:
[0033] X'=X-mμ
[0034] Where X' is the normalized data and mμ is the mean of the features;
[0035] Range normalization: scales the feature values to a specified range [0,1]; similar to the Min-Max normalization algorithm, as shown in the following formula:
[0036] X'=(XX min ) / (X max -X min )
[0037] Where X' is the normalized data, X min and X max These are the minimum and maximum values of the feature, respectively.
[0038] In step S2, population optimization specifically involves: the fitness value of individual xi in the population is f(xi), the probability of selecting individual xi is p(xi), the cumulative probability is q(xi), and the number of solutions is N. The corresponding calculation formula is as follows:
[0039]
[0040]
[0041] A random array is generated, where the values of the elements range from 0 to 1. If the cumulative probability q(xi) is greater than the element in the array, then the individual xi is selected; if it is less than the element, then the next individual xi+1 is compared. This process is repeated N times until an individual is selected.
[0042] The specific data analysis in step S4 is as follows:
[0043] Establish the fitness function:
[0044] G(X)=minΣ(Y r -Y s ) 2 (3)
[0045] In the formula, G(X) is the fitness function of the optimization problem, Yr is the actual yield of wheat, and Ys is the simulated yield of wheat. The yield model of wheat grain growth stage is constructed, and the relevant formulas are as follows:
[0046]
[0047] In formula (4), Ys is the simulated wheat yield, Ps is the number of wheat ears per hectare, Ng is the number of grains per wheat ear, Wg is the dry weight of the wheat grain portion, and Cw is the moisture content of the wheat grain portion.
[0048] PS =P s_plant ×SD (5)
[0049] In equation (5), Ps_plant is the number of ears per wheat plant, and SD is the planting density;
[0050]
[0051] In equation (6), Ng_plant represents the number of grains per wheat plant;
[0052] N g_plant =R g ×W s (7)
[0053] In formula (7), Rg is the number of grains per gram of wheat stem, and Ws is the dry weight of wheat stem at the flowering stage;
[0054] W g =R grain_gfr ×h grain_gfr_Tmean ×f N_grain (8)
[0055] In Equation (8), Rgrain_gfr is the potential grain filling rate, hgrain_gfr_Tmean is the daily average temperature influence factor of grain filling rate, and fN_grain is the nitrogen influence factor of grain filling rate.
[0056]
[0057] In equation (9), hN_poten is the potential grain filling rate under nitrogen limitation, hN_min is the minimum grain filling rate under nitrogen limitation, hN_grain is the influencing factor for grain to produce nitrogen deficiency response, CN is the actual nitrogen content of stems and leaves, CN_min is the minimum nitrogen content, CN_crit is the critical nitrogen content, and fc_N is the CO2 limiting factor.
[0058] W g =min(W g W gm (10)
[0059] In formula (10), Wgm is the maximum dry weight of the grain portion of a single wheat plant; after deriving formulas (4)-(10), the formulas for the wheat yield formation model at the final grain growth stage are as follows:
[0060]
[0061] The beneficial effects of this invention are:
[0062] This invention is a rapid wheat breeding method, which achieves the following significant effects compared with existing technologies:
[0063] 1. Accelerate the breeding cycle: By utilizing population optimization and data analysis, new technologies can significantly shorten the wheat breeding cycle and achieve faster development of new varieties.
[0064] 2. Efficient use of resources: Through data optimization and model prediction, new technologies can utilize resources more effectively and reduce unnecessary waste.
[0065] 3. Enhanced stability: Combined with data analysis, new technologies can provide better environmental adaptability for wheat breeding and improve the stability of new varieties under different geographical and meteorological conditions.
[0066] 4. Precise gene combination: Through data analysis and optimization, new technologies can more accurately select and optimize wheat gene combinations, thereby improving the performance of new varieties.
[0067] 5. Data-driven innovation: New technologies combine genetics, data science, and machine learning to inject data-driven innovation into wheat breeding, improving success rates and prediction accuracy.
[0068] In conclusion, this novel rapid wheat breeding technology is expected to bring about a revolutionary change in wheat breeding, significantly shortening the breeding cycle, improving resource utilization efficiency, enhancing the stability and performance of new varieties, and promoting the sustainable development of agricultural production. Attached Figure Description
[0069] Figure 1 This is a graph showing the relationship between daily average temperature and factors affecting grain filling rate;
[0070] Figure 2 It is the relationship between the critical and minimum nitrogen concentrations in wheat leaves and the growth stage;
[0071] Figure 3 It is the relationship between the critical and minimum nitrogen concentrations in wheat stems and the growth stage;
[0072] Figure 4 This is a graph showing the fitting effect between the simulated and measured values of wheat yield. Detailed Implementation
[0073] The present invention will be further described below with reference to the accompanying drawings and specific embodiments. The illustrative embodiments and descriptions herein are used to explain the present invention, but are not intended to limit the present invention.
[0074] This invention includes the following steps:
[0075] S1: Data Collection and Processing: Collect a large amount of wheat genome data, phenotypic data, and environmental data, including growth conditions, soil type, and weather. Clean, standardize, and normalize the data to ensure data quality and consistency.
[0076] Data cleaning, standardization, and normalization are detailed below:
[0077] Data cleaning:
[0078] Outlier detection: Use Z-scores or box plots to detect outliers and mark them as data points that require further processing;
[0079] Missing value handling: For missing data, use the mean, median, or regression to fill in the missing values;
[0080] Data consistency check: Check whether the date format of the data is correct and whether the text data is consistent;
[0081] Deduplication: Find and delete duplicate data records;
[0082] Data standardization:
[0083] Z-score standardization: Transforms data into a distribution with a mean of 0 and a standard deviation of 1, as shown below:
[0084] Z=(X-μ) / σ
[0085] Where X is the original data, μ is the mean, and σ is the standard deviation;
[0086] Min-Max standardization: Scaling data to a specified interval [0,1]; as shown below:
[0087] X'=(XX min ) / (X max -X min )
[0088] Where X' is the standardized data, X min and X max These are the minimum and maximum values, respectively.
[0089] Robust standardization: using the median and interquartile range to reduce the impact of outliers;
[0090] Data normalization:
[0091] Mean normalization: Subtract the value of each data point from the mean of the feature, as shown in the following formula:
[0092] X'=X-mμ
[0093] Where X' is the normalized data and mμ is the mean of the features;
[0094] Range normalization: scales the feature values to a specified range [0,1]; similar to the Min-Max normalization algorithm, as shown in the following formula:
[0095] X'=(XX min ) / (X max -X min )
[0096] Where X' is the normalized data, X min and X max These are the minimum and maximum values of the feature, respectively.
[0097] S2: Population Optimization: A population selection method is used to optimize the genetic combinations of wheat. Candidate gene combinations are encoded into individuals in a genetic algorithm, and a fitness function is defined to evaluate the performance of each individual. Specifically, the fitness value of individual xi in the population is f(xi), the probability of selecting individual xi is p(xi), the cumulative probability is q(xi), and the number of solutions is N. The corresponding calculation formula is as follows:
[0098]
[0099]
[0100] A random array is generated, where the values of the elements range from 0 to 1. If the cumulative probability q(xi) is greater than the element in the array, then the individual xi is selected; if it is less than the element, then the next individual xi+1 is compared. This process is repeated N times until an individual is selected.
[0101] S3: Genome selection: Based on known gene-phenotype associations, genetic algorithms are used to screen and optimize wheat gene combinations, and machine learning models are used to predict the phenotype of candidate gene combinations.
[0102] S4: Data Analysis: Use data mining and statistical analysis techniques to discover potential gene-phenotype associations, and utilize machine learning algorithms to build predictive models; specifically as follows:
[0103] Establish the fitness function:
[0104] G(X)=minΣ(Y r -Y s ) 2 (3)
[0105] In the formula, G(X) is the fitness function of the optimization problem, Yr is the actual yield of wheat, and Ys is the simulated yield of wheat. The yield model of wheat grain growth stage is constructed, and the relevant formulas are as follows:
[0106]
[0107] In formula (4), Ys is the simulated wheat yield (kg·hm-2), Ps is the number of wheat spikes per hectare (spike / hm2), Ng is the number of grains per wheat spike (grain / spike), Wg is the dry weight of the wheat grain (g), and Cw is the moisture content of the wheat grain (the moisture content of wheat grain at maturity is 20%).
[0108] P S =P s_plant ×SD (5)
[0109] In equation (5), Ps_plant is the number of spikes per wheat plant (spike / plant), and SD is the planting density (plant / hm2).
[0110]
[0111] In equation (6), Ng_plant is the number of grains per wheat plant (grain / plant).
[0112] N g_plant =R g ×W s (7)
[0113] In formula (7), Rg is the number of grains per gram of wheat stem (default value is 25.0), and Ws is the dry weight (g) of wheat stem at the flowering stage, which is measured by the drying method.
[0114] W g =R grain_gfr ×h grain_gfr_Tmean ×f N_grain (8)
[0115] In equation (8), Rgrain_gfr is the potential grain filling rate (the default value is 0.00100 grain / d for the flowering to grain filling stage and 0.00200 grain / d for the grain filling stage), and hgrain_gfr_Tmean is the daily average temperature influence factor of the grain filling rate, with a value between 0 and 1. The value variation curve is shown in the figure. Figure 1 fN_grain is the nitrogen factor affecting grain filling rate; such as Figure 1 As shown.
[0116]
[0117] In equation (9), hN_poten is the potential grain filling rate under nitrogen limitation (default value is 5.50 × 10⁻⁵ ggrain / d), hN_min is the minimum grain filling rate under nitrogen limitation (default value is 1.50 × 10⁻⁵ ggrain / d), hN_grain is the influencing factor for nitrogen deficiency response in grains (default value is 1.00), CN is the actual nitrogen content (%) of stems and leaves, which was measured using the semi-micro Kjeldahl method in the experiment. CN_min is the minimum nitrogen content (%), CN_crit is the critical nitrogen content (%), and the variation curves of CN_crit and CN_min are shown in the figure. Figure 2 fc_N is the CO2 limiting factor (the default value for the stem is 1.00, while for the leaf it is related to the CO2 concentration). Since the effects of nitrogen and CO2 on wheat growth were not considered in this experiment, hN_grain = 1.00 and fc_N = 1.00.
[0118] W g =min(W g W gm (10)
[0119] In equation (10), Wgm is the maximum dry weight of the grain portion of a single wheat plant (default value is 0.04g). After deriving equations (4)-(10), the formulas for the final grain growth stage wheat yield formation model are as follows:
[0120]
[0121] S5: Data optimization: Use gradient descent or Bayesian optimization to adjust the model's hyperparameters to improve model performance, iterate the training of the model, and continuously improve prediction performance based on new data;
[0122] S6: Experimental Validation: The selected gene combination was planted in the experimental field, and the actual performance data was recorded; the data was compared with the model prediction to verify the effectiveness of the model.
[0123] S7: Iterative optimization: Based on experimental results and newly collected data, continuously iterate and improve the model and gene combination, repeating the above steps until a satisfactory wheat variety is achieved.
[0124] Figures 1-3 Simulate the change curves for relevant variables. Figure 1 This is a graph showing the relationship between daily average temperature and factors affecting grain filling rate, specifically the variation curves of hgrain_gfr_Tmean and Tmean. Figure 2 It relates the critical and minimum nitrogen concentrations in wheat leaves to the growth stage. Figure 3This relates the critical and minimum nitrogen concentrations in wheat stems to different growth stages. This invention focuses on stages 6–9, i.e., the flowering to maturity stage.
[0125] The optimized model's RMSE decreased from 363.22 kg·hm⁻² to 57.85 kg·hm⁻², and the NRMSE decreased from 21.78% to 3.47%. The fitting effect between the simulated and measured wheat yield values is as follows: Figure 4 As shown in the figure, after parameter optimization, the simulated values fit the measured values better and are closer to the 1:1 line, showing a very strong consistency between the simulated and measured values.
[0126] The technical solutions of the present invention are not limited to the specific embodiments described above. Any technical modifications made in accordance with the technical solutions of the present invention fall within the protection scope of the present invention.
Claims
1. A method for rapid breeding of wheat, characterized in that, The method comprises the following steps: S1: Data collection and processing: Collect a large amount of wheat genomic data, phenotype data and environmental data, including growth conditions, soil types, weather, data cleaning, standardization and normalization to ensure data quality and consistency; S2: Population optimization: Use population selection method to optimize the genetic combination of wheat, encode the candidate genetic combination into individuals of genetic algorithm, and define fitness function to evaluate the performance of each individual; S3: Genomic selection: Based on the information of known gene-phenotype association, screen and optimize the genetic combination of wheat through genetic algorithm, and use machine learning model to predict the phenotype of candidate genetic combination; S4: Data analysis: Use data mining and statistical analysis techniques to discover potential gene-phenotype associations, and use machine learning algorithms to build prediction models; S5: Data optimization: Adjust the hyperparameters of the model using gradient descent or Bayesian optimization to improve the performance of the model, iteratively train the model, and continuously improve the prediction performance based on new data; S6: Experimental verification: Plant the preferred genetic combination in experimental fields and record the actual performance data; compare with the model prediction to verify the effectiveness of the model; S7: Iterative optimization: Based on the experimental results and newly collected data, continuously iterate and improve the model and genetic combination, repeat the above steps until a satisfactory wheat variety is obtained.
2. A method of rapid breeding of wheat as claimed in claim 1, wherein: The data cleaning, standardization and normalization in step S1 are as follows: Data cleaning: Outlier detection: Use Z-score or box plot to detect outliers and mark them as data points that need further processing; Missing value processing: For missing data, use mean, median or regression to fill in missing values; Data consistency check: Check if the data date format is correct and the text data is consistent; Duplicate data deletion: Find and delete duplicate data records; Data standardization: Z-score standardization: Convert data to a distribution with mean 0 and standard deviation 1, as follows: Z = (X - μ) / σ Where X is the original data, μ is the mean, and σ is the standard deviation; Min-Max standardization: Scale data to a specified interval [0, 1]; as follows: X1' = (X - X 1min ) / (X 1max - X 1min ) where X1' is the normalized data, X 1min and X 1max are the minimum and maximum values, respectively; Robust standardization: Use median and interquartile range to reduce the influence of outliers; Data normalization: Mean normalization: Subtract the mean of the feature from each data point, as follows: X2' = X - mμ Where X2' is the normalized data, and mμ is the mean of the feature; Range normalization: Scale the values of the feature to a specified range [0, 1]; same as Min-Max standardization algorithm, as follows: X3' = (X - X 3min ) / (X 3max - X 3min ) where X3' is the normalized data, X 3min and X 3max are the minimum and maximum values of the feature, respectively.
3. A method of rapid wheat breeding according to claim 2, wherein: The population optimization in step S2 is as follows: The fitness value of individual xi in the population is f(xi), the probability of selecting individual xi is p(xi), the cumulative probability is q(xi), and the number of solutions is N. The corresponding calculation formula is as follows: Randomly generate an array, where the value of each element ranges from 0 to 1. If the cumulative probability q(xi) is greater than the element in the array, individual xi is selected. If it is less than, compare the next individual xi+1, repeat N times, until one individual is selected.
4. A method of rapid wheat breeding according to claim 3, wherein: The data analysis in step S4 is as follows: The fitness function is established as follows: G(X) = min∑(Y r - Y s ) 2 (3) In the formula, G(X) is the fitness function of the optimization problem, Yr is the actual yield of wheat, and Ys is the simulated yield of wheat. The yield model of the wheat grain growth stage is constructed, and the related formula is as follows: In formula (4), Ys is the simulated yield of wheat, Ps is the number of wheat spikes per hectare, Ng is the number of grains per spike, Wg is the dry weight of the wheat grain part, and Cw is the water content of the wheat grain part; P s = P s_plant x SD (5) In formula (5), Ps_plant is the number of spikes of a single wheat plant, and SD is the planting density; In formula (6), Ng_plant is the number of grains of a single wheat plant; N g_plant = R g x W s (7) In formula (7), Rg is the number of grains per gram of the wheat stem part, and Ws is the dry weight of the wheat stem part at the flowering stage; W g = R grain_gfr x h grain_gfr_Tmean x f N_grain (8) In formula (8), Rgrain_gfr is the potential grain filling rate, hgrain_gfr_Tmean is the daily average temperature influence factor of the grain filling rate, and fN_grain is the nitrogen influence factor of the grain filling rate; In formula (9), hN_poten is the potential grain filling rate under nitrogen limitation, hN_min is the minimum grain filling rate under nitrogen limitation, hN_grain is the influence factor of the nitrogen deficiency response of the grain, CN is the actual nitrogen content of the stem and leaf, CN_min is the minimum nitrogen content, CN_crit is the critical nitrogen content, and fc_N is the CO2 limitation factor; W g = min(W g ,W g ) (10) In formula (10), Wgm is the maximum dry weight value of the grain part of a single wheat plant; after deducing formulas (4)-(10), the formula of the final yield formation model of the wheat grain growth stage is as follows:
Citation Information
Patent Citations
Construction method of winter wheat flowering period simulation model based on multi-site genes
CN114496075A
Digitized breeding method for oilseed rape
CN116230097A