Genomic Prediction Accuracy via Optimized Estimation Data Sets

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current genomic prediction methods in plant and animal breeding face challenges in achieving high accuracy due to limitations in the relatedness between training individuals and selection candidates, as well as inefficiencies in data utilization.

Innovation Solution

The method involves constructing an optimized estimation data set by selecting candidates for phenotyping based on their genomic estimated breeding value accuracy, which is higher than that of other candidates, and iteratively adding them to the data set until optimized, thereby improving the accuracy of genomic prediction.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional phenotypic selection or marker-assisted selection is used, then the breeding process is simpler, but the accuracy of selection is lower

Engineering Contradiction:
Improveaccuracy of selectionVSAvoidcomplexity of breeding method
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent performs preliminary genotyping of parents and construction of genomic relationship matrices before phenotypic data is available. This allows the prediction model to be prepared in advance, and phenotypes are only needed to update the prediction, significantly reducing the time and complexity of the breeding process while maintaining high accuracy

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces genomic relationship matrices as an intermediary between traditional phenotypic selection and modern genomic selection. This intermediary structure allows the integration of genotypic information from parents with phenotypic data, enabling accurate predictions without requiring extensive phenotypic data collection on all candidates

Inventive Principle:
Principle #24Intermediary (Mediator)

2Quantity of substance

If all available phenotypes are used in the training data set, then more data is available for estimation, but the accuracy for certain families decreases

Engineering Contradiction:
Improveamount of training dataVSAvoidaccuracy for certain families
Core Design Contradiction:
Quantity of substanceVSMeasurement precision

Solution Approach 1:

The patent applies local quality by allowing different subsets of phenotypic data to be used for different prediction targets. The system can selectively include or exclude certain families' phenotypes based on the specific prediction needs, ensuring that the training data is optimally matched to the prediction target rather than using a uniform approach for all cases

Inventive Principle:
Principle #3Local quality

3Ease of manufacture

If the training data set does not match the prediction target, then data collection is easier, but the accuracy of genomic prediction decreases

Engineering Contradiction:
Improveease of data collectionVSAvoidaccuracy of genomic prediction
Core Design Contradiction:
Ease of manufactureVSMeasurement precision

Solution Approach 1:

The patent performs preliminary genotyping of parents and construction of population-specific genomic relationship matrices before phenotypic data is available. This allows the prediction model to be tailored to specific prediction targets in advance, ensuring that the training data will be appropriately matched when collected, thereby maintaining high accuracy without complicating the data collection process

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS12272429B2Molecular breeding methods
Publication Date: 2025.04.08 PIONEER HI BREED INTERNATIONAL INC
  • US12272429B2 patent drawing
  • US12272429B2 patent drawing
  • US12272429B2 patent drawing

AI summary

Methods to improve the selection of breeding individuals as part of a breeding program are provided in which optimized estimation data sets are constructed by selecting candidates for phenotyping, for which genotypic information is also available, from a candidate set and inputting them into the estimation data set and then evaluating accuracy of genomic estimated breeding values for each candidate (i.e. genomic prediction accuracy). The optimized estimation data set is then used as a model to determine genomic estimated breeding values of breeding individuals based purely on genotypic information.