Near infrared spectroscopy for plant compositional analysis using predictive models

By merging datasets from a subset of plant material with a learned prediction model, the method generates accurate NIRS calibration models for plant composition without the need for wet chemistry analysis, addressing the time-consuming nature of existing methods.

WO2026030314A1PCT designated stage Publication Date: 2026-02-05PIONEER HI BREED INTERNATIONAL INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/039655
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-07-30
Filing Date
2025-07-29
Publication Date
2026-02-05

AI Technical Summary

Technical Problem

The development of calibration models for Near Infrared Spectroscopy (NIRS) in agriculture is time-consuming due to the need for wet chemistry analysis of every sample in large databases, which can contain hundreds or thousands of entries.

Method used

A method for generating calibration models using a combination of a calibration database, a first dataset from a subset of plant material with measured constituent traits, and a learned prediction model to predict constituent traits in a second subset, merging these datasets to create a final model without requiring wet chemistry analysis for each sample.

Benefits of technology

The method produces accurate calibration models that can measure constituent traits in plant material with an accuracy comparable to models developed solely by wet chemistry, reducing the reliance on time-consuming wet chemistry analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IMGF000008_0001
    Figure IMGF000008_0001
  • Figure IMGF000009_0001
    Figure IMGF000009_0001
  • Figure IMGF000014_0001
    Figure IMGF000014_0001
Patent Text Reader

Abstract

The present disclosure provides methods and systems for generating Near Infrared Spectroscopy (NIRS) calibration models to accurately measure the amount of traits of interest in plant samples. The methods and systems reduce the requirement for wet chemistry analysis when compared to traditional approaches.
Need to check novelty before this filing date? Find Prior Art

Description

NEAR INFRARED SPECTROSCOPY FOR PLANT COMPOSITIONAL ANALYSISUSING PREDICTIVE MODELSBACKGROUND

[0001] Near infrared spectroscopy (NIRS) is widely used in agriculture to predict constituents in crops. The most time-consuming step for NIRS is the development of calibration models which compare sample values in a database with hundreds or even thousands of entries by wet chemistry (WC) analysis for a given constituent. Accordingly, there is a need to develop new NIRS methods to reduce the reliance on wet chemistry to generate the calibration models.SUMMARY

[0002] Provided are methods for generating a calibration model for measuring plant composition for a constituent trait of interest comprising providing a calibration database, the calibration database comprising spectra from a population of plant material, measuring the amount of a constituent trait of interest in members of a first subset of the population of plant material to generate a first dataset comprising a spectra and a corresponding amount of the constituent trait of interest for the members of the first subset of the population, inputting spectra of members of a second subset of the population into a learned prediction model to generate a prediction dataset comprising a spectra and a corresponding predicted amount of the constituent trait of interest for the members of the second subset of the population, the learned prediction model trained using the first dataset to learn the relationship of the spectra and corresponding amount of the constituent trait of interest, and merging the first dataset and prediction dataset to produce a final calibration model for the constituent trait of interest. In certain embodiments, the method further comprises validating the final calibration model using a validation dataset, the validation dataset comprising a set of samples with the constituent trait of interest measured by wet chemistry. In certain embodiments, the method further comprises measuring the constituent trait of interest in a test sample using NIRS with the final calibration model.

[0003] Also provided herein are methods for generating a calibration model for measuring plant composition for a constituent trait of interest comprising providing a calibration database, the calibration database comprising spectra from a population of plant material, inputting spectra from a plurality of members of the calibration database into a learned prediction model togenerate a prediction dataset comprising a spectra and a corresponding predicted amount of the constituent trait of interest for the plurality of members of the calibration database, the learned prediction model trained to leam the relationship of the spectra and amount of the constituent trait of interest, and producing a final calibration model for the constituent trait of interest from the prediction dataset. In certain embodiments, the method further comprises validating the final calibration model using a validation dataset, the validation dataset comprising a set of samples with the constituent trait of interest measured by wet chemistry. In certain embodiments, the method further comprises measuring the constituent trait of interest in a test sample using NIRS with the final calibration model.DETAILED DESCRIPTION

[0004] Near Infrared Spectroscopy (NIRS) is widely used in agriculture and other fields to accurately measure the amounts or concentrations of a constituent of interest in samples such as crops. The most time-consuming step for NIRS is the development of calibration models (prediction models) to convert the transmitted and / or reflected light spectra from NIR into an accurate measurement of the amount or concentration of the constituent of interest. Current methods to develop calibration models require an analysis of every sample in a given database, which may contain hundreds or even thousands of entries, by wet chemistry. The present disclosure provides methods for generating a calibration model to accurately measure the amount or concentration of a constituent trait of interest that does not require performing a wet chemistry analysis for each member of a population database.

[0005] NIRS measurements are based on the absorption of light energy in the near-infrared spectrum range (about 780 to 2500 nm) by C — C, C — H, O — H, N — H, S — H and C=O bonds in the organic constituents of the materials being analyzed. The absorption of the light energy is proportional to the concentration of the constituent of interest and the modified light comprising one or more of transmitted and reflected light spectra from the sample (e.g., plant material) can be converted to accurately measure the amounts or concentrations of the constituent trait of interest. “Modified light” as used in the context of this disclosure means light that is transmitted (transmitted light) and / or reflected (reflected light) from a plant material after receiving light from a light source. Transfected light is a combination of reflected and transmitted light and is included in modified light.

[0006] Provided herein are methods for generating a calibration model for measuring plant composition for a constituent trait of interest comprising providing a calibration database, the calibration database comprising spectra from a population of plant material, measuring the amount of a constituent trait of interest in members of a first subset of the population of plant material to generate a first dataset, the first dataset comprising a spectra and a corresponding amount of the constituent trait of interest for the members of the first subset of the population, inputting spectra of members of a second subset of the population into a learned prediction model to generate a prediction dataset, the prediction dataset comprising a spectra and a corresponding predicted amount of the constituent trait of interest for the members of the second subset of the population, and the learned prediction model trained using the first dataset to learn the relationship of the spectra and corresponding amount of the constituent trait of interest, and merging the first dataset and prediction dataset to produce a final calibration model for the constituent trait of interest.

[0007] Also provided are methods for generating a calibration model for measuring plant composition for a constituent trait of interest comprising providing a calibration database, the calibration database comprising spectra from a population of plant material, inputting spectra from a plurality of members of the calibration database into a learned prediction model to generate a prediction dataset comprising a spectra and a corresponding predicted amount of the constituent trait of interest for the plurality of members of the calibration database, the learned prediction model trained to learn the relationship of the spectra and amount of the constituent trait of interest, and producing a final calibration model for the constituent trait of interest from the prediction dataset.

[0008] As used herein, “constituent trait of interest” “constituent of interest” or the like refers to a biochemical or nutritional component for which determining the concentration or amount of is desired. The constituent trait of interest of the methods described herein is not particularly limited as long as the biochemical or nutritional component can be detected by NIR. In certain embodiments of the methods described herein, the constituent trait of interest may be any macroconstituent that is found in plant material. The selection of the constituent trait of interest is reflective of the interests of those involved in determining constituent trait concentrations of plant material. Examples of constituent traits of interest of the methods described herein include, but are not limited to, protein, oil, glucosinolates, water, fatty acids, such as, for example,saturated fatty acids and unsaturated fatty acids including, but not limited to caproic acid, caprylic acid, capric acid, lauric acid, myristic acid, palmitic acid, stearic acid, arachidic acid, behenic acid, lignoceric acid, cerotic acid, oleic acid, linoleic acid, linolenic acid, alpha-linolenic acid, erucic acid and arachidonic acid, or any combination thereof, starch, acid detergent fiber, neutral detergent fiber, carbohydrates, sucrose, sucrosyl-oligosaccharides, such as, for example, stachyose or raffinose, phytate, total dietary fiber (TDF), protein subfractions, such as, for example, cruciferin, napin, glycinin, and conglycinin, forage traits, such as, for example, starch, soluble sugars, lignin, soluble sugars, whole plant or cell wall digestibilities.

[0009] A calibration database in the methods described herein refers to a database comprising spectra from a population of plant material from which the calibration model is developed. The members of the population of the calibration database will vary based on the plant for which measurement of the constituent trait is desired. For example, for measuring oil content in canola plant material the calibration database comprises a population of spectra from canola plant material. The number of members of the population is not particularly limited as long as the population has compositional diversity. In certain embodiments, the population comprises at least 100, 250, 500, 750, 1000, 1250, 1500, 1750, 2000, 2250, 2500, 2750 or 3000 members. In certain embodiments, the population of plant material comprises a plurality of genetic backgrounds, environmental growing conditions, geographic locations, constituent trait ranges, growing seasons, or any combination thereof.

[0010] As used herein, “compositional diversity” refers to variations in the total constituent trait of interest in the population, so that the population has a range of levels for the constituent trait of interest. In certain embodiments of the methods described herein, the population for calibration comprises field grown plants, such as from a diverse genetic background, comprising a compositional diversity, such as, for example, a compositional diversity created using mutation, directed genome modification (e.g., CAS CRISPR editing) and other transgenic techniques. In certain embodiments, the composition diversity created by growing the plants in different geographical locations, under different growth conditions, and / or in different growing seasons.

[0011] In certain embodiments of the methods described herein, a first subset of the population of plant material is analyzed using reference chemistry to determine the amount or concentration of the constituent trait of interest. The number of members of the first subset of the population isnot particularly limited such that the first subset contains at least one less member than the total number of members of the calibration database. In certain embodiments, the first subset comprises less than 70%, 60%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 15%, 10%, 5%, 4%, 3%, 2%, or 1% of the total number of members of the calibration database.

[0012] In certain embodiments of the methods described herein, a second subset of the population of plant material is used to generate the prediction dataset. The number of members of the second subset of the population is not particularly limited. In certain embodiments, the second subset comprises at least 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, or 95% of the total number of members of the calibration database. In certain embodiments of the methods described herein, the members of the second subset of the population are not present in the first subset of the population. In certain embodiments of the methods described herein, less than 50%, 45%, 40%, 35%, 30%, 25%, 20%, 15%, 10%, 5%, 4%, 3%, 2%, or 1% of the members of the second subset of the population are present in the first subset of the population.

[0013] As used herein, a “calibration model” refers to a model developed based on the relationship of NIR measurements and component concentration from a population of samples having diverse compositional ranges. The calibration model can be used to measure component concentration from NIR spectra of a test sample.

[0014] The following provides an example of the method for generating a calibration model for measuring plant composition for a constituent trait of interest. From an initial database of 2000 entries, a representative subset of 100 samples were selected. For the 100 samples, wet chemistry analyses for as many constituents (e.g., constituent A) as needed was performed. A predictive model with these 100 samples for constituent A = Ainit model was built. The predictive model Ainit model was used to predict all of the remaining 1900 samples of the initial database. Optionally, random noise is added to the 1900 predictions (e.g., 1-2 times the error of the model SECV). The two datasets - 100 samples with Awctchem and 1900 samples with Apredictions - are merged to produce a combined database Asum with 2000 entries having 100 samples with Awetchem and 1900 samples with Apredictions. The database ASUm is used to develop a final calibration model. The final calibration model can be validated with a validation database.

[0015] In another example of the method for generating a calibration model for measuring plant composition for a constituent trait of interest a prior model for constituent A is used A = Ainit priormodel. The Ainit prior model is used to predict all samples of the initial spectral database of 2000 samples. Optionally, random noise is added to the 2000 predictions. The 2000 predictions are then used to develop a final calibration model. The final calibration model can be validated with a validation database.

[0016] The plant material for use in the methods described herein is not particularly limited and may be any part of a plant in which measuring a constituent of interest is desired. Examples of plant material for use in the methods described herein include, but is not limited to, seeds, plant forage (including, but not limited to, leaves, flowers, branches, fruit, kernels, ears, cobs, husks, and stalks), roots, shoots, root tips, anthers, ovules, pollen, and grain. Plant material, as used herein, also includes defatted meals and powders produced by grinding plant parts, such as for example, seed powders. Grain is intended to mean the mature seed produced by commercial growers for purposes other than growing or reproducing the species. The plant material may be isolated or taken from any plant species, including, but not limited to, monocots and dicots.

[0017] Examples of plant species include, but are not limited to, maize, Brassica sp. (e.g., B. napus, B. rapa, B.juncea) (also referred to herein as Canola), particularly those Brassica species useful as sources of seed oil, mustard, alfalfa, rice, rye, sorghum, millet (e.g., pearl millet, proso millet, foxtail millet, finger millet), sunflower, safflower, wheat, soybean, tobacco, potato, peanuts, cotton, sweet potato, cassava, coffee, coconut, pineapple, citrus trees, cocoa, tea, banana, avocado, fig, guava, mango, olive, papaya, cashew, macadamia, almond, sugar beets, sugarcane, oats, barley, vegetables, ornamentals, conifers, turf grasses (including cool seasonal grasses and warm seasonal grasses) tomatoes, lettuce, green beans, lima beans, peas, members of the genus Cucumis such as cucumber, cantaloupe, and musk melon (C. meld), azalea, hydrangea, hibiscus, roses, tulips, daffodils, petunias, carnation, poinsettia, guar, locust bean, fenugreek, garden beans, cowpea, mungbean, lima bean, fava bean, lentils, chickpea, chrysanthemum, pines such as loblolly pine, slash pine, ponderosa pine, lodgepole pine, and monterey pine, douglas-fir, western hemlock, sitka spruce, redwood, true firs such as silver fir and balsam fir, and cedars such as Western red cedar and Alaska yellow-cedar, poplar and eucalyptus. In certain embodiments, the plant material is isolated from crop plants such as, for example, maize, soybean, wheat, cotton, sunflower, rice, mustard, sorghum, corn, alfalfa, and sunflower.

[0018] As used herein, "reference chemistry" or “wet chemistry” refers to the values obtained for the measurements of the compositions analyzed herein, using standard wet chemistry referenceanalytical methods or their equivalent. When developing Near Infrared models for constituent analysis harmonized industry standard reference chemistries are often used as the basis for model development. Accordingly, in certain embodiments of the methods described herein, the "standard wet chemistry reference analytical method" used for measuring a constituent trait of interest is a technique performed using industry standard methodologies developed by International and National Accreditation Organizations such as, for example: International Organization for Standardization (ISO); Association of Official Analytical Chemists (AOAC); American Association of Cereal Chemists (AACC); American Oil Chemists Society (AOCS). Non limiting examples of reference chemistry methods are provided in Table 1.Table 1 : Standard Wet Chemistry Reference Analytical Methods

[0019] As used herein a “learned prediction model” refers to a model that has been trained using data (e.g., historical data) to predict likely outcomes. A learned prediction model is usually not a fixed model and is validated or revised regularly to incorporate changes in the underlying data. The type of learned prediction model for use in the methods described herein is not particularly limited and may be any type of learned prediction model that can be trained to learn a relationship of NIR spectra and the composition (e.g., concentration) of the constituent trait of interest. The learned prediction models can be used to generate a calibration model that can accurately predict the concentration of the constituent trait of interest in a test sample using the NIR spectra of the test sample. Similarly, the method of training the learned prediction model is not particularly limited and may be any method known in the art, or described herein, to train a model to learn a relationship of NIR spectra and the composition (e.g., concentration) of the constituent trait of interest.

[0020] In certain embodiments, the learned prediction model comprises a machine learning model. Any suitable machine learning model may be used in the methods and systems described herein. Types of models include without limitation statistical models, such as probability models, regression models, such as partial least squares (PLS) regression, and those involving deep learning, such as supervised, self-supervised, and unsupervised models, or combinations thereof. In certain embodiments, the machine learning model is a classification model, a regression model, a clustering model, a dimensionality reduction model, a distribution model, for example, a multivariate or univariate Gaussian distribution model, or a deep learning model. In certainembodiments, the deep learning model is part of an ensemble model. In certain embodiments, the deep learning model is an ensemble model comprising two or more models. In certain embodiments, the deep learning model is a supervised learning model. The supervised learning model may be a classification or regression model. The machine learning models include support vector machines, neural networks, such as SVM-DA (Support Vector machines) or ANN (Artificial Neural Networks), or deep learning algorithms and the like.

[0021] In certain embodiments, the machine learning model for use in the methods described herein is an artificial neural network (ANN). ANNs are configured to synthesize or learn from a plurality of inputs to produce an output. One or more variables in the algorithms can have weights that are applied to each equation and optimized as the neural network is trained. Based on the amount of training information, the deep learning models or networks get better at producing more helpful outputs. In certain embodiments, the ANN includes a plurality of input factors that may be used to train predicted constituent trait information.

[0022] In certain embodiments, the learned prediction model comprises a linear statistical model. As used herein, a “linear statistical model” is a model that provides a linear relationship between an independent variable (e.g., NIR spectra) and a dependent variable (e.g., concentration of the constituent trait) to predict the concentration in a sample. The type of linear statistical model is not particularly limited and may be any linear statistical model known in the art. In certain embodiments, the linear statistical model comprises a logistic generalized linear prediction model.

[0023] In certain embodiments of the methods described herein, the method further comprises adding random noise to one or more of the predictions in the prediction dataset. Real-world data is inherently noisy due to measurement variation, environmental factors, or the like, such that adding random noise helps to simulate real-world data. As used herein, “random noise” refers to adjusting a prediction in the prediction dataset. In certain embodiments, random noise is added to all of the members of the population. In certain embodiments of the methods described herein, random noise is added to at least 0.5%, 1%, 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, or 75% of the members of the prediction dataset and fewer than 99%, 95%, 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, or 50% of the members of the population. The amount of noise to add to the individual sample is not particularly limited. In certain embodiments of the methods described herein, 0.5-, 1-, 1.5-, 2-, or 2.5-fold of theprediction error is added to one or more members of the prediction dataset. In certain embodiments, the amount of error added corresponds to the magnitude of error that is typical of the reference chemistry methods being used, such that the amount of error added is based on prior calculation knowledge. In certain embodiments, the learned prediction model comprises a linear regression model and the random noise introduces noise in the predictions but does not alter the regression line.

[0024] In certain embodiments of the methods described herein, after the final calibration model for the constituent trait of interest is produced, the method further comprises validating the final calibration model using a validation dataset. In certain embodiments, the validation dataset comprises a set of samples with the constituent trait of interest measured by wet chemistry.

[0025] As used herein, a “validation dataset” refers to a set of samples that are independent from the calibration database (i .e., not part of the population of the calibration database) that have the constituent trait of interest measured by wet chemistry. The number of samples of the validation dataset is not particularly limited. In certain embodiments, the validation dataset comprises at least 25, 50, 75, 100, 150, 200, 250, 300, 350, 400, 450, or 500 samples. In certain embodiments, the samples of the validation dataset are selected to be representative for the constituent of interest. For example, in certain embodiments, the samples of the validation dataset are representative of the range of the constituent in the final calibration model, the genetics of the plant material used to produce the final calibration model, the environments of the plant material used to produce the final calibration model, the growing locations of the plant material used to produce the final calibration model, and / or the growing seasons of the plant material used to produce the final calibration model.

[0026] In certain embodiments of the methods described herein, the final calibration model is validated when the standard error of prediction (SEP) for the constituent trait of interest using NIRS with the final calibration model is within 0.01 %, 0.1 %, 0.5%, 1 %, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11 %, 12%, 13%, 14%, 15% or 20% of the SEP for the constituent trait of interest when measured using NIRS with a control model (e.g., calibration model in which every sample in a given database is measured by wet chemistry).

[0027] In certain embodiments of the methods described herein, the method further comprises measuring the constituent trait of interest in a test sample using NIRS with the final calibration model. In certain embodiments, the amount of the constituent trait of interest is measured usingNIRS with the final calibration model to an accuracy that is within 0.2 wt. %, 0.3 wt. %, 0.4 wt. %, 0.5 wt. %, 0.6 wt. %, 0.7 wt. %, 0.8 wt. %, 0.9 wt. %, 1.0 wt. %, 1.1 wt. %, 1.2 wt. %, 1.3 wt. %, 1.4 wt. %, or 1.5 wt. % of the amount measured using NIRS with a control model (e.g., calibration model in which every sample in a given database is measured by wet chemistry). In certain embodiments, the amount of the constituent trait of interest is measured using NIRS with the final calibration model to an accuracy that is within 0.2 wt. %, 0.3 wt. %, 0.4 wt. %, 0.5 wt. %, 0.6 wt. %, 0.7 wt. %, 0.8 wt. %, 0.9 wt. %, 1.0 wt. %, 1.1 wt. %, 1.2 wt. %, 1.3 wt. %, 1.4 wt. %, or 1.5 wt. % of the amount measured using a standard reference analytical method.

[0028] The moisture content affects the weight percentages of components of the seed, with drier seeds generally having a higher weight percent of the component, such as oil or protein. When comparing NIR-based measurements with standard reference analytical methods, measurements may be taken in each case at the same moisture content, or if measurements are taken at different moisture contents, the values obtained can be corrected to the same moisture content.Measurements can, for example, be taken at or standardized to a moisture content by weight of at least or at least about 0.01 %, 0.1 %, 0.5%, 1 %, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11 %, 12%, 13%, 14%, 15% or 20% and less than or less than about 35%, 30%, 25%, 20%, 19%, 18%, 17%, 16%, 15%, 14%, 13%, 12%, 1 1 %, 10%, 9% or 8%. It is common practice to report constituent concentrations on a Dry Weight Basis. In such cases the dry weight of the sample is that achieved after oven drying using, for example, a method as described in Table 1.

[0029] The following are examples of specific embodiments of some aspects of the invention. The examples are offered for illustrative purposes only and are not intended to limit the scope of the invention in any way.EXAMPLE 1

[0030] This example demonstrates the accuracy of measuring oil, protein, acid detergent fiber (ADF), glucosinolates, oleic acid, and linolenic acid using calibration models developed with prediction datasets.

[0031] The prediction accuracy of calibration models developed using prediction datasets (“New NIRS”) was compared to the prediction accuracy of calibration models developed with a wet chemistry analysis for each member of the calibration dataset (“Normal NIRS”). Constituent traits of interest for these tests ranged from bulk constituents (oil, protein) to minor constituents(oleic acid, linolenic acid, and acid detergent fiber (ADF)) to very difficult to predict constituents (glucosinolates) and for this example canola whole grain was used for the plant material. In this example, a prior calibration model was used for predicting the samples of the calibration database. The prior model used for this example was a proprietary model developed in-house, but similar results are expected using a model obtained from a third party.

[0032] For measuring oil, protein, ADF, glucosinolates, oleic acid, and linolenic acid in canola whole grain using New NIRS a calibration model was developed for each constituent. The calibration database used for producing the calibration model for each constituent contained 1248 spectral entries. The calibration models were built using only NIRS predictions in which partial least squares (PLS) regression was used to predict the constituents from the spectra of the samples in the calibration database. The following meta parameteres using the WinISI software package were applied: wavelengths - 1100nm-2492nm or 1100-1800nm (depending on constituent) at 8nm increments, cross validation groups of 4, scatter correction by multiplicative scatter correction (MSC); and math treatment - second derivative, segment and gap sizes at 5 each.

[0033] For measuring oil, protein, ADF, glucosinolates, oleic acid, and linolenic acid in canola whole grain using Normal NIRS calibration models for each constituent trait were built from a calibration database in which each sample was tested by wet chemistry. The wet chemistry methods used were oil (VDLUFA Methodenbuch III 5.1.3: Bestimmung des Rohfettgehaltes in Olsaaten), glucosinolates (ISO 9167-1; A. Quinsac et al. (1991), J. Assoc. Off. Anal. Chem. 74, 932-939), linolenic acid (AOCS Official Method Ce le-91: Determination of Fatty Acids in Edible Oils and Fats by Capillary GLC), and 200 samples were selected for wet chemistry analysis of ADF (D.7.d. Fibretherm - ADF in Futtermitteln 20-6_2013.doc), protein (AOCS Official Method Ba 4e-93: Generic Combustion Method for Determination of Crude Protein), and oleic acid (AOCS Official Method Ce le-91 : Determination of Fatty Acids in Edible Oils and Fats by Capillary GLC).

[0034] To compare the prediction accuracy of New NIRS and Normal NIRS the calibration models were validated using a validation database containing 321 samples (Oil and Protein) or 305 samples (ADF, Glucosinolates, Oleic acid, Linolenic acid) all with wet chemistry data. As shown in Table 2, the standard error of prediction (SEP) and coefficient of determination (RSQ) were nearly identical for New NIRS and Normal NIRS for each of the traits tested. These resultsindicate that calibration models using prediction datasets are as accurate as models built solely by wet chemistry.Table 2: Comparison of Prediction Accuracy of New NIRS to Normal NIRSWC stands for wet chemistryEXAMPLE 2

[0035] This example demonstrates the accuracy of measuring oil, protein, acid detergent fiber (ADF), glucosinolates, oleic acid, and linolenic acid using calibration models developed with prediction datasets.

[0036] The prediction accuracy of calibration models developed using prediction datasets (“New NIRS”) was compared to the prediction accuracy of calibration models developed with a wet chemistry analysis for each member of the calibration dataset (“Normal NIRS”). Constituenttraits of interest for these tests ranged from bulk constituents (oil, protein) to minor constituents (oleic acid, linolenic acid, and acid detergent fiber (ADF)) to very difficult to predict constituents (glucosinolates) and for this example canola whole grain was used for the plant material. In this example, a no prior calibration model was used for predicting the samples of the calibration database.

[0037] For measuring oil, protein, ADF, glucosinolates, oleic acid, and linolenic acid in canola whole grain using New NIRS, a calibration model was developed for each constituent. The calibration database used for producing the calibration model for each constituent contained 1248 spectra entries. From the calibration database 100 samples were selected for wet chemistry analysis of oil (VDLUFA Methodenbuch III 5.1.3 : Bestimmung des Rohfettgehaltes in Olsaaten), glucosinolates (ISO 9167-1; A. Quinsac et al. (1991), J. Assoc. Off. Anal. Chem. 74, 932-939), linolenic acid (AOCS Official Method Ce le-91 : Determination of Fatty Acids in Edible Oils and Fats by Capillary GLC), and 200 samples were selected for wet chemistry analysis of ADF (D.7.d. Fibretherm - ADF in Futtermitteln 20-6_2013.doc), protein (AOCS Official Method Ba 4e-93: Generic Combustion Method for Determination of Crude Protein), and oleic acid (AOCS Official Method Ce le-91 : Determination of Fatty Acids in Edible Oils and Fats by Capillary GLC) to build a first dataset containing spectra and a corresponding wet chemistry value for each of the samples. A prediction model dataset for each of the constituents was built using the information from the first dataset in which partial least squares (PLS) regression was used to predict the constituents from the spectra of the samples in the calibration database that do not have a corresponding wet chemistry analysis. The following meta parameteres using the WinISI software package were applied: wavelengths - 1100nm-2492nm or 1100-1800nm (depending on constituent) at 8nm increments, cross validation groups of 4, scatter correction by multiplicative scatter correction (MSC); and math treatment - second derivative, segment and gap sizes at 5 each. The values obtained for the samples in the predictive model dataset was merged with the first dataset to provide the final calibration model for each constituent.

[0038] For measuring oil, protein, ADF, glucosinolates, oleic acid, and linolenic acid in canola whole grain using Normal NIRS calibration models for each constituent trait were built from a calibration database in which each sample was tested by wet chemistry.

[0039] To compare the prediction accuracy of New NIRS and Normal NIRS the calibration models were validated using a validation database containing 321 samples (Oil and Protein) or 305 samples (ADF, Glucosinolates, Oleic acid, Linolenic acid) all with wet chemistry data. As shown in Table 3, the standard error of prediction (SEP) and coefficient of determination (RSQ) were nearly identical for New NIRS and Normal NIRS for each of the traits tested. These results indicate that calibration models using prediction datasets are as accurate as calibration models built solely by wet chemistry.Table 3: Comparison of Prediction Accuracy of New NIRS to Normal NIRSWC stands for wet chemistry

[0040] All publications and patent applications in this specification are indicative of the levelof ordinary skill in the art to which this invention pertains. All publications and patent applications are herein incorporated by reference to the same extent as if each individual publication or patent application was specifically and individually indicated by reference.

[0041] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs. Unless mentioned otherwise, the techniques employed or contemplated herein are standard methodologies well known to one of ordinary skill in the art. The materials, methods and examples are illustrative only and not limiting.

[0042] Many modifications and other embodiments of the inventions set forth herein will come to mind to one skilled in the art to which these inventions pertain having the benefit of the teachings presented in the foregoing descriptions and the associated drawings. Therefore, it is to be understood that the inventions are not to be limited to the specific embodiments disclosed and that modifications and other embodiments are intended to be included within the scope of the appended claims. Although specific terms are employed herein, they are used in a generic and descriptive sense only and not for purposes of limitation.

[0043] Units, prefixes and symbols may be denoted in their SI accepted form. Numeric ranges are inclusive of the numbers defining the range.

Claims

We claim:

1. A method for generating a calibration model for measuring plant composition for a constituent trait of interest, the method comprising: a. providing a calibration database, the calibration database comprising spectra from a population of plant material, the population comprising a first subset of members and a second subset of members, the second subset of members comprising at least some members distinct from the first subset of members; b. measuring the amount of the constituent trait of interest using wet chemistry in the first subset of members to generate a first dataset comprising a first plurality of spectra and a plurality of corresponding amounts of the constituent trait of interest for the first subset of the population; c. inputting spectra of the second subset of members into a learned prediction model to generate a prediction dataset comprising a second plurality of spectra and a plurality of corresponding predicted amounts of the constituent trait of interest for the second subset of members, the learned prediction model trained using the first dataset to learn the relationship of the spectra and corresponding amount of the constituent trait of interest; and d. merging the first dataset and prediction dataset to produce a final calibration model for the constituent trait of interest.

2. The method of claim 1, wherein the method further comprises adding random noise to predictions in the prediction dataset.

3. The method of claim 1 or 2, wherein after (d), the method further comprises validating the final calibration model using a validation dataset, the validation dataset comprising a set of samples with the constituent trait of interest measured by wet chemistry.

4. The method of any one of claims 1-3, wherein the final calibration model is validated when the standard error of prediction (SEP) for the constituent trait of interest using NIRS with the final calibration model that is within 20% of the SEP for the constituent trait of interest when measured using a control NIRS model.

5. The method of any one of claims 1 -4, wherein the method further comprises measuring the constituent trait of interest in a test sample using NIRS with the final calibration model.

6. The method of any one of claims 1-5, wherein the constituent trait of interest is selected from the group consisting of oil, protein, glucosinolates, fatty acids, acid detergent fiber, and neutral detergent fiber.

7. The method of any one of claims 1-6, wherein the plant material is a plant seed, plant forage, defatted meals, or seed powder.

8. The method of claim 7, wherein the plant material is from a plant selected from the group consisting of canola, maize, soybean, wheat, cotton, sunflower, rice, mustard and sorghum.

9. The method of any one of claims 1-8, wherein the population of plant material comprises a plurality of genetic backgrounds, environmental growing conditions, geographic locations, constituent trait ranges, growing seasons, or any combination thereof.

10. The method of any one of claims 1-9, wherein the population of plant material comprises at least 500 samples.

11. A method for generating a calibration model for measuring plant composition for a constituent trait of interest, the method comprising: a. providing a calibration database, the calibration database comprising spectra from a population of plant material; b. inputting spectra from a plurality of members of the calibration database into a learned prediction model to generate a prediction dataset comprising the inputted spectra and a plurality of corresponding predicted amounts of the constituent trait of interest for the plurality of members of the calibration database, the learned prediction model trained to learn the relationship of the spectra and amount of the constituent trait of interest; and c. producing a final calibration model for the constituent trait of interest from the prediction dataset.

12. The method of claim 11, wherein the method further comprises adding random noise to predictions in the prediction dataset.