System and method of capturing and analyzing mid-infrared spectral images of pulses

FT-MIR spectroscopy with chemometric models addresses the inefficiencies of traditional phenotyping by enabling rapid, cost-effective, and non-destructive analysis of nutritional traits in dry seeds, facilitating high-throughput breeding programs.

WO2026020166A1PCT designated stage Publication Date: 2026-01-22CLEMSON UNIV RES FOUND
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/038437
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-07-19
Filing Date
2025-07-21
Publication Date
2026-01-22

AI Technical Summary

Technical Problem

Traditional phenotyping methods for nutritional analysis of dry seeds from crops like dry pea, lentil, and chickpea are time-consuming, require significant sample preparation, and destroy multiple seeds, making them unsuitable for early-stage breeding programs where germplasm is scarce.

Method used

Utilizing Fourier Transform Mid-Infrared (FT-MIR) spectroscopy combined with chemometric models to analyze mid-infrared spectral images of plant portions, allowing for rapid, high-throughput phenotyping without chemical pretreatments, and enabling the determination of nutritional traits such as protein quality, fatty acid content, and digestibility.

Benefits of technology

This approach provides instant results with minimal sample destruction, reduces time and cost, and is not confounded by environmental effects, making it suitable for large-scale phenotyping and breeding programs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025038437_22012026_PF_FP_ABST
    Figure US2025038437_22012026_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed herein is a system and method are provided for rapid, high-throughput phenotyping of nutritional traits in plants using mid-infrared (MIR) spectroscopy and chemometric modeling. One or more processors control MIR sensors to capture spectral images of plant material, which are analyzed using a chemometric model, such as partial least squares regression, to determine characteristics including protein, fatty acids, dietary fiber, starch, resistant starch, sulfur-containing amino acids, and digestibility. The method supports non-destructive analysis with minimal sample preparation and is applicable to a wide range of plant species and biological materials. Results can be stored in a local or cloud database and visualized through a graphical user interface. The chemometric models may be trained and validated using K-fold cross validation and can include spectral preprocessing and interpolation filters. This enables rapid, cost-effective nutritional analysis to support plant breeding, food quality assessment, and industrial applications.
Need to check novelty before this filing date? Find Prior Art

Description

SYSTEM AND METHOD OF CAPTURING AND ANALYZING MID-INFRARED SPECTRAL IMAGESOF PULSESGovernment Support

[0001] This invention was made with government support under Grant No. 2019002132, awarded by the United States Agency for International Development (USAID), and Grant No. 2021-02927, awarded by the National Institute of Food and Agriculture’s Organic Agriculture Research and Extension Initiative (NIFA-OREI). The government has certain rights in the invention.Cross-Reference to Related Application(s)

[0002] This application claims the benefit under 35 U.S.C. § 119(e) to U.S. Provisional Application 63 / 673,462, filed July 19, 2024, and entitled “SYSTEM AND METHOD OF CAPTURING AND ANALYZING MID-INFRARED SPECTRAL IMAGES OF PULSES”, which is hereby incorporated herein by reference in its entirety.Field

[0003] The various examples herein relate to infrared spectral image analysisBackg round

[0004] Dry pea (Pisum sativum L. ), lentil (Lens culinaris Medik. ), and chickpea (Cicer arietinum L) are cold season legumes. The dry seeds derived from these are known as pulses. These crops are nutritionally rich in both macronutrients and micronutrients. The macronutrients in the dry seed are carbohydrates (~ 50 - 58 %), proteins (-15 - 22 %), and fats (<10 %). These are quantitative traits and nutritional phenotyping is necessary to understand their distribution and allow for identification of genetic markers for marker-assisted breeding. Rapid and high-throughput phenotyping is highly advantageous in nutritional breeding to reduce the time and money investment of large-scale nutritional analysis. Traditional phenotyping includes chromatographic techniques. Most importantly, reverse phase high performance liquid chromatography (RP- HPLC) with diode array detection (DAD) is used for amino acid and protein analysis. Ion exchange chromatography with pulse amperometry detection (IEC- PAD) is applied with carbohydrate analysis, and fatty acids analysis is performed with gas chromatography with mass spectrometry detection (GC-MS). These techniques require much time and sample preparation prior to analysis. Furthermore, many seeds (~ 10 - 50) are destroyed for analysis, which can be a major breeding constraint at the early stages of trials (before advancing) when germplasm is scarce. Accordingly, traditional phenotyping is no ideal for nutritional phenotyping in breeding programs.Brief Summary

[0005] Discussed herein are various examples for methods of applying chemometric models to midinfrared spectral images of portions of a plant in order to determine various characteristics of at least the portion of the plant, such as digestibility, protein quality / quantity, fatty acid content, and more Acomputing device controls one or more mid-infrared sensors to capture one or more mid-infrared spectral images of at least a portion of a plant. For the purposes of this disclosure, the term mid-infrared spectral images may include any output generated by a mid-infrared sensor or a mid-infrared spectrophotometer, including spectra, spectrum, or a spectral fingerprint. The computing device receives the one or more mid-infrared spectral images and analyzes the one or more mid-infrared spectral images using a chemometric model. The computing device can determine, based at least in part on the analyzing, one or more characteristics of the portion of the plant.

[0006] Fourier transform mid-infrared (FT-MIR) was developed as a high throughput phenotyping tool for nutritional traits in pulse crops. With this approach, the flour of a single seed was sufficient for quantitative analysis without chemical pretreatments. Therefore, no hazardous chemicals were required or generated as waste during the analytical process This technique also saves time and cost per sample with instant results for all the nutritional traits from a single sample scan Commonly, Fourier transform near infrared spectroscopy (FT-NIR) is used in chemometric modeling for high throughput phenotyping. Since, FT-NIR signals (bands) are less selective and always require multivariate analysis (i.e., principal component analysis followed by partial least squares regression (PLSR)) FT-MIR is favored over FT-NIR. The true meaning of the data is hidden in FT-NIR due to the excessive sample penetration by NIR energy causing highly overlapped vibrational overtones. Additionally, FT-MIR solely depends on intense bands of fundamental modes of molecular oscillations which originate from the traits of interest Accordingly, PLSR modeling with MIR data are not confounded by environmental effects (i.e., locations, rainfall etc ) and timely recalibrations aren’t required. Beyond pulses, this technique is feasible with various crops (i.e. cereals) for nutritional analysis and further in pulse and cereal based food industries. Therefore, FT-MIR data is a promising tool for high throughput nutritional phenotyping and food analysis.

[0007] In Example 1 , a method comprises: (i) controlling, by one or more processors, one or more midinfrared sensors to capture one or more mid-infrared spectral images of at least a portion of a plant; (ii) receiving, by the one or more processors, the one or more mid-infrared spectral images; (Hi) analyzing, by the one or more processors, the one or more mid-infrared spectral images using a chemometric model; and (iv) determining, by the one or more processors and based at least in part on the analyzing, one or more characteristics of the portion of the plant.

[0008] Example 2 relates to the method of Example 1 , wherein the portion of the plant comprises a pulse.

[0009] Example 3 relates to the method of Example 2, wherein the pulse remains intact after the one or more mid-infrared spectral images are captured.

[0010] Example 4 relates to the method of any one or more of Examples 1-3, wherein the plant is in a Fabaceae family.

[0011] Example 5 relates to the method of any one or more of Examples 1-4, wherein the plant comprises one of: a dry pea, a lentil, a chickpea, a casein, a sunflower, an oat, a split pea, a kidney bean, wheat, a black eyed pea, any food (e.g., solid or dry feed materials), or any biological material.

[0012] Example 6 relates to the method of any one or more of Examples 1-5, wherein the one or more characteristics comprise any one or more of: a total fatty acid, a total saturated fatty acid, a totalunsaturated fatty acid, a total dietary fiber, a total protein, a total resistant starch, a total starch, a total sulfur containing amino acids, and a digestibility factor

[0013] Example 7 relates to the method of any one or more of Examples 1-6, wherein the chemometric model comprises a partial least squares regression model.

[0014] Example 8 relates to the method of any one or more of Examples 1-7, wherein the one or more characteristics comprise a digestibility factor, and wherein analyzing the one or more mid-infrared spectral images comprises: (i) applying, by the one or more processors, the chemometric model to the one or more mid-infrared spectral images to detect one or more patterns of alpha sheet signals within an amide I band of the one or more mid-infrared spectral images; and (ii) estimating, by the one or more processors, the digestibility factor based on the one or more patterns

[0015] Example 9 relates to the method of any one or more of Examples 1-8, further comprising: (i) storing, by the one or more processors, the one or more characteristics of the portion of the plant in a database; (ii) retrieving, by the one or more processors, historical characteristics of plants of a same species as the plant from the database; (iii) generating, by the one or more processors, a graphical user interface comprising graphical indications of at least one of the one or more characteristics and at least one of the historical characteristics; and (iv) outputting, by the one or more processors and to an output device, the graphical user interface

[0016] Example 10 relates to the method of Example 9, wherein the database comprises a cloud database stored on one or more remote servers.

[0017] Example 11 relates to the method of any one or more of Examples 1-10, wherein the chemometric model includes one or more interpolation filters.

[0018] Example 12 relates to the method of any one of Examples 1-11 , wherein the chemometric model is trained and / or validated using K-fold cross validation.

[0019] Example 13 relates to the method of Example 12, wherein the K-fold cross validation comprises at least five folds.

[0020] Example 14 relates to the method of any one of Examples 12-13, wherein the mid-infrared spectral images are normalized between 0 and 1 prior to analysis.

[0021] Example 15 relates to the method of any one of Examples 12-14, further comprising interpolating the spectral data to achieve a spectral resolution of 1.86 cm-1.

[0022] Example 16 relates to the method of any one of Examples 12-16, wherein the mid-infrared spectral images are captured at a resolution of 2 cm-1 and with a zero-fill factor of 2.

[0023] In Example 17, a computing device comprises one or more processors configured to: (i) control one or more mid-infrared sensors to capture one or more mid-infrared spectral images of at least a portion of a plant; (ii) receive the one or more mid-infrared spectral images; (iii) analyze the one or more mid-infrared spectral images using a chemometric model; and (iv) determine, based at least in part on the analyzing, one or more characteristics of the portion of the plant.

[0024] Example 18 relates to the method of Example 17, wherein the one or more characteristics comprise a digestibility factor, and wherein the one or more processors being configured to analyze the one or more mid-infrared spectral images comprises the one or more processors being configured to: (i) apply the chemometric model to the one or more mid-infrared spectral images to detect one or morepatterns of alpha sheet signals within an amide I band of the one or more mid-infrared spectral images; and (ii) estimate the digestibility factor based on the one or more patterns.

[0025] Example 19 relates to the method of any one of Examples 17-18, wherein the one or more processors are further configured to: (i) store the one or more characteristics of the portion of the plant in a database; (II) retrieve historical characteristics of plants of a same species as the plant from the database; (Hi) generate a graphical user interface comprising graphical indications of at least one of the one or more characteristics and at least one of the historical characteristics; and (iv) output, to an output device, the graphical user interface

[0026] In Example 20, a non-transitory computer-readable storage medium has stored thereon instructions that, when executed, cause one or more processors of a computing device to: (i) control one or more mid-infrared sensors to capture one or more mid-infrared spectral images of at least a portion of a plant; (ii) receive the one or more mid-infrared spectral images; (Hi) analyze the one or more mid-infrared spectral images using a chemometric model; and (iv) determine, based at least in part on the analyzing, one or more characteristics of the portion of the plant.

[0027] While multiple examples are disclosed, still other examples will become apparent to those skilled in the art from the following detailed description, which shows and describes illustrative examples. As will be realized, the various implementations are capable of modifications in various obvious aspects, all without departing from the spirit and scope thereof. Accordingly, the drawings and detailed description are to be regarded as illustrative in nature and not restrictive.Brief Description of the Drawings

[0028] The following drawings are illustrative of particular examples of the present disclosure and therefore do not limit the scope of the invention. The drawings are not necessarily to scale, though examples can include the scale illustrated, and are intended for use in conjunction with the explanations in the following detailed description wherein like reference characters denote like elements. Examples of the present disclosure will hereinafter be described in conjunction with the appended drawings.

[0029] FIG 1 is a block diagram illustrating an example system for capturing mid-infrared spectral images of portions of a plant and analyzing those spectral images, in accordance with one or more techniques of this disclosure.

[0030] FIG 2 is a block diagram illustrating a more detailed example of a computing device configured to perform the techniques described herein

[0031] FIG 3 is a flow diagram illustrating an example method for capturing mid-infrared spectral images of portions of a plant and analyzing those spectral images.

[0032] FIG 4 is a flow diagram illustrating a schematic representation of FT-MIR spectroscopy as applied with high-throughput phenotyping, in accordance with one or more of the techniques of this disclosure.Detailed Description

[0033] The following detailed description is exemplary in nature and is not intended to limit the scope, applicability, or configuration of the techniques or systems described herein in any way. Rather, thefollowing description provides some practical illustrations for implementing examples of the techniques or systems described herein Those skilled in the art will recognize that many of the noted examples have a variety of suitable alternatives

[0034] FIG 1 is a block diagram illustrating an example system 100 for capturing mid-infrared spectral images of portions of a plant and analyzing those spectral images, in accordance with one or more techniques of this disclosure In the example of FIG. 1 , system 100 includes pulse 102, mid-infrared spectrophotometer 104, and computing device 110.

[0035] In the example of system 100, pulse 102 may be any seed derived from a crop where the seed may be analyzed using mid-infrared spectroscopy. Examples of plants where pulse 102 may be derived include a dry pea, a lentil, a chickpea, a casein, a sunflower, an oat, a split pea, a kidney bean wheat, a black eyed pea, any food (e.g., solid or dry feed materials), or any biological material, although pulse 102 may be derived from any other plant with characteristics similar to those listed above or any other plant rich in carbohydrates, proteins, fats, and micronutrients where the dry seed may be analyzed accurately using mid-infrared spectroscopy.

[0036] Mid-infrared spectrophotometer 104 is shown in FIG. 1 as a standalone device that is configured to capture mid-infrared spectral images of pulse 102. In practice, mid-infrared spectrophotometer 104 may be any device that can be controlled to emit mid-infrared energy in order to capture mid-infrared spectral images of pulse 102 and transmit those spectral images to computing device 110, either through a wireless connection, a wired connection, or through a removable hard drive, a removable solid state drive, or other form of removable memory Mid-infrared spectrophotometer 104 may be a standalone device or may be incorporated into computing device 110. For the purposes of this disclosure, the term mid-infrared spectral images may include any output generated by a mid-infrared sensor or a mid-infrared spectrophotometer, including spectra, spectrum, or a spectral fingerprint.

[0037] Computing device 110 may be any computer with the processing power required to adequately execute the techniques described herein. For instance, computing device 110 may be any one or more of a mobile computing device (e.g., a smartphone, a tablet computer, a laptop computer, etc.), a desktop computer, a smarthome component (e.g., a computerized appliance, a home security system, a control panel for home components, a lighting system, a smart power outlet, etc.), an integrated computer system (e.g., integrated into mid-infrared spectrophotometer 104), a vehicle, a wearable computing device (e.g., a smart watch, computerized glasses, a heart monitor, a glucose monitor, smart headphones, etc.), a virtual reality / augmented reality / extended reality (VR / AR / XR) system, a video game or streaming system, a network modem, router, or server system, or any other computerized device that may be configured to perform the techniques described herein.

[0038] In accordance with the techniques of this disclosure, computing device 110 may control midinfrared spectrophotometer 104 to capture one or more mid-infrared spectral images of at least a portion of a plant, such as pulse 102 Computing device 110 may receive the one or more mid-infrared spectral images from mid-infrared spectrophotometer 104 and may analyze the one or more mid-infrared spectral images using a chemometric model. Computing device 110 may determine, based at least in part on the analyzing, one or more characteristics of pulse 102.

[0039] Fourier transform spectroscopic techniques are popular with rapid qualitative and quantitative analysis. In this approach, the plausibility of Fourier transform mid infrared (FT-MIR) spectroscopy for rapid and high throughput phenotyping was studied. The fundamental modes of oscillations associated with the functional groups in the pulse seed flours in the mid IR region (4000 -650 cm-1) were observed and confirmed via chemical tests. For instance, peak shifts of proton - deuterium exchange was observed in the amide I and A bands upon spiking D2O to confirm amide I, II and A bands associated with proteins and protein quality in dry pea, lentil, and chickpea. Carbohydrates were confirmed via analyzing defatted samples followed by protein digestion for above crops and analyzing the remaining peaks with comparisons to the standards (i.e. , corn starch). Fatty acids were spiked into the flour (i.e., chickpea) to confirm the occurrence of total fat associated signals.

[0040] With partial least squares regression (PLSR) chemical data from traditional phenotyping techniques were correlated with spectroscopic data at different regions associated with separate trait of interest. Accordingly, separate chemometric models were developed for proteins, protein quality, total starch, and resistant starch targeting dry pea, lentil and chickpea. FT-MIR spectroscopy was also successful in chemometric modeling for total fatty acids (TFA), total saturated fatty acids (TSFA), a total dietary fiber (TDF), and total saturated fatty acids (TSFA) in chickpea. The challenge of moisture effects (in mid IR region) in model accuracy was mitigated via spectral preprocessing (i.e., derivatization and Stavisky - Golay smoothing etc ) prior to model development for certain traits. All the above models had multivariate regression coefficients (R2) ranging from 0.80 - 0.95 indicating the model accuracy. With leave one out cross validations these models constituted root mean square errors of prediction leading to deviations in the results by 5- 10% from actual values.

[0041] Fourier transform mid Infrared (FT-MIR) spectroscopy has been developed to phenotype nutritional traits that required advancement in pulse breeding. Those nutritional traits included macronutrients such as total protein (TP), total starch (TS), total resistant starch (TRS), total fatty acids (TFA), total dietary fiber (TDF) total saturated fatty acids (TSFA), and total unsaturated fatty acids (TUSFA). Total sulfur containing amino acids (TSAA) and protein digestibility (PDg) were additional traits. TSAA and PDg represented the improvement of protein quality in pulses. Chemometric analysis of TP, TS, TRS, TFA, TDF, TSFA, TUSFA and PDg were developed for chickpea, dry pea, and lentil flours Partial least squares regression (PLSR) was the principle chemometric approach for all the nutritional traits. The tendency to maximize the covariance between spectroscopic and nutritional data through recognizing associated trends of variation influenced the application of PLSR throughout this study.

[0042] Unlike in Fourier transform near infrared (FT-NIR) spectroscopy, FT-MIR data does not require data screening through non-supervised statistical means such as principal component analysis (PCA). Regions of spectral selection were based on the vibrations of associated functional groups of the traits of interest followed by chemical validation through spiking experiments. Therefore, the techniques described herein include selected spectral regions and the combinations of spectral pre-treatments applied with chemometric model development. The presence of amide bands associated with TP and PDg models were confirmed through spiking deuterated water (D2O) followed by monitoring peak shifts associated with proton and deuteron exchange. The peak shifts of amide II (~ 1550 cm1) and A (~ 3270crrr1) bands towards shorter wavenumbers and steady amide I (~ 1650 cm'1) band confirmed the functional groups for developing PLSR models for TP from pulse flour. The steadiness of amide I band was due to beta sheets in pulse proteins which did not significantly exchange with deuterons due to conformational constrains. This phenomenon confirmed its use over PLSR modeling for PDg as beta sheets strictly associated with protein quality. For TS and TRS PLSR models’ functional groups were confirmed via analyzing defatted samples followed by protein digestion and analyzing the remaining peaks with comparisons to the standards (i.e. , corn starch) The ester carbonyl (1763.03 - 1720.17 cm'1), alkyl and vinyl hydrogen bands (3035.92 - 2845.82 cm'1) used to construct TFA, TSFA, and TUSFA PLSR models were confirmed through spiking saturated and unsaturated fatty acid methyl ester standards. This approach was unique with constructing chemometric models for estimating total fatty acids from chickpea flour. The PLSR modeling for TSAA followed a similar approach. Therefore, the identity of amino acids was confirmed by spiking cysteine and methionine standards. Accordingly, the chemical behavior of molecular species associated with nutritional traits were carefully accounted in spectral selection throughout this invention.

[0043] The spectral pre-treatments applied with above PLSR models included normalization, first and second order derivatives, and Savizsky - Golay (SG) smoothing. Normalization encountered error minimization from loose contact between sample and the diamond surface of the attenuated total reflectance (ATR) module. Derivatives followed by SG smoothing provided baseline correction and deconvolution effects for certain spectral bands. For instance, the PLSR model for TP required the first derivative and SG smoothing to mitigate the interference from moisture band at amide I region. Additionally, TFA, TSFA and TUSFA PLSR models required second order derivatization followed by SG smoothing. The ester carbonyl band convoluted with amide I and II bands and, alkyl, and vinyl hydrogen bands were appropriately resolved for chemometric application with the above approach. The second derivatives also decoupled the matrix interferences associated with chickpea total fat analysis due to moisture effects. However, derivatives were not successful with certain sensitive regions. PLSR models for PDg was constructed with normalized spectra without further preprocessing. This was in response to the distortions in derivatized amide I band that prevented Non-linear iterative partial least squares (NIPALS) algorithm of PLSR from recognizing key patterns of beta sheets correlated to phenotypic values.

[0044] The models and algorithms used to execute the techniques described herein may be builty using re-construction with sci-kit learn or tensor flow, such as under a Python 3 platform, although other programming languages and / or methodologies may be utilized. Accordingly, an online cloud system may be implemented into this system, targeting multi-user access with parallel computing. This will provide remote access to interested users to generate phenotypic data of several nutritional traits of interest in high throughput. Consequently, interpolation filters may be added into the PLSR models for rapid computations of phenotypic data for a hassle free and promising workflow

[0045] Fourier transform mid-infrared (FT-MIR) was developed as a high throughput phenotyping tool for nutritional traits in pulse crops. With this approach, the flour of a single seed was sufficient for quantitative analysis without chemical pretreatments. Therefore, no hazardous chemicals were required or generated as waste during the analytical process. This technique also saves time and cost persample with instant results for all the nutritional traits from a single sample scan. Commonly, Fourier transform near infrared spectroscopy (FT-NIR) is used in chemometric modeling for high throughput phenotyping. Since, FT-NIR signals (bands) are less selective and always require multivariate analysis (i.e., principal component analysis followed by partial least squares regression (PLSR)) FT-MIR is favored over FT-NIR. The true meaning of the data is hidden in FT-NIR due to the excessive sample penetration by NIR energy causing highly overlapped vibrational overtones. Additionally, FT-MIR solely depends on intense bands of fundamental modes of molecular oscillations which originate from the traits of interest Accordingly, PLSR modeling with MIR data are not confounded by environmental effects (i.e., locations, rainfall etc ) and timely recalibrations aren’t required. Beyond pulses, this technique is feasible with various crops (i.e. cereals) for nutritional analysis and further in pulse and cereal based food industries.

[0046] A number of specific examples, embodiments, and implementations are described throughout this disclosure to illustrate the principles and applications of the subject matter. These examples are provided for purposes of explanation and are not intended to be exhaustive or to limit the scope of the subject matter in any way. The specific plant species, seed types, sample sizes, numerical values, and parameter ranges described herein are illustrative and should not be construed as limiting Other plant species, biological materials, sample quantities, and operational parameters may be used in accordance with the disclosed methods and systems without departing from the intended scope. Modifications and variations to the described examples will be apparent to those skilled in the art and are considered within the scope of this disclosure.

[0047] In one example, this disclosure provides a system and method for rapid, high-throughput phenotyping of nutritional quality traits in plant materials, with a particular focus on the measurement of in vitro protein digestibility in pulse crops using Fourier-transform mid-infrared (FT-MIR) spectroscopy and chemometric modeling. The method is applicable to a variety of pulse crops, including chickpea (Cicer arietinum L), dry pea (Pisum sativum L), and lentil (Lens culinaris Medik. ), and enables rapid measurement of total protein and sulfur-containing amino acid (SAA) concentrations in plant material. Data quantified using these techniques could further include total fatty acids (TFA), total unsaturated fatty acids (TUSFA), and total saturated fatty acids (TSFA) in chickpea flour using Fourier-transform mid-infrared (FT-MIR) spectroscopy combined with chemometric modeling. The techniques may also quantify total starch (TS) and resistant starch (RS) in pulse crop flours using FT-MIR spectroscopy combined with chemometric modeling. The method is designed to accelerate phenotyping in plant breeding and food analysis by minimizing sample preparation, labor, and chemical use, as well as reducing the time, cost, and sample destruction associated with conventional enzymatic assays.

[0048] The system includes one or more mid-infrared sensors, such as an FT-MIR spectrometer equipped with an attenuated total reflectance (ATR) accessory, which are operably connected to a computing device. The computing device comprises one or more processors configured to control the MIR sensor, receive spectral data, perform chemometric analysis, and determine one or more characteristics of the plant material. The spectral range for data acquisition is typically 650-4000 cm1, and instrument parameters such as resolution and scan number are optimized for each trait and crop. For example, a resolution of 2 or 4 cm’1with 36 to 200 scans per sample are used, depending on thespecific analysis The ATR surface is cleaned between samples to ensure data quality. The ATR surface may be cleaned with HPLC-grade methanol before each measurement to ensure data quality.

[0049] In one example, plant samples, such as dry pea, lentil, or chickpea seeds, are ground into flour with a maximum particle size of approximately 0.5 mm. The flour is stored under controlled temperature and humidity conditions prior to analysis. The method is minimally destructive, requiring only a small amount of sample, such as the flour from a single seed.

[0050] The FT-MIR spectrometer is configured to collect spectral data in the mid-infrared range, typically from 650 to 4000 cm'1, with a focus on the amide I band (1756.81-1586.27 cm'1), which is sensitive to protein secondary structure. Spectral acquisition is performed at a resolution of 2 cm'1, with a zero-fill factor of 2, and using Happ-Genzel apodization. Each sample spectrum is collected with 100 scans, and background spectra are collected with 200 scans at regular intervals to ensure data quality.

[0051] Prior to chemometric analysis, the acquired spectra are preprocessed to enhance model performance and minimize errors. Preprocessing steps may include normalization of absorbance values between 0 and 1 to account for variations in sample contact with the ATR crystal. Additional preprocessing, such as derivatization and Savitzky-Golay smoothing, may be applied as needed to resolve overlapping bands or correct baseline drift.

[0052] Achemometric model, such as partial least squares regression (PLSR), is used to correlate the MIR spectral data with reference measurements of protein digestibility. The PLSR model is constructed using a training set of samples for which both spectral data and reference digestibility values (obtained via standard in vitro protein digestibility corrected amino acid score (PDCAAS) assays) are available. The model is optimized by selecting the appropriate number of latent variables (eigenfactors) to minimize root mean square errors of calibration (RMSEC), cross-validation (RMSECV), and prediction (RMSEP).The chemometric model may be trained and validated using K-fold cross validation or leave- one-out cross validation to ensure robustness and prevent overfitting. The model specifically analyzes the amide I band region, which contains information about protein secondary structures, such as alpha and beta sheets, that are directly related to protein digestibility.

[0053] Once the chemometric model is established, the system can rapidly predict the in vitro protein digestibility of unknown samples by analyzing their MIR spectra. The model detects patterns in the amide I band associated with alpha sheet signals, which are sensitive to enzymatic digestion and thus indicative of protein digestibility. The predicted digestibility values show high correlation with those obtained from traditional wet chemistry assays, but with significantly reduced analysis time and sample destruction.

[0054] The system may include a database for storing and retrieving phenotypic data, including spectral data, predicted digestibility values, and historical records for comparison. A graphical user interface may be provided to display results, visualize trends, and facilitate data management for breeding programs or food analysis laboratories.

[0055] This method enables rapid, cost-effective, and high-throughput phenotyping of protein digestibility and other nutritional traits in pulses and other plant materials The approach is particularly advantageous for breeding programs, where large numbers of samples must be screened efficiently,and for food and feed industries seeking to assess nutritional quality with minimal sample preparation and waste.

[0056] This approach provides non-destructive or minimally destructive analysis, requiring only a small amount of flour and preserving valuable germplasm. The method supports rapid throughput, with each analysis completed in approximately 1-2 minutes per sample Chemometric models provide high accuracy and robustness, with strong correlation to reference methods. The method is broadly applicable and can be extended to other crops and nutritional traits by appropriate model calibration. Integration with data management systems supports storage, retrieval, and visualization of phenotypic data.

[0057] In another example, pulse seeds are collected from breeding programs and ground to a maximum particle size of 0.5 mm. The ground flour is stored under controlled temperature and humidity conditions prior to analysis. For each breeding line, multiple spectra are collected and the most stable spectra are selected for calibration and validation.

[0058] Total nitrogen content is determined using combustion analysis, and SAA concentrations are measured using acid hydrolysis followed by high-performance liquid chromatography (HPLC). These reference values are used to calibrate and validate the chemometric models.

[0059] Spectral preprocessing is performed to enhance the quality and interpretability of the MIR data. Preprocessing steps may include normalization of spectra between 0 and 1 , application of the Savitzky- Golay first-order derivative and smoothing algorithm, and selection of specific spectral regions associated with the traits of interest. For example, the amide A, I, and II bands are used for protein analysis, while C-S and S-CFL stretching and bending bands are used for SAA analysis.

[0060] Partial least squares (PLS) regression may be employed as a chemometric modeling technique The PLS-1 algorithm is used to correlate MIR spectral data with reference concentrations of total protein and SAA. The models are constructed using calibration sets of breeding lines and validated with independent validation sets. The number of PLS factors is optimized for each model to minimize root mean square errors of calibration (RMSEC), cross-validation (RMSECV), and prediction (RMSEP).

[0061] Models are validated using full cross-validation or K-fold cross-validation The predictive performance of the models is assessed by comparing predicted values to actual reference values in the validation set. Statistical tests, such as pooled two-tailed t-tests, are used to confirm that there is no significant difference between actual and predicted means. The models demonstrate high coefficients of determination (R2), low RMSE values, and robust predictive ability across different crops and sample origins. The method enables rapid, nondestructive quantification of total protein and SAA in pulse flours The MIR spectral regions associated with protein functional groups (N-H and C=O) and SAA functional groups (C-S and S-CH3) are used to build reliable chemometric models. The approach is sensitive enough to quantify low-abundance amino acids, such as methionine and cysteine, in complex sample matrices.

[0062] The system supports high-throughput workflows suitable for plant breeding programs and food analysis laboratories. The FT-MIR technique requires minimal sample preparation, does not use hazardous chemicals, and allows for rapid analysis (typically less than a minute per sample). Themethod is cost-effective and does not require highly skilled operators. The compact instrumentation is suitable for deployment in a variety of laboratory and field settings.

[0063] This approach enables large-scale phenotyping of nutritional traits in pulse crops, supporting marker-assisted selection, genomic selection, and quantitative trait loci (QTL) discovery. The method reduces the time and cost associated with traditional wet chemistry techniques and is particularly advantageous for breeding programs that require analysis of thousands of samples. The technique is also applicable to food and feed industries for quality control and product development.

[0064] In yet another example, chickpea seeds may be obtained from breeding programs and ground to a maximum particle size of 0.5 mm. The ground flour is stored at approximately 10°C and 30-40% humidity prior to analysis. For each breeding line, multiple spectra are collected, and the most stable spectra are selected for calibration and validation.

[0065] Fatty acid content is determined using gas chromatography-mass spectrometry (GC-MS) as a reference method. Chickpea flour samples are extracted with hexane, and the fatty acids are converted to methyl esters for analysis. The GC-MS system is operated under standard conditions, and fatty acids are identified and quantified using selective ion monitoring and calibration with known standards.

[0066] Spectral preprocessing is performed to enhance the quality and interpretability of the FT-MIR data. Preprocessing steps include normalization of spectra between 0 and 1 , application of second- order derivatization, and Savitzky-Golay smoothing with a window size of 21 and a third-order polynomial. These steps minimize errors due to sample contact, resolve overlapping bands, and reduce baseline drift and noise. Interpolation may be used to adjust spectral resolution as needed.

[0067] Partial least squares (PLS) regression may be used as the chemometric modeling technique. Separate PLS models are constructed for TFA, TUSFA, and TSFA using calibration sets of breeding lines with both spectral and reference GC-MS data. The models are optimized by selecting the appropriate number of latent variables to minimize root mean square errors of calibration (RMSEC), cross-validation (RMSECV), and prediction (RMSEP). The selected spectral regions for modeling include 2845.82-3035.92 cm'1and 1720.17-1763.03 cm'1, which correspond to the ester carbonyl stretch and C-H stretching vibrations associated with fatty acids.

[0068] The PLS models are validated using leave-one-out cross-validation. The predictive performance of each model is assessed by comparing predicted values to actual GC-MS reference values in the validation set. The models demonstrate high coefficients of determination (R2), low RMSE values, and robust predictive ability for all three fatty acid traits. Statistical tests, such as two-tailed t- tests, confirm that there is no significant difference between actual and predicted means.

[0069] The method enables rapid, nondestructive quantification of TFA, TUSFA, and TSFA in chickpea flour. The FT-MIR spectral regions associated with lipid functional groups provide molecular fingerprints for accurate chemometric modeling. The approach is sensitive enough to quantify fatty acid content in complex sample matrices without the need for chemical pretreatment or hazardous reagents.

[0070] The system supports high-throughput workflows suitable for plant breeding programs and food analysis laboratories. The FT-MIR technique requires minimal sample preparation, does not use hazardous chemicals, and allows for rapid analysis (typically less than a minute per sample). Themethod is cost-effective and does not require highly skilled operators. The compact instrumentation is suitable for deployment in a variety of laboratory and field settings.

[0071] This approach enables large-scale phenotyping of fatty acid traits in chickpea and can be extended to other oilseed crops with appropriate model calibration The method reduces the time and cost associated with traditional wet chemistry techniques and is particularly advantageous for breeding programs that require analysis of thousands of samples The technique is also applicable to food and feed industries for quality control and product development.

[0072] In still yet another example, seeds from pulse crops such as dry pea, chickpea, and lentil are obtained from breeding programs and ground to a maximum particle size of 0.5 mm. The ground flour is stored under controlled temperature and humidity prior to analysis. For each breeding line, multiple spectra are collected, and the most stable spectra are selected for calibration and validation.

[0073] Total starch and resistant starch concentrations are determined using a modified enzymatic assay. Finely ground seed samples are incubated with amyloglucosidase and a-amylase to hydrolyze starch. The resulting fractions are separated, and glucose is quantified using a glucose oxidase / peroxidase colorimetric assay. Calculations are performed to determine the concentrations of nonresistant starch (NRS), RS, and TS in each sample. These reference values are used to calibrate and validate the chemometric models.

[0074] Spectral preprocessing is performed to enhance the quality and interpretability of the FT-MIR data. Preprocessing steps may include normalization of spectra between 0 and 1 , selection of specific spectral regions associated with starch content, and averaging of replicate spectra for each sample. The selected spectral regions for modeling include 880-1476 cm’1for total starch and 1180-1486 cm’1and 3055-3614 cm’1for resistant starch, which correspond to characteristic carbohydrate functional group vibrations

[0075] Partial least squares (PLS) regression is used as the principal chemometric modeling technique Separate PLS models are constructed for TS and RS using calibration sets of breeding lines with both spectral and reference enzymatic assay data. The models are optimized by selecting the appropriate number of latent variables to minimize root mean square errors of calibration (RMSEC), cross-validation (RMSECV), and prediction (RMSEP).

[0076] The PLS models are validated using leave-one-out cross-validation. The predictive performance of each model is assessed by comparing predicted values to actual reference values in the validation set. The models demonstrate high coefficients of determination (R2), low RMSE values, and robust predictive ability for both total and resistant starch across dry pea, chickpea, and lentil samples. Statistical tests confirm that there is no significant difference between actual and predicted means.

[0077] The method enables rapid, nondestructive quantification of TS and RS in pulse flours. The FT- MIR spectral regions associated with carbohydrate functional groups provide molecular fingerprints for accurate chemometric modeling The approach is sensitive enough to quantify starch content in complex sample matrices without the need for chemical pretreatment or hazardous reagents.

[0078] The system supports high-throughput workflows suitable for plant breeding programs and food analysis laboratories. The FT-MIR technique requires minimal sample preparation, does not usehazardous chemicals, and allows for rapid analysis (typically less than a minute per sample). The method is cost-effective and does not require highly skilled operators. The compact instrumentation is suitable for deployment in a variety of laboratory and field settings.

[0079] This approach enables large-scale phenotyping of starch traits in pulse crops and can be extended to other plant species with appropriate model calibration The method reduces the time and cost associated with traditional wet chemistry techniques and is particularly advantageous for breeding programs that require analysis of thousands of samples The technique is also applicable to food and feed industries for quality control and product development.

[0080] In still yet another example, the disclosure describes a high-throughput, nondestructive method for quantifying total dietary fiber (TDF) in pulse crops using FT-MIR spectroscopy in combination with chemometric modeling. The method is designed to address the limitations of traditional dietary fiber analysis, which is time-consuming, expensive, and requires extensive sample preparation and hazardous chemicals.

[0081] A Fourier-transform infrared (FTIR) spectrometer equipped with a diamond attenuated total reflectance (ATR) module is used to acquire mid-infrared spectra from ground pulse flour samples. Spectral data are collected in the range of 650-4000 cm'1under Happ-Genzel apodization. The instrument is typically set to a resolution of 2 or 4 cm'1, with a zero-fill factor of 2 or 4, and 100 to 200 scans per sample and background, respectively. The ATR surface is cleaned with HPLC-grade methanol before each measurement to ensure data quality and prevent cross-contamination.

[0082] Seeds from pulse crops such as chickpea (Cicer arietinum L), dry pea (Pisum sativum L.), and lentil (Lens culinaris Medik.) are obtained from breeding programs and ground to a maximum particle size of 0.5 mm. The ground flour is stored under controlled temperature and humidity prior to analysis. For each calibration sample, multiple subsamples are analyzed, and the resulting spectra are averaged to ensure homogeneity and reproducibility. For validation, flour from a minimal number of seeds (e.g., two) is used to demonstrate the method’s applicability to small sample sizes.

[0083] Total dietary fiber concentrations are determined using a modified version of the Association of Official Analytical Collaboration (AOAC) method 985.29. This enzymatic-gravimetric method involves sequential digestion of flour samples with heat-stable amylase, protease, and amyloglucosidase, followed by precipitation of dietary fiber with ethanol and vacuum filtration. The extracted dietary fiber is weighed, and the resulting values are used as reference data for calibration and validation of the chemometric models.

[0084] Spectral preprocessing is performed to enhance the quality and interpretability of the FT-MIR data. Preprocessing steps include normalization of spectra between 0 and 1 to minimize errors due to sample contact, and application of Savitzky-Golay (SG) smoothing with a window size of 101 and a third-order polynomial. This smoothing reduces noise and simulates the less-structured spectrum of extracted dietary fiber, improving the correlation between spectral data and TDF content. The selected spectral regions for modeling include 650-1800 cm'1and 2771-3700 cm'1, which correspond to characteristic vibrations of polysaccharides, undigested proteins, and minor fat components.

[0085] PLSR is used as the principal chemometric modeling technique. Separate PLSR models are constructed for each crop using calibration sets of samples with both spectral and reference AOACdata. The models are optimized by selecting the appropriate number of latent variables to minimize root mean square errors of calibration (RMSEC), cross-validation (RMSECV), and prediction (RMSEP). Both leave-one-out cross-validation (LOOCV) and K-fold cross-validation (KFCV) are used to validate the models and ensure robust predictive performance.

[0086] The PLSR models are validated by comparing predicted TDF values to actual reference values in independent validation sets. The models demonstrate high coefficients of determination ($RA2$), low RMSE values, and robust predictive ability for TDF in chickpea, dry pea, and lentil. Statistical tests, such as two-tailed t-tests and F-tests, confirm that there is no significant difference between actual and predicted means and variances, supporting the reliability of the method for both large and small sample sizes.

[0087] The method enables rapid, nondestructive quantification of TDF in pulse flours. The FT-MIR spectral regions associated with carbohydrate and protein functional groups provide molecular fingerprints for accurate chemometric modeling The approach is sensitive enough to quantify dietary fiber content in complex sample matrices without the need for chemical pretreatment or hazardous reagents.

[0088] The system supports high-throughput workflows suitable for plant breeding programs and food analysis laboratories. The FT-MIR technique requires minimal sample preparation, does not use hazardous chemicals, and allows for rapid analysis (typically less than a minute per sample). The method is cost-effective and does not require highly skilled operators. The compact instrumentation is suitable for deployment in a variety of laboratory and field settings.

[0089] This approach enables large-scale phenotyping of dietary fiber traits in pulse crops and can be extended to other plant species with appropriate model calibration The method reduces the time and cost associated with traditional wet chemistry techniques and is particularly advantageous for breeding programs that require analysis of thousands of samples The technique is also applicable to food and feed industries for quality control and product development.

[0090] The described method provides a robust, rapid, and nondestructive solution for quantifying total dietary fiber in pulse flours using FT-MIR spectroscopy and chemometric modeling. The approach is validated with high accuracy and reproducibility and is suitable for integration into high-throughput phenotyping and food quality assessment pipelines.

[0091] In still yet another example, the disclosure describes a cloud-based platform designed for high- throughput phenotyping of nutritional quality traits in food crops using Fourier-transform mid-infrared (FT-MIR) spectroscopy and chemometric modeling. The platform enables remote, real-time analysis of large phenotyping datasets, supporting plant breeders, growers, and food analytical laboratories in the rapid assessment of nutritional traits.

[0092] The platform is implemented as a user-friendly web application, developed using the Django web framework and Python 3 (although similar platforms could be developed using other coding languages) Users access the platform through a standard web browser and can register for secure, authenticated access. Upon logging in, users select the crop type (such as dry pea, chickpea, or lentil) and the nutritional trait of interest. The platform supports a range of traits, including total protein, totalstarch, resistant starch, sulfur-containing amino acids, total fatty acids, dietary fiber, moisture, and protein digestibility.

[0093] Users upload FT-MIR spectral data files in formats such as .txt, .csv, or any other file format capable of conveying spectral data in a manner interpretable by the computing devices described herein. Each file contains two columns: wavenumbers and absorbance values. The platform validates the file format and content before proceeding. Once the data are uploaded, the system automatically applies the appropriate spectral preprocessing steps, which may include normalization, first or second derivatives, Savitzky-Golay smoothing, and interpolation to match the spectral resolution required by the chemometric models.

[0094] The platform integrates pre-trained partial least squares regression (PLSR) models for each crop-trait combination. These models are built using calibration datasets and validated using K-fold cross validation to ensure robust and efficient performance. The PLSR models correlate the preprocessed FT-MIR spectra with reference nutritional trait values, enabling accurate prediction of the trait for each sample.

[0095] Spectral preprocessing is tailored to each trait and may involve normalization between 0 and 1, application of first or second derivatives, Savitzky-Golay smoothing with a specified window and polynomial order, and interpolation to achieve the required spectral resolution (e.g., 1.86 cm'1or 0.47 cm'1). The preprocessing pipeline ensures that uploaded spectra are compatible with the calibration data used to train the PLSR models.

[0096] The platform executes all computations on a local or remote server, leveraging optimized vector operations and the Scikit-learn library for efficient data processing The system is designed for horizontal scalability, allowing multiple users to analyze large datasets concurrently without significant delays.

[0097] The PLSR models are validated using K-fold cross validation, which divides the dataset into k subsets and iteratively trains and tests the model on different folds. This approach balances computational efficiency and model robustness, making it suitable for real-time, high-throughput applications The models achieve high coefficients of determination (R2) and low root mean square errors of calibration (RMSEC) and cross-validation (RMSECV), indicating strong predictive performance.

[0098] Benchmarking tests demonstrate that the platform can process thousands of spectra per second, with average processing times for 2,000 spectra typically under one second per user on standard hardware. The system’s memory management and computational efficiency are enhanced by the use of NumPy arrays and optimized Python data structures.

[0099] The platform provides a graphical user interface that guides users through the workflow, from data upload to result generation. Results are displayed in real time and can be downloaded for further analysis. The system supports storage and retrieval of historical data, enabling users to compare current results with previous analyses and track trends over time

[0100] The web-based architecture ensures compatibility with any device capable of running a web browser, providing enhanced accessibility compared to traditional, on-site proprietary software. Theplatform can be deployed on local or cloud servers, supporting both standalone and distributed use cases.

[0101] This cloud-based platform enables rapid, remote, and high-throughput phenotyping of nutritional traits in food crops. The system eliminates the need for on-site chemometric model development and proprietary software, reducing barriers to adoption for users worldwide. The platform is particularly advantageous for plant breeding programs, food and feed industries, and research laboratories that require efficient analysis of large numbers of samples.

[0102] The automated workflow, robust model validation, and scalable architecture support consistent, real-time data analysis for multiple users. The system’s compatibility with FT-MIR spectra from any commercial spectrometer further broadens its applicability.

[0103] The described cloud-based platform provides a scalable, efficient, and user-friendly solution for high-throughput phenotyping of nutritional traits in food crops. By integrating FT-MIR spectroscopy, advanced chemometric modeling, and cloud computing, the system delivers rapid, accurate, and accessible nutritional analysis to support breeding, research, and quality control efforts.

[0104] FIG 2 is a block diagram illustrating a more detailed example of a computing device configured to perform the techniques described herein. Computing device 210 of FIG. 2 is described below as an example of computing device 110 of FIG. 1. FIG. 2 illustrates only one particular example of computing device 210, and many other examples of computing device 210 may be used in other instances and may include a subset of the components included in example computing device 210 or may include additional components not shown in FIG. 2.

[0105] Computing device 210 may be any computer with the processing power required to adequately execute the techniques described herein. For instance, computing device 210 may be any one or more of a mobile computing device (e.g., a smartphone, a tablet computer, a laptop computer, etc.), a desktop computer, a smarthome component (e.g., a computerized appliance, a home security system, a control panel for home components, a lighting system, a smart power outlet, etc.), an integrated computer system, a vehicle, a wearable computing device (e.g., a smart watch, computerized glasses, a heart monitor, a glucose monitor, smart headphones, etc.), a virtual reality / augmented reality / extended reality (VR / AR / XR) system, a video game or streaming system, a network modem, router, or server system, or any other computerized device that may be configured to perform the techniques described herein.

[0106] As shown in the example of FIG. 2, computing device 210 includes user interface components (UIC) 212, one or more processors 240, one or more communication units 242, one or more input components 244, one or more output components 246, and one or more storage components 248. UIC 212 includes display component 202 and presence-sensitive input component 204. Storage components 248 of computing device 210 include communication module 220, analysis module 222, and data store 226.

[0107] One or more processors 240 may implement functionality and / or execute instructions associated with computing device 210 to determine characteristics of portions of plants That is, processors 240 may implement functionality and / or execute instructions associated with computing device 210 to control mid-infrared sensors to capture spectral images of portions of plants and analyze said spectral images with chemometric models to determine one or more characteristics of the plant.

[0108] Examples of processors 240 include any combination of application processors, display controllers, auxiliary processors, one or more sensor hubs, and any other hardware configured to function as a processor, a processing unit, or a processing device, including dedicated graphical processing units (GPUs). Modules 220 and 222 may be operable by processors 240 to perform various actions, operations, or functions of computing device 210. For example, processors 240 of computing device 210 may retrieve and execute instructions stored by storage components 248 that cause processors 240 to perform the operations described with respect to modules 220 and 222. The instructions, when executed by processors 240, may cause computing device 210 to control midinfrared sensors to capture spectral images of portions of plants and analyze said spectral images with chemometric models to determine one or more characteristics of the plant.

[0109] Communication module 220 may execute locally (e.g., at processors 240) to provide functions associated with communicating with external devices, such as mid-infrared sensors or cloud servers that store data. In some examples, communication module 220 may act as an interface to a remote service accessible to computing device 210. For example, communication module 220 may be an interface or application programming interface (API) to a remote server that controls the mid-infrared sensors or manages the cloud interface.

[0110] In some examples, analysis module 222 may execute locally (e.g , at processors 240) to provide functions associated with analyzing mid-infrared spectral images using chemometric models and determining the characteristics of the plants. In some examples, analysis module 222 may act as an interface to a remote service accessible to computing device 210. For example, analysis module 222 may be an interface or application programming interface (API) to a remote server that analyzes mid-infrared spectral images using chemometric models and determines the characteristics of the plants.

[0111] One or more storage components 248 within computing device 210 may store information for processing during operation of computing device 210 (e.g , computing device 210 may store data accessed by modules 220 and 222 during execution at computing device 210). In some examples, storage component 248 is a temporary memory, meaning that a primary purpose of storage component 248 is not long-term storage Storage components 248 on computing device 210 may be configured for short-term storage of information as volatile memory and therefore not retain stored contents if powered off. Examples of volatile memories include random access memories (RAM), dynamic random access memories (DRAM), static random access memories (SRAM), and other forms of volatile memories known in the art.

[0112] Storage components 248, in some examples, also include one or more computer-readable storage media. Storage components 248 in some examples include one or more non-transitory computer-readable storage mediums. Storage components 248 may be configured to store larger amounts of information than typically stored by volatile memory. Storage components 248 may further be configured for long-term storage of information as non-volatile memory space and retain information after power on / off cycles. Examples of non-volatile memories include magnetic hard discs, optical discs, floppy discs, flash memories, or forms of electrically programmable memories (EPROM) or electrically erasable and programmable (EEPROM) memories. In this instance, storage components248 may store ledger information (e.g., in data store 226), such as for blockchain technology. A blockchain ledger is a continuously growing list of records (or blocks) that are linked and secured using cryptographic techniques. As this ledger can become quite large over time, it may be stored on persistent storage to retain data even when computing device 210 is turned off.

[0113] Storage components 248 may store program instructions and / or information (e.g., data) associated with modules 220 and 222 and data store 226. Storage components 248 may include a memory configured to store data or other information associated with modules 220 and 222 and data store 226.

[0114] Communication channels 250 may interconnect each of the components 212, 240, 242, 244, 246, and 248 for inter-component communications (physically, communicatively, and / or operatively). In some examples, communication channels 250 may include a system bus, a network connection, an inter-process communication data structure, or any other method for communicating data.

[0115] One or more communication units 242 of computing device 210 may communicate with external devices via one or more wired and / or wireless networks by transmitting and / or receiving network signals on one or more networks. Examples of communication units 242 include a network interface card (e.g., such as an Ethernet card), an optical transceiver, a radio frequency transceiver, a GPS receiver, a radio-frequency identification (RFID) transceiver, a near-field communication (NFC) transceiver, or any other type of device that can send and / or receive information. Other examples of communication units 242 may include short wave radios, cellular data radios, wireless network radios, as well as universal serial bus (USB) controllers.

[0116] One or more input components 244 of computing device 210 may receive input Examples of input are tactile, audio, and video input. Input components 244 of computing device 210, in one example, include a presence-sensitive input device (e.g., a touch sensitive screen, a PSD), mouse, keyboard, voice responsive system, camera, microphone or any other type of device for detecting input from a human or machine. In some examples, input components 244 may include one or more sensor components (e.g., sensors 252). Sensors 252 may include one or more biometric sensors (e.g., fingerprint sensors, retina scanners, vocal input sensors / microphones, facial recognition sensors, cameras), one or more location sensors (e.g., GPS components, Wi-Fi components, cellular components), one or more temperature sensors, one or more movement sensors (e.g., accelerometers, gyros), one or more pressure sensors (e.g., barometer), one or more ambient light sensors, and one or more other sensors (e.g., infrared proximity sensor, hygrometer sensor, and the like). Other sensors, to name a few other non-limiting examples, may include a radar sensor, a lidar sensor, a sonar sensor, a heart rate sensor, magnetometer, glucose sensor, olfactory sensor, compass sensor, or a step counter sensor.

[0117] One or more output components 246 of computing device 210 may generate output in a selected modality. Examples of modalities may include a tactile notification, audible notification, visual notification, machine generated voice notification, or other modalities Output components 246 of computing device 210, in one example, include a presence-sensitive display, a sound card, a video graphics adapter card, a speaker, a cathode ray tube (CRT) monitor, a liquid crystal display (LCD), a light emitting diode (LED) display, an organic LED (OLED) display, a virtual / augmented / extended reality(VR / AR / XR) system, a three-dimensional display, or any other type of device for generating output to a human or machine in a selected modality.

[0118] UIC 212 of computing device 210 may include display component 202 and presence-sensitive input component 204. Display component 202 may be a screen, such as any of the displays or systems described with respect to output components 246, at which information (e.g., a visual indication) is displayed by UIC 212 while presence-sensitive input component 204 may detect an object at and / or near display component 202.

[0119] While illustrated as an internal component of computing device 210, UIC 212 may also represent an external component that shares a data path with computing device 210 for transmitting and / or receiving input and output For instance, in one example, UIC 212 represents a built-in component of computing device 210 located within and physically connected to the external packaging of computing device 210 (e.g., a screen on a mobile phone). In another example, UIC 212 represents an external component of computing device 210 located outside and physically separated from the packaging or housing of computing device 210 (e.g., a monitor, a projector, etc. that shares a wired and / or wireless data path with computing device 210).

[0120] UIC 212 of computing device 210 may detect two-dimensional and / or three-dimensional gestures as input from a user of computing device 210. For instance, a sensor of UIC 212 may detect a user's movement (e.g., moving a hand, an arm, a pen, a stylus, a tactile object, etc.) within a threshold distance of the sensor of UIC 212. UIC 212 may determine a two or three-dimensional vector representation of the movement and correlate the vector representation to a gesture input (e.g., a handwave, a pinch, a clap, a pen stroke, etc.) that has multiple dimensions. In other words, UIC 212 can detect a multi-dimension gesture without requiring the user to gesture at or near a screen or surface at which UIC 212 outputs information fordisplay. Instead, UIC 212 can detect a multi-dimensional gesture performed at or near a sensor which may or may not be located near the screen or surface at which UIC 212 outputs information for display.

[0121] In accordance with the techniques of this disclosure, communication module 220 may control one or more mid-infrared sensors to capture one or more mid-infrared spectral images of at least a portion of a plant. In some instances, the portion of the plant may be a pulse, or a dry seed derived from the plant. In some examples, the pulse remains intact after the one or more mid-infrared spectral images are captured, thereby preserving the sample.

[0122] In some instances, the plant may be in a Fabaceae family, although the plant may also belong to other classification families. For instance, the plant may be any of a dry pea, a lentil, a chickpea, a casein, a sunflower, an oat, a split pea, a kidney bean wheat, a black eyed pea, any food (e.g., solid or dry feed materials), or any biological material, although any other plant where a pulse or other portion of the plant where aspects of the plant may be determined using mid-infrared imaging may be examined in accordance with the techniques of this disclosure.

[0123] Communication module 220 may receive the one or more mid-infrared spectral images Analysis module 222 may analyze the received one or more mid-infrared spectral images using a chemometric model. In some instances, the chemometric model may be a partial least squares regression model. Additionally or alternatively, the chemometric model may include one or moreinterpolation filters. Additionally or alternatively, the chemometric model is trained and / or validated using K-fold cross validation. In some instances, the K-fold cross validation includes at least five folds. In some such instances, the mid-infrared spectral images are normalized between 0 and 1 prior to analysis. Additionally or alternatively, analysis module 222 may interpolate the spectral data to achieve a spectral resolution of 1.86 cm-1. In some instances, the mid-infrared spectral images are captured at a resolution of 2 cm-1 and with a zero-fill factor of 2.

[0124] Analysis module 222 may determine, based at least in part on the analyzing, one or more characteristics of the portion of the plant. Examples of these characteristics include any one or more of a total fatty acid, a total saturated fatty acid, a total unsaturated fatty acid, a total dietary fiber, a total protein, a total resistant starch, a total starch, a total sulfur containing amino acids, and a digestibility factor (e.g., an analysis of the quantity, types, and structures of the protein and amino acids in the plant).

[0125] For example, when the one or more characteristics include a digestibility factor, and in analyzing the one or more mid-infrared spectral images, analysis module 222 may apply the chemometric model to the one or more mid-infrared spectral images to detect one or more patterns of alpha sheet signals within an amide I band of the one or more mid-infrared spectral images. Analysis module 222 may then estimate the digestibility factor based on the one or more patterns.

[0126] In some examples, communication module 220 may store the one or more characteristics of the portion of the plant in a database, such as data store 226 or one or more remote servers accessible by one or more computing devices (e.g., a cloud database). Communication module 220 may retrieve historical characteristics of plants of a same species as the plant from the database. Analysis module 222 may generate a graphical user interface comprising graphical indications of at least one of the one or more characteristics and at least one of the historical characteristics, thereby displaying a trend or a larger sample of pulses from that plant. Communication module 220 may output, to an output device (e.g., display component 202 or a remote computing device), the graphical user interface.

[0127] FIG 3 is a flow chart illustrating an example mode of operation. The techniques of FIG. 3 may be performed by one or more processors of a computing device, such as system 100 of FIG. 1 and / or computing device 210 illustrated in FIG. 2. For purposes of illustration only, the techniques of FIG. 3 are described within the context of computing device 210 of FIG. 2, although computing devices having configurations different than that of computing device 210 may perform the techniques of FIG. 3.

[0128] In accordance with the techniques of this disclosure, communication module 220 controls one or more mid-infrared sensors to capture one or more mid-infrared spectral images of at least a portion of a plant (302). Communication module 220 receives the one or more mid-infrared spectral images (304). Analysis module 222 analyzes the one or more mid-infrared spectral images using a chemometric model (306). Analysis module 222 determines, based at least in part on the analyzing, one or more characteristics of the portion of the plant (308).

[0129] FIG 4 is a flow diagram illustrating a schematic representation of FT-MIR spectroscopy as applied with high-throughput phenotyping, in accordance with one or more of the techniques of this disclosure. The techniques of FIG. 4 may be performed by one or more processors of a computing device, such as system 100 of FIG. 1 and / or computing device 210 illustrated in FIG. 2. For purposesof illustration only, the techniques of FIG. 4 are described within the context of computing device 210 of FIG. 2, although computing devices having configurations different than that of computing device 210 may perform the techniques of FIG. 4.

[0130] Pulse 402 represents the starting material for the high-throughput phenotyping process. For example, pulse 402 may include seeds derived from various crops, such as dry pea (Pisum sativum L), chickpea (Cicer arietinum L.), lentil (Lens culinaris Medik.), or other pulse crops. In general, these seeds are rich in macronutrients, including carbohydrates, proteins, and fats, as well as micronutrients such as vitamins and minerals. As a result, pulse 402 serves as the biological sample from which nutritional traits are analyzed. The condition of pulse 402 may be representative of the crop’s genetic and environmental variability. In some embodiments, pulse 402 may be intact or partially processed, depending on the requirements of subsequent steps.

[0131] The grinding and milling 404 step involves mechanical processing of pulse 402 to produce a fine, homogeneous flour suitable for spectral analysis. For instance, grinding and milling 404 may be performed using a cyclone grinder, a blade grinder, or other milling equipment capable of reducing particle size to a maximum of 0.5 mm. These operations promote uniformity in the sample, as variations in particle size or composition can influence the accuracy of the Fourier-transform mid-infrared (FT- MIR) spectral data collected in subsequent steps. In an alternative embodiment, the resulting flour is stored under controlled conditions, such as low temperature and humidity, to maintain its chemical and physical properties. This step reduces sample heterogeneity and ensures that the spectral data accurately represent the nutritional composition of pulse 402

[0132] FT-MIR spectral collection 406 involves acquisition of mid-infrared spectral data from ground pulse flour. Specifically, this step is performed using a Fourier-transform mid-infrared spectrometer equipped with an attenuated total reflectance (ATR) accessory. The spectrometer collects spectral data in the mid-infrared range (e.g., typically between 650 and 4000 cm"1) under optimized conditions, such as a resolution of 2 cm"1and a zero-fill factor of 2. For example, the spectral data capture the vibrational modes of molecular functional groups, such as C=O, N-H, and O-H, which are associated with the macronutrients and other chemical components in the pulse flour. In general, FT-MIR spectral collection 406 is a rapid, non-destructive process that generates a distinct spectral fingerprint for each sample, thereby enabling the identification and quantification of nutritional traits.

[0133] Preprocessing 408 involves mathematical and computational treatment of the raw FT-MIR spectral data to enhance quality and suitability for analysis For instance, preprocessing 408 may include normalization, first and second derivatives, Savitzky-Golay smoothing, and spectral interpolation. Normalization adjusts the spectral data to a consistent scale (e.g., between 0 and 1 ) to minimize errors arising from variations in sample thickness or contact with the ATR crystal Derivatives and smoothing algorithms are applied to correct baseline drift, resolve overlapping peaks, and enhance spectral feature resolution. In some embodiments, spectral interpolation adjusts data resolution to match calibration data used in subsequent chemometric modeling. Through these processes, preprocessing 408 refines spectral data by removing noise and artifacts, thereby supporting accurate and reliable analysis.

[0134] Partial least squares (PLS) regression 410 involves application of chemometric modeling to preprocessed spectral data to predict concentrations of nutritional traits in pulse flour. Specifically, PLS regression 410 correlates spectral data (predictor variables) with reference data (response variables) obtained from traditional analytical methods, such as enzymatic assays or chromatography The PLS regression 410 step decomposes the spectral data into latent variables that capture variance associated with nutritional traits of interest. In an alternative example, the model is trained and validated using calibration and validation datasets to ensure high predictive accuracy. Accordingly, nutritional traits that can be predicted using PLS regression 410 include total protein, total starch, resistant starch, total fatty acids, sulfur-containing amino acids, and protein digestibility.

[0135] High-throughput phenotyping 412 represents the culmination of the process, where nutritional traits of pulse flour are quantified and reported. For instance, this step leverages the predictive capabilities of the PLS regression model to analyze large datasets rapidly and with high effectiveness. In general, high-throughput phenotyping 412 facilitates the screening of thousands of samples in a short time, making the process suitable for breeding programs, food quality assessments, and industrial applications The results generated in this step can be applied to identify elite germplasm with desirable nutritional profiles, optimize food formulations, or support research in plant and nutritional sciences. High-throughput phenotyping 412 reduces the time, cost, and labor associated with traditional phenotyping methods, providing a scalable approach for large-scale nutritional analysis.

[0136] Traditional methods for analyzing nutritional traits in pulse crops, such as dry pea, lentil, and chickpea, rely heavily on wet chemistry techniques, including enzymatic assays, chromatography, and mass spectrometry. While these methods provide accurate results, they are time-consuming, labor- intensive, and require significant sample preparation, often destroying the sample in the process. For example, protein digestibility is typically measured using the protein digestibility corrected amino acid score (PDCAAS) assay, which involves extensive enzymatic digestion and chemical analysis, taking up to 24 hours per sample. Similarly, starch and resistant starch analysis require enzymatic hydrolysis and colorimetric assays, which are costly, low-throughput, and unsuitable for large-scale breeding programs. These conventional approaches impose significant bottlenecks in phenotyping workflows, particularly in breeding programs where thousands of samples need to be analyzed rapidly and cost- effectively. Furthermore, the reliance on hazardous chemicals and the generation of chemical waste pose environmental and operational challenges.

[0137] The present approach addresses these limitations by leveraging Fourier-transform mid-infrared (FT-MIR) spectroscopy combined with advanced chemometric modeling techniques, such as partial least squares regression (PLSR). FT-MIR spectroscopy provides a rapid, non-destructive, and high- throughput method for analyzing nutritional traits, including protein digestibility, total starch, resistant starch, total fatty acids, sulfur-containing amino acids, and other macronutrients. Unlike conventional methods, FT-MIR spectroscopy requires minimal sample preparation and can analyze flour derived from a single seed, preserving the sample for further use This methodology employs specialized algorithms and preprocessing techniques, such as spectral normalization, Savitzky-Golay smoothing, and interpolation filters, to improve the accuracy and reliability of the chemometric models. Thesemodels are trained and validated using robust cross-validation techniques, including k-fold cross- validation, to ensure consistent performance across diverse sample matrices.

[0138] Additionally, the described technology may incorporate a cloud-based platform to enable remote, real-time analysis of FT-MIR spectral data. This platform integrates pre-trained PLSR models and supports multi-user access, allowing breeders, growers, and food analysts to upload spectral data and obtain phenotypic results without the need for on-site chemometric expertise. The system architecture is optimized for high-throughput processing, capable of analyzing thousands of spectral datasets in seconds, and is compatible with FT-MIR spectrometers from various manufacturers. By eliminating the need for expensive, proprietary software and enabling remote access, the described technology significantly reduces operational costs and accelerates decision-making in breeding programs and food industries.

[0139] The described solution further addresses the inefficiencies and constraints of conventional phenotyping methods by introducing a rapid, cost-effective, and scalable approach that integrates FT- MIR spectroscopy, advanced chemometric modeling, and cloud-based computational infrastructure. This methodology improves the accuracy and throughput of nutritional trait analysis while broadening access to high-quality phenotyping tools, encouraging advancements in pulse breeding and food processing industries.

[0140] The subject matter described herein is directed to a practical application of FT-MIR spectroscopy and chemometric modeling for the rapid, high-throughput phenotyping of nutritional traits in plant materials. The disclosed methods and systems provide a concrete technological solution to longstanding problems in plant breeding, food analysis, and nutritional quality assessment.

[0141] The described process involves the use of specific hardware components, such as mid-infrared spectrometers and computing devices, to capture, preprocess, and analyze spectral data from biological samples. The method includes tangible steps such as sample preparation, spectral data acquisition, spectral preprocessing, and the application of trained chemometric models to determine physical and chemical characteristics of plant materials. These steps result in a transformation of raw plant material and spectral data into actionable information regarding nutritional content, which is of significant utility in breeding programs, food quality control, and industrial applications.

[0142] Furthermore, the system incorporates a cloud-based computational platform that enables remote, real-time analysis and storage of phenotypic data, further demonstrating a specific and practical application of computer technology The integration of spectral preprocessing, chemometric modeling, and cloud-based data management provides improvements in the speed, accuracy, scalability, and accessibility of nutritional trait analysis, yielding results that cannot be achieved by mental processes or generic computer implementation alone.

[0143] Accordingly, the subject matter as claimed is directed to a specific, technical process that applies and improves technology. The described methods and systems provide a technical solution to a technical problem and produce a useful, concrete, and tangible result

[0144] The techniques described herein include on a system and method for rapid, high-throughput phenotyping of nutritional quality traits in plant materials using FT-MIR spectroscopy combined with advanced chemometric modeling. The approach enables the non-destructive or minimally destructiveanalysis of plant samples, such as pulses, by capturing mid-infrared spectral images and applying PLSR or similar models to determine key nutritional characteristics, including protein, fatty acids, starch, resistant starch, sulfur-containing amino acids, and protein digestibility.

[0145] The system further incorporates automated spectral preprocessing, such as normalization, derivatization, and smoothing, to enhance data quality and model accuracy. Chemometric models are trained and validated using robust cross-validation techniques, including K-fold cross validation, to ensure reliable performance across diverse sample types. The workflow is supported by a cloud-based computational platform that allows remote, real-time analysis and multi-user access, eliminating the need for on-site chemometric expertise and proprietary software.

[0146] This concept provides a practical, scalable, and cost-effective solution for plant breeders, food analysts, and researchers, enabling the rapid screening of large sample sets and supporting data-driven decision-making in breeding programs and food quality assessment. The integration of FT-MIR spectroscopy, chemometric modeling, and cloud-based analytics represents a significant advancement over traditional wet chemistry methods, offering improvements in speed, efficiency, and accessibility for nutritional phenotyping.

[0147] In an example of the techniques described herein, a plant breeding program at a major agricultural research university seeks to develop new chickpea cultivars with improved nutritional profiles, specifically targeting higher protein content, increased resistant starch, and enhanced protein digestibility. Traditionally, the breeding team relied on wet chemistry assays and chromatography to analyze these traits, which required destructive sampling, extensive labor, and weeks of turnaround time for each batch of samples

[0148] To accelerate the breeding process, the program implements the described FT-MIR spectroscopy and cloud-based phenotyping system. Breeders collect seeds from hundreds of chickpea breeding lines grown in field trials. Each seed sample is ground into flour using a standardized milling protocol. The flour samples are then analyzed using a benchtop FT-MIR spectrometer equipped with an ATR module The spectrometer rapidly collects mid-infrared spectra for each sample, typically in less than one minute per sample.

[0149] The spectral data are uploaded directly to the cloud platform via a secure web interface. The platform automatically preprocesses the spectra, applying normalization, smoothing, and interpolation as required, and then applies pre-trained partial least squares regression (PLSR) models to predict total protein, resistant starch, and protein digestibility for each sample The PLSR models have been previously calibrated and validated using reference data from traditional assays and are maintained and updated by the platform administrators.

[0150] Within minutes, the breeders receive a comprehensive report for each breeding line, including predicted values for all targeted nutritional traits. The results are visualized through the platform’s graphical user interface, allowing breeders to compare current data with historical records and identify top-performing lines The cloud-based system supports multi-user access, enabling breeders, nutritionists, and data analysts to collaborate remotely and make informed selection decisions in real time.

[0151] By adopting this system, the breeding program reduces analysis time from weeks to hours, eliminates the need for hazardous chemicals, and preserves valuable seed material for further testing or planting. The rapid, high-throughput workflow enables the team to screen thousands of lines each season, accelerating the development of nutritionally superior chickpea cultivars and supporting the program’s goals for food security and public health.

[0152] Although the various examples have been described with reference to preferred implementations, persons skilled in the art will recognize that changes may be made in form and detail without departing from the spirit and scope thereof.

[0153] It is to be recognized that depending on the example, certain acts or events of any of the techniques described herein can be performed in a different sequence, may be added, merged, or left out altogether (e.g., not all described acts or events are necessary for the practice of the techniques). Moreover, in certain examples, acts or events may be performed concurrently, e.g., through multithreaded processing, interrupt processing, or multiple processors, rather than sequentially

[0154] In one or more examples, the functions described may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functions may be stored on or transmitted over as one or more instructions or code on a computer-readable medium and executed by a hardware-based processing unit. Computer-readable media may include computer-readable storage media, which corresponds to a tangible medium such as data storage media, or communication media including any medium that facilitates transfer of a computer program from one place to another, e.g., according to a communication protocol. In this manner, computer-readable media generally may correspond to (1) tangible computer-readable storage media which is non-transitory or (2) a communication medium such as a signal or carrier wave. Data storage media may be any available media that can be accessed by one or more computers or one or more processors to retrieve instructions, code and / or data structures for implementation of the techniques described in this disclosure. A computer program product may include a computer-readable medium.

[0155] It is contemplated that the various aspects, features, processes, and operations from the various embodiments may be used in any of the other embodiments unless expressly stated to the contrary. Certain operations illustrated may be implemented by a computer executing a computer program product on a non-transient, computer-readable storage medium, where the computer program product includes instructions causing the computer to execute one or more of the operations, orto issue commands to other devices to execute one or more operations.

[0156] By way of example, and not limitation, such computer-readable storage media can comprise RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage, or other magnetic storage devices, flash memory, or any other medium that can be used to store desired program code in the form of instructions or data structures and that can be accessed by a computer. Also, any connection is properly termed a computer-readable medium For example, if instructions are transmitted from a website, server, or other remote source using a coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared, radio, and microwave, then the coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave are included in the definition of medium. It should be understood,however, that computer-readable storage media and data storage media do not include connections, carrier waves, signals, or other transitory media, but are instead directed to non-transitory, tangible storage media. Disk and disc, as used herein, includes compact disc (CD), laser disc, optical disc, digital versatile disc (DVD), floppy disk and Blu-ray disc, where disks usually reproduce data magnetically, while discs reproduce data optically with lasers. Combinations of the above should also be included within the scope of computer-readable media.

[0157] Instructions may be executed by one or more processors, such as one or more digital signal processors (DSPs), general purpose microprocessors, application specific integrated circuits (ASICs), field programmable logic arrays (FPGAs), or other equivalent integrated or discrete logic circuitry. Accordingly, the term “processor,” as used herein may refer to any of the foregoing structures or any other structure suitable for implementation of the techniques described herein. In addition, in some aspects, the functionality described herein may be provided within dedicated hardware and / or software modules configured for encoding and decoding, or incorporated in a combined codec. Also, the techniques could be fully implemented in one or more circuits or logic elements.

[0158] The techniques of this disclosure may be implemented in a wide variety of devices or apparatuses, including a wireless handset, an integrated circuit (IC) or a set of ICs (e.g., a chip set). Various components, modules, or units are described in this disclosure to emphasize functional aspects of devices configured to perform the disclosed techniques, but do not necessarily require realization by different hardware units. Rather, as described above, various units may be combined in a codec hardware unit or provided by a collection of interoperative hardware units, including one or more processors as described above, in conjunction with suitable software and / or firmware.

[0159] Various embodiments of the invention may be implemented at least in part in any conventional computer programming language. For example, some embodiments may be implemented in a procedural programming language (e.g., “C”), or in an object oriented programming language (e.g., “C++”). Other embodiments of the invention may be implemented as a pre-configured, stand-alone hardware element and / or as preprogrammed hardware elements (e.g., application specific integrated circuits, FPGAs, and digital signal processors), or other related components.

[0160] Those skilled in the art should appreciate that such computer instructions can be written in a number of programming languages for use with many computer architectures or operating systems. Furthermore, such instructions may be stored in any memory device, such as semiconductor, magnetic, optical or other memory devices, and may be transmitted using any communications technology, such as optical, infrared, microwave, or other transmission technologies.

[0161] Among other ways, such a computer program product may be distributed as a removable medium with accompanying printed or electronic documentation (e.g., shrink wrapped software), preloaded with a computer system (e.g., on system ROM or fixed disk), or distributed from a server or electronic bulletin board over the network (e.g., the Internet or World Wide Web). In fact, some embodiments may be implemented in a software-as-a-service model (“SAAS”) or cloud computing model. Of course, some embodiments of the invention may be implemented as a combination of both software (e.g., a computer program product) and hardware. Still other embodiments of the invention are implemented as entirely hardware, or entirely software

[0162] While the various systems described above are separate implementations, any of the individual components, mechanisms, or devices, and related features and functionality, within the various system embodiments described in detail above can be incorporated into any of the other system embodiments herein.

[0163] The terms “about” and “substantially,” as used herein, refers to variation that can occur (including in numerical quantity or structure), for example, through typical measuring techniques and equipment, with respect to any quantifiable variable, including, but not limited to, mass, volume, time, distance, wave length, frequency, voltage, current, and electromagnetic field. Further, there is certain inadvertent error and variation in the real world that is likely through differences in the manufacture, source, or precision of the components used to make the various components or carry out the methods and the like. The terms “about” and “substantially” also encompass these variations The term “about” and “substantially” can include any variation of 5% or 10%, or any amount - including any integer - between 0% and 10%. Further, whether or not modified by the term “about” or “substantially,” the claims include equivalents to the quantities or amounts.

[0164] Numeric ranges recited within the specification are inclusive of the numbers defining the range and include each integer within the defined range. Throughout this disclosure, various aspects of this disclosure are presented in a range format. It should be understood that the description in range format is merely for convenience and brevity and should not be construed as an inflexible limitation on the scope of the disclosure. Accordingly, the description of a range should be considered to have specifically disclosed all the possible sub-ranges, fractions, and individual numerical values within that range. For example, description of a range such as from 1 to 6 should be considered to have specifically disclosed sub-ranges such as from 1 to 3, from 1 to 4, from 1 to 5, from 2 to 4, from 2 to 6, from 3 to 6 etc., as well as individual numbers within that range, for example, 1 , 2, 3, 4, 5, and 6, and decimals and fractions, for example, 1.2, 3.8, 1 %, and 4% This applies regardless of the breadth of the range. Although the various embodiments have been described with reference to preferred implementations, persons skilled in the art will recognize that changes may be made in form and detail without departing from the spirit and scope thereof.

[0165] Various examples of the disclosure have been described. Any combination of the described systems, operations, or functions is contemplated These and other examples are within the scope of the following claims.

Claims

ClaimsWhat is claimed is:

1. A method comprising:(i) controlling, by one or more processors, one or more mid-infrared sensors to capture one or more mid-infrared spectral images of at least a portion of a plant;(ii) receiving, by the one or more processors, the one or more mid-infrared spectral images;(iii) analyzing, by the one or more processors, the one or more mid-infrared spectral images using a chemometric model; and(iv) determining, by the one or more processors and based at least in part on the analyzing, one or more characteristics of the portion of the plant.

2. The method of claim 1 , wherein the portion of the plant comprises a pulse.3 The method of claim 2, wherein the pulse remains intact after the one or more midinfrared spectral images are captured.

4. The method of claim 1 , wherein the plant is in a Fabaceae family.

5. The method of claim 1 , wherein the plant comprises one of: a dry pea, a lentil, a chickpea, a casein, a sunflower, an oat, a split pea, a kidney bean, wheat, a black eyed pea, any food, or any biological material.

6. The method of claim 1 , wherein the one or more characteristics comprise any one or more of: a total fatty acid, a total saturated fatty acid, a total unsaturated fatty acid,a total dietary fiber, a total protein, a total resistant starch, a total starch, a total sulfur containing amino acids, and a digestibility factor.

7. The method of claim 1 , wherein the chemometric model comprises a partial least squares regression model.

8. The method of claim 1 , wherein the one or more characteristics comprise a digestibility factor, and wherein analyzing the one or more mid-infrared spectral images comprises:(i) applying, by the one or more processors, the chemometric model to the one or more mid-infrared spectral images to detect one or more patterns of alpha sheet signals within an amide I band of the one or more mid-infrared spectral images; and(ii) estimating, by the one or more processors, the digestibility factor based on the one or more patterns9. The method of claim 1 , further comprising:(i) storing, by the one or more processors, the one or more characteristics of the portion of the plant in a database;(ii) retrieving, by the one or more processors, historical characteristics of plants of a same species as the plant from the database;(iii) generating, by the one or more processors, a graphical user interface comprising graphical indications of at least one of the one or more characteristics and at least one of the historical characteristics; and(iv) outputting, by the one or more processors and to an output device, the graphical user interface.

10. The method of claim 9, wherein the database comprises a cloud database stored on one or more remote servers.11 . The method of claim 1 , wherein the chemometric model includes one or more interpolation filters.

12. The method of claims 1, wherein the chemometric model is trained and / or validated using K-fold cross validation13. The method of claim 12, wherein the K-fold cross validation comprises at least five folds.

14. The method of claim 12, wherein the mid-infrared spectral images are normalized between 0 and 1 prior to analysis.

15. The method of claim 12, further comprising interpolating the spectral data to achieve a spectral resolution of 1.86 cm-1.

16. The method of claim 12, wherein the mid-infrared spectral images are captured at a resolution of 2 cm"1and with a zero-fill factor of 2.

17. A computing device comprising one or more processors configured to:(i) control one or more mid-infrared sensors to capture one or more mid-infrared spectral images of at least a portion of a plant;(ii) receive the one or more mid-infrared spectral images;(iii) analyze the one or more mid-infrared spectral images using a chemometric model; and(iv) determine, based at least in part on the analyzing, one or more characteristics of the portion of the plant.

18. The computing device of claim 17, wherein the one or more characteristics comprise a digestibility factor, and wherein the one or more processors being configured to analyze the one or more mid-infrared spectral images comprises the one or more processors being configured to:(i) apply the chemometric model to the one or more mid-infrared spectral images to detect one or more patterns of alpha sheet signals within an amide I band of the one or more mid-infrared spectral images; and(ii) estimate the digestibility factor based on the one or more patterns.

19. The computing device of claim 17, wherein the one or more processors are further configured to:(i) store the one or more characteristics of the portion of the plant in a database;(ii) retrieve historical characteristics of plants of a same species as the plant from the database;(iii) generate a graphical user interface comprising graphical indications of at least one of the one or more characteristics and at least one of the historical characteristics; and(iv) output, to an output device, the graphical user interface.

20. A non-transitory computer-readable storage medium having stored thereon instructions that, when executed, cause one or more processors of a computing device to:-SO-(i) control one or more mid-infrared sensors to capture one or more mid-infrared spectral images of at least a portion of a plant;(ii) receive the one or more mid-infrared spectral images;(iii) analyze the one or more mid-infrared spectral images using a chemometric model; and(iv) determine, based at least in part on the analyzing, one or more characteristics of the portion of the plant.

Citation Information

Patent Citations

  • Inverse Modeling for Characteristic Prediction from Multi-Spectral and Hyper-Spectral Remote Sensed Datasets

    US20230367272A1