Method for co-designing modeling and monitoring systems
An algorithm addresses the challenges of parameter identifiability in ecophysiological models by optimizing data collection and sensor selection, improving crop trait prediction and management efficiency.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-08-07
- Publication Date
- 2026-04-09
AI Technical Summary
Existing ecophysiological models in crop production face challenges with large numbers of variables and parameters that are difficult or expensive to measure directly or infer indirectly, leading to issues of observability and identifiability, which hinder accurate prediction of crop traits and management decisions.
An algorithm is developed to analyze ecophysiological models and data collection systems, identifying critical parameters for field data collection and suggesting sensor types to train the model, thereby reducing data requirements and improving parameter inference.
The algorithm enhances the accuracy of crop trait predictions and management decisions by determining essential parameters and eliminating non-essential ones, facilitating more efficient breeding programs and crop production management.
Smart Images

Figure US2025041073_09042026_PF_FP_ABST
Abstract
Description
METHOD FOR CO-DESIGNING MODELING AND MONITORING SYSTEMS CROSS-REFERENCE TO RELATED APPLICATIONS
[0001] The current patent application is a non-provisional utility patent application which claims priority benefit of U.S. Provisional Patent Application Serial No. 63 / 680,312 entitled “METHOD FOR CO-DESIGNING MODELING SYSTEMS,” filed August 7, 2024, the entire disclosures of which is incorporated herein by reference. STATEMENT REGARDING FEDERALLY SPONSORED RESEARCH OR DEVELOPMENT
[0002] This invention was made with government support under Contract No. 1826820 awarded by the National Science Foundation. The government has certain rights in the invention. BACKGROUND OF THE INVENTION
[0003] Agricultural crop production is failing to increase fast enough to achieve the levels that will be needed by 2050 for humanity to avoid major food security disruption. This would be true if population growth were the only problem, but there are increasingly adverse compounding factors including climate change, declining water resources, diets shifting toward meat, and competing crop use demands (e.g. bioenergy). Potential ways to ameliorate food shortages include (1) accelerating the genetic gain rates for desirable traits in breeding programs and (2) employing improved in-field crop production management methods such as precision agriculture. Neither strategy is feasible without better behavioral prediction algorithms for novel genetic lines under innovative production management within non-analog environments. Accurate, genetically- informed predictions will enable (1) efficient crossing decisions by breeders, and (2) farmers to estimate the likely outcomes (and, therefore, net profitability) of alternative cultural options (variety selection, tillage methods, fertilizer rates and timing, irrigation programs, pest management, etc.) when making choices for their environmentally unique fields. Such predictive algorithms will effectively compose solutions to the genotype-to-phenotype (G2P) problem, among the very highest priorities in applied biology.
[0004] Two important modeling-based approaches to crop phenotype prediction are, in decreasing order of age, are (1) quantitative genetics, QG; and (2) ecophysiological crop modeling,ECM. The first uses algebraic methods to relate temporal interval endpoint traits to genomic features, most often DNA markers. A strength is that the independent variables are fundamental genotype descriptors. A weakness is that it has difficulty reproducing interaction effects that operate over continuous time. In contrast, ECMs are grounded in differential equations that interweave the mechanistic physiological and environmental continuous time processes that determine a great many plant traits. This strength is the exact complement to the QG weakness just mentioned. ECM weaknesses are that (1) their biotic properties only indirectly link to genetics via large sets of constants that are presumed to be determined by underlying genomic features even though (2) these putative genotype-specific parameters (GSPs) can, in fact, be contaminated with larger or smaller degrees of environmental dependency. Certain models include a statistical significance test for environmental dependence of ECM parameters. Potential GSPs include the lower and upper temperatures supporting development and / or growth, the optimal rate and temperature for vernalization, the maximum achievable leaf area index, and – for complex models – many more. Also needed are the starting values (called “initial conditions”) of the model’s time- dependent “state variables”.
[0005] ECMs and QG have many common and / or complementary features and a small community has expended significant efforts to synergize their respective strengths and weaknesses. The dominant line of attack – under consideration for almost 30 years – has been to use QG models to predict GSP values from DNA or RNA sequence data. A very informative resource is U.S. Patent No.11,985,930, incorporated by reference herein in its entirety. The basic notion is that, given genotype data, one uses a trained QG model to estimate the corresponding GSPs. Those, in turn, are employed by the ECM to predict the line’s breeding value or other phenotypes within either a breeding program’s target population of environments (TPE) or the highly granular environments of a precision agriculture practitioner.
[0006] Unfortunately, a problem with ECM models share is that they often have large numbers of variables and parameters that can be difficult and / or unduly expensive to measure directly or even to infer indirectly from more easily collected data. In general, this problem is referred to either as a lack observability or, when the parameters are the focus of interest, identifiability. Observability / identifiability pertains when it is possible to uniquely infer the values of all the variables / parameters from a given set of observables. Note that if one views parametersas being state variables whose temporal rates of change are identically zero, then the identifiability question is a subset of observability assessment.
[0007] Thus, there is a need for an improved means for ascertaining crop models and monitoring systems. This background discussion is intended to provide information related to the present invention which is not necessarily prior art. SUMMARY OF THE INVENTION
[0008] Embodiments of the current invention address one or more of the above-mentioned problems and provide a distinct advance in the art of system models, data acquisition system configuration, and training system models using improved data acquisition systems.
[0009] System models are tools used to predict time-varying outputs of systems. For example, ecophysiological models are used to predict environmentally-dependent time-varying crop traits. When combined with genetic information ecophysiological models can significantly accelerate genetic gain rates for desirable traits in crop breeding programs by making crossing decisions more efficient. Similarly, when informed by accurate, variety-specific information ecophysiological models can allow farmers to estimate likely outcomes for their environmentally unique fields when making production choices ranging from planting details to nutrient and water management, etc.
[0010] However, a problem with ecophysiological models is that they can include dozens of numeric constants (called “parameters”) whose values must be known. In particular, these parameters include the genetically-related, variety-specific information mentioned above. In some cases, the number of parameters is so great that their direct measurement is not feasible, and certain parameters must, out of necessity, be inferred. Currently, even the inference of parameters from field data is often impossible because many very different and erroneous parameter values produce identical, optimal goodness-of-fit values, thus preventing identification of the correct values. Analogous errors can occur when goodness-of-fit functions are replaced by measures of predicted economic returns and the parameters to be found are production management actions (e.g., fertilizer and irrigation rates and dates).
[0011] Embodiments of the present invention seek to ameliorate this problem either by (1) determining the critical parameters for which field data must be collected (addressing the identifiability problem with ecophysiological models) and / or (2) suggesting which parameters arenonessential and can, therefore, be eliminated from the ecophysiological model, thereby reducing the data requirements for the simplified version. Embodiments of the invention include an inventive algorithm that can be used to analyze an ecophysiological model in view of an actual or proposed data collection system and desired objective, such as minimized crop yield prediction errors or maximized profit. The algorithm can then suggest the field sampling protocol and / or sensor types needed to properly train the ecophysiological model as based upon the parameters that the algorithm has found to be essential.
[0012] A specific application of embodiments of the present invention is in a crop breeding program. Crop breeding programs generally involve growing a genetically diverse population of training individuals and phenotyping those individuals to generate a phenotype training data set. An association training data set is generated by pairing the phenotype training data set with a genotype training data set comprising genetic information across the genome of each training individual. The association training data set is then used to train a model that aims to predict how specified genotypes will perform in targeted environments. A genetically diverse population of breeding individuals are then genotyped to obtain the genotypic marker data. The trained model and the genotypic marker data from the population of breeding individuals are used to predict the performance of offspring from possible breeding pairs. Based on these predictions, breeding pairs are selected that are likely to generate offspring with improvements in one or more desired traits and then crossed.
[0013] The statistical genetic models used for these purposes in most breeding programs have limitations when making predictions for multiple environments – an area wherein ecophysiological models excel. However, the variety-related ecophysiological model parameters are not natively connected to genotype marker data. To overcome this problem, some have viewed such model parameters as phenotypic traits that can be predicted from marker data via statistical genetic models. When successful, offspring performance will become more accurately predictable and breeding pair selection more efficient. However, this approach is stymied when the correct ecophysiological model parameter values cannot be inferred from the types of field data commonly collected, which is often the case and goes unrecognized by practitioners.
[0014] Accordingly, one embodiment of the invention is a method for improving a system model. The method includes selecting a continuous-time model that includes numerically constant model parameters and time-varying variables; selecting an objective function; selecting from themodel parameters a set of model parameters whose numerically constant values are to be estimated; selecting a data acquisition system to supply data for model operation and for estimation of the numerically constant values of the set of model parameters. The data acquisition system includes sensors operable to collect data associated with at least one model parameter of the set. The method further includes assigning temporary values as the numerically constant values of the set of model parameters; generating, via a processor, synthetic data from the model using the temporary values; determining, via the processor, an unidentifiable model parameter of the set presenting an identifiability issue; and displaying, via a user interface in communication with the processor, the unidentifiable model parameter.
[0015] According to another embodiment of the present invention, a method is provided for overcoming the above described problems in relation to plant cross breeding, thus enabling more efficient selection of breeding pairs in a crop breeding program. The method comprises first selecting an ecophysiological model, a proposed data collection system to support inference of the ecophysiological model parameter values, and an objective representing the desired output of the model. The algorithm of embodiments of the present invention then analyzes this material to determine if the proposed data collection system is adequate to train the model; that is, whether it is possible to infer unambiguous variety-related ecophysiological model parameter values for each member of the genetically diverse population of training individuals.
[0016] The algorithm does so by using the model, itself, to generate error-free synthetic data whose structure mirrors that of the proposed data collection system. The mathematics that ground the algorithm then enable determination of whether multiple sets of putative variety-related ecophysiological model parameters values can reproduce the synthetic data’s goodness-of-fit function values.
[0017] If this happens, the algorithm reveals which features of the model or the proposed data collection system facilitated the failure. In addition, the algorithm can identify non-essential parameters within the ecophysiological model, thereby allowing the model to be simplified such as by eliminating one or more parameters therefrom and, concomitantly, reducing the data collection system requirements. Together, the results enable model and / or data collection system amendments to be proposed and reanalyzed. When repeated iterations of the algorithm converge to a successful result, the outcome will inform those responsible for the crop breeding program asto the types of field measurements, and sampling protocols necessary to infer the required sets of ecophysiological model parameter values.
[0018] Once, with the algorithm’s assistance, a suitable ecophysiological model, data collection system, and goodness-of-fit function combination are in hand, the breeding program would be able to infer unique ecophysiological model parameter values for each member of the genetically diverse population of training individuals. After doing so, the resulting set of values for each variety-related parameter would be viewed as a phenotype training set. As described earlier above, each phenotype training data set would then be paired with the genotype training data set comprising genetic information across the genome of each training individual; with each such pairing comprising an association training data set for one of the variety-related parameters. Each variety-related association training data set is then used to train a statistical genetic model, which can be generically referred to as a parameter model, and which aims to predict the value of the corresponding variety-related parameter based on the genetic information for a specified genotype.
[0019] As described earlier, the crop breeding program assembles a genetically diverse population of breeding individuals that are genotyped. Breeding pairs are selected from this population. The selection procedure first uses the parameter models to predict variety-related parameter values for the offspring genotypes of all breeding pairs under consideration. The ecophysiological model is then used to predict how these offspring would perform in the targeted environments, thus suggesting breeding pairs that are likely to generate offspring with improvements in one or more desired traits. Based on these predictions, the most promising breeding pairs are selected and then crossed.
[0020] Another embodiment of the invention is a method of producing plants in a breeding program. The method includes growing a genetically diverse population of training plants; selecting continuous-time ecophysiological crop models that include numerically constant model parameters and time-varying variables; selecting an objective function related to a desired trait; selecting from the model parameters of the ecophysiological crop models a set of model parameters whose numerically constant values are to be estimated; selecting a data acquisition system to supply data for model operation and for estimation of the numerically constant values of the set of model parameters. The data acquisition system includes sensors operable to collect data associated with at least one model parameter of the set of model parameters. The method further includesassigning temporary values as the numerically constant values of the set of model parameters; generating synthetic data from the ecophysiological crop models using the temporary values; determining an unidentifiable model parameter of the set presenting an identifiability issue; adjusting ecophysiological crop models based on the unidentifiable model parameter of the set; iteratively generating synthetic data, determining unidentifiable model parameters, and adjusting the ecophysiological crop models until there are no unidentifiable model parameters in the set to produce modified ecophysiological crop models; inferring unique ecophysiological model parameter values representing phenotypes for each member of the genetically diverse population to generate phenotype training sets; pairing each phenotype training data set with a genotype training data set comprising genetic information across the genome of the corresponding training plant to form an association training data set; training genetic models using the association training data sets; using the resulting genetic models for the parameters to predict variety-related parameter values for the offspring genotypes of all breeding pairs under consideration; using the ecophysiological model to predict the performance of these offspring in the targeted environments; selecting breeding pairs from the genetically diverse population of breeding individuals based on the offspring’s predicted performance for the one or more desired traits in the targeted environments; and crossing the breeding pairs to generate the offspring.
[0021] Another embodiment is a system for training a continuous-time ecophysiological crop model. The system includes a processor and a data acquisition system. The processor is configured to receive the ecophysiological crop model that includes numerically constant model parameters and time-varying variables; receive a selected objective function; receive a set of model parameters whose numerically constant values are to be estimated; assign temporary values as the numerically constant values of the set of model parameters; generate synthetic data from the ecophysiological crop model using the temporary values; determine an unidentifiable model parameters that present an identifiability issue; and adjust the ecophysiological crop model based on the unidentifiable model parameters to produce a modified ecophysiological crop model.
[0022] The data acquisition system is in communication with the processor and includes sensors configured to collect data associated with the set of model parameters of the modified ecophysiological crop model. In one or more embodiments, the processor is configured to receive the data to train the ecophysiological crop model.
[0023] Another embodiment of the invention is a computer-implemented method of improving a continuous-time system model. The computer-implemented method includes receiving, via a processor, the system model that includes constant model parameters and time- varying variables; receiving, via the processor, a selected objective function; receiving, at the processor, a set of model parameters whose numerically constant values are to be estimated selected from the model parameters; receiving, at the processor, a selected data acquisition system to supply data for model operation and for estimation of the numerically constant values of the set of model parameters. The data acquisition system includes a sensor operable to collect data related to at least one model parameter of the set of model parameters. The computer-implemented method further includes assigning, via the processor, temporary values as the numerically constant values of the set of model parameters; generating, via the processor, synthetic data from the model using the temporary values; determining, via the processor, an unidentifiable model parameter of the set that present an identifiability issue ; and displaying, via a user interface in communication with the processor, the unidentifiable model parameter.
[0024] Embodiments of the invention include applications beyond ecophysiological models to other trainable, differential equation-based system models. For example, in one or more embodiments, the algorithms described herein are used to evaluate, enhance, and design effective data collection systems for various biological models such as gene network models and biochemical pathway models. The algorithms are used to evaluate, enhance, and design effective data collection systems for non-biological system models including chemical and biomanufacturing reactor models. For example, in one or more embodiment, a reactor model representing the particular processes by which the reactor produces its products and is used in conjunction with a control system. The reactor model includes formulas describing how variables such as chemical species-specific reaction rates vary depending on features like reactor temperature, system pH, etc. These formulas contain parameters (e.g., one or more reaction rates under constant reference conditions) whose values are included in the model. The algorithms described herein are operable to identify the types of data required to train the reactor model adequately so that, in operation, it produces accurate information describing the reactor’s current internal condition and determine what online control signals are needed to maintain its optimal operation over a practical time horizon. Based on the algorithm’s output, the reactor systemmodel’s training data is then collected. The data is used to train the reactor model and implemented to help control a reactor via the control system.
[0025] This summary is provided to introduce a selection of concepts in a simplified form that are further described below in the detailed description. This summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter. Other aspects and advantages of the current invention will be apparent from the following detailed description of the embodiments and the accompanying drawing figures. BRIEF DESCRIPTION OF DRAWINGS
[0026] Embodiments of the current invention are described in detail below with reference to the attached drawing figures, wherein:
[0027] FIG. 1 is a flowchart depicting exemplary steps of a method according to an embodiment of the present invention;
[0028] FIG. 2 is an exemplary objective function surface for a particular set of modeled observables;
[0029] FIG.3 depicts exemplary time series plots from a demonstration of the ability of an embodiment of the invention to detect unidentifiability;
[0030] FIGS. 4A and 4B depict exemplary two-dimensional contour plots of objective function surfaces from a demonstration of the ability of an embodiment of the invention to remediate unidentifiability;
[0031] FIG. 5 depicts exemplary Eigen-analyses of the Hessian matrix of an objective function;
[0032] FIGS. 6A and 6B depict a flowchart depicting exemplary steps of a method according to another embodiment of the present invention; and
[0033] FIG. 7 is a block diagram depicting selected components of an exemplary environment in which embodiments of the present invention may be implemented.
[0034] The drawing figures do not limit the current invention to the specific embodiments disclosed and described herein. The drawings are not necessarily to scale, emphasis instead being placed upon clearly illustrating the principles of the invention.DETAILED DESCRIPTION OF THE INVENTION
[0035] The following detailed description of the technology references the accompanying drawings that illustrate specific embodiments in which the technology can be practiced. The embodiments are intended to describe aspects of the technology in sufficient detail to enable those skilled in the art to practice the technology. Other embodiments can be utilized and changes can be made without departing from the scope of the current invention. The following detailed description is, therefore, not to be taken in a limiting sense. The scope of the current invention is defined only by the appended claims, along with the full scope of equivalents to which such claims are entitled.
[0036] The focus of embodiment of the present invention is on identifiability because ECM parameter uniqueness is a sine qua non when attempting to link GSPs to genetics via QG models. Put simply, if the parameter estimates do not match the values of the actual physiological traits they quantitate, then there is no way to train a QG model to relate them to actual or prospective genotypes.
[0037] In crop production management decision-making contexts, potential management actions can be rendered parametric by quantitating them as, for instance, by the parcel-specific times and rates of water or chemical applications, seeding rates, planting and harvest times etc. In variable-rate precision agriculture many of these controlled inputs are, like GSPs, essentially continuous.
[0038] From the perspective of farmer confidence, it is a clear desideratum that neighbors using the same varieties and management software in similar environments should get similar recommendations. Therefore, it is important that management recommendations (dates, rates, etc.) also be identifiable.
[0039] With occasional, harmless deviations, there are two major categories of unidentifiability and each of those are subdivide into two subcategories. Practical identifiability exists if it is possible to obtain parameter estimates with adequately narrow confidence limits given uncertainties like the proper model form and / or data collection issues. Structural identifiability, on the other hand, results from the model’s mathematical form, its (for ECMs, environmental) input functions, and the observations used for fitting. It does not exist when different sets of parameter values produce identical model outputs and, when absent, must be corrected before any issues of practical identifiability can be addressed. Subdividing these two categories are local vs. globalidentifiability. The former asks if a point estimate is unique within some parameter space neighborhood that contains it. The second pertains when that estimate is unique within the entire parameter space.
[0040] Turning to FIG.1, the flowchart depicts the steps of an exemplary method 100 of designing a system model and monitoring system. In some alternative implementations, the functions noted in the various blocks may occur out of the order depicted in FIG.1. For example, two blocks shown in succession in FIG.1 may in fact be executed substantially concurrently, or the blocks may sometimes be executed in the reverse order depending upon the functionality involved. In addition, some steps may be optional.
[0041] The steps of the method 100 may be performed with and / or aided by a computing device through the utilization of processors, transceivers, hardware, software, firmware, or combinations thereof. However, some of such actions may be distributed differently among such devices or other devices without departing from the spirit of the present invention. Control of the system may also be partially implemented with computer programs stored on one or more non- transient computer-readable medium(s). The computer-readable medium(s) may include one or more executable programs stored thereon, wherein the program(s) instruct one or more processing elements to perform all or certain of the steps outlined herein. The program(s) stored on the computer-readable medium(s) may instruct processing element(s) to perform additional, fewer, or alternative actions, including those discussed elsewhere herein.
[0042] Referring to step 102, a system model is selected. In one or more embodiments, this step may be implemented by receiving, via one or more processing element, the system model. The system model may be a continuous-time model that includes numerically constant model parameters and time-varying variables. In one or more embodiments, the system model is an ecophysiological crop model. However, the system model may be any type of model for controlling a system and whose outputs are twice-differentiable with respect to the parameters without departing from the scope of the present invention. For example, the system model may be a non-biological system model, a chemical reactor model, a biomanufacturing model, or the like. In one or more embodiments, the reactor model is for controlling operations of a reactor process. Embodiments of the algorithm disclosed herein are operable to analyze the system model to determine if a proposed data collection system is adequate to train the model. For example, in one or more embodiments, the method 100 determines whether it is possible to infer unambiguousvariety-related ecophysiological model parameter values for each member of the genetically diverse population of training individuals.
[0043] Referring to step 104, a set of model parameters whose numerically constant values are to be estimated are selected from the model parameters of the system model. This step may include receiving, at the one or more processing element, the set of model parameters.
[0044] Referring to step 106, an objective function whose outputs are twice-differentiable with respect to the parameters is selected. In one or more embodiments, the objective function may be received, via the processing element. The objective function represents the desired output of the model. For example, in one or more embodiments, the objective function is associated with an economic outcome. The objective function may additionally or alternatively be used by a bi- objective optimization algorithm that determines the model parameters of the set having optimized identifiability and optimized cost of implementing the one or more sensor combinations needed to achieve that identifiability.
[0045] Referring to step 108, a data acquisition system is selected. The data acquisition is selected to supply data for model operation and for estimation of the numerically constant values of the set of model parameters. The data acquisition system includes one or more sensors operable to collect data associated with at least one model parameter of the set of model parameters. In one or more embodiments (as discussed elsewhere herein), the proposed data collection system is selected to support inference of the ecophysiological model parameter values. In one or more embodiments, this step includes receiving, at the processing element, the selected data acquisition system.
[0046] Referring to step 110, temporary values are assigned as the putatively optimal numerically constant values of the set of model parameters. In what follows, these temporary values are denoted by θ. In one or more embodiments, this step includes determining, via the processing element, an optimum point in the objective function. This step may include assigning, via the processing element, estimated / temporary values as the numerically constant values of the set of model parameters. The temporary values may be provided by way of a library of feasible parameters, such as biologically feasible parameters associated with the ecophysiological crop model stored on a memory element of the processing element and / or received via a communication element thereof (such as from a cloud storage location) and / or via a user interface of the processing element. In one or more embodiments, there are at least three ways to select a putative optimal atwhich to apply the test according to embodiments of the present invention: (a) published / temporary values can be used if one has access to the other data (i.e., u, x0, and κ, as defined below) needed to compute g (defined below) and, thus, H (defined below), as discussed below in reference to step 114. A second method, (b) used in the example below, is to set (each) θ to a plausible (e.g., a published) value, compute error-free synthetic values for y (t), and use them to calculate g and H, as discussed below in reference to step 114. This works because θ is, verifiably, optimal for such data. Finally, (c) embodiments of the invention use an appropriate optimizer to estimate a θ from data hypothesized to be suitable and then apply method (b) to assess identifiability, as discussed below in reference to step 114.
[0047] Referring to step 112, synthetic data is generated from the model using the temporary values. In one or more embodiments, the synthetic data is generated using the processing element. It is quite common in modeling contexts for model-generated synthetic data to be additionally perturbed by random values intended to simulate experimental error. In contrast, embodiments of this invention use unperturbed, error-free data so as to detect structural identifiability issues, if present. In one or more embodiments, the processing element comprises a processor configured to generate the synthetic data using serial computing. Alternatively, the processing element may comprise a graphics processing unit, and the synthetic data is generated using parallel computing. However, other methods of generating the data may be used without departing form the scope of the present invention.
[0048] Referring to step 114, one or more model parameters of the set presenting an identifiability issue is determined. Using the synthetic data from step 112, the rate of change of a gradient of the objective function with respect to each of the model parameters of the set at the optimum point selected in step 110 is then determined. This step may include determining, via the processing element, the rate of change of the gradient of the objective function with respect to each of the model parameters of the set at the optimum point. The model parameters for which the objective function’s gradient has a rate of change at or around zero may be designated an unidentifiable parameter. Additionally or alternatively, a user-defined threshold may be implemented for determining unidentifiability of the model parameters. In one or more embodiments, this step includes identifying, via the processing element, the model parameters of the set having the rate of change of the objective function’s gradient below the threshold.
[0049] For example, in one or more embodiments, the starting point for dynamic model observability and identifiability analysis uses the following equations:
[0050] where t is time, x contains ηxmodel state variables, x0holds the corresponding initial conditions, y is comprised of ηy model outputs that are observable, and θ is comprised of ηθ model parameters whose identifiability are to be analyzed. The vector u comprises ηu external model inputs. In control theory contexts, the elements of u are called the uncontrolled inputs.
[0051] The f function contains the differential equations that make up the model. At any given time point it calculates the rates of change (i.e. time derivatives) of all the state variables. Although not shown explicitly in Equation [2], some initial conditions might be unknown and, if so, will be elements of θ. In any particular design, the h function outputs predictions of the observed variables.
[0052] Let U be the scalar objective function quantitating a desirability measure relevant in whatever might be the given context and add the following equation to the system:
[0053] of observation times. In embodiments involving pest / disease / weed control settings, for example, f includes the dynamics of the interfering species. The vector κ contains real and known, fixed values such as (1) observational data (biological, chemical, and / or physical) or, for economics, (2) current prices; insurance, interest, and tax rates, etc. Being so, the κ elements are demonstrably not dependent on the model parameters, θ.
[0054] FIG. 2 depicts an exemplary objective function surface for a particular set of modeled observables, y ; sampling times, t ; and observations (plus other constants), κ ; as plotted over the parameter space , θ . A property of the gradient of a scalar function (often exploited in optimization) is that its negative (series of left-most arrows in FIG. 2) points in the steepestdownhill direction and that its magnitude (arrow length) is the function’s overall rate of change. So, at an optimum (any minimum in FIG.2), the gradient is the ηθ - dimensional zero vector.
[0055] Embodiments of the invention test for the presence of structural unidentifiability. The test asks if, at an optimum (taken herein, without loss ofto be a minimum), there exists a parameter space direction (right-most dark arrow and vector s) in which the gradient remains zero. If option (c) is used in step 110, then given the ever-expanding availability of skillful, population-based global optimization algorithms, it is not unreasonable to assume that if multiple low points exist within θ their presence will at least be hinted at. These will supply candidate locales for applying the test, with even one failure showing the need for an improved model / monitoring system design.
[0056] The rate of change of the gradient of U in direction s is given by Hs where H is the Hessian matrix of U with respect to the elements of θ evaluated at any point of interest. Embodiments of the invention may use at least three ways to calculate the partial derivatives comprising H: (a) symbolically, or via (b) automatic differentiation, or (c) finite differencing differentiation.
[0057] The test asks if, at an optimum in FIG. 2, there is a nonzero s for which Hs = 0? That is, can one be at an optimum, change some parameters, and yet still be at an optimum. From elementary linear algebra this can only happen if det(H), the determinant of H, is zero or, equivalently, one or more of the eigenvalues of H is zero. At a minimum (possibly having a flat neighborhood), Hessians are positive semidefinite, which means their eigenvalues are nonnegative. The number of zero eigenvalues equals the number of independent parameter space directions from a given optimum along which unidentifiability exists. One can characterize situations where one or more eigenvalues are very close to zero (where “very close” entails a judgement call) as being “nearly unidentifiable” and therefore falling below the user-defined threshold. Within the particular set of models, goodness-of-fit measures, and probabilistic assumptions employed in the statistical theory of optimal experimental design, E-optimality (“E” is for “eigenvalue”) seeks schemes that maximize the smallest Hessian eigenvalue. Another method, D-optimality, seeks designs that maximize det(H) .
[0058] D-optimality has a geometric interpretation when described via a 2D example. In the mathematical field of differential geometry, the determinant of the Hessian evaluated at a point is called the Gaussian curvature. A property of this quantity is that shape changes that do not alterwithin-surface distances do not change the Gaussian curvature. For example, a flat sheet of paper has a Gaussian curvature of zero and so does the trough shape in FIG.2 obtained by lifting the two opposed sides of a flat sheet. And, clearly, both such shapes contain regions of unidentifiability. On the other hand, any surface reconfiguration that cannot be achieved without localized stretching or wrinkling will change det(H). An example would be converting a flat sheet into a bowl with a single point at its bottom. For bowls, the Gaussian curvature is positive with sharper curvatures being associated with larger values. And, of course, sharper curvatures equate to a more readily findable optimum.
[0059] These methods are not mutually exclusive. For example, whenever optimization is used – method (c) discussed in reference to step 110 – to obtain a candidate θ, it would be prudent to assume that that the found optimum is only just a good approximation of one. So, as per method (b) discussed in reference to step 110, embodiments of the invention use that θ to generate synthetic data –step 112 – and run the test with them. For example, to evaluate a production decision-making tool embodiments of the invention (1) assume plausible values for the biological components of θ, (2) perform an optimization to find the management actions that maximize the objective function (often called utility by economists) after placing realistic economic values within κ, (3) generate error-free synthetic data using a θ (discussed in reference to step 112) that combines the original biological parameter values with the just-found management actions, and (4) run the identifiability test on the Hessian of U computed from those values.
[0060] Turning back to FIG. 1 and referring to step 116, the one or more unidentifiable model parameter displayed via a user interface in communication with the processing element. This step may include displaying the model parameters in association with their corresponding eigenvalues of the Hessian matrix.
[0061] Referring to step 118, the system model is adjusted based on the one or more parameters determined to present identifiability issues. This step may include removing one or more unidentifiable model parameters of the set from the model to produce a modified system model. Additionally or alternatively, this step may include selecting from the model parameters an additional model parameter that is not in the set, adjusting a type of the one or more of the sensors of the data acquisition system, adding an additional sensor, or capturing, via the data acquisition system, data associated with one or more of the unidentifiable model parameters of the set, and estimating the numerically constant values of the set based the data. As indicated by thefeedback loop, steps 112, 114, and 118 (and optionally step 116) may be repeated until there are no unidentifiable model parameters in the set. Or in other words, these steps may be repeated until the rate of change of the gradient of the objective function with respect to each of the model parameters is above the user-defined threshold.
[0062] By incorporating embodiments of the present invention, the concepts of D- optimality and E-optimality can be utilized in the of system model to improve identifiability by seeking to increase the smallest eigenvalues of H and / or the Gaussian curvature, det(H) evaluated at one or more candidate optima. Embodiments of the invention may include the following strategies for doing so: (a) Altering the sampling plan without changing the types of data collected, (b) inserting known observational methods into the data collection pipeline, (c) developing new algorithms to extract observations from existing pipelines, and (d) innovating novel sensors for the data acquisition system.
[0063] If (b), (c), and (d) measure variables that the model does not already predict then the model is modified to do so by changing h and, possibly, f. Examples of (a) and (b) are given below. An example of (c) would be an improved method of, for example, processing of UAV hyperspectral imagery. Examples of (d) may include, for example, electromagnetic frequency- based sensors, or radar-based approaches.
[0064] As discussed above, embodiments of the invention include applying these strategies manually using iterative trials, each followed by a retest to assess progress. At each step, the eigenvectors corresponding to (nearly) zero eigenvalues can provide suggestions. Each such eigenvector will have at least one nonzero component, with the rest being either zero or nonzero. Local moves paralleling the eigenvector change parameter values corresponding to the nonzero components. However, such moves do not change the gradient from its value of 0. That is, the new point is also an optimum for g. If the eigenvector has only one nonzero value, then the model must be insensitive to the corresponding parameter. If the eigenvector has more than one nonzero element, the corresponding parameter value changes must compensate for each other, leaving the model outputs unaltered. On the other hand, even large moves do not change the parameters having eigenvector components of zero. Therefore, those parameters retain their values from the optimum; i.e., they are not contributing to unidentifiability in that specific parameter space direction.
[0065] Suppose there was a parameter that was practical albeit – perhaps – expensive to measure experimentally. Now consider what would happen to the Hessian if a model was replacedby one wherein that parameter was measured; that is, where its numeric value appeared instead as an element in κ. The new Hessian would have no row or column for that parameter but all its retained elements would have exactly the same values as before. This is because the only change made to any of the equations being differentiated would have been to replace the measured parameter’s symbol with its explicit numeric value. Examining the Gaussian curvature or eigenvalues and vectors of this smaller sized and trivially calculated Hessian will provide immediate information on the usefulness of directly measuring that particular parameter.
[0066] According to embodiments of the invention, an optimal monitoring system design tool divides the model parameters into two sets: (a) ones for which practical (i.e., sufficiently accurate measuring methods existed or that could become measurable with novel but imaginable sensors and (b) ones lacking practical measurement methods. Embodiments would then assign a measurement cost to each of the parameters in set (a). And using an appropriate bi-objective optimization algorithm, seek to find a monitoring system design that (1) maximizes Gaussian curvature and (2) minimizes the total cost of the parameters in set (a) that the solution measures rather than infers.
[0067] One or more embodiments include a method that ignores minimizing the total cost of the parameters. Such embodiments determines which among the remaining (initially all) parameters, would – if measured instead of inferred – maximally increase the Gaussian curvature. While awaiting the addition of a suitable bi-objective optimization algorithm, the outputs of any curvature-only method could at least be manually screened for low-cost options.
[0068] Referring to step 120, once the unidentifiable parameters have been removed from the system model, the model may be used to represent the system and / or facilitate its control.
[0069] The method 100 may include additional, less, or alternate steps and / or device(s), including those discussed elsewhere herein. For example, the method 100 may be applied to ecophysiological crop models for training corresponding genetic models. The method 100 may include growing a training plant; using the temporary values of the ecophysiological crop model to generate a phenotype training set; pairing the phenotype training data set with a genotype training data set comprising genetic information across the genome of the training plant using an association training data set; and training a genetic model using the genotype training data set.
[0070] Embodiments of the invention may be extended to other trainable, differential equation-based system models without departing from the scope of the present invention. Forexample, in one or more embodiments, the algorithms are used to evaluate, enhance, and design effective data collection systems for various biological models such as gene network models and biochemical pathway models. The algorithms are used to evaluate, enhance, and design effective data collection systems for non-biological system models including chemical and biomanufacturing reactor models. For example, in one or more embodiment, a reactor model representing the particular processes by which the reactor produces its products and is used in conjunction with a control system. The reactor model includes formulas describing how variables such as chemical species-specific reaction rates vary depending on features like reactor temperature, system pH, etc. These formulas contain parameters (e.g., one or more reaction rates under constant reference conditions) whose values are included in the model. The algorithms described herein are operable to identify the types of data required to train the reactor model adequately so that, in operation, it produces accurate information describing the reactor’s current internal condition and determine what online control signals are needed to maintain its optimal operation over a practical time horizon. Based on the algorithm’s output, the reactor system model’s training data is then collected. The data is used to train the reactor model and implemented to help control a reactor via the control system. EXAMPLE
[0071] As an illustration of both the unidentifiability test and its remediation, an embodiment of the invention was applied to the model of Poudel et al., “A hierarchical Bayesian approach to dynamic ordinary differential equations modeling for repeated measures data on wheat growth,” Field Crops Research, 283:108549 (2022). A goodness-of-fit function was used that was explicitly intended for time series data. For this a variant of the Nash–Sutcliffe Efficiency coefficient used in hydrology was constructed. Specifically, the function was configured to always give values in the [0,1] interval, with zero being a perfect fit. A simple derivation leads to a Reversed Normalized Nash Sutcliffe Efficiency (RNNSE) measure, denoted by URNNSE. This was used in identifiability analysis.
[0072] FIG.3 depicts time series plots that, between them, illustrate both the presence of unidentifiability and the strategies for its removal. The original model had three state variables (accumulated thermal time; leaf area index, LAI; and biomass) and seven parameters, two of which are relevant to the examples herein (physiological maturity, TTM; and the timing of peak LAI,TTL; both measured in thermal time). Each four-plot column corresponds to one winter wheat planting (among 33) whose simulation was driven by Kansas Mesonet weather data. Because these plots are made using the assumed “true” parameter values and because the synthetic data is error- free, each sample value dot falls exactly on the corresponding predicted model curve.
[0073] The Hessian analysis of URNNSE as determined from the top three plots in the left column gave log10 ^det (H)^ = −5.0273, which is to say the URNNSE response surface is very flat. Following strategies for unidentifiability removal, an additional observable (PhysMat) was added to the model, namely a color change known to happen at thermal time TTM and the number of sampling times was changed. (Note that, consistent with the unmodified model, PhysMat was not observed in the left column of plots, hence the corresponding absence of dots.)
[0074] Consonant with the response surface flatness, the smallest two eigenvalues for the left column of plots are 9.528×10-7and 3.278×10-9. In sharp contrast, the smallest two eigenvalues for the right column of plots are 1.047×10-5and 3.451×10-7and the log of the Gaussian curvature increased to log10 ^det (H)^ = −0.4044 , an identifiability improvement of ca. 42000×. FIG. 4A depicts contour plots whose horizontal and vertical axes are, respectively, scaled multiples of the unit parameter space eigenvectors for the smallest and second-smallest eigenvalues of the above sampling schemes centered on the “true” values. Note that, as shown by the grayscale bars, the vertical values (URNNSE) are in powers of ten.
[0075] The width / height of the X / Y axes was chosen based on the eigenvectors to include prescribed ranges for each parameter. Those ranges were -3 °C to +3 °C for the degree-day base temperature, and ±40% of the “true” value for the remaining parameters. The ranges were arbitrary although the 40% limits do guarantee that the parameters to which they were applied are constrained to have the same sign as the “true” values. Of course, such ranges could be made to reflect, perhaps Bayesian, prior knowledge. As is evident, the right contour plot is magnified relative to the one on the left, indeed, considerably so for the Y axis. Moreover, the range of X values for which 10-3≤ URNNSE ≤ 10-4is much smaller, despite the fact that fewer samples were collected in the right column of time series. Examination of the X and Y eigenvectors for both contour plots reveals that each unit of X is a 0.83 to 0.99 day-°C offset from the “true” TTM value of 1950 days-°C to reach physiological maturity. And each unit of Y has the same offset range from the “true” TTL value of 950 days-°C. Thus, the power of the extra model observable not onlydrastically improves estimates of TTM, but also estimates for TTL, while requiring fewer samples to do so.
[0076] An earlier statement noted that “directly measuring” a parameter’s value is generally more resource- demanding than “inferring” it by fitting the model to observations – assuming, of course, that identifiability pertains. The above paragraphs, however, demonstrate that observing one variable can help improve estimates of another, even when both are marginally identifiable to begin with. This suggests an approach to optimizing the design of a monitoring system.
[0077] A method that ignores minimizing the total cost of the parameters was implemented. It asks, “Which among the remaining (initially all) parameters, would – if measured instead of inferred – maximally increase the Gaussian curvature?” When applied twice to the model whose sampling timeseries plots are in the right column of FIG.3, the two contour plots of FIG.4B resulted. For better graphic resolution the plots are zoomed in from ±40% parameter range limits to ±15% (left) to ±5% (right). The steepness of the URNNSEsurfaces are shown by their multiple order of magnitude changes in short X and Y distances. While awaiting the addition of a suitable bi-objective optimization algorithm, the outputs of any curvature-only method (such as the one used here) could at least be manually screened for low-cost options.
[0078] The first application determined that the best parameter to determine by a focused experiment is TTM. The second application found that, given a known TTM, the best next parameter to directly measure is TTL. The identifiability improvements were substantial – directly measuring TTM increased log10^det (H)^ to 5.8888 and further adding a targeted experimental determination of TTL raises the Gaussian curvature logarithm to 11.0266. This is an identifiability improvement of 16 orders of magnitude compared to the original monitoring system design, which only recorded thermal time, LAI, and biomass (the left contour plot in the first pair).
[0079] These analytic results are in complete and gratifying agreement with the crop modeling community’s long- held belief that phenology parameters (TTM and TTL in this model) are the most important ones to know. Thus, we have demonstrated a working prototype mathematical algorithmic method for co-designing ecophysiological / genomic modeling / monitoring systems to enhance food security.
[0080] FIG. 5 depicts Eigen-analyses of the Hessian matrix for each scenario corresponding to FIGS. 4A and 4B, including the eigenvalues for the x-axis (λx) and y-axis (λy) and the eigenvector components for each parameter being estimated.
[0081] The flow chart of FIGS.6A and 6B depict the steps of an exemplary method 600 of designing an ecophysiological crop model and monitoring system and producing plants in a breeding program. In some alternative implementations, the functions noted in the various blocks may occur out of the order depicted in FIGS. 6A and 6B. For example, two blocks shown in succession in FIGS.6A and 6B may in fact be executed substantially concurrently, or the blocks may sometimes be executed in the reverse order depending upon the functionality involved. In addition, some steps may be optional.
[0082] The steps of the method 600 may be performed and / or aided by a computing device through the utilization of processors, transceivers, hardware, software, firmware, or combinations thereof. However, some of such actions may be distributed differently among such devices or other devices without departing from the spirit of the present invention. Control of the system may also be partially implemented with computer programs stored on one or more non-transient computer-readable medium(s). The computer-readable medium(s) may include one or more executable programs stored thereon, wherein the program(s) instruct one or more processing elements to perform all or certain of the steps outlined herein. The program(s) stored on the computer-readable medium(s) may instruct processing element(s) to perform additional, fewer, or alternative actions, including those discussed elsewhere herein.
[0083] Referring to step 602 depicted in FIG. 6A, a genetically diverse population of training plants are grown. Referring to step 604, a continuous-time ecophysiological crop model is selected. In one or more embodiments the ecophysiological crop model includes numerically constant model parameters, time-varying variables, and external model inputs. Referring to step 606, an objective function related to one or more desired trait of the training plants is selected. Referring to step 608, model parameters of the ecophysiological crop model are selected to form a set of model parameters whose numerically constant values are to be estimated. Individual members of the population of training plants correspond to different sets of values of the selected model parameters. Referring to step 610, a data acquisition system to supply data for model operation and for estimation of the numerically constant values of the set of model parameters is selected. As discussed above, the data acquisition system may include one or more sensorscorresponding to some model parameter of the set of model parameters. The sensors may be any sensor known in the art for detecting various characteristics of the plants, such as, for example, sensors operable to detect leaf area index of the plants, plant height, or the like.
[0084] Referring to step 612, temporary values are assigned as the numerically constant values of the set of model parameters, and referring to step 614, synthetic data is generated using the ecophysiological crop model with the temporary values. To facilitate robustness, embodiments may perform, in sequence, steps 612, 616, 618, and if necessary, repeats of this sequence using multiple groups of temporary values, with each group corresponding to a different member of the population of training plants. When multiple groups of temporary values are utilized, embodiments may assign the sequences of steps corresponding to different groups of training values to different processors for concurrent evaluation.
[0085] Referring to step 616, one or more unidentifiable model parameters of the set are determined as presenting an identifiability issue. This step 616 may include implementing the algorithms discussed elsewhere herein, such as determining an optimum point in the objective function for each mode and then determining a rate of change of a gradient of the objective functions with respect to each of the model parameters of the set at the optimum points. The unidentifiable parameters are identified as the model parameters having a rate of change of the gradient below a threshold. When multiple groups of temporary values are used, embodiments of the invention include identifying parameters as unidentifiable when the rate of change of the gradient for any group falls below a threshold.
[0086] Referring to step 618, the ecophysiological crop model is adjusted based on the unidentifiable model parameters of the set. This may include determining one or more types of sensors of the data acquisition system for tracking one or more types of data based at least in part on the modified ecophysiological crop models of the breeding pairs and installing the types of sensors in the data acquisition system. Additionally or alternatively, this step 618 may include removing the unidentifiable model parameters of the set from the model to produce a modified system model. Additionally or alternatively, this step 618 may include selecting from the model parameters one or more additional model parameters that are not in the set. As indicated by the feedback loop, the synthetic data and gradient calculations can be repeated with each modification iteratively until there are no unidentifiable model parameters in the set to produce modified ecophysiological crop models for the training plants.
[0087] Turning to FIG. 6B in reference to step 620, unique ecophysiological model parameter values representing phenotypes are inferred for each member of the genetically diverse population to generate phenotype training sets. Referring to step 622, each parameter-specific phenotype training data set is paired with a genotype training data set comprising genetic information across the genome of the corresponding training plant using an association training data set. Referring to step 624, parameter genetic models are then trained for each parameter using the association training data sets. Referring to step 626, the genetic modes for each parameter are used to predict the parameter values for the offspring of all breeding pairs under consideration. Referring to step 628, the ecophysiological model is used to predict how offspring having these parameter values would perform with respect to traits that are deemed desirable when grown in the breeding program’s targeted population of environments. Referring to step 630, the resulting predictions are used to select the breeding pairs whose crossings are likely to generate offspring with improvements in one or more desired traits. Referring to step 632, the breeding pairs are crossed in order to generate the offspring.
[0088] The method 600 may include additional, less, or alternate steps and / or device(s), including those discussed elsewhere herein. For example, the method 600 may include growing the offspring with the one or more desired traits. Additionally, the method 600 may include determining one or more production management actions based at least in part on the external model inputs of the modified ecophysiological crop models of the breeding pairs. This step may include producing the production management actions while growing the training plants and / or the offspring. The production management actions may include adjusting a fertilizer application rate, an irrigation application rate, a fertilizer application condition, or an irrigation application condition. The conditions may include the timing applying irrigation and / or fertilization. Additionally or alternatively, the production management actions may be determined based at least in part on data collected by the one or more types of sensors of the data acquisition system in accordance with their modified ecophysiological crop models.
[0089] Turning to FIG. 7, an exemplary system 10 in which embodiments of the present invention may be implemented is depicted. The system 10 is operable to improve and / or train a system model, aid in the design of a data acquisition and / or monitoring system 12, and / or help control operations of a system by providing relevant management action recommendations to anexternal device 32 via wired or wireless communication, such as communication over a network 30.
[0090] The data acquisition system 12 may include one or more sensors 16, 18 for capturing data relevant to model operation and for estimation of the numerically constant values of the set of model parameters developed and / or determined at a computing device 20. The data acquisition system 12 may be installed on site, such as in a reactor, in a field for crop production, or the like.
[0091] The network 30 may be any communication means for facilitating wired and / or wireless connectivity between the devices, which may include but are not limited to computers, smartphones, tablets, servers, and other electronic devices. The network 30 may include architecture for supporting data transmission across various communication protocols and may include local area networks (LANs), wide area networks (WANs), and other suitable network configurations. Furthermore, the network 30 may be operatively connected to the Internet, enabling devices to access a wide range of online resources and services. The network 30 may also incorporate cloud computing environments, allowing for scalable data storage, processing capabilities, and seamless integration with cloud-based applications.
[0092] The external device 32 may be a computer, smartphone, tablet, server, or other electronic device. The external device 32 may be for presenting outputs from the computing device 20 and / or data captured by the data acquisition system 12. The external device 32 may additionally or alternatively be used for controlling operations of the system corresponding to the system model, such as a chemical reactor plant, irrigation system, or the like.
[0093] The computing device 20 is configured to improve and / or train a system model, aid in the design of a data acquisition and / or monitoring system 12, and / or help control operations of a system by providing relevant management action recommendations. The computing device 20 may comprise one or more communication elements 22, one or more memory elements 24, a user interface 26, and one or more processing elements 28.
[0094] In one or more embodiments, the processing element 28 is configured to select a continuous-time model that includes numerically constant model parameters and time-varying variables. The model may be selected according to one or more inputs, such as a type of system to be modeled. Additionally or alternatively, the processing element 28 may be configured toreceive the selected model. In one or more embodiments, the system model is an ecophysiological crop model that includes numerically constant model parameters and time-varying variables;
[0095] In one or more embodiments, processing element 28 may be configured to select an objective function. Similar to the system model, the appropriate objective function may be selected based on one or more inputs received, such as a selection of the desired objective function and / or a selection of a desired trait. For example, a user may select an economic objective function associated with the economics of a system that includes parameters associated with costs. The selection may also be related to a desired output, such as a selection of a desired trait of a plant in a cross-breeding application. In one or more embodiments, the objective function may be used by a bi-objective optimization algorithm that determines optimized the model parameters of the set with having optimized identifiability and optimized cost of implementing the one or more sensor combinations needed to achieve that identifiability. The costs may be derived from costs of implementing one or more sensor combinations in the data acquisition system 12.
[0096] The processing element 28 may also be configured to select from the model parameters a set of model parameters whose numerically constant values are to be estimated. The processing element 28 may be configured to pull all parameters associated with the system model and / or receive a set of model parameters based on user inputs. Additionally or alternatively, the set of model parameters may be selected by identifying model parameters of the set that can be measured using one or more sensor combinations of the data acquisition system 12.
[0097] The processing element 28 may also be configured to receive a selection of a data acquisition system to supply data for model operation and for estimation of the numerically constant values of the set of model parameters. The selection may be made based on the system to modelled and / or the data acquisition system 12 in place, including the sensors 16, 18 operable to collect data associated with at least one model parameter of the set of model parameters.
[0098] The processing element 28 is configured to assign temporary values as the numerically constant values of the set of model parameters. The processor 28 may assign the temporary values based on user inputs and / or pull the temporary values from a library associated with the system model. In one or more embodiments, the temporary values are feasible parameter values associated with the model, such as biologically feasible parameters in the case of an ecophysiological crop model.
[0099] The processing element 28 is configured to generate synthetic data from the model using the temporary values. In one or more embodiments, the processing element 28 comprises one or more processor configured to use serial computing to generate the synthetic data. Alternatively, the processing element 28 may comprise a graphics processing unit that uses parallel computing to generate the synthetic data.
[0100] The processing element 28 is configured to determine one or more unidentifiable model parameter of the set presenting an identifiability issue. In one or more embodiments, this may be accomplished by displaying a representation of the objective function, as depicted in FIG. 4, in which a three-dimensional trough indicates that one or more parameters has identifiability issues. The identification may be relative and / or a threshold identifiability metric may be implemented. In one or more embodiments, the processing element 28 may be configured to determine an identifiability issue by finding an optimum point in the objective function, and determining a rate of change of a gradient of the objective function with respect to each of the model parameters of the set at the optimum point. The processing element 28 may then identify each model parameter of the set that has a rate of change of the gradient below a threshold at the optimum point. The processing element 28 may be configured to determine unidentifiability using other algorithms, including those discussed elsewhere herein.
[0101] In one or more embodiments, upon determination of the presence of identifiability issues, the processing element 28 may be configured to display and / or transmit signal representative of the one or more unidentifiable model parameters. The processing element 28 may also be configured to adjust the system model based upon the determination of unidentifiable parameters. This may include removing the one or more unidentifiable model parameters of the set from the model to produce a modified system model, selecting from the model parameters an additional model parameter that is not in the set, adjusting a type of the one or more of the sensors of the data acquisition system, or adding the one or more sensor combinations corresponding to the optimized model parameters to the data acquisition system. The processing element 28 may be configured to adjust the system model to produce a modified system, and repeat the process of generating synthetic data, determining unidentifiability issues, and modification of the system model until the rate of change of the gradient of the objective function with respect to each of the model parameters is above a user-defined threshold.
[0102] In one or more embodiments, the processing element 28 is configured to receive the data from the data acquisition system 12 to train the modified system model. In embodiments where the system model is an ecophysiological crop model, the processing element 28 may be configured to use the temporary values of the ecophysiological crop model to generate a phenotype training set, and pair the phenotype training data set with a genotype training data set comprising genetic information across the genome of the training plant using an association training data set. The processing element may then be configured to train a genetic model using the genotype training data set.
[0103] Additionally, the processing element 28 may be configured determine one or more management actions based at least in part on the external model inputs of the modified system model. The processing element 28 may be configured to transmit a signal representative of the one or more management actions to the external device 32 through the network 30. For example, the management action may include production management actions, such as a fertilizer application rate, an irrigation application rate, a fertilizer application condition , an irrigation application condition, or the like.
[0104] Throughout this specification, references to “one embodiment”, “an embodiment”, or “embodiments” mean that the feature or features being referred to are included in at least one embodiment of the technology. Separate references to “one embodiment”, “an embodiment”, or “embodiments” in this description do not necessarily refer to the same embodiment and are also not mutually exclusive unless so stated and / or except as will be readily apparent to those skilled in the art from the description. For example, a feature, structure, act, etc. described in one embodiment may also be included in other embodiments, but is not necessarily included. Thus, the current invention can include a variety of combinations and / or integrations of the embodiments described herein.
[0105] Although the present application sets forth a detailed description of numerous different embodiments, it should be understood that the legal scope of the description is defined by the words of the claims set forth at the end of this patent and equivalents. The detailed description is to be construed as exemplary only and does not describe every possible embodiment since describing every possible embodiment would be impractical. Numerous alternative embodiments may be implemented, using either current technology or technology developed after the filing date of this patent, which would still fall within the scope of the claims.
[0106] As used herein, the phrase “and / or,” when used in a list of two or more items, means that any one of the listed items can be employed by itself or any combination of two or more of the listed items can be employed. For example, if a composition is described as containing or excluding components A, B, and / or C, the composition can contain or exclude A alone; B alone; C alone; A and B in combination; A and C in combination; B and C in combination; or A, B, and C in combination.
[0107] The present description also uses numerical ranges to quantify certain parameters relating to various embodiments of the invention. It should be understood that when numerical ranges are provided, such ranges are to be construed as providing literal support for claim limitations that only recite the lower value of the range as well as claim limitations that only recite the upper value of the range. For example, a disclosed numerical range of about 10 to about 100 provides literal support for a claim reciting “greater than or equal to about 10” (with no upper bounds) and a claim reciting “less than or equal to about 100” (with no lower bounds).
[0108] Furthermore, unless otherwise specified, any directional references (e.g., upper, lower, above, below, etc.) are used herein solely for the sake of convenience and should be understood only in relation to each other. For instance, a component might in practice be oriented such that faces referred to as “upper” and “lower” are sideways, angled, inverted, etc. relative to the chosen frame of reference.
[0109] Throughout this specification, plural instances may implement components, operations, or structures described as a single instance. Although individual operations of one or more methods are illustrated and described as separate operations, one or more of the individual operations may be performed concurrently, and nothing requires that the operations be performed in the order illustrated. Structures and functionality presented as separate components in example configurations may be implemented as a combined structure or component. Similarly, structures and functionality presented as a single component may be implemented as separate components. These and other variations, modifications, additions, and improvements fall within the scope of the subject matter herein.
[0110] Certain embodiments are described herein as including logic or a number of routines, subroutines, applications, or instructions. These may constitute either software (e.g., code embodied on a machine-readable medium or in a transmission signal) or hardware. In hardware, the routines, etc., are tangible units capable of performing certain operations and maybe configured or arranged in a certain manner. In example embodiments, one or more computer systems (e.g., a standalone, client or server computer system) or one or more hardware modules of a computer system (e.g., a processor or a group of processors) may be configured by software (e.g., an application or application portion) as computer hardware that operates to perform certain operations as described herein.
[0111] In various embodiments, computer hardware, such as a processing element, may be implemented as special purpose or as general purpose. For example, the processing element may comprise dedicated circuitry or logic that is permanently configured, such as an application- specific integrated circuit (ASIC), or indefinitely configured, such as an FPGA, to perform certain operations. The processing element may also comprise programmable logic or circuitry (e.g., as encompassed within a general-purpose processor or other programmable processor) that is temporarily configured by software to perform certain operations. It will be appreciated that the decision to implement the processing element as special purpose, in dedicated and permanently configured circuitry, or as general purpose (e.g., configured by software) may be driven by cost and time considerations.
[0112] Accordingly, the term “processing element” or equivalents should be understood to encompass a tangible entity, be that an entity that is physically constructed, permanently configured (e.g., hardwired), or temporarily configured (e.g., programmed) to operate in a certain manner or to perform certain operations described herein. Considering embodiments in which the processing element is temporarily configured (e.g., programmed), each of the processing elements need not be configured or instantiated at any one instance in time. For example, where the processing element comprises a general-purpose processor configured using software, the general- purpose processor may be configured as respective different processing elements at different times. Software may accordingly configure the processing element to constitute a particular hardware configuration at one instance of time and to constitute a different hardware configuration at a different instance of time.
[0113] The processing element may include processors, microprocessors (single-core and multi-core), microcontrollers, DSPs, field-programmable gate arrays (FPGAs), analog and / or digital application-specific integrated circuits (ASICs), or the like, or combinations thereof. The processing element may generally execute, process, or run instructions, code, code segments, software, firmware, programs, applications, apps, processes, services, daemons, or the like. Theprocessing element may also include hardware components such as finite-state machines, sequential and combinational logic, and other electronic circuits that can perform the functions necessary for the operation of the current invention. The processing element may be in communication with the other electronic components through serial or parallel links that include address busses, data busses, control lines, and the like.
[0114] Computer hardware components, such as communication elements, memory elements, processing elements, and the like, may provide information to, and receive information from, other computer hardware components. Accordingly, the described computer hardware components may be regarded as being communicatively coupled. Where multiple of such computer hardware components exist contemporaneously, communications may be achieved through signal transmission (e.g., over appropriate circuits and buses) that connect the computer hardware components. In embodiments in which multiple computer hardware components are configured or instantiated at different times, communications between such computer hardware components may be achieved, for example, through the storage and retrieval of information in memory structures to which the multiple computer hardware components have access. For example, one computer hardware component may perform an operation and store the output of that operation in a memory device to which it is communicatively coupled. A further computer hardware component may then, at a later time, access the memory device to retrieve and process the stored output. Computer hardware components may also initiate communications with input or output devices, and may operate on a resource (e.g., a collection of information).
[0115] The memory device or element may include data storage components, such as read- only memory (ROM), programmable ROM, erasable programmable ROM, random-access memory (RAM) such as static RAM (SRAM) or dynamic RAM (DRAM), cache memory, hard disks, floppy disks, optical disks, flash memory, thumb drives, universal serial bus (USB) drives, or the like, or combinations thereof. In some embodiments, the memory element may be embedded in, or packaged in the same package as, the processing element. The memory element may include, or may constitute, a “computer-readable medium”. The memory element may store the instructions, code, code segments, software, firmware, programs, applications, apps, services, daemons, or the like that are executed by the processing element.
[0116] The communication element may generally allow communication with systems and / or external devices. The communication element may include signal or data transmitting andreceiving circuits, such as antennas, amplifiers, filters, mixers, oscillators, digital signal processors (DSPs), and the like. The communication element may establish communication wirelessly by utilizing RF signals and / or data that comply with communication standards such as cellular 2G, 3G, 4G, 5G, or LTE, WiFi, WiMAX, Bluetooth®, BLE, or combinations thereof. The communication element may be in communication with the processing element and the memory element.
[0117] The user interface generally allows the user to utilize inputs and outputs to interact with the device and is in communication with the one or more processing element. Inputs may include buttons, pushbuttons, knobs, jog dials, shuttle dials, directional pads, multidirectional buttons, switches, keypads, keyboards, mice, joysticks, microphones, or the like, or combinations thereof. The outputs of the present invention may include a display and / or any number of additional outputs, such as audio speakers, lights, dials, meters, printers, or the like, or combinations thereof, without departing from the scope of the present invention.
[0118] The various operations of example methods described herein may be performed, at least partially, by one or more processing elements that are temporarily configured (e.g., by software) or permanently configured to perform the relevant operations. Whether temporarily or permanently configured, such processing elements may constitute processing element- implemented modules that operate to perform one or more operations or functions. The modules referred to herein may, in some example embodiments, comprise processing element-implemented modules.
[0119] Similarly, the methods or routines described herein may be at least partially processing element-implemented. For example, at least some of the operations of a method may be performed by one or more processing elements or processing element-implemented hardware modules. The performance of certain of the operations may be distributed among the one or more processing elements, not only residing within a single machine, but deployed across a number of machines. In some example embodiments, the processing elements may be located in a single location (e.g., within a home environment, an office environment or as a server farm), while in other embodiments the processing elements may be distributed across a number of locations.
[0120] Unless specifically stated otherwise, discussions herein using words such as “processing,” “computing,” “calculating,” “determining,” “presenting,” “displaying,” or the like may refer to actions or processes of a machine (e.g., a computer with a processing element andother computer hardware components) that manipulates or transforms data represented as physical (e.g., electronic, magnetic, or optical) quantities within one or more memories (e.g., volatile memory, non-volatile memory, or a combination thereof), registers, or other machine components that receive, store, transmit, or display information.
[0121] As used herein, the terms “comprises,” “comprising,” “includes,” “including,” “has,” “having” or any other variation thereof, are intended to cover a non-exclusive inclusion. For example, a process, method, article, or apparatus that comprises a list of elements is not necessarily limited to only those elements but may include other elements not expressly listed or inherent to such process, method, article, or apparatus.
[0122] The patent claims at the end of this patent application are not intended to be construed under 35 U.S.C. § 112(f) unless traditional means-plus-function language is expressly recited, such as “means for” or “step for” language being explicitly recited in the claim(s).
[0123] Although the technology has been described with reference to the embodiments illustrated in the attached drawing figures, it is noted that equivalents may be employed and substitutions made herein without departing from the scope of the technology as recited in the claims.
[0124] Having thus described various embodiments of the technology, what is claimed as new and desired to be protected by Letters Patent includes the following:
Claims
CLAIMS 1. A method of improving a system model, the method comprising: (a) selecting a continuous-time model that includes numerically constant model parameters and time-varying variables; (b) selecting an objective function; (c) selecting from the model parameters a set of model parameters whose numerically constant values are to be estimated; (d) selecting a data acquisition system to supply data for model operation and for estimation of the numerically constant values of the set of model parameters, the data acquisition system including one or more sensors operable to collect data associated with at least one model parameter of the set of model parameters; (e) assigning temporary values as the numerically constant values of the set of model parameters; (f) generating, via one or more processor, synthetic data from the model using the temporary values; (g) determining, via the one or more processor, one or more unidentifiable model parameter of the set presenting an identifiability issue; and (h) displaying, via a user interface in communication with the one or more processor, the one or more unidentifiable model parameter.
2. The method of claim 1, further comprising: (i) performing at least one of the following steps: removing the one or more unidentifiable model parameter of the set from the model to produce a modified system model, selecting from the model parameters an additional model parameter that is not in the set, adjusting a type of the one or more of the sensors of the data acquisition system, or capturing, via the data acquisition system, data associated with one or more additional model parameters of the set, estimating the numerically constant values of the set based the data.
3. The method of claim 2, further comprising repeating steps (f), (g), and (i) until there are no unidentifiable model parameters in the set.
4. The method of claim 1, further comprising: identifying model parameters of the set that can be associated with one or more sensor combinations; implementing a bi-objective optimization algorithm that determines optimized model parameters of the set with optimized identifiability and optimized cost of implementing the one or more sensor combinations; and adding the one or more sensor combinations corresponding to the optimized model parameters to the data acquisition system.
5. The method of claim 1, wherein: step (e) comprises: determining, via the one or more processor, an optimum point in the objective function; and step (g) comprises:determining, via the one or more processor, a rate of change of a gradient of the objective function with respect to each of the model parameters of the set at the optimum point; and identifying, via the one or more processor, the one or more of the model parameters of the set having a rate of change of the gradient below a threshold.
6. The method of claim 5, further comprising (i) adjusting the system model to produce a modified system, and repeating steps (f), (g), and (i) until the rate of change of the gradient of the objective function with respect to each of the model parameters is above the threshold.
7. The method of claim 1, wherein the system model is an ecophysiological crop model, further comprising: growing a training plant; using the temporary values of the ecophysiological crop model to generate a phenotype training set; pairing the phenotype training data set with a genotype training data set comprising genetic information across the genome of the training plant using an association training data set; and training a genetic model using the association training data set.
8. The method of claim 1, wherein the system model is an ecophysiological crop model, and the objective function is associated with an economic outcome.
9. A method of producing plants in a breeding program, the method comprising: (a) growing a genetically diverse population of training plants; (b) selecting a continuous-time ecophysiological crop model that includes numerically constant model parameters and time-varying variables; (c) selecting an objective function related to one or more desired trait; (d) selecting from the model parameters of the ecophysiological crop model a set of model parameters whose numerically constant values are to be estimated; (e) selecting a data acquisition system to supply data for model operation and for estimation of the numerically constant values of the set of model parameters, the data acquisition system including one or more sensors operable to collect data associated with at least one model parameter of the set of model parameters; (f) assigning temporary values as the numerically constant values of the set of model parameters; (g) generating synthetic data from the ecophysiological crop models using the temporary values; (h) determining one or more unidentifiable model parameter of the set presenting an identifiability issue; (i) adjusting the ecophysiological crop model based on the one or more unidentifiable model parameter of the set; (j) repeating steps (g), (h), and (i) until there are no unidentifiable model parameters in the set to produce a modified ecophysiological crop model; (k) inferring unique ecophysiological model parameter values representing phenotypes for each member of the genetically diverse population to generate phenotype training sets; (l) pairing each phenotype training data set with a genotype training data set comprising genetic information across the genome of the corresponding training plant using an association training data set; (m) using the association training data sets to train genetic models that predict parameter values from genotype data; (n) selecting breeding pairs from the genetically diverse population of breeding individuals by using their genotypes and the genetic models to predict their offspring parametervalues and the ecophysiological models to predict one or more desired traits for the offspring in the targeted environments; and (o) crossing the breeding pairs to generate the offspring.
10. The method of claim 9, further comprising growing the offspring with the one or more desired traits.
11. The method of claim 10, wherein the modified ecophysiological crop model of the breeding pairs includes external model inputs, further comprising determining one or more production management actions based at least in part on the external model inputs of the modified ecophysiological crop model of the breeding pairs.
12. The method of claim 11, wherein the production management actions include at least one of a fertilizer application rate, an irrigation application rate, a fertilizer application condition, or an irrigation application condition.
13. The method of claim 12, wherein step (i) comprises determining one or more types of sensors of the data acquisition system for tracking one or more types of data based at least in part on the modified ecophysiological crop model of the breeding pairs.
14. The method of claim 13, further comprising installing the one or more types of sensors in the data acquisition system.
15. The method of claim 13, wherein determining the one or more production management actions is based at least in part on data collected by the one or more types of sensors.
16. The method of claim 9, wherein step (i) comprises removing the one or more unidentifiable model parameter of the set from the model to produce a modified ecophysiological crop model.
17. The method of claim 9, wherein step (i) comprises selecting from the model parameters an additional model parameter that is not in the set.
18. The method of claim 9, wherein: step (f) comprises: determining an optimum point in the objective function; and step (h) comprises: determining a rate of change of a gradient of the objective function with respect to each of the model parameters of the set at the optimum point; and identifying the one or more of the model parameters of the set having a rate of change of the gradient below a threshold.
19. A system for training a continuous-time ecophysiological crop model, the system comprising: one or more processor configured to: (a) receive the ecophysiological crop model that includes numerically constant model parameters and time-varying variables; (b) receive a selected objective function; (c) receive a set of model parameters whose numerically constant values are to be estimated; (d) assign temporary values as the numerically constant values of the set of model parameters; (e) generate synthetic data from the ecophysiological crop model using the temporary values; (f) determine one or more unidentifiable model parameters that present an identifiability issue; and (g) adjust the ecophysiological crop model based on the unidentifiable model parameters to produce a modified ecophysiological crop model; and a data acquisition system in communication with the one or more processor and comprising one or more sensors configured to collect data associated with the set of model parameters of the modified ecophysiological crop model.
20. The system of claim 19, wherein the one or more processor is configured to receive the data from the data acquisition system to train the modified ecophysiological crop model.
21. The system of claim 19, wherein the processor is configured to transmit a signal representative of an identification of the unidentifiable model parameters.
22. The system of claim 19, wherein step (g) comprises performing at least one of the following: remove the one or more unidentifiable model parameter of the set from the ecophysiological crop model, or select from the model parameters an additional model parameter that is not in the set.
23. The system of claim 19, wherein the one or more processor is configured to repeat steps (e), (f), and (g) until there are no unidentifiable parameters in the set.
24. The system of claim 19, wherein the modified ecophysiological crop model includes external model inputs, and the one or more processor is configured to: determine one or more production management actions based at least in part on the external model inputs of the modified ecophysiological crop models, and transmit a signal representative of the one or more production management actions.
25. A computer-implemented method of improving a continuous-time system model, the computer-implemented method comprising: (a) receiving, via one or more processor, the system model that includes constant model parameters and time-varying variables; (b) receiving, at the one or more processor, a selected objective function; (c) receiving, at the one or more processor, a set of model parameters whose numerically constant values are to be estimated selected from the model parameters; (d) receiving, at the one or more processor, a selected data acquisition system to supply data for model operation and for estimation of the numerically constant values of the set of model parameters, the data acquisition system including one or more sensors operable to collect data associated with at least one model parameter of the set of model parameters; (e) assigning, via the one or more processor, temporary values as the numerically constant values of the set of model parameters; (f) generating, via the one or more processor, synthetic data from the model using the temporary values; (g) determining, via the one or more processor, one or more unidentifiable model parameter of the set that present an identifiability issue; and (h) displaying, via a user interface in communication with the one or more processor, the one or more unidentifiable model parameters.
26. The method of claim 25, wherein: step (e) comprises: determining, via the one or more processor, an optimum point in the objective function; step (g) comprises: determining, via the one or more processor, a rate of change of a gradient of the objective function with respect to each of the model parameters of the set at the optimum point; and identifying the one or more of the model parameters of the set having a rate of change of a gradient below a threshold.
27. The method of claim 26, further comprising: (i) adjusting, via the one or more processor, the system model, and repeating, via the one or more processor, steps (f) through (i) until the rate of change of the gradient of all model parameters are above the threshold.
28. The method of claim 25, wherein step (f) is performed via the one or more processor using serial computing.
29. The method of claim 25, wherein the one or more processor includes a graphics processing unit, and step (f) is performed via the graphics processing unit using parallel computing.