Method to build a digital twin
Patent Information
- Authority / Receiving Office
- EP · EP
- Patent Type
- Applications
- Current Assignee / Owner
- CIRILLO FABIO
- Filing Date
- 2024-06-14
- Publication Date
- 2026-05-06
AI Technical Summary
Current digital twin technologies do not provide a holistic approach to simulating the growth of living systems from seed to harvest-ready state, lacking multi-spectrum detections and holistic analytics, and are not applied to living organisms like plants in a controlled environment for biosynthetic pathway elucidation and optimal cultivation conditions.
A method involving multiple living systems grown under controlled conditions with data collection on various parameters over time, generating a digital twin for simulating growth and determining optimized conditions using statistical analysis, machine learning, and multi-variant simulations for biosynthetic pathway elucidation and digital biomarker identification.
Enables in-depth knowledge of living systems for optimized growth conditions, biosynthetic pathway elucidation, and reduced resource consumption by simulating growth under different conditions in a controlled environment, enhancing precision in plant cultivation and health management.
Smart Images

Figure EP2024066539_02012025_PF_FP_ABST
Abstract
Description
[0001] METHOD TO BUILD A DIGITAL TWIN
[0002] The current invention relates to a method to build a digital twin.
[0003] The evolution of the digital twin technology allowed the next industrial revolution, known as Industry 4.0. Physical objects, transferred as CAD or similar designs into the digital world, paired with movement, material, and logic data allow for simulative use and in-silico experimentation without the necessity to construct those physical objects. As the application of such objects is widely used in engineering, aviation, robotics, buildings / constructions, and agriculture, little to none of this approach is used holistically on living plants.
[0004] W02020201567 and WO2020224779A1 describe a cell-based digital twin system to predict cell responses on tested compounds and thus, isolate a living cell from the organism. Other digital twin solutions used in medicine are described in CN115005981 that use a human digital twin approach for surgical path planning, again isolating a specific aspect from the entire living organism.
[0005] WO2022236064A2 does describe a digital twin system for an industrial environment with sensor readings and data but does not apply the digital twin to a living organism, such as a plant. The publications by Westermann, L. (Bringen Baume mehr Ertrag, 15. Juni 2023) and Verdouw, Cor (twins in smart farming, agricultural systems, Vol. 189, 2021 ) describe the use of a digital twin for precision agriculture using prediction of crop, crop time and other states of the plant in the environment, i.e. in an uncontrolled environment.
[0006] All of these systems won’t allow a holistic approach to the digital twin model creation of a living system from seed, seedling or egg to the harvest-ready living system, with the subsequent derivation of in-depth knowledge on system reactions and feedback to further utilize such a digital twin model for biosynthetic pathway elucidation, digital biomarker, optimal conditions of breeding, cultivation, and / or health sustaining of a living organism, such as a plant. Moreover, none of these systems uses multi-spectrum detections, e.g. UV, VIS, NIR, IR, and / or holistic analytics such as multi-spectrum paired with gene sequencing (genomics), transcriptions (transcriptom ics) or other omics.
[0007] The object of the current invention is the provision of a method to build a digital twin for a living system to simulate the growth of said living system under different conditions in a controlled environment, including hermetically closed environments. The current invention achieves this object with a method with the features of claim 1 .
[0008] A method according to the invention is provided for the building of a digital twin for a living system and for the use of said digital twin. Said living system is an organism, and could preferably be a plant, an algae, a water-borne animal, a fungus, and / or a worm. Said organisms can be cultivated in a small space with high efficiency. Humans and / or larger animals are thus not in the focus of the current method.
[0009] The method comprises at least the following steps:
[0010] Step A: Provision of a number of multiple living systems of the same type, i.e. genus and / or species;
[0011] Said number might be at least 4 living systems of the same genus, such as a specific genus of a plant but optionally of different species. More preferably said number might be at least 16 living systems of the same genus. Said living systems might be clones from the same plant or different species.
[0012] Step B: Growing the living systems of the same genus under different controlled growing conditions;
[0013] Said controlled conditions might be the simulation of a desert, rainforest, normal climate, or the like. Said controlled growing conditions might comprise a different temperature, sunlight, wavelength, type of earth, amount of moisture, such as rain or fog, and further conditions, to simulate natural habitats.
[0014] Step C: Determination of one or more parameters of the living system that is or are subjected to a change over the time period of the growth; wherein multiple data sets are determined and stored, each data set has said parameter or parameters and a time stamp of the time of determination;
[0015] Said parameter might be the movement of a water-borne animal, the height of a plant over a period of growth, the mass growth of the plant over a time period of growth, the growth of substances, such as minerals and / or vitamins and / or active ingredients, over a period of growth and the degradation of substances over a period of growth.
[0016] Said substances might be inside the living system, such as a plant, such as an ingredient or the degradation of nutrients outside or inside the living system. Each parameter is stored with a time stamp. Thus, the parameter is stored together with the specific time within the time period of growth to allow for time-series analysis.
[0017] Step D: Generation of a digital twin for said living system based on said multiple data set representing different growth or health stages of said living system.
[0018] The digital twin represents the digital twin at a specific point of growth, preferably at the start of the growth. Such a digital twin might be a seed or an egg or the like.
[0019] The method is preferably applied under controlled environmental conditions and most preferably applied for a Greenhouse or other forms of an indoor farm, including hermetically closed environments.
[0020] The method may further comprise a Step E representing the determination of optimized growing and / or health-sustaining conditions with respect to at least one parameter of the living system that is subjected to a change over the period of the growth, cultivation and / or health state, said parameter is selected between group 1 :
[0021] I the concentration of an excrement and / or exhaust;
[0022] II the concentration of an ingredient generated by or digested by the living system; and / or
[0023] III the suppression of an ingredient generated by the living system and group 2:
[0024] IV the speed of growth;
[0025] V the amount of biomass;
[0026] VI the movement of a living system;
[0027] VII the amount and / or the mass and / or the form of descendants or offspring, such as fruits and / or blossoms.
[0028] Thus at least two parameters are used, one being selected from group 1 and the second being selected from group 2.
[0029] A speed of growth for the length of the plant often demands a different condition than the mass of fruits in a plant. Different stadiums of growth may also demand other conditions. A seed often prefers warm humid conditions whereas an adult plant may be subjected to fouling under these conditions. By using a digital twin, different growth conditions, such as climate conditions, can be adopted to the different stadiums of growth.
[0030] Said Parameters that are determined in step C might be one or more of the following: i speed of growth of the living system; ii color of the living system; iii form of specific parts of the living system; iv form, color, number and / or size of fruits or other offspring of the living system; v concentration of ingredients that are formed during the growth vi intake of nutrients of the living systems vii movement of the living systems viii excrements and / or exhausts of the living systems
[0031] Said data set may further comprise the growing condition of each living system, especially the growing condition at the time of the determination of said parameter. Each data set has thus at least three values.
[0032] Said data set may further comprise one or multiple health conditions of each living system, especially the lack and / or absence or malnutrition conditions, immobility, death, fouling, and alike.
[0033] The generation of the digital twin for said living system might be further based on said multiple data sets representing different growth stages of said one or more living systems of a different type. Said living system of claim 1 might be rape seed and the living system of a different type might be a mustard plant. Thus, the type of the plant is different although both plants are from the same plant family. Thus, it is likely that rape seed and mustard will grow similar under the same growth conditions. This way, more data is generated making the reaction of the digital twin at different climate conditions more reliable.
[0034] The generation of the digital twin for said living system may be further based on said multiple data set representing at least one health condition of said one or more living systems of a different type, that can be either one or a mixture of mobility, water content, nutrition concentration, and movement of specific parts of the living system.
[0035] Said inventive method may comprise the use, relying on, and / or referring to at least one potentially affecting the process of the growth, development, sprouting, movement, differentiation, propagation, cultivation, multiplication, and / or reproduction of such a living system. The method may further comprise at least one input factor and at least one response, feedback, measurable outcome, and / or vital signal.
[0036] The generation of the digital twin in Step D and / or determination of optimized growing condition in Step E may comprise an analysis, information gathering, logic application, trends, and / or coincidences to elucidate potential and / or specific biomarkers and / or digital biomarkers.
[0037] The analysis in Step D and / or E may be performed by means of statistical analysis, neuronal networks, machine learning, artificial analysis, regression analysis, and / or trending.
[0038] Said generation and / or determination according to Step D and / or Step E may further comprise prospective simulation, retrospective simulation, extrapolation, interpolation, targeting, minimizing, maximizing, ranging, and / or multi-variant simulation for one or multiple desired factors and / or results in the living system.
[0039] The invention is further explained below, relating to Fig. 1 .
[0040] The present invention describes a method of at least one step to create a digital twin of a living system, such as a plant, and derive in-depth knowledge of the digital twin to optimize subsequent procedures such as optimized growth conditions, biosynthetic pathway elucidation, digital biomarkers and / or reduced resource consumption.
[0041] With the use of conventional systems for plant growth, a said plant of a known or unknown genus or species is treated during one or many of its growth stages to observe reactions and feedback of the plant. This one or many treatments as a sequence or in parallel can be done on at least one of said plants and / or multiplied to many of said plant and / or plants of similar or different genus and / or species. The treatment itself is conducted in a controlled environment.
[0042] Such a treatment and / or treatments can generate data, which can be used to form a digital twin. Such data can be used alone or together with literature or data from other sources, such as a previous treatment and / or treatments to form a digital twin of the said plant. Further, the data can come exclusively from either source or be combined with any source to form a digital twin of said plant. As an additional and / or single source of data, the genus of said plants can be used as well whereas not only the genus by means of identifier, such as a genus name, identification number, or any other descriptor can be used but also, either alone or in addition, the genetic code as whole or in part. Such genetic code can come from different sources, such as literature, lab results, and / or other references.
[0043] Therefore, the basis to form the digital twin is data in the form of responses of a plant to a treatment, identifier, genetic code, or literature of plant behavior under given and / or specific conditions.
[0044] This first step of forming a digital twin by input data can be done either quantitatively or qualitatively whereas absolute and / or relative data become important, but not necessary, to refine the digital twin. Absolute data in the context of this invention are measurable data, such as weight, length, width, color, absorbance, pH, humidity, and many more. Relative data in the context of this invention are data relating to treatments and their plant feedback and / or responses, which again can, but not necessarily have to, be absolute data. Such relative data can be, for example, a treatment of higher temperature leading to a color change in the leaves of the plant. Relative data, therefore, can be understood as vectors or pathways for which input leads to an output. Relative data can be global data or local data, depending on the stage of growth the said plant is in. As an example, the vector “wind” can have a minor effect on the plant when the plant has a wooden stem and can easily withstand higher windspeeds of for example 10 m / s whereas a seedling would be kinked and potentially be irreversibly and fatally harmed, thus being a local relative data. An example of global relative data could be that fixation of CO2 is only possible when the light of a specific wavelength is present.
[0045] Such local or global data can also be derived from the genus, identifier, and / or genetic code of said plant. Whereas missing gene sequences would depict global data for the non-existence of a specific gene expression, local data could be the presence of an exon that is only expressed under specific conditions, such as a cascading reaction of signaling molecules.
[0046] Data can further come as continuous data or numerical data and / or as categoric data.
[0047] To form a digital twin of said plant, at least one data input, and an output is needed, forming a treatment. Such input / output pairs can be formed in a specific context, whereas such context can be combined with a linker and / or said vector. With this, a triple of data is formed, often referred to as semantic triples in modern web applications or the W3W approach to the semantic web. As such triples represent the context of local or global relative data but do not necessarily have to be part of it, the digital twin can be enhanced with time-dependent, multi-layer, annihilating, forming, catalyzing, re-enforcing, and / or knock-out feedback loops.
[0048] Vectors can come, in addition or exclusively from physical, biological, or chemical processes and / or from mathematical formulas and algorithms. For example, there is a certain reproduction speed of cells of a specific type that cannot be overcome and sets a limit to the vector. Such limits, either one-sided or two-sided, add to the quality of information of such said vector, adding an additional layer of information, and / or an additional dimension to such a triple.
[0049] Limits, as a term used in this context to describe an additional dimension, can come in different aspects, such as a sequence of events, (e.g. roots must be formed before nutrients can be absorbed via the roots, chromophores must be formed before photosynthesis is possible, CO2 must be present together with light to allow photosynthesis), restriction of resources (e.g. only a certain wavelength can start photosynthesis, depletion of minerals in the water over time), hydrostatic restrictions of water transport in plants, speed of chemical reactions in plants, and / or activation energies of chemical reactions.
[0050] Input data, partially or in whole, is analyzed during the next step. The analysis of input data is done in order to provide a digital twin of the said plant. The analysis step of the present invention needs input factors of at least a type identifier (1 ) and the to- be-studied factors (2) to be varied. The type identifier can be both abstract, by means of a self-assigned identifier, or a very defined identifier, such as a genetic code, e.g. of the said plant to be studied. The main purpose of the identifier in the first step is to allow the analysis to understand that a different or new system, e.g. said plant, is studied. The analysis thus blocks, by means of blocking as in a DoE (Design of Experiment) approach, and described in detail in Taschenbuch der Versuchsplanung, Kleppmann W., 7thed., against another studied plant. Thus, the analysis is systemagnostic and can be used not only for plants but in fact for any living organism, such as prokaryotes, eukaryotes, and / or archaea.
[0051] Blocking in this context can be used in one experiment on the same genus and on different species of that genus in order to run experiments in parallel. As experimental runs, due to the factorial and / or time-dependent nature, would not be economically feasible by means of the number of runs, regular and in literature-known compression algorithms can be used to keep run numbers low, such as fractional factorial designs. The analysis of the input factors can be provided by other statistical means, such as neuronal networks of at least one layer, machine learning algorithms with at least one training set per said plant, and / or artificial intelligence algorithms with at least one training set per said plant, combined in the term machine learning for the sake of reading.
[0052] As living organisms are time-dependent, the analysis has to cover a nearly indefinite number of possible outcomes, which are reduced to computable amounts by means of statistics and time-laps analysis. Thus, the analysis is used to control, by means of a closed-loop, the process by, for example, directing the addition or mixing of the different factors 2 in a combinatorial approach 3. The main difference to regular DoE or combinatorial designs lies in the fact of the time-laps, i.e. a factor could have an influence in a very early stage of the breeding and / or growth of a living organism, but later not or only reduced with no significant effect on the interested response or feedback parameters 4 of the living organism, as described herein as global or local parameters. Therefore, the analysis can make use of a multi-layered DoE 5, a time-sequence DoE, time-dependent machine learning, and / or combinatorial approach to alter the factors 2 and can be paired with machine-learning algorithms 6 to interpret the measured system responses at one- or several-time steps. Specifically, the statistical approaches and / or algorithms are used to analyze the system responses over time to guide and control the process to understand and / or map system feedback (4) to meet a predefined target 7.
[0053] The algorithms 5 and / or 6 are predominantly used to understand parameter effects on the living system from which, with the help of either statistics 8 and / or machinelearning algorithm 9 are used to interpret the data as time-effect and factor-dependent outcomes. These outcomes can be collated into triples and / or triples with vectors.
[0054] These data points, triples and / or triples with vectors can further be compared to biochemical pathways from literature, bench tests, and / or previous responses and / or feedback to determine whether the data and / or signal from the said plant can be aligned with a biomarker and / or digital biomarker. If such a biomarker and / or digital biomarker determination is possible, it can be added for later verification and validation by a specific side-arm of a DoE or other statistical test, either as augmentation of the algorithms 5 and / or 6, or as a new specific test. In the case that a trend, probability, coincidence, or alike is seen for a potential biomarker and / or digital biomarker, which is not yet known and / or fully understood in the literature, the biomarker or digital biomarker can be marked as a potential biomarker and or digital biomarker and added to a virtual biochemical pathway library, and / or triple and / or triple with vector. Such libraries can be verified and / or validated either by statistical experimentation, statistical analysis, or other means of testing including verification by chemical synthesis paired with knock-out strategies and biochemical tests.
[0055] Biosynthetic pathways add to data points, triples and / or triples with vectors for global relative data.
[0056] In the case of statistics 8, the interpretation can be done, for example, by an analysis of variances (ANOVA analysis), which is used in regular factorial DoE approaches for creating a response surface (RSM) to understand maximums and minimums. However, as the living organism is time-dependent, multi-level DoE is needed for which some factors can be clustered in a mixture design whereas others can be clustered in a factorial DoE followed by RSM. This approach is computation-intensive as at each interval of observation, each or some layers of the DoE are to be analyzed per each or some targets 7.
[0057] Once enough data is presented, i.e. at least as many data points as factors studied to vary these factors by at least 2 levels, either as an intermediary result or as the final result, e.g. at harvest for plants, the data can be used to form a digital twin of the studied system. The digital twin can be used for purposes of comparison to other systems, identification of bio-pathways, identification of mutations, phenotypic optimization, and / or optimal growth parameters to meet one or multiple specific targets 7, in the next step.
[0058] The digital twin of a system can then be used to control the process of growth of said plant over time by means of target-to-actual data comparison of analytical, phenotypical, genetic, and / or aspect factors 4. A digital twin can be built per species, and / or per genus allowing pre-set error margins and confidence levels to predict outcomes within a genus. Tolerance analysis can be further applied to define and / or predict a potential specification per factor of said plant.
[0059] Genus-to-genus comparisons can be used to understand bio-pathways for certain ingredients and / or feedback and might allow, depending on the underpinning bio-path- way, to take or omit a bio-pathway to further optimize the growth protocol according to a target 7 of said plant. Further, with known biomarkers and / or digital biomarkers, bio-pathways can be challenged by adding biomarkers as chemicals to the living system to verify and / or validate the pathway as well as or in addition to, force the living system to follow a specific bio-pathway. This can be of importance if the growth cycle of said plant needs to be optimized and / or a certain behavior, feedback, response or biochemical synthesis is needed to be induced in the living system. Of special interest can be the induction of secondary metabolites to increase and / or decrease amounts, compositions and / or formation in specific parts of the living system.
[0060] Previously named targets 7 can be a local / global maximum, a local / global minimum, a certain mixture of ingredients, hormones, aspects, color, dimensions, taste, fragrance, smell, biomass, water content, growth speed, or other chemicals and / or biomarkers, and / or a range of these. Targets can be elaborated as pre-defined targets to reach, phenotypical limits, and / or closed-loop optimization of said living organism.
[0061] The process of the present invention is a two-step approach whereas at least the first step (I) is mandatory to build the digital twin of a living system.
[0062] Step (I) is to vary all potential and / or essential parameters for a given system either as factorial DoE, as mixture DoE, or as a machine-learning algorithm over predefined time laps with predefined time intervals for feedback (4) measurement. The gathered data allow for compiling the digital twin of a species, and / or a genus of a living system. Here, the time laps become decisive of how elaborated the digital twin is by means of statistical data per time interval. The more data is available, the higher is the accuracy and / or shorter time interval for deriving intelligence from the digital twin.
[0063] Step (II) is to use the digital twin to simulate and / or control industrial-scale and / or real-world growth of living systems for different reasons, e.g., production of medicinal herbs with specific ingredient profiles or cultivation of food and feed plants under harsh atmospheric conditions. This second step (II) can be used to simulate in-silico potential outcomes of the said living system when altering the input factors and / or using the digital twin to predict system outcomes. All of these simulations and / or calculations can be supported by either one or many of the calculations, such as accuracy, confidence, tolerance, standard deviations, median, mean, average, quartiles, significance (p-value), fit (F-value), difference in fits, curvature, residuals, studentized residuals, correlation, transformations, effects, power, variance, ANOVA, t-value, standard error, degrees of freedom, coefficient estimates, R-squared, Chi-squared, Cook’s distance, Box-Cox, interaction, perturbation, signal-to-noise, variance inflation, maximum likelihood, restricted maximum likelihood, least significant difference, covariance trace, covariance ratio, influence by DFBETA, minimum, maximum, limit, normal probability, probability, leverage, time, outlier test, and ignored data.
[0064] Step (II) can be augmented, sequentially and / or in parallel, with additional data, which might be gathered over time and / or from different sources, such as additional test results, literature, and / or expert know-how. With this augmentation, the digital twin from Step (I) can be refined, re-calculated, and / or re-adjusted to further meet new targets, increase statistical values described herein, extrapolate and / or interpolate for new insights for responses and / or feedback.
[0065] Step (II) can include the statistical and / or logical calculations and / or comparisons of two or more digital twins for living systems, in order to find similarities and / or differences. This is of special interest to find analogies, potential insights, differences, similarities, and / or identical biomarkers and / or digital biomarkers for biosynthetic pathway elucidation.
[0066] Thus, Step (I) can be understood as an isolated step and / or as a step in a circular and / or looped approach between itself and Step (II) to, over time, expand the dataset of the digital twin without any limit to the loops to a specific digital twin.
[0067] Drawing:
[0068] 1) Data
[0069] 2) Genus 3) Digital Twin
[0070] 4) Biomarker and / or digital biomarker
[0071] 5) Biosynthetic pathway
[0072] 6) Plant optimization / simulation
Claims
Claims1 . Method for building a digital twin for a living system, whereas such a system is an organism, preferably a plant, an algae, a water-borne animal, a fungus, and / or a worm, characterized by at least the following steps:A Provision of a number of multiple living systems of the same type;B Growing the living systems of the same type under different growing conditions;C Determination of one or more parameters of the living system that is or are subjected to a change in over the time period of the growth; wherein multiple data sets are determined and stored, each data set has said parameter or parameters and a time stamp of the time of determination; andD Generation of a digital twin for said living system based on said multiple data set representing different growth stages of said living system.
2. Method according to claim 1 characterized in that the method is applied for controlled environment and further comprises in Step E the determination of optimized growing conditions with respect to at least one parameter of the living system that is subjected to a change over the time period of the growth, said parameter is selected per each group between group 1 :I the concentration of an ingredient generated by or digested by the living system;II the concentration of an excrement and / or exhaust and / or;III the suppression of an ingredient generated by the living system; and group 2:VI the speed of growth;V the amount of biomass;VI the movement of a living system and / or;VII the amount and / or the mass and / or the form of descendants or offspring, such as fruits and / or blossoms.
3. Method according to claim 1 or 2, characterized in that the parameter of step C is selected from the following parameters: i speed of growth of the living system; ii color of the living system; iii form of specific parts of the living system;iv form, color, number and / or size of fruits or other offspring of the living system; v concentration of ingredients in excrements and / or exhausts; vi ability, type and / or direction of movement; vii concentration of ingredients that are formed during the growth viii intake of nutrients of the living systems4. Method according to one of the preceding claims, characterized in that the data set further comprises the growing condition of each living system, especially the growing condition at the time of the determination of said parameter.
5. Method according to one of the preceding claims, characterized in that the generation of the digital twin for said living system is further based on said multiple data set representing different growth stages of said one or more living systems of a different type.
6. Method according to one of the preceding claims, characterized in that the generation of the digital twin for said living system is further based on said multiple data set representing at least one health condition of said one or more living systems of a different type, that can be either one or a mixture of mobility, water content, nutrition concentration, and movement of specific parts of the living system.
7. Method according to one of the preceding claims characterized in that the method comprises the use, relying on, and / or referring to at least one potentially affecting the process of the growth, development, sprouting, differentiation, propagation, cultivation, multiplication, and / or reproduction of such a living system.
8. Method according to one of the preceding claims characterized in that the method comprises at least one input factor and at least one response, feedback, measurable outcome, and / or vital signal.
9. Method according to one of the preceding claims characterized in that Step D and / or Step E further comprises analysis, information gathering, logic application, trends, and / or coincidences to elucidate potential and / or specific biomarkers and / or digital biomarkers.
10. Method according to one of the preceding claims characterized in that the analysis in Step D and / or E is performed by means of statistical analysis, neuronal networks, machine learning, artificial analysis, regression analysis, and / or trending.11 . Method according to one of the preceding claims characterized in that Step D and / or Step E further comprise prospective simulation, retrospective simulation, extrapolation, interpolation, targeting, minimizing, maximizing, ranging, and / or multi-variant simulation.