Digital twinborn growth simulation and optimization method for safflower seedling breeding
By constructing individualized digital twins of safflower plants and combining physical mechanisms with data-driven models, the shortcomings of traditional models in simulating individual differences and environmental interactions are solved. This enables precise growth simulation and optimization of safflower seedlings, generates efficient cultivation control strategies, and supports intelligent propagation of safflower seedlings.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- INST OF TRADITIONAL CHINESE MEDICINE HENAN ACAD OF AGRI SCI
- Filing Date
- 2026-01-15
- Publication Date
- 2026-05-08
AI Technical Summary
Traditional crop growth models are difficult to accurately simulate individual differences and complex environmental interactions, resulting in a lack of scientific basis for optimizing safflower seedling breeding programs. In particular, they are highly sensitive to microenvironmental disturbances in the early developmental stages and cannot reveal implicit interaction patterns.
We construct individualized digital twins of safflower plants that integrate multi-source heterogeneous data, and combine a hybrid modeling mechanism of physical mechanism model and data-driven model to achieve high-fidelity simulation of the entire life cycle of safflower seedlings in complex dynamic environment. We also generate precise cultivation and regulation strategies through multi-objective collaborative optimization.
It enables precise characterization of the individual development trajectory of safflower seedlings, improves the accuracy and timeliness of growth simulation, generates cultivation control instructions that take into account yield, quality and sustainability, and supports the intelligent and refined breeding of safflower seedlings.
Smart Images

Figure CN121997729A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of digital twin and agricultural information technology, specifically relating to a digital twin growth simulation and optimization method for safflower seedling propagation. Background Technology
[0002] With the rapid development of precision agriculture and smart breeding technologies, crop growth simulation plays an increasingly crucial role in seedling propagation, variety selection, and cultivation management. Traditional crop growth models are mostly based on empirical formulas or simplified mechanistic equations, relying on population average parameters to macroscopically describe physiological processes such as photosynthesis, respiration, and nutrient allocation. While these methods possess some predictive ability at the field scale, they struggle to characterize the dynamic response characteristics of individual plants under genotypic differences, microenvironmental fluctuations, and agronomical management interventions. This is particularly true in the seedling propagation of medicinal and economic crops like safflower, where individual plant phenotypic plasticity is high, growth cycles are sensitive, and the effects of temperature, light, water, and fertilizer coupling are significant. Traditional models, lacking a detailed analysis of the genotype-environment-management ternary interaction mechanism, result in systematic deviations between simulation results and actual growth trajectories, making it difficult to support high-precision breeding decisions and individualized propagation regulation.
[0003] Digital twin technology offers a new paradigm for crop growth modeling, its core lying in constructing high-fidelity virtual individuals that map reality to the physical world, interact in real time, and evolve dynamically. For safflower seedling propagation, there is an urgent need to establish digital twins that integrate physiological mechanisms with data-driven approaches to achieve a leap from population averages to precise individual plant measurements. The key to this direction lies in how to preserve the inherent constraints of plant physiological processes while possessing adaptive learning capabilities for complex nonlinear interactions under limited observational data conditions.
[0004] While pure mechanistic models offer interpretability, their fixed parameters and rigid structure make them ill-suited to individual variations and environmental disturbances. Conversely, purely data-driven models (such as deep neural networks), despite their strong fitting capabilities, are prone to overfitting, exhibit poor generalization performance under small sample conditions, and lack physiological constraints, leading to predictions that contradict biological principles. More critically, current methods generally overlook the high sensitivity of safflower seedlings to microenvironmental disturbances during early development, failing to reveal latent interaction patterns (such as the nonlinear enhancement effect of nitrogen response in specific genotypes under low temperature and low light conditions), thus lacking a scientific basis for optimizing breeding programs. Summary of the Invention
[0005] This invention provides a digital twin growth simulation and optimization method for safflower seedling propagation. By constructing an individualized digital twin of safflower plants that integrates multi-source heterogeneous data, and combining a hybrid modeling mechanism of physical mechanism model and data-driven model, it achieves high-fidelity simulation of the entire life cycle of safflower seedlings in complex dynamic environments. Based on the simulation results, it performs multi-objective collaborative optimization to generate precise cultivation control strategies, thereby solving the technical problems of traditional crop growth models relying on empirical formulas and being unable to accurately simulate individual differences and complex environmental interactions.
[0006] This invention provides a digital twin growth simulation and optimization method for safflower seedling propagation, comprising: Acquire multi-dimensional real-time sensing data on individual genotype information, initial phenotypic parameters, and cultivation environment of safflower seedlings; Based on the individual genotype information and initial phenotypic parameters, an individualized three-dimensional geometric morphological skeleton of safflower seedlings was constructed, and its physiological state variables were initialized. The multi-dimensional real-time sensing data and historical environmental time-series data are spatiotemporally aligned and fused to form a unified environment-driven field dataset. A set of mechanistic sub-models was established, which included core physiological processes such as photosynthesis, respiratory metabolism, water transpiration, nutrient absorption and distribution, and organogenesis. Each sub-model was described by a set of differential equations. A data-driven correction module based on a deep temporal neural network is constructed. This module takes the environmental driving field dataset and physiological state variables as input and outputs the real-time compensation amount for the prediction bias of the mechanism sub-model. The output of the mechanism sub-model set is weighted and fused with the compensation amount of the data-driven correction module to generate the updated physiological state variables and three-dimensional geometric morphological parameters of safflower seedlings at the next time step. Based on the updated three-dimensional geometric morphology parameters, the digital twin visualization model of safflower seedlings is dynamically reconstructed, and its internal state database is updated synchronously. A multi-objective optimization problem is set up with the objective functions of maximizing biomass accumulation rate, achieving the threshold of effective medicinal component content, and optimizing water resource utilization efficiency. A genetic algorithm based on non-dominated sorting was used to iteratively optimize four decision variables: irrigation amount, fertilizer ratio, light intensity control range, and temperature and humidity setpoint. The optimized decision variable sequence is transformed into executable cultivation control instructions, which are then sent to the physical planting unit via an IoT execution terminal.
[0007] As one embodiment of the present invention, obtaining the individual genotype information of safflower seedlings specifically includes: performing whole-genome resequencing on safflower seed samples using a high-throughput sequencing platform to obtain a single nucleotide polymorphism site map, and extracting molecular marker combinations related to plant height, number of branches, flowering period, glandular hair density, and key enzyme encoding genes in the hydroxysafflower yellow A synthesis pathway.
[0008] As one embodiment of the present invention, the initial phenotypic parameters include seed weight per thousand seeds, radicle length, cotyledon unfolding angle, and initial moisture content. The multi-dimensional real-time sensing data includes total solar radiation illuminance above the canopy, photosynthetically active radiation flux density, air temperature, relative humidity, carbon dioxide concentration, soil volumetric water content, soil electrical conductivity, and soil temperature.
[0009] As one embodiment of the present invention, the construction of the individualized three-dimensional geometric morphology skeleton of safflower seedlings specifically includes: generating an initial topological structure based on the L-system fractal algorithm, according to the initial branching level, internode length, and leaf tilt angle distribution function; using spherical harmonic functions to parametrically model the leaf contour; and using radial basis function interpolation to map discrete organ measurement point cloud data to a continuous surface model.
[0010] As one embodiment of the present invention, the temporal resolution of the environmental driving field dataset is 10 minutes, the spatial resolution is at the single-plant scale, and its fusion processing includes filling missing data with cubic spline interpolation and removing abnormal mutation points with sliding window midpoint filtering.
[0011] As one embodiment of the present invention, the photosynthesis sub-model in the mechanism sub-model set adopts the Farquhar biochemical model framework, with the inputs being intercellular carbon dioxide concentration, mesophyll conductance, and maximum carboxylation rate, and the output being net photosynthetic rate; the water transpiration sub-model is based on the Penman-Monteith equation and introduces a dynamic feedback term for stomatal conductance; the nutrient absorption sub-model adopts the Michaelis-Menten kinetic equation and couples the root distribution density function with the soil nutrient diffusion coefficient.
[0012] In one embodiment of the present invention, the deep temporal neural network is a bidirectional gated recurrent unit network with 128 hidden layer nodes, an input sequence length of 72 time steps, and an output of the first derivative correction term for each physiological state variable. The network is trained under supervision using historical field observation datasets in the offline stage, and receives real-time data streams and outputs instantaneous compensation quantities in the online stage using a sliding window method.
[0013] As one embodiment of the present invention, the weighted fusion adopts an adaptive weight allocation strategy. The weight coefficients are dynamically adjusted according to the variance of the residual of the mechanism model. When the residual variance exceeds the preset threshold of 0.5, the weight of the data-driven correction module is increased to 0.7; otherwise, it remains at 0.3.
[0014] As one embodiment of the present invention, the objective function of the multi-objective optimization problem is defined as follows: the first objective function is the increase in aboveground dry matter per unit time, the second objective function is that the mass fraction of hydroxysafflower yellow pigment A in the petals is not less than 1.5%, and the third objective function is that the amount of irrigation water consumed per kilogram of dry matter does not exceed 8 liters; the constraints include that the soil moisture content is not less than 60% and not more than 90% of the field capacity, and the average daily temperature is between 15 degrees Celsius and 28 degrees Celsius.
[0015] In one embodiment of the present invention, the population size of the non-dominated sorting genetic algorithm is 200, the crossover probability is 0.9, the mutation probability is 0.1, and the maximum number of generations is 500. The decision variables are encoded as real numbers, the irrigation amount ranges from 0 to 10 liters per day, the nitrogen mass fraction in the nitrogen, phosphorus and potassium fertilizer ratio is between 0.1% and 0.5%, the lower limit of the light intensity control range is not less than 400 micromoles per square meter per second, and the difference between the daytime temperature setpoint and the nighttime temperature setpoint is not less than 6 degrees Celsius.
[0016] As one embodiment of the present invention, the Internet of Things execution terminal includes an electric proportional regulating valve, a Venturi fertilizer applicator, an LED supplemental lighting array, and a wet curtain fan linkage unit. The cultivation control commands include valve opening percentage, fertilizer mother liquor injection rate, supplemental lighting operating current, and fan start / stop cycle duty cycle.
[0017] This invention also provides a digital twin growth simulation and optimization system for safflower seedling propagation, comprising: Individual data acquisition unit is used to acquire multi-dimensional real-time sensing data of individual genotype information, initial phenotypic parameters, and cultivation environment of safflower seedlings; A digital twin construction unit is used to construct an individualized three-dimensional geometric morphological skeleton of safflower seedlings based on the individual's genotype information and initial phenotypic parameters, and to initialize its physiological state variables. The environmental data fusion unit is used to perform spatiotemporal alignment and fusion processing on the multi-dimensional real-time sensing data and historical environmental time-series data to form a unified environmental driving field dataset. The hybrid modeling and computing unit is used to establish a set of mechanistic sub-models that include core physiological processes such as photosynthesis, respiratory metabolism, water transpiration, nutrient absorption and distribution, and organogenesis. It also constructs a data-driven correction module based on a deep temporal neural network, which weights and fuses the outputs of the two to generate updated physiological state variables and three-dimensional geometric morphological parameters of safflower seedlings at the next time step. The visualization rendering unit is used to dynamically reconstruct the digital twin visualization model of safflower seedlings based on the updated three-dimensional geometric morphology parameters, and to update its internal state database simultaneously. The multi-objective optimization decision unit is used to set up a multi-objective optimization problem with the objective functions of maximizing biomass accumulation rate, achieving the threshold of effective medicinal component content, and optimizing water resource utilization efficiency. It uses a genetic algorithm based on non-dominated sorting to iteratively optimize four decision variables: irrigation amount, fertilizer ratio, light intensity control range, and temperature and humidity setpoint. The instruction issuance and execution unit is used to transform the optimized decision variable sequence into executable cultivation control instructions, and then issue them to the physical planting unit through the Internet of Things execution terminal.
[0018] In one embodiment of the present invention, the individual data acquisition unit includes a high-throughput gene sequencer, a laser scanning three-dimensional reconstruction device, a multispectral imager, a micro weather station, and a soil multi-parameter sensor array; the high-throughput gene sequencer has a read length of not less than 150 base pairs, the laser scanning three-dimensional reconstruction device has a spatial resolution of 0.5 mm, and the soil multi-parameter sensor array is deployed to cover a soil layer of 0 to 40 cm in depth, and is arranged in layers at 10 cm intervals in the vertical direction.
[0019] In one embodiment of the present invention, the hybrid modeling computing unit is deployed on an edge computing gateway device, which is equipped with a 16-core central processing unit, 64 gigabytes of memory, a dedicated tensor computing accelerator card, and a real-time Linux kernel operating system with a task scheduling cycle of 100 milliseconds.
[0020] As one embodiment of the present invention, the multi-objective optimization decision unit and the hybrid modeling computing unit exchange data through shared memory. During the optimization iteration process, the fitness evaluation of each generation of individuals calls the simulation interface of the hybrid modeling computing unit. The duration of a single simulation covers the next 72 hours, with a time step of 10 minutes.
[0021] Compared with the prior art, the beneficial effects of the present invention are as follows: 1. This invention, by constructing an individualized digital twin that integrates genotype information, initial phenotypic parameters, and high-resolution environmental perception data, breaks through the simplistic assumption of treating the population as homogeneous individuals in traditional crop models, and achieves accurate characterization of the individual developmental trajectory of safflower seedlings.
[0022] 2. The hybrid modeling mechanism of mechanism model and data-driven model proposed in this invention not only retains the theoretical rigor of plant physiological and ecological processes, but also compensates for the prediction bias caused by model structure uncertainty and parameter drift in real time through deep neural networks, which significantly improves the accuracy and timeliness of growth simulation in complex dynamic environments.
[0023] 3. The multi-objective collaborative optimization framework established in this invention integrates mutually restrictive agronomic objectives such as biomass, medicinal component content, and resource utilization efficiency into the decision-making system. The generated cultivation control instructions take into account yield, quality, and sustainability, solving the problem that traditional experience-based management cannot take multiple objectives into account.
[0024] 4. The system architecture of this invention supports closed-loop control from data acquisition and model calculation to instruction execution, providing a feasible technical path for the intelligent and refined breeding of high-value-added medicinal plants such as safflower. Attached Figure Description
[0025] Figure 1 This is a schematic diagram of the overall technical solution architecture of the digital twin growth simulation and optimization method for safflower seedling propagation proposed in this invention; Figure 2 This is a schematic diagram of the core principle framework of the hybrid modeling mechanism of mechanism model and data-driven model in this invention; Figure 3 This is a flowchart illustrating the logical process framework for the construction of individualized digital twins and the fusion processing of environment-driven fields in this invention. Figure 4 This is a logical flowchart of the multi-objective collaborative optimization decision-making process for generating cultivation regulation strategies in this invention. Figure 5 This is a flowchart illustrating the logical process of dynamic updating and visualization rendering of the digital twin in this invention. Figure 6 This is a schematic diagram of the closed-loop interaction relationship and data flow between the physical planting unit and the digital twin system in this invention. Detailed Implementation
[0026] This invention provides a digital twin growth simulation and optimization method for safflower seedling propagation. Its core lies in constructing an individualized, high-fidelity digital twin of the safflower plant, integrating a physical mechanism model and a data-driven model to achieve precise simulation of the entire life cycle of safflower seedlings under complex dynamic cultivation environments. Based on this, multi-objective collaborative optimization is performed to generate executable, precise cultivation control instructions. The following will describe each step in detail with reference to the technical solution of this invention.
[0027] The method first performs step S1: acquiring multi-dimensional real-time sensing data on the individual genotype information, initial phenotypic parameters, and cultivation environment of safflower seedlings. This step is the fundamental input link of the entire digital twin system, and its data quality directly determines the accuracy of subsequent modeling and optimization. The acquisition of individual genotype information is completed by whole-genome resequencing of single safflower seed samples using a high-throughput sequencing platform, with a read length of no less than 150 base pairs to ensure coverage of key functional regions. After comparing the raw sequencing data with the reference genome, a genome-wide map of single nucleotide polymorphism sites is identified. Based on this map, a combination of molecular markers closely related to the synthesis of key agronomic traits and medicinal components of safflower is extracted, including but not limited to the homologous sequence of the GA20ox gene controlling plant height, TB1-class transcription factors determining the number of primary branches, allelic variations of the FT gene regulating the initial flowering period, MYB family transcription factors affecting glandular trichome density, and gene sequences encoding key enzymes such as chalcone synthase and flavonoid 3-hydroxylase in the hydroxysafflower yellow pigment A synthesis pathway.
[0028] Initial phenotypic parameters were measured non-destructively during the early stages of seed germination or when cotyledons were fully expanded. These parameters included thousand-seed weight, radicle length, cotyledon unfolding angle, and initial moisture content. Thousand-seed weight was determined using a precision electronic balance. Radicle length and cotyledon unfolding angle were calculated from plant point cloud data captured by a laser scanning 3D reconstruction device. Initial moisture content was obtained through near-infrared spectral reflectance inversion. Multi-dimensional real-time sensing data was continuously collected by a sensor network deployed above the canopy and in the root zone of individual plants, with a time resolution of 10 minutes. Sensors above the canopy included a total irradiance meter, a photosynthetically active radiation quantum sensor, an integrated air temperature and humidity probe, and a carbon dioxide infrared analyzer. The root zone sensors were a multi-parameter soil sensor array, arranged vertically at 10-cm intervals in layers from 0 to 40 cm in soil depth, monitoring soil volumetric water content, soil conductivity, and soil temperature in real time. All sensor data were accurately timestamped and transmitted to an edge computing gateway via an industrial fieldbus protocol.
[0029] After completing step S1, step S2 is executed: based on the individual genotype information and initial phenotypic parameters, an individualized three-dimensional geometric morphological skeleton of the safflower seedling is constructed, and its physiological state variables are initialized. This step aims to create a unique, structured digital identity for each safflower individual. The construction of the three-dimensional geometric morphological skeleton is based on the L-system fractal algorithm. The initial commonality of the L-system is determined by the radicle length and cotyledon unfolding angle in the initial phenotypic parameters, and the generation rules are activated by branching-related molecular markers in the genotype information. Specifically, if a highly expressed TB1-type transcription factor allele is detected, the rule inhibiting lateral bud germination is weakened, thereby increasing the probability of branching in the L-system. The initial branching level is set to level one, the internode length is calculated using empirical formulas based on the thousand-grain weight and initial moisture content, and the leaf tilt angle distribution function adopts a Beta distribution, with its shape parameters derived from the cotyledon unfolding angle. The fine geometric modeling of the leaves is parameterized using spherical harmonic functions, representing the leaf contour as a function of radius with respect to polar angle and azimuth angle in spherical coordinates, with its coefficients determined by fitting the leaf point cloud data obtained by laser scanning.
[0030] For organs such as stems and buds, radial basis function interpolation is used to map discrete organ measurement point cloud data onto a smooth, continuous surface model. The initialization of physiological state variables covers multiple dimensions, including carbon and nitrogen metabolic pools, water status, and hormone concentrations, such as leaf chlorophyll content, root vigor index, cell turgor pressure, and basal abscisic acid concentration. These initial values are partly directly mapped from phenotypic parameters and partly indirectly inferred from metabolic pathway-related markers in genotypic information, collectively constituting the initial internal state of the digital twin.
[0031] Then, step S3 is executed: the multi-dimensional real-time sensing data is spatiotemporally aligned and fused with historical environmental time-series data to form a unified environmental driving field dataset. The environmental driving field is the external excitation source driving the evolution of the digital twin, and its spatiotemporal consistency is crucial. Spatiotemporal alignment first addresses the time offset problem caused by different sensors due to sampling frequency and transmission delay. All real-time sensing data streams are matched with the contemporaneous data stored in the historical environmental database based on their timestamps. For missing data points, cubic spline interpolation is used to fill in the gaps. This method ensures the continuity of the first and second derivatives of the interpolation curve and is suitable for the smooth variation characteristics of environmental parameters.
[0032] For anomalous abrupt changes caused by sensor malfunctions or extreme weather, a sliding window median filter is used for filtering. The window width is set to 30 time steps (i.e., 5 hours). When the absolute deviation of a data point from the median of its window exceeds three times the standard deviation of that window, it is considered an anomaly and corrected. Fusion processing integrates heterogeneous data from the canopy and root zones into a unified data structure. This data structure uses the individual plant as the smallest spatial unit, and each time step contains a complete environmental feature vector with fixed dimensions, sequentially arranging total solar irradiance, photosynthetically active radiation flux density, air temperature, relative humidity, carbon dioxide concentration, and volumetric water content, conductivity, and temperature at four soil depths, totaling 21 environmental variables. This environmentally driven field dataset has a spatial resolution of individual plants and a temporal resolution of 10 minutes, providing high-quality input for subsequent hybrid modeling.
[0033] After obtaining the environmental driving field dataset, step S4 is executed: a set of mechanistic sub-models is established, encompassing core physiological processes such as photosynthesis, respiratory metabolism, water transpiration, nutrient absorption and distribution, and organogenesis. Each sub-model is described by a system of differential equations, which constitute the core physiological engine of the digital twin.
[0034] The photosynthesis sub-model adopts the Farquhar biochemical model framework, whose core equation describes the net photosynthetic rate under three conditions: RuBP carboxylation, RuBP oxygenation, and TPU limitation. The model inputs are intercellular carbon dioxide concentration, mesophyll conductance, and the maximum carboxylation rate (Vcmax) derived from genotype information. The water transpiration sub-model is based on the Penman-Monteith equation but introduces a dynamic feedback term for stomatal conductance, determined by leaf water potential and atmospheric saturated vapor pressure difference, to reflect the plant's self-regulatory behavior under water stress.
[0035] The nutrient uptake sub-model employs the Michaelis-Menten kinetic equations to describe the rate of root uptake of major elements such as nitrogen, phosphorus, and potassium. This model is coupled with a root distribution density function, calculated from the current root geometry, and also considers the soil nutrient diffusion coefficient, which is derived from soil electrical conductivity. The respiratory metabolism model distinguishes between maintenance respiration and growth respiration; the former is proportional to biomass, while the latter is related to the rate of new substance synthesis. The organogenesis model dynamically updates the dry matter accumulation of each organ based on the real-time allocation of carbon and nitrogen resources, thereby driving the topological evolution of the L-system's geometric framework. All these sub-models are implemented through a set of coupled ordinary differential equations, with their state variables being the physiological state variables initialized in step S2.
[0036] To overcome the prediction bias caused by parameter uncertainty and lack of dynamic modeling in pure mechanistic models, step S5 is performed: a data-driven correction module based on a deep temporal neural network is constructed. This module takes the environmental driving field dataset and physiological state variables as input and outputs a real-time compensation for the prediction bias of the mechanistic sub-model. This data-driven correction module adopts a bidirectional gated recurrent unit (GRU) network architecture. The network's input layer receives a sliding window of data with a length of 72 time steps (i.e., 12 hours), containing the environmental driving field data and the physiological state variable sequence output by the mechanistic model. The hidden layer consists of 128 bidirectional GRU units, capable of simultaneously capturing the forward and backward dependencies of the time series. The output layer is a fully connected layer with a dimension consistent with the number of physiological state variables, and the output value is the correction term of the first derivative of each physiological state variable at the current time step. During the offline phase before system deployment, the network is trained under supervised supervision using a large amount of historical field observation dataset. The training dataset contains the actual physiological state evolution trajectory of safflower seedlings under different environmental conditions, with the label being the residual between the first derivative of the actual state variable and the predicted derivative of the mechanistic model. During the online operation phase, the module continuously receives the environmental driving field data output from step S3 and the mechanistic model state output from step S4 in a sliding window manner, and calculates and outputs the compensation amount in real time.
[0037] After obtaining the original predictions from the mechanistic model and the data-driven corrections, step S6 is executed: the outputs of the mechanistic sub-model set and the compensation from the data-driven correction module are weighted and fused to generate the updated physiological state variables and three-dimensional geometric morphological parameters of the safflower seedlings at the next time step. The weighted fusion strategy employs an adaptive weight allocation mechanism to balance the theoretical reliability of the mechanistic model with the real-time correction capability of the data-driven model. Let the predicted value of a certain physiological state variable by the mechanistic model be... The correction amount for the data-driven model is The final predicted value after fusion ,in These are adaptive weighting coefficients. Weighting coefficients The value of is dynamically adjusted based on the prediction residual variance of the mechanistic model over a past period. The system continuously monitors the predictive performance of the mechanistic model and calculates its residual variance. .when When the value exceeds the preset threshold of 0.5, it indicates that the mechanism model has a significant bias in the current environment. At this point, w is increased to 0.7 to give the data-driven correction a higher confidence level; when When the value is below this threshold, it indicates that the mechanistic model is performing well. The value is maintained at 0.3 to preserve the physical consistency of the model. Through this fusion mechanism, all updated physiological state variables are obtained. These updated physiological state variables are fed into the organogenesis model to calculate the newly added dry matter distribution in each organ, and based on this, the parameters of the L-system geometric skeleton, such as intersegmental elongation, new leaf area, and branching angle, are updated to generate the three-dimensional geometric morphological parameters for the next time step.
[0038] Based on the updated geometric parameters, step S7 is executed: The digital twin visualization model of the safflower seedlings is dynamically reconstructed based on the updated 3D geometric parameters, and its internal state database is updated synchronously. This step achieves the visualization and data persistence of the digital twin. The reconstruction of the visualization model is completed by a dedicated rendering engine, which reads the latest L-system topology, spherical harmonic function leaf parameters, and radial basis function organ surfaces output from step S6, and generates a high-fidelity 3D mesh model in real time. This model supports visual effects such as lighting, shadows, and materials, and can intuitively display the real-time growth status of the safflower seedlings on virtual reality or 2D monitoring interfaces. Simultaneously, the internal state database is updated synchronously. This database stores all physiological state variables, environmental driving field snapshots, genotype metadata, and current geometric parameters in key-value pairs. The database uses memory-mapped file technology to ensure high-speed data exchange with the hybrid modeling computing unit, providing complete historical state backtracking capabilities for subsequent optimization decisions.
[0039] After the digital twin completes a full state update, the system enters the optimization decision-making phase, executing step S8: A multi-objective optimization problem is defined with the objective functions of maximizing biomass accumulation rate, achieving the threshold for effective medicinal component content, and optimizing water resource utilization efficiency. The mathematical expression of this multi-objective optimization problem is as follows: The first objective function f1 is the increase in aboveground dry matter per unit time, i.e. Where W is the dry weight of the aboveground part. The time step; the second objective function The mass fraction of hydroxysaffron yellow pigment A in the petals is required. ≥1.5%; Third objective function The amount of irrigation water consumed per kilogram of dry matter, i.e. ,Require ≤8 liters / kg. The constraints of this optimization problem include: soil volumetric water content θ must satisfy 60% ≤ θ ≤ 90% of field capacity; daily average temperature... The temperature must be 15 degrees Celsius or less. ≤28 degrees Celsius. The decision variables are four controllable cultivation factors: daily irrigation amount. (Value range: 0 to 10 liters), NPK fertilizer ratio (where the mass fraction of nitrogen is between 0.1% and 0.5%), lower limit of the light intensity control range. (Not less than 400 micromoles per square meter per second), and temperature and humidity setpoints (including daytime temperature setpoints). With nighttime temperature setting The difference is not less than 6 degrees Celsius.
[0040] To address the aforementioned multi-objective optimization problem, step S9 is executed: a genetic algorithm based on non-dominated sorting is used to iteratively optimize four decision variables: irrigation amount, fertilizer ratio, light intensity control range, and temperature and humidity setpoints. The specific configuration of this genetic algorithm is as follows: population size of 200, using real-number encoding to directly represent the four decision variables; simulated binary crossover with a crossover probability of 0.9; polynomial mutation with a mutation probability of 0.1; and a maximum evolutionary generation of 500 generations. During each generation, the algorithm generates 200 candidate solutions (i.e., cultivation strategies). The fitness evaluation of each candidate solution is not calculated using analytical functions, but rather by calling the hybrid modeling simulation interface comprised of steps S4 to S6. Specifically, the decision variables from the candidate solutions are used as environmental control setpoints for the next 72 hours, superimposed on the current environmental driving field, and then the hybrid model is driven to perform a 72-hour (432 time steps) forward simulation. After the simulation, the values of three objective functions are extracted from the results as the fitness of the candidate solution. The non-dominated sorting mechanism ranks all individuals in the population hierarchically based on these three objective values, prioritizing the retention of non-dominated solutions at the forefront. After 500 generations of iteration, the algorithm converges to a set of Pareto optimal solutions, representing the best cultivation strategy under different trade-offs among biomass, medicinal components, and water efficiency.
[0041] Finally, step S10 is executed: the optimized decision variable sequence is transformed into executable cultivation control instructions, which are then distributed to the physical planting units via IoT execution terminals. This step completes the closed-loop control from the digital world to the physical world. The Pareto optimal solution set output by the optimization algorithm typically selects a specific solution based on the current production objective (e.g., prioritizing yield or quality). This solution contains the values of four decision variables at 10-minute intervals over the next 72 hours. These values are then translated into specific equipment control instructions: irrigation amount... The percentage opening of the electrically operated proportional control valve is converted into a value and controlled by a pulse width modulation signal; fertilizer ratio. The fertilizer mother liquor injection rate, converted to a Venturi fertilizer applicator, is executed by a high-precision metering pump; the lower limit of the light intensity control range... The operating current of the LED supplemental lighting array is converted into a constant current drive power supply for regulation; the temperature and humidity setpoints are converted into the duty cycle of the wet curtain fan linkage unit, managed by the PLC controller. All these instructions are sent from the edge computing gateway to the corresponding execution terminal in the greenhouse or field via the Industrial Internet of Things (IIoT) protocol, thereby precisely controlling the actual growth environment of the safflower seedlings and guiding them to develop along the ideal trajectory simulated and optimized by the digital twin. At this point, the complete digital twin growth simulation and optimization cycle ends, and the system immediately enters the next time step, beginning a new round of data acquisition, status updates, and optimization decisions.
Claims
1. A digital twin growth simulation and optimization method for safflower seedling propagation, characterized in that, include: Acquire multi-dimensional real-time sensing data on individual genotype information, initial phenotypic parameters, and cultivation environment of safflower seedlings; Based on the individual genotype information and initial phenotypic parameters, an individualized three-dimensional geometric morphological skeleton of safflower seedlings was constructed, and its physiological state variables were initialized. The multi-dimensional real-time sensing data and historical environmental time-series data are spatiotemporally aligned and fused to form a unified environment-driven field dataset. A set of mechanistic sub-models was established, which included core physiological processes such as photosynthesis, respiratory metabolism, water transpiration, nutrient absorption and distribution, and organogenesis. Each sub-model was described by a set of differential equations. A data-driven correction module based on a deep temporal neural network is constructed. This module takes the environmental driving field dataset and physiological state variables as input and outputs the real-time compensation amount for the prediction bias of the mechanism sub-model. The output of the mechanism sub-model set is weighted and fused with the compensation amount of the data-driven correction module to generate the updated physiological state variables and three-dimensional geometric morphological parameters of safflower seedlings at the next time step. Based on the updated three-dimensional geometric morphology parameters, the digital twin visualization model of safflower seedlings is dynamically reconstructed, and its internal state database is updated synchronously. A multi-objective optimization problem is set up with the objective functions of maximizing biomass accumulation rate, achieving the threshold of effective medicinal component content, and optimizing water resource utilization efficiency. A genetic algorithm based on non-dominated sorting was used to iteratively optimize four decision variables: irrigation amount, fertilizer ratio, light intensity control range, and temperature and humidity setpoint. The optimized decision variable sequence is transformed into executable cultivation control instructions, which are then sent to the physical planting unit via an IoT execution terminal.
2. The digital twin growth simulation and optimization method for safflower seedling propagation according to claim 1, characterized in that, The process of obtaining individual genotype information for safflower seedlings includes: Whole genome resequencing of safflower seed samples was performed using a high-throughput sequencing platform to obtain single nucleotide polymorphism (SNP) site maps, and molecular marker combinations related to plant height, number of branches, flowering period, glandular hair density, and key enzyme encoding genes in the hydroxysafflower yellow A synthesis pathway were extracted.
3. The digital twin growth simulation and optimization method for safflower seedling propagation according to claim 2, characterized in that, The initial phenotypic parameters include seed weight per thousand seeds, radicle length, cotyledon unfolding angle, and initial moisture content. The multi-dimensional real-time sensing data includes total solar radiation illuminance above the canopy, photosynthetically active radiation flux density, air temperature, relative humidity, carbon dioxide concentration, soil volumetric water content, soil electrical conductivity, and soil temperature.
4. The digital twin growth simulation and optimization method for safflower seedling propagation according to claim 3, characterized in that, The construction of the individualized three-dimensional geometric morphological skeleton for safflower seedlings includes: Based on the L-system fractal algorithm, the initial topology is generated according to the initial branch level, internode length, and leaf tilt angle distribution function. The blade profile is parametrically modeled using spherical harmonic functions; The radial basis function interpolation method is used to map discrete organ measurement point cloud data to a continuous surface model.
5. The digital twin growth simulation and optimization method for safflower seedling propagation according to claim 4, characterized in that, The environmental driving field dataset has a temporal resolution of 10 minutes and a spatial resolution of single-plant scale. Its fusion processing includes filling missing data with cubic spline interpolation and removing abnormal mutation points with sliding window midpoint filtering.
6. The digital twin growth simulation and optimization method for safflower seedling propagation according to claim 5, characterized in that, The photosynthesis sub-model in the mechanistic sub-model set adopts the Farquhar biochemical model framework, with the inputs being intercellular carbon dioxide concentration, mesophyll conductance, and maximum carboxylation rate, and the output being net photosynthetic rate; the water transpiration sub-model is based on the Penman-Monteith equation and introduces a dynamic feedback term for stomatal conductance; the nutrient uptake sub-model adopts the Michaelis-Menten kinetic equation and couples the root distribution density function with the soil nutrient diffusion coefficient.
7. The digital twin growth simulation and optimization method for safflower seedling propagation according to claim 6, characterized in that, The deep temporal neural network is a bidirectional gated recurrent unit network with 128 hidden layer nodes, an input sequence length of 72 time steps, and outputs the first derivative correction terms of each physiological state variable. The network is trained in an offline phase using historical field observation datasets under supervision, and in an online phase, it receives real-time data streams in a sliding window manner and outputs instantaneous compensation quantities.
8. The digital twin growth simulation and optimization method for safflower seedling propagation according to claim 7, characterized in that, The weighted fusion adopts an adaptive weight allocation strategy. The weight coefficients are dynamically adjusted according to the variance of the residuals of the mechanism model. When the residual variance exceeds the preset threshold of 0.5, the weight of the data-driven correction module is increased to 0.7; otherwise, it remains at 0.
3.
9. The digital twin growth simulation and optimization method for safflower seedling propagation according to claim 8, characterized in that, The objective functions of the multi-objective optimization problem are defined as follows: the first objective function is the increase in aboveground dry matter per unit time; the second objective function is that the mass fraction of hydroxysafflower yellow pigment A in the petals is not less than 1.5%; and the third objective function is that the amount of irrigation water consumed per kilogram of dry matter does not exceed 8 liters. The constraints include soil moisture content not less than 60% and not more than 90% of field capacity, and daily average temperature between 15 degrees Celsius and 28 degrees Celsius.
10. The digital twin growth simulation and optimization method for safflower seedling propagation according to claim 9, characterized in that, The non-dominated sorting genetic algorithm has a population size of 200, a crossover probability of 0.9, a mutation probability of 0.1, and a maximum number of generations of evolution of 500. The decision variables are encoded using real numbers. The irrigation amount ranges from 0 to 10 liters per day. The nitrogen mass fraction in the nitrogen, phosphorus, and potassium fertilizer ratio is between 0.1% and 0.5%. The lower limit of the light intensity control range is not less than 400 micromoles per square meter per second. The difference between the daytime temperature setpoint and the nighttime temperature setpoint is not less than 6 degrees Celsius.
Citation Information
Cited By
A tree species growth environment dynamic monitoring method and system based on internet of things perception
CN122198586A