Crop planting optimization decision-making method based on digital twinning and reinforcement learning

By constructing a crop-environment digital twin model and using deep reinforcement learning agents for interactive learning, a global optimal control strategy is generated, which solves the problems of static isolation and reactive control strategies in existing technologies and achieves the optimal balance between crop yield and resource utilization.

CN120671935AActive Publication Date: 2025-09-19SICHUAN AGRI UNIV

Patent Information

Application Number
CN202511179101.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-22
Publication Date
2025-09-19
Estimated Expiration
2045-08-22

AI Technical Summary

Technical Problem

Existing crop planting management technologies have problems such as static and isolated control strategies, reactive rather than predictive control strategies, and difficulty in finding the global optimal control strategy, making it impossible to achieve the optimal balance between crop yield and resource utilization.

Method used

A method based on digital twins and reinforcement learning is used to construct a crop-environment digital twin model. Through interactive learning between deep reinforcement learning agents and digital twins, a globally optimal control strategy is generated.

Benefits of technology

It achieves dynamic and adaptive generation of global optimal control strategies, which can maximize yield and resource utilization efficiency during the crop growth cycle, and realize forward-looking predictive control and personalized planting.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120671935A_ABST
    Figure CN120671935A_ABST
Patent Text Reader

Abstract

The invention relates to the field of crop planting optimization decision making, in particular to a crop planting optimization decision making method based on digital twinning and reinforcement learning. According to the technical scheme, the method comprises the steps of collecting greenhouse multi-modal data including greenhouse environment data, crop data, soil data and management data; cleaning, aligning and normalizing the acquired data to form a structured multi-modal time series data set; identifying the growth cycle of the crop based on the collected data; constructing a crop-environment digital twinborn model which is mapped with the physical greenhouse environment in real time; constructing a digital twin simulation space; and reinforcement learning and strategy optimization are carried out in a digital twinborn simulation space, so that a global optimal control strategy is dynamically and adaptively generated. The method is suitable for crop planting optimization decision making.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of crop planting optimization decision-making, and specifically to a crop planting optimization decision-making method based on digital twins and reinforcement learning. Background Art

[0002] With the development of the Internet of Things, big data, and artificial intelligence technologies, precision agriculture has become a key area of ​​modern agriculture. Traditional crop management relies primarily on grower experience or passive, extensive control based on simple sensor thresholds (e.g., turning on fans when the temperature exceeds 30°C or initiating irrigation when soil moisture falls below 50%). These approaches suffer from the following problems: 1) Control strategies are static and isolated, failing to account for the coupled effects of multiple environmental factors (such as temperature, light, humidity, CO2 concentration, and soil nutrients); 2) Control approaches are reactive rather than predictive, often requiring intervention only after crop stress appears, failing to achieve optimal growth; 3) Globally optimal control strategies are difficult to identify, making it difficult to strike an optimal balance between resource consumption (water, fertilizer, and electricity) and crop yield and quality.

[0003] To address the above issues, some studies have begun to apply crop growth models for growth simulation, or use machine learning for pest and disease identification and growth prediction. Although the use of automated systems has improved decision-making capabilities, the following problems still exist.

[0004] Most automated systems use control logic based on fixed thresholds (for example, ventilation is activated when the temperature exceeds 30°C). This approach is reactive and cannot adapt to the dynamic environmental demands of crops at different growth stages (such as seedling, flowering, and fruiting). It also cannot handle the coupled effects of multiple variables, such as the synergistic promotion of photosynthesis by high light and high carbon dioxide concentrations.

[0005] Complex, long-term management cycles (such as integrated water and fertilizer programs throughout the entire growing season) rely heavily on the experience of planting experts. This experience is difficult to quantify, replicate, and optimize, and trial and error is costly and risky when faced with extreme weather or new varieties.

[0006] Simple crop growth mechanism models have fixed parameters and are difficult to accurately reflect the actual growth conditions of a specific variety in a specific greenhouse; while simple data-driven models (such as machine learning predictions) lack biological mechanism constraints and have poor generalization and interpretability.

[0007] Existing technical solutions are mostly open-loop prediction or single-step control, and lack a method for generating a global optimal control strategy that can maximize the final yield or benefit throughout the entire growth cycle. Summary of the Invention

[0008] The purpose of the present invention is to overcome the shortcomings of the existing technology and provide a crop planting optimization decision-making method based on digital twins and reinforcement learning, which realizes the dynamic and adaptive generation of the global optimal control strategy.

[0009] The present invention adopts the following technical solutions to achieve the above-mentioned purpose. The present invention provides a crop planting optimization decision-making method based on digital twins and reinforcement learning, comprising: S1. Collect greenhouse multimodal data, including greenhouse environment data, crop data, soil data, and management data; S2. Clean, align, and normalize the collected data to form a structured multimodal time series dataset; S3. Identify the growth cycle of crops based on the collected data; S4. Build a crop-environment digital twin model that maps the physical greenhouse environment in real time; S5. Build digital twin simulation space; The digital twin is formalized as a Markov decision process defined as the tuple , where represents the state space, represents the action space, represents the state transition probability function, represents the reward function, Represents the discount factor. The Markov decision process is the training environment for reinforcement learning agents. The core is the state transfer function , describes the current state Next action After that, the system moves to the next state probability; S6, reinforcement learning strategy optimization; Through interactive learning between a deep reinforcement learning agent and a digital twin, the state space S is defined as follows: state At the moment is defined as: ; Where, Indicates the current state, Indicates the current time Growth stage, Indicates the current time Cumulative dry matter content, Indicates the current time Leaf area index, Indicates the current time temperature, Indicates the current time humidity, Indicates the current time The amount of light, Indicates the current time The carbon dioxide concentration, Indicates the current time The soil nitrogen content, Indicates the temperature forecast for the next 24 hours. Indicates the predicted light value for the next 24 hours; Definition of action space A: ; Where, Indicates the action to be performed in the current state. Indicates the temperature setting value, Indicates the humidity setting value, Indicates the fill light duration. Indicates the irrigation amount, Indicates fertilizer concentration; Design of reward function R: ; Where, Indicates reward, It represents the single-step dry matter increment simulated by the digital twin and is a direct proxy indicator of yield. Indicates execution of an action The cost of electricity consumed, Indicates the cost of water and fertilizer consumed, Indicates the next state The penalty item generated when the environmental indicators exceed the suitable range of crops, Respectively represent the adjustable weight coefficients of the corresponding items; Strategy Learning: The goal of the agent is to learn a policy , to maximize the expected cumulative reward , represents the offset of a time step, Indicates expected cumulative reward; This is achieved by optimizing the policy network through the PPO algorithm. The core is the iterative update of the value function, following the Bellman optimality equation: ; By conducting simulation training for multiple growth cycles in the digital twin simulation space, the agent eventually learns the optimal action-value function and the optimal strategy , E means in the current state Take action The expectation of the total return that may be obtained in the future under the conditions of

[0010] Furthermore, the optimization decision-making method also includes: S7, closed-loop control and feedback optimization; The trained optimal strategy Deploy and run; Forward control: Every decision cycle, collect the real status ,Will Input policy network , get the optimal action ,Will Parse into specific instructions to drive physical actuators; Feedback optimization: At the same time, the real state collected As new training samples, they are used to update or fine-tune the data-driven correction sub-model online.

[0011] Furthermore, in step S1, collecting greenhouse environment data, crop data, soil data and management data specifically includes: Greenhouse environment data collection: Through temperature and humidity sensors, light intensity sensors and carbon dioxide concentration sensors, real-time data on air temperature, humidity, light and carbon dioxide concentration in the greenhouse are collected; Crop data collection: Using RGB cameras, 3D cameras, hyperspectral cameras, and leaf temperature sensors, the system collects crop height, leaf area index, canopy structure, leaf color, fruit quantity and size, and leaf temperature data. Soil data collection: Soil sensors collect soil pH, conductivity, temperature, humidity, and nutrient content data; Management Data Collection: Manually record planting time, irrigation amount, fertilizer concentration and spraying data through touch screen terminals or mobile apps.

[0012] Furthermore, step S3 specifically includes: The collected crop data and the crop planting start date are input into the pre-trained convolutional neural network. The pre-trained convolutional neural network is used to identify the phenological characteristics of the crops, and then based on the phenological model of effective accumulated temperature, the current growth stage of the crops is comprehensively judged.

[0013] Furthermore, step S4 specifically includes: The crop-environment digital twin model includes the following coupled sub-models; Crop growth mechanism sub-model: The core of the crop growth mechanism sub-model is dry matter accumulation, and the daily dry matter increment is calculated by photosynthesis efficiency: ; Where, represents the daily dry matter increase, represents the efficiency of light energy utilization, (Photosynthetically Active Radiation) means photosynthetically active radiation. represents the extinction coefficient, which describes the degree of light shielding by the canopy. (Leaf Area Index) represents the leaf area index. are the stress correction functions of temperature, carbon dioxide concentration and nitrogen on photosynthesis respectively; Data-driven corrector model: Using real-time collected data, the set parameters of the crop growth mechanism sub-model are corrected online through Kalman filtering or long short-term memory network.

[0014] The beneficial effects of the present invention are: This invention constructs a high-fidelity crop-environment dynamic digital twin that integrates crop growth mechanisms with real-time multimodal data. Based on this, it creatively uses it as a virtual training environment for reinforcement learning agents. Within this digital twin, the agent autonomously explores and masters a long-term, multivariable, and dynamically coordinated global optimal control strategy by learning a strategy designed to maximize cumulative rewards (combined yield and cost) throughout the entire growth period. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 This is a flow chart of a crop planting optimization decision-making method based on digital twins and reinforcement learning provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0016] To make the objectives, technical solutions and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention.

[0017] The present invention provides a crop planting optimization decision-making method based on digital twin and reinforcement learning, such as Figure 1 As shown, specifically including: S1. Collect greenhouse multimodal data, including greenhouse environment data, crop data, soil data, and management data; Greenhouse environment data collection: Through temperature and humidity sensors, light intensity sensors and carbon dioxide concentration sensors, real-time data on air temperature, humidity, light and carbon dioxide concentration in the greenhouse are collected; Crop data collection: Using RGB cameras, 3D cameras, hyperspectral cameras, and leaf temperature sensors, non-destructive data such as plant height, leaf area index, canopy structure, leaf color, fruit number and size, and leaf temperature are collected. Soil data collection: Soil sensors collect soil pH, conductivity, temperature, humidity, and nutrient content (nitrogen, phosphorus, potassium) data; Management Data Collection: Manually record the planting time (year-month-day), irrigation amount (L), fertilizer concentration (g / L), and number of spraying times through the touch screen terminal or mobile app.

[0018] The present invention can also use other sensors such as laser radar to obtain plant morphological data.

[0019] S2. Clean the collected data (remove outliers), align (unify timestamps), and normalize them to form a structured multimodal time series dataset, providing a data foundation for the subsequent construction and synchronization of digital twin models.

[0020] S3. Identify the growth cycle of crops based on the collected data; The RGB camera deployed in the greenhouse regularly collects crop images, combines them with the planting start date, and inputs them into a pre-trained convolutional neural network to identify key phenological characteristics of the crop (such as the number of flower buds and fruits). Combined with the phenological model based on effective accumulated temperature, the current growth stage of the crop is comprehensively judged. (seedling, vegetative growth period, flowering period, fruit setting period, maturity period).

[0021] S4. Build a crop-environment digital twin model that maps the physical greenhouse environment in real time; The crop-environment digital twin model consists of the following coupled sub-models; Crop growth mechanism sub-model: The core of the crop growth mechanism sub-model is dry matter accumulation, and the daily dry matter increment is calculated by photosynthesis efficiency: ; Where, represents the daily dry matter increase, represents the efficiency of light energy utilization, represents photosynthetically active radiation, represents the extinction coefficient, which describes the degree of light shielding by the canopy. represents the leaf area index, are the stress correction functions of temperature, carbon dioxide concentration and nitrogen on photosynthesis (range 0-1); The crop growth mechanism sub-model can be replaced by other mature models, such as WOFOST (World Food Studies, a mechanistic crop growth model), DSSAT (Decision Support System for Agrotechnology Transfer, a comprehensive agricultural technology decision support system), or TOMGRO (Tomato Growth Model, a greenhouse tomato growth model).

[0022] Data-driven corrector model: Using real-time collected data, the parameters of the crop growth mechanism sub-model are corrected online through Kalman filtering or long short-term memory network, such as light energy utilization efficiency. and extinction coefficient , ensuring that model outputs are synchronized with the real world. Long Short-Term Memory (LSTM) networks were chosen for their ability to capture long-term temporal dependencies and learn the dynamic patterns of parameter changes from historical environmental and growth data sequences. This is crucial for simulating processes such as crop growth, which have long-term memory effects.

[0023] For example, the long short-term memory network receives historical environmental data sequences and crop growth data sequences, predicts the model parameter correction amount at the next moment, and thus dynamically adjusts the mechanism model so that it can better fit the growth curve under specific varieties and environments.

[0024] In addition to the long short-term memory network, the present invention can also select other time series prediction models such as GRU (Gated Recurrent Unit) and Transformer.

[0025] S5. Build digital twin simulation space; The digital twin is formalized as a Markov decision process defined as the tuple , where represents the state space, represents the action space, represents the state transition probability function, represents the reward function, Represents the discount factor. The Markov decision process is the training environment for reinforcement learning agents. The core is the state transfer function , describes the state Next action After that, the system moves to the next state probability; in the present invention, this function is determined and calculated by the crop-environment digital twin model.

[0026] S6, reinforcement learning strategy optimization; A deep reinforcement learning agent interacts and learns with the digital twin. In reinforcement learning, the agent represents the subject that performs actions, interacts with the environment and learns. In the present invention, the agent represents a deep neural network model.

[0027] Definition of state space S: state At the moment is defined as: ; Where, Indicates the current state, Indicates the current time Growth stage, Indicates the current time Cumulative dry matter content, Indicates the current time Leaf area index, Indicates the current time temperature, Indicates the current time humidity, Indicates the current time The amount of light, Indicates the current time The carbon dioxide concentration, Indicates the current time The soil nitrogen content, Indicates the temperature forecast for the next 24 hours. Indicates the predicted light value for the next 24 hours; Definition of action space A: ; Where, Indicates the action to be performed in the current state. Indicates the temperature setting value, Indicates the humidity setting value, Indicates the fill light duration. Indicates the irrigation amount, Indicates fertilizer concentration; Design of reward function R: ; Where, Indicates reward, It represents the single-step dry matter increment simulated by the digital twin and is a direct proxy indicator of yield. Indicates execution of an action The cost of electricity consumed, Indicates the cost of water and fertilizer consumed, Indicates the next state The penalty item generated when the environmental indicators exceed the suitable range of crops, Respectively represent the adjustable weight coefficients of the corresponding items; The design of the reward function of the present invention is adjustable. In addition to maximizing yield, other optimization objectives can also be set, such as maximizing quality (such as the sugar-acid ratio), maximizing profit (taking into account both selling price and costs), or minimizing carbon emissions.

[0028] Strategy Learning: The goal of the agent is to learn a policy , to maximize the expected cumulative reward , Represents the offset of a time step, that is represents the reward at step k, Indicates expected cumulative reward; This is achieved by optimizing the policy network through the PPO algorithm. The core is the iterative update of the value function, following the Bellman optimality equation: ; Where, Indicates that the status Take action The best value, Indicates the next state Select the best action , E means in the current state Take action The expectation of the total return that may be obtained in the future under the conditions of

[0029] In the framework of reinforcement learning, the response of the environment (i.e., state transition and reward) is not necessarily completely deterministic, but often random and uncertain. Take exactly the same action , the environment may also transition to multiple different next states , and the probability of each transition occurring is different.

[0030] Therefore, we cannot only consider a single future result, but must consider all possible future results and make a weighted average based on their probability of occurrence. This is the mathematical expectation. The work done.

[0031] In the present invention, this uncertainty is reflected in the following aspects: Biological randomness: Crops are living organisms, and their responses to environmental manipulations have inherent biological randomness. For example, even if the same light, water, and fertilizer (action ), the dry matter increment of a single crop (affecting the next state ) may also fluctuate within a small range.

[0032] Uncertainty in environmental models: Although digital twin models offer high fidelity, they are still approximations of the real world. The models themselves may contain stochastic processes (for example, simulating the random occurrence of disease), or weather forecasts are inherently probabilistic. For example, when executing the "ventilation and cooling" action, the resulting temperature (the next state) will be affected by uncontrollable random factors such as outdoor wind speed.

[0033] Sensor Noise: Measuring the Next State The sensor readings themselves may also contain random noise.

[0034] Uncertainty of rewards: Immediate rewards It may also be uncertain. For example, the electricity cost of performing the fill light action may be affected by the time-of-use electricity price, and the electricity price at some point in the future may be a random variable.

[0035] will expect Expand: Assume that from the state Execute an action Later, it may be transferred to Different next states , the probability of each transition is , and the corresponding reward is .

[0036] Then the Bellman optimality equation can be written more specifically as: ; The above expansion clearly shows that: The value of is the weighted average of the total rewards of all possible futures, and the weight is the total reward of each future (i.e., the transfer to state ) probability of occurrence The introduction of the expectation E reflects that the technical solution of the present invention takes into account the inherent randomness and uncertainty in the agricultural production process. The goal of the intelligent agent is not to obtain the highest reward in a lucky situation, but to learn a strategy that performs best in the long term and on average. Value, representing the current state Execute an action The most stable and reliable long-term value expectation makes the final generated control strategy more robust and practical.

[0037] The present invention prefers the Proximal Policy Optimization (PPO) algorithm because it limits the difference between the new and old policies during policy updates by introducing a pruning objective function, effectively avoiding the training crash problem caused by an excessively large policy update step size, thereby ensuring the stability and convergence efficiency of reinforcement learning in complex agricultural environment models.

[0038] By conducting simulation training for multiple growth cycles in the digital twin simulation space, the agent eventually learns the optimal action-value function and the optimal strategy .

[0039] S7, closed-loop control and feedback optimization; The trained optimal strategy Deploy and run; Forward control: Every decision cycle, collect the real status ,Will Input policy network , get the optimal action ,Will Parse into specific instructions to drive physical actuators; Feedback optimization: At the same time, the real state collected As new training samples, they are used to update or fine-tune the data-driven correction sub-model online.

[0040] The present invention achieves a shift from "passive reaction" to "active optimization": Different from the fixed threshold control of the existing technology, the present invention uses reinforcement learning to perform global optimization in digital twins, and can discover dynamic and coordinated multivariable optimal control strategies that are difficult for human experts to formulate, thereby maximizing crop yield and resource utilization efficiency.

[0041] This invention enables highly accurate, forward-looking predictive control: Digital twins combine mechanisms and data to accurately predict the long-term impact of control actions. Reinforcement learning agents can leverage these predictions to make forward-looking decisions, creating a consistently optimal environment for crop growth.

[0042] This invention provides a safe and low-cost platform for strategy validation: Digital twins provide a virtual testing ground for complex control strategies. Researchers and growers can conduct extensive experiments and optimization safely, quickly, and cost-effectively without impacting the physical crops, significantly reducing decision-making risks and R&D costs.

[0043] This invention achieves highly adaptive and personalized planting: The framework of this invention is universally applicable. By replacing or adjusting the crop and environmental model parameters in the digital twin and retraining, it can quickly adapt to different crop varieties, different greenhouse structures, and different climatic conditions, achieving personalized precision planting with "one crop, one strategy" and "one greenhouse, one strategy."

[0044] The foregoing description is merely a preferred embodiment of the present invention. It should be understood that the present invention is not limited to the form disclosed herein and should not be construed as excluding other embodiments. Rather, the present invention can be used in various other combinations, modifications, and environments and can be modified within the scope of the concept described herein through the above teachings or techniques or knowledge in the relevant field. Modifications and variations made by those skilled in the art that do not depart from the spirit and scope of the present invention are intended to be protected by the appended claims.

Claims

1. A crop planting optimization decision-making method based on digital twins and reinforcement learning, characterized by: include: S1. Collect greenhouse multimodal data, including greenhouse environment data, crop data, soil data, and management data; S2. Clean, align, and normalize the collected data to form a structured multimodal time series dataset; S3. Identify the growth cycle of crops based on the collected data; S4. Build a crop-environment digital twin model that maps the physical greenhouse environment in real time; S5. Build digital twin simulation space; S6. Perform reinforcement learning strategy optimization in the digital twin simulation space.

2. The crop planting optimization decision-making method based on digital twin and reinforcement learning according to claim 1 is characterized in that: The optimization decision-making method also includes: S7, closed-loop control and feedback optimization; The trained optimal strategy Deploy and run; Forward control: Every decision cycle, collect the real status ,Will Input the policy network to get the optimal action ,Will Parse into specific instructions to drive physical actuators; Feedback optimization: At the same time, the real state collected As new training samples, they are used to update or fine-tune the data-driven correction sub-model online.

3. The crop planting optimization decision-making method based on digital twin and reinforcement learning according to claim 1 is characterized in that: In step S1, collecting greenhouse environment data, crop data, soil data and management data specifically includes: Greenhouse environment data collection: Through temperature and humidity sensors, light intensity sensors and carbon dioxide concentration sensors, real-time data on air temperature, humidity, light and carbon dioxide concentration in the greenhouse are collected; Crop data collection: Using RGB cameras, 3D cameras, hyperspectral cameras, and leaf temperature sensors, the system collects crop height, leaf area index, canopy structure, leaf color, fruit quantity and size, and leaf temperature data. Soil data collection: Soil sensors collect soil pH, conductivity, temperature, humidity, and nutrient content data; Management Data Collection: Manually record planting time, irrigation amount, fertilizer concentration and spraying data through touch screen terminals or mobile apps.

4. The crop planting optimization decision-making method based on digital twin and reinforcement learning according to claim 1 is characterized in that: Step S3 specifically includes: The collected crop data and the crop planting start date are input into the pre-trained convolutional neural network. The pre-trained convolutional neural network is used to identify the phenological characteristics of the crops, and then based on the phenological model of effective accumulated temperature, the current growth stage of the crops is comprehensively judged.

5. The crop planting optimization decision-making method based on digital twin and reinforcement learning according to claim 1 is characterized in that: Step S4 specifically includes: The crop-environment digital twin model includes the following coupled sub-models; Crop growth mechanism sub-model: The core of the crop growth mechanism sub-model is dry matter accumulation, and the daily dry matter increment is calculated by photosynthesis efficiency: ; Where, represents the daily dry matter increase, It represents the efficiency of light energy utilization. represents photosynthetically active radiation, represents the extinction coefficient, which describes the degree of light shielding by the canopy. represents the leaf area index, are the stress correction functions of temperature, carbon dioxide concentration and nitrogen on photosynthesis respectively; Data-driven corrector model: Using real-time collected data, the set parameters of the crop growth mechanism sub-model are corrected online through Kalman filtering or long short-term memory network.

6. The crop planting optimization decision-making method based on digital twin and reinforcement learning according to claim 1 is characterized in that: Step S5 specifically includes: The digital twin is formalized as a Markov decision process defined as the tuple , where represents the state space, represents the action space, represents the state transition probability function, represents the reward function, Represents the discount factor. The Markov decision process is the training environment for reinforcement learning agents. The core is the state transfer function , describes the current state Next action Then, transfer to the next state probability.

7. The crop planting optimization decision-making method based on digital twin and reinforcement learning according to claim 6 is characterized in that: Step S6 specifically includes: Through interactive learning between a deep reinforcement learning agent and a digital twin, the state space S is defined as follows: state At the moment is defined as: ; Where, Indicates the current state, Indicates the current time Growth stage, Indicates the current time Cumulative dry matter content, Indicates the current time Leaf area index, Indicates the current time temperature, Indicates the current time humidity, Indicates the current time The amount of light, Indicates the current time The carbon dioxide concentration, Indicates the current time The soil nitrogen content, Indicates the temperature forecast for the next 24 hours. Indicates the predicted light value for the next 24 hours; Definition of action space A: ; Where, Indicates the action to be performed in the current state. Indicates the temperature setting value, Indicates the humidity setting value, Indicates the fill light duration. Indicates the irrigation amount, Indicates fertilizer concentration; Design of reward function R: ; Where, Indicates reward, It represents the single-step dry matter increment simulated by the digital twin and is a direct proxy indicator of yield. Indicates execution of an action The cost of electricity consumed, Indicates the cost of water and fertilizer consumed, Indicates the next state The penalty item generated when the environmental indicators exceed the suitable range of crops, Respectively represent the adjustable weight coefficients of the corresponding items; Strategy Learning: The goal of the agent is to learn a policy , to maximize the expected cumulative reward , represents the offset of a time step, Indicates expected cumulative reward; This is achieved by optimizing the policy network through the PPO algorithm. The core is the iterative update of the value function, following the Bellman optimality equation: ; Where, Indicates that the status Take action The best value, Indicates the next state Select the best action , E means in the current state Take action The expectation of the total return that may be obtained in the future under the conditions of By conducting simulation training for multiple growth cycles in the digital twin simulation space, the agent eventually learns the optimal action-value function and the optimal strategy .

Citation Information

Patent Citations

  • Multi-source federated environment control method for crop cultivation in plant factory

    CN114967626A

  • Intelligent environment control method for crop planting in plant factory

    CN115016413A

  • Farmland irrigation intelligent decision-making system based on digital twinning

    CN115804334A

  • Crop cultivation optimization method based on digital twinning

    CN115880433A

  • Agricultural digital twinning management system combining real digital space and virtual digital space

    CN118799105A

Cited By

  • Agricultural greenhouse intelligent monitoring control system

    CN120928896A

  • An intelligent monitoring and control system for agricultural greenhouse

    CN120928896B

  • Greenhouse environment regulation and control method and system for plant cultivation

    CN121478052A

  • Rice seedling raising shed monitoring method and system

    CN121657801A