Plant factory plant growth environment control system based on deep learning

By using deep learning technology, a plant factory growth environment control system was constructed, which achieved synergistic regulation of short-term and long-term goals, solved the problems of environmental mutation and stress, and realized efficient and stable plant growth environment control.

CN121349023AInactive Publication Date: 2026-01-16HEBEI YIXUE REFRIGERATION TECH CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511619100.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-06
Publication Date
2026-01-16
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In existing plant factory environmental control systems, short-term operational goals and long-term growth goals are difficult to coordinate, high-frequency control lacks long-term predictive guidance, leading to environmental mutations and biological stress, and long-term strategies cannot be corrected based on high-frequency execution data.

Method used

A deep learning-based plant factory growth environment control system is employed to achieve predictive and proactive multi-objective regulation through the collaborative work of data acquisition, state aggregation, low-frequency strategy, cross-scale coupling, and high-frequency control unit. The system comprises a data acquisition unit, a state aggregation unit, a low-frequency strategy unit, a cross-scale coupling unit, and a high-frequency control unit. It utilizes a temporal prediction network and a deep reinforcement learning model to generate a dynamically weighted reward function and the final execution action.

Benefits of technology

It resolves the multi-scale conflict between short-term energy consumption and long-term yield, ensures the stability and predictability of control, reduces the stress of environmental mutations on plants, and realizes proactive and coordinated control of the plant growth environment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121349023A_ABST
    Figure CN121349023A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of plant factory plant growth environment control, in particular to a plant factory plant growth environment control system based on deep learning. The system comprises a data acquisition unit, a state aggregation unit, a low-frequency strategy unit, a cross-scale coupling unit and a high-frequency control unit. According to the system, data of high-frequency environment, low-frequency plants and the like are collected, the core of the system is that a low-frequency strategy unit calculates probability distribution of future key development nodes through a time sequence prediction network, and predictive induction intensity is calculated through a cross-scale coupling unit; the high-frequency control unit takes the intensity as a dynamic weight, a reward function of dynamic weighting is solved by combining a deep reinforcement learning model, and a final action is generated; the conflict problem that a short-term operation target and a long-term growth target in a plant factory are difficult to cooperate is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of plant growth environment control technology in plant factories, specifically a plant growth environment control system for plant factories based on deep learning. Background Technology

[0002] In the field of environmental control in plant factories, system operation involves high-frequency environmental state changes and low-frequency plant growth and development cycles. Existing control strategies generally suffer from multi-scale objective conflicts, namely, the difficulty in coordinating short-term operational objectives with long-term growth objectives. On the one hand, high-frequency control units, lacking long-term predictive guidance, are prone to getting trapped in short-term local optima during optimization. On the other hand, low-frequency long-term strategies are mostly open-loop predictions, unable to be corrected based on actual data from high-frequency execution, leading to discrepancies between the model and reality. Furthermore, the instantaneous decisions of the controller can easily produce drastic actions, triggering environmental mutations and causing biological stress. Therefore, how to achieve predictive and proactive multi-objective coordinated regulation to overcome the conflict between short-term energy consumption and long-term yield, and to ensure the stability of control, is an urgent technical problem to be solved. Summary of the Invention

[0003] To address the aforementioned technical problems, this invention provides a deep learning-based plant growth environment control system for plant factories. Specifically, the technical solution of this invention includes: The data acquisition unit is used to collect high-frequency environmental status data and low-frequency plant status and runtime data. The state aggregation unit is used to aggregate the runtime data to obtain an execution summary; The low-frequency strategy unit is used to receive the low-frequency plant state and the execution summary, and calculate the probability distribution of future key developmental nodes based on the time-series prediction network. A cross-scale coupling unit is used to calculate the predictive induction intensity based on the probability distribution of the key developmental nodes and in combination with a preset urgency weight function. The high-frequency control unit is used to combine the high-frequency environmental state with the predictive induction intensity, calculate the dynamically weighted reward function, and use a deep reinforcement learning model to generate the final execution action based on the dynamically weighted reward function.

[0004] Preferably, the low-frequency plant status includes leaf area index or developmental stage from plant canopy image analysis; the execution summary includes total daily energy consumption, average stress integral, or daily cumulative light intensity.

[0005] Preferably, the process by which the cross-scale coupling unit calculates the predictive induced intensity is as follows: Multiply the probability value of each day in the probability distribution of the key development node by the weight value corresponding to the preset urgency weight function; Perform a weighted summation on the product of the products; The result of the weighted summation is normalized to obtain the predictive induction intensity.

[0006] Preferably, the process by which the high-frequency control unit calculates the dynamically weighted reward function is as follows: Determine the basic reward and the incentive reward; Using the predictive induction intensity as a dynamic weight, the base reward and the induction reward are linearly interpolated to generate the dynamically weighted reward function.

[0007] Preferably, the basic reward is derived from the estimated instantaneous growth rate and the instantaneous energy consumption.

[0008] Preferably, the induced reward is calculated based on the degree of compliance with a preset set of induced rules.

[0009] Preferably, the high-frequency control unit is further used for: Physiological inertial parameters are calculated using an exponential moving average based on the rate of change of the high-frequency environmental state.

[0010] Preferably, the high-frequency control unit is further used for: The constraint factor is calculated by applying an exponential decay function based on the physiological inertia parameters.

[0011] Preferably, the high-frequency control unit is further used for: The deep reinforcement learning model is used to generate high-frequency original actions; The final execution action is generated by combining the high-frequency original action with the constraint factor and through smoothing constraints.

[0012] Preferably, the runtime data includes the instantaneous energy consumption, the estimated instantaneous growth rate, or the physiological inertia parameter.

[0013] Compared with the prior art, the present invention has the following beneficial effects: 1. This system transforms the probability distribution of future key developmental nodes into predictive induction intensity, and uses this as a dynamic weight to dynamically interpolate the basic reward representing short-term efficiency and the induction reward representing long-term goals, thus solving the conflict problem of difficulty in coordinating short-term operational goals and long-term growth goals in plant factories. 2. This system, by setting up a state aggregation unit, processes the runtime data during high-frequency control execution into an execution summary and feeds it back to the low-frequency strategy unit, thus constructing a feedback loop from high-frequency execution to low-frequency prediction. This overcomes the problem of model-reality deviation caused by open-loop prediction in traditional long-term strategies. 3. This system calculates physiological inertia parameters based on the rate of change of environmental conditions through a high-frequency control unit, and solves for constraint factors. It then smooths the original actions generated by the deep reinforcement learning model, effectively avoiding drastic changes in control actions, reducing environmental mutation stress on plants, and ensuring the stability of control. 4. The overall architecture of this system uses a time-series prediction network to make probabilistic predictions of key developmental nodes in the future, and combines it with an urgency weight function to achieve forward-looking regulation. This enables the system to no longer passively respond to the current state, but to achieve predictive, proactive, and multi-objective collaborative control of the plant growth environment. Attached Figure Description

[0014] The present invention will be further explained below with reference to the accompanying drawings and embodiments: Figure 1 This is a structural diagram of the system of the present invention. Detailed Implementation

[0015] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to specific embodiments.

[0016] Example 1: Please see Figure 1 A deep learning-based plant factory plant growth environment control system includes: The data acquisition unit is used to collect high-frequency environmental status data and low-frequency plant status and runtime data. The state aggregation unit is used to aggregate runtime data to obtain an execution summary; The low-frequency strategy unit is used to receive low-frequency plant status and execution summary, and calculate the probability distribution of future key developmental nodes based on the time-series prediction network. A cross-scale coupling unit is used to calculate the predictive induction intensity based on the probability distribution of key developmental nodes and a preset urgency weight function. The high-frequency control unit is used to combine the high-frequency environmental state and the predictive induction intensity to calculate the dynamically weighted reward function, and then use a deep reinforcement learning model to generate the final execution action based on the dynamically weighted reward function.

[0017] A deep learning-based plant factory plant growth environment control system aims to achieve predictive, proactive, and multi-objective coordinated regulation of the plant growth environment. In this embodiment, the system is divided into five collaborative core units, which constitute a complete data and control closed loop. Data acquisition unit: This unit is the foundation of the system's perception and its purpose is to acquire multi-source heterogeneous data required for system operation in real time; High-frequency environmental status: refers to real-time environmental parameters at the minute or second level; in this embodiment, it is achieved through sensors deployed within the plant factory, such as temperature and humidity sensors. Sensors, quantum sensors, and power meters are used to collect data, such as instantaneous temperature and instantaneous pulse. Concentration, instantaneous energy consumption ; Low-frequency plant status refers to phenotypic parameters on a daily or weekly basis that reflect the long-term growth status of plants. In this embodiment, it is obtained by capturing images of the plant canopy with an industrial camera and analyzing them using image processing algorithms, such as leaf area index. Or the current stage of development; Runtime data: refers to intermediate process data generated by the high-frequency control unit during execution, which is used for subsequent aggregation analysis; Low-Frequency Strategy (LFS): This unit is the long-term strategy hub of the system, and its purpose is to predict key developmental nodes in the future of the plant based on historical data and the current state. The probability of occurrence; This unit receives two key inputs: low-frequency plant status provided by the data acquisition unit and an execution summary of the previous cycle provided by the status aggregation unit; This unit is based on a temporal prediction network for computation; in this embodiment, the network is preferably a Long Short-Term Memory (LSTM) network, because it is good at capturing long temporal dependencies in the biological growth process. The output of this unit is not a specific date, but a probability distribution of a key developmental node. For example, a 14-dimensional vector representing the events that occur each day from the next 1 to 14 days. The probability; this probabilistic output, implemented using Softmax, can better handle the inherent uncertainty of organisms; The specific calculation process is as follows: As shown, where, It is a low-frequency state vector that includes low-frequency plant states and execution summary; The hidden state of the LSTM network in the previous time step is obtained by recursive calculation within the LSTM unit; Cross-Scale Coupling Module (CCM): This unit is the core hub connecting the Long Scale Strategy Unit (LFS) and the High Frequency Control Unit (HFCU). Its purpose is to transform the high-dimensional, low-frequency probability distribution output by the LFS into high-frequency, scalar command signals that the HFC can use in real time. The probability distribution of critical developmental nodes received by this unit from the LFS solution. ; This unit incorporates a pre-defined urgency weighting function. The urgency weighting function refers to a function based on prior biological knowledge, whose purpose is to weight upcoming events. Assign higher weights; for example, it can be designed as ,in, It is an integer index representing the number of future days; the technical consideration of this design is that it ensures the control system's ability to detect nearby... Comparison of distant More sensitive; This unit calculates a key intermediate parameter—the predictive induction intensity. This parameter is a scalar in the range [0,1], which quantifies the urgency and necessity of executing the current induction strategy. High-Frequency Control (HFC): This unit is the high-frequency control center of the system, and its purpose is to optimize environmental control actions in real time based on the guidance of predictive induction intensity. This unit receives high-frequency environmental conditions provided by the data acquisition unit and predictive induced intensity calculated by the cross-scale coupling unit. ; The core of this unit is solving the dynamically weighted reward function. A dynamically weighted reward function refers to a function whose internal structure and objective are adjusted according to... Real-time changing reward signals; This unit utilizes a deep reinforcement learning model, preferably a deep deterministic policy gradient (DDPG) model in this embodiment, which is trained and used to make decisions based on this dynamically weighted reward function to generate the final action to be executed. State aggregation unit: This unit is the key to realizing the feedback loop. Its purpose is to summarize and refine the runtime data of HFC during high-frequency execution, so that LFS can correct its long-term prediction in the next cycle. This unit processes runtime data, such as the instantaneous energy consumption generated by HFC. Physiological inertial parameters Aggregation processes, such as summation, mean calculation, and maximum value calculation, are performed within a low-frequency period. The output of this unit is an execution summary. ,For example ; The execution summary is then fed back to the low-frequency policy unit as one of the inputs for its next cycle prediction; The five units mentioned above work together to form a predictive-execution-correction dual-loop coupled architecture; LFS provides long-term predictive guidance for HFC through the cross-scale coupled unit CCM. The signal addresses the limitation of HFC in optimizing only short-term objectives; the runtime data of HFC is fed back to LFS through the state aggregation unit and data acquisition unit, which solves the model-reality bias problem encountered by LFS in open-loop prediction without feedback correction. By constructing a collaborative architecture of units such as LFS, CCM, and HFC, this system overcomes the multi-scale conflict between short-term energy consumption and long-term yield in existing technologies. It no longer passively waits for the plant to enter the next stage, but instead uses probabilistic predictions from LFS and feedforward signals from CCM to achieve precise control over key plant developmental nodes. The predictive and proactive induction, along with the feedback loop from HFC to LFS, ensures the robustness of the control strategy to biological uncertainties.

[0018] Example 2: Low-frequency plant status includes leaf area index or developmental stage from plant canopy image analysis; executive summary includes total daily energy consumption or average stress integral or cumulative daily light intensity.

[0019] This embodiment is a further specification of the input data source in Embodiment 1; Low-frequency plant states, the purpose of which is to provide LFS with intuitive evidence of long-term plant growth accumulation; in this embodiment, it specifically includes: Leaf area index The canopy is captured by the camera in the data acquisition unit, and the image segmentation algorithm is used to calculate the result. It is a key indicator of photosynthetic capacity and biomass accumulation; Developmental stages: Plant morphology is identified using image analysis algorithms to determine its biological stage, such as vegetative growth or flower bud differentiation; The execution summary aims to provide a quantitative summary of the HFC execution performance in the previous cycle for LFS; in this embodiment, it specifically includes: Total daily energy consumption: Instantaneous energy consumption in runtime data, calculated by the state aggregation unit. The result is obtained by summing the integrals over the entire day. Average stress integral: Calculated by averaging the physiological inertial parameters in the runtime data over the entire day using the state aggregation unit, used to quantify the environmental stress experienced by the plant throughout the day; Daily cumulative illumination: The daily cumulative amount of photosynthetically active radiation (PAR) is obtained by integrating the readings of the photon sensor over the entire day using the state aggregation unit; By clearly defining the input plant phenotypes and execution feedback for LFS, the input quality of the long-term prediction model is ensured; Visual features make predictions closer to the actual biological state of plants, while using performance summaries such as energy consumption and stress allows LFS predictions to be dynamically corrected based on actual control effects, thus improving accuracy. Predictive accuracy and adaptability.

[0020] Example 3: The process of calculating the predictive induced intensity using a cross-scale coupled unit is as follows: Step 1: Multiply the probability values ​​of each day in the probability distribution of key development nodes by the weight values ​​corresponding to the preset urgency weight function; Step 2: Perform a weighted summation on the product of the multiplications; Step 3: Normalize the weighted summation result to obtain the predictive induction intensity.

[0021] This embodiment demonstrates how the cross-scale coupling unit CCM in Embodiment 1 calculates the predictive induced intensity. A detailed explanation of the specific algorithm steps; this process is the key to the transformation from prediction to control in this invention; The solution process follows these steps: Multiplication: Obtain the probability distribution vector of key developmental nodes output by the low-frequency policy unit LFS. Its elements express In the future The probability of this happening; simultaneously, obtain the preset urgency weight function. As mentioned above Multiply the corresponding elements of the two to obtain ; Weighted summation: For the summation in step 1 Prediction window, for example Summing all products within a day, i.e. ; Normalization: Apply a standard normalization function to the weighted summation result in step 2. For example, the Sigmoid function or Min-Max scaling can be used to ensure that the output value falls within the [0,1] range; Based on this, predictive induction intensity The calculation formula is specified in this embodiment as follows: ; in, : In the future The probability of an event occurring; determined by the low-frequency strategy unit LFS based on the formula. Calculated; Urgency weighting function; hyperparameters designed based on prior biological knowledge, such as... The aim is to amplify the control weights of adjacent KDPs; Normalization function, standard mathematical function; By employing this original computational method of urgency-weighted summation, this invention reduces the dimensionality of LFS's high-dimensional, uncertain probability predictions and transforms them into a scalar control signal that can accurately quantify the urgency-inducing effect. This signal can more robustly guide HFC when and to what extent to begin switching its control target, a guidance method superior to the traditional argmax or expectation value method.

[0022] Example 4: The process of the high-frequency control unit solving the dynamically weighted reward function is as follows: Determine the basic reward and the incentive reward; By using predictive induction intensity as a dynamic weight, a linear interpolation is performed on the base reward and the induction reward to generate a dynamically weighted reward function.

[0023] The basic reward is derived from the estimated instantaneous growth rate and the instantaneous energy consumption.

[0024] The incentive reward is calculated based on the degree of compliance with the preset incentive rule set.

[0025] This embodiment demonstrates how the high-frequency control unit (HFC) in Embodiment 1 utilizes... Solving the dynamically weighted reward function Specific explanation; this is the core mechanism for resolving multi-scale target conflicts; Dynamically weighted reward function HFC's dynamically weighted reward function It is generated by linear interpolation, or a convex combination, of two sub-objectives—basic reward and induced reward; predictive induced strength. It is innovatively used as the dynamic weight for this interpolation; The calculation formula is specified in this embodiment as follows: ; in, : Predictive induced intensity; based on the formula of the cross-scale coupling unit CCM Calculated; Basic rewards; Inducing rewards; The logic behind this formula is: when In the distant past ( ), HFC's objective automatically focuses on optimizing basic rewards. when Approaching time ( ), HFC's objective automatically switches to optimizing incentive rewards; Basic Rewards Basic Rewards Its purpose is to optimize the daily operating efficiency of plant factories, i.e., the short-term goal; it is based on the estimated instantaneous growth rate and the instantaneous energy consumption calculation. The calculation formula is specified in this embodiment as follows: ; in, Instantaneous growth rate estimate; based on high-frequency environmental conditions using a surrogate model, such as a pre-trained multiple linear regression model. Or a simplified photosynthesis mechanism model, which can be estimated in real time; among which... Instantaneous temperature This refers to the instantaneous carbon dioxide concentration. Instantaneous solar radiation intensity; The intercept constant is... These are the regression coefficients corresponding to instantaneous temperature, instantaneous carbon dioxide concentration, and instantaneous solar radiation intensity, respectively; these coefficients are all calibrated using historical data, for example, a set of possible calibration values. The use of a proxy model instead of a complete mechanistic model is to ensure that the HFC unit can perform high-frequency real-time reward calculations to meet the decision speed requirements of deep reinforcement learning models. Instantaneous energy consumption; measured in real time by the data acquisition unit; : Calibrate hyperparameters; use historical operating data for regression analysis or reinforcement learning training to calibrate them, in order to balance growth benefits and energy costs; for example, data from historical high-yield batches can be used for multiple regression analysis. Photosynthetic contribution and The impact of energy consumption costs on the final dry weight, in order to set As an initial baseline, it is then handed over to the deep reinforcement learning model for online fine-tuning based on the actual running data; Induced rewards Induced rewards Its purpose is to The optimal strategy, i.e. the long-term goal, is strictly implemented during the induction period; it is calculated based on the degree of adherence to the preset set of induction rules. Preset induction rule set refers to a set of triggering rules defined based on the knowledge of agronomic experts. The set of environmental conditions that must be met; for example, the set of rules for inducing flowering in a certain plant. It can be defined as: {daytime temperature} The temperature must be between [24, 26]°C, and the photoperiod must be [missing information]. The threshold for this rule set, such as [24,26]°C, is determined based on the agronomic standards of the target crop or statistical data from historical high-yield experiments. The calculation is specified in this embodiment as follows ,in, It is a function used to quantize the current state. and actions right The degree of compliance, for example, 1 if the rule is met, and -1 if it is not met; Induced reward calibration coefficient, which is a hyperparameter set empirically or calibrated during training; By Long-term prediction signals as The present invention addresses the issue of short-term energy consumption dynamic weighting at the reward function level of the control algorithm by adjusting the weighting of short-term control objectives. Representative and long-term inducement The system eliminates the need for manual switching phases, achieving a smooth, automatic, and proactive transition between energy-saving and induced modes, thereby maximizing energy efficiency and productivity throughout the entire growth cycle. The co-optimal for induction success rate.

[0026] Example 5: The high-frequency control unit is also used for: Physiological inertial parameters are calculated using exponential moving averages based on the rate of change of high-frequency environmental conditions.

[0027] The high-frequency control unit is also used for: The constraint factor is calculated by applying an exponential decay function based on physiological inertia parameters.

[0028] The high-frequency control unit is also used for: Generate high-frequency primitive actions using deep reinforcement learning models; By combining high-frequency original actions and constraint factors, the final execution action is generated through smoothing constraints.

[0029] This embodiment is a detailed description of the motion constraint and smoothing mechanism designed by the high-frequency control unit HFC in Embodiment 1 to prevent drastic changes in control actions; this mechanism is crucial for protecting plants from drastic environmental stresses. Calculate physiological inertial parameters To prevent HFCDDPG from outputting aggressive control actions, this invention introduces an original parameter—the physiological inertia parameter. ; Physiological inertial parameters It refers to a parameter used to quantify the cumulative stress experienced by plants due to the rate of environmental change, rather than quantifying the environmental state itself; This unit is based on the rate of change of high-frequency environmental conditions, such as the rate of temperature change. Calculated using the Exponential Moving Average (EMA) ; The calculation formula is specified in this embodiment as follows: ; in, Physiological inertial parameters, cumulative stress values; : The rate of change of high-frequency environmental conditions; calculated in real time from sensor data of the data acquisition unit, such as ; This invention uses the square of the rate of change as the smoothing object for EMA, which is physically closer to stress or energy and can capture violent fluctuations more sensitively. Calibrate hyperparameters, forgetting factor, and sensitivity coefficient; based on experimental datasets of plant stress responses, e.g., at different rates of environmental change. The following tests the corresponding physiological stress indicators of plants. Determined through data fitting, such as the least squares method, to ensure It can accurately reflect real biological stress; for example, it can set... Value stress indicators The units of measurement are unified; then, through fitting experimental data, if it is found that the influence weight of minute-level stress on the current moment is 90%, then a forgetting factor can be set. ; Solving constraint factors The HFC unit then uses the obtained physiological inertia parameters... The constraint factor is calculated by applying an exponentially decaying function. ; constraint factors It refers to a scalar in the range [0,1], which serves as a constraint coefficient for the original HFC action; The calculation formula is specified in this embodiment as follows: ; in, Physiological inertia; Constraint sensitivity hyperparameter; calibrated based on the control system's safe operation specifications and plant tolerance experimental data. The dimensions are The reciprocal of the dimensions to ensure that the exponent term is dimensionless; for example, according to safety specifications, when cumulative stress... When the preset biosafety threshold is reached, the constraint factor is required. Must decay to From this, we can solve the problem in reverse. ;in, This refers to the biosafety threshold that a plant can tolerate in terms of cumulative stress. This threshold can be determined based on experimental data from specific crops and environments; for example, it can be based on historical data. Set to 5.0; The logic of this formula lies in: when cumulative stress... At a very high level, Constraints are enhanced; when the stress is very low, Restriction lifted; Generate the final execution action The HFC unit utilizes the policy network of its deep reinforcement learning model DDPG to generate high-frequency primitive actions. ; High-frequency primitive motion This refers to the reward function that the DRL model believes can maximize the dynamically weighted reward. Ideally, it should immediately adjust the temperature from 20°C to 25°C. The HFC unit combines this high-frequency original action With the constraint factor of the solution The final execution action is generated through a smoothing constraint formula. ; The calculation formula is specified in this embodiment as follows: ; in, : The high-frequency raw action of the DRL output; Constraint factor, constraint coefficient; The final action executed in the previous moment; The logic behind this formula is: when hour, The DRL model completely takes over control; when hour, High-frequency original actions are rejected, and actions are frozen to prevent drastic changes; By employing a step-by-step mechanism of calculating stress, converting it into constraint coefficients, and applying constraints, this invention establishes a physiological safety constraint layer within the HFC. This mechanism effectively smooths out the drastic, high-frequency control actions that the DRL controller DDPG might take in pursuit of instantaneous optimal rewards, making environmental changes more consistent with the physiological inertia of plants. This significantly reduces growth stress and improves the stability and safety of the system during nonlinear transitions.

[0030] Example 6: Runtime data includes estimates of instantaneous energy consumption or instantaneous growth rate, or physiological inertia parameters.

[0031] This embodiment is a detailed description of the runtime data collected by the data acquisition unit in embodiment 1. This data is the raw material for the state aggregation unit to generate the execution summary and constitutes the feedback path from HFC to LFS. Runtime data refers to intermediate variables that HFC generates or calculates in real time during high-frequency operation, specifically including: Instantaneous energy consumption The base reward is calculated based on real-time measurements from the power meter or estimates based on equipment status. ; Instantaneous growth rate estimate Estimated in real time by a photosynthesis proxy model, used to calculate the base reward. ; Physiological inertial parameters : Calculated in real time by the HFC motion constraint module to quantify instantaneous stress; The state aggregation unit in each low-frequency cycle, such as At the end of the day, all the aforementioned runtime data from that period will be collected, for example, All within the period , , The values ​​are calculated and aggregated, such as by summation, averaging, and maximization, to generate an execution summary. ; By explicitly collecting and aggregating key performance indicators within the HFC, such as energy consumption, growth estimation, and stress, as runtime data, this invention constructs a complete and information-rich feedback loop from short-term execution effects to long-term strategy correction; this enables the low-frequency policy unit (LFS) to... The predictions are not only based on plant phenotypes, but also fully learn and adapt to the actual performance and control costs of HFCs, making the prediction and control strategies of the entire two-scale system more robust, more self-consistent, and closer to the global optimum.

[0032] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention.

Claims

1. A deep learning-based plant factory plant growing environment control system, characterized by, The method comprises the following steps: a data acquisition unit is used to collect high-frequency environmental state and low-frequency plant state and runtime data; a state aggregation unit is used to aggregate the runtime data to obtain an execution summary; a low-frequency policy unit is used to receive the low-frequency plant state and the execution summary, and calculate the probability distribution of future key development nodes based on a time series prediction network; a cross-scale coupling unit is used to calculate a predictive induction intensity based on the probability distribution of the key development nodes and in combination with a preset urgency weight function; a high-frequency control unit is used to combine the high-frequency environmental state and the predictive induction intensity to calculate a dynamically weighted reward function, and generate a final execution action based on the dynamically weighted reward function using a deep reinforcement learning model.

2. The plant factory plant growing environment control system based on deep learning according to claim 1, wherein, The low-frequency plant state includes leaf area index or development stage obtained by plant canopy image analysis; and the execution summary includes daily total energy consumption or average stress integral or daily cumulative illumination.

3. The plant factory plant growing environment control system based on deep learning according to claim 1, wherein, The cross-scale coupling unit calculates the predictive induction intensity in the following manner: multiply each day probability value in the probability distribution of the key development nodes by a weight value corresponding to the preset urgency weight function; perform weighted summation on the multiplied products; normalize the result of the weighted summation to obtain the predictive induction intensity.

4. The plant factory plant growing environment control system based on deep learning according to claim 1, wherein, The high-frequency control unit calculates the dynamically weighted reward function in the following manner: determine a basic reward and an induction reward; perform linear interpolation on the basic reward and the induction reward using the predictive induction intensity as a dynamic weight to generate the dynamically weighted reward function.

5. The plant factory plant growing environment control system based on deep learning according to claim 4, wherein, The basic reward is calculated based on an instantaneous growth rate estimate and an instantaneous energy consumption.

6. The plant factory plant growing environment control system based on deep learning according to claim 4, wherein, The induction reward is calculated based on the degree of compliance with a preset induction rule set.

7. The plant factory plant growing environment control system based on deep learning according to claim 5, wherein, The high-frequency control unit is further used to: calculate a physiological inertia parameter through exponential moving average calculation according to the rate of change of the high-frequency environmental state.

8. The plant factory plant growing environment control system based on deep learning according to claim 7, wherein, The high-frequency control unit is further used to: calculate a constraint factor by applying an exponential decay function according to the physiological inertia parameter.

9. The plant factory plant growing environment control system based on deep learning according to claim 8, wherein, The high-frequency control unit is further used to: generate a high-frequency original action using the deep reinforcement learning model; generate the final execution action through smoothing constraint in combination with the high-frequency original action and the constraint factor.

10. The plant factory plant growing environment control system based on deep learning according to claim 7, wherein, The runtime data includes the instantaneous energy consumption, the instantaneous growth rate estimate, or the physiological inertia parameter.