Green vegetable growth optimization system and method based on SARSA reinforcement learning
By dynamically adjusting the electrical conductivity (EC) based on the SARSA reinforcement learning method and combining it with the cumulative radiant heat product (TEP) and fresh weight data, the real-time and accuracy issues of conductivity adjustment in traditional green vegetable cultivation were solved, achieving real-time automated optimization of the green vegetable growth environment and increasing yield.
Patent Information
- Application Number
- CN202510366509.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-26
- Publication Date
- 2025-09-26
AI Technical Summary
In the existing vegetable cultivation process, the adjustment of electrical conductivity (EC) mainly relies on experience and manual operation, which makes it difficult to control the growth environment of the vegetables in real time and accurately, resulting in the impact on growth efficiency and quality.
A SARSA reinforcement learning method was adopted to dynamically adjust the electrical conductivity (EC) through the state acquisition unit, growth model unit, state definition unit, action selection unit, reward function unit and SARSA update unit. Combined with the cumulative radiant heat product (TEP) and fresh weight data, the ε-greedy strategy and SARSA algorithm were used to optimize the Q-value table to achieve real-time automatic optimization of the growth environment.
It achieves real-time automated optimization of growth environment parameters, significantly shortens the vegetable growth cycle, improves yield stability, accurately controls the fresh weight of vegetables, reduces the incidence of abnormal conditions, and ensures stable operation of the system under different climatic conditions.
Smart Images

Figure CN120706734A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of agricultural technology, and in particular to a vegetable growth optimization system and method based on SARSA reinforcement learning. Background Art
[0002] In existing vegetable cultivation processes, electrical conductivity (EC) adjustment relies primarily on experience and manual manipulation, making it difficult to accurately control the vegetable growth environment in real time, thus affecting growth efficiency and quality. With the development of machine learning technology, reinforcement learning methods can be used to intelligently optimize the vegetable growth process.
[0003] Patent application document CN117152739A discloses a data-driven growth optimization system and optimization method. It uses an encoder model to encode the nutrient solution temperature, photon density, air temperature, and air humidity at multiple time points, and constructs a Gaussian density map from the four obtained feature vectors. A convolutional neural network is then used to extract features from images of hydroponic lettuce at these multiple time points. This allows a Gaussian mixture model to be constructed based on the feature vectors at each time point and the Gaussian density map to achieve consistency between the various response positions of the Gaussian density map and its response feature vectors, thereby achieving a certain degree of matching between the response range of the response feature vector and the target scale of the Gaussian density map. However, this patent fails to fully resolve existing technical problems and fails to meet the requirements of the present invention. Summary of the Invention
[0004] In view of the defects in the prior art, the purpose of the present invention is to provide a vegetable growth optimization system and method based on SARSA reinforcement learning.
[0005] The vegetable growth optimization system based on SARSA reinforcement learning provided by the present invention includes:
[0006] The status acquisition unit obtains the current cumulative radiant heat product TEP and fresh weight data of green vegetables;
[0007] Growth model unit, baseline fresh weight estimated according to the TEP;
[0008] The state definition unit determines the growth state based on the percentage difference between the actual fresh weight and the benchmark fresh weight;
[0009] The action selection unit selects the action to adjust the conductivity based on the ε-greedy strategy according to the growth state;
[0010] Reward function unit, which outputs reward value according to growth status;
[0011] SARSA update unit, which updates the Q value table using the SARSA algorithm;
[0012] The training unit optimizes the Q-value table through multi-cycle interactive training, dynamically decays the exploration rate ε, and adjusts the learning rate and discount factor based on historical rewards.
[0013] Preferably, the formula for estimating the benchmark fresh weight according to the TEP is:
[0014]
[0015] Where TEP stands for cumulative radiant heat product.
[0016] Preferably, the formula for updating the Q value table using the SARSA algorithm is:
[0017] Q(s,a)←Q(s,a)+α[r+γ(s′,a′)-Q(s,a)]
[0018] Among them, s represents the current state, a represents the executed action; s′ represents the future state, a′ represents the future executed action; r is the immediate reward of the current round; γ is the discount factor used to adjust the value of future rewards; α is the learning rate used to update the Q value.
[0019] Preferably, the dynamic attenuation formula based on the ε-greedy strategy is:
[0020] ε=ε0·e -k·t
[0021] The initial value of ε0 is 0.3, k = 0.01, and it decays once every 100 cycles during training, and t is the number of decays.
[0022] Preferably, the training unit includes:
[0023] The parameter adaptation module reduces the learning rate α to 90% of the original value when the Q value fluctuates by more than 10%;
[0024] Convergence judgment module, if the maximum change of the Q value table is less than 0.01 within 20 consecutive cycles, it is judged to be converged and training is terminated.
[0025] The vegetable growth optimization method based on SARSA reinforcement learning provided by the present invention includes:
[0026] Step 1: Obtain the current cumulative radiant heat product TEP and vegetable fresh weight data;
[0027] Step 2: estimating baseline fresh weight based on the TEP;
[0028] Step 3: Determine the growth status based on the percentage difference between the actual fresh weight and the benchmark fresh weight;
[0029] Step 4: Select the action to adjust the conductivity based on the ε-greedy strategy according to the growth state;
[0030] Step 5: Output the reward value according to the growth status;
[0031] Step 6: Update the Q value table using the SARSA algorithm;
[0032] Step 7: Optimize the Q-value table through multi-cycle interactive training, dynamically decay the exploration rate ε, and adjust the learning rate and discount factor based on historical rewards.
[0033] Preferably, the formula for estimating the benchmark fresh weight according to the TEP is:
[0034]
[0035] Where TEP stands for cumulative radiant heat product.
[0036] Preferably, the formula for updating the Q value table using the SARSA algorithm is:
[0037] Q(s,a)←Q(s,a)+α[r+γ(s′,a′)-Q(s,a)]
[0038] Among them, s represents the current state, a represents the executed action; s′ represents the future state, a′ represents the future executed action; r is the immediate reward of the current round; γ is the discount factor used to adjust the value of future rewards; α is the learning rate used to update the Q value.
[0039] Preferably, the dynamic attenuation formula based on the ε-greedy strategy is:
[0040] ε=ε0·e -k·t
[0041] The initial value of ε0 is 0.3, k = 0.01, and it decays once every 100 cycles during training, and t is the number of decays.
[0042] Preferably, the step 7 includes:
[0043] When the Q value fluctuates by more than 10%, the learning rate α is reduced to 90% of the original value;
[0044] If the maximum change in the Q-value table is less than 0.01 within 20 consecutive cycles, it is considered converged and the training is terminated.
[0045] Compared with the prior art, the present invention has the following beneficial effects:
[0046] 1. By introducing the SARSA reinforcement learning algorithm to dynamically adjust the electrical conductivity (EC) value, the problem of low efficiency and poor accuracy caused by traditional reliance on manual experience to adjust EC is solved. This achieves real-time automated optimization of growth environment parameters, significantly shortening the vegetable growth cycle and improving yield stability.
[0047] 2. By integrating the real-time feedback mechanism of cumulative thermal radiant product (TEP) and fresh weight data, the disconnect between environmental parameters and growth status in traditional planting is resolved. This enables precise regulation based on plant physiological indicators, reducing the deviation rate between the fresh weight of green vegetables and the target value to within 5%.
[0048] 3. By designing a multi-level EC adjustment action space based on the ε-greedy strategy, the problem of insufficient adaptability of a single adjustment strategy was solved, and flexible adjustment of the EC value (±0.05-0.1mS / cm) was achieved, avoiding growth stress caused by EC mutations.
[0049] 4. By establishing a state classification mechanism (normal, alert, warning) and a differentiated reward function, this approach addresses the problem of delayed response to growth anomalies in traditional methods, enabling early warning and proactive intervention of growth risks, reducing the incidence of abnormal states by over 40%.
[0050] 5. By combining the benchmark fresh weight model with the Q-value iterative update method, the problem that the static model cannot adapt to dynamic changes in the environment is solved, and the online self-optimization of the growth model is realized, ensuring that the system can operate stably under different climatic conditions. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] Other features, objects and advantages of the present invention will become more apparent upon reading the detailed description of non-limiting embodiments with reference to the following drawings:
[0052] Figure 1 This is the workflow diagram of the vegetable growth optimization system based on SARSA reinforcement learning;
[0053] Figure 2 This is a flow chart of the vegetable growth optimization method based on SARSA reinforcement learning. DETAILED DESCRIPTION
[0054] The present invention will be described in detail below with reference to specific embodiments. The following examples will help those skilled in the art to further understand the present invention, but are not intended to limit the present invention in any form. It should be noted that, for those skilled in the art, several changes and improvements can be made without departing from the scope of the present invention. These all fall within the scope of protection of the present invention.
[0055] Example 1
[0056] like Figure 1As shown in the figure, the vegetable growth optimization system based on SARSA reinforcement learning includes a state acquisition unit, a growth model unit, a state definition unit, an action selection unit, a reward function unit, a SARSA update unit and a training unit.
[0057] The status acquisition unit obtains the current cumulative radiant heat product TEP and fresh weight information of vegetables in real time.
[0058] The cumulative thermal product (Product of thermal effectiveness and PAR, TEP) is a variable that comprehensively reflects the temperature and light accumulation, and is numerically equal to the product of relative thermal effectiveness (RET) and photosynthetically active radiation (PAR).
[0059]
[0060] Where T is the sensor sampling time; RET(T) is the relative thermal effect corresponding to time T, which is between 0 and 1 and dimensionless; T b The lower limit of the growth temperature of leaf lettuce is 7℃; T ob The lower limit of the optimum growth temperature is 18℃; T ou The upper limit of the optimum growth temperature is 25℃; T m The upper limit of the growth temperature of leaf lettuce is 35℃. The dimension of PAR is μmol / m-2·s, which is directly measured by the sensor. Since the sensor records the reading every 5 minutes, the TEP within 5 minutes is 5min Time conversion needs to be considered, as calculated by the following formula:
[0061] TEP 5min =RET(T)×PAR×300
[0062] Therefore, the daily cumulative radiant heat product (DTEP) is TEP 5min The accumulation of:
[0063]
[0064] For growth model units, the baseline fresh weight is estimated based on TEP using the formula:
[0065]
[0066] The state definition unit determines the growth state based on the percentage difference between the actual fresh weight and the benchmark fresh weight, diff_percent. The formula is:
[0067]
[0068] fresh_weight indicates actual fresh weight, and base_weight indicates base fresh weight;
[0069] The specific standards are:
[0070] Normal: diff_percent < 5%;
[0071] Alert: 5% <diff_percent<20%;
[0072] Warning: diff_percent > 20%;
[0073] The action selection unit selects the action to adjust EC based on the ε-greedy strategy. The process is as follows
[0074] 1. Formal definition of ε-greedy strategy:
[0075] Randomly select actions (exploration) with probability ε (usually set to 0.1-0.3);
[0076] Select the optimal action a*=argmaxaQ(s,a) in the current Q-value table with probability 1-ε.
[0077] 2. Action selection process:
[0078] Step 1: Get the Q values of all actions corresponding to the current state s;
[0079] Step 2: Generate a random number r∈[0,1]r∈[0,1];
[0080] If r < ε, randomly select an action from the action set {increase EC 0.1, increase EC 0.05, decrease EC 0.05, decrease EC 0.1, EC unchanged};
[0081] If r ≥ ε, select the action with the largest Q value (if multiple actions have the same Q value, they are executed according to the priority of "EC unchanged > fine-tuning > large-scale adjustment");
[0082] Step 3: After executing the selected action, update the Q-value table based on the environment feedback.
[0083] 3. Strategy optimization design:
[0084] Dynamic attenuation ε value (such as ε=ε0·e -k·t , ε0=0.3,k=0.01), initially focusing on exploration and later gradually leaning towards utilization;
[0085] Set a basic probability weight (+10%) for the "EC unchanged" action to avoid frequent adjustments causing environmental fluctuations.
[0086] The reward function unit calculates the reward based on the growth status. The specific reward value is:
[0087] When the status is 'normal', the reward value is -1;
[0088] When the status is 'alert', the reward value is -20;
[0089] When the status is 'Warning', the reward value is -100.
[0090] SARSA update unit uses the SARSA algorithm to update the Q value. The formula is:
[0091] Q(s,a)←Q(s,a)+α[r+γ(s′,a′)-Q(s,a)]
[0092] s represents the current state, a represents the executed action; s′ represents the future state, a′ represents the future executed action; r represents the immediate reward of the current round; γ is the discount factor used to adjust the value of future rewards; α is the learning rate used to update the Q value;
[0093] The training unit gradually optimizes the growth process through the continuous interaction between the agent and the environment over multiple cycles. The process is as follows:
[0094] 1. Training process
[0095] Step 1: Initialization;
[0096] Construct the Q value table Q(s,a) and initialize it to a zero matrix;
[0097] Set hyperparameters: learning rate α = 0.1, discount factor γ = 0.9, initial exploration rate ε = 0.3.
[0098] Step 2: Single cycle training;
[0099] 1. Environment reset: Initialize the vegetable growth environment (TEP = 0, fresh weight = initial value);
[0100] 2. State observation: Get the current state s t (calculate diff_percent based on TEP and fresh weight);
[0101] 3. Action selection: Select action a according to the ε-greedy strategy t (See action selection unit);
[0102] 4. Execute action: adjust EC value and simulate environmental feedback to get new state s t+1 With instant rewards t ;
[0103] 5. Q value update: Use SARSA algorithm to update Q(st ,a t ), the formula is:
[0104] Q(s t ,a t )←Q(s t ,a t )+α[r t +γQ(s t+1 ,a t+1 )-Q(s t ,a t )]
[0105] State transfer: s t ←s t+1 , repeat steps 2.3 to 2.5 until the terminal state is reached (fresh weight ≥ target value).
[0106] Step 3: Multi-cycle iteration;
[0107] 1. Repeat step 2 for a total of N training cycles (default N = 1000).
[0108] 2. Decay exploration rate every 100 cycles: ε←ε·e -0.01 .
[0109] 3. Optimize the process;
[0110] Parameter Adaptation:
[0111] 1. Dynamically adjust the learning rate α. When the Q value fluctuation range is greater than 10%, reduce α to 90% of the original value.
[0112] 2. Adjust the discount factor γ based on the historical reward mean. If the reward is continuously negative, increase γ (upper limit 0.95) to enhance long-term benefits.
[0113] Convergence judgment:
[0114] 1. When the maximum change in the Q table is less than 0.01 within 20 consecutive cycles, it is considered converged and the training is terminated;
[0115] 2. If convergence has not yet occurred but the maximum number of cycles N has been reached, save the current Q-table as the suboptimal strategy.
[0116] Model iteration:
[0117] 1. Regularly inject noise actions (with a probability of 5%) to verify the robustness of the strategy;
[0118] 2. Fine-tune the baseline fresh weight model parameters based on actual planting data.
[0119] At each step, the agent observes the state, chooses an action based on the policy, receives a reward, and updates the Q-value accordingly. This process continues until it reaches a terminal state (when the fresh weight exceeds the target weight).
[0120] The present invention optimizes the growth process of green vegetables through the coordinated work of the above units, improves the yield and quality of green vegetables, and contributes to the sustainable development of agriculture.
[0121] Example 2
[0122] The present invention also provides a vegetable growth optimization method based on SARSA reinforcement learning, comprising: step 1: obtaining current cumulative radiant heat product TEP and vegetable fresh weight data; step 2: estimating a benchmark fresh weight based on the TEP; step 3: determining a growth state based on a percentage difference between the actual fresh weight and the benchmark fresh weight; step 4: selecting an action for adjusting conductivity based on an ε-greedy strategy according to the growth state; step 5: outputting a reward value according to the growth state; step 6: updating a Q-value table using a SARSA algorithm; and step 7: optimizing the Q-value table through multi-cycle interactive training, dynamically attenuating the exploration rate ε, and adjusting the learning rate and discount factor according to historical rewards.
[0123] The formula for estimating the baseline fresh weight based on the TEP is:
[0124]
[0125] Where TEP stands for cumulative radiant heat product.
[0126] The formula for updating the Q value table using the SARSA algorithm is:
[0127] Q(s,a)←Q(s,a)+α[r+γ(s′,a′)-Q(s,a)]
[0128] Among them, s represents the current state, a represents the executed action; s′ represents the future state, a′ represents the future executed action; r is the immediate reward of the current round; γ is the discount factor used to adjust the value of future rewards; α is the learning rate used to update the Q value.
[0129] The dynamic attenuation formula based on the ε-greedy strategy is:
[0130] ε=ε0·e -k·t
[0131] The initial value of ε0 is 0.3, k = 0.01, and it decays once every 100 cycles during training, and t is the number of decays.
[0132] The step 7 comprises:
[0133] When the Q value fluctuates by more than 10%, the learning rate α is reduced to 90% of the original value;
[0134] If the maximum change in the Q-value table is less than 0.01 within 20 consecutive cycles, it is considered converged and the training is terminated.
[0135] Those skilled in the art will appreciate that, in addition to implementing the system, device, and various modules provided by the present invention in purely computer-readable program code, it is entirely possible to implement the same program in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, embedded microcontrollers, and the like by logically programming the method steps. Therefore, the system, device, and various modules provided by the present invention can be considered a hardware component, and the modules included therein for implementing various programs can also be considered structures within the hardware component; the modules for implementing various functions can also be considered both software programs for implementing the method and structures within the hardware component.
[0136] The above describes specific embodiments of the present invention. It should be understood that the present invention is not limited to the specific embodiments described above, and those skilled in the art may make various changes or modifications within the scope of the claims, which do not affect the essence of the present invention. The embodiments of this application and the features in the embodiments may be combined with each other in any manner unless there is a conflict.
Claims
1. A vegetable growth optimization system based on SARSA reinforcement learning, characterized in that: include: The status acquisition unit obtains the current cumulative radiant heat product TEP and fresh weight data of green vegetables; Growth model unit, baseline fresh weight estimated according to the TEP; The state definition unit determines the growth state based on the percentage difference between the actual fresh weight and the benchmark fresh weight; The action selection unit selects the action to adjust the conductivity based on the ε-greedy strategy according to the growth state; Reward function unit, which outputs reward value according to growth status; SARSA update unit, which updates the Q value table using the SARSA algorithm; The training unit optimizes the Q-value table through multi-cycle interactive training, dynamically decays the exploration rate ε, and adjusts the learning rate and discount factor based on historical rewards.
2. The vegetable growth optimization system based on SARSA reinforcement learning according to claim 1, characterized in that The formula for estimating the baseline fresh weight based on the TEP is: Where TEP stands for cumulative radiant heat product.
3. The vegetable growth optimization system based on SARSA reinforcement learning according to claim 1, characterized in that: The formula for updating the Q value table using the SARSA algorithm is: Q(s,a)←Q(s,a)+α[r+γ(s′,a′)-Q(s,a)] Among them, s represents the current state, a represents the executed action; s′ represents the future state, a′ represents the future executed action; r is the immediate reward of the current round; γ is the discount factor used to adjust the value of future rewards; α is the learning rate used to update the Q value.
4. The vegetable growth optimization system based on SARSA reinforcement learning according to claim 1, characterized in that The dynamic attenuation formula based on the ε-greedy strategy is: ε=ε0·e -k·t The initial value of ε0 is 0.3, k = 0.01, and it decays once every 100 cycles during training, and t is the number of decays.
5. The vegetable growth optimization system based on SARSA reinforcement learning according to claim 1, characterized in that: The training unit comprises: The parameter adaptation module reduces the learning rate α to 90% of the original value when the Q value fluctuates by more than 10%; Convergence judgment module, if the maximum change of the Q value table is less than 0.01 within 20 consecutive cycles, it is judged to be converged and training is terminated.
6. A vegetable growth optimization method based on SARSA reinforcement learning, characterized in that: include: Step 1: Obtain the current cumulative radiant heat product TEP and vegetable fresh weight data; Step 2: estimating baseline fresh weight based on the TEP; Step 3: Determine the growth status based on the percentage difference between the actual fresh weight and the benchmark fresh weight; Step 4: Select the action to adjust the conductivity based on the ε-greedy strategy according to the growth state; Step 5: Output the reward value according to the growth status; Step 6: Update the Q value table using the SARSA algorithm; Step 7: Optimize the Q-value table through multi-cycle interactive training, dynamically decay the exploration rate ε, and adjust the learning rate and discount factor based on historical rewards.
7. The vegetable growth optimization method based on SARSA reinforcement learning according to claim 6, characterized in that The formula for estimating the baseline fresh weight based on the TEP is: Where TEP stands for cumulative radiant heat product.
8. The vegetable growth optimization method based on SARSA reinforcement learning according to claim 6, characterized in that The formula for updating the Q value table using the SARSA algorithm is: Q(s,a)←Q(s,a)+α[r+γ(s′,a′)-Q(s,a)] Among them, s represents the current state, a represents the executed action; s′ represents the future state, a′ represents the future executed action; r is the immediate reward of the current round; γ is the discount factor used to adjust the value of future rewards; α is the learning rate used to update the Q value.
9. The vegetable growth optimization method based on SARSA reinforcement learning according to claim 6, characterized in that The dynamic attenuation formula based on the ε-greedy strategy is: ε=ε0·e -k·t The initial value of ε0 is 0.3, k = 0.01, and it decays once every 100 cycles during training, and t is the number of decays.
10. The vegetable growth optimization method based on SARSA reinforcement learning according to claim 6, characterized in that: The step 7 comprises: When the Q value fluctuates by more than 10%, the learning rate α is reduced to 90% of the original value; If the maximum change in the Q-value table is less than 0.01 within 20 consecutive cycles, it is considered converged and the training is terminated.
Citation Information
Patent Citations
Data-driven growth optimization system and optimization method thereof
CN117152739A