Intelligent irrigation method and system for traditional Chinese medicine seedlings based on reinforcement learning
By employing a reinforcement learning-based intelligent irrigation method for Chinese medicinal herb seedlings, combined with multi-dimensional state perception and online updates, the problem of dynamically balancing seedling survival and effective component accumulation in existing technologies has been solved, achieving dynamic balance and precise irrigation decisions during the seedling cultivation process.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- INST OF MEDICINAL PLANTS YUNNAN ACAD OF AGRI SCI
- Filing Date
- 2026-04-28
- Publication Date
- 2026-07-21
Smart Images

Figure CN122431132A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of agricultural irrigation technology, specifically to a method and system for intelligent irrigation of Chinese medicinal herb seedlings based on reinforcement learning. Background Technology
[0002] In the cultivation of Chinese medicinal herbs, water management directly affects the survival rate of seedlings and the accumulation of active medicinal components. Studies have shown that moderate drought stress can induce the activation of secondary metabolic pathways in medicinal plants, promote the synthesis of active medicinal components, and improve the quality of medicinal materials. However, excessive drought stress leads to root damage. While sufficient water can ensure the survival of seedlings, it inhibits the accumulation of active medicinal components, and excessive irrigation can easily induce root rot.
[0003] Existing irrigation methods for Chinese medicinal herbs mostly employ fixed threshold-triggered irrigation or timed and quantitative irrigation. These methods have the following shortcomings: they cannot dynamically adjust irrigation strategies based on the actual growth stage of the seedlings, the current level of water stress, and the accumulation status of active ingredients; and they cannot achieve a dynamic balance between ensuring seedling survival and promoting the accumulation of active ingredients. Therefore, how to achieve an intelligent irrigation method for Chinese medicinal herb seedlings that can dynamically balance seedling survival and the enhancement of active ingredients is a technical problem that needs to be solved in this field. Summary of the Invention
[0004] This invention provides a reinforcement learning-based intelligent irrigation method for Chinese medicinal herb seedlings, addressing the technical problem that existing technologies cannot achieve a dynamic balance between seedling survival assurance and effective component enhancement in intelligent irrigation of Chinese medicinal herb seedlings. In view of the above problems, this invention provides a reinforcement learning-based intelligent irrigation method and system for Chinese medicinal herb seedlings.
[0005] In a first aspect, the present invention provides a smart irrigation method for Chinese medicinal herb seedlings based on reinforcement learning, the method comprising: Acquire the status sensing parameters of the current seedling cultivation environment, including soil moisture stress, seedling growth stage index and canopy wilting characterization value; Based on the state-aware parameters and the cumulative water stress experienced by the seedlings, the estimated value of the accumulation of effective components in the seedlings is obtained. The state-aware parameters and the effective component accumulation estimate are input into the pre-trained irrigation decision model to obtain irrigation action decisions, wherein the irrigation action decisions include irrigation duration and irrigation intensity. The irrigation action decision control system executes irrigation, and obtains updated status perception parameters and updated cumulative water stress level after irrigation. Based on the updated state perception parameters and the updated cumulative water stress level, the accumulated value of the updated effective components is obtained, and a reward value is calculated. The reward value is then used to update the irrigation decision model online.
[0006] Secondly, the present invention also provides an intelligent irrigation system for Chinese medicinal herb seedlings based on reinforcement learning, the system comprising: The state perception module is used to acquire the state perception parameters of the current seedling cultivation environment. The state perception parameters include soil moisture stress degree, seedling growth stage index and canopy wilting characterization value. The effective component assessment module is used to obtain the estimated value of the current seedling's effective component accumulation based on the state perception parameters and the current cumulative water stress experienced by the seedling. An irrigation decision module is used to input the state-aware parameters and the effective component accumulation estimate into a pre-trained irrigation decision model to obtain irrigation action decisions, wherein the irrigation action decisions include irrigation duration and irrigation intensity. The irrigation execution module is used to control the irrigation equipment to perform irrigation according to the irrigation action decision, and to obtain the updated status perception parameters and the updated cumulative water stress degree after irrigation. The online update module obtains the accumulated value of the effective components for updating based on the update status perception parameters and the cumulative water stress level, calculates the reward value, and uses the reward value to update the irrigation decision model online.
[0007] One or more technical solutions provided in this invention have at least the following technical effects or advantages: First, this invention elevates the irrigation objective for medicinal herb seedlings from simply maintaining moisture to a synergistic optimization of survival assurance and medicinal quality by introducing an estimated value of effective component accumulation as a key input feature for irrigation decisions. Compared to existing technologies that rely solely on soil moisture for threshold-triggered or timed irrigation, this invention dynamically adjusts the irrigation strategy based on the current effective component accumulation status of the seedlings, achieving a dynamic balance between promoting the synthesis of secondary metabolites and preventing excessive drought damage. This solves the technical problem of existing technologies that struggle to simultaneously ensure seedling survival and enhance effective component levels.
[0008] Second, this invention employs a reinforcement learning framework to construct an irrigation decision model and calculates multi-dimensional reward values using updated state perception parameters obtained after irrigation, thereby updating the irrigation decision model online. This closed-loop feedback mechanism enables the irrigation strategy to continuously self-optimize based on actual irrigation results, adapting to dynamic changes in different seedling varieties, growth stages, and environmental conditions. This avoids the shortcomings of fixed-rule or offline-trained irrigation decision models in terms of insufficient adaptability in complex cultivation scenarios, thus improving the accuracy and personalization of irrigation decisions.
[0009] In summary, this invention solves the technical problem that existing technologies cannot achieve a dynamic balance between seedling survival assurance and effective component enhancement in intelligent irrigation of Chinese medicinal herb seedlings through a closed-loop architecture of multi-dimensional state perception, effective component prediction, reinforcement learning decision-making, and multi-objective reward feedback. Attached Figure Description
[0010] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0011] Figure 1 This is a flowchart illustrating the intelligent irrigation method for Chinese medicinal herb seedlings based on reinforcement learning provided in the embodiments of this application. Figure 2 This is a schematic diagram of the structure of the intelligent irrigation system for Chinese medicinal herb seedlings based on reinforcement learning provided in the embodiments of this application; The components represented by each label in the attached diagram are explained as follows: Status perception module 11, Effective ingredient evaluation module 12, Irrigation decision module 13, Irrigation execution module 14, and Online update module 15. Detailed Implementation
[0012] This application provides a method and system for intelligent irrigation of Chinese medicinal herb seedlings based on reinforcement learning, which addresses the technical problem that existing technologies cannot achieve a dynamic balance between seedling survival assurance and effective component enhancement in intelligent irrigation of Chinese medicinal herb seedlings.
[0013] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application.
[0014] It should be noted that the terms "comprising" and "having" are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or server that includes a series of steps or units is not necessarily limited to those steps or units that are explicitly listed, but may include other steps or modules that are not explicitly listed or that are inherent to these processes, methods, products, or devices.
[0015] Example 1 like Figure 1 As shown, this application provides a smart irrigation method for Chinese medicinal herb seedlings based on reinforcement learning, the method comprising: S100: Obtain the status perception parameters of the current seedling cultivation environment, including soil moisture stress degree, seedling growth stage index and canopy wilting characterization value; In the cultivation of Chinese medicinal herb seedlings, accurate perception of moisture status is the foundation for intelligent irrigation decisions. Existing irrigation methods often directly use absolute soil moisture content as the trigger condition, which has two shortcomings: First, different varieties of Chinese medicinal herbs have significantly different water requirements. For example, Panax notoginseng prefers a shady and humid environment, while Astragalus membranaceus is more drought-tolerant; using a fixed moisture content threshold uniformly cannot meet the cultivation needs of multiple varieties. Second, relying solely on a single indicator of soil moisture cannot reflect the growth and development stage of the seedlings and the real-time status of the plants, easily leading to a mismatch between irrigation timing and the actual needs of the seedlings.
[0016] Step S100 in the method provided in this application embodiment includes: Acquire the current seedling cultivation environment's state-sensing parameters, including soil moisture stress, seedling growth stage index, and canopy wilting characterization value, including: The current soil volumetric moisture content is obtained from the soil moisture sensor, and the lower and upper limits of the optimal soil moisture content corresponding to the current seedling are obtained from the pre-established variety and optimal moisture range mapping table. When the current soil volumetric moisture content is less than the minimum optimal soil moisture content, the soil moisture stress degree is obtained by subtracting the current soil volumetric moisture content from the minimum optimal soil moisture content and dividing by the minimum optimal soil moisture content. When the current soil volumetric moisture content is greater than the upper limit of the optimal soil moisture content, the soil moisture stress degree is obtained by subtracting the upper limit of the optimal soil moisture content from the current soil volumetric moisture content and dividing by the upper limit of the optimal soil moisture content. When the current soil volumetric moisture content is between the lower limit and the upper limit of the optimal soil moisture content, the soil moisture stress degree is set to 0. Based on the cumulative number of days from the sowing date to the current date, obtain the seedling growth stage index; Obtain the canopy image of the current seedling and input it into the pre-trained seedling wilting recognition model to obtain the canopy wilting characterization value.
[0017] The specific implementation method is as follows: First, the current soil volumetric moisture content is obtained in real time from soil moisture sensors deployed in the seedling cultivation substrate. Soil moisture sensors can be time domain reflectometers (TDR) or frequency domain reflectometers (FDR), and the burial depth is determined based on the main root distribution layer of the seedling. For example, for rhizomes and other root-type medicinal herbs such as Astragalus membranaceus, the burial depth is 10 to 15 cm; for shallow-rooted varieties such as Panax notoginseng and Dendrobium officinale, the burial depth is 5 to 8 cm.
[0018] Furthermore, based on the current seedling variety information, the optimal soil moisture content lower and upper limits for each variety are retrieved from a pre-established variety-optimal moisture range mapping table. This mapping table is pre-calibrated based on the physiological water requirements of different medicinal herb varieties through field moisture gradient experiments. For example, the optimal soil volumetric moisture content range for Astragalus membranaceus is 18% to 28%, for Panax notoginseng it is 22% to 35%, and for Paris polyphylla it is 20% to 30%. When the actual moisture content is within this range, the water absorption and transpiration of the seedling roots are in a dynamic equilibrium.
[0019] Furthermore, based on the relative relationship between the current soil volumetric moisture content and the optimum moisture range, the soil moisture stress degree is calculated: When the soil volumetric moisture content is lower than the lower limit of the optimum soil moisture content, it indicates that the seedlings are under drought stress. Soil moisture stress degree = (lower limit of optimum soil moisture content - current soil volumetric moisture content) ÷ lower limit of optimum soil moisture content. For example, if the current soil volumetric moisture content of a certain Astragalus seedling is 12%, and its optimum soil moisture content lower limit is 18%, then the soil moisture stress degree = (18% - 12%) ÷ 18% ≈ 0.33.
[0020] When the soil volumetric moisture content exceeds the upper limit of the optimum soil moisture content, it indicates that the seedlings are under waterlogging stress and are at risk of root rot. Soil moisture stress degree = (current soil volumetric moisture content - upper limit of optimum soil moisture content) ÷ upper limit of optimum soil moisture content. For example, if the current soil volumetric moisture content of a Panax notoginseng seedling is 42%, and its optimum soil moisture content upper limit is 35%, then the soil moisture stress degree = (42% - 35%) ÷ 35% = 0.20.
[0021] When the soil volumetric moisture content is between the lower and upper limits of the optimum soil moisture content, it indicates that the moisture state is suitable, and the soil moisture stress degree is taken as 0.
[0022] Furthermore, a seedling growth stage index is obtained based on the cumulative number of days from the sowing date to the current date. The growth period divisions differ among different varieties of Chinese medicinal herbs. Taking Panax notoginseng as an example, it can be divided into the emergence stage (1 to 60 days after sowing), leaf expansion stage (61 to 120 days), rhizome enlargement stage (121 to 180 days), and maturity stage (after 181 days). The seedling growth stage index can be represented by normalized growth days, i.e., the current cumulative number of days divided by the total number of days in the entire growth period of the variety, with a value range of 0 to 1. Discrete stage coding can also be used, for example, assigning a value of 1 to the emergence stage, 2 to the leaf expansion stage, 3 to the rhizome enlargement stage, and 4 to the maturity stage. Normalized growth days are preferred, as they provide a more precise reflection of seedling development progress with continuous values.
[0023] Finally, the canopy image of the current seedling is acquired and input into the pre-trained seedling wilting recognition model to obtain the canopy wilting characterization value. The canopy image can be acquired using a visible light or near-infrared camera mounted above the cultivation facility, with the acquisition frequency synchronized with the irrigation decision cycle, for example, once per hour. The seedling wilting recognition model can be constructed using a convolutional neural network, consisting of an input layer, multiple convolutional layers, multiple pooling layers, a fully connected layer, and an output layer connected sequentially. The input layer receives the pre-processed seedling canopy image; the convolutional layers extract leaf edges, textures, and wilting morphological features from the input image through convolution operations; the pooling layers downsample the feature map output by the convolutional layers, compressing the data while retaining the main features; the fully connected layers map the extracted features to the classification probability of each wilting level; and the output layer outputs the final wilting level classification result.
[0024] In the training phase of the seedling wilting recognition model, a large number of canopy image samples of different varieties of medicinal herb seedlings under different cultivation conditions and with varying degrees of wilting were first collected. Each image sample was manually labeled with a wilting degree label, with the following standards: 0 for upright, fully unfolded leaves; 0.3 for slightly drooping leaf edges; 0.6 for overall drooping leaves with curling tips; and 1 for severely curled, wilted, and collapsed leaves. These four labels were then assigned to four wilting levels as classification labels for the model training. To improve the generalization ability of the seedling wilting recognition model, data augmentation processing was performed on the original image samples, including random scaling, flipping, rotation, and brightness adjustment. During the training of the seedling wilting recognition model, the labeled training set images were used as input, and the wilting level labels were used as supervision signals. The cross-entropy loss function was used to calculate the error between the predicted values and the true labels of the seedling wilting recognition model, and the weight parameters of each layer of the seedling wilting recognition model were iteratively updated using the backpropagation algorithm and gradient descent optimizer. During training, the accuracy of the seedling wilting recognition model is periodically evaluated on the validation set. Training is stopped and the optimal model parameters are saved when the validation set accuracy no longer improves after several consecutive training epochs. After training, the seedling wilting recognition model is independently evaluated using the test set. The model must achieve a preset performance threshold (e.g., above 90%) on the test set before it can be deployed as a pre-trained seedling wilting recognition model for practical application. Collected canopy images are input into the trained seedling wilting recognition model. The model outputs the probability distribution of each wilting level. The labeled value corresponding to each wilting level is used as a weight, multiplied by the corresponding level probability output by the model, and summed to obtain a continuous value between 0 and 1 as the canopy wilting characterization value. A larger value indicates a more severe degree of wilting.
[0025] The following technical effects were achieved through this step: First, a three-dimensional state perception system integrating soil environment, seedling development stage and plant phenotype was constructed, which can more comprehensively reflect the actual water requirements of seedlings compared with single soil moisture monitoring.
[0026] Second, soil moisture stress is quantified by the deviation ratio relative to the optimal moisture range of the variety, which unifies the stress measurement dimensions of different varieties and different moisture deviation directions, and provides standardized input for the subsequent generalization of cross-variety irrigation strategies.
[0027] Third, introducing canopy wilting characterization values as upper-level phenotypic feedback of plant water status can promptly detect plant water shortage signals that soil moisture sensors cannot capture, thus improving the sensitivity and accuracy of status perception.
[0028] S200: Based on the state perception parameters and the cumulative water stress experienced by the current seedling, obtain the estimated value of the effective component accumulation of the current seedling; In the cultivation of Chinese medicinal herbs, a complex nonlinear relationship exists between water stress and the accumulation of active ingredients. Moderate drought stress can induce plants to initiate secondary metabolic pathways, promoting the synthesis and accumulation of medicinal active ingredients such as phenols, saponins, and polysaccharides. However, too mild a stress level fails to effectively stimulate secondary metabolic responses, while excessive stress or prolonged stress can lead to physiological damage to the plants, thus inhibiting the normal synthesis of active ingredients. Existing irrigation methods mostly focus only on ensuring seedling survival, failing to incorporate the accumulation status of active ingredients into decision-making, and thus failing to achieve a dynamic balance between increasing the content of active ingredients through moderate water stress regulation and ensuring seedling survival through moderate irrigation in water management.
[0029] Step S200 in the method provided in this application embodiment includes: Based on the state-aware parameters and the cumulative water stress experienced by the seedlings, the estimated value of the current seedling's effective component accumulation is obtained, including: Obtain the duration of soil moisture stress exceeding the preset stress threshold during each growth stage of the current seedling since sowing, and obtain the cumulative water stress sequence; The cumulative water stress sequence is input into a pre-trained stress-component assessment model to obtain the estimated value output by the stress-component assessment model, which is used as the estimated value of the current effective component accumulation in the seedling.
[0030] The specific implementation method is as follows: First, the duration of soil moisture stress exceeding a preset stress threshold during each growth stage of the current seedlings since sowing is obtained, constructing a cumulative water stress sequence. The preset stress threshold refers to the minimum stress level that can induce the plant to initiate a secondary metabolic response, and can be calibrated through pre-experiments based on the physiological characteristics of different Chinese medicinal herb varieties. For example, for Panax notoginseng, the preset stress threshold can be set to 0.15; for Astragalus membranaceus, the preset stress threshold can be set to 0.12. When the soil moisture stress level calculated in step S100 is greater than the preset stress threshold, it indicates that the current water stress has reached a level sufficient to induce the synthesis of effective components, and the stress duration is then calculated.
[0031] Furthermore, the entire growth period of the seedlings is divided into multiple growth stages, and the division method of each growth stage is consistent with the stage division of the seedling growth stage index in step S100. Taking Panax notoginseng as an example, it is divided into four stages: emergence stage, leaf expansion stage, rhizome enlargement stage, and maturity stage. For each growth stage, the cumulative duration of soil moisture stress exceeding the preset stress threshold within that stage is calculated. The cumulative duration is expressed in hours or days, forming a sequence value corresponding to the number of growth stages, i.e., the cumulative water stress sequence. For example, if a certain Astragalus membranaceus seedling is currently in the rhizome enlargement stage, and its cumulative stress duration during the emergence stage is 20 hours, during the leaf expansion stage it is 55 hours, and during the rhizome enlargement stage it is 30 hours, then the current cumulative water stress sequence of this seedling can be represented as {20, 55, 30}. If a certain growth stage has not yet been experienced or the cumulative stress duration within that stage is zero, then the corresponding sequence value is 0. The cumulative water stress sequence reflects the comprehensive history of the intensity and duration of water stress experienced by seedlings at different developmental stages, and is a core input for assessing the current level of effective ingredient accumulation.
[0032] Finally, the constructed cumulative water stress sequence is input into a pre-trained stress-component assessment model to obtain the model's output estimate of effective component accumulation. The stress-component assessment model is a pre-trained machine learning model used to establish a mapping relationship between the cumulative water stress sequence and the current relative content of effective components in seedlings.
[0033] The construction of the stress-component assessment model includes: Collect multiple sets of historical planting sample data. Each set of sample data includes the cumulative water stress sequence of the sample seedlings, which consists of the duration of soil water stress exceeding the preset stress threshold at each growth stage, and the actual content of effective components obtained by testing the sample seedlings after harvest. A stress-component assessment model is constructed, which takes the cumulative moisture stress sequence as input and the actual content of the effective component as output. The stress-component assessment model is trained using the multiple sets of historical planting sample data until it converges, thus obtaining the completed stress-component assessment model.
[0034] Specifically, multiple sets of historical planting sample data were collected. Each set of sample data originated from the historical planting records of the same variety of medicinal herb seedlings. This included the cumulative water stress sequence of the sample seedlings, comprising the duration for which soil moisture stress exceeded a preset stress threshold at each growth stage, and the actual content of active ingredients obtained after harvesting. The actual content of active ingredients could be determined using standard analytical methods such as high-performance liquid chromatography or ultraviolet spectrophotometry, and expressed as the mass percentage of the target active ingredient per unit mass of medicinal herb. To improve the generalization ability of the stress-component assessment model, sample data covering different planting batches and different stress treatment gradients should be collected as comprehensively as possible.
[0035] Furthermore, a stress-component assessment model is constructed, which takes the cumulative water stress sequence as input and the actual content of the effective component as output. The stress-component assessment model can employ a multiple linear regression model, a support vector regression model, or a shallow fully connected neural network. Preferably, a fully connected neural network with one hidden layer is used, consisting of an input layer, a hidden layer, and an output layer connected sequentially. The number of nodes in the input layer is equal to the length of the cumulative water stress sequence, i.e., equal to the number of growth stage divisions; the hidden layer contains several neurons used to extract the nonlinear mapping relationship between stress time-series features and effective component accumulation; the output layer contains one node, outputting a normalized estimated effective component content, ranging from 0 to 1, representing the ratio of the current effective component accumulation to the theoretical maximum accumulation for this variety. The activation function between layers can be a linear rectified function.
[0036] The stress-component assessment model was trained using the aforementioned multiple sets of historical planting sample data until convergence. During training, the cumulative moisture stress sequence of the samples was used as the input to the stress-component assessment model, and the corresponding actual content of effective components was used as the supervision label. The mean squared error loss function was used to calculate the error between the model's predicted values and the actual contents, and the weight parameters of each layer of the stress-component assessment model were iteratively updated using the backpropagation algorithm and the gradient descent optimizer. During training, the sample data was divided into training and validation sets. Parameters were updated on the training set, and the prediction error of the stress-component assessment model was monitored on the validation set. When the loss function value on the validation set no longer decreased for several consecutive training epochs, the stress-component assessment model was considered to have converged. Training was stopped, the optimal model parameters were saved, and the completed stress-component assessment model was obtained. After training, the prediction accuracy of the stress-component assessment model could be evaluated using an independent test set. The correlation coefficient between the predicted values and the actual values was required to reach a preset threshold, such as 0.85 or higher, before the model could be deployed as a pre-trained stress-component assessment model.
[0037] In practical applications, the cumulative water stress sequence of the current seedlings is input into the trained stress-component assessment model. The model outputs a continuous value between 0 and 1, which is the estimated value of the current accumulation of effective components in the seedlings. The larger the value, the higher the accumulation level of effective components in the seedlings under the current cumulative stress history.
[0038] The following technical effects were achieved through this step: First, by constructing a cumulative water stress sequence, the water stress history at different stages throughout the entire growth period of seedlings can be expressed in a structured and quantitative manner, which can accurately reflect the cumulative effect and stage dependence of water stress on the accumulation of effective components.
[0039] Second, a stress-component assessment model was trained based on historical planting sample data, and a quantitative mapping relationship between the cumulative water stress sequence and the accumulation level of effective components was established, enabling irrigation decisions to predict the current medicinal quality status of seedlings in real time.
[0040] Third, by incorporating the estimated accumulation of effective components into the irrigation decision-making loop, the dual objectives of ensuring seedling survival and enhancing the efficacy of medicinal components are optimized in a coordinated manner, thus solving the technical problem that existing irrigation methods cannot simultaneously achieve both seedling survival and efficacy enhancement.
[0041] S300: Input the state-aware parameters and the effective component accumulation estimate into the pre-trained irrigation decision model to obtain irrigation action decisions, wherein the irrigation action decisions include irrigation duration and irrigation intensity; In intelligent irrigation of medicinal herb seedlings, irrigation decisions need to find the optimal balance among multiple interrelated and even mutually restrictive objectives. On the one hand, irrigation needs to alleviate current water stress, ensure normal seedling growth, and prevent excessive wilting; on the other hand, moderate water stress is a necessary condition for inducing the accumulation of effective components, while excessive irrigation will inhibit the synthesis of secondary metabolites. In addition, seedlings at different growth stages have different sensitivities to water; the same amount of irrigation may lead to waterlogging during the seedling stage, but may be insufficient during the rhizome enlargement stage.
[0042] Step S300 in the method provided in this application embodiment includes: The soil moisture stress degree, the seedling growth stage index, the canopy wilting characterization value, and the effective component accumulation prediction value are combined into a multidimensional state vector; The multidimensional state vector is input into the pre-trained irrigation decision model to obtain the output irrigation action decision, wherein the irrigation action decision includes irrigation duration and irrigation intensity.
[0043] The specific implementation method is as follows: First, the parameters obtained in steps S100 and S200 are combined into a multidimensional state vector. Specifically, the multidimensional state vector consists of the following four components: the first component is the soil moisture stress degree, ranging from 0 to positive infinity, with a larger value indicating more severe water stress; the second component is the seedling growth stage index, ranging from 0 to 1, representing the current growth stage of the seedling; the third component is the canopy wilting characterization value, ranging from 0 to 1, with a larger value indicating a more severe degree of plant wilting; and the fourth component is the effective component accumulation prediction value, ranging from 0 to 1, with a larger value indicating a higher current effective component accumulation level. These four components are then concatenated in a preset order to form a four-dimensional state vector, which serves as the input to the irrigation decision model.
[0044] Furthermore, the constructed multidimensional state vector is input into a pre-trained irrigation decision model to obtain the irrigation action decision output by the model. The irrigation decision model can be a pre-trained deep neural network model used to establish the mapping relationship between the multidimensional state vector and the optimal irrigation action. The irrigation decision model can adopt a fully connected neural network structure containing multiple hidden layers, consisting of an input layer, multiple hidden layers, and an output layer connected sequentially. The input layer has four nodes, corresponding to the four components of the multidimensional state vector; the hidden layers are used to extract the nonlinear combination relationship between state features, and a linear rectified function can be used as the activation function between the hidden layers; the output layer contains two nodes, outputting irrigation duration and irrigation intensity, respectively. The unit of irrigation duration is seconds or minutes, and the unit of irrigation intensity is liters per hour or millimeters per hour, representing the irrigation water volume per unit time.
[0045] The training of the irrigation decision model includes: Multiple sets of historical irrigation records were collected. Each set of historical irrigation event records included sample state perception parameters and sample effective component accumulation prediction values collected before irrigation, as well as sample irrigation action decisions to be evaluated during actual irrigation. Updated soil moisture stress, updated seedling growth stage index, updated canopy wilting characterization value, and updated effective component accumulation value were obtained after irrigation. Fitness analysis was performed on the updated soil moisture stress degree, updated seedling growth stage index, updated canopy wilting characterization value, and updated effective component accumulation value. The irrigation duration and irrigation intensity with the highest fitness were extracted as the sample irrigation action decision. Using the sample state perception parameters and the sample effective component accumulation prediction as inputs, and the sample irrigation actions as supervision labels, the irrigation decision model is trained until convergence, thus obtaining a pre-trained irrigation decision model.
[0046] Specifically, multiple sets of historical irrigation records were collected. Each set of historical irrigation records originated from complete irrigation events recorded during actual planting and included the following: sample state perception parameters collected before irrigation, specifically including sample soil moisture stress, sample seedling growth stage index, and sample canopy wilting characterization value; estimated sample effective component accumulation value before irrigation; sample irrigation action decision to be evaluated during actual irrigation, specifically including sample irrigation duration and sample irrigation intensity; and updated state parameters collected after a preset time interval, such as two hours or four hours after irrigation, specifically including updated soil moisture stress, updated seedling growth stage index, updated canopy wilting characterization value, and updated effective component accumulation value.
[0047] In practical applications, the multidimensional state vector at the current moment is input into the trained irrigation decision model. The model outputs two continuous values, which serve as irrigation duration and irrigation intensity, respectively. The irrigation decision model can dynamically adjust its output irrigation actions based on real-time changes in the input state. For example, when soil moisture stress is high and canopy wilting is significant, the model tends to output a longer irrigation duration and a higher irrigation intensity. Conversely, when the estimated accumulation of effective components is low and the stress level is within a suitable range, the model tends to output a shorter irrigation duration to maintain appropriate stress and promote the synthesis of effective components.
[0048] Fitness analysis was performed on the updated soil moisture stress, updated seedling growth stage index, updated canopy wilting characterization value, and updated effective component accumulation value. The irrigation duration and irrigation intensity with the highest fitness were extracted as the basis for sample irrigation action decisions, including: Calculate the first absolute difference between the updated soil moisture stress degree and the preset ideal stress degree, calculate 1 minus the first absolute difference and multiply by the first weighting coefficient to obtain the first fitness component; Calculate the second absolute difference between the updated seedling growth stage index and the preset ideal stage index, calculate 1 minus the second absolute difference and multiply by the second weighting coefficient to obtain the second fitness component; The third fitness component is obtained by multiplying the updated canopy wilting characterization value by the third weighting coefficient after subtracting the value from the calculation. Multiply the updated effective ingredient accumulation value by the fourth weighting coefficient to obtain the fourth fitness component; The first fitness component, the second fitness component, the third fitness component, and the fourth fitness component are added together to obtain the fitness of the irrigation action decision to be evaluated in the sample. The sample irrigation action decision with the highest overall fitness was selected as the sample irrigation action decision.
[0049] The specific calculation method for fitness analysis is as follows: First, calculate the first absolute difference between the updated soil moisture stress level and the preset ideal stress level. The preset ideal stress level refers to the target stress level that can effectively induce the synthesis of active ingredients without causing physiological damage, and can be preset according to different varieties and growth stages. For example, for the rhizome enlargement stage of Panax notoginseng, the preset ideal stress level can be set to 0.25. Calculate 1, subtract the first absolute difference, and multiply by the first weighting coefficient to obtain the first fitness component. The first fitness component reflects the degree of closeness between the actual stress level after irrigation and the ideal stress level; the closer it is to the ideal value, the higher the fitness.
[0050] Furthermore, the second absolute difference between the updated seedling growth stage index and the preset ideal stage index is calculated. The preset ideal stage index refers to the standard development progress reference value for that growth stage under normal water conditions. The second fitness component is obtained by subtracting the second absolute difference from the calculated value and multiplying it by the second weighting coefficient. The second fitness component reflects the promoting effect of irrigation on seedling development progress.
[0051] The third fitness component is obtained by subtracting the updated canopy wilting characterization value from Calculation 1 and multiplying it by the third weighting coefficient. The third fitness component reflects the effect of irrigation on alleviating plant wilting; the lower the degree of wilting after irrigation, the higher the fitness.
[0052] Finally, the updated effective component accumulation value is multiplied by the fourth weighting coefficient to obtain the fourth fitness component. The fourth fitness component reflects the contribution of irrigation to the accumulation of effective components; the higher the effective component accumulation value after irrigation, the higher the fitness.
[0053] The first, second, third, and fourth weighting coefficients mentioned above can be set according to the priority of the actual breeding objectives. If more emphasis is placed on improving the effective components, the fourth weighting coefficient can be set to a larger value; if more emphasis is placed on ensuring seedling survival, the first and third weighting coefficients can be set to larger values. The sum of the four weighting coefficients can be 1 to normalize the overall fitness.
[0054] The first, second, third, and fourth fitness components are added together, and the sum is taken as the fitness of the irrigation action decision to be evaluated for that sample. The higher the fitness value, the better the overall effect of the irrigation action decision.
[0055] For the same set of pre-irrigation conditions, there may be multiple different sets of samples to evaluate irrigation action decisions and their post-irrigation effects. The sample irrigation action decision with the highest overall fitness is selected as the sample irrigation action decision under that condition.
[0056] The following technical effects were achieved through this step: First, soil moisture stress, seedling growth stage index, canopy wilting characterization value, and effective component accumulation prediction value are combined into a unified multidimensional state vector, enabling the irrigation decision model to simultaneously perceive information from four dimensions: root environment, development stage, plant phenotype, and quality status, thus achieving effective fusion of multi-source heterogeneous information.
[0057] Second, training samples are constructed based on historical irrigation records and fitness analysis. The fitness function comprehensively considers four objectives: stress control, development protection, wilting relief, and component improvement, so that the trained irrigation decision model can output the comprehensive optimal irrigation action under multi-objective constraints.
[0058] Third, the irrigation decision model outputs continuous irrigation duration and intensity values, which, compared to discrete on / off irrigation control, enables more precise water volume regulation and meets differentiated irrigation needs under different conditions.
[0059] S400: Control the irrigation equipment to perform irrigation according to the irrigation action decision, and obtain the updated status perception parameters and the updated cumulative water stress degree after irrigation; In intelligent irrigation of medicinal herb seedlings, the output of irrigation action decisions needs to be physically executed to affect the seedling cultivation environment. The irrigation execution process involves converting the irrigation duration and intensity output by the decision model into control commands for the irrigation equipment and coordinating the collaborative work of multiple execution elements. Furthermore, after irrigation is completed, water infiltration into the soil and root absorption require a certain amount of time. Therefore, it is essential to re-collect state parameters at an appropriate time to accurately reflect the actual effect of the irrigation. If the data is collected too early, the soil moisture has not yet stabilized, and the collected data cannot represent the true state after irrigation; if the data is collected too late, the dynamic change window of the seedling status may be missed.
[0060] Step S400 in the method provided in this application embodiment includes: The irrigation action decision control system executes irrigation by controlling the irrigation equipment and acquiring updated status perception parameters and updated cumulative water stress level after irrigation.
[0061] The specific implementation method is as follows: First, based on the irrigation action decision obtained in step S300, the irrigation equipment is controlled to perform irrigation. The irrigation equipment may include actuators such as solenoid valves, water pumps, drip irrigation pipes, or sprinkler heads. The control system converts the irrigation duration and intensity into corresponding control signals; for example, the irrigation duration is converted into the solenoid valve's opening duration, and the irrigation intensity is converted into the water pump speed or the opening degree of the flow regulating valve. After irrigation is completed, a preset time interval is allowed to allow the irrigation water to fully infiltrate the soil and interact with the seedling roots. The preset time interval can be set according to soil type and irrigation volume; for example, for sandy loam, it can be set to two hours after irrigation; for clay loam, it can be set to three to four hours after irrigation.
[0062] Further, after waiting for a preset time interval, the parameters are re-acquired according to the method in step S100 to obtain updated state perception parameters. Specifically, the updated soil volumetric water content value is re-acquired from the soil moisture sensor, and the updated soil moisture stress degree is obtained according to the calculation method in step S100; the current date is re-acquired, and the updated seedling growth stage index is calculated; the seedling canopy image is re-acquired, input into the pre-trained seedling wilting recognition model, and the updated canopy wilting characterization value is obtained.
[0063] Furthermore, the updated cumulative water stress level is obtained. The method for obtaining the updated cumulative water stress level is as follows: based on the comparison between the updated soil water stress level and a preset stress threshold, it is determined whether the soil is under stress at the current time after irrigation. If the updated soil water stress level is greater than the preset stress threshold, the time from this irrigation to the current time is included in the cumulative stress duration of the current growth stage to obtain the updated cumulative water stress level.
[0064] The following technical effects were achieved through this step: By setting a preset waiting time interval after irrigation, soil moisture is allowed to stabilize before collecting and updating state parameters, avoiding misjudgments caused by improper data collection timing and providing an accurate and reliable data foundation for subsequent reward value calculation. The cumulative water stress level is updated synchronously after irrigation, and changes in stress status after this irrigation are promptly included in the cumulative stress record, ensuring the integrity and continuity of the water stress history throughout the seedling's entire growth period.
[0065] S500: Based on the updated state perception parameters and the updated cumulative water stress level, obtain the updated effective component accumulation value, calculate the reward value, and use the reward value to update the irrigation decision model online.
[0066] In the cultivation of medicinal herb seedlings, dynamic changes in the planting environment, physiological differences among individual seedlings, and fluctuations in varietal characteristics between different batches of seedlings make it difficult for offline pre-trained irrigation decision models to fully cover all state scenarios in real-world applications. If the irrigation decision model remains fixed after offline training, it cannot adapt to long-term changes in environmental conditions and seedling states during cultivation, and its decision-making performance will gradually decline over time. Furthermore, the actual effectiveness of irrigation decisions needs to be evaluated and verified through post-execution state changes. Feeding the execution results back to the irrigation decision model to form a closed-loop optimization is crucial for achieving continuous online updates.
[0067] Step S500 in the method provided in this application embodiment includes: The updated cumulative water stress level is combined with the cumulative water stress sequence to form an updated cumulative water stress sequence; The updated cumulative water stress sequence is input into the stress-component assessment model to obtain the updated effective component accumulation value; The difference between the updated effective component accumulation value and the effective component accumulation estimate is calculated and recorded as the component increase. When the component increase is positive and the soil moisture stress before irrigation is in the preset stress induction range, the component increase is multiplied by the preset component incentive coefficient to obtain the first reward component. Calculate the difference between the updated canopy wilting characterization value and the preset wilting warning threshold to obtain the wilting safety margin; When the wilting safety margin is positive, the wilting safety margin is multiplied by a preset safety reward coefficient to obtain the second reward component; When the wilting safety margin is negative, the wilting safety margin is multiplied by a preset safety penalty coefficient to obtain a second reward component; Calculate the absolute value of the improvement in soil moisture stress before and after irrigation to obtain the stress improvement amount. Divide the stress improvement amount by the product of irrigation duration and irrigation intensity in this irrigation action decision, and multiply by a preset efficiency reward coefficient to obtain the third reward component. If the volumetric moisture content of the updated soil after irrigation is greater than the upper limit of the optimal soil moisture content, the preset root rot risk penalty value is obtained as the fourth reward component; otherwise, the root rot risk penalty value is set to 0. The first reward component, the second reward component, the third reward component, and the fourth reward component are added together, and the sum is taken as the reward value. The state perception parameters before irrigation, the irrigation action decision, the reward value, and the updated state perception parameters after irrigation are combined into a complete experience sample and stored in the experience replay pool. The specific implementation method is as follows: First, obtain the updated cumulative water stress sequence. Combine the updated cumulative water stress level with the cumulative water stress sequence to form an updated cumulative water stress sequence, including: obtaining the cumulative water stress sequence and obtaining the updated cumulative water stress level; adding the updated cumulative water stress level to the corresponding growth stage in the cumulative water stress sequence to form the updated cumulative water stress sequence. Specifically, for example, if a Panax notoginseng seedling is in the root and stem enlargement stage before irrigation, its cumulative water stress sequence is {20, 55, 30}. After irrigation, the newly added cumulative stress duration in this growth stage is 5 hours, then the corresponding sequence value is updated to 30 plus 5 equals 35, and the updated cumulative water stress sequence is {20, 55, 35}.
[0068] Furthermore, the updated cumulative water stress sequence is input into the stress-component assessment model constructed in step S200 to obtain the updated effective component accumulation value output by the model. The updated effective component accumulation value reflects the latest estimated state of effective component accumulation level in seedlings after irrigation.
[0069] Further, the reward value is calculated. The reward value is used to quantitatively evaluate the effectiveness of this irrigation action decision; a higher reward value indicates a better irrigation decision. The reward value is calculated using a multi-component weighted summation method, and the calculation methods for each component are as follows: Calculate the first reward component. Calculate the difference between the updated effective component accumulation value and the estimated effective component accumulation value, denoted as the component increase. When the component increase is positive and the soil moisture stress level before irrigation is within a preset stress induction range, multiply the component increase by a preset component incentive coefficient to obtain the first reward component. The preset stress induction range refers to the range of water stress levels that can effectively induce the synthesis of secondary metabolites, for example, it can be set between 0.15 and 0.40. When the pre-irrigation stress level is within this range, it indicates that the seedlings are in the sensitive period for effective component induction, and the increase in effective components at this time gives a positive reward. If the component increase is negative or the pre-irrigation stress level is not within the preset stress induction range, the first reward component is 0.
[0070] Calculate the second reward component. Calculate the difference between the updated canopy wilting characteristic value and the preset wilting warning threshold to obtain the wilting safety margin. The preset wilting warning threshold refers to the critical value of wilting degree at which irreversible physiological damage to the seedlings is about to occur; for example, it can be set to 0.7. When the wilting safety margin is positive, it indicates that the degree of wilting after irrigation is within a safe range. Multiply the wilting safety margin by the preset safety reward coefficient to obtain the second reward component. When the wilting safety margin is negative, it indicates that the degree of wilting after irrigation has exceeded the warning threshold, and the seedlings face the risk of physiological damage. Multiply the wilting safety margin by the preset safety penalty coefficient to obtain the second reward component. The absolute value of the safety penalty coefficient should be greater than the absolute value of the safety reward coefficient to strengthen the avoidance of excessive wilting risk. For example, the safety reward coefficient can be set to 0.5, and the safety penalty coefficient can be set to -1.0.
[0071] Calculate the third reward component. Calculate the absolute value of the improvement in soil moisture stress before and after irrigation to obtain the stress improvement amount. Divide the stress improvement amount by the product of irrigation duration and irrigation intensity in this irrigation action decision to obtain the stress improvement effect produced by the unit irrigation resource input. Then multiply by a preset efficiency reward coefficient to obtain the third reward component. The third reward component aims to achieve the greatest possible stress improvement effect with the least possible irrigation resource consumption, promoting the efficient use of water resources. If the stress level worsens after irrigation, the stress improvement amount is negative, and the third reward component is set to 0.
[0072] Calculate the fourth reward component. When the post-irrigation soil volumetric moisture content exceeds the upper limit of the optimum soil moisture content, it indicates that the irrigation caused excessive soil moisture accumulation, posing a risk of inducing root rot. A preset root rot risk penalty value is then obtained as the fourth reward component. The root rot risk penalty value can be set to a fixed negative value, such as -0.5. When the post-irrigation soil volumetric moisture content is less than or equal to the upper limit of the optimum soil moisture content, the fourth reward component is set to 0.
[0073] The first reward component, the second reward component, the third reward component, and the fourth reward component are added together, and the sum is used as the reward value corresponding to this irrigation action decision. The positive or negative sign and magnitude of the reward value comprehensively reflect the overall performance of this irrigation in four dimensions: improvement of effective components, prevention and control of wilting risk, improvement of stress efficiency, and rational utilization of water resources.
[0074] Finally, the irrigation decision model is updated online using the reward value. The specific update method is as follows: The state-aware parameters before irrigation, the irrigation action decision, the reward value, and the updated state-aware parameters after irrigation are combined into a complete experience sample and stored in the experience replay pool. The data structure of a complete experience sample can be represented as {state vector before irrigation, irrigation action decision, reward value, state vector after irrigation}. The experience replay pool uses a fixed-capacity queue structure to store experience samples. When the number of stored samples exceeds the preset capacity, the oldest stored sample is automatically removed to ensure that the experience replay pool contains the most recent experience data.
[0075] When the number of complete experience samples stored in the experience replay pool reaches the preset number of update batches, a batch of experience samples is randomly sampled from the experience replay pool, and a reinforcement learning algorithm is used to update the parameters of the irrigation decision model. The preset number of update batches can be set according to the storage and computing capabilities of the actual system, for example, it can be set to 32 or 64. The reinforcement learning algorithm can be a deep Q-network algorithm or a deep deterministic policy gradient algorithm. Taking the deep deterministic policy gradient algorithm as an example, the irrigation decision model acts as a policy network, while maintaining a target network and a value network with the same structure as the policy network. During parameter updates, the temporal difference error is calculated using the sampled experience samples, and the value network parameters are updated using the gradient descent method; then, the policy network parameters are updated using the policy gradient method; the target network parameters are updated softly, slowly approaching the policy network parameters at small percentages every few steps. Through the above online update mechanism, the irrigation decision model can continuously optimize its parameters using the actual execution effect of each irrigation, so that subsequent irrigation decisions gradually approach the optimal strategy.
[0076] The following technical effects were achieved through this step: First, a multi-component reward function was constructed, which includes four dimensions: component enhancement, wilting safety, stress improvement efficiency, and rational water resource utilization. This enables the irrigation decision-making model to comprehensively evaluate irrigation effects from multiple perspectives, avoiding decision-making biases caused by single-indicator evaluation.
[0077] Second, an experience replay pool is used to store historical experience samples, and batch updates are performed through random sampling. This breaks the temporal correlation between continuous irrigation events and improves the stability of irrigation decision model updates and data utilization efficiency.
[0078] Third, it realizes online closed-loop updating of the irrigation decision model, enabling the model to continuously learn and self-evolve from the feedback of each actual irrigation, and gradually adapt to the dynamic changes in the planting environment and the physiological differences of individual seedlings.
[0079] Example 2 like Figure 2As shown, based on the same inventive concept as the intelligent irrigation method for medicinal herb seedlings based on reinforcement learning provided in Embodiment 1, this embodiment of the invention also provides an intelligent irrigation system for medicinal herb seedlings based on reinforcement learning. The system includes a data acquisition and status perception module, an effective component evaluation module, an irrigation decision module, and an online update module. The system includes: The state perception module 11 is used to acquire the state perception parameters of the current seedling cultivation environment. The state perception parameters include soil moisture stress degree, seedling growth stage index and canopy wilting characterization value. The effective component evaluation module 12 is used to obtain the estimated value of the current seedling's effective component accumulation based on the state perception parameters and the current seedling's cumulative water stress level. Irrigation decision module 13 is used to input the state-aware parameters and the effective component accumulation estimate into a pre-trained irrigation decision model to obtain irrigation action decisions, wherein the irrigation action decisions include irrigation duration and irrigation intensity; The irrigation execution module 14 is used to control the irrigation equipment to perform irrigation according to the irrigation action decision, and to obtain the updated status perception parameters and the updated cumulative water stress degree after irrigation. The online update module 15 obtains the accumulated value of the effective components for updating based on the update status perception parameters and the cumulative water stress degree of the update, calculates the reward value, and uses the reward value to update the irrigation decision model online.
[0080] In one embodiment, the state perception module 11 is further configured to acquire state perception parameters of the current seedling cultivation environment, the state perception parameters including soil moisture stress degree, seedling growth stage index, and canopy wilting characterization value, including: The current soil volumetric moisture content is obtained from the soil moisture sensor, and the lower and upper limits of the optimal soil moisture content corresponding to the current seedling are obtained from the pre-established variety and optimal moisture range mapping table. When the current soil volumetric moisture content is less than the minimum optimal soil moisture content, the soil moisture stress degree is obtained by subtracting the current soil volumetric moisture content from the minimum optimal soil moisture content and dividing by the minimum optimal soil moisture content. When the current soil volumetric moisture content is greater than the upper limit of the optimal soil moisture content, the soil moisture stress degree is obtained by subtracting the upper limit of the optimal soil moisture content from the current soil volumetric moisture content and dividing by the upper limit of the optimal soil moisture content. When the current soil volumetric moisture content is between the lower limit and the upper limit of the optimal soil moisture content, the soil moisture stress degree is set to 0. Based on the cumulative number of days from the sowing date to the current date, obtain the seedling growth stage index; Obtain the canopy image of the current seedling and input it into the pre-trained seedling wilting recognition model to obtain the canopy wilting characterization value.
[0081] In one embodiment, the effective component assessment module 12 is further configured to obtain an estimated value of the current seedling's effective component accumulation based on the state-aware parameters and the current cumulative water stress experienced by the seedling, including: Obtain the duration of soil moisture stress exceeding the preset stress threshold during each growth stage of the current seedling since sowing, and obtain the cumulative water stress sequence; The cumulative water stress sequence is input into a pre-trained stress-component assessment model to obtain the estimated value output by the stress-component assessment model, which is used as the estimated value of the current effective component accumulation in the seedling.
[0082] The construction of the stress-component assessment model includes: Collect multiple sets of historical planting sample data. Each set of sample data includes the cumulative water stress sequence of the sample seedlings, which consists of the duration of soil water stress exceeding the preset stress threshold at each growth stage, and the actual content of effective components obtained by testing the sample seedlings after harvest. A stress-component assessment model is constructed, which takes the cumulative moisture stress sequence as input and the actual content of the effective component as output. The stress-component assessment model is trained using the multiple sets of historical planting sample data until it converges, thus obtaining the completed stress-component assessment model.
[0083] In one embodiment, the irrigation decision module 13 is further configured to input the state-aware parameters and the effective component accumulation estimate into a pre-trained irrigation decision model to obtain an irrigation action decision, wherein the irrigation action decision includes irrigation duration and irrigation intensity, including: The soil moisture stress degree, the seedling growth stage index, the canopy wilting characterization value, and the effective component accumulation prediction value are combined into a multidimensional state vector; The multidimensional state vector is input into the pre-trained irrigation decision model to obtain the output irrigation action decision, wherein the irrigation action decision includes irrigation duration and irrigation intensity.
[0084] The training of the irrigation decision model includes: Multiple sets of historical irrigation records were collected. Each set of historical irrigation event records included sample state perception parameters and sample effective component accumulation prediction values collected before irrigation, as well as sample irrigation action decisions to be evaluated during actual irrigation. Updated soil moisture stress, updated seedling growth stage index, updated canopy wilting characterization value, and updated effective component accumulation value were obtained after irrigation. Fitness analysis was performed on the updated soil moisture stress degree, updated seedling growth stage index, updated canopy wilting characterization value, and updated effective component accumulation value. The irrigation duration and irrigation intensity with the highest fitness were extracted as the sample irrigation action decision. Using the sample state perception parameters and the sample effective component accumulation prediction as inputs, and the sample irrigation actions as supervision labels, the irrigation decision model is trained until convergence, thus obtaining a pre-trained irrigation decision model.
[0085] Fitness analysis was performed on the updated soil moisture stress, updated seedling growth stage index, updated canopy wilting characterization value, and updated effective component accumulation value. The irrigation duration and irrigation intensity with the highest fitness were extracted as the basis for sample irrigation action decisions, including: Calculate the first absolute difference between the updated soil moisture stress degree and the preset ideal stress degree, calculate 1 minus the first absolute difference and multiply by the first weighting coefficient to obtain the first fitness component; Calculate the second absolute difference between the updated seedling growth stage index and the preset ideal stage index, calculate 1 minus the second absolute difference and multiply by the second weighting coefficient to obtain the second fitness component; The third fitness component is obtained by multiplying the updated canopy wilting characterization value by the third weighting coefficient after subtracting the value from the calculation. Multiply the updated effective ingredient accumulation value by the fourth weighting coefficient to obtain the fourth fitness component; The first fitness component, the second fitness component, the third fitness component, and the fourth fitness component are added together to obtain the fitness of the irrigation action decision to be evaluated in the sample. The sample irrigation action decision with the highest overall fitness was selected as the sample irrigation action decision.
[0086] In one embodiment, the irrigation execution module 14 is used to combine the updated cumulative water stress level with the cumulative water stress sequence to form an updated cumulative water stress sequence, including: Obtain the cumulative water stress sequence and the updated cumulative water stress level; The updated cumulative water stress level is added to the cumulative water stress sequence corresponding to the growth stage to form an updated cumulative water stress sequence.
[0087] In one embodiment, the online update module 15 is used to obtain the accumulated value of the effective components for updating based on the update status sensing parameters and the cumulative water stress degree of the update, calculate a reward value, and update the irrigation decision model using the reward value, including: The updated cumulative water stress level is combined with the cumulative water stress sequence to form an updated cumulative water stress sequence; The updated cumulative water stress sequence is input into the stress-component assessment model to obtain the updated effective component accumulation value; The difference between the updated effective component accumulation value and the effective component accumulation estimate is calculated and recorded as the component increase. When the component increase is positive and the soil moisture stress before irrigation is in the preset stress induction range, the component increase is multiplied by the preset component incentive coefficient to obtain the first reward component. Calculate the difference between the updated canopy wilting characterization value and the preset wilting warning threshold to obtain the wilting safety margin; When the wilting safety margin is positive, the wilting safety margin is multiplied by a preset safety reward coefficient to obtain the second reward component; When the wilting safety margin is negative, the wilting safety margin is multiplied by a preset safety penalty coefficient to obtain a second reward component; Calculate the absolute value of the improvement in soil moisture stress before and after irrigation to obtain the stress improvement amount. Divide the stress improvement amount by the product of irrigation duration and irrigation intensity in this irrigation action decision, and multiply by a preset efficiency reward coefficient to obtain the third reward component. If the updated soil volumetric moisture content is greater than the upper limit of the optimal soil moisture content after irrigation, the preset root rot risk penalty value is obtained as the fourth reward component; otherwise, the root rot risk penalty value is set to 0. The first reward component, the second reward component, the third reward component, and the fourth reward component are added together, and the sum is taken as the reward value. The state perception parameters before irrigation, the irrigation action decision, the reward value, and the updated state perception parameters after irrigation are combined into a complete experience sample and stored in the experience replay pool. When the number of complete experience samples stored in the experience replay pool reaches the preset number of update batches, random sampling is performed from the experience replay pool, and the parameters of the irrigation decision model are updated using a reinforcement learning algorithm.
[0088] It should be noted that the processes depicted in the accompanying drawings do not necessarily require a specific or sequential order to achieve the desired result. In some implementations, multitasking and parallel processing are possible or may be advantageous.
[0089] The above description is only a preferred embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.
[0090] This specification and accompanying drawings are merely illustrative examples of this application and are intended to cover any and all modifications, variations, combinations, or equivalents within the scope of this application. Clearly, those skilled in the art can make various alterations and modifications to this application without departing from its scope. Therefore, if such modifications and modifications fall within the scope of this application and its equivalents, this application intends to include such modifications and modifications.
Claims
1. A smart irrigation method for Chinese medicinal herb seedlings based on reinforcement learning, characterized in that, include: Acquire the status sensing parameters of the current seedling cultivation environment, including soil moisture stress, seedling growth stage index and canopy wilting characterization value; Based on the state-aware parameters and the cumulative water stress experienced by the seedlings, the estimated value of the accumulation of effective components in the seedlings is obtained. The state-aware parameters and the effective component accumulation estimate are input into the pre-trained irrigation decision model to obtain irrigation action decisions, wherein the irrigation action decisions include irrigation duration and irrigation intensity. The irrigation action decision control system executes irrigation, and obtains updated status perception parameters and updated cumulative water stress level after irrigation. Based on the updated state perception parameters and the updated cumulative water stress level, the accumulated value of the updated effective components is obtained, and a reward value is calculated. The reward value is then used to update the irrigation decision model online.
2. The intelligent irrigation method for medicinal herb seedlings based on reinforcement learning according to claim 1, characterized in that, Acquire the current seedling cultivation environment's state-sensing parameters, including soil moisture stress, seedling growth stage index, and canopy wilting characterization value, including: The current soil volumetric moisture content is obtained from the soil moisture sensor, and the lower and upper limits of the optimal soil moisture content corresponding to the current seedling are obtained from the pre-established variety and optimal moisture range mapping table. When the current soil volumetric moisture content is less than the minimum optimal soil moisture content, the soil moisture stress degree is obtained by subtracting the current soil volumetric moisture content from the minimum optimal soil moisture content and dividing by the minimum optimal soil moisture content. When the current soil volumetric moisture content is greater than the upper limit of the optimal soil moisture content, the soil moisture stress degree is obtained by subtracting the upper limit of the optimal soil moisture content from the current soil volumetric moisture content and dividing by the upper limit of the optimal soil moisture content. When the current soil volumetric moisture content is between the lower limit and the upper limit of the optimal soil moisture content, the soil moisture stress degree is set to 0. Based on the cumulative number of days from the sowing date to the current date, obtain the seedling growth stage index; Obtain the canopy image of the current seedling and input it into the pre-trained seedling wilting recognition model to obtain the canopy wilting characterization value.
3. The intelligent irrigation method for medicinal herb seedlings based on reinforcement learning according to claim 1, characterized in that, Based on the state-aware parameters and the cumulative water stress experienced by the seedlings, the estimated value of the current seedling's effective component accumulation is obtained, including: Obtain the duration of soil moisture stress exceeding the preset stress threshold during each growth stage of the current seedling since sowing, and obtain the cumulative water stress sequence; The cumulative water stress sequence is input into a pre-trained stress-component assessment model to obtain the estimated value output by the stress-component assessment model, which is used as the estimated value of the current effective component accumulation in the seedling.
4. The intelligent irrigation method for medicinal herb seedlings based on reinforcement learning according to claim 3, characterized in that, The construction of the stress-component assessment model includes: Collect multiple sets of historical planting sample data. Each set of sample data includes the cumulative water stress sequence of the sample seedlings, which consists of the duration of soil water stress exceeding the preset stress threshold at each growth stage, and the actual content of effective components obtained by testing the sample seedlings after harvest. A stress-component assessment model is constructed, which takes the cumulative moisture stress sequence as input and the actual content of the effective component as output. The stress-component assessment model is trained using the multiple sets of historical planting sample data until it converges, thus obtaining the completed stress-component assessment model.
5. The intelligent irrigation method for medicinal herb seedlings based on reinforcement learning according to claim 1, characterized in that, The state-aware parameters and the accumulated effective component estimate are input into a pre-trained irrigation decision model to obtain irrigation action decisions, wherein the irrigation action decisions include irrigation duration and irrigation intensity, including: The soil moisture stress degree, the seedling growth stage index, the canopy wilting characterization value, and the effective component accumulation prediction value are combined into a multidimensional state vector; The multidimensional state vector is input into the pre-trained irrigation decision model to obtain the output irrigation action decision, wherein the irrigation action decision includes irrigation duration and irrigation intensity.
6. The intelligent irrigation method for medicinal herb seedlings based on reinforcement learning according to claim 5, characterized in that, The training of the irrigation decision model includes: Multiple sets of historical irrigation records were collected. Each set of historical irrigation event records included sample state perception parameters and sample effective component accumulation prediction values collected before irrigation, as well as sample irrigation action decisions to be evaluated during actual irrigation. Updated soil moisture stress, updated seedling growth stage index, updated canopy wilting characterization value, and updated effective component accumulation value were obtained after irrigation. Fitness analysis was performed on the updated soil moisture stress degree, updated seedling growth stage index, updated canopy wilting characterization value, and updated effective component accumulation value. The irrigation duration and irrigation intensity with the highest fitness were extracted as the sample irrigation action decision. Using the sample state perception parameters and the sample effective component accumulation prediction as inputs, and the sample irrigation actions as supervision labels, the irrigation decision model is trained until convergence, thus obtaining a pre-trained irrigation decision model.
7. The intelligent irrigation method for medicinal herb seedlings based on reinforcement learning according to claim 6, characterized in that, Fitness analysis was performed on the updated soil moisture stress, updated seedling growth stage index, updated canopy wilting characterization value, and updated effective component accumulation value. The irrigation duration and irrigation intensity with the highest fitness were extracted as the basis for sample irrigation action decisions, including: Calculate the first absolute difference between the updated soil moisture stress degree and the preset ideal stress degree, calculate 1 minus the first absolute difference and multiply by the first weighting coefficient to obtain the first fitness component; Calculate the second absolute difference between the updated seedling growth stage index and the preset ideal stage index, calculate 1 minus the second absolute difference and multiply by the second weighting coefficient to obtain the second fitness component; The third fitness component is obtained by multiplying the updated canopy wilting characterization value by the third weighting coefficient after subtracting the value from the calculation. Multiply the updated effective ingredient accumulation value by the fourth weighting coefficient to obtain the fourth fitness component; The first fitness component, the second fitness component, the third fitness component, and the fourth fitness component are added together to obtain the fitness of the irrigation action decision to be evaluated in the sample. The sample irrigation action decision with the highest overall fitness was selected as the sample irrigation action decision.
8. The intelligent irrigation method for medicinal herb seedlings based on reinforcement learning according to claim 3, characterized in that, Based on the updated state perception parameters and the updated cumulative water stress level, the accumulated value of the updated effective components is obtained, and a reward value is calculated. The irrigation decision model is then updated using the reward value, including: The updated cumulative water stress level is combined with the cumulative water stress sequence to form an updated cumulative water stress sequence; The updated cumulative water stress sequence is input into the stress-component assessment model to obtain the updated effective component accumulation value; The difference between the updated effective component accumulation value and the effective component accumulation estimate is calculated and recorded as the component increase. When the component increase is positive and the soil moisture stress before irrigation is in the preset stress induction range, the component increase is multiplied by the preset component incentive coefficient to obtain the first reward component. Calculate the difference between the updated canopy wilting characterization value and the preset wilting warning threshold to obtain the wilting safety margin; When the wilting safety margin is positive, the wilting safety margin is multiplied by a preset safety reward coefficient to obtain the second reward component; When the wilting safety margin is negative, the wilting safety margin is multiplied by a preset safety penalty coefficient to obtain a second reward component; Calculate the absolute value of the improvement in soil moisture stress before and after irrigation to obtain the stress improvement amount. Divide the stress improvement amount by the product of irrigation duration and irrigation intensity in this irrigation action decision, and multiply by a preset efficiency reward coefficient to obtain the third reward component. If the volumetric moisture content of the updated soil after irrigation is greater than the upper limit of the optimal soil moisture content, the preset root rot risk penalty value is obtained as the fourth reward component; otherwise, the root rot risk penalty value is set to 0. The first reward component, the second reward component, the third reward component, and the fourth reward component are added together, and the sum is taken as the reward value. The state perception parameters before irrigation, the irrigation action decision, the reward value, and the updated state perception parameters after irrigation are combined into a complete experience sample and stored in the experience replay pool. When the number of complete experience samples stored in the experience replay pool reaches the preset number of update batches, random sampling is performed from the experience replay pool, and the parameters of the irrigation decision model are updated using a reinforcement learning algorithm.
9. The intelligent irrigation method for medicinal herb seedlings based on reinforcement learning according to claim 8, characterized in that, The updated cumulative water stress level is combined with the cumulative water stress sequence to form an updated cumulative water stress sequence, including: Obtain the cumulative water stress sequence and the updated cumulative water stress level; The updated cumulative water stress level is added to the cumulative water stress sequence corresponding to the growth stage to form an updated cumulative water stress sequence.
10. A smart irrigation system for Chinese medicinal herb seedlings based on reinforcement learning, characterized in that, The system is used to implement the intelligent irrigation method for medicinal herb seedlings based on reinforcement learning as described in any one of claims 1-9, the system comprising: The state perception module is used to acquire the state perception parameters of the current seedling cultivation environment. The state perception parameters include soil moisture stress degree, seedling growth stage index and canopy wilting characterization value. The effective component assessment module is used to obtain the estimated value of the current seedling's effective component accumulation based on the state perception parameters and the current cumulative water stress experienced by the seedling. An irrigation decision module is used to input the state-aware parameters and the effective component accumulation estimate into a pre-trained irrigation decision model to obtain irrigation action decisions, wherein the irrigation action decisions include irrigation duration and irrigation intensity. The irrigation execution module is used to control the irrigation equipment to perform irrigation according to the irrigation action decision, and to obtain the updated status perception parameters and the updated cumulative water stress degree after irrigation. The online update module obtains the accumulated value of the effective components for updating based on the update status perception parameters and the cumulative water stress level, calculates the reward value, and uses the reward value to update the irrigation decision model online.