A big data-based intelligent optimization method for asexual propagation of carex rhizomes

By using big data intelligent optimization methods, combined with multi-source sensing and reinforcement learning, the problems of environmental coupling and temporal dependence in the asexual reproduction of Kentucky bluegrass rhizomes were solved. Stable and safe optimization of water and fertilizer resources and strategy iteration were achieved, improving the consistency of production results and the utilization efficiency of historical data.

CN121980973BActive Publication Date: 2026-07-03LANZHOU UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-04-07
Publication Date
2026-07-03

AI Technical Summary

Technical Problem

During the asexual reproduction of Kentucky bluegrass rhizomes, environmental conditions are highly coupled and regulatory behaviors have obvious time-series dependence. Existing methods are difficult to achieve stable optimization and cross-seasonal strategy iteration under multiple constraints, and lack the structured utilization of high-quality decision sequences from historical reproduction processes.

Method used

We construct an intelligent optimization method based on big data. Through a collaborative system of multi-source field perception, edge computing, and central intelligent decision-making, we uniformly model soil moisture, nutrient supply, light conditions, and historical operation information. We introduce off-policy trajectory prefix modeling and reinforcement learning training to generate high-quality decision trajectory prefixes, thereby achieving collaborative optimization and strategy iteration of water and fertilizer input and agricultural machinery operation.

Benefits of technology

It enhances the continuous perception and time-series response capabilities of reproductive regulation decisions, improves the overall consistency of key output indicators and the utilization rate of historical data, realizes resource optimization under security constraints and rapid convergence of strategy models, and supports stable production across seasons.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121980973B_ABST
    Figure CN121980973B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of propagation regulation, and provides an intelligent optimization method for asexual propagation of early grass rhizomes based on big data, which is applied to a rhizome asexual propagation production scene and runs in a closed-loop regulation system composed of a field perception device, an edge computing gateway, a central intelligent decision server, and water, fertilizer and agricultural machinery execution devices. The method constructs a propagation state vector based on multi-source field perception data, models the rhizome asexual propagation regulation process as a long-term time sequence decision problem with resource and safety constraints, and introduces a prefix-guided reinforcement learning training mechanism by screening historical off-policy trajectories, so as to realize the collaborative optimization of irrigation, fertilization and agricultural machinery operation under the premise of meeting the device executability and agronomic safety constraints, thereby improving the stability and resource utilization efficiency of the asexual propagation of early grass rhizomes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of reproductive regulation technology, and in particular to an intelligent optimization method for asexual propagation of Kentucky bluegrass rhizomes based on big data. Background Technology

[0002] Kentucky bluegrass, a typical cool-season turfgrass and forage grass, is typically produced on a large scale using rhizomatous asexual propagation in turfgrass production bases, forage grass propagation bases, and agricultural research experimental fields. This propagation method is highly sensitive to soil moisture status, nutrient supply structure, light conditions, temperature environment, and the timing of operations such as mowing, mulching, and compaction. The propagation process exhibits characteristics of multi-source environmental factor coupling, multi-decision variable linkage, temporal nonlinear evolution, and significant spatial differences. In actual production, output indicators such as rhizome extension length, number of tillering nodes, seedling density, seedling uniformity, and aboveground biomass dynamically change with the propagation stage and environmental conditions. There is a clear stage dependence and lag effect between water and fertilizer input intensity and operational parameters, making it difficult for a single model or static strategy to uniformly characterize the complete propagation cycle. With the application of field sensing devices and agricultural information systems, multi-source sensor data is gradually accumulating. However, existing methods mostly focus on the analysis and prediction of single environmental indicators or short-term states, lacking systematic modeling of the mapping relationship between reproductive regulation behavior and long-term output returns. In particular, they lack a structured utilization mechanism for high-quality decision sequences in historical reproductive processes, making it difficult to achieve stable optimization and cross-seasonal strategy iteration of the entire process of rhizome asexual reproduction under multiple constraints. Summary of the Invention

[0003] This invention addresses the challenges of highly coupled environmental states, significant time-dependent regulatory behaviors, and high sensitivity of reproductive output to key operational stages during the asexual propagation of Kentucky bluegrass rhizomes. It proposes a big data-based intelligent optimization method for Kentucky bluegrass rhizome asexual propagation. Through a closed-loop system integrating multi-source field sensing, edge computing, and central intelligent decision-making, it unifies the modeling of multi-dimensional information such as soil moisture, nutrient supply, light conditions, meteorological drivers, and operational history. This constructs a reproductive state vector that continuously characterizes the evolution of the reproductive process, formalizing the regulation problem of Kentucky bluegrass rhizome asexual propagation as a long-term time-series decision-making problem with resource and security constraints. Furthermore, it introduces off-strategy trajectories oriented towards the reproductive task. The prefix modeling and guidance mechanism filters and structures the decision trajectories with stable reproductive effects and good spatial consistency in historical production seasons. It extracts the trajectory prefixes of key stages such as irrigation strategy switching, fertilizer ratio adjustment, pruning operation window determination, and stubble height setting, and embeds them as conditional information into the reinforcement learning training process. Through prefix gradient masking and prefix-unprefix task dynamic matching training, the strategy model can maintain its ability to autonomously optimize subsequent decision stages while inheriting high-quality historical reproductive experience. Thus, under the premise of meeting equipment executability and agronomic safety constraints, it can achieve synergistic optimization of water and fertilizer input and agricultural machinery operation, and support cross-seasonal rhizomatous asexual reproduction strategy iteration updates.

[0004] This invention provides an intelligent optimization method for the asexual propagation of Kentucky bluegrass rhizomes based on big data, executed in a closed-loop control system consisting of field sensing devices, edge computing gateways, a central intelligent decision server, and water, fertilizer, and agricultural machinery execution equipment; the method includes the following steps:

[0005] Step S1: Multi-source breeding data collection and calibration: Field sensing devices are deployed in the breeding plots according to a preset grid to collect data on soil moisture content, soil temperature, soil electrical conductivity, nitrate nitrogen, available phosphorus, available potassium, photosynthetically active radiation, air temperature and humidity, and rainfall from the mother plant nursery and propagation nursery. Device number binding, sampling frequency registration, and spatial location calibration are performed on each data source to form raw breeding data with plot ID and batch ID.

[0006] Step S2: Construction of reproductive state vector: On the edge computing gateway side, the original reproductive data is processed by timestamp alignment, outlier removal, missing data completion, unit unification and interval normalization. Based on the sliding window, the cumulative irrigation amount, cumulative fertilizer amount, near window mean, near window change rate and intraday amplitude index are calculated and fused to form a reproductive state vector for rhizome asexual reproduction decision.

[0007] Step S3: Reproduction output reward modeling: During the reproduction cycle, collect the output indicators of rhizome asexual reproduction and form an evaluation quantity. The central intelligent decision server maps the output indicators of rhizome asexual reproduction into reward signals. The reward signals include output improvement items and resource cost penalty items. The resource cost penalty items include water consumption penalty, fertilizer application penalty and threshold safety penalty items. The threshold safety penalty items are triggered when any of the following occurs: soil moisture content exceeds the upper limit threshold, electrical conductivity exceeds the salt damage threshold, or nitrate nitrogen exceeds the leaching risk threshold.

[0008] Step S4: Definition of Reproduction Action Constraints: Define the control variables of water, fertilizer and agricultural machinery execution equipment as reproduction regulation actions, and establish a set of action constraints;

[0009] Step S5: Off-Strategy Trajectory Prefix Generation: The central intelligent decision server constructs an off-strategy trajectory library based on historical production season operation log data and field multi-source sensor records. Each off-strategy trajectory in the library consists of a reproductive state sequence composed of reproductive state vectors at multiple decision moments corresponding to the asexual reproduction process of Kentucky bluegrass rhizomes, arranged in chronological order, and a reproductive action sequence composed of reproductive control actions corresponding to each reproductive state vector, arranged in chronological order. The trajectory periodic reward value calculated from the recorded reward signal is also considered. Each trajectory in the off-strategy trajectory library is filtered according to preset reproductive effect indicators to obtain a set of effective trajectories. Reproductive effect indicators include one or more of the following: rhizome expansion rate, seedling density per unit area, tillering node formation rate, seedling uniformity, and fertilizer application gain. A reproductive task is defined, and the correct trajectory is selected for each reproductive task from the set of effective trajectories, and the correct trajectory is extracted. The breeding action sequence in the track is encoded into an operation label sequence. A candidate prefix length set is constructed based on the operation label sequence. In the candidate prefix length set, the operation label sequence is truncated according to different prefix lengths to obtain the trajectory prefix segment. The trajectory prefix segment is concatenated with the task representation of the corresponding breeding task to obtain the prefixed task representation. A prefix problem set is generated by setting multiple different prefix lengths. The prefix length is determined according to the following rules: a baseline strategy model is established, and the periodic return prediction value of the baseline strategy model under the given prefix segment condition reaches a preset threshold as the prefix selection condition. A segmented candidate prefix length selection mechanism is adopted. The periodic return prediction value of the baseline strategy model under the corresponding prefix condition is calculated in the candidate prefix length set. The candidate prefix length interval where the periodic return prediction value increases significantly is selected as the effective prefix length interval for generating the prefix problem set, and the prefix length is determined.

[0010] Step S6: Prefix-guided policy training: Construct a reinforcement learning model that includes a policy network and a value network; during the training phase, simultaneously use the prefix task set and the unprefix task set as input to perform policy sampling, obtain the corresponding reward and advantage estimation results, and update the policy network parameters through the prefix-guided mechanism and the dynamic training ratio of prefix-unprefix tasks to obtain the trained policy network.

[0011] Step S7: Online Reproduction Control Execution: During the online operation phase, the central intelligent decision server calls the trained strategy network to output the current reproduction control action. After the feasibility of the action constraint set is verified, it is sent to the solenoid valve, frequency converter pump, valve group, fertilizer pump, fertilizer mixer and operation scheduling terminal to complete the execution of irrigation, fertilization, pruning, soil covering and compaction, and write the execution results and output indicators back to the off-strategy trajectory database.

[0012] Furthermore, for each off-strategy trajectory in the off-strategy trajectory library, a reproductive consistency threshold is introduced for screening, based on meeting the reproductive effect index. The reproductive consistency threshold is calculated by jointly calculating the seedling density dispersion index and seedling uniformity dispersion index of different grid units within the same plot. Only when the corresponding off-strategy trajectory meets both the reproductive effect index and the consistency requirement under the constraint of the reproductive consistency threshold is the off-strategy trajectory determined as a valid trajectory and used for subsequent off-strategy prefix construction, thereby ensuring that the selected off-strategy prefix comes from the asexual reproduction process of rhizomes with stable spatial distribution and repeatable results.

[0013] Furthermore, step S6 specifically includes the following:

[0014] Step S61: Obtain the unprefixed task set corresponding to the prefixed question set; Based on the prefixed question set and the unprefixed task set, construct a reinforcement learning model containing a policy network and a value network. The policy network is used to generate subsequent asexual reproduction control actions of Kentucky bluegrass rhizomes under given reproduction task conditions, and the value network is used to estimate the reproduction cycle reward under the current policy conditions, thereby providing a unified model structure foundation for subsequent policy updates based on prefix guidance.

[0015] Step S62: Based on the reinforcement learning model in step S61, during the training phase, policy sampling is performed simultaneously using the prefix problem set and the unprefixed task set as inputs to obtain the corresponding periodic reward and advantage estimation results. When updating the policy for the training samples corresponding to the prefix problem set, a prefix gradient masking mechanism is introduced. A zero-weight mask is set for the loss term corresponding to the trajectory prefix segment, so that the gradient is not backpropagated through the prefix segment. The policy loss, value loss, and entropy regularization term are calculated only for the subsequent decision segments generated after the prefix segment. Combined with the update results of the unprefixed task samples, the parameters of the policy network and the value network are jointly updated, thereby enhancing the effective learning signal under the prefix guidance condition.

[0016] Step S63: During the joint policy update process in step S62, the sampling ratio of the prefix problem set and the unprefixed task set in training is dynamically adjusted based on the progress status of the training phase. In the early stage of training, the sampling ratio of the prefix problem set is increased to improve the efficiency of obtaining high-return samples. In the middle and later stages of training, the sampling ratio of the unprefixed task set is gradually increased to strengthen the autonomous decision-making ability of the learning model under the unprefixed condition. After meeting the preset training round conditions, the policy network trained by prefix guidance and dynamic matching is output as the final decision model of the intelligent optimization method for the asexual reproduction of Kentucky bluegrass rhizomes.

[0017] Furthermore, the strategy trajectory library is stored using a seasonal and plot-based index structure. After each breeding season, the central intelligent decision server applies the same filtering rules to the newly added trajectories to update the set of correct trajectories. Based on the updated set of correct trajectories, a prefix problem set is regenerated for reinforcement learning model training in the next breeding season, thereby achieving iterative optimization of the cross-seasonal rhizomatous asexual reproduction strategy.

[0018] By adopting the above solution, the beneficial effects achieved by the present invention are as follows:

[0019] This invention constructs a multi-source reproductive state vector for the entire process of Kentucky bluegrass rhizome asexual reproduction, achieving a unified representation of soil moisture, nutrient supply, light conditions, meteorological drivers, and historical water and fertilizer application behaviors. This enhances the continuous perception and temporal response capability of reproductive regulation decisions to changes in environmental conditions, solving the problems of fragmented reproductive state characterization, disjointed decision-making basis, and difficulty in covering the complete reproductive cycle in existing methods. As a result, it enhances the overall consistency of regulation of key output indicators such as rhizome expansion, tillering node formation, seedling density, and seedling uniformity, enabling the Kentucky bluegrass rhizome asexual reproduction process to maintain stable and repeatable production effects in different plots and different reproductive stages.

[0020] This invention, by introducing a strategy trajectory prefix modeling and prefix-guided reinforcement learning training mechanism, achieves effective inheritance of key decision segments from historical high-quality reproduction processes. This improves the learning efficiency and decision reliability of the strategy model in the early stages of reproduction and critical operation phases, solving the problems of long reproduction cycles, high costs of online trial and error, and low utilization of historical data. As a result, it enhances the strategy model's ability to quickly converge and generate high-return reproduction control schemes under limited training conditions. This invention enables the full utilization of the big data value accumulated from historical production seasons without increasing production risks, providing stable and controllable intelligent decision support for the asexual reproduction of Kentucky bluegrass rhizomes.

[0021] This invention integrates the improvement of rhizome asexual reproduction output with safety constraints such as water and fertilizer input intensity, salt stress, and nitrogen leaching risk into reward modeling. By combining action constraint sets and prefix gradient masking mechanisms, it achieves synergistic optimization and safety control of water, fertilizer, and agricultural machinery operations, improving resource utilization efficiency and the safety of the reproduction process. It solves the problem of balancing high output targets with resource constraints and safety thresholds, thereby enhancing the stability and applicability of this invention in continuous application across multiple seasons and plots, enabling it to support the intelligent, refined, and sustainable operation of Kentucky bluegrass rhizome asexual reproduction production in the long term. Attached Figure Description

[0022] Figure 1 This is a flowchart illustrating an intelligent optimization method for asexual propagation of Kentucky bluegrass rhizomes based on big data, as proposed in this invention.

[0023] Figure 2 This is a graph showing the training ratio of the prefixed task and the unprefixed task as a function of iteration, as proposed in Example 4. Detailed Implementation

[0024] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.

[0025] Example 1, according to Figure 1 This invention provides an intelligent optimization method for the asexual propagation of Kentucky bluegrass rhizomes based on big data. This method is applied to Kentucky bluegrass rhizome asexual propagation production scenarios in turf production bases, forage propagation bases, or agricultural research experimental fields. It is executed within a closed-loop control system consisting of field sensing devices, an edge computing gateway, a central intelligent decision server, and water, fertilizer, and agricultural machinery execution equipment. The field sensing devices are located in the mother plant nursery and propagation nursery and include soil moisture sensors, soil temperature sensors, soil conductivity sensors, soil nitrate nitrogen sensors, soil available phosphorus sensors, soil available potassium sensors, photosynthetically active radiation sensors, air temperature and humidity sensors, and rainfall meters. The water, fertilizer, and agricultural machinery execution equipment includes solenoid valves, variable frequency pumps, drip / sprinkler irrigation switching valve groups, fertilizer pumps, fertilizer mixers, lawnmowers, soil covering devices, and compaction devices. The edge computing gateway has data acquisition and communication interfaces, and the central intelligent decision server has model training and online inference capabilities. The method includes the following steps:

[0026] Step S1: Multi-source breeding data collection and calibration: Field sensing devices are deployed in the breeding plots according to a preset grid to collect data on soil moisture content, soil temperature, soil conductivity, nitrate nitrogen, available phosphorus, available potassium, photosynthetically active radiation, air temperature and humidity, and rainfall from the mother plant nursery and propagation nursery. Device number binding, sampling frequency registration, and spatial location calibration are performed on each data source to form raw breeding data with plot ID and batch ID. The edge computing gateway acquires the sensor data through RS485 and writes the collected data into the local cache unit, while simultaneously uploading it to the central intelligent decision server via Ethernet.

[0027] Step S2: Construction of the reproductive state vector: On the edge computing gateway side, the original reproductive data is processed by timestamp alignment, outlier removal, missing data completion, unit unification and interval normalization. Based on the sliding window, the cumulative irrigation amount, cumulative fertilizer amount, near window mean, near window change rate and intraday amplitude index are calculated and fused to form a reproductive state vector for rhizome asexual reproduction decision. The reproductive state vector includes soil moisture state, soil temperature state, salt stress state, nitrogen, phosphorus and potassium supply state, light state, air evapotranspiration driven state, rainfall infiltration state, historical water and fertilizer input state and pruning operation state.

[0028] Step S3: Reproductive Output Reward Modeling: During the reproductive cycle, rhizome asexual reproduction output indicators are collected and evaluated. These indicators include the increase in rhizome length, the increase in the number of tillers per unit area, seedling density, seedling uniformity, and the increase in aboveground biomass. The central intelligent decision server maps these indicators to reward signals, which include output enhancement items and resource cost penalties. Resource cost penalties include water consumption penalties, fertilizer application penalties, and threshold safety penalties. Threshold safety penalties are triggered when soil moisture content exceeds the upper limit threshold, electrical conductivity exceeds the salt damage threshold, or nitrate nitrogen exceeds the leaching risk threshold. The increase in rhizome length and the increase in the number of tillers per unit area in the rhizome asexual reproduction output indicators are obtained through periodic sampling and counting by the rhizome observation box and image measurement. Image measurement is performed by the edge computing gateway to perform image segmentation and skeletonization calculations to obtain the rhizome length and the number of branch points.

[0029] Threshold exceeding safety penalty: If, within the current decision-making cycle, the soil moisture content exceeds the preset upper limit threshold, or the soil electrical conductivity exceeds the preset salt damage electrical conductivity threshold, or the soil nitrate nitrogen concentration exceeds the preset leaching risk threshold, then the threshold exceeding safety penalty is triggered, and a negative penalty is applied to the reward signal corresponding to the current decision-making cycle.

[0030] Step S4: Definition of Reproduction Action Constraints: Define the control variables of water, fertilizer and agricultural machinery execution equipment as reproduction regulation actions. Actions include irrigation start / stop, irrigation duration, drip / sprinkler irrigation mode selection, fertilization start / stop, fertilizer dosage, fertilizer solution concentration ratio, fertilization sequence, pruning timing, stubble height, soil covering thickness and compaction intensity. Establish a set of action constraints, which includes pump station flow limit, valve group channel limit, fertilizer pump flow limit, fertilizer solution concentration limit, stubble height upper and lower limits, operation window restrictions and plot access restrictions, so that the strategy output actions meet the requirements of equipment executability and agronomic safety.

[0031] Step S5: Off-strategy trajectory prefix generation: The central intelligent decision server constructs an off-strategy trajectory library based on historical production season operation log data and field multi-source sensor records. Each off-strategy trajectory in the library consists of a sequence of reproductive state vectors at multiple decision points corresponding to the asexual reproduction process of *Poa annua*, arranged in chronological order; and a sequence of reproductive action sequences corresponding to each reproductive state vector, arranged in chronological order. The trajectory's periodic reward value is also linked to the reward signal. Each trajectory in the off-strategy trajectory library is filtered according to preset reproductive effect indicators to obtain a set of effective trajectories. The performance indicators include: root and stem expansion rate, seedling density per unit area, tillering node formation rate, seedling uniformity, and fertilizer application gain. A propagation task is defined; for each propagation task, the correct trajectory is selected from the effective trajectory set, and the propagation action sequence is extracted from that correct trajectory. The propagation action sequence is encoded as an operation label sequence; a candidate prefix length set is constructed based on the operation label sequence; within the candidate prefix length set, the operation label sequence is truncated according to different prefix lengths to obtain trajectory prefix segments; the trajectory prefix segments are concatenated with the corresponding propagation task's task representation to obtain a prefixed task representation; by setting multiple different prefix lengths, a prefix problem set is generated.

[0032] Any off-policy trajectory in the off-policy trajectory library is represented as:

[0033] ;

[0034] in, Indicates the trajectory number; Indicates the distance from the first trajectory in the strategy trajectory library The strategy trajectory is used to describe the continuous decision-making process of a certain Kentucky bluegrass rhizome asexual reproduction task within a complete reproduction cycle or decision-making cycle. Indicates the first The trajectory of the policy is separated in the first step. The reproductive state corresponding to each decision moment Indicates the first The first of the strategies At each decision point, the system considers the reproductive state. The regulatory actions performed Represents state-action pairs;

[0035] The correct trajectory prefix is ​​concatenated with the corresponding task representation of the reproduction task to form a prefixed task representation:

[0036] ;

[0037] in, This indicates a task representation with a prefix. This indicates the concatenation operator. Indicates the task of reproduction. Representing the reproductive task Task representation; Indicates the prefix length. Representation and reproductive tasks The corresponding correct trajectory operation marker sequence prefix, i.e., from the first token to the second token. A truncated fragment of a token;

[0038] The prefix length is determined according to the following rules: A baseline strategy model is established, and the predicted periodic return value of the baseline strategy model under a given prefix segment condition reaches a preset threshold as the prefix selection condition. A segmented candidate prefix length selection mechanism is adopted. The predicted periodic return value of the baseline strategy model under the corresponding prefix condition is calculated in the candidate prefix length set. The candidate prefix length interval where the predicted periodic return value shows a significant jump is selected as the effective prefix length interval for generating the prefix problem set. The prefix length is then determined so that the truncated prefix segment coverage can significantly improve the key strategy triggering stage of subsequent Kentucky bluegrass rhizome asexual reproduction return. The key strategy triggering stage includes the irrigation strategy switching stage, the fertilizer ratio switching stage, the pruning operation window determination stage, and the stubble height determination stage. The formulas used are as follows:

[0039] ;

[0040] in, Represents the baseline strategy model. This indicates that the prefix length is Under the given conditions, the periodic return value predicted by the baseline strategy model; This indicates the subsequent generated decision completion sequence (which can be understood as the action sequence or operation mark sequence generated after the prefix). This indicates that the prefix length is The task representation with prefixes constructed at the time; Indicates the index variable at the decision time. Indicates time The reproductive state vector, Indicates time The regulatory actions; This represents an instant reward function; This indicates the end of the termination time. Represents the expectation operator; This indicates that the prefix length is condition, before The starting point of subsequent decisions after the completion of the first decision; Represents the time from the start of the decision-making process to the end of the decision-making process. The cumulative value of the real-time returns at each moment during the period is used to represent the cumulative returns of the subsequent strategy completion process;

[0041] Step S6: Prefix-guided policy training: Construct a reinforcement learning model that includes a policy network and a value network; during the training phase, simultaneously use the prefix task set and the unprefix task set as input to perform policy sampling, obtain the corresponding reward and advantage estimation results, and update the policy network parameters through the prefix-guided mechanism and the dynamic training ratio of prefix-unprefix tasks to obtain the trained policy network.

[0042] Step S7: Online Reproduction Control Execution: During the online operation phase, the central intelligent decision server calls the trained strategy network to output the current reproduction control action. After the feasibility of the action constraint set is verified, it is sent to the solenoid valve, frequency converter pump, valve group, fertilizer pump, fertilizer mixer and operation scheduling terminal to complete the execution of irrigation, fertilization, pruning, soil covering and compaction, and write the execution results and output indicators back to the off-strategy trajectory database.

[0043] Example 2 is based on Example 1. In this example, for each off-strategy trajectory in the off-strategy trajectory library, a reproductive consistency threshold is introduced for screening based on meeting the reproductive effect index. The reproductive consistency threshold is calculated by jointly calculating the seedling density dispersion index and the seedling uniformity dispersion index of different grid units within the same plot. Only when the corresponding off-strategy trajectory meets both the reproductive effect index and the consistency requirement under the constraint of the reproductive consistency threshold is the off-strategy trajectory determined as a valid trajectory and used for subsequent off-strategy prefix construction, thereby ensuring that the selected off-strategy prefix comes from the asexual reproduction process of rhizomes with stable spatial distribution and repeatable results.

[0044] Example 3, based on Example 1, introduces a time stability threshold for each off-strategy trajectory in the off-strategy trajectory library, in addition to satisfying the reproduction effect index. The time stability threshold is calculated by jointly using the fluctuation index of rhizome extension length increment and the seedling density time change rate index of the same plot in multiple consecutive observation periods, and is used to characterize the stability of rhizome asexual reproduction output in the time dimension. Only when the corresponding off-strategy trajectory simultaneously meets the reproduction effect index and time stability requirements under the constraint of the time stability threshold is the off-strategy trajectory determined as a valid trajectory and used for subsequent off-strategy prefix construction, so that the selected off-strategy prefix comes from the rhizome asexual reproduction process with stable time changes but without constraint on the spatial distribution consistency of different grid units within the same plot.

[0045] Example 4, according to Figure 2 This embodiment is based on Embodiment 2. In this embodiment, step S6 specifically includes the following:

[0046] Step S61: Obtain the unprefixed task set corresponding to the prefixed question set; Based on the prefixed question set and the unprefixed task set, construct a reinforcement learning model containing a policy network and a value network. The policy network is used to generate subsequent asexual reproduction control actions of Kentucky bluegrass rhizomes under given reproduction task conditions, and the value network is used to estimate the reproduction cycle reward under the current policy conditions, thereby providing a unified model structure foundation for subsequent policy updates based on prefix guidance.

[0047] Step S62: Based on the reinforcement learning model in step S61, during the training phase, policy sampling is performed simultaneously using the prefix problem set and the unprefixed task set as inputs to obtain the corresponding periodic reward and advantage estimation results. When updating the policy for the training samples corresponding to the prefix problem set, a prefix gradient masking mechanism is introduced. A zero-weight mask is set for the loss term corresponding to the trajectory prefix segment, so that the gradient is not backpropagated through the prefix segment. The policy loss, value loss, and entropy regularization term are calculated only for the subsequent decision segments generated after the prefix segment. Combined with the update results of the unprefixed task samples, the parameters of the policy network and the value network are jointly updated, thereby enhancing the effective learning signal under the prefix guidance condition.

[0048] Step S63: During the joint policy update process in step S62, the sampling ratio of the prefix problem set and the unprefixed task set in training is dynamically adjusted based on the progress status of the training phase. In the early stage of training, the sampling ratio of the prefix problem set is increased to improve the efficiency of obtaining high-return samples. In the middle and later stages of training, the sampling ratio of the unprefixed task set is gradually increased to strengthen the autonomous decision-making ability of the learning model under the unprefixed condition. After meeting the preset training round conditions, the policy network trained by prefix guidance and dynamic matching is output as the final decision model of the intelligent optimization method for the asexual reproduction of Kentucky bluegrass rhizomes.

[0049] In this embodiment, the scheduling relationship and stage division of the sampling ratio of prefixed tasks and unprefixed tasks within the training iteration range of 0 to 400 training rounds are as follows: Figure 2 As shown; Figure 2 The graph shows the training ratio of prefix tasks and unprefixed tasks as a function of iterations. The horizontal axis represents the number of training iterations, and the vertical axis represents the sample ratio. The blue curve represents the sampling ratio of prefix tasks, and the orange curve represents the sampling ratio of unprefixed tasks. The shaded area in the graph divides the training process into stages: the prefix reinforcement period, the mixed transition period, and the stabilization period. These stages represent the changes in the scheduling weights of prefix and unprefixed tasks by the dynamic training allocation mechanism at different training stages.

[0050] Example 5, based on Example 4, describes how the off-strategy trajectory library is stored using a seasonal and plot-based index structure. After each breeding season, the central intelligent decision server updates the set of correct trajectories by applying the same filtering rules to newly added trajectories, and regenerates a prefix problem set based on the updated set of correct trajectories for training the reinforcement learning model in the next breeding season, thereby achieving iterative optimization of the cross-seasonal rhizomatous asexual reproduction strategy.

[0051] Example 6, based on Example 4, describes a scenario where historical breeding trajectories are centrally stored in a unified trajectory pool. Before the start of a new breeding season, the central intelligent decision server directly generates a prefix question set based on all accumulated off-strategy trajectories in the unified trajectory pool and performs reinforcement learning model training. Instead of independently updating and filtering the trajectory set for different breeding seasons or different plots, the server does not perform incremental filtering and correct trajectory set updates for newly added trajectories after each breeding season. This results in the generated prefix question set containing mixed breeding process trajectories under different seasons and plot conditions, and the strategy update process does not distinguish between seasonal environmental changes and plot differences.

[0052] The present invention and its embodiments have been described above. This description is not restrictive. The accompanying drawings are only one embodiment of the present invention, and the actual structure is not limited thereto. In short, if a person skilled in the art is inspired by this description and designs a similar structure and embodiment without departing from the spirit of the present invention, such design should fall within the protection scope of the present invention.

Claims

1. A big data-based intelligent optimization method for premature grass rhizome asexual reproduction, characterized in that, The method includes the following steps: Step S1: Collect raw breeding data; Step S2: Process the original breeding data to form a breeding state sequence; the breeding state sequence includes soil moisture state, soil temperature state, salt stress state, nitrogen, phosphorus and potassium supply state, light state, air evapotranspiration driven state, rainfall infiltration state, historical water and fertilizer input state and pruning operation state. Step S3: Collect rhizome asexual reproduction output indicators, including the increase in rhizome extension length, the increase in the number of tillering nodes per unit area, seedling density, seedling uniformity, and the increase in aboveground biomass. Map the rhizome asexual reproduction output indicators into reward signals. Step S4: Define reproductive regulation actions and establish a set of action constraints; Step S5: Construct an off-strategy trajectory library; each off-strategy trajectory in the library consists of a breeding state sequence and breeding control actions, and is associated with the trajectory periodic reward value calculated from the recorded reward signal; filter each trajectory in the off-strategy trajectory library to obtain a set of valid trajectories; define breeding tasks, select the corresponding correct trajectory for each breeding task in the set of valid trajectories, and extract the breeding action sequence from the correct trajectory; encode the breeding action sequence into an operation marker sequence, construct a candidate prefix length set based on the operation marker sequence, and truncate the operation marker sequence according to different prefix lengths in the candidate prefix length set to obtain trajectory prefix segments; concatenate the trajectory prefix segments with the corresponding breeding task's task representation to obtain a prefixed task representation, and generate a prefixed problem set by setting multiple different prefix lengths; Step S6: Construct a reinforcement learning model that includes a policy network and a value network, obtain reward and advantage estimation results, and introduce a prefix guidance mechanism and a dynamic training ratio of prefix-unprefix tasks to update the parameters of the policy network, thus obtaining the trained policy network. Step S7: Call the trained policy network to output the current reproduction control action, and after the feasibility of the action constraint set is verified, send it to the solenoid valve, frequency converter pump, valve group, fertilizer pump, fertilizer mixer and operation scheduling terminal, and write the execution results and output indicators back to the off-policy trajectory library. The prefix length is determined according to the following rules: A baseline strategy model is established, and the periodic return prediction value of the baseline strategy model under the given trajectory prefix segment condition reaches a preset threshold as the prefix selection condition. A segmented candidate prefix length selection mechanism is adopted. The periodic return prediction value of the baseline strategy model under the corresponding prefix condition is calculated in the candidate prefix length set. The candidate prefix length interval where the periodic return prediction value jumps is selected as the effective prefix length interval for generating the prefix problem set, and the prefix length is determined.

2. The big data-based intelligent optimization method for the asexual reproduction of the sedge rhizome according to claim 1, characterized in that: The reward signal includes output improvement items and resource cost penalty items. The resource cost penalty items include water consumption penalty, fertilizer application penalty, and threshold safety penalty items. The threshold safety penalty items are triggered when any of the following occurs: soil moisture content exceeds the upper limit threshold, electrical conductivity exceeds the salt damage threshold, or nitrate nitrogen exceeds the leaching risk threshold.

3. The big data-based intelligent optimization method for the asexual reproduction of the sedge rhizome according to claim 1, characterized in that: Step S5 specifically includes: constructing a strategy trajectory library; wherein each strategy trajectory in the strategy trajectory library consists of a reproductive state sequence and reproductive control actions corresponding to the asexual reproduction process of Kentucky bluegrass rhizomes, and is associated with the trajectory periodic reward value calculated by recording the reward signal; filtering each trajectory in the strategy trajectory library according to preset reproductive effect indicators to obtain a set of effective trajectories; defining a reproductive task, selecting the corresponding correct trajectory for each reproductive task in the set of effective trajectories, and extracting the reproductive action sequence from the correct trajectory; encoding the reproductive action sequence into an operation marker sequence, constructing a candidate prefix length set based on the operation marker sequence, and truncating the operation marker sequence according to different prefix lengths in the candidate prefix length set to obtain trajectory prefix segments; concatenating the trajectory prefix segments with the task representation of the corresponding reproductive task to obtain a task representation with a prefix, and generating a prefix problem set by setting multiple different prefix lengths.

4. The big data-based intelligent optimization method for the asexual reproduction of the sedge rhizome according to claim 3, characterized in that: The propagation effect indicators include one or more of the following: rhizome expansion rate, seedling density per unit area, tillering node formation rate, seedling uniformity, and fertilizer application gain.

5. The big data-based intelligent optimization method for the asexual reproduction of the sedge rhizome according to claim 3, characterized in that: For each off-strategy trajectory in the off-strategy trajectory library, a reproductive consistency threshold is introduced for screening based on meeting the reproductive effect index; the reproductive consistency threshold is calculated by jointly calculating the seedling density dispersion index and the seedling uniformity dispersion index of different grid units within the same plot.

6. The big data-based intelligent optimization method for the asexual reproduction of the sedge rhizome according to claim 1, characterized in that: Step S6 specifically includes the following: Step S61: Construct a reinforcement learning model that includes a policy network and a value network, wherein the policy network is used to generate subsequent asexual reproduction control actions of Kentucky bluegrass rhizomes under given reproductive task conditions, and the value network is used to estimate the reproductive cycle reward under the current policy conditions. Step S62: Obtain the unprefixed task set. Based on the reinforcement learning model in step S61, during the training phase, simultaneously use the prefixed problem set and the unprefixed task set as inputs to perform policy sampling and obtain the corresponding periodic reward and advantage estimation results. When updating the policy for the training samples corresponding to the prefixed problem set, a zero-weight mask is set for the loss term corresponding to the trajectory prefix segment through the prefix guidance mechanism, so that the gradient is not backpropagated through the trajectory prefix segment. Only the policy loss, value loss, and entropy regularization term are calculated for the subsequent decision segments generated after the trajectory prefix segment. Combined with the update results of the unprefixed task samples in the prefixed task set, the parameters of the policy network and the value network are jointly updated. Step S63: During the joint policy update process in step S62, the sampling ratio of the prefix problem set and the unprefixed task set in training is dynamically adjusted based on the progress status of the training phase. In the early stage of training, the sampling ratio of the prefix problem set is increased to improve the efficiency of obtaining high-reward samples. In the middle and later stages of training, the sampling ratio of the unprefixed task set is gradually increased to enhance the autonomous decision-making ability of the learning model under the unprefixed condition. After the preset training round conditions are met, the trained policy network is output.

7. The big data-based intelligent optimization method for the asexual reproduction of the sedge rhizome according to claim 1, characterized in that: The departure strategy trajectory library is stored using a seasonal and plot-based index structure.

Citation Information

Patent Citations

  • Crop nitrogen fertilizer management system and method based on multi-agent reinforcement learning

    CN121280165A