A farmland fertilizer and water management method based on knowledge guided training of a large model

CN122736105APending Publication Date: 2026-09-11CHINA AGRI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611200198.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-10
Publication Date
2026-09-11

AI Technical Summary

Technical Problem

这类模型以生物地球化学和作物生理学原理为基础,在单点或田块尺度具有较好的物理可解释性,已被广泛用于作物产量模拟、养分循环评估及环境效应分析,然而,现有过程模型通常针对特定过程或子系统构建,模型结构和参数体系差异较大,同一模型难以同时兼顾作物生长、水分利用和氮素损失等多个关键过程,在大尺度应用时往往依赖大量经验参数和假设,区域泛化能力有限

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122736105A_ABST
    Figure CN122736105A_ABST
Patent Text Reader

Abstract

This invention belongs to the field of agricultural information management technology and provides a method for farmland fertilizer and water management based on a knowledge-guided training model. The method includes the following steps: S1, collecting historical crop data; S2, adjusting crop process model parameters to generate a synthetic model, inputting historical data into the synthetic model to generate simulated data, and generating synthetic data based on the validated synthetic model; S3, using historical and synthetic data to train a multi-task model to form a farmland fertilizer and water management model, constructing multiple dedicated prediction subnetworks, and strengthening the causal relationship between tasks through yield feature coupling; S4, collecting actual meteorological and soil data and inputting them into the farmland fertilizer and water management model to obtain farmland fertilizer and water management strategies. The beneficial effects of this invention are: by jointly driving the model with simulated and historical data, the model possesses both broad scenario adaptability and can realistically reflect the operational characteristics of the farmland system, significantly improving the model's generalization performance in cross-regional and cross-crop applications.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of agricultural information management technology, and in particular relates to a method for farmland fertilizer and water management based on a knowledge-guided training large model. Background Technology

[0002] Currently, technologies related to farmland nutrient management and fertilization decisions mainly include mechanism-based farmland process models, data-driven machine learning models, and integrated methods that combine the two.

[0003] Mechanism-based farmland process models are widely used in agricultural production simulation and management decision-making. The crop growth model APSIM is mainly used to simulate crop growth, development, and yield formation processes. The soil carbon and nitrogen cycle model DNDC focuses on the transformation of carbon and nitrogen elements and greenhouse gas emissions in farmland. The water management model AquaCrop focuses on characterizing the impact of crop water consumption and water stress on yield. These models are based on biogeochemical and crop physiological principles and have good physical interpretability at the single-point or field scale. They have been widely used for crop yield simulation, nutrient cycling assessment, and environmental effect analysis. However, existing process models are usually built for specific processes or subsystems, and the model structures and parameter systems vary greatly. It is difficult for the same model to simultaneously take into account multiple key processes such as crop growth, water use, and nitrogen loss. When applied at a large scale, they often rely on a large number of empirical parameters and assumptions, and their regional generalization ability is limited.

[0004] Data-driven machine learning models have seen rapid development in the agricultural field in recent years. Methods such as random forests and deep neural networks have been used for crop yield prediction, nitrogen loss estimation, and fertilization decision support. Their advantages lie in high computational efficiency and strong fitting ability for complex nonlinear relationships. However, such models are highly dependent on historical measured data and are difficult to maintain stable predictive performance under conditions of insufficient samples or large changes in management scenarios. At the same time, the model results lack clear biophysical constraints, have weak interpretability, and are difficult to use directly for refined farmland management decisions.

[0005] Existing farmland fertilizer and water management technologies suffer from the following problems: First, existing single models or simple model combinations cannot simultaneously reflect the intrinsic relationship between crop yield, water use efficiency, and nitrogen loss. The limited sources of training data for these models result in insufficient predictive ability under complex management scenarios, especially prone to systematic biases in large-scale applications. Second, relying solely on mechanistic models to simulate data can easily introduce structural biases, while relying solely on measured data is limited by sample size and scenario coverage, making it difficult for models to maintain stability under different regions and crop conditions. Third, existing machine learning models lack clear biophysical constraints, making them prone to producing predictive results that do not conform to agronomic common sense under conditions of sparse data or changing scenarios. Their insufficient model stability and interpretability limit their credibility in farmland management decision-making. Summary of the Invention

[0006] In view of this, the present invention aims to propose a farmland fertilizer and water management method based on knowledge-guided training of a large model, in order to solve at least one of the above-mentioned technical problems.

[0007] To achieve the above objectives, the technical solution of the present invention is implemented as follows:

[0008] The first aspect of this invention provides a method for farmland fertilizer and water management based on a knowledge-guided training large model, comprising the following steps:

[0009] S1. Collect historical data on crops;

[0010] S2. Adjust the parameters of the crop process model to generate a synthetic model. Input historical data into the synthetic model to generate simulation data. Validate the synthetic model through simulation data. Generate synthetic data based on the validated synthetic model.

[0011] S3. Use historical data and synthetic data to train a multi-task model to form a farmland fertilizer and water management model. Multiple dedicated prediction subnets are constructed in the multi-task model, and the causal relationship between tasks is strengthened by coupling yield characteristics.

[0012] S4. Collect actual meteorological and soil data and input them into the farmland fertilizer and water management model to obtain farmland fertilizer and water management strategies.

[0013] Furthermore, the historical data of crops in S1 includes meteorological data, soil data, and management measure data;

[0014] Meteorological data includes: daily maximum and minimum temperatures, rainfall, wind speed, relative humidity, and sunshine duration.

[0015] Management data includes: sowing date, harvest date, seeding rate, yield, and management data;

[0016] Soil data include: soil organic carbon, total nitrogen, pH, bulk density, and clay content.

[0017] Furthermore, S2 includes the following steps:

[0018] S21. Divide the historical data of crops into a training set and a validation set;

[0019] S22. Adjust the parameters in the crop process model through trial and error, and input the training set data into the crop process model to obtain simulation data and key variety parameters;

[0020] Process models include crop growth models, soil carbon and nitrogen cycle models, and water management models;

[0021] S23. Use the validation set to validate the simulated data through linear regression or index evaluation method to obtain the key variety parameters that have passed the validation. Adjust the crop process model according to the key variety parameters to form a synthetic model.

[0022] S24. Collect meteorological data and soil data of crops at different locations and input them into the synthetic model to form synthetic data.

[0023] Furthermore, S3 includes the following steps:

[0024] S31. The model input data is divided into categorical factor sets and continuous factor sets. The model output task is defined, and the combination is described as a triple structure. Missing value labeling and mask generation are performed.

[0025] S32. Fit the preprocessed data transformer and save it persistently.

[0026] The categorical and continuous features defined in S33 and S31 are transformed into 256-dimensional input features;

[0027] S34. Based on a multi-head self-attention stacked architecture, a 6-layer Transformer encoder backbone is constructed to extract deep, globally interactive shared feature representations from 256-dimensional input features.

[0028] S35. The shared features obtained in S34 are divided into 8 task feature representations. Cross-task attention is used to perform inter-task information interaction, and weighted fusion is used to obtain the final features.

[0029] S36. Construct 8 dedicated prediction subnets to strengthen the causal relationship between tasks through output feature coupling and add physical constraints;

[0030] S37. Construct a loss function to transform agronomic physical laws and empirical relationships into computable constraint loss terms;

[0031] S38. The trainable range of the backbone layer is controlled by freezing the ratio, and the differentiated learning rate is achieved by grouping parameters. The farmland fertilizer and water management model is formed through iterative training.

[0032] Furthermore, S35 includes the following steps:

[0033] S351. Construct 8 task encoders for each task. The 8 encoders have the same structure but independent parameters, and map the shared representation to the 8 task feature representations.

[0034] S352. Perform cross-task attention interaction, stack the 8 task feature representations along the sequence dimension to form a feature sequence, and apply a multi-head self-attention mechanism to the feature sequence;

[0035] S353. Perform a linear weighted combination of the task feature representation after cross-task attention enhancement and the shared features.

[0036] Furthermore, S36 includes the following steps:

[0037] S361. Construct a dedicated prediction subnet for each prediction task;

[0038] The forecasting tasks include forecasting crop yield, nitrogen fertilizer use efficiency, water use efficiency, nitrous oxide emissions, ammonia volatilization, nitrate leaching, soil carbon sequestration, and greenhouse gas emissions.

[0039] S362. Define the output of a certain intermediate layer in the crop yield prediction subnet as the yield feature, the final task feature of the crop yield prediction subnet as the input, and the final task features of the other special prediction subnets and the yield feature are concatenated along the feature dimension to form a concatenated vector as the input of the corresponding subnet.

[0040] S363, the crop yield prediction subnet is mapped to a 1-dimensional real value through a linear output layer;

[0041] The remaining 7 prediction subnets first obtain initial predictions through a linear output layer, and then apply the Softplus activation function for nonlinear transformation.

[0042] Furthermore, S37 includes the following steps:

[0043] S371. The Huber loss function is used as the basic loss function for each prediction task.

[0044] S372. Transform agronomic physical laws and empirical relationships into calculable constraint loss terms;

[0045] The following constraint losses are included: physical relationship fitting loss, curve slope penalty loss, and prediction internal consistency penalty loss;

[0046] S373. The Huber loss function is weighted and fused with the physical relationship fitting loss, curve slope penalty loss and prediction internal consistency penalty loss.

[0047] A second aspect of the present invention provides a knowledge-guided farmland fertilizer and water management device, comprising:

[0048] The data acquisition module is configured to collect historical data on crops, as well as actual meteorological and soil data.

[0049] The data preprocessing module is configured to adjust the parameters of the crop process model to generate a synthetic model. The historical data collected by the acquisition module is input into the synthetic model to generate simulated data. The synthetic model is verified through the simulated data, and synthetic data is generated based on the verified synthetic model.

[0050] The prediction module is configured to train a multi-task model using historical and synthetic data to form a farmland fertilizer and water management model. Multiple dedicated prediction subnets are built in the multi-task model, and the causal relationship between tasks is strengthened by coupling yield features.

[0051] The actual meteorological and soil data collected by the acquisition module are input into the farmland fertilizer and water management model to obtain farmland fertilizer and water management strategies.

[0052] A third aspect of the present invention provides a server including at least one processor and a memory communicatively connected to the processor, the memory storing instructions executable by the at least one processor, the instructions being executed by the processor to cause the at least one processor to perform the method as described in the first aspect.

[0053] A fourth aspect of the present invention provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the method described in the first aspect.

[0054] Compared with existing technologies, the farmland fertilizer and water management method based on a knowledge-guided training large model described in this invention has the following beneficial effects:

[0055] (1) The farmland fertilizer and water management method based on knowledge-guided training large model described in this invention, through the joint drive of simulated data and historical data, enables the model to have a wide range of scenario adaptability and to truly reflect the operation characteristics of the farmland system, significantly improving the generalization performance of the model in cross-regional and cross-crop applications.

[0056] (2) The farmland fertilizer and water management method based on knowledge-guided training large model described in this invention explicitly introduces the mechanistic knowledge constraints of the farmland system during the training and inference process of the model. It embeds agronomic knowledge such as material conservation, process response directionality and reasonable range of management measures into the model structure design and loss function construction, so that the model not only minimizes the prediction error during the optimization process, but also needs to meet the basic mechanistic laws of the farmland system. The model is guided by knowledge constraints during the training process, so that the prediction results of the model are always within the reasonable range of agronomy. While ensuring the prediction accuracy, it significantly improves the stability, reliability and interpretability of the model. Attached Figure Description

[0057] The accompanying drawings, which form part of this invention, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute an undue limitation of the invention. In the drawings:

[0058] Figure 1 This is a schematic diagram of the method described in an embodiment of the present invention. Detailed Implementation

[0059] It should be noted that, unless otherwise specified, the embodiments and features described in the present invention can be combined with each other.

[0060] In the description of this invention, it should be understood that the terms "center," "longitudinal," "lateral," "upper," "lower," "front," "rear," "left," "right," "vertical," "horizontal," "top," "bottom," "inner," and "outer," etc., indicating orientations or positional relationships based on the orientations or positional relationships shown in the accompanying drawings, are only for the convenience of describing the invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the invention. Furthermore, the terms "first," "second," etc., are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined with "first," "second," etc., may explicitly or implicitly include one or more of that feature. In the description of this invention, unless otherwise stated, "a plurality of" means two or more.

[0061] In the description of this invention, it should be noted that, unless otherwise explicitly specified and limited, the terms "installation," "connection," and "linking" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection of two components. Those skilled in the art will understand the specific meaning of the above terms in this invention based on the specific circumstances.

[0062] The present invention will now be described in detail with reference to the accompanying drawings and embodiments.

[0063] Example 1:

[0064] like Figure 1 As shown, a method for farmland fertilizer and water management based on knowledge-guided training of a large model includes the following steps:

[0065] S1. Collect historical data on crops;

[0066] Historical data on crops includes meteorological data, soil data, and management data;

[0067] Meteorological data includes: daily maximum and minimum temperatures, rainfall, wind speed, relative humidity, and sunshine duration.

[0068] Management data includes: sowing date, harvest date, seeding rate (sowing density), yield, and management data (tillage, irrigation, and fertilization time, method, and amount).

[0069] Soil data include: soil organic carbon, total nitrogen, pH, bulk density, and clay content.

[0070] S2. Adjust the parameters of the crop process model to generate a synthetic model. Input historical data into the synthetic model to generate simulation data. Validate the synthetic model through simulation data. Generate synthetic data based on the validated synthetic model.

[0071] S2 includes the following steps:

[0072] S21. Divide the historical data of crops into a training set and a validation set;

[0073] Specifically, historical data on crops in the same region are divided into training and validation sets;

[0074] S22. Adjust the parameters in the crop process model through trial and error. Input the training set data into the crop process model to obtain the simulation data and parameter adjustment data of key varieties.

[0075] The simulation data includes yield data, nitrogen fertilizer use efficiency data, water use efficiency data, nitrogen loss data, soil carbon sequestration data, and greenhouse gas emission data;

[0076] The process models include the crop growth model APSIM, the soil carbon and nitrogen cycle model DNDC, and the water management model AquaCrop. The simulation results of the three models in the cross-domain can be cross-verified, which enhances the credibility of the synthetic data. The synthetic data has an inherent mechanistic consistency among each stage, which is better than the synthetic data generated independently.

[0077] S23. Use the validation set to validate the simulated data through linear regression or index evaluation method to obtain the key variety parameters that have passed the validation. Adjust the crop process model according to the key variety parameters to form a synthetic model.

[0078] The process of validating linear regression on simulated data is as follows:

[0079] Linear regression judges the fit trend line by combining the points corresponding to the validation set and the simulated data. The closer the trend line is to the 1:1 line, the higher the consistency between the simulated data and the historical data.

[0080] The process of validating the simulated data using the index evaluation method is as follows:

[0081] The indicator evaluation method uses the coefficient of determination (R²) 2 The simulation data were validated using statistical indicators such as normalized root mean square error (NRMSE) and MEA.

[0082] The validated process model encapsulation can generate simulation data in batches for different scenarios.

[0083] The validation set can be used to validate the simulated data through linear regression or index evaluation, or the validation set can be used to validate both linear regression and index evaluation simultaneously, and key variety parameters that meet both criteria can be selected.

[0084] S24. Collect meteorological data and soil data of crops at different locations, and input them into the synthetic model to form synthetic data, as detailed below:

[0085] The study conducted simulations at a 10km × 10km grid scale nationwide, spanning from 2001 to 2022. Input data included climate and soil data. The study comprehensively considered three crop types (rice, wheat, and maize), five fertilizer types (urea, enhanced-efficiency fertilizers, ammonium nitrogen fertilizers, organic fertilizers, and other types of fertilizers), three nitrogen application methods (surface application, deep application, and mixed application), two soil tillage methods (conventional tillage and no-till), two nitrogen fertilizer application frequencies (single and multiple applications), nitrogen application rate, irrigation methods (flood irrigation, drip irrigation, and sprinkler irrigation), and irrigation volume. To ensure the comprehensiveness and representativeness of the synthetic data, 10,000 grid scenarios were randomly selected from each of the three crops for pre-training. Each sampling randomly selected one grid and one scenario from approximately 72,805 grids and 60 management combination scenarios. The 10,000 selected samples are almost evenly distributed across all regions of the country, which can fully represent all scenario combinations and cover agricultural production scenarios under different crops, climates, soils and management measures, thus ensuring that the synthetic database effectively reflects diverse agricultural production conditions.

[0086] Synthetic models can proactively generate simulated data for different scenarios, avoiding uncontrollable extrapolation biases in inference when the model encounters unseen scenarios. The validated synthetic models improve data accuracy, ensure data diversity, and provide data support for subsequent training. The process models within the synthetic models utilize formulas from crop physiology and soil biology; the synthetic data is generated through these formulas, further enhancing data accuracy.

[0087] S3. Use historical data and synthetic data to train a multi-task model to form a farmland fertilizer and water management model. Multiple dedicated prediction subnets are constructed in the multi-task model, and the causal relationship between tasks is strengthened by coupling yield characteristics.

[0088] Based on synthetic data and historical crop data related to crop parameters, a multi-task Transformer soil crop system fertilizer, water and nutrient management model driven by knowledge data hybridization was constructed and trained.

[0089] The mapping relationship is expressed as follows:

[0090] ;

[0091] Where X represents input characteristics such as management, climate, soil, and inputs. This is a joint prediction vector for multiple tasks. These are the model parameters.

[0092] The multi-task joint prediction vector includes: minimizing greenhouse gas emissions, minimizing farmland nutrient loss, minimizing water resource utilization, maximizing crop yield, maximizing farmers' economic benefits, maximizing nitrogen fertilizer use efficiency, maximizing water use efficiency, and maximizing soil carbon sequestration.

[0093] S3 includes the following steps:

[0094] S31. The model input data is divided into categorical factor sets and continuous factor sets. The model output task is defined, and the combination is described as a triple structure. Missing value labeling and mask generation are performed.

[0095] Synthetic data and historical data are used as training input data for the model. The training input data is divided into two feature sets, including:

[0096] Categorical factor set (used to represent management classification information):

[0097] Fertilizer application location, fertilization frequency, fertilizer type, irrigation frequency, irrigation method, tillage method, tillage time, planting density, and sowing date.

[0098] Continuous factor set (used to characterize continuous variables such as climate, soil, and inputs):

[0099] Average temperature during the growing season, average precipitation during the growing season, soil organic carbon content, total nitrogen in the soil, pH, bulk density, clay content, nitrogen application rate, and irrigation amount.

[0100] The model output is an 8-task prediction vector, including: crop yield, nitrogen use efficiency (NUE), water use efficiency (WUE), soil carbon sequestration, environmental losses of three types of reactive nitrogen (nitrous oxide emissions N2O, ammonia volatilization NH3, and nitrate leaching NO3-leaching), and greenhouse gas emissions (GHG).

[0101] Organize each data point into triples:

[0102] : Category feature vector (after discrete encoding);

[0103] Continuous feature vectors (after normalization / standardization);

[0104] The multi-task label vector (normalized / standardized for each task) is expressed as follows:

[0105] ;

[0106] Missing value tagging and mask generation (removing missing tags by task).

[0107] For each sample-task pair, construct a binary mask:

[0108] ;

[0109] All inputs are retained during training, but only the inputs are processed when calculating the loss for task k. The sample positions with a value of 1 are included, thus enabling multiple tasks to use incompletely labeled data for joint training and improving data utilization.

[0110] S32. Fit the preprocessed data transformer and save it persistently.

[0111] S321. Divide the observed data and synthetic data into training set / validation set / test set in a ratio of 8:1:1, and fix the division result with a random seed;

[0112] S322. For each category feature column, convert all values ​​into strings and uniformly map all missing values ​​to specific placeholder strings;

[0113] S323. Stabilized Category Encoder: The category encoder learns to map each unique category string to a unique integer index. After training, the encoder model is stabilized and saved to ensure that the same category value always receives the same integer encoding during subsequent model training and online inference phases.

[0114] The continuous feature scaler performs a fit standardization (e.g., by z-score standardization) and then saves it.

[0115] An independent normalizer is fitted to the output of each task, which is then solidified into a task dictionary-style scaler and trained stably under different task scale differences.

[0116] The categorical and continuous features defined in S33 and S31 are transformed into 256-dimensional input features;

[0117] S331. There are 9 category features, and their category dimensions (number of value categories) are [3, 4, 6, 4, 4, 5, 4, 5, 3]. The embedding dimension of each category feature is 16. Each category feature value will be mapped to a 16-dimensional continuous vector: [16, 16, 16, 16, 16, 16, 16, 16]. A learnable embedding lookup table will be constructed for each category feature.

[0118] S332. For the 9 categories of the sample, look up the table to obtain 9 16-dimensional vectors and concatenate them into 144-dimensional vectors. Embed the 9-dimensional continuous feature vectors and the 144-dimensional category feature vectors to form a 153-dimensional fusion feature representation.

[0119] S333: The 153-dimensional vector is mapped to hidden_dim=256 through a learnable projection layer. The 256-dimensional vector after projection is the standard input feature representation of the Transformer backbone network.

[0120] S34. Based on a multi-head self-attention stacked architecture, a 6-layer Transformer encoder backbone is constructed to extract deep, globally interactive shared feature representations from 256-dimensional input features.

[0121] The self-attention mechanism allows any two input feature locations to interact directly, capturing long-range dependencies between climate, soil, and management factors (such as the impact of the interaction between nitrogen application rate and rainfall on nitrogen leaching).

[0122] Construct a Transformer backbone (multi-head self-attention stack) and perform forward propagation. The number of layers in the Transformer backbone is L=6, the number of attention heads is num_heads=2, and the attention dropout (layers 1 to 4) is [0.1,0.15,0.2,0.25,0.25,0.25]; the QKV projection dimension of each layer is qkv_dim=[16,16,16,16,16,16]; the unified hidden dimension is hidden_dim=256 (the input and output dimensions of all layers are consistent).

[0123] The Transformer backbone is composed of the following:

[0124] Construct L Transformer coded blocks and concatenate them into a backbone. Each coded block contains, in sequence:

[0125] Multihead self-attention quantum layer (MHA);

[0126] Feedforward Network Sublayer (FFN);

[0127] Each sub-layer is followed by: residual connection + layer normalization + Dropout (self-attention sub-layers use attention_dropout), which concatenates L encoded blocks in sequence to form a complete Transformer backbone network.

[0128] Perform a forward propagation through the backbone network, inputting the data into the pre-constructed backbone network for computation. Input the standard input feature representation H0 (dimension 256) into the first encoding block, and then process it through all L encoding blocks in sequence. Each encoding block performs the following operations: multi-head self-attention computation (updating global feature dependencies), feedforward transformation computation (enhancing nonlinear expressive power), residual connection and normalization (stabilizing training). After processing through all L=6 layers, the final backbone shared feature representation H_shared is output.

[0129] S35. The shared features obtained in S34 are divided into 8 task feature representations. Cross-task attention is used to perform inter-task information interaction, and weighted fusion is used to obtain the final features.

[0130] S35 includes the following steps:

[0131] S351. Construct 8 task encoders for each task. The 8 encoders have the same structure but independent parameters. Map the shared representation to the 8 task feature representations, as follows:

[0132] For each task, a lightweight task encoder (linear transformation + dropout = 0.25) is constructed. The input to each encoder is a shared feature representation H_shared, and both the input and output dimensions are 256, generating a dedicated task Ht. The eight task encoders have the same structure but independent parameters. Parameter independence allows each task to learn different feature emphases, while also improving computational efficiency.

[0133] S352. Perform cross-task attention interaction, stack the 8 task feature representations along the sequence dimension to form a feature sequence, and apply a multi-head self-attention mechanism to the feature sequence, as follows:

[0134] The task features are stacked into a sequence: [Ht1,Ht2,Ht3,Ht4,Ht5,Ht6,Ht7,Ht8]. Cross-task multi-head self-attention is applied to this sequence with parameters: number of attention heads = 8, attention dropout = 0.1. This calculation outputs an enhanced sequence of task features and yields an inter-task correlation attention weight matrix (used to identify the information contribution relationship between tasks).

[0135] Cross-task attention enables the model to explicitly learn the relationships between tasks. The output inter-task attention weight matrix can quantify the information contribution relationship between each task, providing interpretability for the model's decisions.

[0136] S353. The task feature representation after cross-task attention enhancement is linearly weighted and combined with the shared features, with a fusion coefficient of {0.7, 0.3}.

[0137] S36. Construct 8 dedicated prediction subnets to strengthen the causal relationship between tasks through output feature coupling and add physical constraints;

[0138] S36 includes the following steps:

[0139] S361. Construct a dedicated prediction subnet for each prediction task;

[0140] The forecasting tasks include forecasting crop yield, nitrogen fertilizer use efficiency, water use efficiency, nitrous oxide emissions, ammonia volatilization, nitrate leaching, soil carbon sequestration, and greenhouse gas emissions.

[0141] The dedicated prediction subnet consists of three sequentially connected fully connected layers, each with a hidden dimension of 64, and is equipped with a 2-head self-attention module for feature refinement within the subnet. The attention dropout is 0.25 and the QKV dimension is 4.

[0142] S362. Define the output of a certain intermediate layer in the crop yield prediction subnet as the yield feature, and the final task feature of the crop yield prediction subnet as the input. The final task features of the other dedicated prediction subnets are concatenated with the yield feature along the feature dimension to form a concatenated vector, which is used as the input of the corresponding subnet, as follows:

[0143] The yield characteristic output by the intermediate layer of the crop yield subnet is represented as F. yield The input for each subnet is:

[0144] Crop Yield Subnet: Directly uses the final features of the yield task as input. .

[0145] Nitrogen fertilizer use efficiency subnet: This subnet connects the final characteristics of the nitrogen fertilizer use efficiency task with F... yield Concatenate along the feature dimension, input .

[0146] Water use efficiency subnet: This subnet connects the final characteristics of the water use efficiency task with F... yield Concatenate along the feature dimension, input .

[0147] Soil carbon sequestration subnetwork: linking the final characteristics of soil carbon sequestration tasks with F yield Concatenate along the feature dimension, input .

[0148] Nitrous oxide emission subnet: This subnet will integrate the final characteristics of nitrous oxide emission targets with F... yield Concatenate along the feature dimension, input .

[0149] Ammonia Volatilization Subnet: Connecting the final characteristics of ammonia volatilization with F yield Concatenate along the feature dimension, input .

[0150] Nitrate leaching sub-network: This sub-network connects the final characteristics of the nitrate leaching task with F... yield Concatenate along the feature dimension, input .

[0151] Greenhouse Gas Emissions Subnet: This subnet will ultimately characterize greenhouse gas emission targets in relation to F... yield Concatenate along the feature dimension, input .

[0152] Yield is the core output of crop systems. Most environmental indicators (nitrogen loss, water use, carbon sequestration) are highly coupled with yield formation processes. By explicitly splicing yield characteristics, other predictive subnets can more accurately extrapolate related indicators.

[0153] This ensures agronomic rationality. Under high-yield scenarios, ammonia volatilization is usually also high. Explicitly inputting yield features means that the model does not have to infer yield from the original features, thus improving prediction accuracy. Only yield will affect other subnets, which is consistent with agronomic causal logic.

[0154] S363, the crop yield prediction subnet is mapped to a 1-dimensional real value through a linear output layer;

[0155] The remaining seven prediction subnets (nitrogen fertilizer use efficiency subnet, water use efficiency subnet, soil carbon sequestration quantum network, nitrous oxide emission subnet, ammonia volatilization subnet, nitrate leaching subnet, and greenhouse gas emission quantum network) first obtain initial predictions through a linear output layer, and then undergo nonlinear transformation using the Softplus activation function. This ensures that the subnet output values ​​are positive, thus conforming to the physical laws of the soil-crop system and improving the accuracy of the predictions.

[0156] S37. Construct a loss function to transform agronomic physical laws and empirical relationships into computable constraint loss terms;

[0157] S37 includes the following steps:

[0158] S371. The Huber loss function is used as the basic loss function for each prediction task, as detailed below:

[0159] The Huber loss function is adopted as the basic loss function for the prediction task. The Huber loss function combines the advantages of mean squared error (MSE) and mean absolute error (MAE), which improves the robustness of the prediction results.

[0160] For prediction error The Huber loss function is defined as follows:

[0161] ;

[0162] in, Actual value For predicted values, For prediction error, This is the error threshold.

[0163] In this embodiment, δ is set to 1.35. When the error is small, the square term is used to encourage accurate prediction, and when the error is large, the linear term is used to suppress the influence of outliers.

[0164] For each task, a binary mask vector is constructed based on the validity of its label. During loss calculation, only the sample positions where the mask is true (the label is valid) are considered. If all samples of a task in a batch are missing labels, the loss term for that task is set to zero.

[0165] The task losses are summed using a weighted average of the weight vectors, with the default weights being [1,1,1,1]. When single-task reinforcement training is needed, the weights of non-target tasks can be reset to zero to switch the training focus.

[0166] S372. Transform agronomic physical laws and empirical relationships into calculable constraint loss terms;

[0167] The following constraint losses are included: physical relationship fitting loss, curve slope penalty loss, and prediction internal consistency penalty loss;

[0168] The formula for the physical relationship fitting loss is as follows:

[0169] By constraining the response relationship between nitrogen application rate (Nrate), water input, and yield and nitrogen loss, the loss function is kept within typical agronomic laws, as detailed below:

[0170] The exponential saturation response of yield to nitrogen / water input levels;

[0171] ;

[0172] ;

[0173] in, Nrate indicates crop yield, and Nrate indicates nitrogen application rate. This represents the amount of water input, where e is the natural constant. This indicates the potential for increased yield from nitrogen application. This indicates the yield without nitrogen application. This indicates the potential for increased yields through irrigation. This indicates the yield without irrigation. and These represent the rate at which yield increases with fertilization and irrigation, respectively.

[0174] Based on crop nitrogen uptake coefficient Estimating the nitrogen surplus:

[0175] ;

[0176] in, Indicates the crop nitrogen uptake coefficient. This indicates a nitrogen surplus.

[0177] The theoretical curve is generated by a typical exponential growth function;

[0178] ;

[0179] in, Reactive nitrogen loss This indicates the amount of active nitrogen lost when no nitrogen is applied. This indicates the rate at which active nitrogen increases with increasing nitrogen surplus.

[0180] The computational model's prediction of virtual samples Compared with the theoretical curve Mean square error:

[0181] ;

[0182] The formula for curve slope penalty loss is as follows:

[0183] By constraining the shape of the local response curve, abnormal fluctuations or anomalous turns can be prevented.

[0184] Virtual input gradient sequence Neighbor prediction Define adjacent slopes:

[0185] ;

[0186] Change in slope:

[0187] ;

[0188] Penalty when local changes fall below an empirical threshold:

[0189] ;

[0190] in, This represents the nitrogen application rate value corresponding to the i-th virtual nitrogen application rate. This represents the predicted target nitrogen loss obtained by the model at this nitrogen application level. This represents the slope of the change in predicted values ​​between two adjacent nitrogen application rates, and it is obtained by the ratio of the difference between adjacent predicted values ​​to the difference in nitrogen application rates. This represents the change between adjacent slopes, where `threshold` is a preset slope change threshold parameter, and `mean` represents the average of all calculation results within the virtual nitrogen application gradient interval. When the value is less than the threshold, a penalty term is applied to the corresponding interval. No penalty is imposed if the value is greater than or equal to the threshold.

[0191] The loss function for predicting internal consistency penalty is as follows:

[0192] Strengthening the intrinsic correlation between related output tasks, the "expected" nitrogen loss curve is derived by fitting a function based on yield series predictions. , compared with the curve directly predicted by the model Apply consistency penalty:

[0193] ;

[0194] S373. The Huber loss function is weighted and fused with the physical relationship fitting loss, curve slope penalty loss, and prediction internal consistency penalty loss. The fusion expression is as follows:

[0195] ;

[0196] in It is an adjustable weight used to control the intensity of knowledge guidance.

[0197] S38. The trainable range of the backbone layer is controlled by freezing the ratio, and the differentiated learning rate is achieved by grouping parameters. The farmland fertilizer and water management model is formed through iterative training.

[0198] The S38 step is as follows:

[0199] S381. Phased training, as detailed below:

[0200] Training general hyperparameters and initializing the frozen dictionary;

[0201] Training epochs: 2000, batch size: 256, base learning rate: 1.0 × 10⁻⁶. -4 Random inactivation rate 0.25, early termination patience value 10.

[0202] The initial freeze configuration dictionary `freeze_config` covers the input embedding layer, input projection layer, Transformer backbone layer set, multi-task subnet, and output header (yield, NUE, WUE, carbon fixation, N2O, NH3, NO3-, GHG). The initial freeze ratio is 0 (all trainable).

[0203] The main layer is partially frozen proportionally.

[0204] When the total number of backbone layers is L and the freezing ratio is r∈[0,1], the number of frozen layers is... Pick:

[0205] ;

[0206] Before freezing Layer parameters do not participate in gradient updates; the remaining layers remain trainable along with other modules that are not completely frozen.

[0207] Parameter grouping and differential learning rate;

[0208] The trainable parameters are divided into four logical parameter groups (the example division satisfies the separation principle of "shared / backbone / task / output"):

[0209] The backbone shares the following parameter set: input embedding layer + input projection layer + all trainable Transformer backbone layers;

[0210] Task encoder group (the task encoder group mentioned in S35);

[0211] Cross-task attention group (Cross-taskTransformer);

[0212] Dedicated prediction subnet and output header group (the dedicated prediction subnet mentioned in S36).

[0213] Configure an independent learning rate for each group; if not explicitly specified, it will fall back to the base learning rate. The fine-tuning phase allows setting a multiplier (e.g., 0.1x, 2x) for specific task groups to achieve reinforcement / refinement learning.

[0214] The optimizer uses Adam and passes in the parameter set and learning rate to achieve joint optimization.

[0215] S382, based on the S381 configuration, performs training according to a preset sequence of multiple stages.

[0216] Each stage loads: training / validation set path, task weight vector, freeze configuration dictionary, batch size, dropout, whether to enable knowledge-guided constraint loss and its target. It applies `freeze_config` to freeze module parameters, initializes Adam by parameter grouping and learning rate, initializes the gradient scaler for mixed-precision training in a GPU environment, and initializes the optimal validation loss record and early stopping counter.

[0217] S383. Each stage is executed in a loop. For each batch, forward inference is performed, and the multi-task master loss (including mask) is calculated. If enabled, the knowledge-guided constraint loss is superimposed. The precision backpropagation and parameter update are mixed. The validation loss is periodically evaluated and calculated on the validation set. The weight with the lowest validation loss is saved. If the validation loss does not improve in consecutive patience rounds, the training of this stage is terminated early, and the optimal model of this stage is retained.

[0218] S384, Output of Final Model and Supporting Materials

[0219] The training iterations were set to 2000 rounds to ensure sufficient and stable model convergence. After the multi-stage training process was completed, a complete knowledge-data hybrid-driven multi-task Transformer model and its accompanying files were output.

[0220] S385. The model's predictive accuracy was evaluated using the coefficient of determination (R²) and root mean square error (RMSE). This was to further verify the model's ability to fit the agronomic laws of other indicators.

[0221] S4. Collect actual meteorological and soil data and input them into the farmland fertilizer and water management model to obtain farmland fertilizer and water management strategies.

[0222] A multi-objective optimization evolutionary algorithm is used to clarify a multi-objective synergistic crop solution that integrates high yield, high efficiency, and emission reduction.

[0223] The adjustment objectives are to minimize greenhouse gas emissions, minimize farmland nutrient loss, minimize water resource utilization, maximize crop yield, and maximize farmers' economic benefits. The optimal combination of management measures is simulated within each grid using the following formula:

[0224]

[0225] Among them, Z 1,2,3 Z represents greenhouse gas emissions from farmland, nutrient loss from farmland, and water resource utilization. 4,5 It indicates crop yield and farmers' economic benefits.

[0226] A knowledge-guided farmland fertilizer and water management device includes:

[0227] The data acquisition module is configured to collect historical data on crops, as well as actual meteorological and soil data.

[0228] The data preprocessing module is configured to adjust the parameters of the crop process model to generate a synthetic model. The historical data collected by the acquisition module is input into the synthetic model to generate simulated data. The synthetic model is verified through the simulated data, and synthetic data is generated based on the verified synthetic model.

[0229] The prediction module is configured to train a multi-task model using historical and synthetic data to form a farmland fertilizer and water management model. Multiple dedicated prediction subnets are built in the multi-task model, and the causal relationship between tasks is strengthened by coupling yield features.

[0230] The actual meteorological and soil data collected by the acquisition module are input into the farmland fertilizer and water management model to obtain farmland fertilizer and water management strategies.

[0231] Beneficial effects:

[0232] 1. In the model building stage, a multi-process mechanism model collaborative simulation mechanism is introduced. Through a unified scenario parameter interface, crop growth model, nitrogen cycle model and water management model are simulated together under the same management scenario. The system generates large-scale pre-training data covering different soil types, climate conditions and management method combinations. Through unified scenario design and index system, multiple models participate in the simulation under the same input conditions, thereby explicitly characterizing the coupling relationship between different processes at the data level. The agricultural big model learns the inherent coupling law between multiple processes in farmland during the training stage, which significantly enhances the model's ability to express complex management scenarios.

[0233] 2. During the training and inference process of the model, the mechanistic knowledge of the farmland system is explicitly introduced as a constraint. Agronomical knowledge such as the conservation of matter, the directionality of process response, and the reasonable range of management measures are embedded into the model structure design and loss function construction. This ensures that the model not only minimizes the prediction error during the optimization process, but also meets the basic mechanistic laws of the farmland system. Guided by knowledge constraints during the training process, the model's prediction results are always within the reasonable range of agronomy. This significantly improves the stability, reliability, and interpretability of the model while ensuring prediction accuracy.

[0234] 3. After completing the pre-training based on multi-process simulation data, the agricultural large model is fine-tuned in stages through transfer learning and incremental learning, so that the model gradually adapts to the operating characteristics of the real farmland system. Unlike the model construction method that only relies on simulation data or measured data, the model is driven by the joint simulation data and historical data, which enables the model to have a wide range of scenario adaptability and to truly reflect the operating characteristics of the farmland system, significantly improving the generalization performance of the model in cross-regional and cross-crop applications.

[0235] Example 2:

[0236] A server includes at least one processor and a memory communicatively connected to the processor, the memory storing instructions executable by the at least one processor to cause the at least one processor to perform the method as described in Embodiment 1.

[0237] Example 3:

[0238] A computer-readable storage medium storing a computer program that, when executed by a processor, implements the method described in Embodiment 1.

[0239] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention, and they should all be covered within the scope of the claims and specification of the present invention.

[0240] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A method for farmland fertilizer and water management based on a knowledge-guided training large model, characterized in that, Includes the following steps: S1. Collect historical data on crops; S2. Adjust the parameters of the crop process model to generate a synthetic model. Input historical data into the synthetic model to generate simulation data. Validate the synthetic model through simulation data. Generate synthetic data based on the validated synthetic model. S3. Use historical data and synthetic data to train a multi-task model to form a farmland fertilizer and water management model. Multiple dedicated prediction subnets are constructed in the multi-task model, and the causal relationship between tasks is strengthened by coupling yield characteristics. S4. Collect actual meteorological and soil data and input them into the farmland fertilizer and water management model to obtain farmland fertilizer and water management strategies.

2. The method for farmland fertilizer and water management based on a knowledge-guided training large model according to claim 1, characterized in that: The historical data of crops in S1 includes meteorological data, soil data, and management measures data; Meteorological data includes: daily maximum and minimum temperatures, rainfall, wind speed, relative humidity, and sunshine duration. Management data includes: sowing date, harvest date, seeding rate, yield, and management data; Soil data include: soil organic carbon, total nitrogen, pH, bulk density, and clay content.

3. The method for farmland fertilizer and water management based on a knowledge-guided training large model according to claim 1, characterized in that, S2 includes the following steps: S21. Divide the historical data of crops into a training set and a validation set; S22. Adjust the parameters in the crop process model through trial and error, and input the training set data into the crop process model to obtain simulation data and key variety parameters; Process models include crop growth models, soil carbon and nitrogen cycle models, and water management models; S23. Use the validation set to validate the simulated data through linear regression or index evaluation method to obtain the key variety parameters that have passed the validation. Adjust the crop process model according to the key variety parameters to form a synthetic model. S24. Collect meteorological data and soil data of crops at different locations and input them into the synthetic model to form synthetic data.

4. The method for farmland fertilizer and water management based on a knowledge-guided training large model according to claim 1, characterized in that, S3 includes the following steps: S31. The model input data is divided into categorical factor sets and continuous factor sets. The model output task is defined, and the combination is described as a triple structure. Missing value labeling and mask generation are performed. S32. Fit the preprocessed data transformer and save it persistently. The categorical and continuous features defined in S33 and S31 are transformed into 256-dimensional input features; S34. Based on a multi-head self-attention stacked architecture, a 6-layer Transformer encoder backbone is constructed to extract deep, globally interactive shared feature representations from 256-dimensional input features. S35. The shared features obtained in S34 are divided into 8 task feature representations. Cross-task attention is used to perform inter-task information interaction, and weighted fusion is used to obtain the final features. S36. Construct 8 dedicated prediction subnets to strengthen the causal relationship between tasks through output feature coupling and add physical constraints; S37. Construct a loss function to transform agronomic physical laws and empirical relationships into computable constraint loss terms; S38. The trainable range of the backbone layer is controlled by freezing the ratio, and the differentiated learning rate is achieved by grouping parameters. The farmland fertilizer and water management model is formed through iterative training.

5. A method for farmland fertilizer and water management based on a knowledge-guided training large model according to claim 4, characterized in that: S35 includes the following steps: S351. Construct 8 task encoders for each task. The 8 encoders have the same structure but independent parameters, and map the shared representation to the 8 task feature representations. S352. Perform cross-task attention interaction, stack the 8 task feature representations along the sequence dimension to form a feature sequence, and apply a multi-head self-attention mechanism to the feature sequence; S353. Perform a linear weighted combination of the task feature representation after cross-task attention enhancement and the shared features.

6. The method for farmland fertilizer and water management based on a knowledge-guided training large model according to claim 5, characterized in that, S36 includes the following steps: S361. Construct a dedicated prediction subnet for each prediction task; The forecasting tasks include forecasting crop yield, nitrogen fertilizer use efficiency, water use efficiency, nitrous oxide emissions, ammonia volatilization, nitrate leaching, soil carbon sequestration, and greenhouse gas emissions. S362. Define the output of a certain intermediate layer in the crop yield prediction subnet as the yield feature, the final task feature of the crop yield prediction subnet as the input, and the final task features of the other special prediction subnets and the yield feature are concatenated along the feature dimension to form a concatenated vector as the input of the corresponding subnet. S363, the crop yield prediction subnet is mapped to a 1-dimensional real value through a linear output layer; The remaining 7 prediction subnets first obtain initial predictions through a linear output layer, and then apply the Softplus activation function for nonlinear transformation.

7. The method for farmland fertilizer and water management based on a knowledge-guided training large model according to claim 1, characterized in that, S37 includes the following steps: S371. The Huber loss function is used as the basic loss function for each prediction task. S372. Transform agronomic physical laws and empirical relationships into calculable constraint loss terms; The following constraint losses are included: physical relationship fitting loss, curve slope penalty loss, and prediction internal consistency penalty loss; S373. The Huber loss function is weighted and fused with the physical relationship fitting loss, curve slope penalty loss and prediction internal consistency penalty loss.

8. A knowledge-guided farmland fertilizer and water management device, characterized in that, include: The data acquisition module is configured to collect historical data on crops, as well as actual meteorological and soil data. The data preprocessing module is configured to adjust the parameters of the crop process model to generate a synthetic model. The historical data collected by the acquisition module is input into the synthetic model to generate simulated data. The synthetic model is verified through the simulated data, and synthetic data is generated based on the verified synthetic model. The prediction module is configured to train a multi-task model using historical and synthetic data to form a farmland fertilizer and water management model. Multiple dedicated prediction subnets are built in the multi-task model, and the causal relationship between tasks is strengthened by coupling yield features. The actual meteorological and soil data collected by the acquisition module are input into the farmland fertilizer and water management model to obtain farmland fertilizer and water management strategies.

9. A server, characterized in that: The method includes at least one processor and a memory communicatively connected to the processor, the memory storing instructions executable by the at least one processor, the instructions being executed by the processor to cause the at least one processor to perform the method as described in any one of claims 1-7.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, it implements the method described in any one of claims 1-7.