Multi-source sensing data driven soil improvement effect intelligent evaluation method and system
By constructing a high-density sensor network and neural network model in the source region, a virtual soil data field is generated. A lightweight evaluation model is trained and additional sensors are deployed, solving the problem of evaluating the soil improvement effect in data-scarce areas and achieving efficient and accurate dynamic monitoring.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- INSTITUTE OF ENVIRONMENT AND SUSTAINABLE DEVELOPMENT IN AGRICULTURE CAAS
- Filing Date
- 2026-03-19
- Publication Date
- 2026-06-12
AI Technical Summary
Existing technologies struggle to conduct high-precision, dynamic assessments of soil improvement effects in target areas with scarce data. Traditional assessment methods rely on periodic sampling and cannot achieve real-time, dynamic monitoring, making cross-regional model migration difficult.
By deploying a high-density sensor network in the source region, a neural network model driven by multi-source sensor data is constructed to generate a virtual soil data field and train a lightweight evaluation model. Based on the uncertainty of prediction, additional physical sensors are deployed to update the model to evaluate the soil improvement effect.
It enables high-precision, dynamic assessment of soil improvement effects in data-scarce regions, improving assessment efficiency, accuracy, and reliability, and optimizing the allocation of monitoring resources.
Smart Images

Figure REF-OBJ-1773775493383-000002 
Figure REF-OBJ-1773775493383-000003
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing technology, specifically to a method and system for intelligent evaluation of soil improvement effects driven by multi-source sensor data. Background Technology
[0002] Soil improvement involves the purposeful regulation and optimization of soil structure, fertility, and biological activity through physical, chemical, or biological measures to enhance land productivity and strengthen ecosystem services. However, assessing the effectiveness of soil improvement is a complex spatiotemporal dynamic problem that faces multiple challenges. First, soil exhibits high spatial heterogeneity and temporal variability, with its physicochemical properties potentially changing drastically at the field scale. Furthermore, the effectiveness of soil improvement measures interacts in complex ways with local meteorological conditions, vegetation growth dynamics, and existing soil baseline values. Traditional assessments rely on periodic, discrete field sampling and laboratory analysis, resulting in sparse data, poor timeliness, and difficulty in capturing the continuous spatiotemporal evolution of soil properties. Moreover, they cannot achieve real-time, dynamic monitoring of the improvement effects over large areas. Second, target areas requiring soil improvement often have sparsely deployed physical sensors or a limited number of sampling points, making it extremely difficult to directly construct high-precision, reliable assessment models in these areas. Multi-source sensor data-driven assessment methods can acquire multi-dimensional data such as soil moisture, temperature, electrical conductivity, and near-surface meteorological and vegetation spectra with unprecedented spatiotemporal resolution. A crucial aspect is how to effectively transfer this data to new target areas with limited data—that is, solving the problem of cross-regional and cross-condition model transfer and adaptation.
[0003] Therefore, current technologies face the challenge of conducting high-precision, dynamic assessments of soil improvement effects in target areas with scarce data. Summary of the Invention
[0004] This application provides a method and system for intelligent evaluation of soil improvement effects driven by multi-source sensor data. It solves the technical problem in the prior art that it is difficult to conduct high-precision and dynamic evaluation of soil improvement effects in target areas with scarce data. It achieves the technical effect of realizing intelligent optimization of monitoring resources and improving the efficiency, accuracy and reliability of soil improvement effect evaluation.
[0005] This application provides a multi-source sensor data-driven intelligent evaluation method for soil improvement effects. The method includes: deploying a high-density sensor network in a source region to acquire soil physicochemical property data, vegetation growth data, meteorological data, and applied improvement measures; constructing a source region training set to train a first evaluation model; acquiring meteorological data of a target region, initial monitoring data from sparsely deployed sensors, and a preset improvement scheme; generating a virtual soil data field of the target region based on the first evaluation model; training a lightweight second evaluation model based on the virtual soil data field; and determining N recommended locations for additional physical sensors in the target region based on the prediction uncertainty of the second evaluation model; updating the second evaluation model based on real monitoring data acquired at the recommended locations; and generating an evaluation result for the soil improvement effect in the target region.
[0006] In a possible implementation, the first evaluation model is trained using a source domain training set, including: constructing a neural network architecture, wherein the input layer of the neural network architecture receives initial soil physicochemical properties, meteorological time-series data, and applied improvement measures, and the output layer contains two branches, the first branch outputs predicted values of vegetation growth indicators, and the second branch outputs inverted values of soil mediating parameters; preprocessing the source domain training set to construct sample labels, each sample label containing the true values of vegetation growth measured in the same period and the true values of soil mediating parameter analysis; and iteratively training the neural network architecture using the source domain training set based on the sample labels until the composite loss function converges to obtain the first evaluation model.
[0007] In a possible implementation, the intelligent evaluation method for soil improvement effect driven by multi-source sensor data includes: the composite loss function is a weighted sum of the loss from the vegetation growth prediction task and the loss from the soil mediator parameter inversion task.
[0008] In a possible implementation, the intelligent evaluation method for soil improvement effect driven by multi-source sensor data includes: soil mediating parameters including soil moisture conductivity, nutrient availability content and microbial activity index.
[0009] In a possible implementation, iteratively training the neural network architecture using the source domain training set further includes: constructing positive sample pairs and negative sample pairs in the source domain training set, wherein the positive sample pairs are data samples collected under the same improvement measures but at different meteorological conditions, and the negative sample pairs are data samples collected under similar meteorological conditions but at different improvement measures; after the feature extraction layer of the neural network architecture, introducing a contrastive learning head to calculate the feature similarity loss of the positive sample pairs and the feature discrimination loss of the negative sample pairs to obtain the contrastive learning loss; and combining the contrastive learning loss with the composite loss function to perform training of the neural network architecture.
[0010] In a possible implementation, meteorological data of the target area, initial monitoring data from sparsely deployed sensors, and a preset improvement scheme are acquired. A virtual soil data field of the target area is generated based on the first evaluation model, including: defining an initial spatial grid according to the deployment network of sparsely deployed sensors; forming an initial input vector for each grid cell according to the meteorological data, the initial monitoring data, and the preset improvement scheme; inputting the initial input vector of each grid cell into the first evaluation model to infer the predicted value corresponding to each grid cell, including the predicted value of virtual soil mediator parameters and the predicted value of virtual vegetation growth response; and using the Kriging spatial interpolation algorithm, with the predicted value corresponding to each grid cell as a known point, combined with the semivariance function model of soil properties, interpolating to generate virtual soil mediator parameters and virtual vegetation growth response parameters that cover the target area and are spatially continuous, thus constituting the virtual soil data field.
[0011] In one possible implementation, the Kriging spatial interpolation algorithm is used. With the predicted value corresponding to each grid cell as a known point, and combined with a semivariance function model of soil properties, virtual soil mediator parameters and virtual vegetation growth response parameters covering the target area and spatially continuously distributed are generated through interpolation. This includes: for each type of soil mediator parameter and vegetation growth response variable to be interpolated, a semivariance function model is fitted by collecting spatial variation sample data; for any unknown spatial point to be interpolated within the target area, the spatial correlation weight between the unknown spatial point and all known points is calculated based on the semivariance function model; using ordinary Kriging interpolation, the spatial correlation weight is linearly combined with the observed values of the corresponding known points to calculate the optimal unbiased estimate of the unknown spatial point, thus completing the interpolation.
[0012] In a possible implementation, a lightweight second evaluation model is trained based on the virtual soil data field, including: constructing an initial architecture of the second evaluation model using a lightweight neural network; employing knowledge distillation technology, using the first evaluation model as the teacher model, the initial architecture as the student model, and the virtual soil data field as the training samples, to train the second evaluation model by mimicking the first evaluation model, thereby generating the second evaluation model; wherein, the distillation loss function of the training process consists of the KL divergence loss between the outputs of the first evaluation model and the second evaluation model, and the mean squared error loss between intermediate layer features.
[0013] In a possible implementation, based on the prediction uncertainty of the second evaluation model, N recommended locations for additional physical sensors to be deployed in the target area are determined, including: using the virtual soil data field, for each spatial location, enabling the built-in Dropout layer of the second evaluation model to perform multiple forward propagations during the prediction phase to obtain a set of predicted values for each spatial location; calculating the statistical variance of the set of predicted values for each spatial location as a quantification index of prediction uncertainty; and selecting N spatial locations whose quantification index of prediction uncertainty is greater than a preset threshold as the N recommended locations.
[0014] This application also provides a multi-source sensor data-driven intelligent evaluation system for soil improvement effects. The system includes: a first evaluation model training module, used to deploy a high-density sensor network in the source region to acquire soil physicochemical property data, vegetation growth data, meteorological data, and applied improvement measures, and construct a source region training set to train the first evaluation model; a virtual soil data field generation module, used to acquire meteorological data of the target region, initial monitoring data from sparsely deployed sensors, and a preset improvement scheme, and generate a virtual soil data field of the target region based on the first evaluation model; a second evaluation model training module, used to train a lightweight second evaluation model based on the virtual soil data field, and determine N recommended locations for additional physical sensors to be deployed in the target region based on the prediction uncertainty of the second evaluation model; and a soil improvement effect evaluation module, used to update the second evaluation model based on the real monitoring data acquired at the recommended locations, and generate an evaluation result of the soil improvement effect in the target region.
[0015] This application proposes a multi-source sensor data-driven intelligent evaluation method and system for soil improvement effects. A high-density sensor network is deployed in the source region to acquire soil physicochemical property data, vegetation growth data, meteorological data, and applied improvement measures. A source region training set is then constructed to train a first evaluation model. Based on the first evaluation model, a virtual soil data field is generated. A lightweight second evaluation model is trained, and N recommended locations for additional physical sensors in the target region are determined. Real monitoring data is acquired to update the second evaluation model, generating an evaluation result for the soil improvement effect in the target region. This solves the technical problem of difficulty in conducting high-precision, dynamic soil improvement effect evaluation in data-scarce target areas, achieving intelligent optimization of monitoring resources and improving the efficiency, accuracy, and reliability of soil improvement effect evaluation. Attached Figure Description
[0016] To more clearly illustrate the technical solutions of the embodiments of this disclosure, the accompanying drawings of the embodiments of this disclosure will be briefly described below. Flowcharts are used in this application to illustrate the operations performed by the system according to the embodiments of this application. It should be understood that the preceding or following operations are not necessarily performed precisely in sequence. Instead, various steps can be processed in reverse order or simultaneously as needed. Furthermore, other operations can be added to these processes, or one or more steps can be removed from these processes.
[0017] Figure 1 This is a schematic diagram of the process for a multi-source sensor data-driven intelligent evaluation method for soil improvement effects provided in an embodiment of this application.
[0018] Figure 2 This is a schematic diagram of the structure of a multi-source sensor data-driven intelligent evaluation system for soil improvement effects provided in an embodiment of this application.
[0019] Figure labeling: First evaluation model training module 10, virtual soil data field generation module 20, second evaluation model training module 30, soil improvement effect evaluation module 40. Detailed Implementation
[0020] To further illustrate the technical means and effects adopted by the present invention in order to achieve the intended purpose, the following detailed description is provided in conjunction with the accompanying drawings and preferred embodiments, based on the specific implementation methods, structures, features and effects of the present invention.
[0021] This application provides an intelligent evaluation method for soil amendment effects driven by multi-source sensor data, such as... Figure 1 As shown, the method includes: Step S100: Deploy a high-density sensor network in the source region to acquire soil physicochemical property data, vegetation growth data, meteorological data, and applied improvement measures, and construct a source region training set to train the first evaluation model.
[0022] Preferably, the source region refers to a specific area with complete data conditions and where soil improvement has been implemented. A high-density sensor network is formed by deploying sensors at multiple sampling points within the source region to collect real-time data on soil physicochemical properties, including but not limited to soil moisture, temperature, pH, electrical conductivity, organic matter content, nitrogen, phosphorus, and potassium nutrient content, and soil texture; vegetation growth data, including vegetation index, leaf area index, biomass, and plant height; meteorological data, including precipitation, air temperature, humidity, wind speed, and solar radiation; and structured operational information on applied improvement measures, including the type of amendment, application rate, application time, tillage method, and irrigation system. The acquired data is then aligned spatiotemporally to form a training sample set with input-output correspondence, i.e., the source region training set. Based on neural network training, a model can be established that maps the complex relationship between improvement measures, environmental conditions, and soil-vegetation response—the first evaluation model. This model can predict changes in key soil mediator parameters and vegetation growth responses based on the input initial soil state, meteorological conditions, and improvement measures.
[0023] Furthermore, step S100 also includes constructing a neural network architecture, wherein the input layer of the neural network architecture receives initial soil physicochemical properties, meteorological time-series data, and applied improvement measures, and the output layer contains two branches, the first branch outputs predicted values of vegetation growth indicators, and the second branch outputs inverted values of soil mediating parameters; preprocessing the source domain training set to construct sample labels, each sample label containing the ground truth values of vegetation growth measured in the same period and the ground truth values of soil mediating parameters; based on the sample labels, iteratively training the neural network architecture using the source domain training set until the composite loss function converges to obtain the first evaluation model.
[0024] Preferably, a neural network architecture is constructed, including an input layer, a hidden layer, and an output layer. The input layer receives three types of structured data: initial soil physicochemical properties, such as humidity, pH, and organic matter content; meteorological time-series data sequences, such as daily temperature, precipitation, and radiation over a past period (e.g., the crop growing season); and amendment measure encoding vectors, such as amendment type (one-hot encoding), application rate, and application time. The hidden layer consists of fully connected layers, convolutional layers, or long short-term memory networks, used to extract and fuse high-level features from the input data. The output layer has a two-branch structure. The first branch is used for vegetation growth prediction, outputting predicted values of vegetation growth indicators, such as leaf area index, biomass, or vegetation index. The second branch is used for soil parameter inversion, outputting inverted values of soil intermediate parameters, such as soil water conductivity, available nitrogen / phosphorus / potassium content, and microbial respiration rate.
[0025] Preferably, the source domain training set is preprocessed, including time synchronization (ensuring that input and output correspond to the same time period) and spatial matching (ensuring that the locations are consistent) of the original sensor data and test data, and handling missing values and outliers. Then, the input data is normalized and standardized, and the time-series meteorological data is organized into feature tensors with fixed time windows. Then, supervised learning samples are constructed, where each sample contains input features and labels. The input features refer to the initial soil properties of a certain spatiotemporal location before or at the beginning of the implementation of the improvement measures, the meteorological time-series data during the implementation period, and the specific description of the applied improvement measures. The sample label refers to the true value that corresponds completely to the input feature in spatiotemporal space and is obtained by actual measurement through sensors or laboratory tests, including the true value of vegetation growth, such as the NDVI measurement value or biomass sampling data of the same period, and the true value of soil mediating parameters, such as soil moisture conductivity, available nutrient content, and microbial activity index obtained by instrument measurement or laboratory chemical analysis of the same period. Then, the input features of the training samples are fed into the neural network. The predicted values of the two branches are calculated, and the weighted sum of the task losses of the two branches is used as the composite loss function to quantify the gap between the predicted values and the true labels. For example, mean squared error (MSE) is used to calculate the error between the predicted vegetation indicators and the measured values, and mean squared error or mean absolute error is used to calculate the error between the predicted soil parameters and the test values. Based on the composite loss value, all trainable parameters in the neural network, including weights and biases, are updated through backpropagation using the Adam gradient descent algorithm to minimize the overall prediction error. The training is repeated multiple times on the entire training set until the value of the composite loss function stabilizes near the minimum point, indicating that the model has fully learned the mapping relationship from "measures and environmental conditions" to "soil and vegetation response". At this point, the training is complete, and a usable first evaluation model is obtained, which can simultaneously and collaboratively predict vegetation performance and soil intrinsic state.
[0026] Furthermore, step S100 also includes the composite loss function comprising a weighted sum of the loss from the vegetation growth prediction task and the loss from the soil mediator parameter inversion task.
[0027] Preferably, the composite loss function is composed of a weighted sum of the loss from the vegetation growth prediction task and the loss from the soil mediator parameter inversion task. The loss from the vegetation growth prediction task is the difference between the predicted and actual measured values of the vegetation growth index output by the model in the first branch, such as mean squared error loss. The loss from the soil mediator parameter inversion task is the difference between the predicted and actual measured values of the soil mediator parameters output by the model in the second branch, such as mean absolute error loss. The weights are positive real numbers used to balance the relative importance and magnitude difference between the two tasks in the total loss. Furthermore, the weights can be adjusted to emphasize learning accuracy according to actual application needs.
[0028] Furthermore, step S100 also includes soil mediating parameters including soil moisture conductivity, nutrient availability, and microbial activity index.
[0029] Preferably, soil water conductivity refers to the flow rate of water through a unit cross-sectional area of soil per unit time under a unit water potential gradient. It characterizes the soil's water movement capacity and directly determines the infiltration, redistribution, and drainage rates of water in the soil. Soil improvement measures such as applying organic fertilizers and deep loosening affect water conductivity by changing soil structure, porosity, and aggregate stability. Changes in water conductivity further regulate root zone water availability, aeration, and solute migration, ultimately leading to a promoting or inhibiting effect on vegetation growth. Nutrient availability refers to the concentration of soil nutrients that plants can directly absorb and utilize in the short term. For nitrogen, this mainly refers to ammonium nitrogen and nitrate nitrogen; for phosphorus, it mainly refers to water-soluble phosphorus and easily desorbed adsorbed phosphorus; and for potassium, it mainly refers to water-soluble potassium and exchangeable potassium. It characterizes the soil nutrient supply intensity and reflects the soil's... The immediate nutrient supply capacity of soil is directly altered by the application of fertilizers, biochar, or microbial agents. These changes primarily manifest in the increase or decrease of available nutrient content, which is the direct source of plant root absorption. The microbial activity index, a comprehensive indicator of the metabolic activity of soil microbial communities measured using biochemical methods, is commonly expressed as the soil's basic respiration rate or specific enzyme activity. It reflects the strength of soil biological function and the potential for organic matter transformation. Soil improvement measures such as adding organic materials and inoculating with mycorrhizal fungi directly affect the substrate, energy supply, and habitat of soil microorganisms, thus altering their activity and driving the decomposition of organic matter, nutrient mineralization, and pollutant degradation in the soil. This releases available nutrients and synthesizes growth-promoting hormones, ultimately influencing plant growth.
[0030] Furthermore, step S100 also includes constructing positive sample pairs and negative sample pairs in the source domain training set, wherein the positive sample pairs are data samples collected under the same improvement measures but at different meteorological conditions, and the negative sample pairs are data samples collected under similar meteorological conditions but at different improvement measures; after the feature extraction layer of the neural network architecture, a contrastive learning head is introduced to calculate the feature similarity loss of the positive sample pairs and the feature discrimination loss of the negative sample pairs to obtain the contrastive learning loss; the contrastive learning loss is combined with the composite loss function to perform training of the neural network architecture.
[0031] Preferably, two data samples are selected from the source domain training set, which must meet the same improvement measure, such as applying the same dose of biochar, but collected under different meteorological conditions, such as one in the rainy summer and the other in the dry autumn. A positive sample pair is constructed, and the model is forced to believe that the feature representations of the two samples should be similar. This guides the model to remove the interference caused by meteorological fluctuations when extracting features, and to focus on capturing the relatively stable soil-vegetation response pattern caused by the improvement measure itself. Alternatively, two data samples are selected from the source domain training set, which must meet similar meteorological conditions, such as being in the spring sowing period with similar temperature and humidity conditions, but using different improvement measures, such as one applying organic fertilizer and the other not applying any amendment. A negative sample pair is constructed, and the model is forced to believe that the feature representations of the two samples should be "different". This guides the model to be able to keenly distinguish the differences in the effects produced by different improvement measures when extracting features. Following the feature extraction layer of the neural network, an auxiliary network module, the contrastive learning head, is added. This head maps the high-dimensional features output by the feature extraction layer to a contrastive feature space more suitable for similarity measurement. For each pair of samples, the similarity of their feature vectors in the contrastive feature space is calculated using cosine similarity. Then, the feature discrimination loss of the negative sample pairs is calculated to obtain the contrastive learning loss, thereby narrowing the feature distance between positive sample pairs and widening the feature distance between negative sample pairs, i.e., increasing the difference between the similarity of positive sample pairs and the similarity of all negative sample pairs. The contrastive learning loss is combined with a composite loss function used for multi-task prediction to form the final training objective function. During training, the prediction accuracy and feature representation objectives are optimized simultaneously by minimizing the final training loss. This includes making the model's predicted values for vegetation and soil parameters as close as possible to the true values, while ensuring that the feature representations learned by the model satisfy the constraints of contrastive learning: features of the same measures but different weather conditions are similar, and features of different measures but the same weather conditions are different. The trained feature representations are more robust and generalizable than features learned solely through prediction tasks, which helps improve the model's performance when transferring to target areas with different weather patterns or different improvement measures.
[0032] Step S200: Obtain meteorological data of the target area, initial monitoring data from sparsely deployed sensors, and preset improvement schemes; generate a virtual soil data field of the target area based on the first evaluation model.
[0033] Step S200 further includes defining an initial spatial grid based on the sparsely deployed sensor network; forming an initial input vector for each grid cell based on the meteorological data, the initial monitoring data, and the preset improvement scheme; inputting the initial input vector of each grid cell into the first evaluation model to infer the predicted value corresponding to each grid cell, including the predicted value of virtual soil mediator parameters and the predicted value of virtual vegetation growth response; and using the Kriging spatial interpolation algorithm, with the predicted value corresponding to each grid cell as a known point, and combining the semivariance function model of soil properties, interpolating to generate virtual soil mediator parameters and virtual vegetation growth response parameters that cover the target area and are spatially continuous, thus constituting the virtual soil data field.
[0034] Preferably, based on the scope of the target area and the location of the sparse sensor network, the entire target area is discretized into regular or irregular initial spatial grids. For each initial spatial grid, an input vector is constructed using meteorological data, initial monitoring data, and a preset improvement scheme, thus forming the initial input vector for each grid cell. The meteorological data refers to the meteorological time series data representing the location of the cell or the region. If there is a sparse sensor in the grid, the initial soil physicochemical properties measured by it are directly used as the initial monitoring data. If there is no sensor in the grid, the initial soil physicochemical properties are estimated from the data of neighboring sensors by simple interpolation or by assigning a regional average value. The preset improvement scheme is the improvement measures planned to be implemented in the cell, such as the type and amount of amendment, as a structured input.
[0035] Preferably, the initial input vector of each grid cell is input into the first evaluation model to infer the predicted value corresponding to each grid cell, including the predicted value of virtual soil mediator parameters, such as the predicted soil moisture conductivity of the cell after the implementation of the improvement scheme, and the predicted value of virtual vegetation growth response, such as the predicted NDVI value of the cell after the implementation of the improvement scheme. Thus, discrete virtual predicted values distributed at the center point of each grid cell are obtained. For each variable requiring interpolation, such as "moisture conductivity" and "NDVI", a semivariogram function model is fitted using existing spatial variability knowledge of the target area or data from the source area to quantify the variation of the variable with increasing distance in space. Then, for any location within the target area that requires higher resolution, Kriging interpolation is performed. Specifically, based on its distance from all known points, the spatial correlation weight between the unknown point and each known point is calculated using a semivariance function model. The weights satisfy the conditions of unbiasedness and optimality. The predicted values of all known points around the unknown point are linearly combined according to the calculated weights to obtain the optimal estimate of the unknown point. Interpolation calculations are repeated for countless points within the area, ultimately resulting in a spatially continuous and smoothly distributed virtual soil mediator parameter field and virtual vegetation growth response field covering the entire target area. Together, they constitute a complete virtual soil data field, thereby overcoming the bottleneck of extremely scarce real data in the initial stage of the target area.
[0036] Furthermore, step S200 also includes: for each type of soil mediating parameter and vegetation growth response variable to be interpolated, by collecting spatial variation sample data, fitting a semivariogram function model respectively; for any unknown spatial point to be interpolated within the target area, calculating the spatial correlation weight between the unknown spatial point and all known points according to the semivariogram function model; using ordinary kriging interpolation, linearly combining the spatial correlation weight with the observation value of the corresponding known point, calculating the optimal unbiased estimate of the unknown spatial point, and completing the interpolation.
[0037] Preferably, for each type of soil mediating parameter and vegetation growth response variable to be interpolated, such as soil moisture conductivity, nitrate nitrogen content, and vegetation index NDVI, a separate spatial structure model is established. The experimental semivariogram is calculated by collecting spatial variation sample data, that is, the semivariogram at different distances between known points is calculated. The semivariogram function model is determined by fitting a spherical model, exponential model, or Gaussian model to describe the mapping relationship in which the variable loses correlation with the increase of spatial distance. For any unknown spatial point to be interpolated in the target area, all known points within a certain range around it are found. The spatial correlation weight between the unknown spatial point and all known points is calculated according to the semivariogram function model. Specifically, the semivariogram between the unknown point and the known points is calculated, and then the semivariogram between each pair of known points is calculated. The Kriging equation system is constructed and solved to determine the weight assigned to the observation value of each known point. The constraints of the equation system are unbiasedness, the sum of all weights is , and optimality, where the variance of the estimation error is minimized. The solution result of the equation is the optimal weight, which represents the spatial correlation weight based on spatial correlation. Ordinary Kriging interpolation is used to linearly combine the spatial correlation weights with the predicted observations of the corresponding known points to calculate the Kriging estimate of the unknown spatial points, which is the optimal unbiased estimate, thus ensuring that the variance of the estimation error is minimized and completing the interpolation. The calculation is repeated for all unknown points in the target area, and finally an estimate is calculated for each point, thereby generating an interpolation surface that covers the entire area, is spatially continuous and smooth, i.e., a virtual data field.
[0038] Step S300: Based on the virtual soil data field, a lightweight second evaluation model is trained, and based on the prediction uncertainty of the second evaluation model, N recommended locations for additional physical sensors to be deployed in the target area are determined.
[0039] Step S300 further includes: constructing an initial architecture for the second evaluation model using a lightweight neural network; employing knowledge distillation technology, using the first evaluation model as the teacher model, the initial architecture as the student model, and the virtual soil data field as the training sample, to train the second evaluation model by imitating the first evaluation model, thereby generating the second evaluation model; wherein, the distillation loss function of the training process consists of the KL divergence loss between the outputs of the first evaluation model and the second evaluation model, and the mean squared error loss between intermediate layer features.
[0040] Preferably, the initial architecture of the second evaluation model is constructed using a lightweight neural network. This architecture has fewer network layers, fewer neurons per layer, or uses more efficient computational modules, such as depthwise separable convolutions. The goal is to minimize model parameters and reduce computational complexity, facilitating rapid deployment and inference on resource-constrained devices (such as edge computing nodes) in the target region. The first evaluation model serves as the teacher model, the initial architecture as the student model, and a virtual soil data field generated in the target region is used as training samples. Knowledge distillation techniques are employed to train the second evaluation model in imitation of the first evaluation model, mimicking the behavior of the teacher model under the same input. Specifically, training samples from the virtual data field are simultaneously input into both the teacher and student models. The teacher model outputs its predictions, such as a probability distribution processed by Softmax or a direct regression prediction, containing complex knowledge learned from the source region, such as inter-category relationships and data smoothness. The student model also outputs its predictions. A distillation loss function is then calculated to measure the difference between the student model's output and the teacher model's output. The distillation loss function in the training process consists of the KL divergence loss between the outputs of the first evaluation model and the mean squared error loss between intermediate layer features. It aims to guide the student model to imitate the teacher model from different levels. The KL divergence measures the difference between the probability distribution of the student model's output and that of the teacher model's output. The KL divergence loss between outputs represents knowledge imitation of the output layer, forcing the student model's final prediction results to approximate the teacher model in terms of overall trend and uncertainty, inheriting the teacher model's judgment "style" and confidence. The mean squared error loss between intermediate layer features represents knowledge imitation of the intermediate feature layer, forcing the student model's internal feature representations to converge with the teacher model, guiding the student model to learn how the teacher model extracts and combines features. The distillation loss function is minimized through gradient descent optimization, updating all parameters of the student model. Training is complete when the loss converges, resulting in a lightweight second evaluation model with performance close to the complex teacher model but with very low size and computational requirements. This enables efficient and lightweight transfer from knowledge-rich regions to data-scarce regions.
[0041] Furthermore, step S300 also includes using the virtual soil data field, for each spatial location, enabling the built-in Dropout layer of the second evaluation model to perform multiple forward propagations during the prediction phase to obtain a set of predicted values for each spatial location; calculating the statistical variance of the set of predicted values for each spatial location as a quantitative index of prediction uncertainty; and selecting N spatial locations whose quantitative index of prediction uncertainty is greater than a preset threshold as the N recommended locations.
[0042] Preferably, the model prediction uncertainty quantification based on the Bayesian approximation idea intelligently guides the decision to deploy additional physical sensors. Specifically, for each spatial location in the virtual soil data field, i.e., each grid point, data representing the environment and measures at that location are input into a pre-trained second evaluation model. Multiple random predictions are performed using Monte Carlo Dropout, including enabling the Dropout layer used by the model during the training phase. That is, each time the model performs forward propagation, it randomly generates a slightly different network structure and repeats the forward propagation multiple times on the same input data. Due to the randomness of Dropout, a slightly different prediction result is obtained each time, ultimately obtaining a set of predicted values for each spatial location. Then, statistical analysis is performed on the multiple predicted values obtained for each location, and the statistical variance of the set of predicted values is calculated as a quantification index of prediction uncertainty. This index measures the degree of dispersion among multiple predicted values. If the model's prediction for a certain location is very certain, and the output is highly consistent regardless of how the network structure changes randomly, the variance is small. Conversely, if the model's prediction for that location is unstable and the results are scattered, the variance is large. High variance indicates that the model lacks knowledge at that location and the prediction reliability is low. After calculating the uncertainty quantification index of all spatial locations within the entire target area, a spatial distribution map of prediction uncertainty is generated. A preset threshold is set, and spatial locations with uncertainty indices greater than the preset threshold are selected, representing the areas where the model prediction is the least reliable and the information gap is the largest. Then, from multiple high uncertainty locations, the top N locations with the highest uncertainty are further selected as recommended points, where N is a positive integer. This intelligently determines the best locations for additional physical sensors to improve the performance of the entire system in the most efficient way.
[0043] Step S400: Based on the real monitoring data obtained at the recommended locations, the second evaluation model is updated to generate an evaluation result of the soil improvement effect in the target area.
[0044] Preferably, physical sensors are additionally deployed at the recommended N locations. After one monitoring cycle, real monitoring data for these locations is acquired, including actual measurements of soil mediator parameters, such as water conductivity and nutrient content obtained from actual laboratory tests, and actual observed values of vegetation growth indicators, such as NDVI and biomass obtained through field sampling or hyperspectral imagery. Then, using the meteorological conditions, initial soil conditions, and improvement plans of the recommended locations as inputs, and the real monitoring data as outputs, the pre-trained second evaluation model is fine-tuned using new, high-quality input-output paired data. For example, some or all parameters of the second evaluation model are unlocked, and a small number of iterations are performed using new data with a low learning rate. Finally, the newly acquired real knowledge is injected into the second evaluation model to better adapt its predictive ability to the local realities of the target area, correcting any errors that may have occurred with the virtual data. To address systematic biases or uncertainties, an updated second assessment model is obtained and used for forward propagation inference across all spatial locations in the entire target area. This model outputs a final predicted value for each location, including the spatial distribution of key intermediate parameters of the improved soil, such as organic matter enhancement rate, salinity reduction rate, and hydraulic conductivity improvement; vegetation growth response results, such as the spatial distribution of expected crop yield, biomass increase, or improvement in vegetation health index; and potentially a comprehensive assessment index integrating multiple indicators, visually displaying the spatial pattern of improvement effectiveness. Furthermore, Monte Carlo Dropout can be used again to calculate the spatial distribution of uncertainties predicted by the updated model, serving as supplementary data for the reliability of the assessment results. Finally, an assessment result of the soil improvement effect in the target area is generated, thereby achieving intelligent optimization of monitoring resources and improving the efficiency, accuracy, and reliability of soil improvement effect assessment.
[0045] In the above text, refer to Figure 1 This paper describes in detail a method for intelligent evaluation of soil amendment effects driven by multi-source sensor data according to embodiments of the present invention. Next, reference will be made to... Figure 2 This invention describes a multi-source sensor data-driven intelligent evaluation system for soil amendment effects according to embodiments of the present invention.
[0046] The multi-source sensor data-driven intelligent evaluation system for soil improvement effects according to embodiments of the present invention addresses the technical problem in existing technologies that make it difficult to conduct high-precision, dynamic evaluation of soil improvement effects in data-scarce target areas. It achieves the technical effect of intelligently optimizing the allocation of monitoring resources and improving the efficiency, accuracy, and reliability of soil improvement effect evaluation. Figure 2 As shown, the intelligent evaluation system for soil improvement effect driven by multi-source sensor data includes: a first evaluation model training module 10, a virtual soil data field generation module 20, a second evaluation model training module 30, and a soil improvement effect evaluation module 40.
[0047] The first evaluation model training module 10 is used to deploy a high-density sensor network in the source area to acquire soil physicochemical property data, vegetation growth data, meteorological data, and applied improvement measures, and construct a source area training set to train the first evaluation model; the virtual soil data field generation module 20 is used to acquire meteorological data of the target area, initial monitoring data from sparsely deployed sensors, and preset improvement schemes, and generate a virtual soil data field of the target area based on the first evaluation model; the second evaluation model training module 30 is used to train a lightweight second evaluation model based on the virtual soil data field, and determine N recommended locations for additional physical sensors to be deployed in the target area based on the prediction uncertainty of the second evaluation model; the soil improvement effect evaluation module 40 is used to update the second evaluation model based on the real monitoring data acquired at the recommended locations, and generate an evaluation result of the soil improvement effect in the target area.
[0048] The specific configuration of the first evaluation model training module 10 will be described in detail below. The first evaluation model training module 10 further includes: constructing a neural network architecture, wherein the input layer of the neural network architecture receives initial soil physicochemical properties, meteorological time-series data, and applied improvement measures; the output layer contains two branches, the first branch outputs predicted values of vegetation growth indicators, and the second branch outputs inverted values of soil mediating parameters; preprocessing the source domain training set to construct sample labels, each sample label containing the true values of vegetation growth measured concurrently and the true values of soil mediating parameter assays; and iteratively training the neural network architecture using the source domain training set based on the sample labels until the composite loss function converges to obtain the first evaluation model.
[0049] The specific configuration of the first evaluation model training module 10 will be described in detail below. The first evaluation model training module 10 further includes: the composite loss function is a weighted sum of the loss from the vegetation growth prediction task and the loss from the soil mediator parameter inversion task.
[0050] The specific configuration of the first evaluation model training module 10 will be described in detail below. The first evaluation model training module 10 further includes soil mediating parameters, including soil moisture conductivity, nutrient availability, and microbial activity index.
[0051] The specific configuration of the first evaluation model training module 10 will be described in detail below. The first evaluation model training module 10 further includes: constructing positive sample pairs and negative sample pairs in the source domain training set, wherein the positive sample pairs are data samples collected under the same improvement measures but at different meteorological conditions, and the negative sample pairs are data samples collected under similar meteorological conditions but at different improvement measures; after the feature extraction layer of the neural network architecture, a contrastive learning head is introduced to calculate the feature similarity loss of the positive sample pairs and the feature discrimination loss of the negative sample pairs, thus obtaining the contrastive learning loss; the contrastive learning loss is combined with the composite loss function to perform training of the neural network architecture.
[0052] The specific configuration of the virtual soil data field generation module 20 will be described in detail below. The virtual soil data field generation module 20 further includes: defining an initial spatial grid based on a sparsely deployed sensor network; forming an initial input vector for each grid cell based on the meteorological data, the initial monitoring data, and the preset improvement scheme; inputting the initial input vector of each grid cell into the first evaluation model to infer the predicted value corresponding to each grid cell, including the predicted value of the virtual soil mediator parameter and the predicted value of the virtual vegetation growth response; and using a Kriging spatial interpolation algorithm, with the predicted value corresponding to each grid cell as a known point, combining it with a semivariance function model of soil properties to interpolate and generate virtual soil mediator parameters and virtual vegetation growth response parameters that cover the target area and are spatially continuous, thus constituting the virtual soil data field.
[0053] The specific configuration of the virtual soil data field generation module 20 will be described in detail below. The virtual soil data field generation module 20 further includes: for each type of soil mediating parameter and vegetation growth response variable to be interpolated, a semivariogram function model is fitted by collecting spatial variation sample data; for any unknown spatial point to be interpolated within the target area, the spatial correlation weight between the unknown spatial point and all known points is calculated according to the semivariogram function model; using ordinary kriging interpolation, the spatial correlation weight is linearly combined with the observed values of the corresponding known points to calculate the optimal unbiased estimate of the unknown spatial point, thus completing the interpolation.
[0054] The specific configuration of the second evaluation model training module 30 will be described in detail below. The second evaluation model training module 30 further includes: constructing an initial architecture for the second evaluation model using a lightweight neural network; employing knowledge distillation technology, using the first evaluation model as the teacher model, the initial architecture as the student model, and the virtual soil data field as the training samples, to train the second evaluation model by mimicking the first evaluation model, thereby generating the second evaluation model; wherein the distillation loss function in the training process consists of the KL divergence loss between the outputs of the first evaluation model and the second evaluation model, and the mean squared error loss between intermediate layer features.
[0055] The specific configuration of the second evaluation model training module 30 will be described in detail below. The second evaluation model training module 30 further includes: using the virtual soil data field, for each spatial location, enabling the built-in Dropout layer of the second evaluation model to perform multiple forward propagations during the prediction phase to obtain a set of predicted values for each spatial location; calculating the statistical variance of the set of predicted values for each spatial location as a quantification index of prediction uncertainty; and selecting N spatial locations whose quantification index of prediction uncertainty is greater than a preset threshold as the N recommended locations.
[0056] The intelligent soil improvement effect evaluation system driven by multi-source sensor data provided in the embodiments of the present invention can execute the intelligent soil improvement effect evaluation method driven by multi-source sensor data provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the method.
[0057] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention in any way. Although the present invention has been disclosed above with reference to preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make some modifications or alterations to the above-disclosed technical content to create equivalent embodiments without departing from the scope of the present invention. Any modifications, equivalent changes, and alterations made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the scope of the present invention.
Claims
1. A method for intelligent evaluation of soil amendment effects driven by multi-source sensor data, characterized in that, include: A high-density sensor network was deployed in the source region to acquire soil physicochemical property data, vegetation growth data, meteorological data, and applied improvement measures, and a source region training set was constructed to train the first evaluation model. Acquire meteorological data of the target area, initial monitoring data from sparsely deployed sensors, and preset improvement schemes, and generate a virtual soil data field of the target area based on the first evaluation model; Based on the virtual soil data field, a lightweight second evaluation model is trained, and based on the prediction uncertainty of the second evaluation model, N recommended locations for additional physical sensors to be deployed in the target area are determined. Based on the real monitoring data obtained at the recommended locations, the second evaluation model is updated to generate an evaluation result on the soil improvement effect in the target area.
2. The intelligent evaluation method for soil improvement effect driven by multi-source sensor data as described in claim 1, characterized in that, Construct the source domain training set to train the first evaluation model, including: A neural network architecture is constructed. The input layer of the neural network architecture receives the initial physicochemical properties of the soil, meteorological time series data, and the applied improvement measures. The output layer contains two branches: the first branch outputs the predicted value of vegetation growth index, and the second branch outputs the inverted value of soil mediating parameters. The source domain training set is preprocessed to construct sample labels. Each sample label includes the true values of vegetation growth measured at the same time and the true values of soil mediating parameters. Based on the sample labels, the neural network architecture is iteratively trained using the source domain training set until the composite loss function converges, thus obtaining the first evaluation model.
3. The intelligent evaluation method for soil improvement effect driven by multi-source sensor data as described in claim 2, characterized in that, The composite loss function is a weighted sum of the loss from the vegetation growth prediction task and the loss from the soil mediator parameter inversion task.
4. The intelligent evaluation method for soil improvement effect driven by multi-source sensor data as described in claim 2, characterized in that, Soil intermediate parameters include soil moisture conductivity, nutrient availability, and microbial activity index.
5. The intelligent evaluation method for soil improvement effect driven by multi-source sensor data as described in claim 2, characterized in that, Iterative training of the neural network architecture using the source domain training set further includes: Positive sample pairs and negative sample pairs are constructed in the source domain training set. Positive sample pairs are data samples collected under the same improvement measures but at different meteorological conditions, while negative sample pairs are data samples collected under similar meteorological conditions but at different improvement measures. After the feature extraction layer of the neural network architecture, a contrastive learning head is introduced to calculate the feature similarity loss of the positive sample pair and the feature discrimination loss of the negative sample pair, thus obtaining the contrastive learning loss. The contrastive learning loss is combined with the composite loss function to train the neural network architecture.
6. The intelligent evaluation method for soil improvement effect driven by multi-source sensor data as described in claim 1, characterized in that, Acquire meteorological data of the target area, initial monitoring data from sparsely deployed sensors, and a preset improvement plan; generate a virtual soil data field of the target area based on the first evaluation model, including: Based on the sparsely deployed sensor network, an initial spatial grid is defined, and an initial input vector for each grid cell is formed based on the meteorological data, the initial monitoring data, and the preset improvement scheme. The initial input vector of each grid cell is input into the first evaluation model, and the predicted value corresponding to each grid cell is obtained by reasoning, including the predicted value of virtual soil mediator parameter and the predicted value of virtual vegetation growth response. Using the Kriging space interpolation algorithm, with the predicted value corresponding to each grid cell as a known point, and combined with the semivariance function model of soil properties, virtual soil mediating parameters and virtual vegetation growth response parameters that cover the target area and are spatially continuous are interpolated to form the virtual soil data field.
7. The intelligent evaluation method for soil improvement effect driven by multi-source sensor data as described in claim 6, characterized in that, Using the Kriging spatial interpolation algorithm, with the predicted value corresponding to each grid cell as a known point, and combined with the semivariance function model of soil properties, virtual soil mediator parameters and virtual vegetation growth response parameters covering the target area and spatially continuously distributed are interpolated and generated, including: For each type of soil mediating parameter and vegetation growth response variable to be interpolated, a semivariogram function model is fitted by collecting spatial variation sample data. For any unknown spatial point to be interpolated within the target area, the spatial correlation weight between the unknown spatial point and all known points is calculated according to the semivariance function model. Ordinary Kriging interpolation is used to linearly combine the spatial correlation weights with the observations of the corresponding known points to calculate the optimal unbiased estimate of the unknown spatial points, thus completing the interpolation.
8. The intelligent evaluation method for soil improvement effect driven by multi-source sensor data as described in claim 1, characterized in that, Based on the virtual soil data field, a lightweight second evaluation model is trained, including: The initial architecture of the second evaluation model is constructed using a lightweight neural network; Using knowledge distillation technology, the first evaluation model is used as the teacher model, the initial architecture is used as the student model, and the virtual soil data field is used as the training sample to train the second evaluation model by imitating the first evaluation model, thereby generating the second evaluation model. The distillation loss function in the training process consists of the KL divergence loss between the outputs of the first evaluation model and the second evaluation model, and the mean square error loss between intermediate layer features.
9. The intelligent evaluation method for soil improvement effect driven by multi-source sensor data as described in claim 8, characterized in that, Based on the prediction uncertainty of the second evaluation model, N recommended locations for additional physical sensors to be deployed in the target area are determined, including: Using the virtual soil data field, for each spatial location, the built-in Dropout layer of the second evaluation model is enabled to perform multiple forward propagations during the prediction phase to obtain a set of predicted values for each spatial location. Calculate the statistical variance of the set of predicted values for each spatial location as a quantitative indicator of prediction uncertainty; N spatial locations with a prediction uncertainty quantification index greater than a preset threshold are selected as the N recommended locations.
10. A multi-source sensor data-driven intelligent evaluation system for soil amendment effects, characterized in that, The system is used to implement the intelligent evaluation method for soil amendment effect driven by multi-source sensor data as described in any one of claims 1 to 9, and the system comprises: The first evaluation model training module is used to deploy a high-density sensor network in the source region to acquire soil physicochemical property data, vegetation growth data, meteorological data, and applied improvement measures, and to construct a source region training set to train the first evaluation model. The virtual soil data field generation module is used to acquire meteorological data of the target area, initial monitoring data from sparsely deployed sensors, and preset improvement schemes, and generate a virtual soil data field of the target area based on the first evaluation model. The second evaluation model training module is used to train a lightweight second evaluation model based on the virtual soil data field, and to determine N recommended locations for additional physical sensors to be deployed in the target area based on the prediction uncertainty of the second evaluation model. The soil improvement effect evaluation module is used to update the second evaluation model based on the real monitoring data obtained at the recommended locations, and generate an evaluation result of the soil improvement effect in the target area.