Method for predicting natural abundance of ammonia volatile nitrogen isotope based on environment and soil factors
By improving the Keras Sequential Sequential model, combining the feature-level self-attention module and the DeepSHAP interpreter, the candidate index set is optimized, which solves the problem of low efficiency of traditional ammonia volatile nitrogen isotope analysis methods, and achieves more efficient and accurate prediction of natural nitrogen isotope abundance.
Patent Information
- Application Number
- CN202510570710.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-06
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2045-05-06
AI Technical Summary
The traditional ammonia volatile nitrogen isotope analysis method is too inefficient, making it difficult to predict the dynamic changes in farmland ammonia volatility in real time, and there is a lack of efficient algorithms that comprehensively consider environmental and soil factors.
The improved method based on Keras Sequential sequential model is adopted, and the characteristic-level self-attention module and time decay factor are combined with the DeepSHAP interpreter and the preset shap threshold to optimize the candidate index set to construct an index-optimized natural abundance prediction model.
The prediction efficiency and accuracy of the model are improved, and the problem of excessive size and low efficiency of conventional pre-trained models is solved, which realizes the reduction of model parameters and the integration of correlation between data.
Smart Images

Figure CN120087567A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the cross - technical field of agricultural environmental science and isotope ecology, and particularly relates to a prediction method for the natural abundance of ammonia - volatilized nitrogen isotopes based on environmental and soil factors. Background Art
[0002] Ammonia, as a major alkaline gas in the atmospheric environment and an important component of reactive nitrogen, plays a key role in atmospheric chemical processes and soil nitrogen cycling. However, after a large amount of ammonia enters the environment, it has various negative impacts on human health and ecosystem functions. For example, ammonia in the air easily reacts with acidic substances such as sulfur dioxide and nitrogen oxides to form important precursors of PM2.5, exacerbating air pollution and causing serious harm to human health. In addition, ammonia enters terrestrial and marine ecosystems through atmospheric nitrogen deposition, which may cause soil acidification, water eutrophication, and a reduction in biodiversity. Therefore, clarifying and quantifying the contribution of farmland systems to atmospheric ammonia is the basis for reasonable emission reduction. The natural abundance technique of nitrogen isotopes is an important tool for identifying and quantifying the sources of ammonia in the atmosphere, and can reflect the sources, transformations, and fates of ammonia in ecosystems.
[0003] However, traditional nitrogen isotope analysis methods mainly rely on laboratory isotope mass spectrometry. This method is not only time - consuming and laborious, but also difficult to predict the dynamic changes during the ammonia volatilization process in farmland in real - time. Currently, the calculation method for the natural abundance value of nitrogen isotopes during the ammonia volatilization process in farmland is not yet perfect, and there is a lack of an efficient algorithm that can comprehensively consider environmental and soil factors. Summary of the Invention
[0004] In order to overcome the deficiencies of the prior art, the purpose of the present invention is to provide a prediction method for the natural abundance of ammonia - volatilized nitrogen isotopes based on environmental and soil factors, and solve the problem of the too - low efficiency of traditional ammonia - volatilized nitrogen isotope analysis methods.
[0005] To achieve the above - mentioned purpose, the present invention provides the following solution: A prediction method for the natural abundance of ammonia - volatilized nitrogen isotopes based on environmental and soil factors, comprising: Determine a candidate index set, and perform data collection and data pre - processing in a target research area according to the candidate index set to obtain a candidate index data set; Collect the natural abundance value of ammonia - volatilized nitrogen isotopes in the farmland of the target research area by using an isotope mass spectrometry method to obtain a natural abundance data set; Build a Keras Sequential model, insert a feature-level self-attention module between the input layer and the first hidden layer of the Keras Sequential model, add a time decay factor before the Dropout layer of the Keras Sequential model, replace the original bias in the activation function of the hidden layer of the Keras Sequential model with a fixed bias and a variable bias according to the fixed amount and variable amount in the candidate index set, and judge the enablement of the fixed bias and the variable bias through an indicator function to obtain an improved Keras Sequential model; Build a loss function, and iteratively train the improved Keras Sequential model using the candidate index data set and the natural abundance data set according to the loss function and the early stopping method to obtain an original natural abundance prediction model; Use the DeepSHAP interpreter to calculate the shap value of each candidate index in the original natural abundance prediction model to obtain a candidate index shap data set; Filter the candidate index shap data set using a preset shap threshold to obtain the updated candidate index set; Train the improved Keras Sequential model using the candidate index data set, the data in the natural abundance data set corresponding to the candidate index set to obtain an index-optimized natural abundance prediction model; Use a preset error function to calculate the error degree of the index-optimized natural abundance prediction model with respect to the original natural abundance prediction model to obtain an error ratio; When the error ratio is less than a preset error threshold, determine the index-optimized natural abundance prediction model as the nitrogen isotope natural abundance prediction model; Use the nitrogen isotope natural abundance prediction model to predict a target farmland area to obtain a prediction result.
[0006] Preferably, the iterative training process of the Keras Sequential model includes: Convert the candidate index data set into a tensor form to obtain input data; Use the feature-level self-attention module to perform linear transformation, attention weight calculation, weighted fusion, residual connection and normalization processing on the input data to obtain an attention output; Input the attention output into the hidden layer for abstraction processing to obtain a feature tensor; Use the Dropout layer to perform dynamic retention and random masking processing on the feature tensor to obtain a sparsified tensor; Calculate the sparsified tensor using the remaining hidden layer and the Dropout layer of the Keras Sequential model in sequence to obtain a feature extraction result; Input the feature extraction result into a fully connected layer for ammonia volatilization nitrogen isotope abundance value mapping to obtain an iterative prediction result.
[0007] Preferably, the expression of the activation function is: ; where, is the output of the activation function; is the activation function; is the weight parameter; is the eigenvalue; is the fixed bias; is the variable bias; is the indicator function; is the index type corresponding to the eigenvalue; represents a fixed index; represents a variable index.
[0008] Preferably, the expression of the error function is: ; where, is the error ratio; is the prediction result of the original natural abundance prediction model; is the prediction result of the index-optimized natural abundance prediction model; is the natural abundance measurement value.
[0009] Preferably, the determination of the candidate index set includes: air temperature, atmospheric humidity, soil pH value, soil ammonium nitrogen concentration, soil nitrate nitrogen concentration, soil type, land use type, fertilization type, nitrogen fertilizer application rate, temperature change rate, humidity change rate, ammonium nitrogen concentration change rate, nitrate nitrogen concentration change rate, precipitation, wind speed, light intensity, soil humidity, and soil conductivity.
[0010] Preferably, the range of the shap threshold is from 0.1 to 0.2.
[0011] Preferably, the data preprocessing includes: outlier removal, linear interpolation, mode filling, and data standardization.
[0012] Preferably, the expression of the feature-level self-attention module is: ; where, is the attention output; The layer normalization operation; The normalization exponential function; is the input data matrix; , , are the query matrix, the key matrix, and the value matrix respectively; is the output projection weight matrix; is the scaling factor.
[0013] Preferably, the expression of the Dropout layer is: ; where is the output of the Dropout layer; is the input tensor; is the random mask matrix; is the retention probability; represents the element-wise multiplication operation.
[0014] Preferably, the acquisition devices of the candidate index set include: a glass electrode pH sensor, a digital ammonia and nitrate nitrogen measurement sensor, a resistive soil moisture sensor, and a conductivity sensor.
[0015] The present invention discloses the following technical effects: The present invention provides a method for predicting the natural abundance of ammonia volatilization nitrogen isotope based on environmental and soil factors. By optimizing the candidate index set through the index shap data, the problems of the too large volume and low efficiency of the conventional pre-training model are solved, and the reduction of model parameters is realized; by improving the Keras Sequential model, the problems that the existing Keras Sequential model lacks a synchronous training strategy for different indexes and lacks data correlation processing are solved, and the fusion of data correlation and different training strategies for different types of indexes are realized. Brief Description of the Drawings
[0016] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required to be used in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0017] Figure 1 is the schematic diagram of the prediction process for the natural abundance of ammonia volatilization nitrogen isotope based on environmental and soil factors provided by the embodiments of the present invention; Figure 2 is the schematic diagram of the training process of the improved Keras Sequential model provided by the embodiments of the present invention; Figure 3 Scatter plot distribution of predicted values and measured values of the pre-trained model provided by the embodiments of the present invention; Figure 4 Scatter plot distribution of predicted values and measured values of the multiple linear model provided by the embodiments of the present invention; Figure 5 Schematic diagram of importance evaluation provided by the embodiments of the present invention. Detailed implementation manners
[0018] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0019] The purpose of the present invention is to provide a prediction method for the natural abundance of ammonia volatilization nitrogen isotope based on environmental and soil factors, so as to solve the problem of too low efficiency of traditional ammonia volatilization nitrogen isotope analysis methods.
[0020] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific implementation manners.
[0021] Figure 1 Schematic diagram of the prediction process for the natural abundance of ammonia volatilization nitrogen isotope based on environmental and soil factors provided by the embodiments of the present invention. As Figure 1 shown, the present invention provides a prediction method for the natural abundance of ammonia volatilization nitrogen isotope based on environmental and soil factors, including: Step 100: Determine a candidate index set, and collect and preprocess data in the target research area according to the candidate index set to obtain a candidate index data set; Step 200: Use the isotope mass spectrometry analysis method to collect the natural abundance values of ammonia volatilization nitrogen isotope in the farmland of the target research area to obtain a natural abundance data set; Step 300: Build a Keras Sequential model, insert a feature-level self-attention module between the input layer and the first hidden layer of the Keras Sequential model, add a time decay factor before the Dropout layer of the Keras Sequential model, and replace the original bias in the activation function of the hidden layer of the Keras Sequential model with a fixed bias and a variable bias according to the fixed quantity and variable quantity in the candidate index set, and judge the enablement of the fixed bias and the variable bias through an indicator function to obtain an improved Keras Sequential model; Step 400: Construct a loss function, and iteratively train the improved Keras Sequential model using the candidate index dataset and the natural abundance dataset according to the loss function and the early stopping method to obtain an original natural abundance prediction model; Step 500: Use the DeepSHAP interpreter to calculate the SHAP values of each candidate index in the original natural abundance prediction model to obtain a candidate index SHAP dataset; Step 600: Filter the candidate index SHAP dataset using a preset SHAP threshold to obtain the updated candidate index set; Step 700: Train the improved Keras Sequential model using the data in the candidate index dataset, the natural abundance dataset corresponding to the candidate index set to obtain an index-optimized natural abundance prediction model; Step 800: Calculate the error degree of the index-optimized natural abundance prediction model with respect to the original natural abundance prediction model using a preset error function to obtain an error ratio; Step 900: When the error ratio is less than a preset error threshold, determine the index-optimized natural abundance prediction model as the nitrogen isotope natural abundance prediction model; Step 1000: Use the nitrogen isotope natural abundance prediction model to predict a target farmland area to obtain a prediction result.
[0022] Reference Figure 2 , the iterative training process of the Keras Sequential model includes: Step 401: Convert the candidate index dataset into a tensor form to obtain input data; Step 402: Use the feature-level self-attention module to perform linear transformation, attention weight calculation, weighted fusion, residual connection, and normalization processing on the input data to obtain an attention output; Step 403: Input the attention output into the hidden layer for abstraction processing to obtain a feature tensor; Step 404: Use the Dropout layer to perform dynamic retention and random masking processing on the feature tensor to obtain a sparsified tensor; Step 405: Successively use the remaining hidden layers and the Dropout layer of the Keras Sequential model to calculate the sparsified tensor to obtain a feature extraction result; Step 406: Input the feature extraction result into a fully connected layer to map the ammonia volatilization nitrogen isotope abundance value to obtain an iterative prediction result.
[0023] Furthermore, the expression of the activation function is: ; where is the output of the activation function; is the activation function; is the weight parameter; is the eigenvalue; is the fixed bias; is the variable bias; is the indicator function; is the index type corresponding to the eigenvalue; represents the fixed index; represents the variable index.
[0024] Specifically, the expression of the error function is: ; where is the error ratio; is the prediction result of the original natural abundance prediction model; is the prediction result of the index-optimized natural abundance prediction model; is the measured value of natural abundance.
[0025] Optionally, the determination of the candidate index set includes: air temperature, atmospheric humidity, soil pH value, soil ammonium nitrogen concentration, soil nitrate nitrogen concentration, soil type, land use type, fertilization type, nitrogen fertilizer application rate, temperature change rate, humidity change rate, ammonium nitrogen concentration change rate, nitrate nitrogen concentration change rate, precipitation, wind speed, light intensity, soil humidity, and soil conductivity.
[0026] Specifically, the range of the shap threshold is from 0.1 to 0.2.
[0027] Preferably, the data preprocessing includes: outlier removal, linear interpolation, mode filling, and data standardization.
[0028] Furthermore, the expression of the feature-level self-attention module is: ; where is the attention output; represents the layer normalization operation; represents the normalization exponential function; is the input data matrix; 、 、 are the query matrix, key matrix, and value matrix respectively; is the output projection weight matrix; is the scaling factor.
[0029] Specifically, the expression of the Dropout layer is: ; in, is the output of the Dropout layer; is the input tensor; is a random mask matrix; To retain the probability; Represents an element-wise multiplication operation.
[0030] Preferably, the collection equipment of the candidate indicator set includes: a glass electrode pH sensor, a digital ammonia nitrogen and nitrate nitrogen measurement sensor, a resistive soil moisture sensor and a conductivity sensor.
[0031] Further, Figure 3 It is a scatter plot of the predicted value and the measured value of the pre-trained model. Figure 4 It is a scatter plot of the predicted values and measured values of the multivariate linear model. By comparison, it can be determined that the pre-trained model of this embodiment is more accurate than the multivariate linear model in predicting the natural abundance of nitrogen isotopes in farmland ammonia volatilization.
[0032] Specifically, data collection and preprocessing: obtain the natural abundance values of ammonia volatilization nitrogen isotopes in the target area, and simultaneously collect the corresponding environmental factor data set and soil factor data set; the environmental factors include temperature, humidity, etc., and the soil factors include soil pH value, ammonium nitrogen concentration, nitrate nitrogen concentration, soil type, etc.; perform missing value processing, standardization or normalization operations on the data, and construct training sets and test sets. The test set is used to verify the prediction performance of the two models, and the determination coefficient ( ), root mean square error (RMSE) and mean absolute error (MAE) indicators to compare model accuracy.
[0033] Furthermore, data preprocessing included: outlier removal, using the box plot method or the 3σ principle, defining data that exceeded 1.5 times the upper and lower interquartile range or exceeded the mean ±3 times the standard deviation as outliers; samples with missing values exceeding the threshold (≥20%) were removed, and the remaining missing values were interpolated using the multiple interpolation method (MICE), with the number of interpolation iterations set to 10 and the convergence tolerance set to 0.001; continuous variables were standardized by Z-score, using the formula: in, For standard output, For the data to be processed, is the mean, is the standard deviation; One-Hot Encoding is performed on categorical variables.
[0034] Preferably, the time decay function is used to dynamically adjust the neuron retention probability to alleviate the noise accumulation problem in time series data. Its calculation formula is: Here, t represents the time step.
[0035] Optionally, this embodiment chooses to directly obtain data such as temperature and humidity from the weather report website while ensuring that the temporal resolution and spatial coverage match the farmland prediction requirements. Static indicators such as soil type need to be collected once at the beginning of the prediction, while dynamic indicators such as ammonium nitrogen concentration need to be predicted at a high frequency (once every hour).
[0036] Optionally, tensor conversion: the preprocessed structured data is converted into a three-dimensional tensor (number of samples × time steps × number of features). For example, hourly data for 5 consecutive days forms a 120×18 tensor slice. The time dimension generates sequence samples through a sliding window, with the window length set to 24 (representing a single-day cycle) and the step size set to 6 to achieve overlapping sampling to enhance data utilization.
[0037] Specifically, in the linear projection layer, the input tensor first passes through three independent fully connected layers to generate query, key, and value matrices respectively. The projection dimension is set to 1 / 4 of the original number of features to achieve a balance between computational efficiency and information retention. For example, the weight matrices W_Q, W_K, and W_V of the 18-dimensional feature projected into the 4-dimensional space have dimensions of 18×4.
[0038] Further, the attention weights are calculated by matrix multiplication to obtain the similarity matrix between features: the query matrix is multiplied by the transposed key matrix, and the scaling factor d is taken as the square root of the projection dimension. The obtained similarity matrix is normalized to generate the attention weight matrix. Each element of this matrix reflects the strength of association between two features in the current context.
[0039] Furthermore, feature fusion and residual connection: the value matrix is weighted and summed using attention weights to obtain context-aware feature representation. The output projection layer maps the fused features back to the original dimension. To prevent information loss, the original input is added to the attention output, and then the gradient propagation is stabilized through layer normalization. This process enables the model to autonomously identify key feature combinations, such as the synergistic effect of soil pH and ammonium nitrogen concentration.
[0040] Specifically, in the model initialization stage, a type mapping table is established according to the characteristic properties: static indicators (such as soil type) are marked as , dynamic indicators (such as temperature change rate) are marked as . Each input feature enters the network with a type label.
[0041] Preferably, two sets of bias parameters are set in the hidden layer activation function: (fixed bias) and (variable bias). During forward propagation, the corresponding bias term is selected and enabled according to the feature type label. The indicator function is essentially a conditional judgment gate. When the feature belongs to , it participates in the calculation; when it belongs to , it takes effect. This mechanism enables the network to establish a stable representation for static features and maintain a flexible response to dynamic features.
[0042] Optionally, a smaller learning rate (1 / 10 of the model learning rate) is adopted for the fixed bias to prevent excessive adjustment of the parameters related to static features; an adaptive learning rate optimizer (such as Nadam) is used for the variable bias to quickly capture the change law of dynamic indicators. The gradients of the two types of biases are calculated separately during backpropagation to ensure training stability.
[0043] Preferably, a time decay Dropout layer is designed to retain the probability decay curve: the initial retention probability is set to 0.8, the decay coefficient is 0.01, t represents the number of training epochs (time steps), and after each epoch is completed, the global retention probability drops by about 1%. This design allows more neurons to participate in learning in the initial stage and gradually sparsifies in the later stage to inhibit overfitting.
[0044] Furthermore, dynamic mask generation is performed, and a new random mask matrix is generated for each training, with the element values being 0 or 1, and the occurrence probability being determined by the current p(t). In the first 100 epochs, p(t) slowly drops from 0.8 to 0.29, and the network maintains a strong learning ability; after 300 epochs, p(t) approaches 0.04, forming a highly sparse connection pattern.
[0045] Specifically, during the inference stage processing, the Dropout layer is turned off during testing, but the average activation intensity of each layer of neurons in the training stage is retained. By recording the typical mask patterns at the end of training, the weights are scaled and compensated during inference to maintain the consistency of the output magnitude.
[0046] Preferably, the Huber loss function is adopted as the loss function. This function adaptively switches between the mean squared error (MSE) and the mean absolute error (MAE), with the set threshold being 1.0. When the prediction deviation is less than the threshold, MSE is used to promote convergence, and when it is greater than the threshold, it switches to MAE to enhance robustness.
[0047] Optionally, monitor the smoothed loss value on the validation set: calculate the moving average of the loss for the last 10 epochs, and trigger early stopping when the average loss no longer decreases for three consecutive times. To prevent misjudgment due to occasional fluctuations, set the patience coefficient to 15 epochs, allowing the loss to fluctuate within the threshold range.
[0048] Specifically, use the Nadam optimizer (Adam + NAG momentum), with an initial learning rate of 3 , combined with a triangular cyclic learning rate schedule: it periodically changes between 1 and 3 every 5 epochs to promote escaping from local optima. The batch size is set to 64, taking into account both memory efficiency and gradient estimation stability.
[0049] Furthermore, after completing the initial training, perform perturbation analysis on the validation set samples. Each feature is perturbed by ±10% in turn, and the change amplitude of the model output is recorded. Calculate the shap value through Monte Carlo sampling to quantify the marginal contribution of each feature to the prediction result, and map the shap value to a percentage system, referring to Figure 5 , where NH3 represents ammonia concentration, soil_pH represents soil pH value, soil_NH4 represents soil ammonium nitrogen concentration; soil_NO3 represents soil nitrate nitrogen concentration, Gleyi-Stagnic Anthrosol represents Gleyi-Stagnic Anthrosol, soil_humidity represents soil humidity, incubation_days represents the number of measurement days; N_application represents nitrogen fertilizer application level; Yellow-brown soil represents Yellow-brown soil; Forest represents forest; Vineyard represents vineyard farmland.
[0050] The environmental factors include temperature, humidity, etc., and the soil factors include soil pH value, ammonium nitrogen concentration, nitrate nitrogen concentration, soil type, etc.; perform missing value processing, standardization or normalization operations on the data.
[0051] Furthermore, the threshold screening strategy sets the absolute shap value threshold to 60% of the maximum value (which can be adjusted according to the data scale), and retains the top 60% of the features in terms of contribution. Bundle and screen strongly correlated feature groups (such as temperature and temperature change rate) to avoid misdeleting related features. The screened feature set must simultaneously meet the following conditions: the shap value of a single feature is greater than the threshold, and at least one representative item is retained within the feature group.
[0052] Optionally, select 3 test areas with different climate types outside the target area and use transfer learning for fine-tuning: freeze the weights of the first 3 layers of the network and only train the last fully connected layer. The fine-tuning period is limited within 20 epochs to verify the environmental adaptability of the model.
[0053] Preferably, for online prediction deployment: convert the final model into the TensorRT format and deploy it on edge computing devices. Design a two-level caching mechanism: real-time data is first stored in a circular buffer, and model inference is triggered every time the data is full for 1 day. The output results are uploaded to the cloud platform through the LoRa wireless module to achieve unattended prediction in the field.
[0054] Furthermore, conduct visual monitoring during the training process. Regularly export the self-attention matrix of the first layer and draw a heat map of feature correlation. By analyzing the evolution trend of the attention weights, verify whether the model has learned the expected physical and chemical relationships; record the activation value distribution of each hidden layer and monitor whether there is gradient disappearance / explosion. Ideally, the values after ReLU activation should show a right-skewed distribution; perform time series analysis on the SHAP values to identify the periodic changes in the contribution degrees of key features.
[0055] The beneficial effects of the present invention are as follows: The present invention optimizes the candidate index set through the shap data of the index, realizes the reduction of model parameters, and improves the prediction efficiency of the model; by improving the model with Keras Sequential, it realizes the fusion of the correlations between data and different training strategies for different types of indexes, solves the overfitting risk during the model training process, and improves the prediction accuracy of the model for natural abundance.
[0056] In this specification, each embodiment is described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. For the same or similar parts among the embodiments, reference can be made to each other.
[0057] In this article, specific examples are used to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, based on the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present invention.
Claims
1. A method for predicting the natural abundance of nitrogen isotopes from ammonia volatilization based on environmental and soil factors, characterized in that: include: Determine a candidate indicator set, and perform data collection and data preprocessing in a target research area according to the candidate indicator set to obtain a candidate indicator data set; The natural abundance values of nitrogen isotopes in farmland ammonia volatilization in the target study area are collected by isotope mass spectrometry analysis to obtain a natural abundance data set; Construct a Keras Sequential model, insert a feature-level self-attention module between the input layer and the first hidden layer of the Keras Sequential model, add a time decay factor before the Dropout layer of the Keras Sequential model, replace the original bias in the activation function of the hidden layer of the Keras Sequential model with a fixed bias and a variable bias according to the fixed amount and the variable amount in the candidate indicator set, and determine the activation of the fixed bias and the variable bias through an indicator function to obtain a Keras Sequential improved model; Constructing a loss function, and iteratively training the Keras Sequential improved model using the candidate indicator dataset and the natural abundance dataset according to the loss function and the early stopping method to obtain an original natural abundance prediction model; Using the DeepSHAP interpreter to calculate the shap value of each candidate indicator in the original natural abundance prediction model, and obtain a candidate indicator shap data set; Filtering the candidate indicator shap data set using a preset shap threshold to obtain an updated candidate indicator set; The Keras Sequential improved model is trained using the candidate indicator data set, the data in the natural abundance data set and the data corresponding to the candidate indicator set to obtain an indicator-optimized natural abundance prediction model; Calculating the error degree of the index-optimized natural abundance prediction model with respect to the original natural abundance prediction model using a preset error function to obtain an error ratio; When the error ratio is less than a preset error threshold, determining the index-optimized natural abundance prediction model as a nitrogen isotope natural abundance prediction model; The nitrogen isotope natural abundance prediction model is used to predict the target farmland area to obtain a prediction result.
2. The method for predicting the natural abundance of nitrogen isotopes from ammonia volatilization based on environmental and soil factors according to claim 1, characterized in that: The iterative training process of the Keras Sequential model includes: Convert the candidate indicator data set into a tensor form to obtain input data; Using the feature-level self-attention module to perform linear transformation, attention weight calculation, weighted fusion, residual connection and normalization on the input data to obtain an attention output; Inputting the attention output into the hidden layer for abstract processing to obtain a feature tensor; The Dropout layer is used to dynamically retain and randomly mask the feature tensor to obtain a sparse tensor; The remaining hidden layers and the Dropout layers of the Keras Sequential model are used in sequence to calculate the sparsified tensor to obtain a feature extraction result; The feature extraction result is input into the fully connected layer to map the ammonia volatilization nitrogen isotope abundance value to obtain an iterative prediction result.
3. The method for predicting the natural abundance of nitrogen isotopes of ammonia volatilization based on environmental and soil factors according to claim 1, characterized in that: The expression of the activation function is: ; in, is the output of the activation function; is the activation function; is the weight parameter; is the characteristic value; is the fixed bias; biasing the variation; is the indicative function; is the indicator type corresponding to the characteristic value; Indicates a fixed indicator; Indicates the change indicator.
4. The method for predicting the natural abundance of nitrogen isotopes from ammonia volatilization based on environmental and soil factors according to claim 1, characterized in that: The expression of the error function is: ; in, is the error ratio; is the prediction result of the original natural abundance prediction model; Optimizing the prediction results of the natural abundance prediction model for the indicator; is the natural abundance value.
5. The method for predicting the natural abundance of nitrogen isotopes of ammonia volatilization based on environmental and soil factors according to claim 1, characterized in that: The candidate indicator set determined includes: air temperature, atmospheric humidity, soil pH value, soil ammonium nitrogen concentration, soil nitrate nitrogen concentration, soil type, land use type, fertilization type, nitrogen fertilizer application amount, temperature change rate, humidity change rate, ammonium nitrogen concentration change rate, nitrate nitrogen concentration change rate, precipitation, wind speed, light intensity, soil moisture and soil conductivity.
6. The method for predicting the natural abundance of nitrogen isotopes from ammonia volatilization based on environmental and soil factors according to claim 1, characterized in that: The shap threshold ranges from 0.1 to 0.
2.
7. The method for predicting the natural abundance of nitrogen isotopes from ammonia volatilization based on environmental and soil factors according to claim 1, characterized in that: The data preprocessing includes: outlier removal, linear interpolation, mode filling and data standardization.
8. The method for predicting the natural abundance of nitrogen isotopes from ammonia volatilization based on environmental and soil factors according to claim 2, characterized in that: The expression of the feature-level self-attention module is: ; in, is the attention output; Representation layer normalization operation; represents the normalized exponential function; is the input data matrix; , , They are query matrix, key matrix, and value matrix respectively; is the output projection weight matrix; is the scaling factor.
9. The method for predicting the natural abundance of nitrogen isotopes from ammonia volatilization based on environmental and soil factors according to claim 2, characterized in that: The expression of the Dropout layer is: ; in, is the output of the Dropout layer; is the input tensor; is a random mask matrix; To retain the probability; Represents an element-wise multiplication operation.
10. The method for predicting the natural abundance of nitrogen isotopes of ammonia volatilization based on environmental and soil factors according to claim 5, characterized in that: The collection equipment of the candidate indicator set includes: a glass electrode pH sensor, a digital ammonia nitrogen and nitrate nitrogen measurement sensor, a resistive soil moisture sensor and a conductivity sensor.
Citation Information
Patent Citations
Geological disaster risk assessment method and system fusing random forest and attention
CN116167617A
Distortion correction face recognition large-angle recognition algorithm
CN118522062A
Carbon accounting process factor correction system based on LSTM network model
CN118608162A
Prediction method for corrosion rate in natural gas pipeline
CN119202558A
Pesticide residue prediction method based on LSTM variant
CN119359084A
Cited By
Rapid distinguishing method for predicting nitrogen source produced by flooded soil through combination of in-situ observation and model
CN120741822A