A prediction method for the natural abundance of ammonia volatilization nitrogen isotopes based on environmental and soil factors
By constructing the improved Keras Sequential model and the DeepSHAP interpreter optimization index set, the problem of low efficiency of traditional nitrogen isotope analysis methods is solved, and efficient and accurate prediction of natural nitrogen isotope abundance is achieved.
Patent Information
- Application Number
- CN202510570710.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-06
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-05-06
AI Technical Summary
Traditional nitrogen isotope analysis methods are time-consuming and labor-intensive, and it is difficult to predict the dynamic changes in farmland ammonia volatility in real time, and there is a lack of efficient algorithms that comprehensively consider environmental and soil factors.
The Keras Sequential sequential model was constructed, the feature-level self-attention module was inserted, the time decay factor was added, the candidate index set was optimized using the DeepSHAP interpreter, the loss function and early stop method were constructed for iterative training, and the model was optimized to predict the natural abundance of nitrogen isotopes.
The reduction of model parameters is achieved, the prediction efficiency is improved, the risk of overfitting is solved, and the accuracy and real-time prediction of natural abundance of nitrogen isotopes is improved.
Smart Images

Figure CN120087567B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the cross - technical field of agricultural environmental science and isotope ecology, and particularly to a prediction method for the natural abundance of ammonia - volatilized nitrogen isotopes based on environmental and soil factors. Background Art
[0002] As a major alkaline gas in the atmospheric environment and an important component of reactive nitrogen, ammonia plays a key role in atmospheric chemical processes and soil nitrogen cycling. However, after a large amount of ammonia enters the environment, it has various negative impacts on human health and ecosystem functions. For example, ammonia in the air is highly reactive with acidic substances such as sulfur dioxide and nitrogen oxides, forming important precursors of PM2.5, exacerbating air pollution and causing serious harm to human health. In addition, ammonia enters terrestrial and marine ecosystems through atmospheric nitrogen deposition, which may cause soil acidification, water eutrophication, and a reduction in biodiversity. Therefore, clarifying and quantifying the contribution of farmland systems to atmospheric ammonia is the basis for reasonable emission reduction. The natural abundance technique of nitrogen isotopes is an important tool for identifying and quantifying the sources of ammonia in the atmosphere, and can reflect the sources, transformations, and fates of ammonia in ecosystems.
[0003] However, traditional nitrogen isotope analysis methods mainly rely on laboratory isotope mass spectrometry. This method is not only time - consuming and laborious, but also difficult to predict the dynamic changes during the ammonia volatilization process in farmland in real - time. Currently, the calculation method for the natural abundance value of nitrogen isotopes during the ammonia volatilization process in farmland is not yet perfect, lacking an efficient algorithm that can comprehensively consider environmental and soil factors. Summary of the Invention
[0004] In order to overcome the deficiencies of the prior art, the purpose of the present invention is to provide a prediction method for the natural abundance of ammonia - volatilized nitrogen isotopes based on environmental and soil factors, and to solve the problem of the too - low efficiency of traditional ammonia - volatilized nitrogen isotope analysis methods.
[0005] To achieve the above - mentioned purpose, the present invention provides the following solution:
[0006] A prediction method for the natural abundance of ammonia - volatilized nitrogen isotopes based on environmental and soil factors, comprising:
[0007] Determine a candidate index set, and collect and pre - process data in a target research area according to the candidate index set to obtain a candidate index data set;
[0008] Collect the natural abundance value of ammonia - volatilized nitrogen isotopes in the farmland of the target research area by using an isotope mass spectrometry analysis method to obtain a natural abundance data set;
[0009] Build a Keras Sequential model, insert a feature-level self-attention module between the input layer and the first hidden layer of the Keras Sequential model, add a time decay factor before the Dropout layer of the Keras Sequential model, replace the original bias in the activation function of the hidden layer of the Keras Sequential model with a fixed bias and a variable bias according to the fixed quantity and variable quantity in the candidate index set, and determine the enablement of the fixed bias and the variable bias through an indicator function to obtain an improved Keras Sequential model;
[0010] Build a loss function, and iteratively train the improved Keras Sequential model using the candidate index data set and the natural abundance data set according to the loss function and the early stopping method to obtain an original natural abundance prediction model;
[0011] Use the DeepSHAP interpreter to calculate the shap values of each candidate index in the original natural abundance prediction model to obtain a candidate index shap data set;
[0012] Filter the candidate index shap data set using a preset shap threshold to obtain the updated candidate index set;
[0013] Train the improved Keras Sequential model using the data corresponding to the candidate index data set, the natural abundance data set, and the candidate index set to obtain an index-optimized natural abundance prediction model;
[0014] Calculate the error degree of the index-optimized natural abundance prediction model with respect to the original natural abundance prediction model using a preset error function to obtain an error ratio;
[0015] When the error ratio is less than a preset error threshold, determine the index-optimized natural abundance prediction model as the nitrogen isotope natural abundance prediction model;
[0016] Use the nitrogen isotope natural abundance prediction model to predict a target farmland area to obtain a prediction result.
[0017] Preferably, the iterative training process of the Keras Sequential model includes:
[0018] Convert the candidate index data set into a tensor form to obtain input data;
[0019] The input data is subjected to linear transformation, attention weight calculation, weighted fusion, residual connection and normalization processing by using the feature-level self-attention module to obtain an attention output;
[0020] The attention output is input into the hidden layer for abstraction processing to obtain a feature tensor;
[0021] The Dropout layer is used to dynamically retain and randomly mask the feature tensor to obtain a sparsified tensor;
[0022] The remaining hidden layers and Dropout layers of the Keras Sequential model are successively used to calculate the sparsified tensor to obtain a feature extraction result;
[0023] The feature extraction result is input into a fully connected layer for ammonia volatilization nitrogen isotope abundance value mapping to obtain an iterative prediction result.
[0024] Preferably, the expression of the activation function is:
[0025]
[0026] where f(x i ) is the output of the activation function; ReLU(·) is the activation function; w i is a weight parameter; x i is a feature value; b fixed is the fixed bias; b var is the variable bias; I(·) is an indicator function; is the index type corresponding to the feature value; T0 represents a fixed index; T1 represents a variable index.
[0027] Preferably, the expression of the error function is:
[0028]
[0029] where E is the error ratio; is the prediction result of the original natural abundance prediction model; is the prediction result of the index-optimized natural abundance prediction model; y n is the natural abundance measurement value.
[0030] Preferably, the determination of the candidate index set includes: air temperature, atmospheric humidity, soil pH value, soil ammonium nitrogen concentration, soil nitrate nitrogen concentration, soil type, land use type, fertilization type, nitrogen fertilizer application rate, temperature change rate, humidity change rate, ammonium nitrogen concentration change rate, nitrate nitrogen concentration change rate, precipitation, wind speed, light intensity, soil humidity and soil conductivity.
[0031] Preferably, the range of the shap threshold is from 0.1 to 0.2.
[0032] Preferably, the data preprocessing includes: outlier removal, linear interpolation, mode filling, and data standardization.
[0033] Preferably, the expression of the feature-level self-attention module is:
[0034]
[0035] where Output is the attention output; LayerNorm(·) represents the layer normalization operation; Softmax(·) represents the normalized exponential function; X is the input data matrix; W Q , W K , W V are the query matrix, key matrix, and value matrix respectively; W O is the output projection weight matrix; d is the scaling factor.
[0036] Preferably, the expression of the Dropout layer is:
[0037]
[0038] where Dropout(X t , t) is the output of the Dropout layer; X t is the input tensor; M t is the random mask matrix; p(t) is the retention probability; ⊙ represents the element-wise multiplication operation.
[0039] Preferably, the acquisition devices of the candidate index set include: a glass electrode pH sensor, a digital ammonia nitrogen and nitrate nitrogen measurement sensor, a resistive soil moisture sensor, and a conductivity sensor.
[0040] The present invention discloses the following technical effects:
[0041] The present invention provides a prediction method for the natural abundance of ammonia volatilization nitrogen isotopes based on environmental and soil factors. By optimizing the candidate index set through index shap data, the problems of the overly large size and low efficiency of conventional pre-trained models are solved, and the reduction of model parameters is achieved; by improving the model with Keras Sequential, the problems of the existing Keras Sequential model lacking a synchronous training strategy for different indicators and lacking data correlation processing are solved, and the fusion of data correlations and different training strategies for different types of indicators are achieved. Description of the Drawings
[0042] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the accompanying drawings required for use in the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.
[0043] Figure 1 Schematic diagram of the prediction process of the natural abundance of ammonia volatilization nitrogen isotope based on environmental and soil factors provided by the embodiment of the present invention;
[0044] Figure 2 Schematic diagram of the training process of the improved Keras Sequential model provided by the embodiment of the present invention;
[0045] Figure 3 Scatter plot distribution diagram of the predicted value and the measured value of the pre-trained model provided by the embodiment of the present invention;
[0046] Figure 4 Scatter plot distribution diagram of the predicted value and the measured value of the multiple linear model provided by the embodiment of the present invention;
[0047] Figure 5 Schematic diagram of importance evaluation provided by the embodiment of the present invention. Detailed implementation manners
[0048] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0049] The purpose of the present invention is to provide a method for predicting the natural abundance of ammonia volatilization nitrogen isotope based on environmental and soil factors, and solve the problem of too low efficiency of traditional ammonia volatilization nitrogen isotope analysis methods.
[0050] To make the above objects, features and advantages of the present invention more obvious and understandable, the following will further describe the present invention in detail with reference to the accompanying drawings and specific implementation manners.
[0051] Figure 1 Schematic diagram of the prediction process of the natural abundance of ammonia volatilization nitrogen isotope based on environmental and soil factors provided by the embodiment of the present invention. As Figure 1 shown, the present invention provides a method for predicting the natural abundance of ammonia volatilization nitrogen isotope based on environmental and soil factors, including:
[0052] Step 100: Determine the candidate index set, and perform data collection and data preprocessing in the target research area according to the candidate index set to obtain a candidate index data set;
[0053] Step 200: Use the isotope mass spectrometry analysis method to collect the natural abundance value of farmland ammonia volatilization nitrogen isotope in the target research area to obtain a natural abundance data set;
[0054] Step 300: Build a Keras Sequential sequential model, insert a feature-level self-attention module between the input layer and the first hidden layer of the Keras Sequential sequential model, add a time decay factor before the Dropout layer of the Keras Sequential sequential model, and according to the fixed quantity and variable quantity in the candidate index set, replace the original bias in the activation function of the hidden layer of the Keras Sequential sequential model with a fixed bias and a variable bias, and judge the enabling of the fixed bias and the variable bias through an indicator function to obtain an improved Keras Sequential model;
[0055] Step 400: Build a loss function, and use the candidate index data set and the natural abundance data set to iteratively train the improved Keras Sequential model according to the loss function and the early stopping method to obtain an original natural abundance prediction model;
[0056] Step 500: Use the DeepSHAP interpreter to calculate the shap value of each candidate index in the original natural abundance prediction model to obtain a candidate index shap data set;
[0057] Step 600: Use a preset shap threshold to filter the candidate index shap data set to obtain the updated candidate index set;
[0058] Step 700: Use the candidate index data set, the natural abundance data set, and the data corresponding to the candidate index set to train the improved Keras Sequential model to obtain an index-optimized natural abundance prediction model;
[0059] Step 800: Use a preset error function to calculate the error degree of the index-optimized natural abundance prediction model with respect to the original natural abundance prediction model to obtain an error ratio;
[0060] Step 900: When the error ratio is less than a preset error threshold, determine the index-optimized natural abundance prediction model as the nitrogen isotope natural abundance prediction model;
[0061] Step 1000: Use the natural abundance prediction model of nitrogen isotopes to predict the target farmland area and obtain a prediction result.
[0062] Reference Figure 2 , the iterative training process of the Keras Sequential model includes:
[0063] Step 401: Convert the candidate index data set into a tensor form to obtain input data;
[0064] Step 402: Use the feature-level self-attention module to perform linear transformation, attention weight calculation, weighted fusion, residual connection, and normalization processing on the input data to obtain an attention output;
[0065] Step 403: Input the attention output into the hidden layer for abstraction processing to obtain a feature tensor;
[0066] Step 404: Use the Dropout layer to perform dynamic retention and random masking processing on the feature tensor to obtain a sparsified tensor;
[0067] Step 405: Successively use the remaining hidden layers and Dropout layers of the Keras Sequential model to calculate the sparsified tensor to obtain a feature extraction result;
[0068] Step 406: Input the feature extraction result into the fully connected layer for ammonia volatilization nitrogen isotope abundance value mapping to obtain an iterative prediction result.
[0069] Furthermore, the expression of the activation function is:
[0070]
[0071] where f(x i ) is the output of the activation function; ReLU(·) is the activation function; w i is the weight parameter; x i is the eigenvalue; b fixed is the fixed bias; b var is the variable bias; I(·) is the indicator function; is the index type corresponding to the eigenvalue; T0 represents the fixed index; T1 represents the variable index.
[0072] Specifically, the expression of the error function is:
[0073]
[0074] where E is the error ratio; is the prediction result of the original natural abundance prediction model; is the prediction result of the index-optimized natural abundance prediction model; y n is the measured value of natural abundance.
[0075] Optionally, the determination of the candidate index set includes: air temperature, atmospheric humidity, soil pH value, soil ammonium nitrogen concentration, soil nitrate nitrogen concentration, soil type, land use type, fertilization type, nitrogen fertilizer application rate, temperature change rate, humidity change rate, ammonium nitrogen concentration change rate, nitrate nitrogen concentration change rate, precipitation, wind speed, light intensity, soil humidity, and soil conductivity.
[0076] Specifically, the range of the SHAP threshold is from 0.1 to 0.2.
[0077] Preferably, the data preprocessing includes: outlier removal, linear interpolation, mode filling, and data standardization.
[0078] Furthermore, the expression of the feature-level self-attention module is:
[0079]
[0080] where Output is the attention output; LayerNorm(·) represents the layer normalization operation; Softmax(·) represents the normalized exponential function; X is the input data matrix; W Q , W K , W V are the query matrix, key matrix, and value matrix respectively; W O is the output projection weight matrix; d is the scaling factor.
[0081] Specifically, the expression of the Dropout layer is:
[0082]
[0083] where Dropout(X t , t) is the output of the Dropout layer; X t is the input tensor; M t is the random mask matrix; p(t) is the retention probability; ⊙ represents the element-wise multiplication operation.
[0084] Preferably, the acquisition devices for the candidate index set include: glass electrode pH sensor, digital ammonia and nitrate nitrogen measurement sensor, resistive soil humidity sensor, and conductivity sensor.
[0085] Furthermore, Figure 3 is the scatter plot of the predicted values and measured values of the pre-trained model, Figure 4It is a scatter plot of the predicted values and measured values of the multivariate linear model. By comparison, it can be determined that the pre-trained model of this embodiment is more accurate than the multivariate linear model in predicting the natural abundance of nitrogen isotopes in farmland ammonia volatilization.
[0086] Specifically, data collection and preprocessing: obtain the natural abundance values of ammonia volatilization nitrogen isotopes in the target area, and simultaneously collect the corresponding environmental factor data set and soil factor data set; the environmental factors include temperature, humidity, etc., and the soil factors include soil pH value, ammonium nitrogen concentration, nitrate nitrogen concentration, soil type, etc.; perform missing value processing, standardization or normalization operations on the data, and construct training sets and test sets. The test set is used to verify the prediction performance of the two models, and the coefficient of determination (R 2 ), root mean square error (RMSE) and mean absolute error (MAE) indicators to compare model accuracy.
[0087] Furthermore, data preprocessing includes: outlier removal, using the box plot method or the 3σ principle, defining data that exceeds 1.5 times the upper and lower interquartile range or exceeds the mean ± 3 times the standard deviation as outliers; samples with missing values exceeding the threshold (≥20%) are removed, and the remaining missing values are interpolated using the multiple interpolation method (MICE), setting the number of interpolation iterations to 10 times and the convergence tolerance to 0.001; Z-score standardization of continuous variables, the formula is:
[0088]
[0089] Among them, z is the standard output, x is the data to be processed, μ is the mean, and σ is the standard deviation; One-Hot Encoding is performed on categorical variables.
[0090] Preferably, the time decay function is used to dynamically adjust the neuron retention probability to alleviate the noise accumulation problem in time series data. Its calculation formula is:
[0091] p(t)=0.8·e -0.01t
[0092] Here, t represents the time step.
[0093] Optionally, this embodiment chooses to directly obtain data such as temperature and humidity from the weather report website while ensuring that the temporal resolution and spatial coverage match the farmland prediction requirements. Static indicators such as soil type need to be collected once at the beginning of the prediction, while dynamic indicators such as ammonium nitrogen concentration need to be predicted at a high frequency (once every hour).
[0094] Optionally, tensor conversion: The preprocessed structured data is converted into a three-dimensional tensor (number of samples × number of time steps × number of features). For example, hourly data for 5 consecutive days forms a tensor slice of 120×18. The time dimension generates sequence samples through a sliding window, with the window length set to 24 (representing a single-day cycle) and the step size set to 6 to achieve overlapping sampling and enhance data utilization.
[0095] Specifically, for the linear projection layer, the input tensor first passes through three independent fully connected layers to generate Query, Key, and Value matrices respectively. The projection dimension is set to 1 / 4 of the original number of features, achieving a balance between computational efficiency and information retention. For example, the weight matrices W_Q, W_K, and W_V for projecting 18-dimensional features into a 4-dimensional space have dimensions of 18×4.
[0096] Furthermore, for attention weight calculation, a similarity matrix between features is obtained through matrix multiplication: the query matrix is multiplied by the transposed key matrix, and the scaling factor d is taken as the square root of the projection dimension. The resulting similarity matrix is normalized to generate an attention weight matrix. Each element of this matrix reflects the association strength between two features in the current context.
[0097] Even further, for feature fusion and residual connection: The value matrix is weighted and summed using the attention weights to obtain a context-aware feature representation. The output projection layer maps the fused features back to the original dimension. To prevent information loss, the original input is added to the attention output, and then layer normalization is used to stabilize gradient propagation. This process enables the model to autonomously identify key feature combinations, such as the synergistic effect between soil pH and ammonium nitrogen concentration.
[0098] Specifically, in the model initialization stage, a type mapping table is established according to the feature properties: Static indicators (such as soil type) are marked as T0, and dynamic indicators (such as temperature change rate) are marked as T1. Each input feature enters the network with a type label.
[0099] Preferably, two sets of bias parameters are set in the activation function of the hidden layer: b fixed (fixed bias) and b var (variable bias). During forward propagation, the corresponding bias term is selected and enabled according to the feature type label. The indicator function I(·) is essentially a conditional judgment gate. When the feature belongs to T0, b fixed participates in the calculation; when it belongs to T1, b var becomes effective. This mechanism enables the network to establish a stable representation for static features and maintain a flexible response to dynamic features.
[0100] Optionally, a smaller learning rate (1 / 10 of the model learning rate) is adopted for the fixed bias to prevent excessive adjustment of the parameters related to static features; an adaptive learning rate optimizer (such as Nadam) is used for the variable bias to quickly capture the changing rules of dynamic metrics. The gradients of the two types of biases are calculated separately during backpropagation to ensure training stability.
[0101] Preferably, a time decay Dropout layer is designed to retain the probability decay curve: the initial retention probability is set to 0.8, the decay coefficient is 0.01, and t represents the number of training epochs (time steps). After each epoch is completed, the global retention probability decreases by approximately 1%. This design allows more neurons to participate in learning in the initial stage and gradually sparsifies in the later stage to suppress overfitting.
[0102] Furthermore, dynamic mask generation, a new random mask matrix M is generated for each training t , the element values are 0 or 1, and the occurrence probability is determined by the current p(t). In the first 100 epochs, p(t) slowly decreases from 0.8 to 0.29, and the network maintains a strong learning ability; after 300 epochs, p(t) approaches 0.04, forming a highly sparse connection pattern.
[0103] Specifically, during the inference stage processing, the Dropout layer is turned off during testing, but the average activation intensity of each layer of neurons in the training stage is retained. By recording the typical mask patterns at the end of training, the weights are scaled and compensated during inference to maintain the consistency of the output magnitude.
[0104] Preferably, the Huber loss function is adopted as the loss function. This function adaptively switches between the mean squared error (MSE) and the mean absolute error (MAE). The threshold is set to 1.0. When the prediction deviation is less than the threshold, MSE is used to promote convergence, and when it is greater than the threshold, it switches to MAE to enhance robustness.
[0105] Optionally, the smoothed loss value is monitored on the validation set: calculate the moving average of the losses in the last 10 epochs, and early stopping is triggered when the average loss does not decrease for 3 consecutive times. To prevent misjudgment due to occasional fluctuations, the patience coefficient is set to 15 epochs, allowing the loss to fluctuate within the threshold range.
[0106] Specifically, the Nadam optimizer (Adam + NAG momentum) is used, and the initial learning rate is 3e -4 , combined with a triangular cyclic learning rate schedule: it periodically changes between 1e -4 and 3e -4 every 5 epochs to promote jumping out of local optima. The batch size is set to 64, taking into account both memory efficiency and the stability of gradient estimation.
[0107] Further, after completing the initial training, perturbation analysis is performed on the validation set samples. Each feature is perturbed by ±10% in turn, and the change amplitude of the model output is recorded. The SHAP value is calculated through Monte Carlo sampling to quantify the marginal contribution of each feature to the prediction result, and the SHAP value is mapped to a percentage system for reference Figure 5 , where NH3 represents ammonia concentration, soil_pH represents soil pH value, soil_NH4 represents soil ammonium nitrogen concentration; soil_NO3 represents soil nitrate nitrogen concentration, Gleyi-Stagnic Anthrosol represents gleyic-stagnic anthrosol, soil_humidity represents soil humidity, incubation_days represents the number of measurement days; N_application represents nitrogen fertilizer application level; Yellow-brown soil represents yellow-brown soil; Forest represents forest; Vineyard represents grape-growing farmland.
[0108] The environmental factors include temperature, humidity, etc., and the soil factors include soil pH value, ammonium nitrogen concentration, nitrate nitrogen concentration, soil type, etc.; missing value processing, standardization or normalization operations are performed on the data.
[0109] Furthermore, the threshold screening strategy sets the absolute SHAP value threshold to 60% of the maximum value (which can be adjusted according to the data scale), and retains the top 60% of the features in terms of contribution. Bundle screening is performed on strongly correlated feature groups (such as temperature and temperature change rate) to avoid misdeleting associated features. The screened feature set needs to satisfy both: the SHAP value of a single feature is greater than the threshold, and at least one representative item is retained within the feature group.
[0110] Optionally, three test areas with different climate types are selected outside the target area, and transfer learning is used for fine-tuning: the weights of the first three layers of the network are frozen, and only the last fully connected layer is trained. The fine-tuning period is limited within 20 epochs to verify the environmental adaptability of the model.
[0111] Preferably, for online prediction deployment: the final model is converted into the TensorRT format and deployed on edge computing devices. A two-level caching mechanism is designed: real-time data is first stored in a circular buffer, and model inference is triggered every time 1 day's worth of data is filled. The output result is uploaded to the cloud platform through a LoRa wireless module to achieve unattended prediction in the field.
[0112] Furthermore, the training process is visually monitored. The self-attention matrix of the first layer is regularly exported, and a feature correlation heatmap is plotted. By analyzing the evolution trend of the attention weights, it is verified whether the model has learned the expected physical and chemical relationships; the activation value distributions of each hidden layer are recorded to monitor whether there is gradient vanishing / explosion. Ideally, the values after ReLU activation should show a right-skewed distribution; time series analysis is performed on the SHAP values to identify the periodic changes in the contribution degrees of key features.
[0113] The beneficial effects of the present invention are as follows:
[0114] The present invention optimizes the candidate index set through the SHAP data of the index, realizes the reduction of model parameters, and improves the prediction efficiency of the model; by improving the model with Keras Sequential, it realizes the fusion of the correlations between data and different training strategies for different types of indexes, solves the overfitting risk in the model training process, and improves the prediction accuracy of the model for natural abundance.
[0115] The various embodiments in this specification are described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. The same or similar parts among the embodiments can be referred to each other.
[0116] Specific examples are used in this article to elaborate on the principles and implementation manners of the present invention. The descriptions of the above embodiments are only used to help understand the method of the present invention and its core idea; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present invention.
Claims
1. A prediction method for the natural abundance of ammonia volatilization nitrogen isotope based on environmental and soil factors, characterized in that, Including: Determine a candidate index set, and perform data collection and data preprocessing in the target research area according to the candidate index set to obtain a candidate index data set; Collect the natural abundance value of ammonia volatilization nitrogen isotope in the farmland of the target research area by using the isotope mass spectrometry analysis method to obtain a natural abundance data set; Construct a Keras Sequential model, insert a feature-level self-attention module between the input layer and the first hidden layer of the Keras Sequential model, add a time decay factor before the Dropout layer of the Keras Sequential model, and replace the original bias in the activation function of the hidden layer of the Keras Sequential model with a fixed bias and a variable bias according to the fixed quantity and variable quantity in the candidate index set, and judge the enabling of the fixed bias and the variable bias through an indicator function to obtain an improved Keras Sequential model; Construct a loss function, and iteratively train the improved Keras Sequential model by using the loss function and the early stopping method with the candidate index data set and the natural abundance data set to obtain an original natural abundance prediction model; Use a DeepSHAP interpreter to calculate the shap value of each candidate index in the original natural abundance prediction model to obtain a candidate index shap data set; Filter the candidate index shap data set by using a preset shap threshold to obtain the updated candidate index set; Train the improved Keras Sequential model by using the candidate index data set, the natural abundance data set and the data corresponding to the candidate index set to obtain an index-optimized natural abundance prediction model; Calculate the error degree of the index-optimized natural abundance prediction model with respect to the original natural abundance prediction model by using a preset error function to obtain an error ratio; When the error ratio is less than a preset error threshold, determine the index-optimized natural abundance prediction model as the nitrogen isotope natural abundance prediction model; Use the nitrogen isotope natural abundance prediction model to predict the target farmland area to obtain a prediction result.
2. The prediction method of the natural abundance of ammonia-volatilized nitrogen isotopes based on environmental and soil factors according to claim 1, wherein The iterative training process of the Keras Sequential model includes: Convert the candidate index data set into a tensor form to obtain input data; Perform linear transformation, attention weight calculation, weighted fusion, residual connection and normalization processing on the input data by using the feature-level self-attention module to obtain an attention output; Input the attention output into the hidden layer for abstraction processing to obtain a feature tensor; Perform dynamic retention and random masking processing on the feature tensor by using the Dropout layer to obtain a sparsified tensor; Successively calculate the sparsified tensor by using the remaining hidden layer and the Dropout layer of the Keras Sequential model to obtain a feature extraction result; Input the feature extraction result into the fully connected layer for ammonia volatilization nitrogen isotope abundance value mapping to obtain an iterative prediction result.
3. A prediction method for the natural abundance of ammonia volatilization nitrogen isotope based on environmental and soil factors according to claim 1, characterized in that The expression of the activation function is: where f(x i ) is the output of the activation function; ReLU(·) is the activation function; w i is the weight parameter; x i is the eigenvalue; b fixed is the fixed bias; b var is the variable bias; I(·) is the indicator function; is the index type corresponding to the eigenvalue; T0 represents the fixed index; T1 represents the variable index.
4. A prediction method for the natural abundance of ammonia volatilization nitrogen isotope based on environmental and soil factors according to claim 1, characterized in that, The expression of the error function is: where E is the error ratio; is the prediction result of the original natural abundance prediction model; is the prediction result of the index-optimized natural abundance prediction model; y n is the measured value of natural abundance.
5. A prediction method for the natural abundance of ammonia volatilization nitrogen isotopes based on environmental and soil factors according to claim 1, characterized in that, The determined candidate index set includes: air temperature, atmospheric humidity, soil pH value, soil ammonium nitrogen concentration, soil nitrate nitrogen concentration, soil type, land use type, fertilization type, nitrogen fertilizer application rate, temperature change rate, humidity change rate, ammonium nitrogen concentration change rate, nitrate nitrogen concentration change rate, precipitation, wind speed, light intensity, soil humidity, and soil conductivity.
6. The prediction method of the natural abundance of ammonia volatilization nitrogen isotope based on environmental and soil factors according to claim 1, characterized in that, The range of the shap threshold is from 0.1 to 0.
2.
7. A prediction method for the natural abundance of ammonia volatilization nitrogen isotopes based on environmental and soil factors according to claim 1, characterized in that, The data preprocessing includes: outlier removal, linear interpolation, mode filling, and data standardization.
8. A prediction method for the natural abundance of ammonia volatilization nitrogen isotopes based on environmental and soil factors according to claim 2, characterized in that, The expression of the feature-level self-attention module is: Among them, Output is the attention output; LayerNorm(·) represents the layer normalization operation; Softmax(·) represents the normalized exponential function; X is the input data matrix; W Q , W K , W V are the query matrix, the key matrix, and the value matrix respectively; W O is the output projection weight matrix; d is the scaling factor.
9. A method for predicting the natural abundance of ammonia volatilization nitrogen isotope based on environmental and soil factors according to claim 2, characterized in that, The expression of the Dropout layer is: where Dropout(X t , t) is the output of the Dropout layer; X t is the input tensor; M t is a random mask matrix; p(t) is the retention probability; ⊙ represents the element-wise multiplication operation.
10. A method for predicting the natural abundance of ammonia volatilization nitrogen isotope based on environmental and soil factors according to claim 5, characterized in that The acquisition devices for the candidate index set include: glass electrode pH sensor, digital ammonia and nitrate nitrogen measurement sensor, resistive soil humidity sensor, and conductivity sensor.
Citation Information
Patent Citations
Geological disaster risk assessment method and system fusing random forest and attention
CN116167617A
Distortion correction face recognition large-angle recognition algorithm
CN118522062A