Soil organic matter content prediction method based on two-channel model

By using a 2C-Net model with a dual-channel model in soil organic matter content prediction, using the feature extraction and fusion of spectral data and climatic topographic data, the problem that existing methods fail to make full use of the connection between multiple bands of spectral data is solved, and higher prediction accuracy is achieved.

CN119989442AActive Publication Date: 2025-05-13HEILONGJIANG UNIV
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510059492.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-15
Publication Date
2025-05-13
Estimated Expiration
2045-01-15

AI Technical Summary

Technical Problem

Existing soil organic matter content prediction methods fail to make full use of the potential connections between multiple bands in the spectral data, resulting in insufficient prediction accuracy.

Method used

A method for predicting soil organic matter content based on a dual-channel model is proposed. The spectral data and climatic topographic data are extracted and fused through the 2C-Net model. The spectral data and climatic topographic data are processed respectively by using the temporal feature extraction channel and the spatial feature extraction channel, and the fusion prediction is performed through the prediction head.

Benefits of technology

By fully utilizing the correlation between multiple bands in spectral data, the accuracy of soil organic matter content prediction is significantly improved, surpassing the performance of existing mainstream machine learning and deep learning models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure QLYQS_2
    Figure QLYQS_2
  • Figure QLYQS_3
    Figure QLYQS_3
  • Figure BDA0005242428960000053
    Figure BDA0005242428960000053
Patent Text Reader

Abstract

The invention discloses a soil organic matter content prediction method based on a two-channel model, and relates to the field of soil organic matter mapping based on a deep learning model. The invention aims to solve the problem that the existing modeling method in the field of soil organic matters cannot fully utilize the potential relation among a plurality of wavebands in spectral data. The method comprises the following steps: acquiring spectral data and climate and terrain environment covariant data of a region to be predicted, and inputting the two types of data into a trained 2C-Net model to obtain a soil organic matter content prediction result; the 2C-Net model comprises a time feature extraction channel, a spatial feature extraction channel and a prediction head; the time feature extraction channel is used for capturing features between each wave band and multiple wave bands in the spectral data in a time dimension to obtain time channel output features; the spatial feature extraction channel is used for carrying out spatial dimension modeling on climate and terrain environment covariant data to obtain spatial channel output features; and the prediction head fuses the time channel output features and the space channel output features to predict the soil organic matter content. The method is used for predicting the soil organic matter content.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of soil organic matter mapping based on deep learning models, and in particular to a soil organic matter content prediction method based on a dual-channel model. Background Art

[0002] Soil organic matter content prediction is one of the important research directions of digital soil mapping, and its main goal is to estimate the soil organic matter content within the target area. In the past decade, classic soil organic matter content prediction methods based on machine learning and deep learning have made significant progress. However, the accuracy of traditional soil organic matter content prediction methods based on machine learning is not high, so how to improve the accuracy of soil organic matter content prediction has become a research focus in this field.

[0003] In 2001, Breiman first proposed a machine learning algorithm based on decision trees, named Random Forests (RF). Subsequently, Grimm et al. first applied RF to the field of digital soil mapping. Due to its ease of use and higher prediction accuracy than other methods, RF has become one of the most popular models in the field of digital soil mapping in the past.

[0004] With the rise of deep learning in recent years, researchers have gradually noticed that introducing deep learning into digital soil mapping tasks can achieve higher prediction accuracy. Padarian et al. introduced CNN (Convolutional Neural Network) into the soil organic carbon estimation task, and improved the mapping accuracy with the help of the neighborhood pixel information of the covariate, and provided an effective framework for the subsequent digital soil mapping task. Wang et al. introduced LSTM into the soil cover classification task and combined it with spectral data for time series modeling, successfully surpassing the accuracy of mainstream machine learning models such as RF. Xu et al. used 3D-CNN (Three-Dimension CNN) to model point cloud data and spectral data, taking into account the three-dimensional spatial characteristics, and improved the prediction accuracy of the model compared with CNN. Zhang et al. combined CNN with LSTM to model environmental covariates and spectral data respectively, and achieved good results in the soil organic carbon mapping task. Although the digital soil mapping task based on deep learning has achieved remarkable results, the existing research work represented by the above research only models multiple bands within the spectral data separately, and does not fully utilize the correlation between multiple bands within the spectral data, resulting in a lot of room for improvement in the prediction accuracy of the model. Summary of the invention

[0005] The purpose of the present invention is to solve the problem that the existing modeling methods in the field of soil organic matter mapping fail to fully utilize the potential connections between multiple bands in spectral data, and propose a soil organic matter content prediction method based on a dual-channel model.

[0006] A method for predicting soil organic matter content based on a dual-channel model. The specific process is as follows:

[0007] Obtain spectral data, climate and terrain data of the area to be predicted, input the above two types of data into the trained 2C-Net model, and obtain the prediction results of soil organic matter content (for simplicity, the "2C-Net model" below is equivalent to the "dual-channel model");

[0008] The trained 2C-Net model is obtained in the following way:

[0009] Step S1: Divide the spectral data, climate and terrain data of the pixel unit at the location of the sample into a training set and a test set according to a data ratio of 7:3;

[0010] Step S2: train the 2C-Net model using the training set and test it using the test set to obtain a trained 2C-Net model;

[0011] The 2C-Net model includes: a temporal feature extraction channel, a spatial feature extraction channel, and a prediction head;

[0012] The time feature extraction channel is used to capture the characteristics of each band in the spectral data and the characteristics between multiple bands in the time dimension, and obtain the output feature y of the time feature extraction channel t , where t represents the time dimension;

[0013] The spatial feature extraction channel is used to model the climate and terrain data in spatial dimensions, and obtain the spatial feature extraction channel output feature y s , where s represents the spatial dimension;

[0014] The prediction head outputs the time channel feature y t And the spatial channel output feature y s Fusion, to predict soil organic matter content.

[0015] Furthermore, the temporal feature extraction channel includes: an encoder, a MFFM (Multispectral Feature Fusion Module) decoder, and an output layer;

[0016] The encoder inputs the encoded vector set containing spectral information and coordinate information into the decoder;

[0017] The MFFM decoder uses the header retrieval module to inject the output of the encoder into the retrieval header of the MFFM decoder, and then uses the header fusion module to fuse the features of the retrieval header, and then sends the fused multi-sequence features to the output layer;

[0018] The output layer further enhances the expression capability of the model by performing linear and nonlinear transformations on the multi-sequence features decoded by the MFFM decoder.

[0019] Furthermore, the encoder includes: a value embedding module VE (Value Embedding) and a coordinate embedding module CE (Coordinate Embedding); wherein the value embedding module VE is used to map the spectral data into a vector set; the coordinate embedding module CE is used to map the longitude and latitude information into a vector and add it to the vector set; the working process of the value embedding module VE and the coordinate embedding module CE is specifically as follows:

[0020] Step 1: The value embedding module VE uses a linear layer and an activation function to map the spectral data X into a feature vector set x. i , specifically:

[0021] x i =ReLU(Linear(X))

[0022] Among them, i∈[1,N], N represents the number of samples trained simultaneously;

[0023] Step 2: The coordinate embedding module CE uses a linear layer and an activation function to map the longitude and latitude information P into a vector form, and then uses an addition operation to embed it with the vector set x processed by the value embedding module VE. i Fusion, become w, specifically:

[0024] w=x i +ReLU(Linear(P))

[0025] Among them, i∈[1,N], N represents the number of samples trained simultaneously.

[0026] Furthermore, the MFFM decoder includes: a header initialization module, a header retrieval module, and a header fusion module; the MFFM decoder uses the header initialization module to obtain an initialized retrieval header, and then uses the header retrieval module to inject the output of the encoder into the retrieval header of the MFFM decoder, and then uses the header fusion module to perform feature fusion on the retrieval header, and finally sends the fused multi-sequence features to the output layer.

[0027] Furthermore, the header initialization module obtains the initialized retrieval header {h i,1,h i,2 ,…,h i,B}, specifically:

[0028] {h i,1 ,h i,2 ,…,h i,B}=Rearrange(ReLU(Linear(P)))

[0029] Where i∈[1,N], N represents the number of samples trained simultaneously, B represents the number of retrieval heads, and B is set to be the same as the number of bands of the spectral data.

[0030] Furthermore, the head retrieval module includes: a cross attention module Cross Attention; the head retrieval module takes the retrieval head as the query Q, dimensionally aligns the output of the encoder with the retrieval head to obtain multi-sequence features, and uses the encoder information as the key K and value V to inject the encoder information into the retrieval head through cross attention.

[0031] Furthermore, the head fusion module includes: a router mechanism; the router mechanism integrates the two cross attention modules, collects the input vectors of I router mechanisms using R intermediate vectors, and distributes them to the output vectors of O router mechanisms, and the vector dimension D R >D I =D O , specifically:

[0032] Step 1: Input B retrieval headers (B=I) with encoder information obtained by the header retrieval module into the router mechanism as key K and value V; declare R intermediate vectors as query Q, and aggregate the retrieval header information containing B band information in the spectral data into R intermediate vectors through cross attention:

[0033] Vector R =CrossAttention(Vector R ,Vector I ,Vector I )

[0034] Step 2: Use the R intermediate vectors obtained in Step 1 as the key K and value V, and the B search heads as the query Q. Use cross attention to redistribute the R intermediate vector information that integrates the B search head information to the B search heads:

[0035] Vector I =CrossAttention(Vector I ,Vector R ,Vector R )

[0036] Furthermore, the output layer uses a layer normalization operation to ensure that the distribution of the intermediate features in the vector space remains consistent, uses two linear layers and an activation function to perform linear and nonlinear transformations on the intermediate features, and uses the Dropout operation and residual connection to reduce the risk of overfitting and gradient disappearance in the neural network; then the output of the output layer is reshaped and then sent to the fully connected layer to obtain the abstract feature y output of the temporal feature extraction channel t , specifically:

[0037] y t =FC(Reshape(OutputLayer(u)))

[0038] Among them, u represents the intermediate feature of the input to the output layer, Reshape represents the reshaping operation, and FC represents the fully connected layer.

[0039] Furthermore, the spatial feature extraction channel includes: Diverse Convolutional Architecture (DCA) Block 1 and DCA Block 2; the DCA Block 1 uses convolution operations with different kernel sizes to preliminarily extract pixel features within a 5×5 range near the sampling point, and then uses the maximum pooling layer to preliminarily compress the important information in the pixel features; the DCA Block 2 uses convolution operations with different kernel sizes to further extract compressed pixel features, and then uses the maximum pooling layer to further compress the important information in the pixel features; the spatial feature extraction channel uses the Diverse Convolutional Architecture (DCA) to model climate and terrain data, specifically:

[0040] Step 1. Obtain the pixel data within the 5×5 pixel unit range around the sampling point in the climate and terrain data as M "images", each "image" is 5×5 in size, where M is the number of climate and terrain data; stack the M "images" of each sample and input them into DCA Block 1 for feature extraction, where DCA Block 1 consists of 4 convolutional layers and 1 maximum pooling layer. The kernel sizes of the 4 convolutional layers are 3×3, 3×3, 4×4, and 1×1 respectively. Each convolutional layer will be followed by a ReLU activation function, and the kernel size of the maximum pooling layer is 2×2;

[0041] Step 2: Send the result of Step 1 to DCA Block 2 for further feature extraction. DCABlock2 consists of 2 convolutional layers and 1 maximum pooling layer, with kernel sizes of 2×2 and 1×1 respectively. Each convolutional layer is followed by a ReLU activation function, and the kernel size of the maximum pooling layer is 2×2.

[0042] Step 3: Reshape the output of Step 2 and then send it to the fully connected layer to obtain the abstract feature y output by the spatial feature extraction channel s , specifically:

[0043] y s =FC(Reshape(v))

[0044] Among them, v represents the output feature of Step2, Reshape represents the reshaping operation, and FC represents the fully connected layer.

[0045] Furthermore, the prediction head extracts the output y of the temporal feature extraction channel t , and the output y of the spatial feature extraction channel s Concatenate (Concat), then input the concatenated features into the fully connected layer (FC), and then use the ReLU activation function to obtain the final prediction result Specifically:

[0046]

[0047] Furthermore, the optimization target of the 2C-Net model is calculated by the Huber loss function:

[0048]

[0049] in, is the loss value, y is the actual result, is the final prediction result, δ is a non-negative constant used as a conditional threshold.

[0050] The beneficial effects of the present invention are:

[0051] The present invention proposes a soil organic matter content prediction method based on a dual-channel model. The method is based on a 2C-Net model. In the process of soil organic matter content mapping, the time feature extraction channel is first used to extract features and interactively fuse information from multiple bands of spectral data to achieve global modeling of spectral information in the time dimension; at the same time, the spatial feature extraction channel is used to model climate and terrain data; finally, the output features of the two channels are fused for prediction, and the Huber loss function is used to optimize the model's perception of abnormal values ​​in the soil organic matter content mapping task to calculate the model loss, thereby obtaining the optimization target of the next iteration of the model.

[0052] The temporal feature extraction channel of the 2C-Net model proposed in the present invention uses an encoder to independently model multiple bands of spectral data in the temporal dimension, and then uses an MFFM decoder to perform information interactive fusion on them, thereby achieving the purpose of global modeling of spectral data in the temporal dimension.

[0053] The spatial feature extraction channel of the 2C-Net model proposed in the present invention can use the multivariate convolutional architecture DCA to model climate and terrain data in the spatial dimension, where the multivariate convolutional architecture DCA includes two blocks, which can alternately use 1×1 convolution kernels and other convolution kernels to enrich and abstract features. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] Figure 1 A model training framework diagram of a soil organic matter content prediction method based on a dual-channel model according to an embodiment of the present invention;

[0055] Figure 2 The overall framework diagram of the model of the soil organic matter content prediction method based on the dual-channel model provided by the present invention;

[0056] Figure 3 A schematic diagram of the structure of a router mechanism in an embodiment of the present invention;

[0057] Figure 4 A model reasoning framework diagram of a soil organic matter content prediction method based on a dual-channel model according to an embodiment of the present invention;

[0058] Figure 5 The predicted scatter plots for the comparison models are shown in Figure 2.

[0059] Figure 6 This is the mapping result diagram for the comparison model; DETAILED DESCRIPTION

[0060] Figure 1 FIG. 4 is a flow chart of a model training method for predicting soil organic matter content based on a dual-channel model according to an embodiment of the present invention. Figure 1As shown, the model training method of the soil organic matter content prediction method based on the dual-channel model of the embodiment of the present invention includes: step S110, inputting the spectral data in the training set, the training set comes from the training part in the data set, the data set is composed of spectral data, climate and terrain data and measured data of soil organic matter content at the locations of 576 sampling points, 70% of the data is used for training by random division, and the remaining 30% is used for testing; step S120, inputting the climate and terrain data in the training set, this step is performed simultaneously with step S110; step S130, standardizing the spectral data; step S140, standardizing the climate The spectral data is standardized with the terrain data; in step S150, the time feature extraction channel is used to extract features from the spectral data to obtain the abstract features of the spectral data in the time dimension; in step S160, the spatial feature extraction channel is used to extract features from the climate and terrain data to obtain the abstract features of the climate and terrain data in the spatial dimension; in step S170, the information extracted from the two channels is combined, and prediction is performed using the prediction head to obtain the predicted value of the model after one iteration; in step S180, the parameters of the model are updated using the loss function and the gradient optimizer until the model reaches the best fitting state and the optimal parameters of the overall model are obtained.

[0061] In the field of soil organic matter content mapping, it is usually necessary to model and analyze the soil organic matter content from both the time dimension and the space dimension. The change value of spectral reflectance over a period of time can reflect the physical and chemical properties of the soil from the time dimension, thereby helping to indirectly predict the soil organic matter content. Climate and topographic factors can reflect the impact of humans and nature on the soil from the spatial dimension, thereby establishing the data distribution of soil organic matter content from the spatial dimension. The reflectance values ​​of the soil in different spectral bands of spectral data over a period of time are regarded as multiple time series, and the time feature extraction channel of the model is used to model them, and combined with the spatial abstract features obtained through the spatial feature extraction channel in the model, so as to generate the model's prediction results for the soil organic matter content. The present invention focuses on the correlation between the reflectances of different spectral bands, so that the model can make full use of the information of different spectral bands of spectral data and perform global modeling on them, thereby enhancing the prediction accuracy of the model.

[0062] Figure 2 The overall framework diagram of the model of the soil organic matter content prediction method based on the dual-channel model provided by the present invention. Figure 2 As shown in the figure, the 2C-Net model mainly consists of three parts: temporal feature extraction channel, spatial feature extraction channel, and prediction head. Now we will explain the model prediction process in detail in conjunction with the overall framework diagram of the model:

[0063] The time extraction channel includes: an encoder, an MFFM decoder, and an output layer. The different bands of spectral data are stacked and sent to the time feature extraction channel, and standardized spectral data is obtained after preprocessing. The standardized spectral data is sent to the encoder for encoding. The encoder includes a value embedding module and a coordinate embedding module. The value embedding module includes a linear layer and an activation function. Through the value embedding module, the standardized spectral data will be mapped to a four-dimensional vector space. The dimensions of the four-dimensional vector space are batch size, number of spectral data bands, spectral data time step, and hidden dimension. The coordinate embedding module includes a linear layer, an activation function, and an addition operation. Through the coordinate embedding module, the longitude and latitude information will also be mapped to a four-dimensional vector space and added to the spectral data mapped to the four-dimensional vector space to obtain standardized spectral data with position encoding. The standardized spectral data with position encoding is then sent to the decoder for decoding. The decoder includes a header initialization module, a header retrieval module, and a header fusion module. The head initialization module includes a linear layer, an activation function and a rearrangement operation. Through the head initialization module, the longitude and latitude information will be mapped again to a new four-dimensional vector space, and then rearranged into a three-dimensional vector with the same number of bands as the spectral data. The dimensions of the three-dimensional vector are batch size × number of bands, spectral data time step, and hidden dimension. After the head initialization module, retrieval heads with the same number of bands as the spectral data are obtained. Since these retrieval heads and the standardized spectral data with position encoding in the encoder are all or partially mapped by longitude and latitude, each retrieval head can retrieve a feature vector of a certain band with similarity. Subsequently, the retrieval head and encoder information are sent to the head retrieval module. Through the head retrieval module, each band information of the spectral data will be independently injected into the corresponding retrieval head. Subsequently, the retrieval header with injected band information is sent to the head fusion module, such as Figure 3 As shown in the figure, through the router mechanism in the header fusion module, the information in each retrieval header will be interactively fused with each other, so as to achieve the purpose of global modeling of different bands of spectral data.

[0064] The router mechanism includes: two cross-attention mechanisms. First, the intermediate features of the model are input into the router mechanism and sent to the first cross-attention mechanism as key K and value V; temporarily declare an intermediate vector with a dimension greater than the router mechanism input, and send it to the first cross-attention mechanism as query Q. After the first cross-attention mechanism, the input information of the router mechanism is distributed to the intermediate vector for storage. Then the intermediate vector is used as the key K and value V, and the input vector of the router mechanism is sent to the second cross-attention mechanism as Q. After the second cross-attention mechanism, the information stored in the intermediate vector of the router mechanism is distributed to the output vector of the router mechanism. The dimensions D of the three vectors of the router mechanism R>D I =D O , where D R The intermediate vector dimension of the router mechanism, D I and D O Respectively represent the input and output vector dimensions of the router mechanism. Through the router mechanism, the intermediate features of the model input to the router mechanism undergo sufficient information interaction and fusion, achieving global modeling of the intermediate features of the model.

[0065] The output of the decoder is then sent to the output layer for linear and nonlinear mapping. The output layer includes a layer normalization operation, two linear layers, an activation function, a Dropout operation, and a residual connection. Finally, the output of the output layer is reshaped and sent to the fully connected layer to obtain the abstract features of the temporal feature extraction channel output.

[0066] The spatial feature extraction channel includes two DCABlocks. After the climate and terrain data are stacked and input into the spatial feature extraction channel, the unit pixel where the sampling point is located is first expanded outward to a 5×5 neighborhood to form a series of small "images", which are then stacked and standardized. The standardized 5×5 small "images" are sent to the DCA to perform feature extraction using convolution operations with different kernel sizes. The DCA includes two blocks, and its basic parameters are shown in Table 1.

[0067] Table 1 Composition structure and key parameters of DCA

[0068]

[0069] After the convolution and pooling operation of the first block, the original small "image" is "reduced" to 3×3, and after the convolution and pooling operation of the second block, it is "reduced" to 1×1. That is, after DCA, the information within the neighborhood range of the original 5×5 unit pixel is abstracted from the original 5×5 neighborhood to the 1×1 unit pixel size. This operation can consider the soil pixels in a small range near the sampling point. Finally, the abstracted 1×1 features are reshaped and sent to the fully connected layer to obtain the abstract features output by the spatial feature extraction channel.

[0070] The prediction head includes: a connection operation and a fully connected layer. The connection operation refers to concatenating vectors according to the last dimension of the vector at this time, such as Figure 2 As shown, the concatenated vector is sent to the fully connected layer to obtain the final output of the model.

[0071] Figure 4 FIG. 4 is a model reasoning framework diagram of a soil organic matter content prediction method based on a dual-channel model according to an embodiment of the present invention. Figure 4As shown, the model reasoning method of the soil organic matter content prediction method based on the dual-channel model of the embodiment of the present invention includes: step S210, inputting the spectral data of the area to be predicted, the spectral data of the area to be predicted refers to the spectral data of all pixel units in the area to be predicted, and is no longer the spectral data of scattered pixels in the training process. Subsequently, similar to the training process, the spectral data of the area to be predicted is stacked and sent to the time feature extraction channel of the model; step S220, inputting the climate and terrain data of the area to be predicted, the climate and terrain data of the area to be predicted is similar to step S210, and refers to the climate and terrain data of all pixel units in the area to be predicted, and is no longer the climate and terrain data of scattered pixels in the training process. Subsequently, similar to the training process, the climate and terrain data of the area to be predicted are stacked and sent to the spatial feature extraction channel of the model; step S230, using the model with optimal parameters obtained by training to perform prediction; step S240, obtaining the prediction result, the prediction result is a TIFF raster file containing the prediction value of each pixel unit.

[0072] Embodiment: In order to verify the beneficial effects of the present invention, the following experiments were carried out:

[0073] The experimental study described in this embodiment was carried out on an NVIDIA RTX 2080ti GPU, using multispectral spectral data, climate and terrain data, and measured data on the organic matter content of soil samples. They were divided into training sets and test sets in a ratio of (7:3), where 70% of the data was used to form the training set and 30% of the data was used to form the test set for model training and reasoning.

[0074] This embodiment uses cloud-free Sentinel-2 multispectral image data obtained through the Google Earth Engine (GEE) platform as spectral data, and the data covers the period from April to November 2021. The image data provides multispectral remote sensing information, which can be used to analyze the ecological and environmental characteristics of the surface.

[0075] This embodiment uses 30-meter resolution digital elevation model (DEM) data obtained through the geospatial data cloud platform, temperature and precipitation data obtained through the "China Ecological Environment Database", and river network datum data extracted through SAGA terrain analysis software as climate and terrain data.

[0076] This embodiment uses the organic matter content data of 576 soil samples in the study area measured by laboratory, and uses it as the true value supervision model training and verification model.

[0077] This example conducts multiple experiments based on the above data set to verify the main indicators of the model, including root mean square error (RMSE), mean absolute error (MAE), mean square error (MSE), determination coefficient (R 2 ).

[0078] This example uses CNN-LSTM as the baseline method, selects the encoder / MFFM decoder architecture to model multiple bands in the spectral data in the time dimension, and DCA Blocks to model the climate and terrain data in the spatial dimension. The model is trained for 2000 epochs to iterate the optimal parameters, and the learning rate is reduced to half of the original after every 500 epochs. Table 2 compares the 2C-Net model with existing models.

[0079] Table 2 Comparison with existing mainstream machine learning models and deep learning models in the field of digital soil mapping on the same dataset

[0080]

[0081] The 2C-Net model uses CNN-LSTM as the baseline and is compared with the mainstream machine learning models and deep learning models used in the current digital soil mapping tasks. Among them, CNN-GRU is the implementation that replaces the LSTM in the baseline with the RNNs model GRU that also uses the gating mechanism for supplementary reference verification.

[0082] According to Table 2, among the mainstream machine learning modeling methods for digital soil mapping, Random Forest achieved the highest accuracy: RMSE was 0.955, MAE was 0.610, MSE was 0.912, and R 2 is 0.451; among the mainstream deep learning modeling methods for digital soil mapping, the baseline CNN-LSTM has an RMSE of 0.982, a MAE of 0.610, and a MSE of 0.912. 2 The RMSE, MAE, and MSE of the CNN-GRU model used as a reference and supplementary verification are 0.977, 0.678, and 0.955, respectively. 2 The 2C-Net proposed in this paper achieves the highest accuracy in all indicators: RMSE is 0.884, MAE is 0.581, MSE is 0.781, R 2It is 0.524, which significantly surpasses the mainstream modeling methods in current digital soil mapping, including the CNN-LSTM baseline method and Random Forest. Figure 5 The prediction scatter distribution results of RandomForest, CNN-LSTM, CNN-GRU and 2C-Net models are shown: the narrower the light and dark strips are, the better the prediction effect is; the m value represents the slope of the fitting, the closer it is to 1, the better the prediction effect is; the b value represents the intercept of the fitting, the closer it is to 0, the better the prediction effect is; Figure 5 It can be seen that the 2C-Net model is optimal under the above conditions. Figure 6 The mapping results between Random Forest, CNN-LSTM, CNN-GRU and 2C-Net models are shown. Figure 6 As can be seen from the figure, the mapping results of 2C-Net are more detailed and accurate than those of other models.

[0083] In order to verify the effectiveness of each key component of 2C-Net, the present invention conducted an ablation experiment on the 2C-Net model.

[0084] As shown in Table 3, three key components in the model are gradually removed to obtain different variants: MFFM decoder, multivariate convolutional architecture (DCA) and coordinate embedding module (CE). All experiments adopt the same implementation as the experiments in Table 2.

[0085] Table 3 Experimental results obtained by gradually dismantling key components of the model

[0086] Model RMSE↓ MAE↓ MSE↓ <![CDATA[R 2 ↑]]> 2C-Net without MFFM DCA CE 1.112 0.824 1.237 0.347 2C-Net without MFFM DCA 1.004 0.720 1.009 0.386 2C-Net without MFFM 0.996 0.677 0.992 0.402 2C-Net 0.884 0.581 0.781 0.524

[0087] By gradually introducing the MFFM decoder, the multivariate convolutional architecture (DCA) and the coordinate embedding module (CE), the present invention can significantly improve the prediction performance of the model. According to the experimental results in Table 3, with the introduction of these three key components one by one, the prediction ability of the model is gradually enhanced, thereby verifying the effectiveness and complementarity of each component in improving the performance of the model. Specifically, the MFFM decoder optimizes the information decoding process, the multivariate convolutional architecture (DCA) enhances the diversity of feature extraction, and the coordinate embedding module (CE) improves the model's ability to model spatial position relationships.

[0088] As shown in Table 4, the MSE loss function in the baseline is replaced by the Huber loss function. In addition, the MAE loss function commonly used in regression tasks is used as a reference for comparison. All experiments are implemented in the same way as the experiments in Table 2.

[0089] Table 4 Experimental results using different loss functions

[0090] Loss Function RMSE↓ MAE↓ MSE↓ <![CDATA[R 2 ↑]]> MAE Loss 0.925 0.615 0.856 0.478 MSE loss 0.922 0.635 0.851 0.482 Huber loss function 0.884 0.581 0.781 0.524

[0091] As can be seen from Table 4, the introduction of the Huber loss function can obtain the highest model accuracy, proving the effectiveness of choosing the Huber loss function as the loss function. The Huber loss function combines the advantages of MSE loss and MAE loss. Compared with MSE loss, the Huber loss function is more robust in the face of outliers; compared with MAE loss, the Huber loss function has a faster convergence speed and can learn quickly when the error is small, avoiding the learning efficiency problem caused by too small gradient.

[0092] Therefore, it can be seen that the temporal feature extraction channel of the 2C-Net model proposed in the present invention uses an encoder to independently model multiple bands of spectral data in the temporal dimension, and then uses the MFFM decoder to interactively fuse them, so as to achieve the purpose of global modeling of spectral data in the temporal dimension; at the same time, the spatial feature extraction channel of the 2C-Net model proposed in the present invention uses the multivariate convolutional architecture DCA to model climate and terrain data in the spatial dimension; finally, the 2C-Net model proposed in the present invention fuses the outputs of the temporal feature extraction channel and the spatial feature extraction channel, and uses the Huber loss function to enhance the model's perception of abnormal values ​​in the soil organic matter content mapping task; finally, after multiple model iterations, the optimal model parameters are obtained, and it is used to perform the organic matter content prediction task to obtain high-precision organic matter content mapping results.

Claims

1. A method for predicting soil organic matter content based on a dual-channel model, characterized in that: The specific process of the method is: Obtain the spectral data, climate and topographic environmental covariate data of the area to be predicted and input them into the trained dual-channel model (for simplicity, the "2C-Net model" below is equivalent to the "dual-channel model") to obtain the prediction results of soil organic matter content; The trained 2C-Net model is obtained in the following way: Step S1: Divide the pixel data (spectral data, climate and terrain data) of the sample location into a training set and a test set according to a data ratio of 7:3; Step S2: Use the training set and Huber loss function to train the 2C-Net model, and use the test set to test it to obtain the trained 2C-Net model; The 2C-Net model includes: a temporal feature extraction channel, a spatial feature extraction channel, and a prediction head; The time feature extraction channel is used to capture the characteristics of each band in the spectral data and the characteristics between multiple bands in the time dimension, and obtain the output feature y of the time feature extraction channel t , where t represents the time dimension; The spatial feature extraction channel is used to model the spatial dimension of the climate and terrain environment covariate data to obtain the output feature y of the spatial feature extraction channel s , where s represents the spatial dimension; The prediction head outputs the time channel feature y t And the spatial channel output feature y s Fusion, to predict soil organic matter content.

2. The method for predicting soil organic matter content based on a dual-channel model according to claim 1, characterized in that: The temporal feature extraction channel includes: an encoder, a MFFM (Multispectral Feature Fusion Module) decoder, and an output layer; The encoder inputs the encoded vector set containing spectral information and coordinate information into the decoder; The MFFM decoder uses the header retrieval module to inject the output of the encoder into the retrieval header of the MFFM decoder, and then uses the header fusion module to fuse the features of the retrieval header, and then sends the fused multi-sequence features to the output layer; The output layer further enhances the expression capability of the model by performing linear and nonlinear transformations on the multi-sequence features decoded by the MFFM decoder.

3. The method for predicting soil organic matter content based on a dual-channel model according to claim 2, characterized in that: The encoder includes: a value embedding module VE (Value Embedding) and a coordinate embedding module CE (Coordinate Embedding), and the specific structure and function are as follows: The value embedding module VE uses a linear layer and an activation function to map the spectral data X into a feature vector set x i , specifically: x i =ReLU(Linear(X)) Among them, i∈[1,N], N represents the number of samples trained simultaneously; The coordinate embedding module CE uses a linear layer and an activation function to map the latitude and longitude information P into a vector form, and then uses an addition operation to add it to the vector set x processed by the value embedding module VE. i Fusion, become w, specifically: w=x i +ReLU(Linear(P)) Among them, i∈[1,N], N represents the number of samples trained simultaneously.

4. The method for predicting soil organic matter content based on a dual-channel model according to claim 2, characterized in that: The MFFM decoder includes: a header initialization module, a header retrieval module, and a header fusion module. The specific structure and function are as follows: The head initialization module uses a linear layer, an activation function and a rearrangement operation to map the latitude and longitude information P of the sample into a set of initialized vectors {h i,1 ,h i,2 ,…,h i,B }, this set of vectors is called the search head, specifically: {h i,1 ,h i,2 ,…,h i,B }=Rearrange(ReLU(Linear(P))) Where i∈[1,N], N represents the number of samples trained simultaneously, B represents the number of retrieval heads, and B is set to be the same as the number of bands of the spectral data; The head retrieval module includes: a cross attention module Cross Attention, which can take the retrieval head as a query Q, then align the output of the encoder in claim 2 with the retrieval head to obtain a multi-sequence feature, and use it as a key K and a value V, and inject the information of the encoder into the retrieval head through cross attention; The head fusion module includes a router mechanism; the router mechanism integrates two cross attention modules, collects information of I vectors input into the router mechanism using R intermediate vectors, and distributes it to O output vectors of the router mechanism, where the vector dimension D R >D I =D O ; The working process of the head fusion module is specifically as follows: Step B1: Input B retrieval headers (B=I) with encoder information obtained by the header retrieval module into the router mechanism as key K and value V; declare R intermediate vectors as query Q, and aggregate the retrieval header information containing B band information in the spectral data into R intermediate vectors through cross attention, specifically: Vector R =CrossAttention(Vector R ,Vector I ,Vector I ) Step B2: Use the R intermediate vectors obtained in step B1 as the key K and value V, and the B search heads as the query Q. Use cross attention to redistribute the R intermediate vector information that fuses the B search head information to the B search heads. Specifically: Vector I =CrossAttention(Vector I ,Vector R ,Vector R )。 5. The method for predicting soil organic matter content based on a dual-channel model according to claim 2, characterized in that: The output layer uses a layer normalization operation to ensure that the distribution of intermediate features in the vector space remains consistent, uses two linear layers and an activation function to perform linear and nonlinear transformations on the intermediate features, and uses Dropout operation and residual connection to reduce the risk of overfitting and gradient disappearance in the neural network; then the output of the output layer is reshaped and sent to the fully connected layer to obtain the abstract feature y output of the temporal feature extraction channel t , specifically: y t =FC(Reshape(OutputLayer(u))) Among them, u represents the intermediate feature of the input to the output layer, Reshape represents the reshaping operation, and FC represents the fully connected layer.

6. The method for predicting soil organic matter content based on a dual-channel model according to claim 1, characterized in that: The spatial feature extraction channel includes: Diverse Convolutional Architecture (DCA) Block 1 and DCA Block 2; The DCA Block 1 uses convolution operations with different kernel sizes to preliminarily extract pixel features within a 5×5 range near the sampling point, and then uses a maximum pooling layer to preliminarily compress important information in the pixel features; The DCA Block 2 further extracts compressed pixel features using convolution operations with different kernel sizes, and then further compresses important information in the pixel features using a maximum pooling layer.

7. The method for predicting soil organic matter content based on a dual-channel model according to claim 6, characterized in that: The working process and structure information of DCA Block 1 are as follows: The pixel data within the 5×5 pixel unit range around the sampling point in the acquired climate and terrain data are taken as M "images", each of which is 5×5 in size, where M is the number of climate and terrain data; the M "images" of each sample are stacked and input into DCA Block 1, where the convolutional layers are used to gradually extract features from the pixel data, and the maximum pooling layer is used to compress the features. DCA Block 1 consists of 4 convolutional layers and 1 maximum pooling layer. The kernel sizes of the 4 convolutional layers are 3×3, 3×3, 4×4, and 1×1, respectively. Each convolutional layer is followed by a ReLU activation function, and the kernel size of the maximum pooling layer is 2×2.

8. The method for predicting soil organic matter content based on a dual-channel model according to claim 6, characterized in that: The working process and structure information of DCA Block 2 are as follows: The output of DCA Block 1 described in claim 7 is used as the input of DCA Block 2, and the convolution layer of DCA Block 2 is used to further extract features of the pixel data, and the maximum pooling layer is used to further compress the features, wherein DCA Block 2 is composed of 2 convolution layers and 1 maximum pooling layer, and the kernel sizes are 2×2 and 1×1 respectively, each convolution layer is followed by a ReLU activation function, and the kernel size of the maximum pooling layer is 2×2.

9. The method for predicting soil organic matter content based on a dual-channel model according to claim 8, characterized in that: The output of DCA Block 2 described in claim 8 is reshaped and then sent to the fully connected layer to obtain the abstract feature y output by the spatial feature extraction channel s , specifically: y s =FC(Reshape(v)) Among them, v represents the output feature of DCABlock 2 described in claim 8, Reshape represents a reshaping operation, and FC represents a fully connected layer.

10. A method for predicting soil organic matter content based on a dual-channel model according to claim 5, 9, characterized in that: The function of the prediction head is to extract the output y of the temporal feature extraction channel. t , and the output y of the spatial feature extraction channel s Concatenate (Concat) and input the concatenated features into the fully connected layer (FC), and then activate them using the ReLU activation function to obtain the final prediction result Specifically: The Huber loss function is then used to calculate the model loss value of the current iteration, specifically: in, is the loss value, y is the actual result, is the final prediction result, δ is a non-negative constant used as a conditional threshold; Huber function is used to calculate the prediction results The loss of the true value y is used to obtain the optimization target of the next iteration of the model, and then the model is iteratively optimized.

Citation Information

Patent Citations

  • Hyperspectral soil available boron content prediction method

    CN116578851A

  • Soil nutrient inversion method, electronic equipment and storage medium

    CN117688835A

  • Soil organic matter content determination method and system and electronic equipment

    CN117907244A

  • Corn yield prediction method and system based on deep learning attention mechanism

    CN118521008A

  • Soil humidity prediction method of attention coding and decoding LSTM (Long Short Term Memory) model based on physical process

    CN118940160A