A cultivated land soil ph estimation method based on multi-source remote sensing and spatial weighting
By using multi-source remote sensing data and spatial weighting technology, a weighted random forest model was constructed, which solved the spatial heterogeneity problem in soil pH estimation and achieved high-precision soil pH estimation and dynamic monitoring of arable land.
Patent Information
- Application Number
- CN202511440769.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-10
- Publication Date
- 2025-12-30
- Estimated Expiration
- 2045-10-10
AI Technical Summary
Existing technologies do not fully consider the spatial heterogeneity of soil in the estimation of pH in arable land, resulting in insufficient estimation accuracy of the model in areas with significant differences or uneven spatial distribution of samples.
By employing a spatial weighting method combining multi-source remote sensing data, a spatial weighted feature matrix of multi-source remote sensing environmental factor data is constructed to enhance the input of the random forest model and improve the model's ability to perceive and adapt to the spatial distribution of soil pH.
It improves the spatial continuity and prediction accuracy of farmland soil pH estimation, and is suitable for high-precision estimation and dynamic monitoring in multiple regions and at multiple scales.
Smart Images

Figure CN120908419B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of digital soil mapping and arable land soil environmental monitoring technology, specifically involving a method for estimating arable land soil pH based on multi-source remote sensing and spatial weighting. Background Technology
[0002] Soil pH is a crucial physicochemical parameter for assessing soil acidity and alkalinity, significantly influencing crop growth, soil nutrient availability, and agricultural management decisions. Obtaining high-precision, high-resolution spatial distribution information of soil pH is particularly critical in farmland management and precision agriculture. With the rapid development of remote sensing and geospatial information technologies, integrating multi-source remote sensing environmental factors with machine learning models to achieve spatial estimation of soil properties has become an important research direction in digital soil mapping. Existing studies primarily employ optical remote sensing, topographic, and meteorological factors to construct soil pH estimation models, with fewer systematically incorporating radar remote sensing data. Radar factors can indirectly reflect soil moisture, texture, and surface roughness, exhibiting a potential response to pH. Furthermore, commonly used modeling methods, including support vector machines, random forests, and gradient boosting trees, mainly focus on the statistical correlation between environmental factors and soil properties, failing to fully consider the spatial heterogeneity of soil. This is especially problematic in areas with significant differences in farmland soil or uneven spatial distribution of samples, where model estimation accuracy is easily affected. Summary of the Invention
[0003] The purpose of this invention is to address the aforementioned problems in the existing technology by providing a method for estimating farmland soil pH based on multi-source remote sensing and spatial weighting. This method can improve the model's ability to perceive the spatial distribution of soil pH and its adaptability to spatially heterogeneous regions, and is suitable for high-precision estimation and dynamic monitoring of farmland soil pH at the regional scale.
[0004] The above-mentioned objectives of the present invention are achieved by the following technical means:
[0005] A method for estimating soil pH in arable land based on multi-source remote sensing and spatial weighting includes the following steps:
[0006] Step 1: Set up multiple sampling points in the target area. Based on the global land cover dataset, collect soil samples at a set depth from each sampling point in the target area, determine the actual pH value of the soil samples collected from each sampling point, and record the latitude and longitude coordinates of each sampling point.
[0007] Step 2: Based on the latitude and longitude coordinates of each sampling point, extract the multi-source remote sensing environmental factor data of the corresponding location. The multi-source remote sensing environmental factor data includes the mean of multiple types of secondary feature data.
[0008] Step 3: Normalize the mean values of various secondary feature data of each sampling point extracted in Step 2, and construct the original feature matrix. The actual pH values of soil samples from each sampling point were used to construct a tag matrix. ;
[0009] Step 4: Construct a spatially weighted feature matrix based on the latitude and longitude coordinates between sampling points. and the original feature matrix Spatial weighted fusion is performed to obtain the enhanced feature matrix. ;
[0010] Step 5: Set the loss function to enhance the feature matrix. The random forest model is constructed into a weighted random forest model by taking the random forest model as input, and the weighted random forest model is trained to obtain the trained weighted random forest model.
[0011] Step 6: Set up sampling points in the target area to be tested, and execute steps 1 to 5 in sequence to obtain the trained weighted random forest model of the target area to be tested, and then use the enhanced feature matrix of the target area to obtain the model. Input the data to obtain the prediction matrix of soil pH values for the target area.
[0012] As mentioned above, in step 4, a spatially weighted eigenma matrix is constructed using a Gaussian kernel function. The Gaussian kernel function set in step 5 Candidate intervals are traversed with a set step size, and the reconstructed enhanced feature matrix is calculated at each step. A weighted random forest model is trained, using the root mean square error (RMSE) loss function. The minimum RMSE is used as the evaluation metric to determine the optimal model. After training, the model parameters are saved to obtain the trained weighted random forest model.
[0013] The parameter is the set range of influence of the adjustment space.
[0014] As described above, step 2 specifically includes the following steps:
[0015] Multi-source remote sensing environmental factor data includes multiple types of primary feature data, namely optical vegetation index, radar backscattering coefficient, meteorological factors, and topographic factors;
[0016] The optical vegetation index is calculated as follows: images of the target area within a set time range are selected, and cloud and cloud shadow interference are removed respectively. Then, atmospheric impedance vegetation index, soil-regulated vegetation index, normalized difference vegetation index, normalized difference water index, normalized difference red edge vegetation index, conversion vegetation index, normalized difference cloud index, difference vegetation index, enhanced vegetation index, and ratio vegetation index are calculated as two types of feature data.
[0017] The calculated vegetation indices are synthesized using pixel-level time mean values to obtain the mean values of various vegetation indices for each pixel within a set time period. The mean values of various vegetation indices for each sampling point are then extracted to form the optical vegetation index features of the corresponding sampling point.
[0018] The radar backscattering coefficient is calculated as follows: Select radar images, extract secondary feature data such as VV polarization, VH polarization, and the ratio of VV polarization to VH polarization within a set time period, after correcting all radar images, perform pixel-level time mean synthesis of VV polarization, VH polarization, and the ratio of VV polarization to VH polarization within the set time period to obtain the mean of VV polarization, VH polarization, and the ratio of VV polarization to VH polarization for each pixel within the target time period, and then extract the mean of VV polarization, the mean of VH polarization, and the mean of the ratio of VV polarization to VH polarization for the pixels where the latitude and longitude coordinates of the sampling point are located to form the radar backscattering coefficient of the corresponding sampling point;
[0019] The meteorological factors are calculated as follows: the original meteorological data of the target area for a set time period is obtained using a public climate dataset, and then resampled to a set resolution and aligned with the latitude and longitude coordinates of each sampling point to obtain the resampled meteorological data of the target area.
[0020] The annual temperature data, annual precipitation data, and annual soil moisture data of the target area are extracted from the resampled meteorological data, and the mean values are calculated to obtain the annual average temperature, annual average precipitation, and annual average soil moisture of the target area. The annual average temperature, annual average precipitation, and annual average soil moisture of each sampling point are extracted to form the meteorological factors of the corresponding sampling point.
[0021] The terrain factors are calculated as follows: Based on the digital elevation model, spatial interpolation is performed on the digital elevation model to a set resolution. In the cloud computing platform, the terrain analysis module extracts the elevation, slope, aspect, and terrain shadow index of each sampling point in the spatially interpolated digital elevation model to form the terrain factors of the corresponding sampling points.
[0022] The original feature matrix as described in step 3 above , The number of sampling points. The original feature matrix represents the number of types of secondary feature data included in the multi-source remote sensing environmental factor data. Each row of data is a feature data vector composed of the mean values of various secondary feature data of a sampling point;
[0023] The tag matrix , It is the set of real numbers.
[0024] As described above, step 4 specifically includes the following steps:
[0025] Step 4.1: Calculate the Euclidean distance between each sampling point and other sampling points based on the latitude and longitude coordinates of all sampling points, and construct a spatial distance matrix. ;
[0026] Step 4.2: Construct the spatial weight matrix using a Gaussian kernel function. And for the spatial weight matrix Normalization is performed to obtain the normalized spatial weight matrix. ;
[0027] Step 4.3: The normalized spatial weight matrix is obtained through matrix multiplication. With the original feature matrix Multiplying them together yields the spatially weighted eigenvalue matrix. , Spatial weighted feature matrix The first in The line is the first Spatial weighted feature vector of each sampling point.
[0028] As described above, spatial distance matrix Specifically, it is constructed in the following way:
[0029] Calculate the Euclidean distance between each sampling point and other sampling points based on the latitude and longitude coordinates of all sampling points. Construct a spatial distance vector for each sampling point according to its sampling point index. Then, construct a spatial distance matrix from all sampling point spatial distance vectors in sequence according to their sampling point indices. , Sampling points With sampling points The Euclidean distance between them and These are all the serial numbers of the sampling points. , Spatial distance matrix each The line is the first The Euclidean distance between each sampling point and other sampling points, when When the Euclidean distance is 0, the distance between the two sides is zero.
[0030] As described above, the spatial weight matrix Calculated based on the following formula:
[0031]
[0032] For the spatial weight matrix Normalization is performed on the spatial weight matrix. The spatial weights of each row are normalized so that the sum of the spatial weights of each sampling point is 1, based on the following formula:
[0033]
[0034] In the formula, Sampling points With sampling points Spatial weights between them.
[0035] The enhanced feature matrix as described in step 4 above To transform the original feature matrix With spatial weighted characteristic matrix Column-wise assembly:
[0036]
[0037] Enhanced feature matrix .
[0038] A computer device includes a memory and a processor, the memory storing a computer program, the processor executing the computer program to implement the steps of the method described above.
[0039] A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the method described above.
[0040] Compared with the prior art, the present invention has the following advantages:
[0041] The method of this invention comprehensively utilizes multi-source remote sensing environmental information, including optical, radar, topographic, and meteorological data, to improve the response sensitivity to changes in farmland soil pH; by introducing a spatial weighting mechanism, it enhances the model's adaptability to the spatial heterogeneity of farmland; and by strengthening the feature matrix... The weighted random forest model constructed by using it as input to the random forest model improves the spatial continuity and prediction accuracy of farmland soil pH estimation results. The method of this invention has good versatility and scalability, and is suitable for rapid estimation of farmland soil pH in multiple regions and at multiple scales. Attached Figure Description
[0042] Figure 1 This is a schematic flowchart of the method of the present invention;
[0043] Figure 2 This is a verification diagram of the soil pH inversion model for cultivated land in a certain area of Hubei Province obtained based on the RF model of the present invention;
[0044] Figure 3This is a verification diagram of the soil pH inversion model for cultivated land in a certain area of Hubei Province obtained based on the SWRF model of the present invention. Detailed Implementation
[0045] To facilitate understanding and implementation of the present invention by those skilled in the art, the present invention will be further described in detail below with reference to embodiments. The embodiments described herein are for illustration and explanation only and are not intended to limit the present invention.
[0046] Example 1:
[0047] like Figure 1 As shown, a method for estimating the pH of arable land soil based on multi-source remote sensing and spatial weighting includes the following steps:
[0048] This embodiment uses cultivated land in a certain area of Hubei Province as the object to estimate the pH of cultivated land soil;
[0049] Step 1: Set up multiple sampling points in the target area. Based on the 10-meter resolution WorldCover global land cover dataset, collect soil samples at a set depth from each sampling point in the target area, measure the pH value of the soil samples collected from each sampling point, and record the latitude and longitude coordinates of each sampling point.
[0050] In this embodiment, based on the 10-meter resolution WorldCover global land cover dataset released by the European Space Agency (ESA) in 2022, 526 typical farmland sampling points were set up within a certain region's farmland area. Soil samples from the surface layer (0.1–20 cm depth) of each sampling point were collected, and soil pH data were obtained through laboratory testing. The latitude and longitude coordinates of each sampling point were recorded simultaneously.
[0051] Step 2: Acquire multi-source remote sensing environmental factor data: Based on the Google Earth Engine (GEE) cloud computing platform, automatically extract multi-source remote sensing environmental factor data for each sampling point according to its latitude and longitude coordinates, and combine the mean values from different time series to ensure the representativeness and timeliness of the factors. This includes the following steps:
[0052] Multi-source remote sensing environmental factor data includes multiple types of primary feature data, namely optical vegetation index, radar backscattering coefficient, meteorological factors, and topographic factors;
[0053] Step 2.1: Extract optical vegetation indices. Optical vegetation indices include multiple types of secondary feature data, namely Atmospheric Impedance Vegetation Index (ARVI), Soil-Adjusted Vegetation Index (SAVI), Normalized Difference Vegetation Index (NDVI), Normalized Difference Water Index (NDWI), Normalized Difference Red Edge Vegetation Index (NDRE), Transformed Vegetation Index (TVI), Normalized Difference Cloud Index (NDCI), Difference Vegetation Index (DVI), Enhanced Vegetation Index (EVI), and Ratio Vegetation Index (RVI). The specific steps include:
[0054] Step 2.1.1: Based on multispectral remote sensing data from the Sentinel-2 (Level-2A) satellite (Sentinel-1 is an Earth observation satellite launched by the European Space Agency (ESA), mainly used to acquire various data on the Earth's surface; Level-2A data is atmospherically corrected and can be directly used for surface reflectance analysis), select all available images of the target area from March 1, 2024 to May 31, 2024, and after removing cloud and cloud shadow interference, calculate various vegetation indices of the optical vegetation index;
[0055] Step 2.1.2: Perform pixel-level time mean synthesis on the various vegetation indices calculated in Step 2.1.1 (i.e., perform time mean synthesis on all images of various vegetation indices of each pixel within a set time range (March 1, 2024 to May 31, 2024 in this embodiment) to obtain the mean of various vegetation indices of each pixel within the set time range, and extract the mean of various vegetation indices of each sampling point to form the optical vegetation index features of the corresponding sampling point.
[0056] Step 2.2: Extract the radar backscattering coefficient. The radar backscattering coefficient includes multiple types of secondary feature data, namely VV polarization (referring to the polarization direction of both the electromagnetic waves transmitted and received by the radar being vertical), VH polarization (the polarization direction of the electromagnetic waves transmitted by the radar is vertical, while the polarization direction of the received electromagnetic waves is horizontal), and the ratio of VV polarization to VH polarization (by comparing the VV and VH polarization signals, land cover classification and vegetation monitoring can be performed more effectively). Specifically, this includes the following steps:
[0057] Step 2.2.1: Select Sentinel-1 SAR radar imagery (Sentinel-1 is an Earth observation satellite launched by the European Space Agency (ESA); SAR stands for Synthetic Aperture Radar), and extract the VV polarization, VH polarization, and the ratio of VV polarization to VH polarization from March 1, 2024 to May 31, 2024 to reflect the surface structure characteristics and soil water content.
[0058] Step 2.2.2: After all radar images have been corrected (the correction process in this embodiment includes radiometric correction and geometric correction), the VV polarization, VH polarization, and the ratio of VV polarization to VH polarization in the target time period (March 1, 2024 to May 31, 2024 in this embodiment) are synthesized using pixel-level time averages. The average of VV polarization, VH polarization, and the ratio of VV polarization to VH polarization for each pixel in the target time period is obtained. Then, the average of VV polarization, VH polarization, and the ratio of VV polarization to VH polarization for the pixel where the latitude and longitude coordinates of the sampling point are located is extracted to form the radar backscattering coefficient of the corresponding sampling point.
[0059] Step 2.3: Extract meteorological factors, which include multiple types of secondary feature data, namely annual average temperature, annual average precipitation, and annual average soil moisture, specifically:
[0060] Step 2.3.1: Use the TERRACLIMATE public climate dataset (global land surface monthly climate and climate water balance dataset) to obtain the raw meteorological data of the target area for the whole year of 2024 (January 1 to December 31). Since the raw meteorological data is usually at a resolution of 1km or less, it is resampled to a resolution of 10 meters using bicubic interpolation and aligned with the latitude and longitude coordinates of each sampling point to obtain the resampled meteorological data of the target area.
[0061] Step 2.3.2: Extract the annual temperature data, annual precipitation data, and annual soil moisture data from the resampled meteorological data of the target area, and calculate the mean values to obtain the annual average temperature, annual average precipitation, and annual average soil moisture of the target area. Extract the annual average temperature, annual average precipitation, and annual average soil moisture of each sampling point to form the meteorological factors of the corresponding sampling point.
[0062] Step 2.4: Extract terrain factors. Terrain factors include multiple types of secondary feature data, namely elevation, slope, aspect, and hillshade. Specifically, based on the SRTM (Space Shuttle Radar Topography) 30m resolution digital elevation model (DEM), the digital elevation model is spatially interpolated to a 10m resolution. In the cloud computing platform, the terrain analysis module extracts the elevation, slope, aspect, and hillshade of each sampling point in the spatially interpolated digital elevation model to form the terrain factors of the corresponding sampling point.
[0063] This step involves spatial alignment and scale unification processing of multi-source remote sensing environmental factor data (i.e., resampling to 10 meters) to provide consistent and complete original feature data for subsequent model construction.
[0064] Step 3: Normalize the mean values of the various secondary feature data extracted in Step 2, and process the normalized mean values of the various secondary feature data for each sampling point into feature data vectors for the corresponding sampling points. Randomly divide the feature data vectors of all sampling points into training datasets and validation datasets. Finally, construct the original feature matrix from the feature data vectors of each sampling point in the training dataset. The actual pH values of the soil samples collected at each sampling point constitute the tag vector. , It is the set of real numbers;
[0065] In this embodiment The number of sampling points in the training dataset. The original feature matrix represents the number of types of secondary feature data included in the multi-source remote sensing environmental factor data. Each row of data is a feature data vector of a sampling point;
[0066] In this embodiment, the multi-source remote sensing environmental factor data extracted in step 2 are all continuous variables. Therefore, the Min-Max normalization method is used to normalize the mean of 10 vegetation indices, the mean of VV polarization, the mean of VH polarization, the mean of the ratio VV / VH, the annual average temperature, the annual average precipitation, the annual average soil moisture, the elevation, the slope, the aspect, and the topographic shadow index, scaling them to the [0,1] interval. This can reduce the interference of variable dimension differences on modeling and improve the stability and efficiency of model training. If the variables are categorical, one-hot encoding is used for normalization.
[0067] This invention acquires measured data of soil pH in cultivated land within a target area and its corresponding spatial coordinates; it integrates multi-source remote sensing environmental factor data related to soil pH, wherein the introduction of radar backscattering coefficient enhances the responsiveness to soil structure and moisture conditions.
[0068] Step 4: To enhance the model's adaptability to spatial heterogeneity, this invention introduces a spatial weighting mechanism. Based on the geographic spatial distance (latitude and longitude coordinates) between sampling points, a spatially weighted feature matrix is constructed, and the original feature matrix is adjusted according to the spatially weighted feature matrix. Spatial weighted fusion includes the following steps:
[0069] Step 4.1: Construct a spatial distance matrix based on the latitude and longitude coordinates of all sampling points. Specifically, the Euclidean distance formula is used to calculate the spatial distance between each sampling point and other sampling points. The Euclidean distances between each sampling point and other sampling points are then used to construct spatial distance vectors for the corresponding sampling points according to their sampling point indices. All spatial distance vectors of the sampling points are then used to construct a spatial distance matrix in order of sampling point indices. , Spatial distance matrix The elements in, where, and These are all the serial numbers of the sampling points. , , Indicates sampling point With sampling points Euclidean distance between them, spatial distance matrix each The line represents the first The Euclidean distance between each sampling point and other sampling points, when When the Euclidean distance is 0, the distance between the two sides is zero.
[0070] Step 4.2: Based on the spatial distance matrix The spatial weight matrix is constructed using a Gaussian kernel function. To improve the numerical stability of the weights, the spatial weight matrix is modified. The spatial weights of each row are normalized so that the sum of the spatial weights at each sampling point is 1, resulting in a normalized spatial weight matrix. ;
[0071] Spatial weight matrix Calculated based on the following formula:
[0072] (1)
[0073] In the formula, Sampling points With sampling points Spatial weights between them The parameter representing the range of influence of the set adjustment space;
[0074] Normalization is based on the following formula:
[0075] (2)
[0076] Step 4.3: Construct the spatially weighted feature matrix The normalized spatial weight matrix is obtained through matrix multiplication. With the original feature matrix Multiplying them together yields the spatially weighted eigenvalue matrix. ,Right now:
[0077] (3)
[0078] In the formula, the spatial weighted characteristic matrix The first in The line is the first The spatially weighted feature vector of each sampling point represents the feature vector of each sampling point. The comprehensive feature obtained by taking the center as the reference and performing a weighted average after considering the spatial proximity of all other sampling points includes information from the surrounding sampling points.
[0079] Step 4.4: Construct an enhanced feature matrix that integrates spatial perception information. : The original feature matrix With spatial weighted characteristic matrix Concatenate the columns to obtain the enhanced feature matrix. Based on the following formula:
[0080] (4)
[0081] Enhanced feature matrix It also includes the original feature information of each sampling point and the weighted information of its spatial neighborhood, thereby enhancing the model's ability to express geospatial variations.
[0082] Step 5: Enhance the feature matrix As input to the random forest model, the random forest model is constructed into a weighted random forest model (SWRF model), and then a loss function is set (the label matrix of the measured soil pH value is calculated). The weighted random forest model is trained by analyzing the loss between the predicted soil pH value and the predicted value matrix output by the model, and then determining the optimal model. The values are then saved, and the model parameters are obtained to obtain the trained weighted random forest model;
[0083] Training a weighted random forest model specifically involves determining the optimal Gaussian kernel function. The value is determined using a grid search combined with cross-validation, within the set range. Within the candidate interval, traverse the interval with a set step size, repeating steps 4.2 to 4.4 (i.e., reconstructing the enhanced feature matrix) for each iteration. Train a weighted random forest model and determine the optimal model by minimizing the root mean square error (RMSE). value;
[0084] This embodiment is set in Within the candidate interval [0.01, 0.5], the optimal value is obtained by traversing the interval with a step size of 0.01. The value is 0.08.
[0085] Model evaluation: Construct the original feature matrix from the feature data vectors of each sampling point in the validation dataset. Then, perform step 4 to obtain the enhanced feature matrix corresponding to the validation dataset. A random forest model (RF model) was used as a control group to verify the original feature matrix corresponding to the dataset. This serves as input to the random forest model; it is used to validate the augmented feature matrix corresponding to the dataset. This is the input for the random forest model; the model parameters are set consistently, including: number of decision trees: 500; maximum tree depth: 42; other parameters use default values.
[0086] Step 6: Set up sampling points in the target area to be tested, and execute steps 1 to 4 in sequence to obtain the enhanced feature matrix of the target area to be tested. Then, perform step 5 to obtain the trained weighted random forest model of the target region to be tested, and then use the enhanced feature matrix of the target region to obtain the model. Input the data to obtain the prediction matrix of soil pH values for the target area.
[0087] To verify the effectiveness of the spatial adaptive weighting mechanism, the pH of the topsoil (0.1-20cm) in a certain area of Hubei Province, predicted based on multi-source remote sensing environmental factor data, was used as the reference standard. The pH values were calculated from R... 2 The effectiveness of the method of this invention in predicting the pH of topsoil (0.1-20cm) in a certain area of Hubei Province was evaluated using three evaluation indicators: coefficient of determination (RCD), RMSE, and MAE (mean absolute error).
[0088]
[0089] in, R represents 2The percentage improvement, or the percentage decrease in RMSE error, or the percentage decrease in MAE error. This indicates the accuracy of the pH prediction of the topsoil (0.1-20cm) in a certain area of Hubei Province using the SWRF model of this invention. This indicates the accuracy of the pH prediction for the topsoil (0.1-20cm) of cultivated land in a certain area of Hubei Province using the RF model.
[0090] Table 1 shows the prediction comparison results between the RF model and the SWRF model.
[0091]
[0092] The results show that the method of this invention significantly improves the accuracy of the pH prediction of the topsoil (0-15cm) in a certain area of Hubei Province compared with the RF model, and the modeling accuracy R... 2 An improvement of 30.9% (from 0.463 to 0.606), a decrease of 14.3% in RMSE error (from 0.791 to 0.678), and a decrease of 17.5% in MAE error (from 0.618 to 0.510), such as Figure 2 and Figure 3 As shown;
[0093] In terms of the three evaluation metrics, the SWRF model proposed in this invention outperforms the RF model, indicating that the introduction of the spatial weight matrix significantly improves the model's prediction accuracy and spatial adaptability.
[0094] In one embodiment, a computer device is also provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above method embodiments.
[0095] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon that, when executed by a processor, implements the steps in the above method embodiments.
[0096] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above method embodiments.
[0097] It should be noted that the embodiments described in this invention are merely illustrative of the spirit of the invention. Those skilled in the art to which this invention pertains can make various modifications or additions to the described embodiments or use similar methods to substitute them, without departing from the spirit of the invention or exceeding the scope defined by the appended claims.
Claims
1. A method for estimating cultivated soil pH based on multi-source remote sensing and spatial weighting, characterized in that, The method comprises the following steps: Step 1, arranging a plurality of sampling points in a target area, collecting soil samples of a set depth of each sampling point in the target area based on a global land cover dataset, measuring the actual pH value of the soil samples collected at each sampling point, and recording the longitude and latitude coordinates of each sampling point; Step 2, extracting multi-source remote sensing environmental factor data of the corresponding position according to the longitude and latitude coordinates of each sampling point, wherein the multi-source remote sensing environmental factor data comprises mean values of a plurality of secondary feature data; Step 3, normalize the mean value of each type of secondary feature data of each sampling point extracted in step 2, and construct into an original feature matrix The actual pH value of the soil sample of each sampling point is constructed into a label matrix ; Step 4, based on the latitude and longitude coordinates between the sampling points, a spatially weighted feature matrix is constructed , and the original feature matrix is spatially weighted and fused to obtain an enhanced feature matrix ; Step 5, setting a loss function to enhance the feature matrix The random forest model is constructed as a weighted random forest model for the input of the random forest model, the weighted random forest model is trained, and the trained weighted random forest model is obtained. Step 6, layout sampling points in the target area to be measured, sequentially execute steps 1~5 to obtain a trained weighted random forest model of the target area to be measured, input the enhanced feature matrix of the target area to be measured into the trained weighted random forest model to obtain a prediction matrix of the soil pH value of the target area to be measured. Input, obtain the prediction matrix of the soil pH value of the target area to be measured. The step 4 is to construct a spatially weighted feature matrix by a Gaussian kernel function The step 5 is to set a Gaussian kernel function The step 6 is to traverse the candidate interval with a set step length, and calculate a reconstructed enhanced feature matrix each time The step 7 is to train a weighted random forest model, adopt a root mean square error loss function as a loss function, and determine an optimal value by taking the root mean square error as a judgment index, save model parameters after training, and obtain the trained weighted random forest model a parameter for a set adjustment space influence range; The step 2 specifically comprises the following steps: The multi-source remote sensing environmental factor data comprises a plurality of primary feature data, which are respectively optical vegetation index, radar backscatter coefficient, meteorological factor, and terrain factor; The optical vegetation index is calculated by the following method: selecting images of a set time range in the target area, and respectively removing cloud and cloud shadow interference, and respectively calculating atmospheric impedance vegetation index, soil-adjusted vegetation index, normalized difference water index, normalized difference red edge vegetation index, conversion vegetation index, normalized difference cloud index, difference vegetation index, enhanced vegetation index, and ratio vegetation index, which are secondary feature data; Pixel-level time mean value synthesis is performed on each type of vegetation index to obtain the mean value of each type of vegetation index of each pixel in the set time period, and the mean value of each type of vegetation index of the pixel where each sampling point is located is extracted to form the optical vegetation index feature of the corresponding sampling point; The radar backscatter coefficient is calculated by the following method: selecting radar images, extracting VV polarization, VH polarization, and the ratio of VV polarization and VH polarization in the set time period, and performing pixel-level time mean value synthesis on the VV polarization, VH polarization, and the ratio of VV polarization and VH polarization in the set time period after correction processing of all radar images to obtain the mean value of VV polarization, VH polarization, and the ratio of VV polarization and VH polarization of each pixel in the target time period, and then extracting the mean value of VV polarization, the mean value of VH polarization, and the mean value of the ratio of VV polarization and VH polarization of the pixel where the sampling point is located to form the radar backscatter coefficient of the corresponding sampling point; The meteorological factor is calculated by the following method: obtaining original meteorological data of the target area in the set time period by using a public climate dataset, resampling to a set resolution, and aligning with the longitude and latitude coordinates of each sampling point to obtain resampled meteorological data of the target area; Extracting annual temperature data, annual precipitation data, and annual soil moisture data in the resampled meteorological data of the target area, respectively, and calculating the mean values to obtain the annual average temperature, the annual average precipitation, and the annual average soil moisture of the target area, and extracting the annual average temperature, the annual average precipitation, and the annual average soil moisture of the position where each sampling point is located to form the meteorological factor of the corresponding sampling point; The terrain factor is calculated by: based on a digital elevation model, spatially interpolating the digital elevation model to a set resolution, extracting the elevation, slope, aspect, and terrain shadow index of each sampling point in the spatially interpolated digital elevation model by a terrain analysis module in a cloud computing platform to form the terrain factor of the corresponding sampling point.
2. The method according to claim 1, wherein, The original feature matrix of step 3 , is the number of sampling points, is the number of secondary feature data included in the multi-source remote sensing environmental factor data, the original feature matrix Each row of data of the original feature matrix is a feature data vector composed of the mean values of various types of secondary feature data of a sampling point. The label matrix , is the set of real numbers. 3.The method of claim 1, wherein, The step 4 specifically includes the following steps: Step 4.1, calculate the Euclidean distance between each sampling point and other sampling points based on the latitude and longitude coordinates of all sampling points, and construct a spatial distance matrix ; Step 4.2, constructing a spatial weight matrix using a Gaussian kernel function , and normalizing the spatial weight matrix to obtain a normalized spatial weight matrix ; Step 4.3: The normalized spatial weight matrix is obtained through matrix multiplication. With the original feature matrix Multiplying them together yields the spatially weighted eigenvalue matrix. , Spatial weighted feature matrix The first in The line is the first Spatial weighted feature vector of each sampling point.
4. The method according to claim 3, wherein, The spatial distance matrix is constructed in particular by Calculate the Euclidean distance between each sampling point and other sampling points based on the latitude and longitude coordinates of all sampling points. Construct a spatial distance vector for each sampling point according to its sampling point index. Then, construct a spatial distance matrix from all sampling point spatial distance vectors in sequence according to their sampling point indices. , Sampling points With sampling points The Euclidean distance between them and These are all the serial numbers of the sampling points. , Spatial distance matrix each The line is the first The Euclidean distance between each sampling point and other sampling points, when When the Euclidean distance is 0, the distance between the two sides is zero.
5. The method according to claim 4, wherein, The spatial weight matrix is calculated based on the following equation: ; normalizing the spatial weight matrix normalizing each row of the spatial weight matrix each sample point to 1 based on the following equation: ; wherein is the spatial weight between sample point and sample point .
6. The method according to claim 1, wherein, the enhanced feature matrix of step 4 to obtain the original feature matrix the spatially weighted feature matrix concatenating by column: ; enhanced feature matrix . 7.A computer device, comprising a memory and a processor, wherein the memory stores a computer program, and the computer device is configured to perform the method according to any one of claims 1-6 when the computer program is executed by the processor. The processor implements the steps of the method of any one of claims 1 to 6 when executing the computer program.
8. A computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program, when executed by the processor, implements the steps of the method of any one of claims 1 to 6.
Citation Information
Patent Citations
Soil mineral binding state organic carbon prediction method and device based on random forest and environmental variables
CN115758270A
Cultivated land change monitoring method and system based on remote sensing satellite data
CN117935081A