A method for monitoring salinized arable land based on spectral analysis
By integrating spectral analysis and model fusion technologies, the limitations of traditional soil salinity monitoring methods in terms of accuracy and coverage have been overcome, enabling high-precision monitoring of saline-alkali farmland information, guiding crop planting, and improving crop yield and quality in saline-alkali areas.
Patent Information
- Application Number
- CN202510362492.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-26
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2045-03-26
AI Technical Summary
Traditional methods for monitoring soil salinity are limited by manpower and resources, making it difficult to achieve large-scale, high-precision monitoring.
A method for monitoring salinized farmland based on spectral analysis was adopted. By integrating multiple spectral indices and fusing multispectral data with ground hyperspectral data, feature selection and soil salinity inversion were performed using random forest and support vector machine models, and dynamic monitoring was carried out in conjunction with Sentinel-2 satellite data.
It significantly improves the accuracy of soil salinity inversion, provides guidance for crop planting, avoids planting salt-sensitive crops in severely salinized areas, and improves crop yield and quality.
Smart Images

Figure CN120160995B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of soil salinization technology, and in particular to a method for monitoring salinized farmland based on spectral analysis. Background Technology
[0002] Traditional methods for monitoring soil salinity mainly rely on ground-based measurement data. However, due to limitations in manpower and resources, it is difficult to achieve large-scale, high-precision monitoring. Therefore, a method for monitoring salinized farmland based on spectral analysis is proposed. Summary of the Invention
[0003] The purpose of this invention is to address the shortcomings of existing technologies by proposing a method for monitoring information on salinized farmland based on spectral analysis.
[0004] In order to achieve the above object, the present invention adopts the following technical solutions:
[0005] A method for monitoring salinized arable land based on spectral analysis includes the following steps:
[0006] Step 1: Data Preparation: Collect and process data covering the target area. The data includes ground data and satellite data. The ground data includes soil samples, hyperspectral data, crop planting data, and autumn irrigation information. The satellite data includes multispectral data, land use data, and elevation data.
[0007] Step 2: Spectral Index Calculation and Fusion: The salinity index (such as derivative, logarithm, etc.) after ground hyperspectral transformation is fused with the salinity index calculated by Sentinel-2 multispectral analysis at the same time through univariate linear regression. The satellite multispectral index is fused with the hyperspectral index to generate a hyperspectral-multispectral fused index dataset. Based on the transformed data, a feature variable library is constructed. The feature variable library includes salinity index, vegetation index, and environmental variables, integrating 72 salinity indices, 6 vegetation indices (NDVI, EVI, CRSI, etc.), and 3 environmental variables (elevation, autumn irrigation status, crop type), for a total of 81 parameters.
[0008] Step 3: Feature Variable Selection: Select key features from the feature variable library, evaluate feature importance based on the random forest model, calculate the relative importance of each feature, and select key variables based on cumulative importance ≥ 85% to form the final feature set;
[0009] Step 4: Support Vector Machine Model Construction and Training: Divide the dataset into a training set (70%) and a validation set (30%) according to the proportions. Use a radial basis function kernel and optimize the penalty parameter and kernel parameter through grid search. Input the final feature set and the measured soil salinity values to train the support vector machine regression model.
[0010] Step 5: Salinity Inversion and Accuracy Verification: The SVM model is applied to Sentinel-2 imagery and fusion index data to output predicted soil salinity values. Based on the predicted soil salinity values, a spatial distribution map of soil salinity in the Hetao Irrigation District is generated, and the salinization level of the cultivated land is classified in combination with the inversion results.
[0011] Step Six: Dynamic Monitoring: Acquire Sentinel-2 images quarterly, repeat steps up to Step Five, and generate a dynamic map of salinization changes. Combine autumn irrigation records with crop growth to assess the rate of salt accumulation and the effectiveness of control measures. Based on the salinization level (mild / moderate / severe), develop differentiated salt leaching irrigation plans. Combine planting structure data to recommend salt-tolerant crops (such as sunflowers) or areas for application of soil conditioners.
[0012] The above further includes:
[0013] Furthermore, in step one, by using a grid-based sampling method in the cultivated land of the Hetao Irrigation District, combined with adjustments to roads, soil types, and planting structures, topsoil samples from 0cm to 20cm were collected. Simultaneously, the hyperspectral reflectance in the range of 350nm to 2500nm was measured using an SR-3500 ground object spectrometer, spectral transformation was performed, and a fusion spectral index was generated. Whiteboard calibration was performed before each measurement, and the average of three measurements was taken. Soil samples were then analyzed to obtain soil salinity data.
[0014] The crop planting data was constructed based on Sentinel-2 satellite imagery to create an NDVI (Normalized Difference Vegetation Index) time series. Combined with Savitzky-Golay filtering and a decision tree hierarchical classification model, the crop planting structure was extracted (classification accuracy 92.1%).
[0015] The autumn irrigation information was obtained by extracting irrigated areas using the MNDWI (Modified Normalized Water Body) index and distinguishing between irrigated and non-irrigated plots using a threshold method. The MNDWI index is expressed as follows: Wherein, G represents the reflectance of the green band, corresponding to the spectral characteristics of vegetation and water bodies, and MIR represents the reflectance of the middle infrared band, which is sensitive to water bodies and used to distinguish water bodies from dry surfaces.
[0016] Further, in step one, the multispectral data is used to acquire Sentinel-2MSI images, calculate the Normalized Difference Vegetation Index (NDVI), Enhanced Vegetation Index (EVI), and Canopy Salt Response Index (CRSI) vegetation indices, and extract the maximum / minimum values of vegetation indices from June to September to eliminate cloud interference.
[0017] The land use data uses an ESRI 10m resolution land use map, and the "crops" pixels are used to define the scope of cultivated land.
[0018] By integrating Copernicus DEM data, surface elevation information was obtained.
[0019] Furthermore, the calculation formulas for the Normalized Difference Vegetation Index (NDVI), Enhanced Vegetation Index (EVI), and Canopy Salinity Response Index (CRSI) are expressed as follows:
[0020] NDVI = (NIR - R) / (NIR + R);
[0021] EVI=G×(NIR-R) / (NIR+6R+7.5B+1);
[0022]
[0023] maxNDVI=max(NDVIt1,NDVIt2,…,NDVItn);
[0024] minNDVI = min(NDVIt1, NDVIt2, ..., NDVItn);
[0025] Where NIR represents the reflectance of the Near Infrared Band, R represents the reflectance of the Red Band, G represents the reflectance of the Green Band, B represents the reflectance of the Blue Band, and NDVI... t1 NDVI t2 , ..., NDVI tn The values represent NDVI at different time points. maxNDVI represents the NDVI of the crop under optimal growth conditions, reflecting the vegetation potential in areas with less salinization. minNDVI represents the NDVI of the crop under the most severe stress, which is directly related to the degree of salinization. The larger the difference between the two (maxNDVI-minNDVI), the more significant the seasonal impact of salinization on crop growth.
[0026] Further, in step two, the process of fusing the salinity index (such as derivative, logarithm, etc.) after ground hyperspectral transformation with the salinity index calculated by Sentinel-2 multispectral analysis at the same time through univariate linear regression, and fusing the satellite multispectral index with the hyperspectral index to generate a hyperspectral-multispectral fused index dataset, includes the following steps:
[0027] A univariate linear regression analysis was performed on the salinity index after ground hyperspectral transformation and the salinity index calculated by Sentinel-2 multispectral analysis. The salinity index is expressed as...
[0028]
[0029] Among them, BI represents the brightness index, reflecting the brightness characteristics of the soil surface. This index can effectively capture the brightness differences in salinized areas. R represents the reflectance of the red band, corresponding to the spectral characteristics of minerals such as iron oxide in the soil. NIR represents the reflectance of the near-infrared band, which is related to soil moisture and organic matter content. S3 aims to enhance the sensitivity of salt to the green-red band combination, while suppressing the influence of atmospheric scattering or light changes through the blue band. It is suitable for the identification of salt frost in arid areas. G represents the reflectance of the green band, which is related to vegetation cover and some salt characteristics. B represents the reflectance of the blue band. The reflectance of the blue and red bands is used to detect salt frost on the soil surface; S5 utilizes the high sensitivity of the blue and red bands to salt, combined with green band normalization to reduce interference from vegetation cover, and is suitable for shallow salinization monitoring; S6 represents the salinity index 6, which can separate the mixed spectral signals of salt and vegetation, and is suitable for complex scenarios where vegetation cover and salinization coexist; SI2 can comprehensively characterize the spatial distribution of soil salt and is suitable for large-scale dynamic monitoring of salinization; SI3 targets color changes caused by surface salt (such as salt frost whitening or crusting darkening) and is used to quickly identify salinization patches.
[0030] A regression equation was established, with the salinity index calculated by Sentinel-2 multispectral analysis as the independent variable and the salinity index after ground hyperspectral transformation as the dependent variable. The regression equation was expressed as MNSI = a * MNSI s2 +b, where MNSI represents the salinity index after hyperspectral transformation of the ground, MNSI s2 The salinity index is calculated using Sentinel-2 multispectral analysis, where a and b are regression coefficients obtained using the least squares method.
[0031] The Sentinel-2 multispectral data were converted into a fusion index corresponding to the ground hyperspectral data using a regression equation.
[0032] The fusion index obtained by univariate linear regression is combined with the spatial location information of Sentinel-2 multispectral imagery to generate a hyperspectral multispectral fusion index dataset.
[0033] Furthermore, in step three, the process of selecting key features from the feature variable library includes the following steps:
[0034] Random forest model training: Using soil salinity data as the dependent variable and the feature variables (72 salinity indices, 6 vegetation indices, and 3 environmental variables) in the feature variable library as independent variables, a random forest model is constructed. The parameters of the random forest model are set to generate 200 decision trees, and 10-fold cross-validation is used to evaluate the performance.
[0035] Feature importance calculation: The Gini importance metric is used to evaluate feature importance. In each decision tree, feature splitting is measured by the reduction in Gini impurity, as shown in the formula:
[0036]
[0037] Gini parent Gini left Gini right The Gini impurity of the parent node, left child node, and right child node, respectively, N. parent N left N right These represent the number of samples for the corresponding nodes;
[0038] The final importance of each feature is the sum of its Gini impurity reductions across all trees, after normalization, expressed as follows: Where T is the total number of decision trees and M is the total number of features;
[0039] Feature ranking and cumulative filtering: Features are ranked from highest to lowest importance to obtain the sequence {f1, f2, ..., f...}. N} Calculate cumulative importance Where N represents the number of features;
[0040] Select one that satisfies Cumulative Importance k The smallest feature subset {f1,f2,...,f} with ≥85% k} as the final feature set;
[0041] Key output variables: The filtered feature set contains the key variables that contribute the most to soil salinity inversion, which reduces data redundancy and retains more than 85% of the effective information, thus optimizing the model's generalization ability.
[0042] Furthermore, in step four, the specific steps for training the support vector machine regression model are as follows:
[0043] Data partitioning: The filtered feature set and the dataset of measured soil salinity values were divided into training set and validation set in a 7:3 ratio. Stratified random sampling was used during the partitioning to ensure that the distribution of salinization levels in the training set and validation set was consistent and to avoid data bias.
[0044] Data standardization: The input data is processed using Z-score standardization, and the processing is expressed as follows: Where μ is the mean of the features and σ is the standard deviation, ensuring that the mean of each feature is 0 and the variance is 1;
[0045] Model initialization: A support vector regression model is selected, and the radial basis function is used as the kernel function, with the expression K(x) = K(x). i ,x j )=exp(-γ||x i -x j || 2 The objective function of the model is Where C is the penalty parameter, controlling the balance between model complexity and training error; γ is the kernel function parameter, determining the distribution of data mapped to the high-dimensional space; ξ i As slack variables, some samples are allowed to deviate from the hyperplane;
[0046] Parameter optimization: Hyperparameters C and γ are optimized using grid search combined with 5-fold cross-validation. The parameter grid is defined as C∈{0.1,1,10,100}, γ∈{0.001,0.01,0.1,1}. The parameter combinations are traversed, and the mean squared error of the validation set is used as the evaluation metric. The mean squared error is expressed as... Choose the parameter combination that minimizes RMSE as the optimal solution;
[0047] Model training: The standardized training set data is input into the support vector machine regression model, and the optimal parameters are used for training to fit the relationship between the feature variables and soil salinity.
[0048] Model validation: The model performance is evaluated using a validation set, and the coefficient of determination and mean squared error are calculated. The formula for calculating the coefficient of determination is expressed as follows:
[0049] Experimental results show that the validation set R^2 = 0.64 and RMSE = 5.41, and the model's generalization ability is better than that of Random Forest (RF) and Partial Least Squares Regression (PLSR).
[0050] Results Interpretation: SVM captures the nonlinear relationship between multi-source data (spectral indices, vegetation indices, environmental variables) and soil salinity by maximizing the margin and kernel trick. Its lower overfitting risk (compared to RF) and higher explanatory power (compared to PLSR) make it suitable for salinization inversion tasks in the complex environment of the Hetao Irrigation District.
[0051] Further, in step five, the step of generating a spatial distribution map of soil salinity in the Hetao Irrigation Area based on predicted soil salinity values, and classifying the salinity levels of cultivated land based on the inversion results, includes the following steps:
[0052] Data preprocessing: The raster data of soil salinity prediction values output by the SVM model are overlaid with the vector data of cultivated land area in the Hetao Irrigation District to extract the predicted soil salinity values within the cultivated land area.
[0053] Spatial distribution map drawing: Use GIS software (such as ArcGIS) to convert the extracted raster data of farmland soil salinity prediction values into a spatial distribution map, and use different colors or symbols to represent different salinity levels according to the magnitude of the soil salinity prediction values;
[0054] Salinization classification standard: Referring to the soil salinization classification method, the predicted soil salinity is divided into five levels: non-salinized, slightly salinized, moderately salinized, severely salinized, and saline soil.
[0055] Implementation of grading: Based on the predicted soil salinity value, each farmland raster unit is divided into the corresponding salinity level. The grading can be implemented using the reclassification tool in GIS software or a custom script.
[0056] The present invention has the following beneficial effects:
[0057] 1. In this invention, by integrating and applying multiple spectral indices and fusing multispectral data with ground hyperspectral data, the accuracy of soil salinity inversion is significantly improved. The soil salinity information obtained through inversion can provide guidance for crop planting, avoiding the planting of salt-sensitive crops in severely salinized areas, thereby improving crop yield and quality.
[0058] 2. In this invention, targeted monitoring is achieved by combining land use data, crop planting data and autumn irrigation information. The NDVI time series is constructed using Sentinel-2 multi-temporal data, and the max / min vegetation index of the main growth period of crops is extracted to reduce short-term interference. Attached Figure Description
[0059] Figure 1 This is a flowchart illustrating the steps of a method for monitoring salinized farmland based on spectral analysis proposed in this invention. Detailed Implementation
[0060] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0061] Please see Figure 1 As shown, this invention is a method for monitoring salinized arable land based on spectral analysis, comprising the following steps:
[0062] Step 1: Data Preparation: Collect and process data covering the target area. The data includes ground data and satellite data. The ground data includes soil samples, hyperspectral data, crop planting data, and autumn irrigation information. The satellite data includes multispectral data, land use data, and elevation data.
[0063] Step 2: Spectral Index Calculation and Fusion: The salinity index (such as derivative, logarithm, etc.) after ground hyperspectral transformation is fused with the salinity index calculated by Sentinel-2 multispectral analysis at the same time through univariate linear regression. The satellite multispectral index is fused with the hyperspectral index to generate a hyperspectral-multispectral fused index dataset. Based on the transformed data, a feature variable library is constructed. The feature variable library includes salinity index, vegetation index, and environmental variables, integrating 72 salinity indices, 6 vegetation indices (NDVI, EVI, CRSI, etc.), and 3 environmental variables (elevation, autumn irrigation status, crop type), for a total of 81 parameters.
[0064] Step 3: Feature Variable Selection: Select key features from the feature variable library, evaluate feature importance based on the random forest model, calculate the relative importance of each feature, and select key variables based on cumulative importance ≥ 85% to form the final feature set;
[0065] Step 4: Support Vector Machine Model Construction and Training: Divide the dataset into a training set (70%) and a validation set (30%) according to the proportions. Use a radial basis function kernel and optimize the penalty parameter and kernel parameter through grid search. Input the final feature set and the measured soil salinity values to train the support vector machine regression model.
[0066] Step 5: Salinity Inversion and Accuracy Verification: The SVM model is applied to Sentinel-2 imagery and fusion index data to output predicted soil salinity values. Based on the predicted soil salinity values, a spatial distribution map of soil salinity in the Hetao Irrigation District is generated, and the salinization level of the cultivated land is classified in combination with the inversion results.
[0067] Step Six: Dynamic Monitoring: Acquire Sentinel-2 images quarterly, repeat steps up to Step Five, and generate a dynamic map of salinization changes. Combine autumn irrigation records with crop growth to assess the rate of salt accumulation and the effectiveness of control measures. Based on the salinization level (mild / moderate / severe), develop differentiated salt leaching irrigation plans. Combine planting structure data to recommend salt-tolerant crops (such as sunflowers) or areas for application of soil conditioners.
[0068] In one embodiment, in step one, by using a grid-based sampling method in the cultivated land of the Hetao Irrigation District, combined with adjustments to roads, soil types, and planting structures, topsoil samples from 0cm to 20cm were collected. Simultaneously, the hyperspectral reflectance in the range of 350nm to 2500nm was measured using an SR-3500 ground object spectrometer. Spectral transformation was performed to generate a fused spectral index. Whiteboard calibration was performed before each measurement. The average of three measurements was taken, and the quartile method was used to identify and remove outliers in the soil salinity samples, reducing the impact of extreme values on the inversion results. Soil samples were then tested to obtain soil salinity data.
[0069] The spectral transformation includes the first derivative (R'), the second derivative (R”), the reciprocal (1 / R), the reciprocal first derivative ((1 / R)'), the reciprocal second derivative ((1 / R)”), the logarithm (lgR), the logarithmic first derivative ((lgR)'), the logarithmic second derivative ((lgR)”), the square root (R^0.5), the square root first derivative ((R^0.5)'), and the square root second derivative ((R^0.5)”).
[0070] The crop planting data was constructed based on Sentinel-2 satellite imagery to create an NDVI (Normalized Difference Vegetation Index) time series. Combined with Savitzky-Golay filtering and a decision tree hierarchical classification model, the crop planting structure was extracted (classification accuracy 92.1%).
[0071] The autumn irrigation information was obtained by extracting irrigated areas using the MNDWI (Modified Normalized Water Body) index and distinguishing between irrigated and non-irrigated plots using a threshold method. The MNDWI index is expressed as follows: Wherein, G represents the reflectance of the green band, corresponding to the spectral characteristics of vegetation and water bodies, and MIR represents the reflectance of the middle infrared band, which is sensitive to water bodies and used to distinguish water bodies from dry surfaces.
[0072] In one embodiment, in step one, the multispectral data is used to acquire Sentinel-2MSI images, calculate the Normalized Difference Vegetation Index (NDVI), Enhanced Vegetation Index (EVI), and Canopy Salt Response Index (CRSI) vegetation indices, and extract the maximum / minimum values of vegetation indices from June to September to eliminate cloud interference.
[0073] The land use data uses an ESRI 10m resolution land use map, and the "crops" pixels are used to define the scope of cultivated land.
[0074] By integrating Copernicus DEM data, surface elevation information was obtained.
[0075] In one embodiment, the calculation formulas for the Normalized Difference Vegetation Index (NDVI), Enhanced Vegetation Index (EVI), and Canopy Salinity Response Index (CRSI) are expressed as follows:
[0076] NDVI = (NIR - R) / (NIR + R);
[0077] EVI=G×(NIR-R) / (NIR+6R+7.5B+1);
[0078]
[0079] maxNDVI=max(NDVIt1,NDVIt2,…,NDVItn);
[0080] minNDVI = min(NDVIt1, NDVIt2, ..., NDVItn);
[0081] Where NIR represents the reflectance of the Near Infrared Band, R represents the reflectance of the Red Band, G represents the reflectance of the Green Band, B represents the reflectance of the Blue Band, and NDVI... t1 NDVI t2 , ..., NDVI tn The values represent NDVI at different time points. maxNDVI represents the NDVI of the crop under optimal growth conditions, reflecting the vegetation potential in areas with less salinization. minNDVI represents the NDVI of the crop under the most severe stress, which is directly related to the degree of salinization. The larger the difference between the two (maxNDVI-minNDVI), the more significant the seasonal impact of salinization on crop growth.
[0082] In one embodiment, in step two, the process of fusing the salinity index (such as derivative, logarithm, etc.) after ground hyperspectral transformation with the salinity index calculated by Sentinel-2 multispectral analysis at the same time through univariate linear regression, and fusing the satellite multispectral index with the hyperspectral index to generate a hyperspectral-multispectral fused index dataset, includes the following steps:
[0083] A univariate linear regression analysis was performed on the salinity index after ground hyperspectral transformation and the salinity index calculated by Sentinel-2 multispectral analysis. The salinity index is expressed as...
[0084]
[0085] Among them, BI represents the brightness index, reflecting the brightness characteristics of the soil surface. This index can effectively capture the brightness differences in salinized areas. R represents the reflectance of the red band, corresponding to the spectral characteristics of minerals such as iron oxide in the soil. NIR represents the reflectance of the near-infrared band, which is related to soil moisture and organic matter content. S3 aims to enhance the sensitivity of salt to the green-red band combination, while suppressing the influence of atmospheric scattering or light changes through the blue band. It is suitable for the identification of salt frost in arid areas. G represents the reflectance of the green band, which is related to vegetation cover and some salt characteristics. B represents the reflectance of the blue band. The reflectance of the blue and red bands is used to detect salt frost on the soil surface; S5 utilizes the high sensitivity of the blue and red bands to salt, combined with green band normalization to reduce interference from vegetation cover, and is suitable for shallow salinization monitoring; S6 represents the salinity index 6, which can separate the mixed spectral signals of salt and vegetation, and is suitable for complex scenarios where vegetation cover and salinization coexist; SI2 can comprehensively characterize the spatial distribution of soil salt and is suitable for large-scale dynamic monitoring of salinization; SI3 targets color changes caused by surface salt (such as salt frost whitening or crusting darkening) and is used to quickly identify salinization patches.
[0086] A regression equation was established, with the salinity index calculated by Sentinel-2 multispectral analysis as the independent variable and the salinity index after ground hyperspectral transformation as the dependent variable. The regression equation was expressed as MNSI = a * MNSI s2 +b, where MNSI represents the salinity index after hyperspectral transformation of the ground, MNSI s2 The salinity index is calculated using Sentinel-2 multispectral analysis, where a and b are regression coefficients obtained using the least squares method.
[0087] The Sentinel-2 multispectral data were converted into a fusion index corresponding to the ground hyperspectral data using a regression equation.
[0088] The fusion index obtained by univariate linear regression is combined with the spatial location information of Sentinel-2 multispectral imagery to generate a hyperspectral multispectral fusion index dataset.
[0089] In one embodiment, step three, the step of filtering key features from the feature variable library, includes the following steps:
[0090] Random forest model training: Using soil salinity data as the dependent variable and the feature variables (72 salinity indices, 6 vegetation indices, and 3 environmental variables) in the feature variable library as independent variables, a random forest model is constructed. The parameters of the random forest model are set to generate 200 decision trees, and 10-fold cross-validation is used to evaluate the performance.
[0091] Feature importance calculation: The Gini importance metric is used to evaluate feature importance. In each decision tree, feature splitting is measured by the reduction in Gini impurity, as shown in the formula:
[0092]
[0093] Gini parent Gini left Gini right The Gini impurity of the parent node, left child node, and right child node, respectively, N. parent N left N right These represent the number of samples for the corresponding nodes;
[0094] The final importance of each feature is the sum of its Gini impurity reductions across all trees, after normalization, expressed as follows: Where T is the total number of decision trees and M is the total number of features;
[0095] Feature ranking and cumulative filtering: Features are ranked from highest to lowest importance to obtain the sequence {f1, f2, ..., f...}. N} Calculate cumulative importance Where N represents the number of features;
[0096] Select one that satisfies Cumulative Importance k The smallest feature subset {f1,f2,...,f} with ≥85% k} as the final feature set;
[0097] Key output variables: The filtered feature set contains the key variables that contribute the most to soil salinity inversion, which reduces data redundancy and retains more than 85% of the effective information, thus optimizing the model's generalization ability.
[0098] In one embodiment, the specific steps for training the support vector machine regression model in step four are as follows:
[0099] Data partitioning: The filtered feature set and the dataset of measured soil salinity values were divided into training set and validation set in a 7:3 ratio. Stratified random sampling was used during the partitioning to ensure that the distribution of salinization levels in the training set and validation set was consistent and to avoid data bias.
[0100] Data standardization: The input data is processed using Z-score standardization, and the processing is expressed as follows: Where μ is the mean of the features and σ is the standard deviation, ensuring that the mean of each feature is 0 and the variance is 1;
[0101] Model initialization: A support vector regression model is selected, and the radial basis function is used as the kernel function, with the expression K(x) = K(x). i ,x j )=exp(-γ||x i -x j || 2 The objective function of the model is Where C is the penalty parameter, controlling the balance between model complexity and training error; γ is the kernel function parameter, determining the distribution of data mapped to the high-dimensional space; ξ i As slack variables, some samples are allowed to deviate from the hyperplane;
[0102] Parameter optimization: Hyperparameters C and γ are optimized using grid search combined with 5-fold cross-validation. The parameter grid is defined as C∈{0.1,1,10,100}, γ∈{0.001,0.01,0.1,1}. The parameter combinations are traversed, and the mean squared error of the validation set is used as the evaluation metric. The mean squared error is expressed as... Choose the parameter combination that minimizes RMSE as the optimal solution;
[0103] Model training: The standardized training set data is input into the support vector machine regression model, and the optimal parameters are used for training to fit the relationship between the feature variables and soil salinity.
[0104] Model validation: The model performance is evaluated using a validation set, and the coefficient of determination and mean squared error are calculated. The formula for calculating the coefficient of determination is expressed as follows:
[0105] Experimental results show that the validation set R^2 = 0.64 and RMSE = 5.41, and the model's generalization ability is better than that of Random Forest (RF) and Partial Least Squares Regression (PLSR).
[0106] Results Interpretation: SVM captures the nonlinear relationship between multi-source data (spectral indices, vegetation indices, environmental variables) and soil salinity by maximizing the margin and kernel trick. Its lower overfitting risk (compared to RF) and higher explanatory power (compared to PLSR) make it suitable for salinization inversion tasks in the complex environment of the Hetao Irrigation District.
[0107] In one embodiment, step five, which involves generating a spatial distribution map of soil salinity in the Hetao Irrigation District based on predicted soil salinity values and classifying the salinity levels of cultivated land based on the inversion results, includes the following steps:
[0108] Data preprocessing: The raster data of soil salinity prediction values output by the SVM model are overlaid with the vector data of cultivated land area in the Hetao Irrigation District to extract the predicted soil salinity values within the cultivated land area.
[0109] Spatial distribution map drawing: Use GIS software (such as ArcGIS) to convert the extracted raster data of farmland soil salinity prediction values into a spatial distribution map, and use different colors or symbols to represent different salinity levels according to the magnitude of the soil salinity prediction values;
[0110] Salinization classification standard: Referring to the soil salinization classification method, the predicted soil salinity is divided into five levels: non-salinized, slightly salinized, moderately salinized, severely salinized, and saline soil.
[0111] Implementation of grading: Based on the predicted soil salinity value, each farmland raster unit is divided into the corresponding salinity level. The grading can be implemented using the reclassification tool in GIS software or a custom script.
[0112] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A method for monitoring salinized arable land based on spectral analysis, characterized in that, Includes the following steps: Step 1: Data Preparation: Collect and process data covering the target area. The data includes ground data and satellite data. The ground data includes soil samples, hyperspectral data, crop planting data, and autumn irrigation information. The satellite data includes multispectral data, land use data, and elevation data. Step 2: Spectral Index Calculation and Fusion: The salinity index after ground hyperspectral transformation and the salinity index calculated by Sentinel-2 multispectral analysis at the same time are fused by univariate linear regression. The satellite multispectral index and the hyperspectral index are fused to generate a hyperspectral-multispectral fused index dataset. Based on the transformed data, a feature variable library is constructed, which includes salinity index, vegetation index and environmental variables. Step 3 Feature selection: Key features are selected from the feature variable library, feature importance is evaluated based on the random forest model, the relative importance of each feature is calculated, and the final feature set is formed. Step 4: Support Vector Machine Model Construction and Training: Divide the selected feature set and the dataset of measured soil salinity values into training and validation sets according to the proportion. Use radial basis function kernels and optimize the penalty parameters and kernel parameters through grid search. Input the final feature set and measured soil salinity values to train the support vector machine regression model. Step 5: Salinity Inversion and Accuracy Verification: The SVM model is applied to Sentinel-2 imagery and fusion index data to output predicted soil salinity values. Based on the predicted soil salinity values, a spatial distribution map of soil salinity in the Hetao Irrigation District is generated, and the salinization level of the cultivated land is classified in combination with the inversion results. Step Six: Dynamic Monitoring: Acquire Sentinel-2 images quarterly, repeat steps one to five to generate a dynamic change map of salinization, combine autumn irrigation records with crop growth to assess the rate of salt accumulation and the effectiveness of control measures, and formulate a salt leaching irrigation plan based on the salinization level. In step two, the process of fusing the salinity index after ground hyperspectral transformation with the salinity index calculated by Sentinel-2 multispectral analysis at the same time through univariate linear regression, and fusing the satellite multispectral index with the hyperspectral index to generate a hyperspectral-multispectral fused index dataset, includes the following steps: A univariate linear regression analysis was performed on the salinity index after ground hyperspectral transformation and the salinity index calculated by Sentinel-2 multispectral analysis. The salinity index is expressed as... Among them, BI represents the brightness index, reflecting the brightness characteristics of the soil surface; R represents the reflectance in the red band, corresponding to the spectral characteristics of minerals in the soil; NIR represents the reflectance in the near-infrared band, which is related to soil moisture and organic matter content; S3 is suitable for the identification of salt frost in arid areas; G represents the reflectance in the green band, which is related to vegetation cover and some salinity characteristics; B represents the reflectance in the blue band, used to detect salt frost on the soil surface; S5 is suitable for shallow salinization monitoring; S6 represents the salinity index 6; SI2 is suitable for large-scale dynamic monitoring of salinization; SI3 is used for rapid identification of salinization patches. A regression equation was established, with the salinity index calculated by Sentinel-2 multispectral analysis as the independent variable and the salinity index after ground hyperspectral transformation as the dependent variable. The regression equation was expressed as MNSI = a * MNSI s2 +b, where MNSI represents the salinity index after hyperspectral transformation of the ground, MNSI s2 The salinity index is calculated using Sentinel-2 multispectral analysis, where a and b are regression coefficients obtained using the least squares method. The Sentinel-2 multispectral data were converted into a fusion index corresponding to the ground hyperspectral data using a regression equation. The fusion index obtained by univariate linear regression is combined with the spatial location information of Sentinel-2 multispectral imagery to generate a hyperspectral multispectral fusion index dataset.
2. The method for monitoring salinized arable land based on spectral analysis according to claim 1, characterized in that, In step one, by using a grid-based sampling method in the cultivated land of the Hetao Irrigation District, combined with adjustments to roads, soil types and planting structures, topsoil samples from 0cm to 20cm were collected. Simultaneously, the hyperspectral reflectance in the range of 350nm to 2500nm was measured using an SR-3500 ground object spectrometer, spectral transformation was performed, a fusion spectral index was generated, and soil samples were tested to obtain soil salinity data. The crop planting data is based on Sentinel-2 satellite imagery to construct NDVI time series, and combined with Savitzky-Golay filtering and decision tree hierarchical classification model to extract crop planting structure; The autumn irrigation information is extracted using the Improved Normalized Water Extraction Index (MNDWI) to identify irrigated areas, and a threshold method is used to distinguish between irrigated and non-irrigated plots. The Improved Normalized Water Extraction Index (MNDWI) is expressed as... Wherein, G represents the reflectance of the green band, corresponding to the spectral characteristics of vegetation and water bodies, and MIR represents the reflectance of the mid-infrared band, which is sensitive to water bodies and used to distinguish water bodies from dry surfaces.
3. The method for monitoring salinized arable land based on spectral analysis according to claim 1, characterized in that, In step one, the multispectral data is used to acquire Sentinel-2MSI images and calculate the normalized vegetation index, enhanced vegetation index, and canopy salt response index. The land use data uses a land use map, and the "crops" pixels are selected to define the scope of cultivated land; By integrating Copernicus DEM data, surface elevation information was obtained.
4. The method for monitoring salinized arable land based on spectral analysis according to claim 3, characterized in that, The calculation formulas for the Normalized Difference Vegetation Index (NDVI), Enhanced Vegetation Index (EVI), and Canopy Salinity Response Index (CRSI) are expressed as follows: NDVI = (NIR - R) / (NIR + R); EVI=G×(NIR-R) / (NIR+6R+7.5B+1); maxNDVI=max(NDVIt1,NDVIt2,...,NDVItn); minNDVI = min(NDVIt1,NDVIt2,...,NDVItn); Wherein, NIR represents the reflectance of the near-infrared band, R represents the reflectance of the red band, G represents the reflectance of the green band, B represents the reflectance of the blue band, NDVIt1, NDVIt2, ..., NDVItn represent the NDVI values at different time points, maxNIVI represents the NDVI of the crop under optimal growth conditions, reflecting the vegetation potential of areas with less salinization, and minNDVI represents the NDVI of the crop under the most severe stress, which is directly related to the degree of salinization.
5. The method for monitoring salinized arable land based on spectral analysis according to claim 1, characterized in that, In step three, the process of selecting key features from the feature variable library includes the following steps: Random forest model training: Using soil salinity data as the dependent variable and the feature variables in the feature variable library as independent variables, a random forest model is constructed. The parameters of the random forest model are set to generate 200 decision trees, and 10-fold cross-validation is used to evaluate the performance. Feature importance calculation: The Gini importance metric is used to evaluate feature importance. In each decision tree, feature splitting is measured by the reduction in Gini impurity, as shown in the formula: Gini parent Gini left Gini right The Gini impurity of the parent node, left child node, and right child node, respectively, N. parent N left N right These represent the number of samples for the corresponding nodes; The final importance of each feature is the sum of its Gini impurity reductions across all trees, after normalization, expressed as follows: Where T is the total number of decision trees and M is the total number of features; Feature ranking and cumulative filtering: Features are ranked from highest to lowest importance to obtain the sequence {f1, f2, ..., f...}. N } Calculate cumulative importance Where N represents the number of features; Select one that satisfies Cumulative Importance k The smallest feature subset {f1,f2,...,f} with ≥85% k } as the final feature set; Output key variables: The filtered feature set contains the key variables that contribute the most to soil salinity inversion.
6. The method for monitoring salinized arable land based on spectral analysis according to claim 1, characterized in that, In step four, the specific steps for training the support vector machine regression model are as follows: Data partitioning: The filtered feature set and the dataset of measured soil salinity values were divided into training set and validation set in a 7:3 ratio, and stratified random sampling was used for partitioning. Data standardization: The input data is processed using Z-score standardization, and the processing is expressed as follows: Where X is the input data, μ is the feature mean of the input data, and σ is the standard deviation of the input data; Model initialization: A support vector regression model is selected, and the radial basis function is used as the kernel function, with the expression K(x) = K(x). i ,x j )=exp(-γ||x i -x j || 2 The objective function of the model is Where w is the weight vector; C is the penalty parameter, controlling the balance between model complexity and training error; γ is the kernel function parameter, determining the distribution of data mapped to the high-dimensional space; ξ is the slack variable, allowing some samples to deviate from the hyperplane. i Let be the slack variable for the i-th sample; Parameter optimization: Hyperparameters C and γ are optimized using grid search combined with 5-fold cross-validation. The parameter grid is defined as C∈{0.1, 1, 10, 100}, γ∈{0.001, 0.01, 0.1, 1}. The parameter combinations are traversed, and the mean squared error of the validation set is used as the evaluation metric. The mean squared error is expressed as... Choose the parameter combination that minimizes RMSE as the optimal solution; Model training: The standardized training set data is input into the support vector machine regression model, and the optimal parameters are used for training to fit the relationship between the feature variables and soil salinity. Model validation: The model performance is evaluated using a validation set, and the coefficient of determination and mean squared error are calculated. The formula for calculating the coefficient of determination is expressed as follows: Interpretation of results: SVM captures the nonlinear relationship between multi-source data and soil salinity by maximizing the interval and kernel trick, making it suitable for salinization inversion tasks in the complex environment of the Hetao Irrigation District.
7. The method for monitoring salinized arable land based on spectral analysis according to claim 1, characterized in that, In step five, the specific steps for generating a spatial distribution map of soil salinity in the Hetao Irrigation District based on predicted soil salinity values and classifying the salinity levels of cultivated land in conjunction with the inversion results are as follows: Data preprocessing: The raster data of soil salinity prediction values output by the SVM model are overlaid with the vector data of cultivated land area in the Hetao Irrigation District to extract the predicted soil salinity values within the cultivated land area. Spatial distribution map drawing: The extracted raster data of farmland soil salinity prediction values are converted into a spatial distribution map. Different colors or symbols are used to represent different salinity levels according to the magnitude of the soil salinity prediction values. Salinization classification standard: Soil salinity prediction values are divided into five levels: non-salinization, slight salinization, moderate salinization, severe salinization, and saline soil; Grading implementation: Based on the predicted soil salinity values, each farmland grid unit is classified into the corresponding salinization level.
Citation Information
Patent Citations
Early warning method and system for soil secondary salinization degree
CN119168831A
Method for improving satellite data inversion soil salinity by using unmanned aerial vehicle data
CN119314035A
Method for jointly estimating soil profile salinity by using time-series remote sensing image
US20240288602A1