Saline-alkali cultivated land information monitoring method based on spectral analysis
Through the integration of multiple spectral indices and data based on spectral analysis, the support vector machine regression model is trained, which solves the problem that traditional soil salt monitoring methods are difficult to achieve large-area and high-precision monitoring, and realizes high-precision inversion and dynamic monitoring of soil salt, provides targeted guidance, and improves crop yield and quality.
Patent Information
- Application Number
- CN202510362492.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-26
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2045-03-26
AI Technical Summary
Traditional soil salinity monitoring methods are difficult to achieve large-area and high-precision monitoring, and cannot effectively solve the dynamic changes in salinized farmland information and the precise division of salinization levels.
Using a saline-alkali cultivated land information monitoring method based on spectral analysis, a characteristic variable library is constructed, key features are screened, and the support vector machine regression model is trained to achieve high-precision inversion and dynamic monitoring of soil salt by integrating multiple spectral indices and multispectral data with ground hyperspectral data.
It significantly improves the accuracy of soil salt inversion, can dynamically monitor salinization, provide targeted guidance, and avoids the cultivation of salt-sensitive crops in areas with severe salinization, thereby improving crop yield and quality.
Smart Images

Figure CN120160995A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of soil salinization, and particularly to a method for monitoring information of saline-alkali cultivated land based on spectral analysis. Background Art
[0003] Traditional soil salinity monitoring methods mainly rely on ground measured data. Limited by manpower and material resources, it is difficult to achieve large-area and high-precision monitoring. Therefore, a method for monitoring information of saline-alkali cultivated land based on spectral analysis is proposed. Summary of the Invention
[0004] The purpose of the present invention is to solve the deficiencies in the prior art, and a method for monitoring information of saline-alkali cultivated land based on spectral analysis is proposed.
[0005] To achieve the above purpose, the present invention adopts the following technical solutions:
[0006] A method for monitoring information of saline-alkali cultivated land based on spectral analysis, comprising the following steps:
[0007] Step 1: Data preparation: Collect and process data covering the target area. The data includes ground data and satellite data. The ground data includes soil samples, hyperspectral data, crop planting data, and autumn irrigation information. The satellite data includes multispectral data, land use data, and elevation data;
[0008] Step 2: Spectral index calculation and fusion: Fuse the salt index (such as derivative, logarithm, etc.) after transformation of ground hyperspectral with the salt index calculated from Sentinel-2 multispectral in the same period through unary linear regression, and fuse the satellite multispectral index with the hyperspectral index to generate a high-multispectral fusion index dataset. Based on the transformed data, construct a feature variable library. The feature variable library contains salt index, vegetation index, and environmental variables. Integrate 72 salt indices, 6 vegetation indices (NDVI, EVI, CRSI, etc.), and 3 environmental variables (elevation, autumn irrigation status, crop type), a total of 81 parameters;
[0009] Step 3: Feature variable selection: Screen out key features from the feature variable library, evaluate the feature importance based on the random forest model, calculate the relative importance of each feature, and screen key variables according to the cumulative importance ≥ 85% to form the final feature set;
[0010] Step 4: Support vector machine model construction and training: Divide the dataset into a training set (70%) and a validation set (30%) according to a ratio, adopt the radial basis function kernel, optimize the penalty parameter and kernel parameter through grid search, input the final feature set and the measured value of soil salinity, and train the support vector machine regression model;
[0011] Step 5: Salt inversion and accuracy verification: Apply the SVM model to Sentinel-2 images and fusion index data, output the predicted soil salt values, generate the spatial distribution map of cultivated land soil salt in the Hetao Irrigation Area based on the predicted soil salt values, and divide the cultivated land salinization grades in combination with the inversion results;
[0012] Step 6: Dynamic monitoring: Obtain Sentinel-2 images quarterly, repeat Steps 1 to 5, generate the dynamic change map of salinization, evaluate the salt accumulation rate and the effect of treatment measures in combination with autumn irrigation records and crop growth, and formulate a differentiated salt leaching irrigation plan according to the salinization grades (mild / moderate / severe). Combine the planting structure data to recommend salt-tolerant crops (such as sunflowers) or areas for applying soil conditioners.
[0013] The above further includes:
[0014] Further, in Step 1, by using the grid sampling method in the cultivated land of the Hetao Irrigation Area, combined with road, soil type, and planting structure adjustment, collect surface soil samples of 0 cm - 20 cm, synchronously use the SR-3500 ground object spectrometer to measure the hyperspectral reflectance in the range of 350 nm - 2500 nm, perform spectral transformation to generate the fusion spectral index, perform whiteboard calibration before each measurement, take the average of three measurements, analyze the soil samples, and obtain the soil salt data;
[0015] The crop planting data constructs an NDVI (Normalized Difference Vegetation Index) time series based on Sentinel-2 satellite images, combines the Savitzky-Golay filter and the decision tree hierarchical classification model to extract the crop planting structure (classification accuracy 92.1%);
[0016] The autumn irrigation information extracts the irrigation area by using the MNDWI (Modified Normalized Difference Water Index), and uses the threshold method to distinguish between irrigated and non-irrigated plots. The MNDWI (Modified Normalized Difference Water Index) is expressed as where G represents the reflectance of the Green Band, corresponding to the spectral characteristics of vegetation and water bodies, and MIR represents the reflectance of the Middle Infrared Band, which is sensitive to water bodies and is used to distinguish water bodies from dry land surfaces.
[0017] Further, in Step 1, for the multispectral data, obtain the Sentinel-2 MSI image, calculate vegetation indices such as the Normalized Difference Vegetation Index (NDVI), Enhanced Vegetation Index (EVI), and Canopy Salt Response Index (CRSI), and extract the maximum / minimum values of the vegetation indices from June to September to eliminate cloud interference;
[0018] For the land use data, an ESRI land use map with a resolution of 10 m is used to screen the "crop" pixels to define the cultivated land area;
[0019] By integrating Copernicus DEM data, the surface elevation information is obtained.
[0020] Furthermore, the calculation formulas of the vegetation indices such as the Normalized Difference Vegetation Index (NDVI), Enhanced Vegetation Index (EVI), and Canopy Salt Response Index (CRSI) are expressed as
[0021] NDVI = (NIR - R) / (NIR + R);
[0022] EVI = G×(NIR - R) / (NIR + 6R + 7.5B + 1);
[0023]
[0024] maxNDVI = max(NDVIt1, NDVIt2,..., NDVItn);
[0025] minNDVI = min(NDVIt1, NDVIt2,..., NDVItn);
[0026] Among them, NIR represents the reflectance of the Near Infrared Band, R represents the reflectance of the Red Band, G represents the reflectance of the Green Band, B represents the reflectance of the Blue Band, NDVI t1 , NDVI t2 ,..., NDVI tn represents the NDVI values at different time points, maxNDVI; represents the NDVI of the crop in the best growth state, reflecting the vegetation potential in the less saline area, minNDVI represents the NDVI of the crop when the stress is most severe, directly related to the degree of salinization. The larger the difference between the two (maxNDVI - minNDVI), the more significant the seasonal impact of salinization on crop growth.
[0027] Furthermore, in step two, the salt index (such as derivative, logarithm, etc.) after the ground hyperspectral transformation and the salt index calculated by Sentinel-2 multispectral at the same period are fused through simple linear regression. The satellite multispectral index and the hyperspectral index are fused to generate a high-multispectral fusion index dataset, including the following steps:
[0028] Perform simple linear regression analysis on the salt index after the ground hyperspectral transformation and the salt index calculated by Sentinel-2 multispectral. The salt index is expressed as
[0029]
[0030] Among them, BI represents the brightness index, which reflects the brightness characteristics of the soil surface. This index can effectively capture the brightness differences in salinized areas. R represents the reflectance of the red band, corresponding to the spectral characteristics of minerals such as iron oxide in the soil. NIR represents the reflectance of the near-infrared band, which is related to soil moisture and organic matter content. S3 aims to enhance the sensitivity of salts to the green-red band combination and suppress the influence of atmospheric scattering or light changes through the blue band, and is applicable to the identification of salt frost in arid areas. G represents the reflectance of the green band, which is related to vegetation cover and some salt characteristics. B represents the reflectance of the blue band, which is used to detect salt frost on the soil surface. S5 utilizes the high sensitivity of the blue and red bands to salts and combines green band normalization to reduce the interference of vegetation cover, and is applicable to shallow salinization monitoring. S6 represents the salt index 6, which can separate the mixed spectral signals of salts and vegetation and is applicable to complex scenarios where vegetation cover and salinization coexist. SI2 can comprehensively characterize the spatial distribution of soil salts and is applicable to large-scale dynamic monitoring of salinization. SI3 targets the color changes caused by surface salts (such as salt frost whitening or crust darkening) and is used to quickly identify salinized patches;
[0031] A regression equation is established, taking the salt index calculated from Sentinel-2 multispectral data as the independent variable and the salt index after ground hyperspectral transformation as the dependent variable. The regression equation is expressed as MNSI = a * MNSI s2 + b, where MNSI represents the salt index after ground hyperspectral transformation, and MNSI s2 represents the salt index calculated from Sentinel-2 multispectral data. a and b are regression coefficients, which are solved by the least squares method;
[0032] Through the regression equation, the Sentinel-2 multispectral data is converted into a fusion index corresponding to the ground hyperspectral data;
[0033] The fusion index obtained by unary linear regression is combined with the spatial position information of the Sentinel-2 multispectral image to generate a high-multispectral fusion index dataset.
[0034] Furthermore, in step three, screening out the key features from the feature variable library includes the following steps:
[0035] Random forest model training: Using soil salinity data as the dependent variable and the feature variables (72 salinity indices, 6 vegetation indices, and 3 environmental variables) within the feature variable library as independent variables, a random forest model is constructed. The parameters of the random forest model are set to generate 200 decision trees, and 10-fold cross-validation is used to evaluate the performance;
[0036] Feature importance calculation: Gini importance is used to evaluate the feature importance. In each decision tree, the split of a feature is measured by the reduction in Gini impurity, and the formula is
[0037]
[0038] where Gini parent , Gini left , Gini right are the Gini impurities of the parent node, left child node, and right child node respectively, and N parent , N left , N right are the number of samples of the corresponding nodes respectively;
[0039] The final importance of each feature is the sum of the reduction in Gini impurity in all trees and is normalized. The normalization process is expressed as where T is the total number of decision trees and M is the total number of features;
[0040] Feature sorting and cumulative screening: Sort the features from high to low importance to obtain the sequence {f1, f2,..., f N}}, and calculate the cumulative importance where N represents the number of features;
[0041] Select the smallest feature subset {f1, f2,..., f k} that satisfies CumulativeImportance k ≥ 85% as the final feature set;
[0042] Output of key variables: The screened feature set contains the key variables that contribute the most to soil salinity inversion, which not only reduces data redundancy but also retains more than 85% of the effective information, optimizing the generalization ability of the model.
[0043] Furthermore, in step four, the specific steps for training the support vector machine regression model are as follows:
[0044] Data partitioning: The screened feature set and the dataset of measured soil salinity values are partitioned into a training set and a validation set in a 7:3 ratio. Stratified random sampling is used during partitioning to ensure that the distribution of salinization levels in the training set and the validation set is consistent, avoiding data bias;
[0045] Data standardization: The input data is processed using Z-score standardization, which is expressed as where μ is the feature mean and σ is the standard deviation, ensuring that each feature has a mean of 0 and a variance of 1;
[0046] Model initialization: Select the support vector regression model, and the kernel function uses the radial basis function, whose expression is K(x i ,x j ) = exp(-γ||x i -x j || 2 ), and the model objective function is where C is the penalty parameter, controlling the balance between model complexity and training error; γ is the kernel function parameter, determining the distribution of data mapped to the high-dimensional space; ξ i is the slack variable, allowing some samples to deviate from the hyperplane;
[0047] Parameter optimization: The hyperparameters C and γ are optimized through grid search combined with 5-fold cross-validation. Define the parameter grid C ∈ {0.1, 1, 10, 100}, γ ∈ {0.001, 0.01, 0.1, 1}, traverse the parameter combinations, and use the mean squared error of the validation set as the evaluation index. The mean squared error is expressed as Select the parameter combination that minimizes the RMSE as the optimal solution;
[0048] Model training: Input the standardized training set data into the support vector machine regression model and train it using the optimal parameters to fit the relationship between the feature variables and soil salinity;
[0049] Model validation: Use the validation set to evaluate the model performance, calculate the coefficient of determination and the mean squared error. The calculation formula for the coefficient of determination is expressed as
[0050] The experimental results show that the R^2 of the validation set is 0.64 and the RMSE is 5.41. The generalization ability of the model is better than that of the random forest (RF) and partial least squares regression (PLSR);
[0051] Result interpretation: SVM captures the non-linear relationship between multi-source data (spectral index, vegetation index, environmental variables) and soil salinity through maximizing the margin and kernel trick. Its lower overfitting risk (compared with RF) and higher interpretability (compared with PLSR) make it suitable for the salt-affected inversion task in the complex environment of the Hetao Irrigation Area.
[0052] Furthermore, in step five, the steps of generating the spatial distribution map of cultivated land soil salinity in the Hetao Irrigation Area based on the predicted value of soil salinity and dividing the salt-affected grades of cultivated land in combination with the inversion results are as follows:
[0053] Data preprocessing: Perform overlay analysis on the raster data of soil salinity prediction values output by the SVM model and the vector data of the cultivated land scope in the Hetao Irrigation Area to extract the soil salinity prediction values within the cultivated land scope;
[0054] Spatial distribution map drawing: Use GIS software (such as ArcGIS) to convert the raster data of the extracted cultivated land soil salinity prediction values into a spatial distribution map, and use different colors or symbols to represent different salinity levels according to the magnitude of the soil salinity prediction values;
[0055] Salinization level classification standard: Refer to the soil salinization classification method to divide the soil salinity prediction values into five levels: non-salinized, slightly salinized, moderately salinized, severely salinized, and saline soil;
[0056] Implementation of level classification: According to the magnitude of the soil salinity prediction values, divide each cultivated land grid unit into the corresponding salinization level, and the level classification can be achieved using the reclassification tool or custom script in GIS software.
[0057] The present invention has the following beneficial effects:
[0058] 1. In the present invention, by integrating and applying a variety of spectral indices and using the fusion of multi-spectral data and ground hyperspectral data, the accuracy of soil salinity inversion is significantly improved. The soil salinity information obtained through inversion can provide guidance for crop planting, avoid planting salt-sensitive crops in areas with severe salinization, and thus improve the yield and quality of crops.
[0059] 2. In the present invention, by combining land use data, crop planting data, and autumn irrigation information, targeted monitoring is achieved. The NDVI time series is constructed using Sentinel-2 multi-temporal data, and the max / min vegetation indices in the main growth period of crops are extracted to reduce short-term interference. BRIEF DESCRIPTION OF THE DRAWINGS
[0060] Figure 1 It is a step diagram of a method for monitoring saline-alkali cultivated land information based on spectral analysis proposed by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0061] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0062] Please refer to Figure 1 As shown, the present invention is a method for monitoring saline-alkali cultivated land information based on spectral analysis, including the following steps:
[0063] Step 1: Data Preparation: Collect and process data covering the target area. The data includes ground data and satellite data. The ground data includes soil samples, hyperspectral data, crop planting data, and autumn irrigation information. The satellite data includes multispectral data, land use data, and elevation data;
[0064] Step 2: Spectral Index Calculation and Fusion: Fuse the salt index (such as derivative, logarithm, etc.) after the transformation of ground hyperspectral data and the salt index calculated from Sentinel-2 multispectral data of the same period through simple linear regression. Fuse the satellite multispectral index and the hyperspectral index to generate a high-multispectral fusion index dataset. Based on the transformed data, construct a feature variable library. The feature variable library contains salt indices, vegetation indices, and environmental variables. Integrate 72 salt indices, 6 vegetation indices (NDVI, EVI, CRSI, etc.), and 3 environmental variables (elevation, autumn irrigation status, crop type), a total of 81 parameters;
[0065] Step 3: Feature Variable Selection: Screen out key features from the feature variable library, evaluate the feature importance based on the random forest model, calculate the relative importance of each feature, and screen key variables according to the cumulative importance ≥ 85% to form the final feature set;
[0066] Step 4: Support Vector Machine Model Construction and Training: Divide the dataset into a training set (70%) and a validation set (30%) according to a certain proportion. Adopt the radial basis function kernel, optimize the penalty parameter and the kernel parameter through grid search, input the final feature set and the measured soil salt value, and train the support vector machine regression model;
[0067] Step 5: Salt Inversion and Accuracy Verification: Apply the SVM model to Sentinel-2 images and fusion index data, output the predicted soil salt value, generate a spatial distribution map of the cultivated land soil salt in the Hetao Irrigation Area according to the predicted soil salt value, and divide the cultivated land salinization level in combination with the inversion results;
[0068] Step 6: Dynamic Monitoring: Obtain Sentinel-2 images quarterly, repeat Steps 1 to 5, generate a dynamic change map of salinization, evaluate the salt accumulation rate and the effect of treatment measures in combination with autumn irrigation records and crop growth conditions, formulate a differentiated salt leaching irrigation plan according to the salinization level (mild / moderate / severe), and recommend salt-tolerant crops (such as sunflowers) or areas for applying soil conditioners in combination with planting structure data.
[0069] In one embodiment, in step one, by using the grid sampling method in the cultivated land of the Hetao Irrigation District, combined with road, soil type and planting structure adjustment, surface soil samples of 0 cm - 20 cm are collected. At the same time, a SR-3500 ground object spectrometer is used to measure the hyperspectral reflectance in the range of 350 nm - 2500 nm, perform spectral transformation, generate a fused spectral index. Before each measurement, whiteboard calibration is carried out, and the average value of three measurements is taken. The quartile method is used to identify and eliminate outliers in the soil salinity samples, reducing the influence of extreme values on the inversion result. The soil samples are tested to obtain soil salinity data;
[0070] The spectral transformation includes first derivative (R'), second derivative (R''), reciprocal (1 / R), reciprocal first derivative ((1 / R)'), reciprocal second derivative ((1 / R)''), logarithm (lgR), logarithm first derivative ((lgR)'), logarithm second derivative ((lgR)''), square root (R^0.5), square root first derivative ((R^0.5)') and square root second derivative ((R^0.5)'');
[0071] The crop planting data is based on the Sentinel-2 satellite images to construct an NDVI (Normalized Difference Vegetation Index) time series, combined with the Savitzky-Golay filter and the decision tree hierarchical classification model to extract the crop planting structure (classification accuracy 92.1%);
[0072] The autumn irrigation information extracts the irrigation area by using the MNDWI (Modified Normalized Difference Water Index), and uses the threshold method to distinguish between irrigated and non-irrigated plots. The MNDWI (Modified Normalized Difference Water Index) is expressed as where G represents the reflectance of the Green Band, corresponding to the spectral characteristics of vegetation and water bodies, and MIR represents the reflectance of the Middle Infrared Band, which is sensitive to water bodies and is used to distinguish water bodies from dry land surfaces.
[0073] In one embodiment, in step one, for the multispectral data, Sentinel-2 MSI images are obtained, the Normalized Difference Vegetation Index (NDVI), Enhanced Vegetation Index (EVI), and Canopy Salinity Response Index (CRSI) vegetation indices are calculated, and the maximum / minimum values of the vegetation indices from June to September are extracted to eliminate cloud interference;
[0074] For the land use data, an ESRI 10 m resolution land use map is used to screen the "crop" pixels to define the cultivated land range;
[0075] By integrating Copernicus DEM data, surface elevation information is obtained.
[0076] In one embodiment, the calculation formulas of the vegetation indices, namely the Normalized Difference Vegetation Index (NDVI), the Enhanced Vegetation Index (EVI), and the Canopy Salt Response Index (CRSI), are expressed as
[0077] NDVI = (NIR - R) / (NIR + R);
[0078] EVI = G × (NIR - R) / (NIR + 6R + 7.5B + 1);
[0079]
[0080] maxNDVI = max(NDVIt1, NDVIt2,..., NDVItn);
[0081] minNDVI = min(NDVIt1, NDVIt2,..., NDVItn);
[0082] wherein, NIR represents the reflectance of the Near Infrared Band, R represents the reflectance of the Red Band, G represents the reflectance of the Green Band, B represents the reflectance of the Blue Band, NDVI t1 , NDVI t2 , …, NDVI tn represents the NDVI values at different time points, maxNDVI represents the NDVI of the crop in the optimal growth state, reflecting the vegetation potential in the less saline areas, and minNDVI represents the NDVI of the crop under the most severe stress, directly related to the degree of salinization. The larger the difference between the two (maxNDVI - minNDVI), the more significant the seasonal impact of salinization on crop growth.
[0083] In one embodiment, in step two, the salt index (such as derivative, logarithm, etc.) after the transformation of the ground hyperspectral and the salt index calculated from the Sentinel-2 multispectral data of the same period are fused through simple linear regression, and the satellite multispectral index and the hyperspectral index are fused to generate a hyperspectral-multispectral fusion index dataset, including the following steps:
[0084] Perform a simple linear regression analysis on the salt index after the transformation of the ground hyperspectral and the salt index calculated from the Sentinel-2 multispectral data. The salt index is expressed as
[0085]
[0086] Among them, BI represents the brightness index, which reflects the brightness characteristics of the soil surface. This index can effectively capture the brightness differences in salinized areas. R represents the reflectance of the red band, corresponding to the spectral characteristics of minerals such as iron oxide in the soil. NIR represents the reflectance of the near-infrared band, which is related to soil moisture and organic matter content. S3 aims to enhance the sensitivity of salts to the green-red band combination, and at the same time suppress the influence of atmospheric scattering or light changes through the blue band, and is applicable to the identification of salt frost in arid areas. G represents the reflectance of the green band, which is related to vegetation cover and some salt characteristics. B represents the reflectance of the blue band, which is used to detect salt frost on the soil surface. S5 utilizes the high sensitivity of the blue and red bands to salts, combined with green band normalization, to reduce the interference of vegetation cover, and is applicable to shallow salinization monitoring. S6 represents the salinity index 6, which can separate the mixed spectral signals of salts and vegetation, and is applicable to complex scenarios where vegetation cover and salinization coexist. SI2 can comprehensively characterize the spatial distribution of soil salts and is applicable to large-scale dynamic monitoring of salinization. SI3 targets the color changes caused by surface salts (such as salt frost whitening or crust darkening) for rapid identification of salinized patches.
[0087] A regression equation is established, taking the salinity index calculated from Sentinel-2 multispectral data as the independent variable and the salinity index after ground hyperspectral transformation as the dependent variable. The regression equation is expressed as MNSI = a * MNSI s2 + b, where MNSI represents the salinity index after ground hyperspectral transformation, and MNSI s2 represents the salinity index calculated from Sentinel-2 multispectral data. a and b are regression coefficients, which are solved by the least squares method.
[0088] Through the regression equation, the Sentinel-2 multispectral data is converted into a fusion index corresponding to the ground hyperspectral data.
[0089] Combining the fusion index obtained by unary linear regression with the spatial position information of the Sentinel-2 multispectral image, a high-multispectral fusion index dataset is generated.
[0090] In one embodiment, in step three, screening out the key features from the feature variable library includes the following steps:
[0091] Random forest model training: Using soil salinity data as the dependent variable and the feature variables (72 salinity indices, 6 vegetation indices, 3 environmental variables) inside the feature variable library as the independent variables, a random forest model is constructed. The parameters of the random forest model are set to generate 200 decision trees, and 10-fold cross-validation is used to evaluate the performance.
[0092] Feature importance calculation: The Gini importance is used to evaluate the feature importance. In each decision tree, the split of a feature is measured by the reduction in Gini impurity, and the formula is
[0093]
[0094] where Gini parent , Gini left , Gini right are the Gini impurities of the parent node, the left child node, and the right child node respectively, and N parent , N left , N right are the number of samples of the corresponding nodes respectively;
[0095] The final importance of each feature is the sum of the reduction in Gini impurity in all trees and is normalized, and the normalization is expressed as where T is the total number of decision trees and M is the total number of features;
[0096] Feature sorting and cumulative screening: The features are sorted from high to low according to importance to obtain the sequence {f1, f2,..., f N}}, and the cumulative importance is calculated where N represents the number of features;
[0097] Select the smallest feature subset {f1, f2,..., f k} that satisfies CumulativeImportance k ≥ 85% as the final feature set;
[0098] Output of key variables: The screened feature set contains the key variables that contribute the most to soil salinity inversion, which not only reduces data redundancy but also retains more than 85% of the effective information and optimizes the generalization ability of the model.
[0099] In one embodiment, in step four, the specific steps of training the support vector machine regression model are as follows:
[0100] Data division: The screened feature set and the dataset of measured soil salinity values are divided into a training set and a validation set in a ratio of 7:3. Stratified random sampling is used during the division to ensure that the distribution of salinization levels in the training set and the validation set is consistent and avoid data deviation;
[0101] Data standardization: The input data is processed using Z-score standardization, and the processing is expressed as where μ is the feature mean and σ is the standard deviation, ensuring that the mean of each feature is 0 and the variance is 1;
[0102] Model initialization: Select the support vector regression model with the radial basis function as the kernel function, and its expression is K(x i ,x j ) = exp(-γ||x i -x j || 2 ). The objective function of the model is where C is the penalty parameter that controls the balance between model complexity and training error; γ is the kernel function parameter that determines the distribution of data mapped to the high-dimensional space; ξ i is the slack variable that allows some samples to deviate from the hyperplane;
[0103] Parameter optimization: Optimize the hyperparameters C and γ through grid search combined with 5-fold cross-validation. Define the parameter grid C ∈ {0.1, 1, 10, 100}, γ ∈ {0.001, 0.01, 0.1, 1}, traverse the parameter combinations, and use the mean squared error of the validation set as the evaluation index. The mean squared error is expressed as Select the parameter combination that minimizes the RMSE as the optimal solution;
[0104] Model training: Input the standardized training set data into the support vector machine regression model and train it using the optimal parameters to fit the relationship between the feature variables and soil salinity;
[0105] Model validation: Use the validation set to evaluate the model performance, calculate the coefficient of determination and the mean squared error. The calculation formula of the coefficient of determination is expressed as
[0106] The experimental results show that for the validation set, R^2 = 0.64, RMSE = 5.41, and the generalization ability of the model is better than that of the random forest (RF) and partial least squares regression (PLSR);
[0107] Result interpretation: SVM captures the non-linear relationship between multi-source data (spectral index, vegetation index, environmental variables) and soil salinity through maximizing the margin and kernel trick. Its lower overfitting risk (compared with RF) and higher interpretability (compared with PLSR) make it suitable for the salt-affected inversion task in the complex environment of the Hetao Irrigation Area.
[0108] In one embodiment, in step five, the steps of generating the spatial distribution map of the cultivated land soil salinity in the Hetao Irrigation Area according to the predicted soil salinity value and dividing the cultivated land salinization grade in combination with the inversion result are as follows:
[0109] Data preprocessing: Perform overlay analysis on the raster data of the predicted soil salinity value output by the SVM model and the vector data of the cultivated land scope in the Hetao Irrigation Area, and extract the predicted soil salinity value within the cultivated land scope;
[0110] Spatial distribution map drawing: Use GIS software (such as ArcGIS) to convert the raster data of the predicted values of arable land soil salinity into a spatial distribution map. According to the magnitude of the predicted values of soil salinity, different colors or symbols are used to represent different salinity levels;
[0111] Salinization level classification standard: Refer to the soil salinization classification method, and divide the predicted values of soil salinity into five levels: non-salinized, slightly salinized, moderately salinized, severely salinized, and saline soil;
[0112] Implementation of level classification: According to the magnitude of the predicted values of soil salinity, each arable land grid unit is classified into the corresponding salinization level. The level classification can be achieved using the reclassification tool or custom script in GIS software.
[0113] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A method for monitoring salinization of cultivated land based on spectral analysis, characterized in that: The following steps are involved: Step 1: Data preparation: Collect and process data covering the target area, including ground data and satellite data. The ground data includes soil samples, hyperspectral data, crop planting data, and autumn irrigation information. The satellite data includes multispectral data, land use data, and elevation data. Step 2: Spectral index calculation and fusion: The salt index after ground hyperspectral transformation and the salt index calculated by Sentinel-2 multispectral in the same period are fused through univariate linear regression, and the satellite multispectral index is fused with the hyperspectral index to generate a hyper-multispectral fusion index dataset. Based on the transformed data, a feature variable library is constructed, which includes salt index, vegetation index and environmental variables. Step 3 : Feature variable selection: select key features from the feature variable library, evaluate feature importance based on the random forest model, calculate the relative importance of each feature, and form the final feature set; Step 4: Support vector machine model construction and training: Divide the data set into training set and validation set in proportion, use radial basis function kernel, optimize penalty parameters and kernel parameters through grid search, input the final feature set and measured soil salinity value, and train the support vector machine regression model; Step 5: Salinity inversion and accuracy verification: Apply the SVM model to Sentinel-2 images and fusion index data, output the soil salinity prediction value, generate the spatial distribution map of soil salinity in the Hetao irrigation area according to the soil salinity prediction value, and divide the salinization level of cultivated land based on the inversion results; Step 6: Dynamic monitoring: Acquire Sentinel-2 images quarterly, repeat steps 1 to 5, generate a dynamic change map of salinization, combine autumn irrigation records with crop growth, evaluate the salt accumulation rate and the effectiveness of control measures, and formulate a salt irrigation plan based on the salinization level.
2. The method for monitoring salinization of cultivated land based on spectral analysis according to claim 1, characterized in that: In step 1, the grid point method was used in the cultivated land of the Hetao Irrigation Area, and combined with the adjustment of roads, soil types and planting structures, 0cm-20cm surface soil samples were collected, and the SR-3500 ground feature spectrometer was used to measure the high spectral reflectance in the range of 350nm-2500nm, perform spectral transformation, generate fused spectral index, test soil samples, and obtain soil salinity data; The crop planting data is constructed based on Sentinel-2 satellite images to construct NDVI time series, and the crop planting structure is extracted by combining Savitzky-Golay filtering and decision tree hierarchical classification model; The autumn irrigation information is extracted by using the MNDWI index to extract the irrigation area, and the threshold method is used to distinguish between irrigated and non-irrigated plots. The MNDWI index is expressed as Among them, G represents the reflectance of the green band, which corresponds to the spectral characteristics of vegetation and water bodies; MIR represents the reflectance of the mid-infrared band, which is sensitive to water bodies and is used to distinguish water bodies from dry surfaces.
3. The method for monitoring salinization of cultivated land based on spectral analysis according to claim 1, characterized in that: In step 1, the multispectral data, Sentinel-2MSI imagery is acquired, and the normalized vegetation index, enhanced vegetation index, and canopy salt response index vegetation index are calculated; The land use data adopts the land use map and selects the "crop" pixels to define the scope of cultivated land; The surface elevation information is obtained by integrating Copernicus DEM data.
4. The method for monitoring salinization of cultivated land based on spectral analysis according to claim 3, characterized in that: The calculation formulas of the normalized difference vegetation index (NDVI), enhanced vegetation index (EVI), and canopy salt response index (CRSI) vegetation index are expressed as follows: NDVI = (NIR-R) / (NIR+R); EVI=G×(NIR-R) / (NIR+6R+7.5B+1); maxNDVI=max(NDVIt1,NDVIt2,…,NDVItn); minNDVI = min(NDVIt1, NDVIt2, ..., NDVItn); Among them, NIR represents the reflectivity of the near-infrared band, R represents the reflectivity of the red band, G represents the reflectivity of the green band, B represents the reflectivity of the blue band, NDVIt1, NDVIt2, …, NDVItn represent the NDVI values at different time points, maxNDVI represents the NDVI of crops under the best growth state, reflecting the vegetation potential in areas with less salinization, and minNDVI represents the NDVI of crops when the stress is the most serious, which is directly related to the degree of salinization.
5. The method for monitoring salinization of cultivated land based on spectral analysis according to claim 1, characterized in that: In step 2, the salt index after ground hyperspectral transformation and the salt index calculated by Sentinel-2 multispectral in the same period are fused through univariate linear regression, and the satellite multispectral index is fused with the hyperspectral index to generate a hyper-multispectral fusion index data set, including the following steps: A linear regression analysis was performed on the salt index after ground hyperspectral transformation and the salt index calculated by Sentinel-2 multispectral. The salt index is expressed as Among them, BI represents the brightness index, which reflects the brightness characteristics of the soil surface; R represents the reflectance of the red band, which corresponds to the spectral characteristics of minerals in the soil; NIR represents the reflectance of the near-infrared band, which is related to the soil moisture and organic matter content; S3 is suitable for the identification of salt frost in arid areas; G represents the reflectance of the green band, which is related to vegetation coverage and partial salt characteristics; B represents the reflectance of the blue band, which is used to detect salt frost on the soil surface; S5 is suitable for shallow salinization monitoring; S6 represents salt index 6; SI2 is suitable for dynamic monitoring of large-scale salinization; SI3 is used to quickly identify salinization patches; A regression equation is established, with the salt index calculated by Sentinel-2 multispectral as the independent variable and the salt index after ground hyperspectral transformation as the dependent variable. The regression equation is expressed as MNSI = a*MNSI s2 +b, where MNSI represents the salt index after ground hyperspectral transformation, MNSI s2 represents the salt index calculated by Sentinel-2 multispectral, a and b are regression coefficients, solved by the least squares method; Through the regression equation, the Sentinel-2 multispectral data are converted into a fusion index corresponding to the ground hyperspectral data; The fusion index obtained by univariate linear regression is combined with the spatial position information of Sentinel-2 multispectral images to generate a hyper-multispectral fusion index dataset.
6. The method for monitoring salinization of cultivated land based on spectral analysis according to claim 1, characterized in that: In step three, the key features are screened out from the feature variable library, including the following steps: Random forest model training: soil salinity data is used as the dependent variable, and the characteristic variables in the characteristic variable library are used as independent variables to construct a random forest model. The parameters of the random forest model are set to generate 200 decision trees, and 10-fold cross validation is used to evaluate the performance; Feature importance calculation: Gini importance is used to evaluate feature importance. In each decision tree, the split of the feature is measured by the reduction of Gini impurity. The formula is Among them Gini parent ,Gini left ,Gini right are the Gini impurities of the parent node, left child node, and right child node, respectively, N parent ,N left ,N right are the number of samples corresponding to the nodes respectively; The final importance of each feature is the sum of its Gini impurity reduction in all trees, and is normalized, which is expressed as Among them, T is the total number of decision trees, and M is the total number of features; Feature sorting and cumulative screening: Sort the features from high to low in importance to get the sequence {f1,f2,...,f N }, calculate the cumulative importance Where N represents the number of features; Select the one that satisfies Cumulative Importance k ≥85% of the minimum feature subset {f1,f2,...,f k } as the final feature set; Output key variables: The filtered feature set contains the key variables that contribute most to soil salinity inversion.
7. The method for monitoring salinization of cultivated land based on spectral analysis according to claim 1, characterized in that: In step 4, the specific steps of training the support vector machine regression model are: Data division: The screened feature set and the data set of measured soil salinity values were divided into a training set and a validation set in a ratio of 7:3, using stratified random sampling; Data standardization: The input data is processed using Z-score standardization, which is expressed as Among them, μ is the characteristic mean and σ is the standard deviation; Model initialization: Select the support vector regression model, and the kernel function uses the radial basis function, which is expressed as K(x i ,x j ) = exp(-γ||x i -x j || 2 ), the model objective function is Among them, C is the penalty parameter, which controls the balance between model complexity and training error; γ is the kernel function parameter, which determines the distribution of data mapped to high-dimensional space; ξ i It is a slack variable, which allows some samples to deviate from the hyperplane; Parameter optimization: Optimize the hyperparameters C and γ through grid search combined with 5-fold cross validation, define the parameter grid C∈{0.1,1,10,100}, γ∈{0.001,0.01,0.1,1}, traverse the parameter combination, and use the mean square error of the validation set as the evaluation indicator. The mean square error is expressed as Select the parameter combination that minimizes RMSE as the optimal solution; Model training: The standardized training set data is input into the support vector machine regression model, and the optimal parameters are used for training to fit the relationship between the characteristic variables and soil salinity; Model validation: Use the validation set to evaluate the model performance and calculate the coefficient of determination and mean square error. The coefficient of determination calculation formula is expressed as Result interpretation: SVM captures the nonlinear relationship between multi-source data and soil salinity by maximizing the margin and kernel technique, making it suitable for salinization inversion tasks in the complex environment of the Hetao irrigation area.
8. The method for monitoring salinization of cultivated land based on spectral analysis according to claim 1, characterized in that: In step 5, the spatial distribution map of soil salinity in the cultivated land of the Hetao irrigation area is generated according to the predicted value of soil salinity, and the salinization grade of the cultivated land is divided according to the inversion result. The steps are as follows: Data preprocessing: The soil salinity prediction value raster data output by the SVM model is overlaid with the cultivated land range vector data of the Hetao Irrigation District to extract the soil salinity prediction value within the cultivated land range; Spatial distribution map drawing: The extracted raster data of cultivated land soil salinity prediction values are converted into a spatial distribution map. Different colors or symbols are used to represent different salinity levels according to the size of the soil salinity prediction values. Salinization classification standard: The predicted soil salinity value is divided into five levels: non-salinization, mild salinization, moderate salinization, severe salinization and saline soil; Implementation of grade classification: According to the predicted value of soil salinity, each cultivated land grid unit is divided into the corresponding salinization grade.
Citation Information
Patent Citations
Hyperspectral remote sensing judgment method for degree of soil salinization
CN109738380A
Inversion method for soil salinity of Yellow River delta based on Landsat8
CN111783288A
Method for random forest inversion of cultivated land salinization based on GEE cloud platform
CN118095499A
Early warning method and system for soil secondary salinization degree
CN119168831A
Method for improving satellite data inversion soil salinity by using unmanned aerial vehicle data
CN119314035A
Cited By
Salinization remote sensing inversion method fusing crop information and data preprocessing
CN120974081A
Saline-alkali cultivated land salt spot extraction method and system based on multispectral unmanned aerial vehicle
CN121453693A
A method and system for extracting salt marks in saline-alkali farmland based on a multispectral unmanned aerial vehicle
CN121453693B