Bare soil salinity inversion method fusing partial least squares and random forest

The integration of PLS and RF models enhances soil salinity inversion accuracy and reliability, addressing the limitations of existing methods by improving prediction performance in various geographical conditions.

CN120318700APending Publication Date: 2025-07-15NANJING TECH UNIV

Patent Information

Application Number
CN202510204173.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-24
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

The existing soil salt inversion methods have problems with insufficient accuracy and poor generalization capabilities, especially in remote sensing data processing, which is difficult to effectively deal with nonlinear relationships and noise.

Method used

Using the method of fusion partial least squares and random forest, the spectral reflectivity and salinity spectral index were obtained through Sentinel-2 satellite data, a fusion model containing 24-dimensional characteristic variables was established, and a linear regression and random forest model were combined for training and evaluation, and the bare soil salt content was obtained.

Benefits of technology

It significantly improves the accuracy and stability of soil salt inversion, can better handle nonlinear relationships and noise in remote sensing data, and improves the generalization ability of salt inversion, especially in low-salt and medium-salt areas.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120318700A_ABST
    Figure CN120318700A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of remote sensing monitoring, solves the technical problems that an existing soil salinity inversion method is insufficient in precision and poor in generalization ability, and particularly relates to a bare soil salinity inversion method fusing partial least squares and a random forest. Comprising the following steps: acquiring spectral reflectivity data and salinity spectral index of a multiband range and spatial resolution, and preprocessing to establish a digital orthoimage; removing a water body and a vegetation coverage area in the digital orthoimage by using a normalized vegetation index to obtain bare soil data containing 24-dimensional feature variables; establishing a fusion model and carrying out training evaluation; and performing prediction by taking Sentinel-2 satellite remote sensing data as input of the fusion model to obtain the salt content of the bare soil. According to the method, the precision, reliability and generalization ability of soil salinity inversion can be effectively improved, compared with a traditional single model, the nonlinear relation and noise in remote sensing data can be better processed, and the generalization ability of salinity inversion is remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of remote sensing monitoring, and particularly relates to a method for retrieving bare soil salinity by fusing partial least squares and random forest. Background Art

[0002] Soil salinization is an important global environmental problem that seriously affects agricultural production and the ecological environment. Traditional soil salinity monitoring methods rely on ground sampling and chemical analysis, which are inefficient and costly. With the development of remote sensing technology, satellite images have become the main means of soil salinity monitoring. In particular, the Sentinel-2 satellite has become an ideal data source for large-scale soil salinity inversion due to its advantages such as high spatio-temporal resolution and easy access. However, existing soil salinity inversion methods such as multiple linear regression and BP neural network still have problems such as insufficient accuracy and poor generalization ability. Therefore, how to improve the accuracy and stability of soil salinity inversion has become an urgent technical problem to be solved. Summary of the Invention

[0003] Aiming at the deficiencies of the prior art, the present invention provides a method for retrieving bare soil salinity by fusing partial least squares and random forest, which solves the technical problems of insufficient accuracy and poor generalization ability existing in the existing soil salinity inversion methods.

[0004] To solve the above technical problems, the present invention provides the following technical solution: A method for retrieving bare soil salinity by fusing partial least squares and random forest, the method comprising the following steps:

[0005] S1. Obtain spectral reflectance data and salinity spectral indices within a multi-band range and spatial resolution from Sentinel-2 satellite remote sensing data, and perform preprocessing to establish a digital orthophoto image containing original multi-dimensional feature data;

[0006] S2. Use the normalized difference vegetation index to remove water bodies and vegetation-covered areas in the digital orthophoto image to obtain bare soil data containing 24-dimensional feature variables;

[0007] S3. Establish a fusion model and perform training and evaluation. The fusion model includes a linear regression model for inputting 24-dimensional feature variables and outputting an initial prediction result, and a random forest model for outputting the surface soil salt content according to 25-dimensional feature variables;

[0008] S4. Use Sentinel-2 satellite remote sensing data as the input of the fusion model for prediction to obtain the bare soil salt content;

[0009] S5. Use a single machine learning model to perform predictions respectively with 24-dimensional feature variables, and compare with the predicted bare soil salt content to evaluate the performance of the fusion model.

[0010] Further, in step S2, specifically, water bodies with NDVI < 0 and vegetation-covered areas with NDVI > 0.15 in the digital orthophoto image are removed. The calculation formula for the normalized difference vegetation index NDVI is:

[0011]

[0012] where NDVI represents the normalized difference vegetation index; NIR and R represent the reflectance of the near-infrared band and the red band, respectively.

[0013] Further, the expression of the fusion model is:

[0014]

[0015] where y is the dependent variable; ρ i is the explanatory variable of the i-th variable; w i is the model coefficient of the i-th variable; w0 is the constant term; n is the number of variables; is the final prediction result; T i (x) is the prediction result of the i-th decision tree; N is the number of decision trees.

[0016] Further, in step S3, the specific process of training and evaluating the fusion model includes the following steps:

[0017] S31. Use the original multi-dimensional feature data as the input of the linear regression model, and perform regression analysis on the original multi-dimensional feature data to obtain the initial prediction result;

[0018] S32. Combine the initial prediction result with the original multi-dimensional feature data to form a 25-dimensional feature variable, and use the 25-dimensional feature variable as the input of the random forest model to obtain the surface soil salt content as the predicted value;

[0019] S33. Obtain the soil salt content information of the on-site sampling as the actual value;

[0020] S34. Use the true value and the predicted value to evaluate the fusion model until convergence.

[0021] Further, in step S33, the specific process includes:

[0022] At multiple sampling points, soil samples with a depth of 0 - 20 cm are collected respectively. After the soil samples are dried and ground, a soil extract is prepared in a ratio of 1:5 and left to stand for 24 h. Then, its conductivity is measured, and the soil salt content SSC of the soil sample is calculated according to the conductivity. The calculation formula for the soil salt content SSC is:

[0023] SSC = (0.2882 × EC + 0.0183) × 100%

[0024] Among them, EC is the electrical conductivity.

[0025] Further, in step S34, the root mean square error RMS, the coefficient of determination R 2 , the mean bias Bias, and the standard deviation STD are used as the evaluation indicators for the fusion model, and the calculation formulas are as follows:

[0026]

[0027] Among them, y i is the actual value; is the predicted value; n is the number of soil samples; is the mean value of the actual values; x i is the data point; is the mean value of the data points.

[0028] Further, the number of decision trees n of the random forest model is 20, and the minimum number of leaves is 8.

[0029] By means of the above technical solution, the present invention provides a bare soil salt inversion method that combines partial least squares and random forest, and at least has the following beneficial effects:

[0030] 1. The soil salt inversion method proposed by the present invention can effectively improve the accuracy, reliability, and generalization ability of soil salt inversion. Compared with traditional single models, it can better handle the non-linear relationships and noises in remote sensing data, and significantly improve the generalization ability of salt inversion.

[0031] 2. The prediction accuracy of the fusion model proposed by the present invention in different salt ranges is significantly better than that of traditional single models, especially excellent in the low-salt and medium-salt regions, and can be widely applied to soil salt monitoring in different geographical environments.

[0032] 3. The soil salt inversion method proposed by the present invention is applicable to the monitoring and management of soil salinization in large-scale regions, especially in regions with serious soil salinization such as coastal beaches and arid and semi-arid regions. In addition, it can also be popularized and applied to fields such as agricultural production, land management, and environmental protection, providing a scientific basis for land improvement and ecological restoration. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] The drawings described herein are used to provide a further understanding of the present application, and constitute a part of the present application. The schematic embodiments and descriptions thereof of the present application are used to explain the present application and do not constitute an improper limitation of the present application. In the drawings:

[0034] Figure 1 is the flow chart of the bare soil salt inversion method in the present invention;

[0035] Figure 2This is a comparison chart of the results of the PLS model in the training set and the test set in the present invention;

[0036] Figure 3 This is a comparison chart of the results of the RF model in the training set and the test set in the present invention;

[0037] Figure 4 This is a comparison chart of the results of the BP model in the training set and the test set in the present invention;

[0038] Figure 5 This is a comparison chart of the results of the fusion model in the training set and the test set in the present invention. Detailed implementation manners

[0039] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific implementation manners. Thereby, the implementation process of how the present application uses technical means to solve technical problems and achieve technical effects can be fully understood and implemented accordingly.

[0040] In order to use the fusion model to predict the bare soil salt content through Sentinel-2 satellite data and thereby realize the inversion of soil salinity in this embodiment, please refer to Figures 1 - 5 , this embodiment proposes a method for inverting bare soil salinity by fusing partial least squares and random forest, which effectively improves the accuracy and reliability of soil salinity inversion. Compared with traditional single models, it can better handle the non-linear relationships and noises in remote sensing data and significantly enhances the generalization ability of salinity inversion. As Figure 1 shown, this method includes the following steps:

[0041] S1. Obtain spectral reflectance data and salinity spectral indices with multi-band ranges and spatial resolutions from Sentinel-2 satellite remote sensing data, and perform preprocessing to establish a digital orthophoto image containing original multi-dimensional feature data.

[0042] In this embodiment, considering the collection time of soil samples, Sentinel-2 multispectral data on December 24, 2023, and January 1, 2024 (https: / / dataspace.copernicus.eu / ) is used, and its band ranges and spatial resolutions are shown in Table 1.

[0043] Table 1 Band ranges and spatial resolutions of the Sentinel-2 satellite remote sensing data used

[0044] Physical Band Central Wavelength (nm) Pixel Resolution (m) B2 490 10 B3 560 10 B4 665 10 B5 705 20 B6 740 20 B7 783 20 B8 842 10 B8A 865 20 B9 945 60 B11 1610 20 B12 2190 20

[0045] To ensure data quality and analysis accuracy, in the data preprocessing stage, first, bilinear interpolation is used to resample all band data to a spatial resolution of 10 m; then, using the radiometric correction coefficients provided by Sentinel-2, the image data is converted into standardized physical unit reflectance, and atmospheric correction is performed through the Sen2Cor tool; finally, after processing such as registration, fusion, and mosaicking, a digital orthophoto map (DOM) is generated.

[0046] As an effective indicator for capturing soil salinity information, the salinity spectral index is also commonly used for salinity inversion. Therefore, in addition to the above 11 spectral band data, this embodiment also selects common spectral indices as the input of the modeling data to improve the inversion accuracy of the model. The calculation formulas of each spectral index are shown in Table 2.

[0047] Sentinel-2 satellite salinity spectral indices and their calculation formulas used in Table 2

[0048]

[0049] In Table 2, B, G, R, NIR, SWIR-1, and SWIR-2 represent the reflectances of different bands in Sentinel-2, specifically as follows: B = B2 (490 nm), G = B3 (560 nm), R = B4 (665 nm), NIR = B8 (842 nm), SWIR-1 = B11 (1610 nm), SWIR-2 = B12 (2190 nm).

[0050] S2. Use the normalized difference vegetation index to remove water bodies and vegetated areas in the digital orthophoto map to obtain bare soil data containing 24-dimensional feature variables. In this embodiment, based on the normalized difference vegetation index (NDVI), water bodies (NDVI < 0) and areas with relatively dense vegetation cover (NDVI > 0.15) within the study area are removed. The calculation formula of NDVI is:

[0051]

[0052] where NDVI represents the normalized difference vegetation index; NIR and R represent the reflectances of the near-infrared band (B8, 842 nm) and the red band (B4, 665 nm), respectively.

[0053] S3. Establish a fusion model and conduct training and evaluation. The fusion model includes a linear regression model for inputting 24-dimensional feature variables and outputting initial prediction results, and a random forest model for outputting the surface soil salt content based on 25-dimensional feature variables. The fusion model proposed in this embodiment is a combination of a linear regression model based on partial least squares (PLS) and a random forest model (RF), where the number of decision trees n selected for the random forest model (RF) is 20, and the minimum number of leaves is 8. Finally, the fusion model (PLS-RF) is obtained through training.

[0054] Based on Sentinel-2 satellite remote sensing data in this embodiment, spectral reflectance data of 11 bands are preprocessed and extracted. At the same time, 13 salinity spectral indices sensitive to salinity are calculated. The 24-dimensional feature variables of 24 independent variables are used as the input of the linear regression model, and the surface soil salt content is used as the output of the linear regression model to form a multi-dimensional feature dataset. The multi-dimensional feature dataset is randomly divided into a training set and a test set in a ratio of 6:4 for subsequent training evaluation and performance verification of the random forest model.

[0055] PLS regression establishes a regression model by projecting the dependent variable and explanatory variables, and can explain the change of the response variable by finding the linear combination with the largest covariance. It is a statistical method that combines principal component analysis and multiple linear regression, suitable for dealing with multi-collinearity problems and high-dimensional data, and can effectively extract useful soil salinity information from hyperspectral data. The expression of the linear regression model is:

[0056]

[0057] where y is the dependent variable; ρ i is the explanatory variable of the i-th variable; w i is the model coefficient of the i-th variable; w0 is the constant term; n is the number of variables.

[0058] RF is a method of ensemble learning. By randomly sampling data and features, multiple decision trees are constructed and their prediction results are combined for classification or regression. RF performs excellently in dealing with high-dimensional data and has strong generalization ability. Through calculation, in this embodiment, the number of decision trees n is 20, and the minimum number of leaves is 8. The expression of the random forest model is:

[0059]

[0060] where, is the final prediction result; T i (x) is the prediction result of the i-th decision tree; N is the number of decision trees.

[0061] In this embodiment, two models are trained using different inputs and finally a fusion model is obtained. The specific process includes the following steps:

[0062] S31. Use the original multi-dimensional feature data as the input of the linear regression model, and perform regression analysis on the original multi-dimensional feature data to obtain the initial prediction result;

[0063] S32. Combine the initial prediction result and the original multi-dimensional feature data to form a 25-dimensional feature variable, and use the 25-dimensional feature variable as the input of the random forest model to obtain the surface soil salt content as the predicted value;

[0064] S33. Obtain the soil salt content information of on-site sampling as the actual value. The specific process includes: collecting soil samples at multiple sampling points with a depth of 0-20 cm, drying and grinding the soil samples, preparing a soil extract at a ratio of 1:5 and standing for 24 h, measuring its conductivity, and calculating the soil salinity SSC of the soil sample according to the conductivity.

[0065] In this embodiment, soil samples with a depth of 0-20 cm at 110 sampling points were collected on December 24, 2023 and January 1, 2024 respectively. The sampling points are evenly distributed in the study area to make it globally representative in the study area. After the soil samples are dried and ground, a soil extract is prepared at a ratio of 1:5 (mass unit: g) and left standing for 24 h, and its conductivity (EC, Ms / cm) is measured. The soil salinity SSC (%) is calculated using the formula. The calculation formula of the soil salinity SSC is:

[0066] SSC = (0.2882 × EC + 0.0183) × 100%

[0067] where EC is the conductivity.

[0068] S34. Use the true value and the predicted value to evaluate the fusion model until convergence. In this embodiment, the root mean square error RMS, the coefficient of determination R 2 , the mean bias Bias, and the standard deviation STD are used as the evaluation indicators of the fusion model. Specifically, the root mean square error RMS represents the accuracy of the fusion model; the coefficient of determination R 2 reflects the fitting degree of the fusion model; the mean bias Bias represents the systematic bias and is used to evaluate the overall deviation degree of the fusion model; the standard deviation STD represents the random error and is used to measure the stability and reliability of the fusion model. The calculation formulas are as follows:

[0069]

[0070] where y i is the actual value; is the predicted value; n is the number of soil samples; is the mean of the actual values; x i is a data point; is the mean of the data points.

[0071] In this embodiment, a linear regression model is first used to perform a regression analysis on the original 24-dimensional feature variables to obtain an initial prediction result. Subsequently, the initial prediction result and the original 24-dimensional feature variables are jointly used to form 25-dimensional feature variables, which are used as the input of the random forest model. The output of the random forest model is still the surface soil salt content. A fusion model is obtained through training and evaluation. The fusion model retains the key information of the original data through PLS, and the powerful modeling ability of the random forest model can effectively process multi-dimensional data and noise, mine potential laws, improve the overall prediction accuracy and generalization ability of the fusion model, and help to better capture complex soil salt information.

[0072] S4. Use the Sentinel-2 satellite remote sensing data as the input of the fusion model to predict the bare soil salt content. Here, the Sentinel-2 satellite remote sensing data can be directly used to predict the bare soil salt content, and no more details will be elaborated here.

[0073] S5. Use a single BP neural network or a random forest network model to perform predictions respectively using 24-dimensional feature variables, and compare with the predicted bare soil salt content to evaluate the performance of the fusion model.

[0074] As a further explanation, the BP neural network is a multi-layer feedforward neural network, which is trained through the backpropagation algorithm. Optimizing the network parameters (weights and biases) minimizes the global error. In theory, the BP neural network can approximate any complex function, has a powerful non-linear fitting ability, and is suitable for various tasks such as regression and classification. In this embodiment, through trial calculations, the error target of the BP neural network is determined to be 1e-6, the learning rate is 0.001, the number of hidden layer nodes is 6, and the activation function is tansig. The expression of the output of the BP neural network is:

[0075]

[0076] where y is the final prediction result; x i is the input feature; f is the activation function; w i is the weight; b is the bias.

[0077] The results of the conventional linear regression model (PLS) in the training set and the test set in this embodiment are shown in Table 3:

[0078] Table 3 PLS Modeling and Test Results

[0079] <![CDATA[R 2 > Bias RMS STD Training Set 0.515 0 0.459 0.459 Test Set -0.220 0.017 0.636 0.635

[0080] As can be seen from Table 3, the R of the model on the training set 2 is 0.515, indicating that the fitting effect of PLS on the training set is average and it cannot effectively capture the characteristics of the modeling data set. Its Bias is 0, indicating that the model does not show obvious systematic bias on the training set. Both RMS and STD are 0.459, indicating that the prediction error is relatively small, showing that the fitting of the model has a certain stability. Regarding the test set, R 2 is -0.220, indicating that the model fails to effectively fit the data on the prediction set and there may be an overfitting problem. The Bias of the test set is 0.017, indicating that there is no obvious systematic bias in the PLS predicted values. Its RMS and STD are 0.636 and 0.635 respectively, which are larger than those of the training set, indicating that the prediction ability of the model is average. The specific situations of each training sample and test sample are as Figure 2 shown.

[0081] The results of the conventional random forest model (RF) on the training set and test set are shown in Table 4:

[0082] Table 4 RF Modeling and Test Results

[0083] <![CDATA[R 2 > Bias RMS STD Training Set 0.497 -0.012 0.467 0.467 Test Set 0.322 0.052 0.474 0.471

[0084] As can be seen from Table 4, the R of the RF model on the training set 2 is 0.497, indicating that the fitting degree of the modeling result and the true value is relatively high and it can better explain the change trend of the data. RMS, Bias and STD are 0.467, -0.012, 0.467 respectively, proving that the model has high precision. Compared with the training set, the test set results are relatively average. Among them, the Bias is 0.052, indicating that the model has a systematic bias. RMS, R 2 , STD are 0.474, 0.322, 0.471, indicating that when facing unknown samples, the prediction ability of the model is average. The specific situations of each training sample and test sample are as Figure 3 shown.

[0085] The results of the conventional BP neural network on the training set and test set are shown in Table 5:

[0086] Table 5 BP Modeling and Test Results

[0087] <![CDATA[R 2 > Bias RMS STD Training Set 0.213 -0.071 0.585 0.580 Test Set 0.305 0.016 0.480 0.480

[0088] As can be seen from Table 5, the R of the BP model on the training set 2is 0.213, indicating that the model fitting effect is not ideal and the ability to capture the data change trend is limited. RMS, Bias, and STD are 0.585, -0.071, and 0.580 respectively, proving that the modeling accuracy of the model is low and there is a certain degree of underestimation of the salt content. Compared with the training set, RMS, R 2 , Bias, and STD on the test set are 0.480, 0.305, 0.016, and 0.480 respectively, indicating that the generalization ability of the model in the test set is limited. The specific situations of each training sample and test sample are as Figure 4 shown.

[0089] The inversion results of the fusion model (PLS-RF) proposed in this embodiment are shown in Table 6:

[0090] Table 6 PLS-RF Modeling and Test Results

[0091] <![CDATA[R 2 > Bias RMS STD Training Set 0.615 -0.014 0.409 0.409 Test Set 0.562 0.047 0.381 0.378

[0092] As can be seen from Table 6, the training results and test results of the fusion model (PLS-RF) combining PLS and RF are shown in Table 7. Among them, the evaluation index RMS of the training set is 0.409, R 2 is 0.615, Bias is -0.014, and STD is 0.409, which have certain improvements compared with the single PLS and RF. The above results prove that the model can better capture the data characteristics on the training set, has strong fitting ability, and the overall prediction accuracy is relatively high. On the test set, the fusion model also shows strong fitting ability and relatively small overall deviation. Among them, the R 2 of the model is 0.562, which decreases slightly compared with the training set, but can still better reflect the change trend of the data. And RMS, Bias, and STD in the test set are 0.381, 0.047, and 0.378 respectively, indicating that the model has good stability and generalization, and can ensure the accuracy and reliability of the prediction results to a certain extent. The specific situations of each training sample and test sample are as Figure 5 shown.

[0093] To fully discuss the prediction accuracy of each model, Table 7 shows the performance (Bias, RMS, STD) of each model in different salt content intervals (extremely low salt content: 0 - 0.5; low salt content: 0.5 - 1; medium salt content: 1 - 1.5; high salt content: >1.5) among all 110 samples.

[0094] Table 7 Comparison of Model Accuracies in Each Salt Content Interval

[0095]

[0096] In summary, the fusion model (PLS-RF) exhibits excellent performance in all salinity ranges. Specifically, in the low salinity range, Bias is 0.183, RMS is 0.3, and STD is 0.23; in the medium salinity range, Bias is -0.11, RMS is 0.24, and STD is 0.21. The above results indicate that the model has high prediction accuracy and stability in both the low salinity and medium salinity ranges. Among them, in the very low salinity range, Bias is 0.351, RMS is 0.4, and STD is 0.19; in the high salinity range, Bias is -0.57, RMS is 0.61, and STD is 0.23. Although the performance of all models decreases in the very low salinity and high salinity ranges, the fusion model still maintains high accuracy and stability and can better adapt to different salinity conditions.

[0097] In contrast, the PLS model performs well in the medium salinity range (Bias, RMS, and STD are -0.11, 0.32, and 0.3 respectively), but has low accuracy in the very low salinity and low salinity ranges (Bias is 0.327, RMS is 0.56 in the very low salinity range; Bias is 0.26, RMS is 0.490 in the low salinity range). The STD error range of the RF model is (0.19 - 0.23), indicating that the model has good stability and adaptability in most ranges, but has poor prediction accuracy when facing high salinity samples (Bias is -0.7, RMS is 0.75). The Bias of the BP model in the high salinity area is -0.75, RMS is 0.88, and STD is 0.47, proving that the model has low accuracy and unstable results under high salinity conditions.

[0098] In summary, under extreme conditions of very low salinity and high salinity, the performance of models generally decreases, which may be related to the instability of the models caused by large data noise under extreme conditions. Compared with conventional models, the fusion model can provide stable and reliable results, and the fusion model exhibits high accuracy and stability in each salinity range.

[0099] To discuss the spatial distribution of the accuracy of each model, Table 8 shows the performance of conventional models (PLS, RF, BP) and the fusion model (PLS-RF) in different longitude and latitude ranges. The results show that the fusion model has obvious advantages, and the fusion model PLS-RF also shows the best accuracy in each range.

[0100] Table 8 Comparison of model accuracy in space

[0101]

[0102]

[0103] In terms of longitude, the systematic deviations of PLS-RF in each interval are relatively small (Bias values are 0.091, -0.17, and 0.018 respectively), and the RMS and STD values (0.235, 0.452, and 0.458) in different longitude ranges are better than those of other models, indicating its stable performance at different longitudes. In terms of latitude, the performance of Bias and STD of PLS-RF at different latitudes is also relatively stable. Among them, the Bias error range is (-0.196 to 0.109), the RMS error range is (0.169 - 0.555), and the STD error range is (0.155 to 0.555), showing its strong adaptability and robustness. Specifically, the modeling strategy of the fusion model has achieved good inversion results in different salinity levels and different geographical distributions.

[0104] The main innovation of this embodiment lies in proposing a fusion method that combines the partial least squares method (PLS) with the random forest (RF) model, effectively improving the accuracy and reliability of soil salinity inversion. Compared with traditional single models, this embodiment can better handle the non-linear relationships and noises in remote sensing data, significantly enhancing the generalization ability of salinity inversion.

[0105] Through experimental verification, the prediction accuracy of the fusion model in different salinity intervals is significantly better than that of traditional single models, especially excellent in low-salinity and medium-salinity regions, and can be widely applied to soil salinity monitoring in different geographical environments.

[0106] The soil salinity inversion method proposed by the present invention is applicable to the monitoring and management of soil salinization in large-scale regions, especially in regions with severe soil salinization such as coastal tidal flats and arid and semi-arid areas. In addition, this method can also be extended and applied to fields such as agricultural production, land management, and environmental protection, providing a scientific basis for land improvement and ecological restoration.

[0107] Those of ordinary skill in the art can understand that all or part of the steps in implementing the methods of the above embodiments can be completed by instructing relevant hardware through a program. Therefore, this application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, this application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes.

[0108] The above embodiments have introduced the present invention in detail. Specific examples are used in this text to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present invention.

Claims

1. A bare soil salinity inversion method integrating partial least squares and random forest, characterized in that, The method includes the following steps: S1. Obtain spectral reflectance data and salinity spectral indices within multiple band ranges and spatial resolutions from Sentinel-2 satellite remote sensing data, and perform preprocessing to establish a digital orthophoto image containing original multi-dimensional feature data; S2. Use the normalized difference vegetation index to remove water bodies and vegetation-covered areas in the digital orthophoto image to obtain bare soil data containing 24-dimensional feature variables; S3. Establish a fusion model and conduct training and evaluation. The fusion model includes a linear regression model for inputting 24-dimensional feature variables and outputting an initial prediction result, and a random forest model for outputting the surface soil salt content based on 25-dimensional feature variables; S4. Use Sentinel-2 satellite remote sensing data as the input of the fusion model for prediction to obtain the bare soil salt content; S5. Use a single machine learning model to perform predictions respectively with 24-dimensional feature variables, and compare with the predicted bare soil salt content to evaluate the performance of the fusion model.

2. The bare soil salinity inversion method according to claim 1, characterized in that In step S2, specifically, water bodies with NDVI < 0 and vegetation-covered areas with NDVI > 0.15 in the digital orthophoto image are removed. The calculation formula for the normalized difference vegetation index NDVI is: where NDVI represents the normalized difference vegetation index; NIR and R respectively represent the reflectances of the near-infrared band and the red band.

3. The bare soil salinity inversion method according to claim 1, characterized in that, The expression of the fusion model is: where y is the dependent variable; ρ i is the explanatory variable of the i-th variable; w i is the model coefficient of the i-th variable; w0 is the constant term; n is the number of variables; is the final prediction result; T i (x) is the prediction result of the i-th decision tree; N is the number of decision trees.

4. The bare soil salinity inversion method according to claim 1, characterized in that, In step S3, the specific process of training and evaluating the fusion model includes the following steps: S31. Use the original multi-dimensional feature data as the input of the linear regression model, and perform regression analysis on the original multi-dimensional feature data to obtain an initial prediction result; S32. Combine the initial prediction result with the original multi-dimensional feature data to form 25-dimensional feature variables, and use the 25-dimensional feature variables as the input of the random forest model to obtain the surface soil salt content as the predicted value; S33. Obtain the soil salt content information of on-site sampling as the actual value; S34. Use the true value and the predicted value to evaluate the fusion model until convergence.

5. The bare soil salinity inversion method according to claim 4, characterized in that, In step S33, the specific process includes: Collect soil samples at multiple sampling points at a depth of 0 - 20 cm. After drying and grinding the soil samples, prepare a soil extract at a ratio of 1:5 and let it stand for 24 hours. Measure its conductivity, and calculate the soil salt content SSC of the soil sample according to the conductivity. The calculation formula for the soil salt content SSC is: SSC = (0.2882 × EC + 0.0183) × 100% where EC is the conductivity.

6. The bare soil salinity inversion method according to claim 1, wherein In step S34, the root mean square error RMS, coefficient of determination R 2 , mean bias Bias, and standard deviation STD are used as fusion model evaluation metrics, and the calculation formulas are as follows: Among them, y i is the actual value; is the predicted value; n is the number of soil samples; is the mean of the actual values; x i is the data point; is the mean of the data points.

7. The bare soil salinity inversion method according to claim 1, characterized in that The number of decision trees n of the random forest model is 20, and the minimum number of leaves is 8.

Citation Information

Patent Citations

  • Soil water content inversion method based on multi-model ensemble learning

    CN111678866A

  • Method for improving satellite data inversion soil salinity by using unmanned aerial vehicle data

    CN119314035A

Cited By

  • Multi-crop nitrogen content prediction method and system in large-scale environment

    CN121033698A

  • Soil salinity inversion method based on multi-time scale remote sensing feature fusion

    CN121121376A

  • Method for realizing high-precision vegetation water content index inversion by fusing GNSS-IR (Global Navigation Satellite System-Infrared Spectroscopy) and hyperspectral data

    CN121364283A

  • GNSS-R soil salinity inversion method and system based on cross-modal fusion

    CN121859581A