Multi-source remote sensing data soil salt content inversion method and system based on stacking model
By combining ground, UAV, and satellite spectral data with a stacked model, a soil salinity prediction model was constructed, which solved the efficiency and accuracy problems of existing soil salinity measurement methods and achieved efficient and accurate soil salinization monitoring.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- NINGXIA UNIVERSITY
- Filing Date
- 2026-03-30
- Publication Date
- 2026-04-28
AI Technical Summary
Existing methods for determining soil salinity are labor-intensive, costly, and have poor timeliness, making it difficult to support large-scale dynamic monitoring. Furthermore, uncertainties exist when fusing multi-source remote sensing data, affecting the accuracy of quantitative analysis of soil salinization.
A multi-source remote sensing data soil salinity inversion method based on a stacked model was adopted. Monitoring points were selected by Bayesian skryzin interpolation algorithm, and spectral feature indices were constructed by combining ground, UAV and satellite spectral data. Soil salinity was predicted using random forest, support vector machine and extreme gradient boosting model.
It effectively reduces model uncertainty, improves the monitoring accuracy and reliability of dynamic changes in soil salinization, and can accurately extract the salinity of saline soil in arid and semi-arid regions.
Smart Images

Figure CN121933706A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of soil salinity prediction technology, specifically to a method and system for inverting soil salinity based on multi-source remote sensing data using a stacked model. Background Technology
[0002] Global climate change and intensified human activities have led to increasingly severe soil salinization, which seriously damages the physical and chemical properties of soil and poses a serious threat to sustainable agricultural development and ecological environment security. Existing methods for predicting soil salinity are affected by data quality, so there is an urgent need to implement effective soil salinization monitoring and control, which is of great practical significance for ensuring regional food security and optimizing land resource allocation.
[0003] Traditional methods for determining soil salinity, such as field sampling combined with laboratory conductivity measurements, while highly accurate, have inherent limitations such as large workload, high cost, and poor timeliness, making it difficult to support the needs of large-scale dynamic monitoring. The rapid development of remote sensing technology has provided a new approach. High-resolution remote sensing images can capture large-scale salinity change trends, while multi-source remote sensing image fusion technology has made significant progress in the quantitative analysis of soil salinization by integrating the complementary advantages of different data sources such as satellite and UAV images, effectively improving the accuracy and reliability of monitoring.
[0004] Using remote sensing technology for salinity inversion still faces significant uncertainties and challenges. The core issue lies in the inherent differences between data from different sources, such as satellite and UAV imagery, in terms of spatial resolution, spectral resolution, and timeliness. These differences directly affect the accuracy of fused data and may lead to local inversion biases. Therefore, how to rationally select, process, and fuse multi-source data to effectively reduce uncertainties and improve data availability and fusion effectiveness has become a key bottleneck for accurate regional salinity inversion.
[0005] The information disclosed in the background section is only intended to enhance the understanding of the background of this disclosure, and therefore may include information that does not constitute prior art known to those skilled in the art. Summary of the Invention
[0006] The purpose of this invention is to provide a method and system for inverting soil salinity based on multi-source remote sensing data using a stacked model, in order to solve the problems mentioned in the background art.
[0007] To achieve the above objectives, the present invention provides the following technical solution: A method for inverting soil salinity from multi-source remote sensing data based on a stacked model, comprising the following steps: The spatial range of the research target area is clearly defined. With its boundary as a constraint, several target monitoring points are systematically set up within the research target area. The expected information gain of each target monitoring point is calculated using the Bayesian Skrigin interpolation algorithm. Several target monitoring points are selected as sample sampling points based on the expected information gain. Soil electrical conductivity data of the sample sampling points are extracted, and the soil salinity of each sample sampling point is calculated based on the soil electrical conductivity. Ground spectral data, UAV spectral data, and satellite remote sensing spectral data were collected from each sampling point. Based on the satellite remote sensing spectral data as the spectral band alignment benchmark, the ground spectral data and UAV spectral data were resampled to align the three types of spectral data in the spectral band range. Based on the aligned spectral data of the three types, for any sample sampling point, the spectral reflectance of that point under different characteristic bands of each type of spectral data is extracted, and spectral indices are constructed based on the spectral reflectance at each sample sampling point to obtain several spectral characteristic indices for each sample sampling point. Based on the spectral characteristic index and the soil salinity of the corresponding sample points, the Pearson correlation coefficient between the spectral characteristic index and the corresponding salinity is calculated under different types of spectral data using a correlation analysis algorithm. The salinity-sensitive variable is then selected based on the Pearson correlation coefficient. Using a stacked model, with salinity-sensitive variables as input and corresponding soil salinity as labels, a soil salinity prediction model was constructed to determine the soil salinity at points within the target area.
[0008] Furthermore, the target monitoring point is composed of a basic monitoring point and a second basic monitoring point. The specific method for systematically deploying several basic monitoring points within the area is as follows: The target area was divided into several environmental response units of equal area. Several basic monitoring points were randomly selected in each environmental response unit, and the soil salinity of the basic monitoring points was obtained. Within each environmental response unit, several potential monitoring points are randomly selected. Based on the baseline monitoring points and their corresponding soil salinity, a Bayesian Skrigin interpolation model is constructed to calculate the expected information gain for each potential monitoring point. The specific formula for the expected information gain is as follows: In the formula, For the j-th potential monitoring point, Let Kriging's prediction variance be the variance of the j-th potential monitoring point. For balancing parameters, Let be the location coordinates of the i-th basic monitoring point. Let J be the location coordinates of the j-th potential monitoring point. Let be the Euclidean distance between the j-th potential monitoring point and the i-th basic monitoring point, where i is the index of the basic monitoring point and j is the index of the potential monitoring point. The total number of basic monitoring points; Where the Kriging prediction variance of the j-th potential monitoring point The specific method for obtaining the data is as follows: it is obtained through the mutation function of the Bayesian skrygian interpolation model, wherein the mutation function is determined by setting the parameters of the mutation function, and the parameters of the mutation function include nugget value, sill value and range.
[0009] Furthermore, potential monitoring points are selected as second basic monitoring points based on the expected information gain of each potential monitoring point. The specific method includes: setting the number of potential monitoring points to be added in each environmental response unit, and the number of added points is less than the number of potential monitoring points randomly determined in the corresponding environmental response unit; randomly selecting several groups of potential monitoring points from the environmental response unit based on the number of added points; arranging and combining the potential monitoring point groups of each environmental response unit to generate several individuals; constructing an initial population; constructing a fitness function with the goal of maximizing the total expected information gain of the individuals; determining the individual with the largest fitness function through a genetic algorithm; and using the potential monitoring points in the potential monitoring point group corresponding to this individual as the second basic monitoring points. The specific formula upon which the fitness function is constructed is: In the formula, F is the fitness function value. Let be the expected information gain of the p-th potential monitoring point in the individual, where p is the index of the potential monitoring point in the individual, and P is the total number of potential monitoring points in the individual.
[0010] Furthermore, the specific logic for selecting several sample points based on the set target monitoring points is as follows: All secondary basic monitoring points were used as sample sampling points, and several points were randomly selected from the basic monitoring points to determine all sample sampling points. Soil electrical conductivity data were obtained from these sample sampling points. The specific formula used to calculate the soil salinity at each sample sampling point based on the soil electrical conductivity is as follows: In the formula, Soil salinity This refers to the soil electrical conductivity.
[0011] Furthermore, the specific method for collecting ground spectral data, UAV spectral data, and satellite remote sensing spectral data at each sample sampling point is as follows: a spectral acquisition time period is set, and ground spectral data, UAV spectral data, and satellite remote sensing spectral data at each sample sampling point are collected during the spectral acquisition time period. The logic for resampling ground spectral data and UAV spectral data to align the three types of spectral data in the spectral band range is as follows: preprocessing satellite remote sensing spectral data with radiometric calibration and atmospheric correction, and preprocessing ground spectral data and UAV spectral data with Savitzky Golay convolution smoothing. Based on the three types of preprocessed spectral data, and according to the band range of the satellite remote sensing spectral data, the three types of spectral data are aligned in terms of spectral frequency. The specific steps are as follows: Based on the preprocessed ground spectral data and UAV spectral data, spectral resampling is performed to obtain the reflectance values of the ground spectral data and UAV spectral data in different characteristic bands. The characteristic bands include the blue light band, green light band, red light band, and near-infrared band.
[0012] Furthermore, based on the spectral reflectance at each sample sampling point, the spectral characteristic index of each sample sampling point is calculated. The spectral characteristic index includes salinity index, first salinization index, second salinization index, third salinization index, fourth salinization index, soil-regulated vegetation index, greenness difference vegetation index, atmospheric resistance vegetation index, canopy salinity response vegetation index, expansion difference vegetation index, and normalized vegetation index. The specific formulas used to calculate the salinity index, the first salinization index, the second salinization index, the third salinization index, and the fourth salinization index are as follows: In the formula, , , and These are the first salinization index, the second salinization index, the third salinization index, and the fourth salinization index, respectively. Salt content index The spectral reflectance in the blue light band. The spectral reflectance is in the red light band. The spectral reflectance in the green light band, This represents the spectral reflectance value in the near-infrared band.
[0013] Furthermore, the logic underlying the determination of each salt content-sensitive variable is as follows: Based on three types of spectral data, the spectral characteristic index of the corresponding type of spectral data is calculated respectively. For the spectral characteristic index of any type of spectral data, the Pearson correlation coefficient between the spectral characteristic index and the corresponding salt content at different sample sampling points is calculated. The average Pearson correlation coefficient is calculated for the same spectral characteristic index. This average Pearson correlation coefficient is used as the sensitivity coefficient of the spectral characteristic index for the type of spectral data. Determine the sensitivity coefficient of any spectral feature index for different types of spectral data, compare the sensitivity coefficients of the spectral feature index for different types of spectral data, select the type of spectral data corresponding to the largest sensitivity coefficient, and use the spectral feature index as the salt content sensitive variable for the corresponding type of spectral data; traverse all spectral feature indices to determine their spectral data type affiliation.
[0014] Furthermore, the soil salinity prediction model is constructed through a stacked model, specifically including a base layer and a meta-model. The base layer includes three homogeneous learners for ground spectral data, UAV spectral data, and satellite spectral data. The meta-model uses random forest, extreme gradient boosting, and support vector machine. The three types of homogeneous learners are used to receive salinity-sensitive variables from ground spectral data, UAV spectral data, and satellite spectral data, respectively, and output corresponding predicted values of soil salinity. The predicted values of soil salinity from each homogeneous learner are merged to form the meta-feature vector of each homogeneous learner. The meta-feature vector of each homogeneous learner is combined with the actual soil salinity value to form meta-training data, which serves as the input to the meta-model. The output of the meta-model is the comprehensive predicted value of soil salinity. Input the salinity-sensitive variables from ground spectral data, UAV spectral data, and satellite spectral data into the trained soil salinity prediction model to determine the soil salinity of all image pixels in the target area.
[0015] This invention also provides a system for inverting soil salinity from multi-source remote sensing data based on a stacked model. This system is used to execute the aforementioned method for inverting soil salinity from multi-source remote sensing data based on a stacked model, and includes: The sample data acquisition module is used to define the spatial range of the research target area. With its boundary as a constraint, several target monitoring points are systematically set up within the research target area. The expected information gain of each target monitoring point is calculated using the Bayesian Skrigin interpolation algorithm. Several target monitoring points are selected as sample sampling points based on the expected information gain. Soil conductivity data of the sample sampling points are extracted, and the soil salinity of each sample sampling point is calculated based on the soil conductivity. The feature spectrum acquisition module is used to collect ground spectral data, UAV spectral data and satellite remote sensing spectral data at each sample sampling point, and resamples the ground spectral data and UAV spectral data based on the satellite remote sensing spectral data as the spectral band alignment reference to align the three types of spectral data in the spectral band range. The feature parameter calculation module is used to extract the spectral reflectance of each sample point in different feature bands based on the aligned spectral data of the three types. It also constructs spectral indices based on the spectral reflectance of each sample point to obtain several spectral feature indices for each sample point. The sensitive variable screening module is used to calculate the Pearson correlation coefficient between the spectral characteristic index and the corresponding soil salinity at the corresponding sample sampling point based on the spectral characteristic index and the corresponding salinity under different types of spectral data through a correlation analysis algorithm, and to screen the salinity sensitive variable based on the Pearson correlation coefficient. The salinity prediction module is used to construct a soil salinity prediction model by using a stacked model, taking salinity-sensitive variables as input to the model, and using the corresponding soil salinity as a label, in order to determine the soil salinity at points within the research target area.
[0016] Compared with the prior art, the beneficial effects of the present invention are: The proposed method for predicting soil salinization based on a stacked model uses ground spectral data, UAV spectral data, and satellite spectral data as the base model. The meta-model is constructed using random forest, support vector machine, and extreme gradient boosting. By combining and stacking multiple models and data sources, the method effectively reduces the model uncertainty during single-model inversion, providing a stable foundation of data for accurately monitoring the dynamic changes of soil salinization. It also effectively solves the problem of low accuracy in soil salinity prediction models based on low-quality data, and can accurately extract the salinity of saline soils in arid and semi-arid regions. Attached Figure Description
[0017] Figure 1 This is a schematic diagram of the overall method flow of the present invention; Figure 2 A line graph showing the performance comparison of root mean square error between stacked models and single models; Figure 3 A line graph comparing the mean absolute error performance of stacked models versus single models; Figure 4 A comparison chart of the coefficient of determination performance of stacked models versus single models; Figure 5 This is a schematic diagram of the overall system structure of the present invention. Detailed Implementation
[0018] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to specific embodiments.
[0019] It should be noted that, unless otherwise defined, the technical or scientific terms used in this invention should have the ordinary meaning understood by one of ordinary skill in the art to which this invention pertains. The terms "first," "second," and similar terms used in this invention do not indicate any order, quantity, or importance, but are merely used to distinguish different components. Terms such as "comprising" or "including" mean that the element or object preceding the word encompasses the elements or objects listed following the word and their equivalents, without excluding other elements or objects. Terms such as "connected" or "linked" are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. Terms such as "upper," "lower," "left," and "right" are used only to indicate relative positional relationships; when the absolute position of the described object changes, the relative positional relationship may also change accordingly.
[0020] Example: Please see Figures 1-4 The present invention provides a technical solution: A method for inverting soil salinity from multi-source remote sensing data based on a stacked model, comprising the following steps: Step 1: Define the spatial range of the research target area. Using its boundary as a constraint, systematically deploy several target monitoring points within the research target area. Calculate the expected information gain of each target monitoring point using the Bayesian Skrigin interpolation algorithm. Select several target monitoring points as sample sampling points based on the expected information gain. Extract soil electrical conductivity data from the sample sampling points and calculate the soil salinity of each sample sampling point based on the soil electrical conductivity.
[0021] The spatial range of the research target area is clearly defined. With its boundary as a constraint, several target monitoring points are systematically set up within the research target area. Based on the set target monitoring points, several sample sampling points are selected, the soil electrical conductivity data of the sample sampling points are extracted, and the soil salinity of each sample sampling point is calculated based on the soil electrical conductivity. The target monitoring points are composed of basic monitoring points and second basic monitoring points. The specific method for systematically deploying several basic monitoring points in the area is as follows: The target area was divided into several environmental response units of equal area. Several basic monitoring points were randomly selected in each environmental response unit, and the soil salinity of the basic monitoring points was obtained. The size of the environmental response unit is set according to a certain proportion based on the total area of the research target area. The size of a single environmental response unit is generally set between 5% and 10% of the total area of the research target area. Several basic monitoring points are randomly selected in the environmental response unit. The specific selection can be based on the different terrain, soil type and vegetation coverage within the coverage unit. Generally, 5-20 basic monitoring points are selected. Within each environmental response unit, several potential monitoring points are randomly selected, with locations chosen at a suitable distance from the baseline monitoring points to cover a wider area and ensure data comprehensiveness. Generally, 3-15 potential monitoring points are set, adjusted according to the size of the environmental response unit. Based on the baseline monitoring points and their corresponding soil salinity, a Bayesian Skrigin interpolation model is constructed to calculate the expected information gain for each potential monitoring point. The specific formula for the expected information gain is as follows: In the formula, For the j-th potential monitoring point, Let Kriging's prediction variance be the variance of the j-th potential monitoring point. For balancing parameters, Let be the location coordinates of the i-th basic monitoring point. Let J be the location coordinates of the j-th potential monitoring point. Let be the Euclidean distance between the j-th potential monitoring point and the i-th basic monitoring point, where i is the index of the basic monitoring point and j is the index of the potential monitoring point. The total number of basic monitoring points; It should be noted that the expected information gain of the j-th potential monitoring point... Indicators used to measure the value of potential monitoring points, when A higher value indicates that the j-th potential monitoring point has high monitoring value. Monitoring it can significantly improve the understanding and prediction of soil salinity changes in the area. In practice, a higher value... This indicates that monitoring at this point can yield more valuable information, thereby guiding subsequent monitoring decisions and resource allocation.
[0022] in Indicates the first The Kriging prediction variance for each potential monitoring point reflects the degree of uncertainty in the spatial location model of that point. A larger variance indicates greater uncertainty in the prediction of soil salinity at that monitoring point, and therefore higher potential observational value. and Proportional; Using the inverse square of the distance, the closer the base monitoring point is, the greater its contribution to the expected information gain. This setting is to emphasize the importance of spatially close base points to the information gain of potential points, because the information provided by close points is usually more relevant. It is a balancing parameter used to weigh the variance of Kriging predictions against the distance information of baseline monitoring points, and is adjusted by... The relative importance of the two can be controlled. The balance parameter can be set according to expert experience. This balance parameter is an empirical parameter and generally takes a value between 0 and 1. Where the Kriging prediction variance of the j-th potential monitoring point The specific method for obtaining the variogram is as follows: it is obtained through the variogram function of the Bayesian Kriging interpolation model. The variogram function is determined by setting its parameters, which include nugget value, sill value, and range. The specific steps include: selecting a suitable variogram function model (common models include spherical models, exponential models, and Gaussian models); setting the parameters of the variogram function model, including the nugget value, sill value, and initial range values; calculating the variogram function; collecting soil salinity data from basic monitoring points and determining their spatial coordinates; calculating the distance between all basic monitoring points and calculating the variogram value for each pair of monitoring points; fitting the selected variogram function model using statistical methods such as least squares; determining the nugget value, sill value, and range; constructing a Kriging equation using the variogram function, typically including the influence weights of basic monitoring points on potential monitoring points; calculating the covariance matrix between basic monitoring points based on the variogram function; substituting the variogram function result into the Kriging equation to obtain the Kriging prediction variance of the potential monitoring points.
[0023] The method for selecting potential monitoring points as second basic monitoring points based on the expected information gain of each potential monitoring point includes: setting the number of potential monitoring points to be added in each environmental response unit, and the number of added points being less than the number of potential monitoring points randomly determined in the corresponding environmental response unit. The number of potential monitoring points to be added in each environmental response unit is generally set to between 30% and 70% of the number of potential monitoring points in each environmental response unit. Based on the number of added points, several groups of potential monitoring points are randomly selected from the environmental response unit. The potential monitoring point groups of each environmental response unit are arranged and combined to generate several individuals, and an initial population is constructed. A fitness function is constructed with the goal of maximizing the total expected information gain of the individuals. The individual with the largest fitness function is determined by a genetic algorithm, and the potential monitoring points in the potential monitoring point group corresponding to this individual are used as the second basic monitoring points. The specific formula upon which the fitness function is constructed is: In the formula, F is the fitness function value. Let be the expected information gain of the p-th potential monitoring point in the individual, where p is the index of the potential monitoring point in the individual, and P is the total number of potential monitoring points in the individual.
[0024] A higher value indicates that the j-th potential monitoring point has high monitoring value. Monitoring it can significantly improve the understanding and prediction of soil salinity changes in the area. Therefore, [the following is a possible interpretation:] Maximizing the potential monitoring point group is the optimization objective, and the optimal group is selected by genetic algorithm. Since genetic algorithm is a common optimization algorithm, it will not be elaborated here. Randomly selecting monitoring points without considering the spatial distribution of existing monitoring points within the region fails to effectively utilize existing monitoring point information and cannot scientifically assess the value of potential monitoring points, resulting in a relatively random distribution that is difficult to cover spatial variation patterns. By calculating the expected information gain and combining the Kriging prediction variance with the distance relationship with the basic monitoring points, potential monitoring points that can bring higher information gain are prioritized as the second basic monitoring points. This deployment method can more efficiently expand the information coverage and reduce redundant sampling points. In addition, in the random selection method, the selection of potential monitoring points is isolated, and there is no optimization strategy between them, which may lead to uneven distribution of monitoring points or even omission of important areas. By using a genetic algorithm to globally optimize the combination of potential monitoring points within the environmental response unit, the goal is to maximize the total expected information gain. In the genetic algorithm, through operations such as crossover, mutation, and selection of individuals, the optimal combination of potential monitoring points can be searched in multidimensional space to finally determine the second basic monitoring points.
[0025] The specific logic for selecting several sample points based on the set target monitoring points is as follows: All secondary basic monitoring points were used as sample sampling points, and several points were randomly selected from the basic monitoring points to determine all sample sampling points. Soil electrical conductivity data were obtained from these sample sampling points. The specific formula used to calculate the soil salinity at each sample sampling point based on the soil electrical conductivity is as follows: In the formula, Soil salinity Soil electrical conductivity; soil salinity. The unit is g / kg, soil electrical conductivity The unit is us / cm.
[0026] Step 2: Collect ground spectral data, UAV spectral data and satellite remote sensing spectral data from each sampling point, and resample the ground spectral data and UAV spectral data based on the satellite remote sensing spectral data as the spectral band alignment reference to align the three types of spectral data in the spectral band range.
[0027] The specific method for collecting ground spectral data, UAV spectral data and satellite remote sensing spectral data at each sample sampling point is as follows: Set a spectral acquisition time period, and collect ground spectral data, UAV spectral data and satellite remote sensing spectral data at each sample sampling point during the spectral acquisition time period; The logic for resampling ground spectral data and UAV spectral data to align the three types of spectral data in the spectral band range is as follows: preprocessing satellite remote sensing spectral data with radiometric calibration and atmospheric correction, and preprocessing ground spectral data and UAV spectral data with Savitzky Golay convolution smoothing. Based on the three types of preprocessed spectral data, and according to the band range of the satellite remote sensing spectral data, the three types of spectral data are aligned in terms of spectral frequency. The specific steps are as follows: Based on the preprocessed ground spectral data and UAV spectral data, spectral resampling is performed to obtain the reflectance values of the ground spectral data and UAV spectral data in different characteristic bands. The characteristic bands include the blue light band, green light band, red light band, and near-infrared band.
[0028] The specific method for spectral resampling is as follows: Interpolation is used to calculate reflectance values within a specific wavelength range. For each characteristic band, including blue, green, red, and near-infrared bands, its center wavelength is determined. For each characteristic band, interpolation is used to calculate the reflectance values of the ground spectrum and the UAV spectrum at that wavelength. For each center wavelength, the two nearest actual wavelengths are found, and then the corresponding reflectance values are calculated using the interpolation formula.
[0029] The process involves several steps: spectral resampling unifies the acquired image data to the same spectral bands for subsequent analysis; radiometric calibration converts the raw digital values of remote sensing images into physically meaningful radiance values, based on the sensor's calibration file, by converting the raw digital values into radiance or reflectance values; atmospheric correction eliminates atmospheric effects and obtains the true reflectance of ground objects, using radiometric transmission models such as MODTRAN and 6S models for correction, employing meteorological data such as temperature, humidity, and air pressure as input to simulate atmospheric conditions, and processing the data using open-source tools or software packages.
[0030] Savitzky-Golay convolution smoothing is used to remove noise, smooth spectral curves, and improve the signal-to-noise ratio. The specific steps include: selecting an appropriate window size and polynomial order, typically with a window size of 5-11 points and a polynomial order of 2 or 3; using the Savitzky-Golay convolution smoothing formula to output the smoothed spectral values.
[0031] Based on the three types of preprocessed spectral data, and according to the band range of the satellite remote sensing spectral data, the three types of spectral data are aligned in terms of spectral frequency. The specific steps are as follows: Based on the preprocessed ground spectral data and UAV spectral data, spectral resampling is performed to obtain the reflectance values of the ground spectral data and UAV spectral data in the blue, green, red, and near-infrared bands; For the reflectance values of the ground spectral data, UAV spectral data, and satellite remote sensing spectral data in the blue, green, red, and near-infrared bands, the reflectance values of different wavelengths within any band are specifically calculated to obtain the average reflectance value at different wavelengths within the band range, and this average reflectance value is used as the reflectance value of the corresponding spectral data in each characteristic band.
[0032] Step 3: Based on the aligned spectral data of the three types, for any sample sampling point, extract the spectral reflectance of that point in different characteristic bands under each type of spectral data, and construct spectral indices based on the spectral reflectance at each sample sampling point to obtain several spectral characteristic indices for each sample sampling point.
[0033] The spectral characteristic index of each sample sampling point is calculated based on the spectral reflectance at each sampling point. The spectral characteristic index includes salinity index, first salinization index, second salinization index, third salinization index, fourth salinization index, soil-regulated vegetation index, greenness difference vegetation index, atmospheric impedance vegetation index, canopy salinity response vegetation index, expansion difference vegetation index, and normalized vegetation index. The specific formulas used to calculate the salinity index, the first salinization index, the second salinization index, the third salinization index, and the fourth salinization index are as follows: In the formula, , , and These are the first salinization index, the second salinization index, the third salinization index, and the fourth salinization index, respectively. Salt content index The spectral reflectance in the blue light band. The spectral reflectance is in the red light band. The spectral reflectance in the green light band, This represents the spectral reflectance value in the near-infrared band.
[0034] Soil-regulated vegetation index, greenness difference vegetation index, atmospheric resistance vegetation index, canopy salinity-responsive vegetation index, expansion difference vegetation index, and normalized vegetation index are all common characteristic parameters, and the formulas used for their calculation are as follows: In the formula, To regulate the vegetation index in the soil, The vegetation index is the difference in greenness. To expand the difference in vegetation index, Normalized Difference Vegetation Index (NDVI) The vegetation index is the response of canopy salinity. This is the atmospheric impedance vegetation index.
[0035] Step 4: Based on the spectral characteristic index and the soil salinity of the corresponding sample points, calculate the Pearson correlation coefficient between the spectral characteristic index and the corresponding salinity under different types of spectral data using a correlation analysis algorithm, and screen the salinity-sensitive variables based on the Pearson correlation coefficient.
[0036] The logic behind determining the specific sensitivity variables for each salt content is as follows: Based on three types of spectral data, the spectral characteristic index of the corresponding type of spectral data is calculated respectively. For the spectral characteristic index of any type of spectral data, the Pearson correlation coefficient between the spectral characteristic index and the corresponding salt content at different sample sampling points is calculated. The average Pearson correlation coefficient is calculated for the same spectral characteristic index. This average Pearson correlation coefficient is used as the sensitivity coefficient of the spectral characteristic index for the type of spectral data. The sensitivity coefficient of any spectral feature index to different types of spectral data is determined. The sensitivity coefficients of the spectral feature index to different types of spectral data are compared, and the spectral data type corresponding to the largest sensitivity coefficient is selected. This spectral feature index is then used as the salt content sensitivity variable for the corresponding type of spectral data. All spectral feature indices are iterated to determine their spectral data type affiliation. If the largest sensitivity coefficient corresponds to multiple types of spectral data, the spectral feature index is simultaneously incorporated into the salt content sensitivity variable for the corresponding type of spectral data. The calculation method of the Pearson correlation coefficient is a conventional existing technique, and the specific technique will not be elaborated here.
[0037] Specifically, the salinity-sensitive variables in ground spectral data include the salinity index, the first salinization index, the third salinization index, and the fourth salinization index; The salinity-sensitive variables for the spectral data of drones include the second salinization index, the third salinization index, the fourth salinization index, and the salinity index; The salinity-sensitive variables in satellite spectral data include soil-regulated vegetation index, greenness difference vegetation index, atmospheric resistance vegetation index, canopy salinity-responsive vegetation index, expansion difference vegetation index, and normalized vegetation index.
[0038] Step 5: Using a stacked model, with the salinity-sensitive variable as the model input and the corresponding soil salinity as the label, construct a soil salinity prediction model to determine the soil salinity at points within the research target area.
[0039] Specifically, the framework is implemented using a stacking model, with the base layer integrating three homogeneous learners for ground spectral data, UAV spectral data, and satellite spectral data. The soil salinity prediction model is constructed using a stacked model, specifically including a base layer and a meta-model. The base layer includes three homogeneous learners for ground spectral data, UAV spectral data, and satellite spectral data. The meta-model uses random forest, extreme gradient boosting, and support vector machine. The three types of homogeneous learners are used to receive salinity-sensitive variables from ground spectral data, UAV spectral data, and satellite spectral data, respectively, and output corresponding predicted values of soil salinity. The predicted values of soil salinity from each homogeneous learner are merged to form the meta-feature vector of each homogeneous learner. The meta-feature vector of each homogeneous learner is combined with the actual soil salinity value to form meta-training data, which serves as the input to the meta-model. The output of the meta-model is the comprehensive predicted value of soil salinity. The specific steps for forming the meta-feature vectors of each homogeneous learner include: determining the total prediction variance of the base layer learner by comprehensively considering the data source quality uncertainty parameter and the model fitting uncertainty parameter. The formula used to calculate the data source quality uncertainty parameter is as follows: In the formula, The uncertainty parameter of the data source quality is input to the base layer learner for group b. The signal-to-noise ratio of the data input to the base layer learner for group b. The cloud coverage rate of the satellite imagery is input into the base layer learner for group b, specifically the percentage of the imagery obscured by clouds. This refers to the perspective of the drone's orbital observation, specifically the angle between the sensor onboard the drone and the ground normal. b represents the maximum observation angle; b is the index of the input data set.
[0040] The cloud coverage rate of satellite imagery is calculated using a cloud detection algorithm, the specific observation angle of the UAV orbit is determined by actual settings, and the maximum observation angle is obtained by the maximum observation angle allowed by the sensor.
[0041] The uncertainty parameter for model fitting is specifically the standard deviation between the predicted and actual values of the base layer learner. The formula used to comprehensively determine the total prediction variance of the base layer learners is as follows: In the formula, Fit uncertainty parameters to the model using the base layer learner data input in group b. and Here are the weighting coefficients, where and and All are greater than 0; Detailed settings The reason is that the standard deviation between the predicted and actual values of the base layer learner directly reflects the actual prediction accuracy of the prediction model, while the uncertainty parameter of the data source quality indirectly represents its impact on the prediction accuracy. Therefore, setting... .
[0042] Based on the total variance of the predictions from the base layer learners, meta-training data is generated using an attention-weighted approach. The specific formula for calculating the attention weights is as follows: In the formula, The attention weights for the b-th input data. This is the uncertainty penalty coefficient, typically between 0.1 and 1; By using attention weights to fuse input data from different training rounds, a meta-feature vector is obtained.
[0043] The specific training steps include: randomly selecting 2 data points from every 10 sampling points of salinity as the validation set, and using the remaining data as the training set; inputting a stacked model in Python, using the selected sensitive variables as independent variables and the measured soil salinity as the dependent variable, with ground spectral data, UAV spectral data, and satellite spectral data forming the base model, and the meta-model consisting of random forest, support vector machine, and extreme gradient boosting to construct a soil salinity prediction model; continuously optimizing the parameters of the random forest, support vector machine, and extreme gradient boosting, including the number of regression trees, minimum number of leaves, kernel function type, regularization coefficient (C), kernel width (gamma), learning rate, subsample ratio, etc., to finally obtain the optimal model parameter settings; the specific steps include: using the Python environment to build a stacked model... The model consists of selected sensitive variables as independent variables and measured soil salinity as the dependent variable. The base model is composed of ground spectral data, UAV spectral data, and satellite spectral data to provide comprehensive predictions based on multi-source information. Meanwhile, the meta-model selects random forest, support vector machine, and extreme gradient boosting algorithms to improve prediction accuracy. Cross-validation, such as k-fold cross-validation, is used to evaluate the model's performance under different parameter settings. The training set is divided into several subsets, and one subset is used for validation in turn, while the other k subsets are used for training to obtain the average performance of the model under each parameter configuration, reducing the impact of randomness on model evaluation. Based on the performance parameters of the validation set, including mean squared error and R² score, the optimal parameters are selected. Specifically, the model parameters that minimize the mean squared error and maximize the R² score are recorded as the optimal model parameters. The optimal regression tree has 100 trees, a minimum number of leaves of 2, a regularization coefficient of 1, a kernel width of 0.142857143, epsilon=0.05 (allowing 5% absolute error without penalty), a learning rate of 0.3, max_depth=2, subsample=1, colsample_bytree=1, n_estimators=100, an objective function of reg:squarederror, and a random seed of random_state=42.
[0044] The soil salinity prediction model, after being trained, inputs salinity-sensitive variables from ground spectral data, UAV spectral data, and satellite spectral data to determine the soil salinity of all image pixels in the target area. The base model is constructed using ground spectral data, UAV spectral data, and satellite spectral data, while the meta-model is constructed using random forest, support vector machine, and extreme gradient boosting. By combining and stacking multiple models and data sources, the model uncertainty during single-model inversion is effectively reduced, providing a stable basic data guarantee for accurately monitoring the dynamic changes of soil salinization.
[0045] Please see Figure 5 The present invention also provides a soil salinity inversion system based on a stacked model for multi-source remote sensing data. This system is used to execute the aforementioned method for inverting soil salinity based on a stacked model for multi-source remote sensing data, and includes: The sample data acquisition module is used to define the spatial range of the research target area. With its boundary as a constraint, several target monitoring points are systematically set up within the research target area. The expected information gain of each target monitoring point is calculated using the Bayesian Skrigin interpolation algorithm. Several target monitoring points are selected as sample sampling points based on the expected information gain. Soil conductivity data of the sample sampling points are extracted, and the soil salinity of each sample sampling point is calculated based on the soil conductivity. The feature spectrum acquisition module is used to collect ground spectral data, UAV spectral data and satellite remote sensing spectral data at each sample sampling point, and resamples the ground spectral data and UAV spectral data based on the satellite remote sensing spectral data as the spectral band alignment reference to align the three types of spectral data in the spectral band range. The feature parameter calculation module is used to extract the spectral reflectance of each sample point in different feature bands based on the aligned spectral data of the three types. It also constructs spectral indices based on the spectral reflectance of each sample point to obtain several spectral feature indices for each sample point. The sensitive variable screening module is used to calculate the Pearson correlation coefficient between the spectral characteristic index and the corresponding soil salinity at the corresponding sample sampling point based on the spectral characteristic index and the corresponding salinity under different types of spectral data through a correlation analysis algorithm, and to screen the salinity sensitive variable based on the Pearson correlation coefficient. The salinity prediction module is used to construct a soil salinity prediction model by using a stacked model, taking salinity-sensitive variables as input to the model, and using the corresponding soil salinity as a label, in order to determine the soil salinity at points within the research target area.
[0046] The above formulas are all dimensionless calculations. The formulas are derived from software simulations based on a large amount of collected data to obtain the most recent real-world results. The preset parameters in the formulas are set by those skilled in the art according to the actual situation.
[0047] The above embodiments can be implemented, in whole or in part, by software, hardware, firmware, or any other combination thereof. When implemented in software, the above embodiments can be implemented, in whole or in part, as a computer program product. Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution.
[0048] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0049] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application.
Claims
1. A method for inverting soil salinity based on multi-source remote sensing data using a stacked model, characterized in that, The specific steps include: The spatial range of the research target area is clearly defined. With its boundary as a constraint, several target monitoring points are systematically set up within the research target area. The expected information gain of each target monitoring point is calculated using the Bayesian Skrigin interpolation algorithm. Several target monitoring points are selected as sample sampling points based on the expected information gain. Soil electrical conductivity data of the sample sampling points are extracted, and the soil salinity of each sample sampling point is calculated based on the soil electrical conductivity. Ground spectral data, UAV spectral data, and satellite remote sensing spectral data were collected from each sampling point. Based on the satellite remote sensing spectral data as the spectral band alignment benchmark, the ground spectral data and UAV spectral data were resampled to align the three types of spectral data in the spectral band range. Based on the aligned spectral data of the three types, for any sample sampling point, the spectral reflectance of that point under different characteristic bands of each type of spectral data is extracted, and spectral indices are constructed based on the spectral reflectance at each sample sampling point to obtain several spectral characteristic indices for each sample sampling point. Based on the spectral characteristic index and the soil salinity of the corresponding sample points, the Pearson correlation coefficient between the spectral characteristic index and the corresponding salinity is calculated under different types of spectral data using a correlation analysis algorithm. The salinity-sensitive variable is then selected based on the Pearson correlation coefficient. Using a stacked model, with salinity-sensitive variables as input and corresponding soil salinity as labels, a soil salinity prediction model was constructed to determine the soil salinity at points within the target area.
2. The method for inverting soil salinity based on a stacked model using multi-source remote sensing data according to claim 1, characterized in that: The target monitoring points are composed of basic monitoring points and second basic monitoring points. The specific method for systematically deploying several basic monitoring points in the area is as follows: The target area was divided into several environmental response units of equal area. Several basic monitoring points were randomly selected in each environmental response unit, and the soil salinity of the basic monitoring points was obtained. Within each environmental response unit, several potential monitoring points are randomly selected. Based on the baseline monitoring points and their corresponding soil salinity, a Bayesian Skrigin interpolation model is constructed to calculate the expected information gain for each potential monitoring point. The specific formula for the expected information gain is as follows: In the formula, For the j-th potential monitoring point, Let Kriging's prediction variance be the variance of the j-th potential monitoring point. For balancing parameters, Let be the location coordinates of the i-th basic monitoring point. Let J be the location coordinates of the j-th potential monitoring point. Let be the Euclidean distance between the j-th potential monitoring point and the i-th basic monitoring point, where i is the index of the basic monitoring point and j is the index of the potential monitoring point. The total number of basic monitoring points; Where the Kriging prediction variance of the j-th potential monitoring point The specific method for obtaining the data is as follows: it is obtained through the mutation function of the Bayesian skrygian interpolation model, wherein the mutation function is determined by setting the parameters of the mutation function, and the parameters of the mutation function include nugget value, sill value and range.
3. The method for inverting soil salinity based on a stacked model using multi-source remote sensing data according to claim 2, characterized in that: The method for selecting potential monitoring points as second basic monitoring points based on the expected information gain of each potential monitoring point includes: setting the number of potential monitoring points to be added in each environmental response unit, and the number of added points being less than the number of potential monitoring points randomly determined in the corresponding environmental response unit; randomly selecting several groups of potential monitoring points from the environmental response unit based on the number of added points; arranging and combining the potential monitoring point groups of each environmental response unit to generate several individuals; constructing an initial population; constructing a fitness function with the goal of maximizing the total expected information gain of the individuals; determining the individual with the largest fitness function through a genetic algorithm; and using the potential monitoring points in the potential monitoring point group corresponding to this individual as the second basic monitoring points. The specific formula upon which the fitness function is constructed is: In the formula, F is the fitness function value. Let be the expected information gain of the p-th potential monitoring point in the individual, where p is the index of the potential monitoring point in the individual, and P is the total number of potential monitoring points in the individual.
4. The method for inverting soil salinity based on a stacked model using multi-source remote sensing data according to claim 3, characterized in that: The logic for selecting several sample points based on the set target monitoring points is as follows: All secondary basic monitoring points were used as sample sampling points, and several points were randomly selected from the basic monitoring points to determine all sample sampling points. Soil electrical conductivity data were obtained from these sample sampling points. The specific formula used to calculate the soil salinity at each sample sampling point based on the soil electrical conductivity is as follows: In the formula, Soil salinity This refers to the soil electrical conductivity.
5. The method for inverting soil salinity based on a stacked model using multi-source remote sensing data according to claim 4, characterized in that: The specific method for collecting ground spectral data, UAV spectral data and satellite remote sensing spectral data at each sample sampling point is as follows: a spectral acquisition time period is set, and ground spectral data, UAV spectral data and satellite remote sensing spectral data at each sample sampling point are collected respectively during the spectral acquisition time period; The logic for resampling ground spectral data and UAV spectral data to align the three types of spectral data in the spectral band range is as follows: preprocessing satellite remote sensing spectral data with radiometric calibration and atmospheric correction, and preprocessing ground spectral data and UAV spectral data with Savitzky Golay convolution smoothing. Based on the three types of preprocessed spectral data, and according to the band range of the satellite remote sensing spectral data, the three types of spectral data are aligned in terms of spectral frequency. The specific steps are as follows: Based on the preprocessed ground spectral data and UAV spectral data, spectral resampling is performed to obtain the reflectance values of the ground spectral data and UAV spectral data in different characteristic bands. The characteristic bands include the blue light band, green light band, red light band, and near-infrared band.
6. The method for inverting soil salinity based on a stacked model using multi-source remote sensing data according to claim 1, characterized in that: The spectral characteristic index of each sample sampling point is calculated based on the spectral reflectance at each sampling point. The spectral characteristic index includes salinity index, first salinization index, second salinization index, third salinization index, fourth salinization index, soil-regulated vegetation index, greenness difference vegetation index, atmospheric impedance vegetation index, canopy salinity response vegetation index, expansion difference vegetation index, and normalized vegetation index. The specific formulas used to calculate the salinity index, the first salinization index, the second salinization index, the third salinization index, and the fourth salinization index are as follows: In the formula, , , and These are the first salinization index, the second salinization index, the third salinization index, and the fourth salinization index, respectively. Salt content index The spectral reflectance in the blue light band. The spectral reflectance is in the red light band. The spectral reflectance in the green light band, This represents the spectral reflectance value in the near-infrared band.
7. The method for inverting soil salinity based on a stacked model using multi-source remote sensing data according to claim 6, characterized in that: The logic behind determining the specific sensitivity variables for each salt content is as follows: Based on three types of spectral data, the spectral characteristic index of the corresponding type of spectral data is calculated respectively. For the spectral characteristic index of any type of spectral data, the Pearson correlation coefficient between the spectral characteristic index and the corresponding salt content at different sample sampling points is calculated. The average Pearson correlation coefficient is calculated for the same spectral characteristic index. This average Pearson correlation coefficient is used as the sensitivity coefficient of the spectral characteristic index for the type of spectral data. Determine the sensitivity coefficient of any spectral feature index for different types of spectral data, compare the sensitivity coefficients of the spectral feature index for different types of spectral data, select the type of spectral data with the largest sensitivity coefficient, and use the spectral feature index as the salt content sensitive variable for the corresponding type of spectral data; Traverse all spectral feature indices to determine their spectral data type affiliation.
8. The method for inverting soil salinity based on a stacked model using multi-source remote sensing data according to claim 7, characterized in that: The soil salinity prediction model is constructed using a stacked model, specifically including a base layer and a meta-model. The base layer includes three homogeneous learners for ground spectral data, UAV spectral data, and satellite spectral data. The meta-model uses random forest, extreme gradient boosting, and support vector machine. The three types of homogeneous learners are used to receive salinity-sensitive variables from ground spectral data, UAV spectral data, and satellite spectral data, respectively, and output corresponding predicted values of soil salinity. The predicted values of soil salinity from each homogeneous learner are merged to form the meta-feature vector of each homogeneous learner. The meta-feature vector of each homogeneous learner is combined with the actual soil salinity value to form meta-training data, which serves as the input to the meta-model. The output of the meta-model is the comprehensive predicted value of soil salinity. The specific steps for forming the meta-feature vectors of each homogeneous learner include: determining the total prediction variance of the base layer learner by comprehensively considering the data source quality uncertainty parameter and the model fitting uncertainty parameter. The formula used to calculate the data source quality uncertainty parameter is as follows: In the formula, The uncertainty parameter of the data source quality is input to the base layer learner for group b. The signal-to-noise ratio of the data input to the base layer learner for group b. The cloud coverage rate of the satellite imagery is input to the base layer learner for group b, specifically the percentage of the imagery obscured by clouds. This refers to the perspective of the drone's orbital observation, specifically the angle between the sensor carried by the drone and the ground normal. b is the maximum observation angle; is the index of the input data set. The uncertainty parameter for model fitting is specifically the standard deviation between the predicted and actual values of the base layer learner. The formula used to comprehensively determine the total prediction variance of the base layer learners is as follows: In the formula, Fit uncertainty parameters to the model using the base layer learner data input in group b. and Here are the weighting coefficients, where and and All are greater than 0; Based on the total variance of the predictions from the base layer learners, meta-training data is generated using an attention-weighted approach. The specific formula for calculating the attention weights is as follows: In the formula, The attention weights for the b-th input data. This represents the uncertainty penalty coefficient. By using attention weights to fuse input data from different training rounds, a meta-feature vector is obtained. Input the salinity-sensitive variables from ground spectral data, UAV spectral data, and satellite spectral data into the trained soil salinity prediction model to determine the soil salinity of all image pixels in the target area.
9. A soil salinity inversion system based on multi-source remote sensing data using a stacked model, characterized by the meta-model: The aforementioned system for inverting soil salinity from multi-source remote sensing data based on a stacked model is used to execute the method for inverting soil salinity from multi-source remote sensing data based on a stacked model as described in any one of claims 1-8, comprising: The sample data acquisition module is used to define the spatial range of the research target area. With its boundary as a constraint, several target monitoring points are systematically set up within the research target area. The expected information gain of each target monitoring point is calculated using the Bayesian Skrigin interpolation algorithm. Several target monitoring points are selected as sample sampling points based on the expected information gain. Soil conductivity data of the sample sampling points are extracted, and the soil salinity of each sample sampling point is calculated based on the soil conductivity. The feature spectrum acquisition module is used to collect ground spectral data, UAV spectral data and satellite remote sensing spectral data at each sample sampling point, and resamples the ground spectral data and UAV spectral data based on the satellite remote sensing spectral data as the spectral band alignment reference to align the three types of spectral data in the spectral band range. The feature parameter calculation module is used to extract the spectral reflectance of each sample point in different feature bands based on the aligned spectral data of the three types. It also constructs spectral indices based on the spectral reflectance of each sample point to obtain several spectral feature indices for each sample point. The sensitive variable screening module is used to calculate the Pearson correlation coefficient between the spectral characteristic index and the corresponding soil salinity at the corresponding sample sampling point based on the spectral characteristic index and the corresponding salinity under different types of spectral data through a correlation analysis algorithm, and to screen the salinity sensitive variable based on the Pearson correlation coefficient. The salinity prediction module is used to construct a soil salinity prediction model by using a stacked model, taking salinity-sensitive variables as input to the model, and using the corresponding soil salinity as a label, in order to determine the soil salinity at points within the research target area.
Citation Information
Patent Citations
Coastal saline region soil salinity content estimating method
CN108680509A
Satellite spectrum-based saline soil conductivity estimation method
CN112949681A
Method for improving satellite data inversion soil salinity by using unmanned aerial vehicle data
CN119314035A
Soil organic matter remote sensing inversion method based on stacking model
CN119323170A
Soil water and salt dynamic monitoring and quality evaluation method based on multispectral remote sensing
CN121208306A