Remote sensing monitoring method for soil salinization in semi-arid region based on small sample

By constructing an enhanced training sample set and an endmember elastic perturbation mechanism, the overfitting problem of remote sensing monitoring technology under small sample conditions was solved, the accuracy and reliability of soil salinization monitoring were improved, and efficient monitoring of complex land surfaces was achieved.

CN122265860APending Publication Date: 2026-06-23JILIN JIANZHU UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
JILIN JIANZHU UNIVERSITY
Filing Date
2026-03-16
Publication Date
2026-06-23

AI Technical Summary

Technical Problem

Existing remote sensing monitoring technologies are prone to overfitting and are difficult to generalize under small sample conditions. Traditional physical models are difficult to adapt to the complex and variable nature of the surface spectrum, resulting in poor spatial continuity of soil salinization monitoring results.

Method used

By acquiring multispectral remote sensing images and ground-based measured data, an enhanced training sample set is constructed. Using a fully constrained linear spectral unmixing and kernel support vector regression model, combined with an endmember elastic perturbation mechanism, a soil salinity distribution map is generated.

Benefits of technology

It improves the model's generalization ability under small sample conditions, enhances the inversion accuracy for complex mixed pixel regions, and provides a quality evaluation and post-processing system for inversion results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122265860A_ABST
    Figure CN122265860A_ABST
Patent Text Reader

Abstract

This application relates to the fields of remote sensing image processing and environmental monitoring technology, and discloses a remote sensing monitoring method for soil salinization in semi-arid regions based on small samples. The method extracts regional baseline endmembers and performs fully constrained linear spectral unmixing to obtain the baseline reconstruction error. This error is used to screen high-confidence unlabeled pixels to generate pseudo-labels, constructing an enhanced training sample set. An ensemble regression model capable of outputting predicted mean and variance is trained. For pixels with high predicted variance, an endmember elastic perturbation mechanism is triggered, iteratively optimizing under the condition of satisfying physical reconstruction error constraints to obtain the endmember combination and inversion results that minimize the predicted variance. Finally, the results are integrated to generate a salinity distribution map. This invention solves the problem of small-sample training through a physical-statistical coupling strategy and uses adaptive perturbation to correct spectral variations in mixed pixels, effectively improving the inversion accuracy and reliability of results in complex surface environments.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of remote sensing image processing and environmental monitoring technology, specifically a remote sensing monitoring method for soil salinization in semi-arid regions based on small samples. Background Technology

[0002] Soil salinization in semi-arid regions is a key issue restricting regional agricultural development and damaging the ecological environment. Traditional manual field sampling methods are time-consuming, labor-intensive, and have limited coverage, making them unsuitable for large-scale, dynamic monitoring. Remote sensing technology, with its advantages of macroscopic, rapid, and non-contact detection, has become an important means of obtaining regional soil salinity information.

[0003] Existing remote sensing monitoring technologies mainly fall into two application paths: statistical empirical regression and physical spectral unmixing. Statistical empirical regression focuses on uncovering the correlation between remote sensing band reflectance, vegetation indices, and ground-measured salinity data. It establishes inversion models using algorithms such as multiple linear regression, support vector machines, or random forests to directly map spectral signals to salinity content. Physical spectral unmixing, on the other hand, is based on linear mixture model theory, assuming that pixel spectra are linear combinations of the spectra of basic endmembers such as soil and vegetation. This method solves for the abundance ratio of each endmember, removes vegetation interference, extracts bare soil information, and thus indirectly infers the degree of soil salinization.

[0004] However, existing technologies have limitations in practical applications. Obtaining ground-based measured samples is difficult and costly, resulting in a limited amount of data available for model training. Purely data-driven statistical models are prone to overfitting under small sample conditions, exhibiting strong memorization of training data but weak generalization ability in unknown regions, leading to poor spatial continuity of inversion results. Furthermore, the surface environment is highly heterogeneous, with spectral curves of similar land features drifting due to soil moisture and roughness. Traditional physical models often use globally fixed endmember spectra for calculation, ignoring the spectral variability of endmembers. This rigid assumption makes it difficult for the model to adapt to local microenvironmental changes, resulting in decreased accuracy in mixed pixel decomposition. Therefore, this invention provides a remote sensing monitoring method for soil salinization in semi-arid regions based on small samples to address the shortcomings of existing technologies. Summary of the Invention

[0005] To address the shortcomings of existing technologies, this invention provides a remote sensing monitoring method for soil salinization in semi-arid regions based on small samples. This method solves the problem that existing monitoring technologies suffer from limited training samples due to the high cost of acquiring ground-measured soil salinity data. Existing physical models rely on fixed endmembers, making it difficult to accurately describe the complex and variable surface spectral characteristics and prone to systematic biases. Furthermore, purely data-driven statistical regression models are prone to overfitting under small sample conditions and lack physical interpretability, making it difficult to assess the reliability of prediction results.

[0006] To achieve the above objectives, the present invention provides the following technical solution: a remote sensing monitoring method for soil salinization in semi-arid regions based on small samples, comprising the following steps: Acquire multispectral remote sensing images and ground-measured data of the target monitoring area and preprocess them to generate a surface reflectance matrix and extract the regional baseline endmember spectrum. Fully constrained linear spectral unmixing is performed on the pixels in the surface reflectance matrix based on the regional reference endmember spectrum to obtain the reference abundance vector and reference reconstruction error of the pixels; Based on the baseline reconstruction error, high-confidence unlabeled pixels are selected and pseudo-labels are generated. An enhanced training sample set is constructed by combining the ground-measured data. Multiple kernel support vector regression base models are trained based on an enhanced training sample set, and an ensemble regression model is constructed to output the predicted mean and predicted variance. The features of the pixel to be inverted, including the baseline abundance vector and the baseline reconstruction error, are input into the integrated regression model. When the prediction variance exceeds the threshold, the endmember elastic perturbation mechanism is triggered. Under the condition of satisfying the baseline reconstruction error constraint, the endmember combination that minimizes the prediction variance and the corresponding inversion result are found. By integrating the inversion results of all pixels, a distribution map of soil salinity in the target monitoring area is generated.

[0007] Preferably, the step of acquiring multispectral remote sensing images and ground-measured data of the target monitoring area, performing preprocessing, generating a surface reflectance matrix, and extracting regional reference endmember spectra further includes: Radiometric calibration and atmospheric correction are performed on the multispectral remote sensing image to generate a surface reflectance matrix representing physical validity, and a minimum noise separation transformation is performed on the surface reflectance matrix to construct a feature space. Within the feature space, the frequency of pixels falling into the projection extrema is calculated using the pure pixel index algorithm, high-purity candidate pixels are screened, and the physical morphological characteristics of the spectral curve are analyzed in conjunction with an N-dimensional visualization interactive tool. From the high-purity candidate pixels, bare soil end-members, vegetation end-members, and dark ground feature end-members are established. The reference spectra of the bare soil end-members, vegetation end-members, and dark ground feature end-members are combined to construct an initial reference end-member matrix as the reference end-member spectrum of the target area.

[0008] Preferably, the step of performing fully constrained linear spectral unmixing on the pixels in the surface reflectance matrix based on the regional reference endmember spectrum to obtain the reference abundance vector and reference reconstruction error of the pixels further includes: A linear mixture model is constructed, which expresses the observed spectrum of a pixel as the sum of the linear combination of the reference end-member spectrum of the region and the corresponding abundance vector and the residual vector; Under the conditions of forcibly introducing the abundance non-negativity constraint and the abundance sum being one constraint, the solution of the benchmark abundance vector is transformed into a quadratic programming problem, and the benchmark abundance vector is obtained by using the effective set method. The model reconstructed spectrum is obtained by reprojecting the reference abundance vector using the regional reference endmember spectrum. The root mean square error between the pixel observation spectrum and the model reconstructed spectrum is calculated. The root mean square error is defined as the dynamic reference for judging physical distortion, i.e., the reference reconstruction error.

[0009] Preferably, the step of filtering high-confidence unlabeled pixels based on the benchmark reconstruction error and generating pseudo-labels, and constructing an enhanced training sample set in conjunction with the ground-measured data, further includes: The ground-based measured data is cleaned and spatially matched based on statistical criteria to construct an initial ground-based measured sample set, and a physical confidence screening function is defined. Using the physical confidence screening function, pixels whose baseline reconstruction error is less than the first quartile of the full image reconstruction error and whose abundance summation satisfies the abundance summation tolerance constraint are retained as high-confidence unlabeled pixels. Calculate the spectral similarity between the high-confidence unlabeled pixel and the samples in the initial ground-measured sample set. Add a small amount of Gaussian white noise to the label values ​​of the nearest neighbor samples that meet the similarity threshold and then add them to the high-confidence unlabeled pixel to generate a pseudo-labeled sample. Finally, merge the pseudo-labeled sample with the initial ground-measured sample set.

[0010] Preferably, the step of training multiple kernel support vector regression basis models based on an enhanced training sample set to construct an ensemble regression model that outputs the predicted mean and predicted variance further includes: Construct a hybrid feature vector containing the baseline abundance vector and the baseline reconstruction error, and perform standardization processing on the hybrid feature vector to obtain the training feature set; A bootstrap sampling strategy is used to perform random sampling with replacement from the training feature set to generate multiple independent training subsets, and a kernel support vector regression base model is trained independently for each training subset; The standardized feature vectors of the pixels to be tested are input in parallel to all trained kernel support vector regression base models to obtain a set of predicted values. The arithmetic mean of the set of predicted values ​​is calculated as the predicted mean, and the variance of the set of predicted values ​​is calculated as the predicted variance.

[0011] Preferably, the step of triggering the endmember elastic perturbation mechanism when the prediction variance exceeds the threshold further includes: Monitor the prediction variance output by the integrated regression model. If the prediction variance is greater than a preset consensus threshold, the pixel is determined to be a high-uncertainty pixel, and an elastic perturbation matrix is ​​introduced. The elastic perturbation matrix is ​​superimposed on the regional reference endmember spectrum to construct a dedicated endmember matrix, and the Euclidean norm of each endmember perturbation vector in the elastic perturbation matrix is ​​restricted not to exceed a specific proportion of the reference spectrum norm. Based on the dedicated endmember matrix, the new abundance vector and the new reconstruction error are recalculated, and a global cost function containing a statistical consistency term, a physical fidelity term and a regularization term is constructed. The physical fidelity term is used to constrain the new reconstruction error from significantly exceeding the baseline reconstruction error.

[0012] Preferably, the step of finding the endmember combination that minimizes the prediction variance and the corresponding inversion result under the condition of satisfying the benchmark reconstruction error constraint further includes: The global cost function is iteratively optimized using the particle swarm optimization algorithm. After each position update, the elastic boundary constraint is enforced using the boundary constraint function to search for the optimal perturbation matrix. The final eigenvector is calculated based on the optimal perturbation matrix, and the final eigenvector is input into the ensemble regression model to obtain the final prediction variance. If the final prediction variance is less than or equal to the consensus threshold, the corrected prediction mean is output as the inversion result; if the final prediction variance is greater than the consensus threshold, the cell is marked as an anomalous feature and a null mask is output.

[0013] Preferably, the step of integrating the inversion results of all pixels to generate a soil salinity distribution map of the target monitoring area further includes: Based on the row and column index of the original remote sensing image, the discrete inversion results are remapped into a two-dimensional spatial matrix, and a three-channel tensor structure containing a salinity matrix, an uncertainty matrix, and a state mask matrix is ​​constructed. Based on the ratio of the inversion results to the predicted variance, a normalized confidence index is constructed to generate a quality assessment layer for quantifying data quality. Read the geographic transformation parameters and projection coordinate system information of the original remote sensing image, write the geographic transformation parameters into the header file of the salinity matrix and the quality assessment layer, and output a soil salinity distribution map with geographic attributes.

[0014] Preferably, after generating the soil salinity distribution map of the target monitoring area, the method further includes the following steps: Discrete encoding is performed on the pixel states in the state mask matrix, and the pixel states include high confidence regions, correction regions, and abnormal regions; An adaptive median filter is constructed, wherein the adaptive median filter dynamically adjusts the noise judgment conditions according to the pixel state, and a relaxed deviation threshold is used for smoothing in the high confidence area; A strict deviation threshold is applied to the correction area to preserve texture details, and the abnormal area is forcibly retained as null value, thus completing the denoising and fidelity preservation processing of the salt content matrix.

[0015] Preferably, the global cost function is constructed by weighted summation of statistical consistency term, physical fidelity term and regularization term; The statistical consistency term is used to represent the degree of deviation of the new prediction variance from the preset target variance, so that the endmember parameters evolve in the direction of reaching a consensus in the integrated regression model; The physical fidelity term is used to represent the degree of deviation of the new reconstruction error from the reference reconstruction error, so as to ensure that the perturbed endmember combination can restore the observed spectrum; The regularization term, based on the Frobenius norm of the elastic perturbation matrix, is used to penalize the perturbation magnitude in order to preferentially select the solution closest to the baseline state.

[0016] This invention provides a remote sensing monitoring method for soil salinization in semi-arid regions based on small samples. It has the following beneficial effects: 1. This invention utilizes the baseline reconstruction error extracted by fully constrained linear spectral unmixing as a physical constraint index to screen high-confidence unlabeled samples from images and generate pseudo-labels, constructing an enhanced training sample set. In situations where ground-based measured data is scarce, the prior knowledge of the physical model expands the feature distribution of the training data, mitigating the overfitting phenomenon that easily occurs when training statistical regression models under small sample conditions, and improving the model's generalization ability to unsampled regions.

[0017] 2. This invention targets regions where the prediction variance of the integrated regression model exceeds a threshold. Under the premise of satisfying the physical reconstruction error constraint, it adaptively fine-tunes the regional reference endmember spectrum. This overcomes the shortcomings of traditional methods that use fixed endmember spectra to accurately characterize the spatial heterogeneity of the land surface. It can find the endmember combination that minimizes the prediction variance while maintaining reasonable physical meaning, thereby improving the inversion accuracy of complex mixed pixel regions.

[0018] 3. This invention establishes a quality evaluation and post-processing system for inversion results by integrating a regression model to simultaneously output the predicted mean and predicted variance. The adaptive median filter constructed based on the predicted variance can dynamically adjust the filtering strategy according to the pixel's confidence level, effectively suppressing random noise while preserving surface texture details in high-confidence areas. This solves the problems of traditional inversion methods, such as difficulty in evaluating result reliability and difficulty in balancing denoising and fidelity preservation. Attached Figure Description

[0019] Figure 1 This is a system architecture diagram of the present invention; Figure 2 This is a flowchart of the method steps of the present invention; Figure 3 This is a schematic diagram of the end-member elastic perturbation and adaptive inversion logic of the present invention; Figure 4 This invention provides a soil salinization grading diagram and quality identification diagram for the Hetao Irrigation District. Figure 5 Figure 1 is a schematic diagram of the comparative test data of the present invention; wherein, Figure (a) is the spectrum of the extracted regional reference endmember, Figure (b) is a comparison diagram of salinity inversion accuracy, and Figure (c) is a diagram of the adaptive inversion optimization convergence process. Detailed Implementation

[0020] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0021] See attached document Figure 1 , Figure 1 This is a system architecture diagram according to an embodiment of the present invention. The present invention provides a remote sensing monitoring system for soil salinization in semi-arid regions based on small samples. The system includes: a data preprocessing module, a baseline feature extraction module, a sample enhancement module, an integrated modeling module, an adaptive inversion module, and a mapping and quality control module.

[0022] The data preprocessing module receives multispectral remote sensing image data and measured soil salinity data of the target monitoring area. This module performs radiometric calibration and atmospheric correction on the multispectral remote sensing image data to generate a surface reflectance image matrix. Simultaneously, it cleans and performs spatial coordinate matching on the measured soil salinity data to construct an initial measured sample set. The output of the data preprocessing module is connected to the inputs of both the baseline feature extraction module and the sample enhancement module.

[0023] The baseline feature extraction module is used to extract regional baseline end-member spectra from the surface reflectance image matrix based on the pure pixel index algorithm. The regional baseline end-member spectra include bare soil end-members, vegetation end-members, and dark ground cover end-members. This baseline feature extraction module uses a fully constrained linear spectral unmixing algorithm to calculate the baseline abundance vector and baseline reconstruction error of each pixel in the image, and establishes the physical baseline state data of the pixels.

[0024] The sample augmentation module receives the surface reflectance image matrix and the initial ground-measured sample set. This module maps the spectral data of the image to a high-dimensional Hilbert space (or feature space) and calculates the spectral similarity (such as kernel spectral angle or Euclidean distance) between unlabeled pixels and samples in the initial ground-measured sample set. High-confidence samples are selected based on a preset similarity threshold, and pseudo-labels are generated to construct an augmented training sample set.

[0025] The integrated modeling module connects to the sample augmentation module and the baseline feature extraction module. This integrated modeling module employs a bootstrap sampling strategy to resample the augmented training sample set, training multiple independent kernel support vector regression base models. This module constructs an ensemble regression model, which receives the abundance vector of pixels and the reconstruction error, and outputs the predicted mean and variance of soil salinity.

[0026] The adaptive inversion module, connected to both the baseline feature extraction module and the ensemble modeling module, monitors the prediction variance of the ensemble regression model. When the prediction variance exceeds a preset threshold, it triggers an endmember elastic perturbation mechanism. Under the premise of satisfying the physical reconstruction error constraint, this adaptive inversion module adjusts the input endmembers through an iterative optimization method to find the endmember combination that minimizes the prediction variance of the ensemble regression model, and outputs the optimized inversion result.

[0027] The mapping and quality control module, connected to the adaptive inversion module, maps the inversion results of the entire region into a raster image, generating a distribution map of soil salinization levels. Simultaneously, this module generates a quality label map containing baseline confidence, corrected confidence, and low confidence classifications based on the processing status of each pixel in the adaptive inversion module.

[0028] See attached document Figure 2 , Figure 2 This is a flowchart of a method according to an embodiment of the present invention. The present invention provides a remote sensing monitoring method for soil salinization in semi-arid regions based on small samples, comprising the following steps: S1. Preprocessing is performed using multispectral remote sensing images and ground-measured data to generate a surface reflectance matrix and extract regional baseline endmember spectra. S2, perform fully constrained linear spectral unmixing on the pixels in the surface reflectance matrix to obtain the reference abundance vector and reference reconstruction error of the pixels; S3 maps the image spectral data to the kernel space and calculates spectral similarity, filters high-confidence unlabeled pixels to generate pseudo-labels, and constructs an enhanced training sample set; S4, based on the enhanced training sample set, trains multiple kernel support vector regression base models to construct an ensemble regression model that can output the predicted mean and predicted variance; S5. Input the features of the pixel to be inverted into the integrated regression model. When the prediction variance exceeds the threshold, trigger the endmember elastic perturbation. Under the condition of satisfying the baseline reconstruction error constraint, find the endmember combination that minimizes the prediction variance and the corresponding inversion result. S6 integrates the inversion results of all pixels to generate a soil salinity distribution map and an inversion quality label map of the target area.

[0029] In this embodiment, step S1 mainly undertakes the tasks of cleaning and correcting multi-source monitoring data and initializing the physical model. Due to the strong heterogeneity of the surface and the variable content of atmospheric aerosols in semi-arid regions, directly using the original images will lead to systematic biases in the inversion results. Therefore, it is necessary to construct a standardized surface reflectance dataset and establish a regionally representative spectral endmember library based on it.

[0030] To ensure the physical validity of remote sensing inversion, the acquisition of monitoring data must adhere to strict spatiotemporal matching principles. When selecting Sentinel-2 MSI or Landsat-8 OLI multispectral imagery of the target monitoring area, imagery data with cloud cover below 10% and an imaging time interval of no more than 15 days between the ground sampling time should be preferred to minimize errors introduced by changes in land cover and atmospheric transients.

[0031] For the original multispectral imagery, this embodiment employs the FLAASH algorithm based on a radiative transfer model for processing. This process not only converts the dimensionless digital quantization (DN) values ​​into physically meaningful surface reflectance but also compensates for aerosol scattering effects common in semi-arid regions, constraining pixel values ​​to... Within the physically effective range. The corrected image data is reconstructed into a two-dimensional surface reflectance matrix. The mathematical expression of this matrix is ​​as follows: ; in, This indicates the number of spectral bands in the image (e.g., the 10 effective bands of Sentinel-2). This indicates the total number of valid pixels within the image coverage area. Indicates the first Each pixel in Spectral reflectance vectors in each band Representation matrix dimensional space, This represents the vector transpose sign. Here, we need to consider the matrix... Singular value detection is performed. If there are column vectors with all zeros or invalid padding values, they are removed before subsequent calculations to prevent numerical overflow.

[0032] Simultaneously, for ground-measured soil salinity data, in addition to the conventional GPS coordinate projection transformation, data cleaning based on statistical criteria is also required. Specifically, the Laida criterion (…) is adopted. (Principle) Outlier sampling points deviating from the mean by more than 3 standard deviations are removed. Based on the spatial resolution of the image, the coordinates of the sampling points are matched with the coordinates of the image pixel centers to construct a nearest neighbor database. The initial ground-measured sample set of each sample ,in This corresponds to the measured value of soil electrical conductivity. Indicates the first Image metadata corresponding to each sample Indicates the sample's sequence index. This indicates the total number of valid samples contained in the sample set.

[0033] Given the prevalence of mixed pixels, establishing accurate "clean" endmembers is crucial for unmixing the physical model. This embodiment does not directly use the standard spectral library, but instead employs the Clean Pixel Index (PPI) algorithm to adaptively extract endmembers from the current image to eliminate sensor response differences. Considering the band correlation and noise interference commonly found in high-dimensional spectral data, the surface reflectance matrix is ​​first analyzed before performing PPI. Perform minimum noise separation (MNF) transformation. By analyzing the eigenvalue distribution, retain the top-ranked components with a signal-to-noise ratio (SNR) greater than a threshold (e.g., 1.5). Using MNF components as the feature space, random noise can be effectively suppressed to prevent interference with endmember recognition.

[0034] In the high-dimensional simplex geometry constructed in the MNF feature space, pure pixels are typically located at the vertices of the data cloud. By generating a large number of random vectors and projecting data points onto these vectors, the frequency (PPI count) of each pixel falling into the projection extremum is counted. In this embodiment, a PPI count threshold is set (e.g., retaining the top 5% of pixels with the highest PPI counts) to filter out high-purity candidate pixels. Combined with an N-dimensional visualization interactive tool, the following three types of baseline endmembers are established based on the physical morphological characteristics of the spectral curves: Bare Soil End Yuan The mean value of regions with high PPI counts and normalized differential vegetation index (NDVI) less than 0.1 in the image was selected. Their physical spectral characteristics show a continuous upward trend in reflectance from the visible light to the shortwave infrared band, with no significant water absorption band.

[0035] Vegetation end element The average PPI count and NDVI greater than 0.6 were selected from salt-tolerant vegetation areas in the image. Their spectral curves needed to exhibit a "red edge" effect, i.e., a strong absorption valley near 670 nm and a strong reflection peak near 800 nm, to ensure sensitivity to halophyte vegetation.

[0036] Dark-colored ground features : Select the spectral mean of the water body or the deep shaded region. In the physical model, this endmember not only represents the water body, but also acts as a "light absorber" to absorb the spectral brightness attenuation caused by topographic shading and soil moisture during subsequent linear unmixing, preventing it from being incorrectly attributed to changes in soil composition.

[0037] Based on the above extraction results, an initial baseline end-member matrix is ​​constructed. This initial baseline endmember matrix is ​​not merely a data concatenation, but also the physical reference origin for all subsequent perturbation operations. Defined as: ; The column vectors of the matrix correspond to the reference spectra of bare soil, vegetation, and dark ground features, respectively. To ensure the numerical stability of the subsequent FCLSU algorithm, it is necessary to... Perform a condition number check; if the condition number is too large (e.g., exceeding 10),... 3 This indicates severe collinearity among endmembers (e.g., similar spectra between bare soil and withered grass). In this case, it is necessary to fine-tune the endmember selection in N-dimensional space until the matrix satisfies the benignity requirement.

[0038] In this embodiment, step S2 aims to establish the baseline physical state of the observed pixels. By introducing a fully constrained linear spectral mixture model, the system deconstructs the high-dimensional spectral data into endmember abundance distributions with clear geoscientific significance, and calculates the baseline reconstruction error to quantify the linear model's interpretability of the current surface scene. This error index will subsequently serve as a rigid constraint boundary for determining whether endmember elastic perturbations have undergone "physical drift".

[0039] Based on the baseline endmember matrix extracted in step S1 This embodiment employs a linear mixture model to describe the radiative response mechanism of surface mixed pixels in semi-arid regions. This linear mixture model is based on the macroscopic mixture assumption, which states that photons within the sensor's instantaneous field of view primarily undergo single reflections, and surface reflectance can be considered as the area-weighted sum of the reflectances of bare soil, vegetation, and dark ground features. For any observed pixel in the image... Its mathematical expression is: ; in, for A dimensional baseline matrix; This is the corresponding abundance vector, representing the area proportion of each endmember within the current pixel; This is the residual vector, which includes sensor thermal noise and small spectral variations not included in the model. This represents the spectral reflectance vector of any observed pixel in the image.

[0040] To ensure that the unmixing results conform to the actual physical state of the Earth's surface, this embodiment forcibly introduces two physical constraints in the model solution: Abundance Nonnegativity Constraint (ANC): Set (For all) This excludes non-physical cases involving negative areas; Abundance sum-of-ones constraint (ASC): Set This ensures that the sum of the proportions of each component closes within the total field of view of the pixel.

[0041] Under the above constraints, solve for the baseline abundance vector. The process is transformed into a standard quadratic programming (QP) problem. This is because the objective function's Hessian matrix... Given that the endmembers are linearly independent and the matrix is ​​a positive semi-definite matrix, this optimization problem belongs to the category of convex optimization, which theoretically guarantees the existence and uniqueness of the global optimal solution.

[0042] ; ; in, It is a 3×3 identity matrix, where 1 represents a vector consisting entirely of 1s. This represents the value of the variable that minimizes the objective function. Represents the reference end-member matrix, This represents the abundance vector to be solved.

[0043] In practical implementation, to prevent matrix Due to excessively high endmember spectral similarity leading to ill-conditioning, the condition number needs to be pre-calculated. If the condition number exceeds a preset threshold (e.g., 10), further calculation is required. 4 Then, the Tikhonov regularization term is introduced. ( Take the smaller value, such as 10 6 The solution process is stable and numerically sound. The preferred algorithm is the effective set method, which has higher computational efficiency and numerical stability than the interior point method when dealing with small-scale quadratic programming problems with inequality constraints, and can accurately handle boundary cases with an abundance of 0.

[0044] Obtain baseline abundance Then, the model is reprojected using the reference endmember matrix to obtain the reconstructed spectrum. At this point, the goodness of fit of the baseline physical model to the current real observation data is quantified by calculating the root mean square error (RMSE), denoted as . : ; in, Total number of bands ( ), and Representing the observed pixel and the first pixel respectively The terminal in the first Reflectivity of the band No. The area percentage of each end-cell within the current pixel.

[0045] This indicates the degree to which the current pixel deviates from the "ideal linear mixing assumption," representing the baseline reconstruction error. In semi-arid salinized regions, the lower... This indicates that the pixels are mainly composed of a mixture of simple ground features; while the abnormally high This often suggests the presence of complex nonlinear radiative transfer processes in the area (such as multiple scattering from salt crystals) or the existence of anomalous features not included in the end-member library. Therefore, this embodiment... Defined as a dynamic benchmark for judging "physical distortion": any optimization operation aimed at reducing prediction uncertainty must not produce a reconstruction error that significantly exceeds [the threshold]. The defined physical confidence region.

[0046] In this embodiment, step S3 constructs and expands the training sample set. Given the high cost of ground field sampling in semi-arid regions and the extremely limited number of measured samples (typically only a few dozen), directly training the regression model is highly susceptible to overfitting. Therefore, this step establishes a "physically guided-statistical augmentation" sample construction mechanism. Utilizing the physical baseline features obtained in step S2, high-confidence samples conforming to physical laws are selected from a massive number of unlabeled pixels, forming an enhanced training set together with the measured data.

[0047] Before building a basic sample library, it is necessary to solve the problem of non-correspondence between ground point measurement data and remote sensing image pixels in the spatiotemporal dimensions.

[0048] For the time dimension, this embodiment sets a strict time matching window. Only images taken before and after the remote sensing image imaging time were selected. Within a day (preferred in this embodiment) Ground samples were collected over 7 days. The reason for choosing this threshold is that the surface salinity of soil in semi-arid regions is highly dynamic due to the influence of rainfall and evaporation. An interval of more than 7 days may cause the surface spectral characteristics to "decouple" from the measured salinity values, introducing uncontrollable label noise.

[0049] Regarding the spatial dimension, considering the positioning error of handheld GPS devices (typically 3-5 meters) and the geometric correction residuals of images, the value of a single pixel may not accurately represent the habitat of the sampling point.

[0050] For the Each ground sampling point, and its corresponding image feature vector By using the coordinates of that point ( Extraction of a 3×3 neighborhood window centered on ) ; in, The pixel features (including abundance and physical residuals) extracted in step S2. and This represents the relative position offset within the neighborhood window. This represents the coordinates of the neighboring cells involved in the calculation. This represents the mean filtering coefficient. This mean filtering operation not only smooths out sensor noise but also physically simulates the spatial integration effect of mixed pixels, ensuring data scale consistency.

[0051] To augment the sample in the absence of real labels, this embodiment filters unlabeled pixels from the image as "pseudo-samples," based on the physical reconstruction error generated by the fully constrained linear unmixing in step S2. .

[0052] Define the physical confidence screening function : ; in, This represents the physical reconstruction error of the pixel calculated in step S2. This represents the sum of abundance values. To determine the reconstruction error threshold, this embodiment sets it to the first quartile (Q1) of the full-image reconstruction error, which means only the top 25% of pixels with the best physical model fit are retained. The abundance summation tolerance is set to 0.05. Only when the physical model can reconstruct the observed spectrum with minimal error does its eigenvector have value as a training sample.

[0053] The selected high-confidence pixels still lack salt labels. In this embodiment, a "feature space neighborhood propagation" strategy is used to generate pseudo labels.

[0054] For any unlabeled cell that passes the screening Find the measured sample with the closest Euclidean distance (or spectral angular distance) in the MNF feature space or the normalized spectral space. If the distance between the two Less than the preset similarity threshold Then it is considered that the pixel With sample They are in similar physical and chemical environments.

[0055] The value is set to the 10th percentile of the pairwise distance distribution among all measured samples to ensure that only pixels with extremely similar spectra can inherit the label.

[0056] At this time, the pixel Add it to the training set and assign it a pseudo-label. : ; in, The true salinity value is the nearest neighbor measured sample. It is Gaussian white noise. This represents the variance of the Gaussian distribution.

[0057] Here, a small amount of noise is introduced. (Values) The 1% (1%) is to prevent multiple spurious samples from completely replicating the same label, which could cause the regression model to "collapse" or overfit at specific numerical points. This is achieved by introducing data perturbation to enhance the generalization robustness of the model.

[0058] The measured sample set after spatiotemporal matching With the generated pseudo-label sample set Merge and construct the final enhanced training sample set. .

[0059] Before inputting the data into the subsequent model, the data must be Z-score standardized to eliminate the influence of dimensions and ensure that features of different physical quantities have equal weights during the optimization process.

[0060] For the feature matrix Each column in (i.e., the first column) (Each feature dimension) is used to first calculate its global statistical distribution parameters. Feature mean With characteristic standard deviation The definition is as follows: ; ; in, The total number of samples, For the first The sample at the th The original numerical values ​​of the dimensions. Based on the above statistical parameters, the following standardization transformation is performed on each element of the matrix to obtain the input vector: ; in, These are the standardized eigenvalues; To prevent the denominator from approaching 0, a numerical stability constant (taken as 10 in this embodiment) is used. -6 ), used to handle the extreme case where the variance is zero due to completely identical eigenvalues.

[0061] For the statistical parameter set calculated in step S3 This will be solidified and saved as an inherent property of the model, and passed as a system constant to the inference stage in step S6. When subsequently inverting new images, the same set of properties must be strictly applied. Standardization is performed to ensure that the data distribution in the training space is aligned with that inference space.

[0062] See attached document Figure 3 In this embodiment, step S4 is used to construct an integrated regression system with uncertainty quantification capabilities. This system aims to overcome the shortcomings of single regression models, which are prone to overfitting and cannot perceive their own prediction risks under semi-arid and complex surface conditions. By deeply fusing the physical baseline features extracted in step S2 with the enhanced sample set expanded in step S3, a nonlinear mapping from "physical state" to "soil salinity" is established, and statistical differences are used to constrain the confidence level of the inversion results.

[0063] In the preparation phase of regression model training, it is necessary to construct an input feature vector that can comprehensively characterize the state of surface salinization. In this embodiment, a hybrid feature vector with clear physical meaning is constructed. For the enhanced training sample set For any sample in the dataset, its input is defined as a concatenation of the physical abundance vector and the baseline reconstruction error: ; in, Represents a mixed feature vector. Represents an abundance vector. Indicates the transpose symbol. Indicates the baseline reconstruction error. The generated vector It contains 4 numerical components.

[0064] The abundance of bare soil and vegetation directly responds to changes in surface material composition under salt stress, while As a residual of a linear physical model, it implicitly contains nonlinear spectral distortion and micro-topographic texture information caused by salt crystallization. The combination of the two allows the model to learn both macroscopic compositional patterns and microscopic residual correction information.

[0065] Given that Support Vector Regression (SVR) involves distance metrics and is highly sensitive to scale differences in features, Z-score normalization is performed before inputting feature vectors into the model. Specifically, the mean of each feature dimension in the training set is calculated. and standard deviation ( ), and transform the data. ,in represents the original feature value, and represents the standardized feature value. This process is consistent with the standardization logic described in step S3, and the above statistics... These parameters need to be permanently saved as model parameters, and the same set must be applied when retrieving new image pixels in subsequent operations. Standardization is performed to ensure the consistency of the feature distribution space.

[0066] To achieve statistical estimation of model prediction uncertainty, this embodiment introduces a bootstrap sampling mechanism to artificially introduce data perturbation and set the ensemble size parameter. (This embodiment is preferred) That is, to construct 100 independent base learners, for the th Base Model From the total sample size of Perform random sampling with replacement, generating samples of the same size. training subset .

[0067] According to statistical principles, each subset On average, the samples contain approximately 63.2% original unique samples, with the remaining approximately 36.8% unselected. This random variation in sample composition forces the various base models to form slightly different decision boundaries in the feature space, thus laying the data foundation for quantifying uncertainty through the degree of divergence between models.

[0068] For each training subset In this embodiment, a kernel support vector regression (KSVR) base model is trained independently. KSVR (Kinetic-Simplified Dynamics Research) exhibits excellent generalization ability by mapping low-dimensional nonlinear problems to a high-dimensional feature space and finding the optimal hyperplane. Its solution can be used to handle the following constrained optimization problems: ; ; in, Indicates the bias term. Indicates the current training subset The total number of samples in the sample, Indicates the sample index. and They represent the first The true label value and input feature vector of each sample Represents a nonlinear mapping function. For the weight vector, As slack variables, to ensure the model's prediction accuracy and robustness, this embodiment employs a five-fold cross-validation combined with a grid search strategy to jointly optimize the following key hyperparameters: Penalty coefficient The search range is set to [2]. -5 ,twenty one 5 [This is used to balance model complexity and training error; larger ones...] Focus on training accuracy, smaller Pay attention to generalization ability.

[0069] nuclear parameters : Coordinating radial basis kernel function Use, among which Represents the radial basis kernel function. This represents the feature vector, and the search range is set to [2]. -15 ,2 3 [This controls the looseness of the sample distribution in high-dimensional space.]

[0070] Insensitive loss coefficient Set to 0.1 times the standard deviation of the training labels to ignore small prediction errors and enhance sparsity.

[0071] After solving the dual problem using the Sequence Minimum Optimization (SMO) algorithm, we obtain... A trained base model .

[0072] In the inversion stage, for the standardized feature vector of the pixel to be measured... Input it in parallel to In each base model, a set of predicted values ​​is obtained. , Indicates the first A pre-trained set of base model functions. Based on this set, the system outputs the predicted mean of soil salinity. As the final inversion result, the prediction variance is calculated. As a confidence indicator: ; ; in, This represents the ensemble forecast mean. This represents the total number of base models. Indicates the base model index. Indicates the first The predicted values ​​of the base models This represents the prediction variance. It is calculated here. This quantifies the unfamiliarity of the current input data within the model's knowledge domain. If Below the preset consensus threshold (For example This indicates that the base models have reached a consensus, and the results are reliable; if This means that the input features fall into the sparse region of the training samples, and the statistical model fails.

[0073] In this embodiment, step S5 constitutes the control feedback loop of the entire inversion system. This step addresses the "high uncertainty" (i.e., prediction variance) marked in step S4 by constructing a two-way coupling mechanism between the physical model and the statistical model. This addresses the challenging pixel inversion problem. When the statistical model exhibits high divergence regarding the physical features of the current input, it is inferred that the actual endmember spectrum of that pixel may have undergone local drift relative to the baseline state of the entire map (e.g., influenced by soil moisture content or micro-topographical shading). Therefore, the system proactively initiates an "endmember elastic perturbation" process, fine-tuning the endmember parameters within the elastic domain allowed by physical laws, searching for the optimal solution that simultaneously satisfies "minimum physical reconstruction error" and "maximum statistical prediction consensus."

[0074] For the target pixel to be optimized, this embodiment introduces an elastic perturbation matrix. This is used to characterize the local spatiotemporal variation of the endmember spectrum. At this point, the dedicated endmember matrix for that pixel... Defined as the baseline end-member matrix Superposition with the disturbance term: ; in, The endmember matrix representing the target pixel. Represents the reference end-member matrix, Here is the perturbation matrix. Indicates the first The perturbation vector of each endmember.

[0075] To prevent algorithms from generating "mathematical pseudo-solutions" that violate geoscientific common sense in order to cater to data (such as distorting vegetation spectra into water spectra), strict physical elastic boundary constraints must be imposed on the perturbation amplitude. Specifically, this embodiment restricts the Euclidean norm of each endmember perturbation vector from exceeding a specific proportion of its reference spectral norm. : ; in, Indicates the first The perturbation vector of each endmember This represents the proportional coefficient of the elastic constraint. Indicates the first The reference spectral vector of each endmember, The symbol is a mathematical notation, meaning that the constraint applies to all three endmembers simultaneously.

[0076] In this embodiment, parameters The value range is set to [0.05, 0.1]. This threshold setting is based on the fact that the spectral variation of similar land features in nature (caused by soil texture, vegetation water content, etc.) usually does not exceed 10% of its intrinsic spectral energy; if it exceeds this range, it should be regarded as a change in land feature category rather than a variation of the same type.

[0077] The optimal perturbation matrix is ​​found through an optimization algorithm. To evaluate the merits of any candidate perturbation matrix, this embodiment constructs a non-convex global cost function. This function deeply couples the goodness of fit of the physical model with the confidence level of the statistical model. The calculation process is as follows: Will Substituting the fully constrained linear spectral unmixing (FCLSU) algorithm described in step S2, solve for the new abundance vector for this pixel. and physical reconstruction error ; Construct the updated feature vector Then input it into the ensemble SVR model trained in step S4 to obtain the new prediction variance. Based on the above intermediate variables, the total cost function is defined as follows: ; in, Represents the total cost function. Represents the endmember perturbation matrix. To prevent small constants with a denominator of zero (take...) ), Indicates the target variance. This represents the new prediction variance. Denotes the square of the Frobenius norm of a matrix. It is the Frobenius norm; For the statistical consistency term (first term): with the pre-defined target variance Based on this, the end-member parameters are forced to evolve in a direction in which the integrated model can reach a consensus. This is a channel for back-injecting data-driven prior knowledge into the physical model.

[0078] For the physical fidelity term (second term): based on the initial reference error Based on this, we can ensure that the perturbed endmember combination can still accurately reproduce the observed spectrum, and prevent the failure of physical fitting due to excessive pursuit of statistical consistency.

[0079] For regularization (the third term): According to Occam's razor, when the effects are similar, excessive perturbations are penalized, and the solution closest to the baseline state is selected first.

[0080] Weighting: The weighting coefficient is a dimensionless coefficient, and in this embodiment, it is preferably set to... The primary objective of this step is to reduce the uncertainty of the forecast (caused by...). (Dominant), secondly, to maintain physical rationality (by (constraints), and finally the magnitude of the control variables (by...) (Fine-tuning).

[0081] Given the objective function The internal nesting of the FCLSU quadratic programming solver and the SVR nonlinear regression model leads to its dependence on the independent variable. Not only is it non-convex, but its gradient is also non-analytical, rendering traditional gradient descent methods inapplicable. Therefore, this embodiment employs a particle swarm optimization algorithm for the search, with specific parameter configurations tailored to this scenario.

[0082] In practice, the particle swarm size is set. Maximum number of iterations Each particle Position matrix This represents a potential perturbation matrix. Iterative updates follow the following speed With position Evolutionary equation: ; ; in, Indicates the first Individual particles The velocity matrix at the next iteration Indicates the first A particle in the current The velocity matrix at the next iteration This represents the individual's historical best position. Indicates the first A particle in the current The position matrix at the next iteration This indicates the best position in the global history.

[0083] Inertia weight The learning factor is linearly reduced from 0.9 to 0.4 using a linear decreasing strategy. This setting gives the algorithm strong global exploration ability in the early stages, preventing it from getting trapped in local optima; and strong local exploitation ability in the later stages, improving the accuracy of the solution. All are set to 2.0, balancing particles towards their own historical best. and group historical best Learning step size; boundary constraint function After each position update, the aforementioned flexible boundary constraints are enforced. If the perturbation amplitude in a certain dimension exceeds the limit, it is truncated to the boundary value. .

[0084] After the iterative optimization is completed, the perturbation matrix corresponding to the globally optimal particle is extracted. And calculate the final feature vector. At this point, the final logical decision is made: the calculation is based on... Integrated prediction variance .

[0085] like This indicates that, after adaptive fine-tuning, the physical model has successfully matched the characteristics of the observed data, and the statistical model has established a high-confidence understanding of the sample. At this point, the corrected predicted mean is output. As a result of the inversion.

[0086] like This indicates that even with maximum physical elastic perturbation, the pixel cannot be reasonably interpreted by the current model system. This situation usually corresponds to non-soil features (such as man-made structures or water bodies) or extremely anomalous areas in the image. In this case, the system marks the pixel as an "anomalous feature" and outputs a null mask for manual interpretation.

[0087] In this embodiment, step S6 serves as the output terminal of the entire inversion system, performing data reconstruction from "discrete pixel extrapolation" to "continuous spatial mapping." This step aims to map the point-like inversion results output in step S5 back to geospatial space and simultaneously generate a visualized quality assessment layer. This is achieved by constructing a multidimensional raster dataset that includes predicted values, statistical uncertainties, and processing status.

[0088] After completing the adaptive inversion of all effective pixels in the image, the system first uses the row and column indices of the original remote sensing image. The discrete inversion sequence stored in memory is remapped into a two-dimensional spatial matrix. This embodiment constructs a three-channel tensor structure as the output carrier, specifically comprising the following three independent matrix layers: Salt content matrix : Stores the final predicted mean for each pixel location This matrix directly reflects the spatial distribution pattern of surface soil salinization.

[0089] Uncertainty matrix : Store the prediction standard deviation of the corresponding pixel This matrix quantifies the inference dispersion of the model under different land cover conditions and serves as the underlying basis for evaluating data reliability.

[0090] State mask matrix This records the final processing status of each pixel in the inversion process. To facilitate subsequent metadata analysis, this embodiment employs the following discrete encoding logic: Code0 (Background Area): Non-soil areas (such as water bodies or impermeable surfaces) were not involved in the inversion.

[0091] Code1 (High Reliability Region): Pixels that pass the variance test directly in step S4 represent that the model has mature prior knowledge of this type of habitat.

[0092] Code2 (Correction Area): Pixels that meet the convergence condition after the endmember elastic perturbation is performed in step S5, representing difficult samples recovered through physical adaptive correction.

[0093] Code3 (Outlier Zone): Pixels whose variance still exceeds the limit after S5 optimization represent ground features that the current model system cannot explain and require manual intervention.

[0094] To provide users with an intuitive and horizontally comparable quality metric across the entire map, this embodiment constructs a quality evaluation layer based on the ratio of predicted values ​​to uncertainties. .

[0095] Considering that simple coefficients of variation are prone to divergence in low-value regions and lack defined upper and lower bounds, this embodiment designs a normalized confidence index to perform physical nonnegativity constraint processing. For any valid pixel... Its quality rating The calculation is as follows: ; in, To prevent extremely small constants with denominators of zero (taken as 10) -6 Coefficient 2 corresponds to the 95% confidence interval ( The statistical assumptions of ) Represents a cell The final predicted value, Represents a cell The standard deviation of the prediction This represents the numerical stability constant.

[0096] By using the signal-to-noise ratio (SNR) concept from analog signal processing, the unbounded variance is mapped to... The bounded interval. When the prediction standard deviation When it approaches 0, Approaching 100% indicates an extremely high confidence level; when Much larger than predicted hour, A value close to 0 indicates that the results are overwhelmed by noise. This metric provides a uniform standard for measuring data quality across regions with different salinity levels.

[0097] Because pixel-by-pixel inversion ignores the spatial continuity of ground features, isolated numerical abrupt changes (i.e., salt-and-pepper noise) often exist in the original result image. Therefore, this embodiment addresses this issue. Perform adaptive median filtering based on state masks. The response logic of the filter's kernel function dynamically depends on... The status codes in the text are used to balance the conflict between "denoising" and "fidelity preservation," for... Calculate the median of the 3×3 neighborhood centered at the center. and the median absolute deviation (MAD): If the central pixel status code is 1 (high confidence area): the judgment condition is lenient. As long as the pixel value is consistent with... If the deviation exceeds twice the MAD, it is judged as noise and used as... Replacement to smooth out random errors.

[0098] If the center pixel status code is 2 (correction zone): the judgment condition is strict. Since the value of this point is obtained through high-cost optimization, it implies special micro-terrain features. Replacement is only performed when the deviation exceeds 3 times MAD, so as to preserve the real texture details to the maximum extent.

[0099] If the central cell status code is 3 (abnormal area): no interpolation padding is performed, and it is forcibly retained as a null value (NaN). This strategy aims to prevent erroneous interpolation from masking the true anomalous properties of the surface (such as forest clearings or various man-made structures) and to ensure the accuracy of patch boundaries.

[0100] To assign geographic attributes to the matrix and make it compatible with GIS systems, georeferencing is required. Those skilled in the art can use the GDAL library to read the six geographic transformation parameters (including the latitude and longitude of the top left corner, pixel resolution, and rotation angle) and projection coordinate system information (such as WGS84UTM zonal projection) of the original remote sensing image.

[0101] The system writes the aforementioned geographic metadata into the reconstructed... , and The header file is ultimately packaged into a multi-band GeoTIFF format file, outputting a thematic map of soil salinity in semi-arid regions that combines physical accuracy and statistical reliability.

[0102] To better understand the technical solution of this invention, the following description is based on a specific application scenario.

[0103] Remote sensing data: Sentinel-2A multispectral image, imaged on April 15, 2025 (after spring irrigation, low crop cover, and obvious surface salinity accumulation), cloud cover <5%.

[0104] Ground-based measured data: Fifty surface soil samples (0-20cm) were collected during the same period, and the soil electrical conductivity (EC) values ​​were obtained by testing, ranging from 0.5dS / m to 25dS / m.

[0105] Specific implementation steps: 1. Data Preprocessing and Benchmark Establishment: The system receives Sentinel-2 images and obtains a surface reflectance matrix of 10 bands after FLAASH atmospheric correction. Using MNF noise reduction and PPI algorithm, three benchmark endmembers are extracted: bare soil endmember: the average spectrum (high reflectance) of severely saline-alkali wasteland is selected; vegetation endmember: the spectrum of a small number of early-maturing salt-tolerant plants (such as red willow) is selected; dark endmember: the spectrum of water bodies in the Yellow River irrigation canal is selected.

[0106] Fully constrained linear spectral unmixing (FCLSU) was performed to obtain the baseline reconstruction error map of the entire image. The results showed that the reconstruction error (RMSE) of 90% of the regions was less than 0.02, but the error of some mixed pixels was relatively large.

[0107] 2. Sample Augmentation: Since there are only 50 actual samples, the system calculates the similarity between unlabeled pixels and actual samples in the kernel feature space, and selects 200 high-confidence pixels with low reconstruction error and spectral similarity <0.1 (Euclidean distance) to generate pseudo-labels. The final training set is expanded to 250 samples (50 actual samples + 200 pseudo-labels).

[0108] 3. Ensemble modeling: Bootstrap resampling was used to train 100 kernel support vector regression (K-SVR) base models. The models output the predicted mean and the predicted variance (uncertainty) on the validation set.

[0109] 4. Adaptive Inversion Optimization: For mixed pixels (transition zones) located at the edge of farmland, the initial prediction variance of the ensemble model is as high as 0.15 (exceeding the threshold of 0.05), indicating that the model is inaccurate. The system triggers an endmember elastic perturbation mechanism, and the particle swarm optimization (PSO) algorithm fine-tunes the endmember spectrum within a range of ±10%.

[0110] After 30 iterations, while keeping the increase in physical reconstruction error Rnew below 5%, the prediction variance was reduced to 0.03. The corrected salinity value was adjusted from the initial 8.5 dS / m to 12.1 dS / m, which is more consistent with the actual situation in the transition zone.

[0111] 5. Mapping and Output: Generate a soil salinization classification map of the Hetao Irrigation District (non-salinized, slightly, moderately, heavily, and saline soil). Simultaneously generate a quality label map, with red areas marked as "requiring manual verification" (abnormal areas), accounting for only 2% of the entire map.

[0112] Experimental Verification: To verify the effectiveness of this embodiment, it was compared with three traditional methods in an experiment: Comparative Example A (SR-SI): Traditional linear regression method based on spectral indices (such as the salinity index SI). Comparative Example B (SVR): Standard SVR model trained directly using the original measured samples (50 samples), without physical constraints. Comparative Example C (FCLSU): Linear unmixing based solely on the physical model, directly using the abundance of salinity endmembers to represent content.

[0113] Evaluation Metric: R 2 (Coefficient of determination): the higher the better; RMSE (Root Mean Square Error): the lower the better.

[0114] The experimental data are shown in Table 1: Table 1: Comparison data of experimental methods ; According to the appendix Figure 4 Appendix Figure 5 The experimental data in Table 1 show that: The monitoring system proposed in this embodiment outperforms traditional methods in terms of both inversion accuracy and stability. By combining physical model constraints with statistical sample enhancement strategies, the system effectively overcomes the common problem of model overfitting under conditions of scarce ground-measured samples, increasing the determination coefficient of the inversion results to 0.89 and significantly reducing errors. This method can effectively distinguish the spectral differences between water bodies, healthy vegetation, and saline soil, eliminate interference from environmental factors such as soil moisture and artificial facilities, and improve the accuracy and reliability of monitoring results in complex mixed surface environments.

Claims

1. A remote sensing monitoring method for soil salinization in semi-arid regions based on small samples, characterized in that, Includes the following steps: Acquire multispectral remote sensing images and ground-measured data of the target monitoring area and preprocess them to generate a surface reflectance matrix and extract the regional baseline endmember spectrum. Fully constrained linear spectral unmixing is performed on the pixels in the surface reflectance matrix based on the regional reference endmember spectrum to obtain the reference abundance vector and reference reconstruction error of the pixels; Based on the baseline reconstruction error, high-confidence unlabeled pixels are selected and pseudo-labels are generated. An enhanced training sample set is constructed by combining the ground-measured data. Multiple kernel support vector regression base models are trained based on an enhanced training sample set, and an ensemble regression model is constructed to output the predicted mean and predicted variance. The features of the pixel to be inverted, including the baseline abundance vector and the baseline reconstruction error, are input into the integrated regression model. When the prediction variance exceeds the threshold, the endmember elastic perturbation mechanism is triggered. Under the condition of satisfying the baseline reconstruction error constraint, the endmember combination that minimizes the prediction variance and the corresponding inversion result are found. By integrating the inversion results of all pixels, a distribution map of soil salinity in the target monitoring area is generated.

2. The remote sensing monitoring method for soil salinization in semi-arid regions based on small samples according to claim 1, characterized in that, The process of acquiring multispectral remote sensing images and ground-measured data of the target monitoring area, performing preprocessing, generating a surface reflectance matrix, and extracting regional baseline endmember spectra further includes: Radiometric calibration and atmospheric correction are performed on the multispectral remote sensing image to generate a surface reflectance matrix representing physical validity, and a minimum noise separation transformation is performed on the surface reflectance matrix to construct a feature space. Within the feature space, the frequency of pixels falling into the projection extrema is calculated using the pure pixel index algorithm, high-purity candidate pixels are screened, and the physical morphological characteristics of the spectral curve are analyzed in conjunction with an N-dimensional visualization interactive tool. From the high-purity candidate pixels, bare soil end-members, vegetation end-members, and dark ground feature end-members are established. The reference spectra of the bare soil end-members, vegetation end-members, and dark ground feature end-members are combined to construct an initial reference end-member matrix as the reference end-member spectrum of the target area.

3. The remote sensing monitoring method for soil salinization in semi-arid regions based on small samples according to claim 1, characterized in that, The step of performing fully constrained linear spectral unmixing on pixels in the surface reflectance matrix based on regional reference endmember spectra to obtain the reference abundance vector and reference reconstruction error of the pixels further includes: A linear mixture model is constructed, which expresses the observed spectrum of a pixel as the sum of the linear combination of the reference end-member spectrum of the region and the corresponding abundance vector and the residual vector; Under the conditions of forcibly introducing the abundance non-negativity constraint and the abundance sum being one constraint, the solution of the benchmark abundance vector is transformed into a quadratic programming problem, and the benchmark abundance vector is obtained by using the effective set method. The model reconstructed spectrum is obtained by reprojecting the reference abundance vector using the regional reference endmember spectrum. The root mean square error between the pixel observation spectrum and the model reconstructed spectrum is calculated. The root mean square error is defined as the dynamic reference for judging physical distortion, i.e., the reference reconstruction error.

4. The remote sensing monitoring method for soil salinization in semi-arid regions based on small samples according to claim 1, characterized in that, The process of filtering high-confidence unlabeled pixels based on benchmark reconstruction error and generating pseudo-labels, combined with the ground-measured data to construct an enhanced training sample set, further includes: The ground-based measured data is cleaned and spatially matched based on statistical criteria to construct an initial ground-based measured sample set, and a physical confidence screening function is defined. Using the physical confidence screening function, pixels whose baseline reconstruction error is less than the first quartile of the full image reconstruction error and whose abundance summation satisfies the abundance summation tolerance constraint are retained as high-confidence unlabeled pixels. Calculate the spectral similarity between the high-confidence unlabeled pixel and the samples in the initial ground-measured sample set. Add a small amount of Gaussian white noise to the label values ​​of the nearest neighbor samples that meet the similarity threshold and then add them to the high-confidence unlabeled pixel to generate a pseudo-labeled sample. Finally, merge the pseudo-labeled sample with the initial ground-measured sample set.

5. The remote sensing monitoring method for soil salinization in semi-arid regions based on small samples according to claim 1, characterized in that, The method of training multiple kernel support vector regression basis models based on an enhanced training sample set, and constructing an ensemble regression model that outputs the predicted mean and predicted variance, further includes: Construct a hybrid feature vector containing the baseline abundance vector and the baseline reconstruction error, and perform standardization processing on the hybrid feature vector to obtain the training feature set; A bootstrap sampling strategy is used to perform random sampling with replacement from the training feature set to generate multiple independent training subsets, and a kernel support vector regression base model is trained independently for each training subset; The standardized feature vectors of the pixels to be tested are input in parallel to all trained kernel support vector regression base models to obtain a set of predicted values. The arithmetic mean of the set of predicted values ​​is calculated as the predicted mean, and the variance of the set of predicted values ​​is calculated as the predicted variance.

6. The remote sensing monitoring method for soil salinization in semi-arid regions based on small samples according to claim 1, characterized in that, The step of triggering the endmember elastic perturbation mechanism when the prediction variance exceeds the threshold further includes: Monitor the prediction variance output by the integrated regression model. If the prediction variance is greater than a preset consensus threshold, the pixel is determined to be a high-uncertainty pixel, and an elastic perturbation matrix is ​​introduced. The elastic perturbation matrix is ​​superimposed on the regional reference endmember spectrum to construct a dedicated endmember matrix, and the Euclidean norm of each endmember perturbation vector in the elastic perturbation matrix is ​​restricted not to exceed a specific proportion of the reference spectrum norm. Based on the dedicated endmember matrix, the new abundance vector and the new reconstruction error are recalculated, and a global cost function containing a statistical consistency term, a physical fidelity term and a regularization term is constructed. The physical fidelity term is used to constrain the new reconstruction error from significantly exceeding the baseline reconstruction error.

7. The remote sensing monitoring method for soil salinization in semi-arid regions based on small samples according to claim 1, characterized in that, The step of finding the endmember combination that minimizes the prediction variance and the corresponding inversion result under the condition of satisfying the benchmark reconstruction error constraint further includes: The global cost function is iteratively optimized using the particle swarm optimization algorithm. After each position update, the elastic boundary constraint is enforced using the boundary constraint function to search for the optimal perturbation matrix. The final eigenvector is calculated based on the optimal perturbation matrix, and the final eigenvector is input into the ensemble regression model to obtain the final prediction variance. If the final prediction variance is less than or equal to the consensus threshold, the corrected prediction mean is output as the inversion result; if the final prediction variance is greater than the consensus threshold, the cell is marked as an anomalous feature and a null mask is output.

8. The remote sensing monitoring method for soil salinization in semi-arid regions based on small samples according to claim 1, characterized in that, The step of integrating the inversion results of all pixels to generate a distribution map of soil salinity in the target monitoring area further includes: Based on the row and column index of the original remote sensing image, the discrete inversion results are remapped into a two-dimensional spatial matrix, and a three-channel tensor structure containing a salinity matrix, an uncertainty matrix, and a state mask matrix is ​​constructed. Based on the ratio of the inversion results to the predicted variance, a normalized confidence index is constructed to generate a quality assessment layer for quantifying data quality. Read the geographic transformation parameters and projection coordinate system information of the original remote sensing image, write the geographic transformation parameters into the header file of the salinity matrix and the quality assessment layer, and output a soil salinity distribution map with geographic attributes.

9. The remote sensing monitoring method for soil salinization in semi-arid regions based on small samples according to claim 8, characterized in that, After generating the soil salinity distribution map of the target monitoring area, the following steps are also included: Discrete encoding is performed on the pixel states in the state mask matrix, and the pixel states include high confidence regions, correction regions, and abnormal regions; An adaptive median filter is constructed, wherein the adaptive median filter dynamically adjusts the noise judgment conditions according to the pixel state, and a relaxed deviation threshold is used for smoothing in the high confidence area; A strict deviation threshold is applied to the correction area to preserve texture details, and the abnormal area is forcibly retained as null value, thus completing the denoising and fidelity preservation processing of the salt content matrix.

10. The remote sensing monitoring method for soil salinization in semi-arid regions based on small samples according to claim 6, characterized in that, The global cost function is constructed by weighted summation of statistical consistency term, physical fidelity term, and regularization term; The statistical consistency term is used to represent the degree of deviation of the new prediction variance from the preset target variance, so that the endmember parameters evolve in the direction of reaching a consensus in the integrated regression model; The physical fidelity term is used to represent the degree of deviation of the new reconstruction error from the reference reconstruction error, so as to ensure that the perturbed endmember combination can restore the observed spectrum; The regularization term, based on the Frobenius norm of the elastic perturbation matrix, is used to penalize the perturbation magnitude in order to preferentially select the solution closest to the baseline state.