A hyperspectral soil component inversion method and system for heterogeneous small samples
Patent Information
- Application Number
- CN202611023353.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-10
- Publication Date
- 2026-10-09
AI Technical Summary
[0003]本发明的主要目的是提供一种面向异构小样本的高光谱土壤成分反演方法及系统,旨在解决传统方法适应性不足的问题
[0015]效果:形成一条稳健的任务适配路径。既充分利用了预训练学到的光谱先验知识,又通过真实样本主导保证了预测的可靠性,最后引入筛选后的辅助样本弥补了分布不均的缺陷,最终训练出高性能、高泛化能力的单目标成分反演模型。
Smart Images

Figure CN122889147A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of agricultural information technology, specifically to a method and system for inverting hyperspectral soil composition for heterogeneous small samples. Background Technology
[0002] Currently, hyperspectral soil composition inversion utilizes hyperspectral remote sensing data of soil to establish a quantitative mapping relationship between spectral characteristics and the content of target components such as heavy metals and minerals in the soil, which is a core technology for precision agriculture and environmental monitoring. However, Existing methods, whether based on statistical learning, struggle to capture complex nonlinear relationships, or based on deep learning, perform poorly under data scarcity. Conventional data augmentation methods (such as resampling and noise perturbation) cannot specifically supplement samples under scarce conditions, resulting in data that differs significantly from the true distribution and is prone to introducing noise. Therefore, a hyperspectral soil composition inversion method and system for heterogeneous small samples is needed to address these issues. Summary of the Invention
[0003] The main objective of this invention is to provide a method and system for inverting hyperspectral soil composition for heterogeneous small samples, aiming to solve the problem of insufficient adaptability of traditional methods.
[0004] The technical solution proposed in this invention is: a method for inverting hyperspectral soil composition for heterogeneous small samples, comprising the following steps: S1: Using geospatial location as an index, establish a unified data interface covering observation modes from satellites, drones, and ground laboratories to obtain raw spectral data. Map the raw spectral data to a common band space and perform standardization and normalization processing to obtain a standardized dataset. S2: Construct a pre-trained model based on a two-layer masking mechanism using a standardized dataset to obtain a spectral characterization backbone network. The pre-trained model has a dedicated encoder architecture and a multi-dimensional reconstruction loss function. S3: Establish a directional sample augmentation method based on a conditional generation model, and construct a targeted and spectrally realistic auxiliary spectral library based on the directional sample augmentation method. The conditional generation model is based on the dedicated encoder architecture to encode real spectra into high-level features. S4: Based on the aforementioned spectral characterization backbone network as a feature extractor, and connected to a regression prediction head, a soil composition inversion model is constructed; the soil composition inversion model is trained using a three-stage progressive transfer fine-tuning strategy. Phase 1: Establishing the mission boundaries; Phase Two: Controlled Adaptation of High-Level Representations; Phase 3: Fine-tuning with real and auxiliary samples.
[0005] Preferably, step S3 includes: The condition generation model receives the encoded vector of the descriptive condition unit (such as "high concentration, forest land, satellite mode"), generates spectral features that meet the descriptive conditions, and then decodes them into auxiliary spectra in a unified band space.
[0006] Preferably, the auxiliary spectra are screened in multiple dimensions, including: Spectral similarity screening: Calculate the spectral angular distance and cosine similarity between the generated auxiliary spectrum and the average spectrum of the real spectrum under the same conditions. If the similarity exceeds the second threshold, it is discarded.
[0007] Preferably, multi-dimensional screening of auxiliary spectra also includes: Condition consistency screening: Use a small discriminator to verify whether the concentration range predicted by the generated auxiliary spectrum is consistent with the target conditions.
[0008] Preferably, multi-dimensional screening of auxiliary spectra also includes: Feature space proximity filtering: In the pre-trained feature space, the features of the generated sample must be sufficiently close to the features of the nearest real sample in the corresponding conditional unit.
[0009] Preferably, multi-dimensional screening of auxiliary spectra also includes: Diversity control screening: Limiting the number of samples retained within the same conditional unit.
[0010] Preferably, the training of the soil composition inversion model using a three-stage progressive migration fine-tuning strategy includes: Completely freeze all parameters of the pre-trained backbone network and train the regression head using only real labeled samples.
[0011] Specifically, by utilizing the most reliable real supervision signals, a preliminary and stable prediction boundary is quickly determined for the model.
[0012] Preferably, the step of training the soil composition inversion model using a three-stage progressive migration fine-tuning strategy further includes: Unfreeze the last few layers (e.g., 2) of the pre-trained backbone network that are responsible for high-level feature aggregation (such as the last two global modules), and continue fine-tuning them using only real samples, along with the regression head. In this step, parameters are added to maintain regularization to prevent the pre-trained representation from being corrupted.
[0013] The pre-trained general spectral features are then adapted in a controlled manner to specific soil composition inversion tasks, guided by real data.
[0014] Preferably, the step of training the soil composition inversion model using a three-stage progressive migration fine-tuning strategy further includes: Simultaneously, real samples and generated auxiliary samples are used for collaborative fine-tuning, while always maintaining the real samples as the dominant ones (e.g., real: auxiliary = 3:1). Implement differentiated supervision: For real samples, use strict regression loss (such as Huber loss); for auxiliary samples, while using regression loss, add interval consistency constraints between the soft labels of auxiliary samples and the corresponding predicted concentration intervals to acknowledge the non-determinism of the soft labels of auxiliary samples. The loss weights for the auxiliary samples were gradually increased from 0 to 0.4 during training.
[0015] Results: A robust task adaptation path is formed. It fully utilizes the spectral prior knowledge learned from pre-training, ensures the reliability of predictions by using real samples, and finally introduces selected auxiliary samples to compensate for uneven distribution, ultimately training a high-performance, high-generalization-ability single-objective component inversion model.
[0016] A hyperspectral soil composition inversion system for heterogeneous small samples, employing the method as described in any one of claims 1-9; comprising: The data consistency module is used for: multi-source heterogeneous data consistency processing: using geospatial location as an index, a unified data interface covering satellite, UAV and ground laboratory observation modes is established to obtain raw spectral data, the raw spectral data is mapped to a common band space and standardized and normalized to obtain a standardized dataset; The self-supervised pre-training module is used to: construct a pre-trained model based on a two-layer masking mechanism from a standardized dataset to obtain a spectral characterization backbone network, wherein the pre-trained model has a dedicated encoder architecture and a multidimensional reconstruction loss function; An auxiliary sample generation module is used to: establish a directional sample enhancement method based on a conditional generation model, and construct a targeted and spectrally realistic auxiliary spectral library based on the directional sample enhancement method, wherein the conditional generation model is based on the dedicated encoder architecture to encode real spectra into high-level features; The transfer fine-tuning training module is used to: construct a soil composition inversion model based on the spectral characterization backbone network as a feature extractor and connected to a regression prediction head; and train the soil composition inversion model using a three-stage progressive transfer fine-tuning strategy. Phase 1: Establishing the mission boundaries; Phase Two: Controlled Adaptation of High-Level Representations; Phase 3: Fine-tuning with real and auxiliary samples.
[0017] The above technical solution can achieve the following beneficial effects: By establishing a standardized data processing flow, this invention unifies spectra and labels from different sources into a comparable common band and rule system, laying a data foundation for joint modeling. We designed an innovative self-supervised pre-training method to learn a general spectral representation that is sensitive to local weak response features, robust and transferable from massive unlabeled spectra, in order to alleviate the label scarcity problem. By using a conditional generation model, auxiliary samples that are spectrally realistic and meet the conditions are generated for scarce conditional units in labeled data under the combination of dimensions such as "concentration range, background type, and observation mode", instead of blindly expanding the data, thus effectively improving the balance of the training set distribution. The design employs a three-stage migration fine-tuning strategy: first, freeze the model to establish boundaries; then, thaw it to adapt to higher levels; and finally, introduce auxiliary information in a collaborative manner. This strategy ensures that the model achieves an optimal balance between general representations, core supervision signals, and auxiliary information, ultimately enabling high-precision and high-stability soil composition inversion. Attached Figure Description
[0018] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the structures shown in these drawings without creative effort.
[0019] Figure 1 This is a flowchart of the first embodiment of the hyperspectral soil composition inversion method for heterogeneous small samples proposed in this invention; Figure 2 Flowchart for the process of standardizing multi-source heterogeneous spectra and labels; Figure 3 A diagram of a self-supervised pre-training framework for weak response features; Figure 4 Generate an overall framework diagram for the auxiliary spectral model based on the conditions. Detailed Implementation
[0020] It should be understood that the specific embodiments described herein are for illustrative purposes only and are not intended to limit the scope of the invention.
[0021] This invention proposes a method and system for inverting hyperspectral soil composition for heterogeneous small samples.
[0022] Example 1 As attached Figure 1 As shown, the hyperspectral soil composition inversion method for heterogeneous small samples proposed in this invention includes the following steps: S1: Using geospatial location as an index, establish a unified data interface covering observation modes from satellites, drones, and ground laboratories to obtain raw spectral data. Map the raw spectral data to a common band space and perform standardization and normalization processing to obtain a standardized dataset. Specifically, a unified index and data interface are established: using geospatial location (point or region) as the index, an observation modality set covering "satellite (S), UAV (U), and ground laboratory (G)" is established, and the modality availability of each data point is marked by a mask vector to achieve unified management of multi-source data.
[0023] The spectral normalization process involves determining the common effective wavelength range (e.g., 466–940 nm) for all platform data and mapping all original spectra to this unified wavelength space using methods such as resampling and interpolation. During this process, quality control measures such as signal-to-noise ratio calibration, smoothing, and standard normal transformation are performed on the original data for each mode.
[0024] Label normalization: Unit conversion, detection limit processing, and outlier removal are performed on ingredient labels from various sources to ensure that labels from different sources are physically comparable, and the effectiveness of each ingredient is identified using a label availability mask.
[0025] This integrates heterogeneous data such as satellite imagery, UAV imagery, laboratory spectroscopy, and external databases into a standardized data system with a unified format, strong comparability, and clear structure, providing reliable input for all subsequent models.
[0026] S2: Construct a pre-trained model based on a two-layer masking mechanism using a standardized dataset to obtain a spectral characterization backbone network. The pre-trained model has a dedicated encoder architecture and a multi-dimensional reconstruction loss function. Specifically, to address the problem of difficulty in learning effective features under scarce labels, we utilize large-scale unlabeled spectra for self-supervised pre-training to learn general spectral representations.
[0027] The core idea is to learn the essential features (such as absorption peaks, continuous spectrum shape, etc.) contained in the spectral sequence by having a pre-trained model "guess" the masked part of the spectrum without relying on any artificial labels, with particular attention to the local structure related to weak responses.
[0028] Among them, innovative training strategies include: 1. Two-layer masking design: The unified spectrum is divided into continuous "spectral tokens". The first layer, the "primary mask", masks multiple consecutive tokens (e.g., 40%), forcing the model to learn cross-band contextual dependencies; the second layer, the "secondary mask", randomly masks a small number of bands (e.g., 15%) within the visible tokens, forcing the model to focus on local, fine-grained spectral structures, thereby enhancing the ability to capture weak response signals.
[0029] 2. Dedicated encoder architecture: Employs a hybrid encoder combining local convolution and global state space. Local convolutional layers (1D CNN) are used to extract local absorption / reflection features; global state space layers (such as SSM, Transformer) model the long-term dependencies of the entire spectral sequence. The combination of the two aims to simultaneously address weak local responses and global spectral shape.
[0030] 3. Multidimensional reconstruction loss function: Spectral reconstruction loss: measures the model's ability to reconstruct the obscured portion.
[0031] Derivative Consistency Loss: Constrains the consistency of the first derivative (local slope) of the reconstructed spectrum with that of the original spectrum, enhancing the modeling of subtle changes in spectral shape.
[0032] Continuity constraint loss: suppresses non-physical high-frequency noise in the reconstructed spectrum to ensure a smooth and reasonable spectral shape.
[0033] Multi-view consistency loss: encourages the encoded feature representation to remain consistent after the same spectrum undergoes different transformations and masks, thereby improving the stability of the model representation.
[0034] Results: A high-quality spectral characterization backbone network is obtained, which is extremely sensitive to local fine-grained changes in the spectrum (i.e. weak response signals), providing a powerful initialization for downstream small-sample tasks.
[0035] S3: Establish a directional sample augmentation method based on a conditional generation model, and construct a targeted and spectrally realistic auxiliary spectral library based on the directional sample augmentation method. The conditional generation model is based on the dedicated encoder architecture to encode real spectra into high-level features. S4: Based on the aforementioned spectral characterization backbone network as a feature extractor, and connected to a regression prediction head, a soil composition inversion model is constructed; the soil composition inversion model is trained using a three-stage progressive transfer fine-tuning strategy. Phase 1: Establishing the mission boundaries; Phase Two: Controlled Adaptation of High-Level Representations; Phase 3: Fine-tuning with real and auxiliary samples.
[0036] Preferably, step S3 includes: The condition generation model receives the encoded vector of the descriptive condition unit (such as "high concentration, forest land, satellite mode"), generates spectral features that meet the descriptive conditions, and then decodes them into auxiliary spectra in a unified band space.
[0037] Preferably, the auxiliary spectra are screened in multiple dimensions, including: Spectral similarity screening: Calculate the spectral angle distance and cosine similarity between the generated auxiliary spectrum and the average spectrum of the real spectrum under the same conditions. If the values exceed the second threshold (the second threshold is: spectral angle distance (SAD) less than 0.15 and cosine similarity greater than 0.90), they are removed.
[0038] Preferably, multi-dimensional screening of auxiliary spectra also includes: Condition Consistency Screening: A small discriminator is used to verify whether the concentration range predicted by the generated auxiliary spectrum is consistent with the target condition (the target condition is the concentration range value. For example, if the condition unit corresponding to the generated auxiliary spectrum is "high concentration", its target concentration range can be set to [80, 120] mg / kg. During screening, a pre-trained small discriminator network is used to predict the generated spectrum. If the predicted concentration value falls within this range, the spectrum passes the screening.
[0039] Preferably, multi-dimensional screening of auxiliary spectra also includes: Feature space proximity screening: In the pre-trained feature space, the features of the generated sample must be sufficiently close to the features of the nearest real sample in the corresponding conditional unit ("sufficiently close" is specifically defined as: in the feature space encoded by the pre-trained model, the Euclidean distance between the generated auxiliary sample features and all real sample features in the corresponding conditional unit must be less than 1.5 times the average Euclidean distance between the real sample features in that conditional unit).
[0040] Preferably, multi-dimensional screening of auxiliary spectra also includes: Diversity control screening: Limit the number of samples retained within the same condition unit (the limit is: for the same scarce condition unit, the final number of auxiliary samples retained shall not exceed 3 times the number of real samples in that condition unit, and shall not exceed 50).
[0041] Specifically, by screening auxiliary spectra from multiple dimensions, a high-quality, targeted, and realistic auxiliary spectrum library is constructed, which accurately supplements the sparse "shortcomings" in the original training set, rather than indiscriminately amplifying the dataset.
[0042] Preferably, the training of the soil composition inversion model using a three-stage progressive migration fine-tuning strategy includes: Completely freeze all parameters of the pre-trained backbone network and train the regression head using only real labeled samples.
[0043] Specifically, by utilizing the most reliable real supervision signals, a preliminary and stable prediction boundary is quickly determined for the model.
[0044] Preferably, the step of training the soil composition inversion model using a three-stage progressive migration fine-tuning strategy further includes: Unfreeze the last few layers (e.g., 2) of the pre-trained backbone network that are responsible for high-level feature aggregation (such as the last two global modules), and continue fine-tuning them using only real samples, along with the regression head. In this step, parameters are added to maintain regularization to prevent the pre-trained representation from being corrupted.
[0045] The pre-trained general spectral features are then adapted in a controlled manner to specific soil composition inversion tasks, guided by real data.
[0046] Preferably, the step of training the soil composition inversion model using a three-stage progressive migration fine-tuning strategy further includes: Simultaneously, real samples and generated auxiliary samples are used for collaborative fine-tuning, while always maintaining the real samples as the dominant ones (e.g., real: auxiliary = 3:1). Implement differentiated supervision: For real samples, use strict regression loss (such as Huber loss); for auxiliary samples, while using regression loss, add interval consistency constraints between the soft labels of auxiliary samples and the corresponding predicted concentration intervals to acknowledge the non-determinism of the soft labels of auxiliary samples. The loss weights for the auxiliary samples were gradually increased from 0 to 0.4 during training.
[0047] Results: A robust task adaptation path is formed. It fully utilizes the spectral prior knowledge learned from pre-training, ensures the reliability of predictions by using real samples, and finally introduces selected auxiliary samples to compensate for uneven distribution, ultimately training a high-performance, high-generalization-ability single-objective component inversion model.
[0048] A hyperspectral soil composition inversion system for heterogeneous small samples, applying the aforementioned method; comprising: The data consistency module is used for: multi-source heterogeneous data consistency processing: using geospatial location as an index, a unified data interface covering satellite, UAV and ground laboratory observation modes is established to obtain raw spectral data, the raw spectral data is mapped to a common band space and standardized and normalized to obtain a standardized dataset; The self-supervised pre-training module is used to: construct a pre-trained model based on a two-layer masking mechanism from a standardized dataset to obtain a spectral characterization backbone network, wherein the pre-trained model has a dedicated encoder architecture and a multidimensional reconstruction loss function; An auxiliary sample generation module is used to: establish a directional sample enhancement method based on a conditional generation model, and construct a targeted and spectrally realistic auxiliary spectral library based on the directional sample enhancement method, wherein the conditional generation model is based on the dedicated encoder architecture to encode real spectra into high-level features; The transfer fine-tuning training module is used to: construct a soil composition inversion model based on the spectral characterization backbone network as a feature extractor and connected to a regression prediction head; and train the soil composition inversion model using a three-stage progressive transfer fine-tuning strategy. Phase 1: Establishing the mission boundaries; Phase Two: Controlled Adaptation of High-Level Representations; Phase 3: Fine-tuning with real and auxiliary samples.
[0049] Example 2 As attached Figure 2 As shown, this is a flowchart of the multi-source heterogeneous spectrum and label unification process. The unification process revolves around four levels: unified point indexing, unified spectral structure, unified labeling rules, and overall dataset organization. For paired sample points in the study area, a unified index is established using the sampling point as the basic unit, and its observation modality availability is recorded. For external spectral libraries and regional unlabeled samples, their original sample organization forms are retained and incorporated into a unified data expression framework after unification. Secondly, intramodal quality control, cross-platform common band mapping, and numerical scale standardization are implemented for multi-source spectra to ensure comparability of spectra from different platforms in the same band space. Then, unit conversion, detection limit processing, outlier control, and availability masking are implemented for multi-component labels to ensure that supervisory information can be organized under the same rules. Finally, point-level labeled samples and regional unlabeled remote sensing spectra are jointly incorporated into a unified real data system.
[0050] Example 3 As attached Figure 3 As shown, this is a diagram of a self-supervised pre-training framework for weak response features. Building upon the traditional self-supervised reconstruction approach, this paper further constructs a self-supervised pre-trained model for weak response feature extraction, providing a foundation for adaptation to downstream inversion tasks. The training framework proposes a two-layer masking strategy, an encoder-decoder architecture based on local convolutional enhancement and global sequence modeling, and a customized reconstruction loss. It learns a stable, transferable, and sensitive unified representation backbone network to local fine-grained spectral shape changes on a large number of real unlabeled spectra.
[0051] Example 4 As attached Figure 4 As shown, this is the overall framework diagram of the conditional generation auxiliary spectral model. First, discrete conditional units are defined in the set of labeled spectrum-soil composition pairs, and corresponding conditional description vectors are constructed. Then, the pre-trained encoder is used to map the real spectrum to the high-level representation space, and the latent variable distribution is learned under conditional constraints. Finally, controllable auxiliary spectral samples are generated through conditional decoding, conditional consistency discrimination, and feature manifold screening, and an auxiliary spectral library for subsequent supervised enhancement is constructed.
[0052] Example 5 Taking remote sensing inversion of the heavy metal zinc (Zn) in soil as an example: 1. Data preparation and consistency: The experiment was conducted using paired satellite-UAV sampling points in the study area, regional unlabeled remote sensing spectra, and two external ground-based spectral libraries. It integrated data from 67 monitoring points in the study area (including Zhuhai-1 OHS satellite, UAV spectra, and Zn content labels), 279 external Zn-Pb-Cd database samples, and over 80 million regional unlabeled remote sensing spectra.
[0053] Externally labeled data include 279 subsamples from the Zn-Pb-Cd soil spectral database and 19,967 ground spectral samples from the LUCAS Topsoil Survey.
[0054] Regarding unlabeled data, the two Zhuhai-1 OHS images contain a total of 51,131,328 satellite spectral samples. After deducting 67 paired points, this results in 51,131,261 regional-level unlabeled satellite spectral samples. The four UAV images, after being converted according to a unified coverage area and resolution, contain approximately 30,674,513 spectral samples. After deducting 67 paired points, this results in approximately 30,674,446 regional-level unlabeled UAV spectral samples. The aforementioned unlabeled samples are mainly used for self-supervised pre-training. The 67 paired sample points in the study area are mainly used for supervised verification and downstream task adaptation. The two external ground spectral libraries are mainly used for label supplementation, introduction of spectroscopic priors, and auxiliary sample construction.
[0055] All spectra were uniformly mapped to 466-940nm (69 common bands), and the Zn label unit was uniformly set to mg / kg. Missing values and outliers were processed to form a standardized dataset.
[0056] Several new indices for uniformity analysis are introduced, namely spectral angular distance (SAD), to measure the similarity of two spectra in the overall direction. The smaller the value, the closer the spectral shapes are.
[0057] In addition to the overall test set metrics, two specific evaluation subsets were set up: a high-value / boundary region test subset and a scarce background category test subset. The former is used to test the model's fitting ability in high-concentration areas, boundary areas, and extreme value areas, while the latter is used to test the model's generalization stability in weakly covered background categories such as woodland, grassland / shrubland, etc.
[0058] 2. Self-supervised pre-training: Using over 80 million unlabeled remote sensing spectral data (after sampling), a model incorporating a two-layer mask and a local-global hybrid encoder was constructed and trained.
[0059] The AdamW optimizer was used, with a cosine decay learning rate, and the training lasted for 200 epochs.
[0060] After training, the encoder is fixed and saved as a general spectral characterization backbone for subsequent tasks.
[0061] 3. Build the Zn auxiliary library: Analysis of the distribution of 346 valid Zn-tagged samples revealed that the "conditional unit" of "high concentration, forest background, satellite observation" had only 3 real samples, far below the threshold (e.g., 8), and was marked as a scarce unit.
[0062] Using this condition as a constraint, 50 candidate spectra are generated.
[0063] The screening process included calculating the spectral angular distance to the average spectrum of three real samples and checking using an interval discriminator. Approximately 15 high-quality auxiliary spectra were ultimately retained, and a reliable "soft" Zn concentration label (e.g., 40 mg / kg) was generated for each auxiliary spectrum by combining the statistical information of the real label from its parent unit. These were then added to the Zn-specific auxiliary library.
[0064] 4. Three-stage training of the Zn inversion model: Phase I: Initialize the model with a pre-trained encoder and freeze all its parameters, then train the regression for the first 30 rounds using only 70% of the real Zn samples.
[0065] Phase II: Unfreeze the last two layers of the encoder, still using only real samples, and continue training for 50 rounds with a moderate decay of the learning rate and the addition of weight decay constraints.
[0066] Phase III: Mix real samples and auxiliary samples for Zn at a ratio of approximately 3:1. Real samples are supervised using Huber loss, while auxiliary samples are supervised using MSE loss combined with soft labels. Gradually increase the weight of auxiliary samples in the total loss (from 0.1 to 0.4) and co-fine-tune the entire network (unfreezing the high layers and the regression head) until convergence.
[0067] Application effect verification: On the overall test set: To verify the practical competitiveness of our complete approach compared to traditional statistical learning methods and common deep learning methods, we further designed an external strong baseline comparison experiment. The results are shown in Table 1. Traditional methods such as PLSR, SVR, RF, and XGBoost can capture certain statistical relationships, but their overall performance is significantly weaker than the deep learning approach. This indicates that relying solely on shallow statistical relationships is insufficient to fully characterize the complex nonlinear continuous spectral structure in hyperspectral soil composition inversion. While purely supervised deep models such as 1D-CNN and Transformer outperform traditional methods, they are still limited by training instability and insufficient supervision under small sample sizes and uneven distribution conditions, resulting in overall results lower than the pre-trained transfer learning approach. In contrast, our complete approach achieves the best results in R2, RMSE, MAE, RPD, and stability metrics, indicating that its advantage does not solely stem from increased model complexity, but rather from the spectral priors provided by self-supervised pre-training, the stable task adaptation brought by staged transfer fine-tuning, and the effective supplementation of scarce conditional units by selected auxiliary samples.
[0068] Table 1. Comparison results of different models in soil composition inversion task On key subsets (such as high-concentration areas or forest background test points): The advantage of the complete scheme in this study compared with the comparative method is not evenly distributed across all samples, but is more concentrated in high-value areas, boundary areas, and background categories with insufficient samples such as woodland and grassland. The prediction accuracy is more significantly improved, indicating that its supplementary enhancement strategy for scarce units is effective.
[0069] Table 2. Specific evaluation results for scarcity intervals and scarcity background categories. Under strict data partitioning (such as by spatial blocks): the performance advantage of this method is still maintained, and its performance is more stable (with smaller standard deviation) in multiple experiments, which verifies the robustness and generalization ability of the model.
[0070] To verify that the performance advantage of our proposed scheme does not solely stem from the similarity of neighboring samples under random partitioning, we further designed generalization validation experiments under different data partitioning strategies. Table 3 shows that when switching from random partitioning to a more stringent spatial block partitioning, the performance of all methods decreases to some extent, indicating that this setting places higher demands on the model's generalization ability. However, our complete scheme maintains optimal performance under both partitioning methods, with a relatively smaller performance decline. This demonstrates that the spectral representations and task adaptation capabilities learned by our proposed scheme do not solely rely on the similarity of locally adjacent samples, but possess stronger cross-regional generalization capabilities.
[0071] Table 3 Comparison of model performance under different data partitioning strategies The sequence numbers of the above embodiments of the present invention are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.
[0072] The embodiments of the present invention have been described above with reference to the accompanying drawings. However, the present invention is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of the present invention without departing from the spirit and scope of the claims. All of these forms are within the protection scope of the present invention.
Claims
1. A method for inverting hyperspectral soil composition for heterogeneous small samples, characterized in that, Includes the following steps: S1: Using geospatial location as an index, establish a unified data interface covering observation modes from satellites, drones, and ground laboratories to obtain raw spectral data. Map the raw spectral data to a common band space and perform standardization and normalization processing to obtain a standardized dataset. S2: Construct a pre-trained model based on a two-layer masking mechanism using a standardized dataset to obtain a spectral characterization backbone network. The pre-trained model has a dedicated encoder architecture and a multi-dimensional reconstruction loss function. S3: Establish a directional sample augmentation method based on a conditional generation model, and construct a targeted and spectrally realistic auxiliary spectral library based on the directional sample augmentation method. The conditional generation model is based on the dedicated encoder architecture to encode real spectra into high-level features. S4: Based on the spectral characterization backbone network as a feature extractor and connected to the regression prediction head, a soil composition inversion model is constructed; the soil composition inversion model is trained using a three-stage progressive migration fine-tuning strategy.
2. The method for inverting hyperspectral soil composition for heterogeneous small samples according to claim 1, characterized in that, Step S3 includes: The condition generation model receives the encoding vector of the condition description unit, generates spectral features that meet the description conditions, and then decodes them into auxiliary spectra in a unified band space.
3. The method for inverting hyperspectral soil composition for heterogeneous small samples according to claim 2, characterized in that, The auxiliary spectra were screened in multiple dimensions, including: Spectral similarity screening: Calculate the spectral angular distance and cosine similarity between the generated auxiliary spectrum and the average spectrum of the real spectrum under the same conditions. If the similarity exceeds the second threshold, it is discarded.
4. The method for inverting hyperspectral soil composition for heterogeneous small samples according to claim 3, characterized in that, Multi-dimensional screening of auxiliary spectra also includes: Condition consistency screening: Use a small discriminator to verify whether the concentration range predicted by the generated auxiliary spectrum is consistent with the target conditions.
5. The method for inverting hyperspectral soil composition for heterogeneous small samples according to claim 4, characterized in that, Multi-dimensional screening of auxiliary spectra also includes: Feature space proximity filtering: In the pre-trained feature space, the features of the generated sample must be sufficiently close to the features of the nearest real sample in the corresponding conditional unit.
6. The method for inverting hyperspectral soil composition for heterogeneous small samples according to claim 5, characterized in that, Multi-dimensional screening of auxiliary spectra also includes: Diversity control screening: Limiting the number of samples retained within the same conditional unit.
7. The method for inverting hyperspectral soil composition for heterogeneous small samples according to claim 1, characterized in that, The training of the soil composition inversion model using a three-stage progressive migration fine-tuning strategy includes: Completely freeze all parameters of the pre-trained backbone network and train the regression head using only real labeled samples.
8. The method for inverting hyperspectral soil composition for heterogeneous small samples according to claim 1, characterized in that, The method of training the soil composition inversion model using a three-stage progressive migration fine-tuning strategy also includes: Unfreeze the preset number of back-end nodes in the pre-trained backbone network that are responsible for high-level feature aggregation, and continue fine-tuning them together with the regression head using only real samples.
9. The method for inverting hyperspectral soil composition for heterogeneous small samples according to claim 1, characterized in that, The method of training the soil composition inversion model using a three-stage progressive migration fine-tuning strategy also includes: Simultaneously, real samples and generated auxiliary samples are used for collaborative fine-tuning, while always maintaining the dominance of real samples; Differential supervision is implemented: For real samples, a strict regression loss is used; for auxiliary samples, while using regression loss, an interval consistency constraint is added between the soft label of the auxiliary sample and the corresponding predicted concentration interval to acknowledge the non-determinism of the soft label of the auxiliary sample. The loss weights for the auxiliary samples are gradually increased during training.
10. A hyperspectral soil composition inversion system for heterogeneous small samples, characterized in that, Applying the method as described in any one of claims 1-9; comprising: The data consistency module is used for: multi-source heterogeneous data consistency processing: using geospatial location as an index, a unified data interface covering satellite, UAV and ground laboratory observation modes is established to obtain raw spectral data, the raw spectral data is mapped to a common band space and standardized and normalized to obtain a standardized dataset; The self-supervised pre-training module is used to: construct a pre-trained model based on a two-layer masking mechanism from a standardized dataset to obtain a spectral characterization backbone network, wherein the pre-trained model has a dedicated encoder architecture and a multidimensional reconstruction loss function; An auxiliary sample generation module is used to: establish a directional sample enhancement method based on a conditional generation model, and construct a targeted and spectrally realistic auxiliary spectral library based on the directional sample enhancement method, wherein the conditional generation model is based on the dedicated encoder architecture to encode real spectra into high-level features; The migration fine-tuning training module is used to: construct a soil composition inversion model based on the spectral characterization backbone network as a feature extractor and connected to a regression prediction head; and train the soil composition inversion model using a three-stage progressive migration fine-tuning strategy.