Method for predicting soil heavy metal content by using spectral data and social environment factor data

By combining spectral preprocessing and dimensionality reduction techniques with social and environmental factor data, and employing a geographically grouped cross-validation method, the problem of unstable accuracy in soil heavy metal content prediction in cross-regional applications was solved, achieving efficient and stable cross-regional prediction and monitoring.

CN121963953APending Publication Date: 2026-05-01TIANJIN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
TIANJIN UNIV
Filing Date
2026-01-20
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing methods for predicting soil heavy metal content based on spectral data have unstable prediction accuracy when applied across regions and scenarios, and are easily affected by various factors, making it difficult to achieve large-scale, high-frequency monitoring. Furthermore, existing methods lack effective cross-regional generalization capabilities.

Method used

By combining spectral preprocessing and dimensionality reduction/embedding techniques with social environmental factor data, a predictive model is constructed. A cross-validation method based on geographical grouping is adopted to improve the model's cross-regional generalization ability and interpretability.

Benefits of technology

It effectively improves the prediction accuracy and stability of the model, enhances the application value of the model in environmental supervision, and has good cross-regional adaptability and versatility, applicable to spectral data of different bands and spatial resolutions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121963953A_ABST
    Figure CN121963953A_ABST
Patent Text Reader

Abstract

The invention discloses a method for predicting soil heavy metal content by using spectral data and social environment factor data, which comprises the following steps: S1, acquiring a training sample set, each training sample comprises position information (longitude and latitude, region codes or other attributes capable of reflecting spatial positions) and spectral data and social environment data corresponding to the sample; s2, data fusion: based on the position information or other associable sample attributes, carrying out matching fusion on the spectral data and the social environment data to obtain fusion feature data; social environment factors are fused to supplement key information of pollution sources and migration, and model prediction precision and operation stability are effectively improved. By means of a spectrum dimension reduction and embedding technology, the problem of high-dimensional data colinearity is relieved, and the over-fitting risk of the model is reduced; a cross validation method based on geographic grouping is adopted, so that the model validation process is closer to a cross-regional actual deployment scene; meanwhile, the feature contribution degree is output, and the practical application value in environment supervision work is improved.
Need to check novelty before this filing date? Find Prior Art

Description

A method for predicting soil heavy metal content using spectral data and social environmental factor data Technical Field

[0001] This invention relates to the fields of environmental monitoring, remote sensing spectral analysis and machine learning devices, specifically a method for predicting soil heavy metal content using spectral data and social environmental factor data. Background Technology

[0002] Traditional soil heavy metal detection relies on laboratory chemical analysis, which is costly and difficult to implement on a large scale and at high frequency. Hyperspectral reflectance can characterize the absorption and scattering properties of materials. However, existing methods for predicting soil heavy metal content based on spectral data still have the following problems: the response signals of heavy metals in the spectrum are usually weak and are easily affected by multiple factors such as moisture, organic matter, mineral composition, and human activities. This leads to unstable prediction accuracy and insufficient generalization ability of models built solely based on spectral features when applied across regions and scenarios. Social environmental factors (meteorology, emissions, economy, land cover, etc.) can characterize the intensity of pollution sources and migration conditions. Reasonable integration can help improve prediction accuracy and stability. At the same time, spatial correlation can lead to high validation rates for conventional random partitioning. Geographically based grouped cross-validation is needed to more realistically evaluate cross-regional generalization performance. Summary of the Invention

[0003] The purpose of this invention is to provide a method for predicting soil heavy metal content using spectral data and social environmental factor data, so as to solve the problems mentioned in the background art.

[0004] By adopting the above technical solutions, cross-regional generalization and interpretability are improved through spectral preprocessing and dimensionality reduction / embedding, social environment preprocessing and encoding, and cross-validation based on geographic grouping.

[0005] Compared with existing technologies, the beneficial effects of this invention are: by integrating social and environmental factors to supplement key information on pollution sources and migration, the accuracy of model prediction and operational stability are effectively improved; by using spectral dimensionality reduction and embedding techniques, the problem of collinearity in high-dimensional data is alleviated, and the risk of model overfitting is reduced; by adopting a cross-validation method based on geographical grouping, the model validation process is made closer to the actual deployment scenario across regions; at the same time, the output feature contribution value significantly enhances the interpretability of the model and improves its practical application value in environmental supervision.

[0006] Compared with existing technologies, the beneficial effects of this invention are: the method does not rely on a specific remote sensing platform or a single data source, and is applicable to spectral data with different band configurations and spatial resolutions; at the same time, it is not limited to a specific national or regional scale, and can be implemented under multi-regional and multi-country data conditions, exhibiting good cross-regional adaptability and versatility. Furthermore, this invention is not only applicable to the prediction of heavy metal element content, but can also be extended to the quantitative inversion or prediction of other pollutants or environmental indicators. Attached Figure Description

[0007] Figure 1 is a schematic diagram of the process of the present invention; Figure 2 is a schematic diagram of the model prediction of metal As in the present invention. Detailed Implementation

[0008] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0009] Please refer to Figure 1. This invention provides a technical solution: a method for predicting soil heavy metal content using spectral data and social environmental factor data, comprising the following steps: S1, obtaining a training sample set, each training sample including location information (latitude and longitude, regional code, or other attributes reflecting spatial location) and corresponding spectral data and social environmental data; the spectral data is a hyperspectral reflectance sequence such as visible / near-infrared / shortwave infrared; the social environmental data can be associated from measured, surveyed, or public databases through geographic coordinates or other sample attributes, and includes at least one or more of the following: meteorological factors, topographic factors, air pollution emission or concentration factors, economic factors, and land cover factors; the heavy metal elements include at least As, Cd, Co, Cr, Cu, Hg, Mn, Ni, Pb, and Sb. One or more of the following, but not limited to: S2, Data fusion: Based on the location information or other associative sample attributes, the spectral data and social environment data are matched and fused to obtain fused feature data; The fusion process involves aligning the two tables according to the associative fields to obtain the spectral vector and social environment feature vector corresponding to the same sample. When the social environment table contains a heavy metal content field with the same name or type as the target, it can be removed from the input features before modeling, retaining only the exogenous social environment features that are not the target, to ensure that label leakage is prevented; Social environment fields with a high missing ratio are removed first, and the remaining numerical features are filled with the mean / median; One-hot encoding, target encoding, or other equivalent encoding methods can be used for categorical features to achieve missing values ​​and quality control; S3, Spectral preprocessing: The spectral data is cleaned and enhanced, including smoothing, derivative, scattering correction and standardization, or other equivalent spectral preprocessing, to obtain spectral preprocessed features; The spectral dimensionality reduction or embedding module is partial least squares, principal component analysis, sparse reduction The data can be processed using any of the following methods: S4, Social Environment Feature Processing: Cleaning and standardizing the social environment data, including missing value imputation, numerical scaling, and encoding of categorical features, or other equivalent data preprocessing methods, to obtain social environment features; S5, Constructing a Predictive Model: Mapping the spectral preprocessed features to spectral latent variables of a preset dimension, mapping the social environment features to latent variables of a preset dimension, concatenating the spectral embedding and the social environment embedding, inputting the regression head, and outputting the predicted value of the target element content; The fusion regression can be any of the following: gradient boosting decision tree regression model, random forest regression model, regularized linear regression model, or support vector regression model, and other equivalent regression or prediction models can be selected according to the data scale and performance requirements; The social environment data includes at least any five of the following features: temperature, precipitation, relative humidity, altitude, carbon dioxide emissions, nitrogen oxides, sulfur dioxide, and PM2.5. Indicators include GDP, population, road-related indicators, agricultural input indicators, mineral trade indicators, and mineral resource dependence indicators. Unique-hot or target coding can be applied to country or region codes, land cover categories, etc., and these features are not limited to the above. S6. Model training: The model is trained using regression loss and outputs a predicted value for the content of at least one heavy metal element. S7. Cross-validation: The prediction model is evaluated and parameters are selected. Optionally, geographical coordinates can be used as the grouping basis to adapt to cross-regional or cross-scenario applications. S8. Prediction output: The sample to be tested undergoes processing from S2 to S4 and is input into the trained prediction model. The predicted result for the heavy metal content of the sample to be tested is output. The prediction output includes prediction summary and interpretation / derivation. The predictions of each fold validation set are summarized to generate cross-validation out-of-sample predictions (OOF), and the overall index is calculated to achieve prediction summary. A feature list, training configuration, and result file are output. Explanatory results can be exemplarily calculated using the "correlation between prediction and features" to calculate the Top-K contribution of spectral bands and social environmental features. For models that output gating / attention weights, their statistics can be summarized for interpretation. .

[0010] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0011] Please refer to Figure 2. This invention provides a technical solution: a method for predicting soil heavy metal content using spectral data and social environmental factor data. The prediction work can be implemented through a program, which is as follows: Reading and Alignment: Read two CSV tables. Automatically retrieve and sort the spectral columns by wavelength from the spectral table according to the prefix (e.g., spc.) to form a spectral matrix X_spec. Perform missing column removal and deduplication on the social environmental table and align it with the spectral table to obtain fused data; Social Environmental Feature Construction: Exclude location information fields and fields with the same name / type as heavy metal content from the social environmental table, retain only exogenous social environmental features, retain numerical features and fill missing values ​​with the mean to form a social environmental matrix X_soci; Grouping and Coordinate Extraction: If there is a country / region code field, it can be used as a grouping label for cross-validation. At the same time, extract lat / lon from the fused data to form a coordinate matrix coords; Spectral Preprocessing: SNV (sample normalization) and SG can be optionally enabled. Derivatives (configurable window length, polynomial order, and derivative order) enhance effective information absorption and suppress baseline and scattering effects; Target processing: Monotonic transformations (e.g., reciprocal transformations) can be performed on target values ​​to mitigate long-tailed distributions and restore the original dimensions at output; Cross-validation training: The model is trained using K-fold cross-validation (five-fold, ten-fold, etc.); Leave-one-out cross-validation (LOCO) can also be selected for more rigorous cross-region evaluation, where each fold uses only the training fold's fit standardizer and missing value imputation to avoid data leakage; Cross-validation training: The model is trained using K-fold cross-validation (five-fold, ten-fold, etc.) for more rigorous cross-region evaluation, where each fold uses only the training fold's fit standardizer and missing value imputation to avoid data leakage; Deep learning training details: Regression loss is used, and the optimizer can be AdamW... In conjunction with learning rate scheduling, the number of training epochs, batch size, mixed precision, and early stopping strategy can be set, and the optimal weights on the validation set can be saved; Prediction summarization: The predictions of each fold validation set are summarized to generate out-of-sample predictions for cross-validation, and the overall index is calculated; Interpretation and export: Feature list, training configuration, and result files can be output. The interpretive results can be exemplarily calculated by using the "correlation between prediction and feature" to calculate the Top-K contribution of spectral bands and social environment features. For models that output gating / attention weights, their statistics can be summarized for interpretation.

[0012] When the prediction model is a neural network model, the following is an optional programmatic implementation: Five-fold cross-validation (KFold): Example command: python code / 0_main.py --model A_fusion --target as --cv_method kfold --folds 5 --epochs 400 --batch_size 1024 --lr 3e-4 Optional parameters: --no-snv, --no-sg, --sg-window 11 --sg-poly 2 --sg-deriv 1, --precision 16-mixed. Ten-fold cross-validation: Set --folds to 10. Leave-one-out cross-validation: When the data contains grouping fields such as country / region codes, example command: python code / 0_main.py --model A_fusion --target as --cv_method loco.

[0013] Configuration file: Training hyperparameters can be uniformly managed by specifying a YAML / JSON file via --config, and command line overriding of parameters with the same name is supported.

Claims

1. A method for predicting soil heavy metal content using spectral data and social environmental factor data, characterized in that: Includes the following steps: S1. Obtain a training sample set. Each training sample includes location information (latitude and longitude, area code, or other attributes that reflect spatial location) and spectral data and social environment data corresponding to the sample. S2. Fuse data. Based on the location information or other related sample attributes, match and fuse the spectral data and social environment data to obtain fused feature data. S3. Spectral preprocessing: Cleaning and enhancing spectral data, examples include smoothing, derivative, scattering correction and normalization, or other equivalent spectral preprocessing to obtain spectral preprocessing features; S4. Social environment feature processing: Cleaning and standardizing social environment data, including missing value imputation, numerical scaling, and encoding of categorical features, or other equivalent data preprocessing methods, to obtain social environment features; S5. Constructing a prediction model: Mapping the spectral preprocessing features to spectral latent variables of a preset dimension, mapping the social environment features to latent variables of a preset dimension, concatenating the spectral embedding and the social environment embedding, inputting the regression head, and outputting the predicted value of the target element content; S6. Model training: Training the model with regression loss and outputting the predicted value of the content of at least one heavy metal element; S7. Cross-validation: Evaluating and selecting parameters for the prediction model, which can be five-fold or ten-fold random cross-validation, or grouping according to geographical location, or grouping according to a certain feature clustering, to adapt to cross-regional or cross-scenario applications; S8. Prediction output: Performing the processing of S2 to S4 on the test sample and inputting it into the trained prediction model, outputting the predicted result of the heavy metal content of the test sample.

2. The method for predicting soil heavy metal content using spectral data and social environmental factor data according to claim 1, characterized in that: In step S1, the spectral data is a hyperspectral reflectance sequence such as visible / near-infrared / shortwave infrared; the social environment data can be associated from measured, surveyed or public databases through geographic coordinates or other sample attributes, and includes at least one or more of the following: meteorological factors, topographic factors, air pollution emission or concentration factors, economic factors, and land cover factors.

3. The method for predicting soil heavy metal content using spectral data and social environmental factor data according to claim 1, characterized in that: In step S2, the fusion process involves aligning multiple tables according to their associative fields (e.g., using record identifiers for inner joins and deduplicating associated fields) to obtain the spectral vector and social environment feature vector corresponding to the same sample. When the social environment table contains a heavy metal content field with the same name or type as the target, it can be removed from the input features before modeling, retaining only the exogenous social environment features that are not the target, thus preventing label leakage. Social environment fields with excessively high missing values ​​are removed first, and the remaining numerical features are filled using methods such as mean / median. For categorical features (such as country / region codes, land cover types), one-hot coding, target coding, or other equivalent coding methods can be used to achieve missing values ​​and quality control.

4. The method for predicting soil heavy metal content using spectral data and social environmental factor data according to claim 1, characterized in that: In step S3, sample normalization is performed on the spectral vector of each sample to suppress scattering differences. Savitzky-Golay processing is used for smoothing. SG smoothing is performed on the spectral vector and the first or second derivative is calculated to enhance absorption characteristics and suppress baseline drift. SNV or MSC can be combined to suppress scattering and baseline drift, and other equivalent spectral preprocessing combinations can be used. The preprocessed spectral features are standardized to form the final spectral input matrix X_spec.

5. The method for predicting soil heavy metal content using spectral data and social environmental factor data according to claim 1, characterized in that: In step S5, the prediction model is constructed by a spectral dimensionality reduction / embedding module, a social environment feature extraction module, and a fusion regression module. The spectral dimensionality reduction / embedding module maps preprocessed spectral features to spectral latent variables of a preset dimension. The social environment feature extraction module extracts representations from social environment features and maps them to social environment latent variables or feature representations of a preset dimension. The fusion regression module fuses the spectral latent variables with the social environment latent variables or feature representations and outputs the predicted value of the target heavy metal content.

6. The method for predicting soil heavy metal content using spectral data and social environmental factor data according to claim 1, characterized in that: In step S6, the model is trained using regression loss; the learning rate, batch size and number of training rounds can be set, and an early stopping strategy can be adopted; the heavy metal content can be left unchanged, or a monotonic transformation such as the reciprocal can be used to alleviate the long-tail distribution, and the model can be restored to the original scale after training.

7. The method for predicting soil heavy metal content using spectral data and social environmental factor data according to claim 1, characterized in that: In step S7, five-fold or ten-fold cross-validation is used to evaluate the model performance, and the model structure and hyperparameters are selected accordingly. Optionally, grouped cross-validation can be introduced: for example, grouping by country / region with one group left to be validated, or constructing spatial grids / clustering groups by geographic coordinates and then performing grouped cross-validation to better reflect cross-regional deployment scenarios. Evaluation metrics may include R2, RMSE, MAE, etc.

8. The method for predicting soil heavy metal content using spectral data and social environmental factor data according to claim 1, characterized in that: In step S8, the prediction output includes prediction summarization and interpretation / derivation. Predictions from each fold validation set are summarized to generate cross-validated out-of-sample predictions (OOF), and an overall index is calculated to achieve prediction summarization. This is achieved by outputting a feature list, training configuration, and result file. Interpretive results include explanations based on feature contribution. Feature contribution can be obtained through the "correlation between prediction and feature" to obtain the Top-K contribution of spectral bands and social environment features, and / or through perturbation- or gradient-based feature attribution methods. Tree models can use contribution calculation based on Shapley values, and neural networks can use integrated gradient or spectral band occlusion methods to obtain key band contributions. For models that output gating / attention weights, the gating / attention weights can be statistically summarized for interpretation. The interpretation / derivation outputs at least one of the following: a list of key bands and their contribution, a ranking of social environment factor contributions, and a regional summary of contribution or weight.