Rapid detection method and system for heavy metals in mining area soil
By performing band filtering and dimensionality reduction on the spectral data of soil samples from mining areas, and combining a support vector machine classifier and a spatial distribution model, the complexity and noise interference problems of traditional soil heavy metal detection have been solved, enabling rapid and accurate pollution assessment and personalized remediation strategies.
Patent Information
- Application Number
- CN202511824555.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-05
- Publication Date
- 2026-03-13
AI Technical Summary
Traditional methods for detecting heavy metals in soil are complex, time-consuming, and unsuitable for large-scale monitoring. Existing spectral analysis techniques suffer from signal noise interference and difficulty in accurately identifying heavy metal speciation. Soil pollution assessment in mining areas neglects spatial distribution differences.
By acquiring spectral data from soil samples in the mining area, performing band screening and dimensionality reduction, using a support vector machine classifier to identify heavy metal morphologies, and combining this with a spatial distribution model to simulate the spatial distribution patterns of heavy metals, pollution level assessment results and remediation pathways are generated.
It enables rapid and accurate detection and pollution assessment of heavy metals in soil, providing high-precision pollution level assessment results and personalized remediation strategies, thus improving detection efficiency and accuracy.
Smart Images

Figure CN121656162A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of soil testing technology, and in particular to a rapid detection method and system for heavy metals in mining soil. Background Technology
[0002] Traditional methods for detecting heavy metals in soil, such as chemical analysis, while providing high-precision results, are complex, time-consuming, costly, and unsuitable for large-scale monitoring and real-time detection. With the development of spectral analysis technology, rapid detection methods based on spectral data have gradually become a research hotspot. Spectral analysis can quickly obtain relevant information about heavy metals in soil by interpreting the spectral response of a sample. However, existing spectral analysis techniques have some problems in application, such as signal noise interference, data redundancy, and difficulty in accurately identifying heavy metal speciation. Therefore, a new method is needed to improve the accuracy and efficiency of detection by effectively processing, reducing dimensionality, denoising, and classifying spectral data.
[0003] Furthermore, the spatial distribution of soil pollution in mining areas exhibits significant regional characteristics, and traditional pollution assessment methods often overlook these spatial variations. By introducing spatial distribution models and optimization algorithms, polluted areas can be simulated more accurately, providing a scientific basis for soil remediation. Therefore, this application provides an integrated rapid detection and pollution assessment technology for heavy metals in soil by combining spectral analysis, classification methods, and spatial distribution models. Summary of the Invention
[0004] This application provides a method and system for rapid detection of heavy metals in mining soil, which improves the efficiency and accuracy of rapid detection of heavy metals in mining soil.
[0005] Firstly, this application provides a rapid detection method for heavy metals in mining area soil, the rapid detection method for heavy metals in mining area soil comprising: Spectral data of soil samples from the mining area are obtained, the spectral data are preliminarily processed to generate an initial dataset, the initial dataset is dimensionality reduced, a simplified feature set is extracted, and the features are sorted. The simplified feature set is denoised to generate a clean signal sequence. Feature parameters related to the heavy metal speciation in the soil are extracted from the clean signal sequence, and heavy metal speciation identifiers are obtained based on a classification method. Based on the heavy metal morphology identifier, a differentiation standard is constructed and compared with the sample to be tested to obtain the morphology category of the sample to be tested. By combining the morphological categories with the spatial distribution model, the spatial distribution pattern of heavy metals in the mining area soil is simulated, the model parameters of the spatial distribution model are optimized to generate pollution level assessment results, and an identification report and regional remediation path are generated based on the pollution level assessment results.
[0006] Combining the first aspect, an initial dataset is generated, including: The spectral data is filtered by bands to obtain multiple characteristic peak bands. The reflectance value of any of the characteristic peak bands is calculated and a feature matrix is generated. The feature matrix is preprocessed to generate a first dataset. The completeness of the first dataset is checked. If the completeness meets a first preset standard, the first dataset is set as the initial dataset. Otherwise, the first dataset is interpolated and then set as the initial dataset.
[0007] In conjunction with the first aspect, the extraction of a simplified feature set and the ranking of features include: The initial dataset is standardized based on the feature matrix to generate a standard dataset. Principal component analysis is performed on the standard dataset to obtain principal component vectors. The standard dataset is then projected onto the principal component vectors to generate a low-dimensional feature set. This low-dimensional feature set is then set as the simplified feature set. The variance contribution rate of the simplified feature set is obtained based on the morphological characteristics of heavy metals, and the simplified feature set is sorted based on the magnitude of the variance contribution rate.
[0008] In conjunction with the first aspect, generating a pure signal sequence includes: The signal in the simplified feature set is decomposed into multiple scales using adaptive wavelet basis functions. The noise intensity of the decomposed signal is calculated, and the distribution characteristics of the noise intensity are evaluated using the maximum entropy criterion. The denoising threshold is dynamically adjusted based on the distribution characteristics, and the signal of the simplified feature set is subjected to secondary denoising processing based on the denoising threshold. The denoised signal is subjected to local weighted regression smoothing to generate the clean signal sequence.
[0009] In conjunction with the first aspect, the acquisition of heavy metal morphology identifiers based on classification methods includes: The peak parameters and slope parameters in the pure signal sequence are extracted and set as a morphology-specific parameter set. A support vector machine classifier is used to train and match the morphology-specific parameter set. The support vector machine classifier determines the unique identifier of any heavy metal morphology, obtains the classification result of the unique identifier, and if the accuracy of the classification result meets the second preset standard, the unique identifier is retained and set as the heavy metal morphology identifier.
[0010] In conjunction with the first aspect, obtaining the morphological category of the sample to be detected includes: Multidimensional feature analysis was performed on the morphological identifiers of heavy metals to construct multidimensional differentiation criteria and generate morphological classification templates. The spectral input data of the sample to be detected is obtained, the spectral input data is preprocessed and the feature parameters to be analyzed are extracted, the feature parameters to be analyzed are compared with the morphological classification template, and the matching degree is output. If the matching degree is higher than the preset threshold, the morphology classification template outputs the morphology category of the sample to be detected and records the comparison result of the morphology category.
[0011] In conjunction with the first aspect, the generated pollution level assessment results include: Based on the heavy metal morphology identifier, obtain the morphological category data of various heavy metals in the classified mining area soil samples, and generate category sample data; The sample data of the aforementioned categories are input into the spatial distribution model, and the spatial interpolation method is used to simulate the spatial distribution pattern of different heavy metal forms in the soil of the mining area, and the spatial distribution feature parameters are extracted. An iterative optimization algorithm is used to adjust the spatial distribution feature parameters to obtain the fitting accuracy of the spatial distribution model, and the model parameters are optimized based on the fitting accuracy. The optimized spatial distribution model calculates the pollution level assessment results of the morphological category based on the pollution load index.
[0012] In conjunction with the first aspect, the generation of the identification report and regional remediation pathway based on the pollution level assessment results includes: Based on the risk coefficient and the pollution level assessment results, high-risk area data is extracted, and the high-risk area data is summarized based on geographic information to generate the identification report; Based on the heavy metal morphology distribution information in the identification report, a remediation strategy mapping model is constructed. The remediation strategy mapping model matches data from any high-risk area with a remediation method to generate a remediation path. A feasibility analysis is performed on the repair path. If the feasibility meets the third preset standard, the repair path is set as the regional repair path.
[0013] In conjunction with the first aspect, the construction of the repair strategy mapping model includes: Based on the multidimensional feature analysis method, the morphological distribution information of heavy metals is associated with influencing factors to create a multidimensional remediation strategy mapping model. The multidimensional repair strategy mapping model is trained based on a machine learning algorithm, and the trained multidimensional repair strategy mapping model is set as the repair strategy mapping model.
[0014] Secondly, this application provides a rapid detection system for heavy metals in mining area soil, the rapid detection system for heavy metals in mining area soil comprising: The data processing module is used to acquire spectral data of soil samples from the mining area, perform preliminary processing on the spectral data to generate an initial dataset, perform dimensionality reduction processing on the initial dataset, extract a simplified feature set, and sort the features. The classification module is used to denoise the simplified feature set, generate a clean signal sequence, extract feature parameters related to the heavy metal speciation in the soil from the clean signal sequence, and obtain heavy metal speciation identifiers based on the classification method. The detection module is used to construct a differentiation standard based on the heavy metal morphology identifier and compare it with the sample to be detected to obtain the morphology category of the sample to be detected. The assessment module is used to combine the morphological categories with the spatial distribution model to simulate the spatial distribution pattern of heavy metals in the mining area soil, optimize the model parameters of the spatial distribution model to generate pollution level assessment results, and generate identification reports and regional remediation paths based on the pollution level assessment results.
[0015] The technical solution provided in this application enables rapid identification of heavy metal morphologies in soil through the acquisition and processing of soil sample spectral data. By performing band filtering, feature extraction, and a support vector machine (SVM)-based classification method on the spectral data, the morphology of heavy metals can be accurately determined. By combining heavy metal morphology categories with a spatial distribution model, the spatial distribution patterns of heavy metals in mining area soils can be simulated. Through spatial interpolation methods and iterative optimization algorithms, the model accurately reflects the spatial differences in soil pollution and provides high-precision pollution level assessment results. Multidimensional feature analysis, principal component analysis (PCA), and wavelet transform techniques are used to reduce the dimensionality and denoise the data, resulting in purer signals and effectively improving data quality and the accuracy of subsequent analysis. During denoising, the combination of the maximum entropy criterion and wavelet transform adaptively adjusts the denoising threshold to retain the most effective signal. Based on the pollution assessment results, the system can automatically generate a pollution identification report and construct a remediation strategy mapping model based on heavy metal morphology distribution information. It selects the most suitable remediation method for each polluted area, choosing the optimal remediation path according to different pollution conditions in the area and conducting feasibility analysis to ensure the efficiency and operability of the remediation plan. Attached Figure Description
[0016] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0017] Figure 1This is a schematic diagram of one embodiment of a rapid detection method for heavy metals in mining soil according to the present application. Figure 2 This is a schematic diagram of one embodiment of the heavy metal morphology recognition process in this application. Figure 3 This is a schematic diagram of one embodiment of the pollution assessment and remediation process in this application. Figure 4 This is a schematic diagram of one embodiment of a rapid detection system for heavy metals in mining soil according to the present application. Detailed Implementation
[0018] This application provides a method and system for rapid detection of heavy metals in mining soil. The terms "first," "second," "third," "fourth," etc. (if present) in the specification, claims, and accompanying drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments described herein can be implemented in a sequence other than that illustrated or described herein. Furthermore, the terms "comprising" or "having" and any variations thereof are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0019] For ease of understanding, the specific process of the embodiments of this application is described below. Please refer to [link / reference]. Figure 1 One embodiment of a rapid detection method for heavy metals in mining area soil, as described in this application, includes: Step S101: Obtain spectral data of soil samples from the mining area, perform preliminary processing on the spectral data to generate an initial dataset, perform dimensionality reduction processing on the initial dataset, extract a simplified feature set, and sort the features.
[0020] It is understood that the executing entity of this application can be a rapid detection device for heavy metals in mining soil, or it can be a terminal or a server; the specific implementation is not limited here. This application's embodiment uses a server as an example for illustration.
[0021] Specifically, soil samples were collected from the mining area, and raw spectral data of the soil samples were obtained using spectral scanning equipment (such as Fourier transform infrared spectrometer (FTIR) or near-infrared spectrometer (NIR). These spectral data recorded the reflectance of different wavelengths in the soil, providing signals related to the content and speciation of heavy metals in the soil. Different heavy metal speciations in the soil exhibit significant differences in reflectance characteristics across different wavelengths.
[0022] After acquiring the raw spectral data, preliminary processing is performed, including instrument noise removal, background removal, and signal normalization. Baseline correction techniques (such as polynomial fitting methods) are used to remove background noise, ensuring data accuracy and consistency. Through standardization, all data are converted to a uniform dimension, allowing reflectance data from different bands to be analyzed on the same scale, thus eliminating the influence of dimensional differences between different bands.
[0023] The pre-processed spectral data underwent band filtering to select characteristic bands closely related to the speciation of heavy metals. For example, for characteristic bands of heavy metals such as lead and cadmium (e.g., 600-800 nm and 1800-2000 nm), data in these bands were filtered and extracted to generate a feature matrix, which was then integrated into an initial dataset. This dataset includes key signals reflecting the speciation of heavy metals in the soil.
[0024] Principal Component Analysis (PCA) and other dimensionality reduction algorithms are used to reduce the dimensionality of the initial dataset, extracting the most effective features for heavy metal morphology discrimination and reducing redundant information. Dimensionality reduction can compress high-dimensional data into a low-dimensional feature set while retaining the most important information in the original data, thus providing a simplified dataset for subsequent feature ranking and analysis.
[0025] Step S102: Denoise the simplified feature set to generate a clean signal sequence, extract feature parameters related to the heavy metal morphology in the soil from the clean signal sequence, and obtain heavy metal morphology identifiers based on the classification method.
[0026] Specifically, the simplified feature set after dimensionality reduction is denoised to remove irrelevant signals introduced by environmental noise or instrument errors. Wavelet transform is used to decompose the signal at multiple scales through adaptive wavelet basis functions. This process effectively separates signals from noise at different scales and removes low-frequency noise. The selection of wavelet transform is optimized based on the spectral characteristics of the soil samples to ensure effective noise filtering.
[0027] The denoised signal is evaluated using the maximum entropy criterion to assess the noise intensity distribution characteristics. Based on these characteristics, the denoising threshold is dynamically adjusted to further improve the denoising effect. If the noise intensity is high, a second denoising process is performed to ensure signal purity. Finally, local weighted regression smoothing is used to remove any residual noise and generate a clean signal sequence.
[0028] Feature parameters related to heavy metal morphology are extracted from the clean signal sequence, including peak values, slopes, and spectral features. These feature parameters reflect the different spectral representations of various heavy metal morphologies, serving as the basis for classification. A Support Vector Machine (SVM) classifier is used to train and match these features. SVM achieves efficient differentiation of different heavy metal morphologies by maximizing the boundaries between categories. If the classification results meet the preset accuracy criteria, the validity of the heavy metal morphology label is confirmed, and a specific heavy metal morphology category label is assigned.
[0029] Step S103: Construct a differentiation standard based on heavy metal morphology identifiers and compare it with the sample to be tested to obtain the morphology category of the sample to be tested.
[0030] Specifically, heavy metal speciation markers were identified, and corresponding differentiation criteria were constructed based on these markers. This standard defines the distribution range of different heavy metal speciations in the feature space and provides a basis for subsequent comparisons of the samples to be tested. For the soil samples to be tested, their spectral data were processed using the same preprocessing procedures (band selection, data standardization, etc.) to generate a feature set for the samples to be tested.
[0031] The similarity is calculated by comparing the features of the sample to be tested with established classification templates. The comparison process uses algorithms such as Euclidean distance or cosine similarity to quantitatively measure the degree of similarity between the sample and each morphological classification template. If the matching degree between the sample and a certain morphological category template is higher than a preset threshold, the sample is confirmed to belong to the corresponding heavy metal morphological category. This step ensures that the heavy metal morphology of each sample can be accurately identified.
[0032] Step S104: Combine morphological categories with spatial distribution models to simulate the spatial distribution patterns of heavy metals in mining area soil, optimize the model parameters of the spatial distribution model to generate pollution level assessment results, and generate identification reports and regional remediation paths based on the pollution level assessment results.
[0033] Specifically, after obtaining the morphological categories of the samples to be tested, a spatial distribution model is used to simulate the spatial distribution of heavy metals in the soil. By using spatial interpolation techniques (such as Kriging interpolation), the morphological category data is combined with the spatial location of the soil samples to simulate the distribution patterns of different heavy metal forms in mining area soils. Through iterative optimization of the model parameters, the model's fitting accuracy to actual soil samples is improved. The optimization process uses an iterative optimization algorithm to adjust the parameters of the distribution model, improving its accuracy in predicting the spatial distribution of heavy metals in soil.
[0034] The optimized model generates a pollution level assessment, reflecting the pollution levels in different areas of the mining area soil. This assessment allows for the development of appropriate remediation pathways for mining area pollution control. Based on the pollution level assessment, a detailed identification report is generated. The report includes information such as the pollution level of each area, the distribution of different heavy metal speciations, and their environmental impact. The remediation pathway is developed based on factors including pollution level, heavy metal speciation, soil characteristics, and resource availability. The remediation pathway is automatically generated using a remediation strategy mapping model, ensuring that the most suitable remediation method is selected for each high-risk area, taking into account factors such as the feasibility, cost, and resource consumption of the remediation method, ultimately generating a remediation implementation plan.
[0035] In one specific embodiment, generating an initial dataset includes: (1) Perform band filtering on the spectral data to obtain multiple characteristic peak bands, calculate the reflectance value of any characteristic peak band and generate a feature matrix.
[0036] (2) Preprocess the feature matrix to generate the first dataset, check the completeness of the first dataset. If the completeness meets the first preset standard, set the first dataset as the initial dataset. Otherwise, set the first dataset as the initial dataset after interpolation.
[0037] Specifically, the raw spectral data undergoes band screening. The goal of this process is to extract bands associated with different heavy metal speciations in the soil. The screening criteria are based on previously known characteristics of heavy metal speciations in the spectrum (e.g., lead's characteristic absorption band is 600-800 nm, and cadmium may appear in the 1800-2000 nm band), identifying bands containing significant absorption peaks. These bands provide potential characteristic data that can distinguish different heavy metal speciations. Using the screened characteristic band data, the reflectance value for each band is calculated. Reflectance values are typically calculated by comparing the reflected light intensity of the soil sample in each band with the reflected light intensity of a standard reference light source. These reflectance values reflect the spectral response of heavy metals in the soil and form the basis for subsequent data processing. The reflectance values of each band are then organized into a feature matrix, where each column represents the reflectance data for a different band, and each row represents the spectral characteristics of a different sample.
[0038] After the feature matrix is generated, preprocessing is performed. The goal of preprocessing is to remove irrelevant interference signals and noise, and to ensure data quality and consistency. For example, signal processing methods such as baseline correction and filtering are used to remove noise generated by spectral scanning equipment, environmental conditions, or the acquisition process. The reflectance values for each band are standardized to ensure that all data are within the same dimension range. For example, the mean value of each band's reflectance value is subtracted and divided by the standard deviation, so that the reflectance data for each band has zero mean and unit variance, eliminating the influence of differences in the dimensions of data from different bands.
[0039] Check each row (i.e., each sample) of the feature matrix for missing reflectance values. If missing data exists, determine whether the absence affects subsequent analysis. If missing reflectance values exist in the feature matrix (e.g., spectral data for certain bands cannot be acquired), interpolation methods (such as linear interpolation, sample mean imputation, etc.) are used to fill in these missing values. The interpolation method extrapolates the missing data points based on known data points, ensuring the data integrity of the feature matrix. If the integrity meets the first preset standard, this dataset is directly used as the initial dataset for subsequent dimensionality reduction and classification processing.
[0040] In one specific embodiment, extracting a simplified feature set and ranking the features includes: (1) Standardize the initial dataset based on the feature matrix to generate a standard dataset.
[0041] (2) Perform principal component analysis on the standard dataset to obtain the principal component vectors. Project the standard dataset onto the principal component vectors to generate a low-dimensional feature set. Set the low-dimensional feature set as a simplified feature set.
[0042] (3) Obtain the variance contribution rate of the simplified feature set based on the morphological characteristics of heavy metals, and sort the simplified feature set based on the magnitude of the variance contribution rate.
[0043] Specifically, for each band in the initial dataset (i.e., each column of the feature matrix), the mean reflectance value of that band is calculated. and standard deviation Standardization is performed based on the first formula, which is: Where x is the original reflectance value, It is the mean value for that band. x1 is the standard deviation of this band, and x1 is the standardized data.
[0044] Principal Component Analysis (PCA) is a commonly used data dimensionality reduction method that extracts the most representative features from high-dimensional data, reduces redundant information, and retains most of the variance information in the data. It involves calculating the covariance matrix of a standard dataset, performing eigenvalue decomposition on the covariance matrix to obtain eigenvalues and eigenvectors, and selecting the top k principal components as principal component vectors based on the magnitude of the eigenvalues.
[0045] The contribution of each principal component to data variability is assessed by calculating its variance contribution rate. The variance contribution rate is calculated by dividing each eigenvalue by the sum of all eigenvalues. For example, the variance contribution rate p is calculated using the second formula: ,in, Let be the eigenvalue of the i-th principal component vector, and k be the number of principal component vectors selected. Based on the magnitude of the variance contribution rate, the features in the simplified feature set are sorted, retaining those features with larger contributions. Redundant information in the original dataset is removed, while retaining the features most useful for heavy metal morphology discrimination.
[0046] In one specific embodiment, generating a clean signal sequence includes: (1) Use adaptive wavelet basis functions to perform multi-scale decomposition on the signal in the simplified feature set, calculate the noise intensity of the decomposed signal, and use the maximum entropy criterion to evaluate the distribution characteristics of the noise intensity.
[0047] (2) The denoising threshold is dynamically adjusted based on the distribution characteristics, and the signal of the simplified feature set is subjected to secondary denoising processing based on the denoising threshold.
[0048] (3) Perform local weighted regression smoothing on the denoised signal to generate a clean signal sequence.
[0049] Specifically, soil samples from different mining areas may exhibit varying characteristics at different scales. Therefore, selecting a highly adaptable wavelet basis function can more accurately extract signal features. Through wavelet transform, the signal is decomposed into different scales (frequency bands), each corresponding to different signal information and noise components. The decomposition result contains multiple sub-signals (low-frequency and high-frequency signals). The maximum entropy criterion evaluates the distribution relationship between signal and noise by analyzing the entropy value of the signal after wavelet decomposition. High entropy values generally indicate more complex signal information, while low entropy values indicate stronger noise components.
[0050] The denoising threshold is dynamically determined based on the noise intensity and entropy value. If the noise intensity is high, the threshold is set lower to ensure that more noise is filtered out; if the noise is low, the threshold is set higher to avoid accidentally deleting signals. After determining the denoising threshold, the decomposed signal undergoes secondary denoising processing. Wavelet reconstruction is then used to resynthesize the denoised signal, resulting in a noise-removed signal sequence.
[0051] Even after secondary denoising, the signal may still exhibit minor fluctuations and irregularities, necessitating further smoothing to ensure its smoothness and continuity. Local weighted regression smoothing is a common signal optimization technique that removes noise while preserving signal characteristics. For example, weighted regression methods calculate a weighted average of the signal based on the distance and weights of data points within the neighborhood, ensuring the smoothed signal retains good local characteristics. During weighted regression, the weights of data points are adjusted according to their distance from the current point; closer data points have higher weights, while farther data points have lower weights, effectively removing high-frequency noise and preserving low-frequency signal characteristics. Through local weighted regression smoothing, a final clean signal sequence is generated, retaining the main characteristics of heavy metals in the soil.
[0052] In one specific embodiment, obtaining heavy metal speciation identifiers based on a classification method includes: (1) Extract the peak parameters and slope parameters from the pure signal sequence and set them as a morphology-specific parameter set. Use a support vector machine classifier to train and match the morphology-specific parameter set.
[0053] (2) The support vector machine classifier determines the unique identifier of any heavy metal morphology and obtains the classification result of the unique identifier. If the accuracy of the classification result meets the second preset standard, the unique identifier is retained and set as the heavy metal morphology identifier.
[0054] Specifically, peak detection is performed on a clean signal sequence to identify significant fluctuations in the signal. Peaks represent characteristic signals of heavy metal morphologies (such as absorption or reflection peaks). For example, different heavy metals like lead, cadmium, and copper exhibit different peak characteristics in their spectra. For each peak, parameters such as its position (wavelength or wavenumber), height (reflectance value), and width (peak width) are extracted. The slope parameter is an important feature reflecting the rate of signal change. By calculating the slope of the signal in different bands (i.e., the signal gradient), the trend of signal change can be obtained. For example, some heavy metal morphologies may exhibit rapid increases or decreases in their spectra; this change is quantified by the slope parameter. The extracted peak parameters and slope parameters together constitute a set of morphology-specific parameters.
[0055] A training dataset was created based on known soil sample data from mining areas. The input data for each sample consisted of a set of morphology-specific parameters (including peak and slope parameters), and the output label was the corresponding heavy metal morphology (e.g., lead, cadmium, copper). This training dataset was used to train a Support Vector Machine (SVM) classifier, enabling it to learn the differences between different heavy metal morphologies. The peak and slope parameters for each sample were integrated into a feature vector, which served as the input to the SVM. During training, labeled data of known heavy metal morphologies, such as oxidized lead samples, were used to fit the parameter set, and the kernel function, such as the radial basis function, was adjusted to handle nonlinear morphological differences. This training effectively handles high-dimensional parameter sets, improves generalization ability in complex soil environments, and enhances classification accuracy. For parameter sets of unknown samples, the classifier calculated the distance to the training hyperplane; if the distance was positive, it was matched to a specific morphology, such as a commutative state. This matching process is fast and accurate, beneficial for real-time soil detection.
[0056] If the matching degree between the sample to be tested and a certain heavy metal speciation category is higher than a preset threshold, the SVM classifier will output that the sample belongs to that category. For example, if the matching degree of the sample to be tested is highest with the "lead" speciation category, then the sample will be classified as "lead". The classification result is an identifier for the heavy metal speciation, indicating the heavy metal speciation in the soil sample.
[0057] The classification results are compared with the known labels of the actual samples to evaluate the accuracy of the SVM classifier. If the accuracy of the classification results meets the second preset standard (e.g., the classification accuracy reaches 95% or higher), the classification results are confirmed as valid, and the classification results are retained as valid heavy metal morphology identifiers to generate heavy metal morphology identifiers. If the accuracy of the classification results is lower than the preset standard, it can be optimized through further model tuning or supplementing training data.
[0058] In one specific embodiment, obtaining the morphological category of the sample to be detected includes: (1) Perform multidimensional feature analysis on heavy metal morphology identifiers, construct multidimensional differentiation criteria, and generate morphology classification templates.
[0059] (2) Obtain the spectral input data of the sample to be tested, extract the feature parameters to be analyzed after preprocessing the spectral input data, compare the feature parameters to be analyzed with the morphological classification template, and output the matching degree.
[0060] (3) If the matching degree is higher than the preset threshold, the morphology classification template outputs the morphology category of the sample to be detected and records the comparison result of the morphology category.
[0061] Specifically, Figure 2This is a flowchart for heavy metal speciation identification. Based on the acquired heavy metal speciation identifiers, multidimensional feature analysis is performed on the sample characteristics. According to the spectral characteristics of different heavy metal speciations, their performance in multiple feature dimensions (such as reflectance, absorption peaks, slope, etc.) is analyzed. These multidimensional features (reflectance, absorption peaks, slope, etc.) are constructed into a multidimensional feature space. For example, the features of each heavy metal speciation are used as dimensions of this feature space, and the position of each soil sample in this space represents the characteristic distribution of its heavy metal speciation. In this way, different heavy metal speciations will form different distribution patterns in the feature space. A multidimensional discrimination criterion is set based on the distribution of each heavy metal speciation in the multidimensional feature space to determine whether a sample belongs to a specific heavy metal speciation. For example, a separating hyperplane or decision boundary can be created based on algorithms such as cluster analysis or support vector machines (SVM) to separate the data of different heavy metal speciations, defining the "boundary" of each heavy metal speciation in the feature space.
[0062] The spectral input data of the samples to be tested undergoes preprocessing to ensure its quality and consistency. Preprocessing includes noise removal, baseline correction, and data standardization. From the spectral data of the samples, analytical feature parameters related to heavy metal speciation are extracted, such as peak value, slope, and absorption bands. These analytical feature parameters are compared with a pre-constructed morphology classification template. The template contains a feature dataset of known heavy metal speciations, representing the position of each heavy metal speciation in a multi-dimensional feature space. During the comparison, the similarity between the sample and the template is calculated using algorithms such as Euclidean distance or cosine similarity to evaluate the matching degree between the sample and each heavy metal speciation template. If the matching degree is higher than a preset threshold (e.g., 0.85), the sample is confirmed to belong to the heavy metal speciation category corresponding to that template, which is the comparison result. For example, if the sample has the highest matching degree with the "lead" template and exceeds the threshold, the sample is determined to belong to the lead speciation.
[0063] In one specific embodiment, generating a pollution level assessment result includes: (1) Based on the heavy metal morphology identification, obtain the morphological category data of various heavy metals in the classified mining area soil samples and generate category sample data.
[0064] (2) Input the category sample data into the spatial distribution model, use the spatial interpolation method to simulate the spatial distribution pattern of different heavy metal forms in the mining area soil, and extract the spatial distribution characteristic parameters.
[0065] (3) Use iterative optimization algorithm to adjust the spatial distribution feature parameters, obtain the fitting accuracy of the spatial distribution model, and optimize the model parameters based on the fitting accuracy.
[0066] (4) The optimized spatial distribution model calculates the pollution level assessment results of the morphological categories based on the pollution load index.
[0067] Specifically, each soil sample from the mining area is assigned a label in the heavy metal speciation classification, such as lead, cadmium, and copper. Each soil sample includes not only its spectral data and speciation identifier, but also its spatial location information (such as GPS coordinates and area code). By combining the spatial location of the sample with the heavy metal speciation category, category sample data is generated.
[0068] Kriging spatial interpolation was employed to infer the pollution status of unsampled areas based on the spatial location and heavy metal speciation of known sample points. The spatial distribution of different heavy metal speciations in the mining area soil was simulated. For example, lead and cadmium concentrations might be higher in some areas, while copper concentrations might be lower. The interpolation algorithm estimated the heavy metal concentrations and distributions in other areas of the mining area based on existing data, generating a spatial distribution map of different heavy metal speciations within the mining area. Spatial distribution characteristic parameters included the spatial gradient of heavy metal concentrations (reflecting the rate of change of heavy metals in space), pollution hotspots (identified through density analysis or cluster analysis of the most severely polluted areas), and spatial heterogeneity (reflecting the differences in soil pollution across different areas).
[0069] In iterative optimization algorithms, feature data extracted from spatial distribution models (such as spatial distribution gradients and pollution hotspots) are used to construct a training dataset. Each sample includes known soil heavy metal speciation data (as input) and corresponding pollution load and spatial distribution feature parameters (as output labels). A deep neural network (DNN) or convolutional neural network (CNN) model is constructed. The input of this model is the spatial distribution feature data of the soil sample, and the output is the predicted spatial distribution parameters (such as pollution load and heavy metal density). The network structure typically contains multiple hidden layers, each performing a nonlinear transformation on the features, ultimately outputting optimized spatial distribution parameters. During neural network training, backpropagation is used for parameter tuning, with mean squared error (MSE) or cross-entropy loss function as the optimization objective. During training, the network continuously adjusts its weights to minimize the error between the output spatial distribution parameters and the actual pollution data. Cross-validation is used to ensure that the deep neural network exhibits good generalization ability on different datasets through multiple training and validation cycles. Deep learning and reinforcement learning algorithms are combined, and the decision-making process is optimized using deep neural networks within the framework of reinforcement learning. Through joint training, the system can automatically adjust the parameters of the spatial distribution model during the optimization process, and learn how to select the best repair path.
[0070] In each optimization iteration, key parameters of the spatial distribution model are adjusted (e.g., weights in the spatial interpolation method, distribution smoothness, spatial correlation of the data, etc.). By comparing the fitting accuracy before and after optimization, it can be determined whether the model has reached its optimal state. The optimization algorithm evaluates and adjusts parameters by minimizing errors (e.g., mean squared error, MSE) or maximizing the goodness of fit. By calculating the model's fitting accuracy after each iteration, the model accuracy is improved, and the optimal model parameters for the spatial distribution model are obtained.
[0071] Using the pollution load index formula, the heavy metal concentration in each region is compared with standard values to comprehensively calculate the degree of soil pollution. Based on the pollution load index, the pollution level of different mining areas is assessed. Areas with higher pollution levels are marked as high-risk areas, and these areas require priority for pollution remediation; this is the pollution level assessment result.
[0072] In one specific embodiment, generating an identification report and regional remediation pathway based on pollution level assessment results includes: (1) Based on the risk coefficient and pollution level assessment results, extract high-risk area data, and generate an identification report based on the geographic information to summarize the high-risk area data.
[0073] (2) Based on the heavy metal morphology distribution information in the identification report, a remediation strategy mapping model is constructed. The remediation strategy mapping model matches the data of any high-risk area with the remediation method to generate a remediation path.
[0074] (3) Conduct a feasibility analysis on the repair path. If the feasibility meets the third preset standard, then set the repair path as a regional repair path.
[0075] Specifically, Figure 3 This is a flowchart for pollution assessment and remediation. The risk coefficient is calculated by comparing the concentration of various heavy metals in the soil with environmental standards. The risk coefficient for each area is calculated based on the pollution load index, which comprehensively considers factors such as heavy metal concentration, toxicity, and environmental persistence. Areas with high risk coefficients indicate more severe heavy metal pollution. By calculating the risk coefficient for each area, high-risk areas are screened, typically those with a pollution load index exceeding a certain threshold (e.g., 0.5 or 1.0). Using Geographic Information System (GIS) technology, data from each high-risk area is aggregated, including pollution level, heavy metal types, and spatial distribution, to generate an identification report. This report details the pollution level, heavy metal speciation (e.g., lead, cadmium, copper), spatial distribution patterns (e.g., pollution gradient, pollution range), and corresponding environmental impacts for each high-risk area.
[0076] A remediation strategy mapping model is constructed based on the distribution of heavy metal speciation, pollution levels, and soil characteristics (such as soil pH and organic matter content) in each high-risk area. This model selects the most suitable remediation measure for each polluted area by associating the pollution characteristics of each area with possible remediation methods. For example, if lead pollution in a region is mainly concentrated in the exchangeable state, chemical stabilization remediation may be suitable; if cadmium is mainly in the residual state, phytoremediation or soil sealing methods can be considered.
[0077] A detailed feasibility assessment is conducted for each remediation path. The assessment includes the implementation difficulty of the remediation method, the required time, the required resources (such as remediation agents, equipment, and labor), the economic costs, and the expected results after remediation. Using economic or resource optimization models, a cost-benefit analysis is performed on each remediation scheme to ensure that the selected path can achieve the expected results within an acceptable timeframe and budget. The feasibility of the remediation path is verified against a third pre-set standard (such as budget constraints, time requirements, and resource availability). If the feasibility analysis results meet the pre-set standards, the applicability of the remediation path is confirmed.
[0078] Based on the feasibility analysis results, the optimal remediation path, known as the regional remediation path, is selected. A detailed implementation plan is then generated based on this confirmed regional remediation path. The implementation plan includes remediation steps, required resources, timeline, budget, personnel arrangements, and risk control measures.
[0079] In one specific embodiment, constructing a remediation strategy mapping model includes: (1) Based on the multidimensional feature analysis method, the distribution information of heavy metal morphology is associated with influencing factors, and a multidimensional remediation strategy mapping model is created.
[0080] (2) The multidimensional repair strategy mapping model is trained based on machine learning algorithms, and the trained multidimensional repair strategy mapping model is set as the repair strategy mapping model.
[0081] Specifically, the distribution information of heavy metal speciation (e.g., exchangeable, residual, dissolved, etc.) in different soil samples is combined with soil physicochemical properties (e.g., pH, organic matter content, soil acidity / alkalinity) and external environmental conditions (e.g., precipitation, temperature, etc.) for analysis. Through multidimensional feature analysis, the relationship between heavy metal speciation and soil characteristics and environmental conditions can be revealed, forming a multidimensional dataset and creating a multidimensional remediation strategy mapping model.
[0082] Multidimensional remediation strategy mapping models are trained using machine learning algorithms such as decision trees, support vector machines (SVM), or random forests, based on a large amount of known soil sample data from mining areas, to learn the correspondence between various heavy metal speciations and remediation methods. For example, the model can learn how to automatically recommend appropriate remediation methods (such as chemical stabilization, phytoremediation, or physical containment) based on factors such as soil pH, heavy metal speciation, and pollution levels. The trained model can automatically output the optimal remediation strategy based on the input heavy metal speciation distribution and soil characteristic data.
[0083] The trained multidimensional remediation strategy mapping model will automatically match and generate targeted remediation paths based on the actual characteristics of the soil in the mining area to be treated. Each remediation path includes the selected remediation method, required resources, implementation steps, and time schedule. This model not only improves the efficiency and accuracy of remediation strategy selection but also optimizes the applicability of remediation strategies according to different soil conditions in mining areas, ensuring optimal remediation results. Therefore, the trained multidimensional remediation strategy mapping model is designated as the remediation strategy mapping model.
[0084] The above describes a rapid detection method for heavy metals in mining soil according to embodiments of this application. The following describes a rapid detection system for heavy metals in mining soil according to embodiments of this application. Please refer to [link / reference]. Figure 4 One embodiment of the rapid detection system for heavy metals in mining area soil according to this application includes: The data processing module 201 is used to acquire spectral data of soil samples from the mining area, perform preliminary processing on the spectral data to generate an initial dataset, perform dimensionality reduction processing on the initial dataset, extract a simplified feature set, and sort the features.
[0085] The classification module 202 is used to denoise the simplified feature set, generate a clean signal sequence, extract feature parameters related to the heavy metal morphology in the soil from the clean signal sequence, and obtain heavy metal morphology identifiers based on the classification method.
[0086] The detection module 203 is used to construct a differentiation standard based on the heavy metal morphology identifier and compare it with the sample to be tested to obtain the morphology category of the sample to be tested.
[0087] The assessment module 204 is used to combine morphological categories with spatial distribution models to simulate the spatial distribution patterns of heavy metals in mining area soil, optimize the model parameters of the spatial distribution model to generate pollution level assessment results, and generate identification reports and regional remediation paths based on the pollution level assessment results.
[0088] Through the collaborative efforts of the aforementioned components, the technical solution provided in this application enables the rapid identification of heavy metal forms in soil by acquiring and processing spectral data from soil samples. By performing band filtering, feature extraction, and a classification method based on Support Vector Machine (SVM) on the spectral data, the forms of heavy metals can be accurately determined. By combining heavy metal form categories with a spatial distribution model, the spatial distribution patterns of heavy metals in mining area soils can be simulated. Through spatial interpolation methods and iterative optimization algorithms, the model can accurately reflect the spatial differences in soil pollution and provide high-precision pollution level assessment results. Multidimensional feature analysis, principal component analysis (PCA), and wavelet transform techniques are used to reduce the dimensionality and denoise the data, resulting in purer signals and effectively improving data quality and the accuracy of subsequent analysis. During the denoising process, the combination of the maximum entropy criterion and the wavelet transform method adaptively adjusts the denoising threshold to retain the effective signal to the maximum extent. Based on pollution assessment results, the system can automatically generate pollution identification reports and construct a remediation strategy mapping model by combining heavy metal speciation distribution information. It can select the most suitable remediation method for each polluted area, choose the optimal remediation path according to different pollution conditions of the area, and conduct feasibility analysis to ensure the efficiency and operability of the remediation plan.
[0089] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0090] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0091] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.
Claims
1. A rapid detection method for heavy metals in mining area soil, characterized in that, The method for rapid detection of heavy metals in mining area soil includes: Spectral data of soil samples from the mining area are obtained, the spectral data are preliminarily processed to generate an initial dataset, the initial dataset is dimensionality reduced, a simplified feature set is extracted, and the features are sorted. The simplified feature set is denoised to generate a clean signal sequence. Feature parameters related to the heavy metal speciation in the soil are extracted from the clean signal sequence, and heavy metal speciation identifiers are obtained based on a classification method. Based on the heavy metal morphology identifier, a differentiation standard is constructed and compared with the sample to be tested to obtain the morphology category of the sample to be tested. By combining the morphological categories with the spatial distribution model, the spatial distribution pattern of heavy metals in the mining area soil is simulated, the model parameters of the spatial distribution model are optimized to generate pollution level assessment results, and an identification report and regional remediation path are generated based on the pollution level assessment results.
2. The rapid detection method for heavy metals in mining area soil according to claim 1, characterized in that, The generation of the initial dataset includes: The spectral data is filtered by bands to obtain multiple characteristic peak bands. The reflectance value of any of the characteristic peak bands is calculated and a feature matrix is generated. The feature matrix is preprocessed to generate a first dataset. The completeness of the first dataset is checked. If the completeness meets a first preset standard, the first dataset is set as the initial dataset. Otherwise, the first dataset is interpolated and then set as the initial dataset.
3. The rapid detection method for heavy metals in mining area soil according to claim 2, characterized in that, The extraction of the simplified feature set and the ranking of the features include: The initial dataset is standardized based on the feature matrix to generate a standard dataset. Principal component analysis is performed on the standard dataset to obtain principal component vectors. The standard dataset is then projected onto the principal component vectors to generate a low-dimensional feature set. This low-dimensional feature set is then set as the simplified feature set. The variance contribution rate of the simplified feature set is obtained based on the morphological characteristics of heavy metals, and the simplified feature set is sorted based on the magnitude of the variance contribution rate.
4. The rapid detection method for heavy metals in mining area soil according to claim 1, characterized in that, The generation of the pure signal sequence includes: The signal in the simplified feature set is decomposed into multiple scales using adaptive wavelet basis functions. The noise intensity of the decomposed signal is calculated, and the distribution characteristics of the noise intensity are evaluated using the maximum entropy criterion. The denoising threshold is dynamically adjusted based on the distribution characteristics, and the signal of the simplified feature set is subjected to secondary denoising processing based on the denoising threshold. The denoised signal is subjected to local weighted regression smoothing to generate the clean signal sequence.
5. The rapid detection method for heavy metals in mining area soil according to claim 1, characterized in that, The method for obtaining heavy metal morphology identifiers based on classification includes: The peak parameters and slope parameters in the pure signal sequence are extracted and set as a morphology-specific parameter set. A support vector machine classifier is used to train and match the morphology-specific parameter set. The support vector machine classifier determines the unique identifier of any heavy metal morphology, obtains the classification result of the unique identifier, and if the accuracy of the classification result meets the second preset standard, the unique identifier is retained and set as the heavy metal morphology identifier.
6. The rapid detection method for heavy metals in mining area soil according to claim 1, characterized in that, The step of obtaining the morphological category of the sample to be detected includes: Multidimensional feature analysis was performed on the morphological identifiers of heavy metals to construct multidimensional differentiation criteria and generate morphological classification templates. The spectral input data of the sample to be detected is obtained, the spectral input data is preprocessed and the feature parameters to be analyzed are extracted, the feature parameters to be analyzed are compared with the morphological classification template, and the matching degree is output. If the matching degree is higher than the preset threshold, the morphology classification template outputs the morphology category of the sample to be detected and records the comparison result of the morphology category.
7. The rapid detection method for heavy metals in mining area soil according to claim 1, characterized in that, The pollution level assessment results include: Based on the heavy metal morphology identifier, obtain the morphological category data of various heavy metals in the classified mining area soil samples, and generate category sample data; The sample data of the aforementioned categories are input into the spatial distribution model, and the spatial interpolation method is used to simulate the spatial distribution pattern of different heavy metal forms in the soil of the mining area, and the spatial distribution feature parameters are extracted. An iterative optimization algorithm is used to adjust the spatial distribution feature parameters to obtain the fitting accuracy of the spatial distribution model, and the model parameters are optimized based on the fitting accuracy. The optimized spatial distribution model calculates the pollution level assessment results of the morphological category based on the pollution load index.
8. The rapid detection method for heavy metals in mining area soil according to claim 7, characterized in that, The generation of an identification report and regional remediation pathway based on the pollution level assessment results includes: Based on the risk coefficient and the pollution level assessment results, high-risk area data is extracted, and the high-risk area data is summarized based on geographic information to generate the identification report; Based on the heavy metal morphology distribution information in the identification report, a remediation strategy mapping model is constructed. The remediation strategy mapping model matches data from any high-risk area with a remediation method to generate a remediation path. A feasibility analysis is performed on the repair path. If the feasibility meets the third preset standard, the repair path is set as the regional repair path.
9. A rapid detection method for heavy metals in mining area soil according to claim 8, characterized in that, The construction of the repair strategy mapping model includes: Based on the multidimensional feature analysis method, the morphological distribution information of heavy metals is associated with influencing factors to create a multidimensional remediation strategy mapping model. The multidimensional repair strategy mapping model is trained based on a machine learning algorithm, and the trained multidimensional repair strategy mapping model is set as the repair strategy mapping model.
10. A rapid detection system for heavy metals in mining area soil, characterized in that, The rapid detection system for heavy metals in mining area soil includes: The data processing module is used to acquire spectral data of soil samples from the mining area, perform preliminary processing on the spectral data to generate an initial dataset, perform dimensionality reduction processing on the initial dataset, extract a simplified feature set, and sort the features. The classification module is used to denoise the simplified feature set, generate a clean signal sequence, extract feature parameters related to the heavy metal speciation in the soil from the clean signal sequence, and obtain heavy metal speciation identifiers based on the classification method. The detection module is used to construct a differentiation standard based on the heavy metal morphology identifier and compare it with the sample to be detected to obtain the morphology category of the sample to be detected. The assessment module is used to combine the morphological categories with the spatial distribution model to simulate the spatial distribution pattern of heavy metals in the mining area soil, optimize the model parameters of the spatial distribution model to generate pollution level assessment results, and generate identification reports and regional remediation paths based on the pollution level assessment results.