Digital fingerprint spectrum driven multi-component heavy metal intelligent detection method
By employing a digital fingerprint spectroscopy-driven intelligent detection method for multi-component heavy metals, combined with combinatorial chemical probes and deep learning models, this method solves the problems of high cost, complex operation, and limited spectral resolution of traditional heavy metal detection equipment. It achieves rapid and high-precision multi-component heavy metal detection, making it suitable for rapid on-site detection.
Patent Information
- Application Number
- CN202511408150.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-29
- Publication Date
- 2025-11-21
AI Technical Summary
Traditional heavy metal detection methods suffer from high equipment costs, complex operation, difficulty in meeting the needs of simultaneous multi-component detection, and limited spectral resolution capabilities, which cannot effectively resolve nonlinear coupling effects in multi-component mixed systems, resulting in low detection accuracy and difficulty in meeting the requirements of real-time performance and rapid response.
A digital fingerprint spectroscopy-driven intelligent detection method for multi-component heavy metals is developed. This method utilizes a closed-loop technical path of combined chemical probe color development, data augmentation, global modeling, and lightweight system integration. The process includes combined chemical probe configuration, UV-Vis spectral acquisition, digital fingerprint spectral characterization, deep learning model development, and lightweight system integration to achieve rapid and accurate detection of multi-component heavy metal concentrations.
It enables rapid, high-precision, and low-cost multi-component heavy metal detection, overcoming the limitations of traditional methods such as high equipment cost, complex operation, and limited spectral resolution. It is suitable for rapid on-site detection, reducing detection time and hardware investment, and improving detection efficiency and accuracy.
Smart Images

Figure CN120992531A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of analytical chemistry and instrumental analysis technology, specifically to a digital fingerprint spectroscopy-driven intelligent detection method for multi-component heavy metals. Background Technology
[0002] Heavy metal pollution has become a major environmental issue of global concern. Many heavy metal elements, such as mercury, cadmium, and lead, are highly biotoxic, can persist in the environment for a long time, and accumulate through the food chain, posing a serious threat to human health and ecosystems. For example, mercury poisoning can cause neurological diseases, while cadmium pollution can lead to health problems such as kidney damage. Therefore, accurate determination of heavy metal content in the environment and various samples is of great significance for pollution control, risk assessment, and protection of public health. However, traditional heavy metal detection methods have many limitations. For example, although atomic absorption spectrometry and inductively coupled plasma mass spectrometry are highly accurate, they require expensive instruments and equipment, resulting in huge upfront costs. In addition, these methods have complex operating procedures, usually involving cumbersome sample pretreatment steps and instrument operation, and long detection cycles, making it difficult to meet the needs of real-time rapid detection. More importantly, traditional methods often can only analyze single heavy metals and cannot meet the requirements of simultaneous detection of multiple heavy metals in real-world environments. Furthermore, traditional spectroscopic analysis techniques are limited by linear model assumptions, making it difficult to effectively resolve the nonlinear interactions between heavy metal ions in multi-component mixed systems and the resulting spectral overlap effects, thus leading to low detection accuracy and limited applicability.
[0003] Currently, while traditional heavy metal detection methods can achieve high accuracy, they often rely on expensive and complex instruments and equipment, and their detection procedures are cumbersome and time-consuming. These methods exhibit significant shortcomings when faced with simultaneous analysis of multiple heavy metal components. Particularly in complex environmental samples, traditional techniques are easily constrained by insufficient sensitivity, significant cross-interference, and limited spectral resolution, thus affecting the accuracy and applicability of the detection. Therefore, they fail to meet the comprehensive demands of modern detection methods for real-time performance, rapid response, efficient analysis, and low-cost operation.
[0004] The main problems with existing technologies and corresponding improvement measures are as follows:
[0005] (1) Insufficient probe performance: Traditional probes are mostly designed for single heavy metals, and the color reaction has low specificity. In the case of multiple components coexisting, the spectral signal is easily subject to cross interference, especially when detecting low concentrations, the sensitivity is poor. In addition, in order to adapt to the detection needs of different metal ions, the reaction conditions need to be frequently adjusted, which makes it difficult to meet the requirements of simultaneous detection of multiple heavy metals in complex environmental samples (such as sewage and soil leachate).
[0006] (2) Limited spectral analysis capability: Traditional spectral analysis techniques are only suitable for extracting linear absorption features of single components and cannot effectively analyze nonlinear coupling effects (such as signal superposition and feature overlap) in multi-component mixed systems. For complex features of full band and high dimension in digital fingerprint spectrum, traditional methods lack sufficient characterization capability, thus limiting the further improvement of detection accuracy.
[0007] (3) Data and model limitations: The amount of spectral data obtained through traditional experimental methods is small, the diversity is insufficient, and the standardization is low, which makes it difficult to meet the needs of deep learning models for large-scale training data. This can easily lead to model overfitting or poor generalization ability. At the same time, existing chemometric models rely on manual feature engineering and have weak ability to analyze global features in digital fingerprint spectra. They cannot achieve end-to-end synchronous prediction of multi-component concentrations. In addition, the lack of efficient model integration schemes makes it difficult for these models to be adapted to low-cost portable devices.
[0008] (4) Equipment cost and portability barriers: Traditional multi-component detection equipment is bulky and complex to operate. It usually requires professional personnel to operate in a laboratory environment. The equipment cost is as high as hundreds of thousands or even millions of yuan, making it difficult to promote to rapid on-site detection scenarios (such as emergency monitoring of rivers and online monitoring of industrial wastewater). This high cost and complex operation process limit its real-time and mobile detection capabilities and cannot meet the diverse needs in practical applications. Summary of the Invention
[0009] The purpose of this invention is to provide a digital fingerprint spectroscopy-driven intelligent detection method for multi-component heavy metals. Through a closed-loop technical path of "combined chemical probe color development - data enhancement - global modeling - lightweight system integration", it systematically solves the problems of traditional heavy metal detection methods in terms of simultaneous multi-component detection, sensitivity improvement, insufficient spectral resolution, and high equipment cost.
[0010] This invention is achieved through the following technical solution:
[0011] This invention is a digital fingerprint spectroscopy-driven intelligent detection method for multi-component heavy metals, comprising the following steps:
[0012] (1) Preparation of multi-component heavy metal dispersion system combined chemical probes: Based on the chemical characteristics of the target heavy metal ions (e.g., antimony, iron, nickel, cadmium, copper, etc.), colorimetric reagents containing different coordination groups are selected and combined into combined chemical probes through orthogonal experimental design. Then, these combined chemical probes are mixed with single heavy metal solutions respectively. By comparing and analyzing the reaction effects of different probe combinations, their response sensitivity, selectivity and stability to the target metal ions are examined. Finally, probe combinations that can produce obvious differential signals are screened out, thereby ensuring effective signal separation and low interference analysis in the detection of multi-metal mixed samples.
[0013] (2) Ultraviolet-visible spectral acquisition of multi-component heavy metal dispersion system: Different types of heavy metals are mixed in a random combination manner and a large number of multi-component heavy metal mixed simulation samples are prepared. Then, a chemical probe is added to the prepared multi-component heavy metal mixed simulation sample to induce a specific color reaction. Next, the UV-VIS absorption curve data of the multi-component heavy metal dispersion system sample is collected in batches using a UV-VIS spectrometer.
[0014] (3) High-dimensional characterization of UV-Vis digital fingerprint spectrum based on stoichiometric transformation: Based on the acquired UV-VIS absorption curve dataset of multi-component heavy metal dispersion samples, one-dimensional spectral data is transformed into two-dimensional digital fingerprint spectrum through a specific algorithm to realize high-dimensional characterization of spectral signals and provide rich feature inputs for subsequent deep learning modeling.
[0015] (4) Data standardization and dataset preparation: In order to improve the generalization performance of the model, systematic preprocessing and enhancement operations were performed on the original spectral data. By introducing dynamic Gaussian noise, spectral mixing and concentration gradient perturbation, a variety of samples were generated, thereby effectively expanding the scale of the dataset. At the same time, the spectral features were standardized to eliminate the influence of different dimensions. In view of the skewed distribution of heavy metal concentration, the natural logarithmic transformation was used to narrow the numerical range of high concentration samples, further improving the analytical accuracy and sensitivity of the model in the low concentration region.
[0016] (5) Development of deep learning model for multi-component heavy metal dispersion system. Through deep learning model, end-to-end design is adopted to realize full-process automation from input of raw spectral data to output of multi-component concentration. The model input is a standardized digital fingerprint spectral feature vector. After capturing the local correlation pattern of the spectrum through multiple two-dimensional convolutional layers, dual attention calculation of channel and space is inserted. After compressing the feature dimension through global average pooling, the real concentration value is output by the fully connected layer.
[0017] In terms of training strategy, the data is divided into training set and test set. The training set is used for model validation and dynamic hyperparameter tuning. The Adam optimizer is used with mean absolute error (MAE) as the loss function.
[0018] (6) Model evaluation: Root mean square error (RMSE) and coefficient of determination (R²) were used. 2 The model performance is evaluated using the mean absolute error (MAE). When the evaluation index meets the preset threshold, the model is used for intelligent detection of multi-component heavy metal concentration.
[0019] Coefficient of determination (R) 2 The formula is as follows:
[0020]
[0021] Among them, R 2 As the coefficient of determination, The summation symbol represents summing over samples 1 to N, where N is the total number of samples, and x... i This represents the true value of the independent variable for the i-th sample. Let y be the sample mean of the independent variable x. i This represents the true value of the dependent variable for the i-th sample. Let y be the sample mean of the independent variable y;
[0022] The formula for the root mean square error (RMSE) is as follows:
[0023]
[0024] Where RMSE is the root mean square error. n is the sample size, and y i This represents the true value of the dependent variable for the i-th sample. This represents the model prediction value for the i-th sample. The symbol represents the summation over samples 1 to N, where N is the total number of samples.
[0025] The formula for Mean Absolute Error (MAE) is as follows:
[0026]
[0027] Where MAE is the mean absolute error, n is the sample size, and y i This represents the true value of the dependent variable for the i-th sample. This represents the model prediction value for the i-th sample. Let be the absolute value of the residual of the i-th sample. The summation symbol represents summing over samples 1 to N, where N is the total number of samples.
[0028] (7) Development of intelligent detection platform for concentration of multi-component heavy metal dispersion: To ensure the operability and flexibility of the method, this invention develops an interactive prediction platform based on a well-trained deep learning model. This platform can directly predict the concentration of metals contained in the sample end-to-end by directly inputting the digital fingerprint spectrum of the heavy metal mixed solution.
[0029] Furthermore, in step (1), antimony, iron, nickel, cadmium, and copper were first selected as target analytes, and multiple combined chemical probes were screened based on their chemical properties. Subsequently, the prepared combined chemical probes were mixed with a single metal solution, and the colorimetric reaction effect was observed. By comparing and analyzing the color difference values corresponding to each combined chemical probe, combined chemical probe 1 with the best color difference performance was selected.
[0030] Further, in step (2), Sb, Fe, Ni, Cd, and Cu metals were randomly mixed in volume to achieve different concentration gradients, and the mixed solution was stored in a test tube. The metal mixture solution in the test tube was then transferred to a 96-well plate, and a pre-screened combined chemical probe 1 was added. The samples were then optically characterized using a microplate reader, with a wide detection range of 230-780 nm (covering the UV-Vis characteristic absorption region), and high-resolution spectral acquisition was performed at 2 nm intervals. During the scanning process, absorbance data for 276 wavelength nodes were calculated using the absorbance calculation formula for each sample well, constructing a two-dimensional absorbance matrix (number of samples × number of wavelengths). The absorbance calculation formula is as follows:
[0031]
[0032] in, The summation symbol indicates a linear superposition of the five terms (i is the index of the component, from 1 to 5), and A(λ) is the absorbance at wavelength λ, ∈ i (λ) is the molar absorptivity of the i-th heavy metal, c i Let be the concentration of the i-th metal, and l be the optical path length.
[0033] Furthermore, in step (3), the obtained absorbance data is subjected to spectral transformation, resulting in 2700 Gram difference field (GADF), 2700 Gram sum field (GASF), and 2700 wavelet transform (CWT) maps, ultimately forming three types of spectral data. The formulas for Gram difference field (GADF), Gram sum field (GASF), and wavelet transform (CWT) are as follows:
[0034]
[0035] Among them, GADF i,j and GASF i,jThe element in the i-th row and j-th column of the matrix represents the collaborative relationship between time points i and j. and The normalized time series elements (the values at the i and j-th time points, ranging from [0,1]);
[0036]
[0037] Among them, W f (a, b) are wavelet transform coefficients, a is the scaling parameter, b is the translation parameter, f(t) is the continuous signal to be analyzed, φ is the conjugate function, and t is the integration variable.
[0038] Furthermore, in step (4), the original spectral dataset and concentration data are first augmented by adding Gaussian noise and combining it with spectral mixing techniques. Then, the absorbance data is standardized, and subsequently, the natural logarithmic transformation is used to improve the skewed distribution of the data. The Gaussian noise formula, spectral mixing formula, absorbance data standardization formula, and natural logarithmic transformation formula are as follows:
[0039] A noisy =A ornig +N(b,a 2 (8)
[0040] Among them, A noisy For the data after adding noise, A orig The original spectral data are given, and N is Gaussian noise with mean b and standard deviation a.
[0041] A mixed =αA1+(1-α)A2(α~(0.2-0.8)) (9)
[0042] Among them, A mixed The mixed spectral data are represented by A1 and A2, which are two randomly selected spectral samples, and α is the mixing ratio (between 0.2 and 0.8).
[0043]
[0044] Among them, X scaled For the standardized data, X is the original input data (features or labels), μ is the mean absorbance, and σ is the standard deviation;
[0045] y log =In(1+y) (11)
[0046] Where y is the original concentration value, y log This is the value after logarithmic transformation.
[0047] Further, the deep learning model in step (5) is a CNN-CBAM model, which includes: the input layer receives a two-dimensional digital fingerprint map, and extracts local features sequentially through a 3×3 convolutional layer, a batch normalization layer, and a ReLU activation layer; a CBAM attention module is inserted, which first calculates the inter-channel dependency through channel attention, and then generates spatial position weights through spatial attention; after compressing the feature dimension through a global average pooling layer, the concentration values of 5-10 heavy metals are output through a fully connected layer;
[0048] In terms of training strategy, stratified random sampling was used to divide the original dataset into training, validation, and test sets in an 8:1:1 ratio. The Adam optimizer (learning rate 1×10⁻⁶) was selected for optimization. -4 The model uses the mean absolute error (MAE) as the loss function and improves generalization ability through a triple mechanism: dynamic learning rate decay, early stopping mechanism, and Dropout regularization (ratio 0.1). The training batch size is set to 16, and the maximum number of iterations is limited to 200. Finally, the trained CNN-CBAM model can directly predict the concentration of each metal in the solution based on the absorbance of the metal mixture solution.
[0049] Furthermore, in step (6), in the metal concentration prediction task, the study used R... 2 The performance of the model is evaluated using the MAE and RMSE metrics, with preset thresholds for the evaluation metrics: R 2 ≥0.9, RMSE≤1.5, MAE≤1.0, when the model linearly fits the scatter plot of the predicted and actual values on the test set with R0.9. 2 The model is deemed qualified if the above thresholds are met.
[0050] Furthermore, the present invention also provides a multi-component heavy metal intelligent detection system for implementing the above method. The system includes:
[0051] Spectral acquisition module: includes a UV-VIS spectrometer and a 96-well plate. The spectrometer is used to acquire the UV-Vis absorption curve of the heavy metal mixed sample after the addition of combined chemical probes, and the 96-well plate is used to hold the sample and carry out the colorimetric reaction.
[0052] Data processing module: used to convert one-dimensional absorption curves into two-dimensional digital fingerprint maps and perform data augmentation and standardization processing.
[0053] Model inference module: Embedded with a trained CNN-CBAM deep learning model, it receives digital fingerprint data and outputs the concentration of multiple heavy metal components.
[0054] Visualization and Interaction Module: A graphical user interface built with PyQt5 supports spectral data import, fingerprint spectrum generation, concentration prediction result display, and export.
[0055] Furthermore, the system is compatible with portable devices, supports offline inference and local data storage, and can be applied to the simultaneous detection of multiple heavy metals (antimony, iron, nickel, cadmium, and copper) at sewage treatment plants, river monitoring stations, and industrial pollution sources, with a detection cycle of ≤10 minutes.
[0056] The present invention has the following beneficial effects:
[0057] 1. This invention possesses rapid and high-precision detection capabilities. It employs orthogonal experiments to screen composite colorimetric systems, selecting multiple chromogenic agents and optimizing reaction conditions (such as pH, concentration ratio, and reaction time) based on the characteristics of multi-component heavy metals. Utilizing the specific complexation reaction between the chromogenic agents and heavy metals, different heavy metals generate differentiated colorimetric signals in the mixed solution, forming a distinguishable spectral response substrate. The orthogonal screening process, reaction condition optimization method, and components of the composite colorimetric system enable rapid and high-precision detection, overcoming the limitations of traditional methods that are cumbersome and have long detection cycles. It can simultaneously analyze the concentrations of multiple heavy metal components without complex pretreatment, significantly improving detection efficiency and accuracy, and meeting the rapid analysis needs of samples from complex environments.
[0058] 2. This invention possesses strong anti-interference and adaptability. Through chemical probe optimization and data augmentation techniques, it effectively suppresses multi-component cross-interference, enhances spectral signal discrimination, improves the model's robustness to instrument fluctuations and low-concentration samples, and achieves stable detection across the entire concentration range. It converts the one-dimensional spectral sequence of multi-component heavy metal solutions into two-dimensional spectra. Through polar coordinate mapping and trigonometric function (sine / cosine) encoding, it preserves the amplitude and phase information of the spectral data, forming an image-based representation containing inter-wavelength correlation features. This method overcomes the limitations of traditional linear superposition of spectra, enhancing the nonlinear discrimination of multi-component signals through the spatial structure of the spectra, providing high-dimensional feature input for deep learning. The spectral image conversion algorithm includes polar coordinate mapping rules, amplitude normalization methods, and spectra resolution parameter settings. It utilizes trigonometric function encoding to implicitly preserve wavelength position information in spectra generation. The invention also includes a linked process for multi-component heavy metal solution spectral acquisition and spectra conversion.
[0059] 3. This invention features low cost and intelligence, integrating a lightweight end-to-end system that is compatible with low-cost portable devices, reducing hardware investment and operational barriers. It supports one-click automated detection, allowing even non-professionals to quickly obtain visualized results, thus driving the upgrade of detection technology towards on-site and intelligent operation. The system embeds the spectrum conversion module and pre-trained model into low-cost portable devices, integrating spectral acquisition, spectrum generation, model inference, and result visualization functions to achieve an automated "sample in - result out" detection process. The system's modular design adapts to different scenarios. The hardware is compatible with miniaturized spectrometers, while the software provides a one-click operation interface and visualized reports, lowering the professional operation threshold. The lightweight system architecture, combining hardware and software, includes an integrated solution for the spectrum conversion engine and model inference module; a functional design for the visualization analysis platform; and an embedded deployment solution suitable for rapid on-site detection, including low-power processor adaptation, offline inference functionality, and local data storage mechanisms.
[0060] 4. This invention constructs an end-to-end model based on deep learning algorithms. After inputting a spectral image, it extracts local spectral features, uses pooling layers to capture global contextual relationships, and fully connected layers to achieve simultaneous prediction of multi-component heavy metal concentrations. The model automatically analyzes spatial relationships and complex nonlinear patterns in the spectral image, overcoming the limitations of traditional linear spectral analysis. It utilizes a deep learning model architecture and a multi-level feature extraction method for high-dimensional image-based spectral data.
[0061] Of course, any product implementing this invention does not necessarily need to achieve all of the advantages described above at the same time. Attached Figure Description
[0062] Figure 1 A flowchart of a multi-component intelligent detection method for heavy metals;
[0063] Figure 2 Diagram of deep learning model architecture;
[0064] Figure 3 Screening diagram of combinatorial chemical probes for multi-component heavy metal dispersions;
[0065] Figure 4 This is an example diagram of spectrum conversion;
[0066] Figure 5 A scatter plot of the Gram angle difference field;
[0067] Figure 6 A scatter plot of Gram's angle and the field;
[0068] Figure 7 This is a scatter plot of wavelet transform.
[0069] Figure 8 A complete operation process diagram;
[0070] Figure 9 A detailed evaluation index table is provided. Detailed Implementation
[0071] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0072] Please see Figure 1-9 This invention provides a technical solution: a digital fingerprint spectroscopy-driven intelligent detection method for multi-component heavy metals. This example uses a five-element multi-component mixed heavy metal (antimony, iron, nickel, cadmium, and copper) as the detection target. It employs a combination of chemical probe color development and superimposed spectral acquisition. Gram angle difference field, Gram angle sum field, and wavelet transform are used to perform feature mining processing on the spectral data. Specifically, the Gram angle difference field and Gram angle sum field map the data to polar coordinate space to capture periodicity and trends. Wavelet transform decomposes local features through multi-scale analysis, and the optimal data feature expression form is selected through comparison. Finally, a CNN-CBAM model is used for deep learning modeling, successfully achieving rapid and synchronous prediction of multi-component heavy metal concentrations. This method significantly reduces detection costs and time while maintaining high accuracy, verifying the practicality and scalability of the technology. The specific implementation steps are as follows:
[0073] (1) Preparation of multi-component heavy metal dispersion system combined chemical probes: Based on the chemical characteristics of the target heavy metal ions (e.g., antimony, iron, nickel, cadmium, copper, etc.), colorimetric reagents containing different coordination groups are selected and combined into combined chemical probes through orthogonal experimental design. Then, these combined chemical probes are mixed with single heavy metal solutions respectively. By comparing and analyzing the reaction effects of different probe combinations, their response sensitivity, selectivity and stability to the target metal ions are examined. Finally, probe combinations that can produce significantly different signals are screened out, thereby ensuring effective signal separation and low-interference analysis in the detection of multi-metal mixed samples.
[0074] First, antimony, iron, nickel, cadmium, and copper were selected as target analytes, and several combined chemical probes were screened based on their chemical properties. Then, the prepared combined chemical probes were mixed with single-metal solutions, and their colorimetric reactions were observed. By comparing and analyzing the color difference values of each combined chemical probe, combined chemical probe 1 with the best color difference performance was selected.
[0075] (2) Ultraviolet-visible spectral acquisition of multi-component heavy metal dispersion system: different types of heavy metals are mixed in a random combination manner and a large number of multi-component heavy metal mixed simulation samples are prepared. Then, a chemical probe is added to the prepared multi-component heavy metal mixed simulation sample to induce a specific color reaction. Next, the UV-VIS absorption curve data of the multi-component heavy metal dispersion system sample is collected in batches using a UV-VIS spectrometer.
[0076] Sb, Fe, Ni, Cd, and Cu metals were randomly mixed in volume to achieve different concentration gradients, and the mixed solutions were stored in test tubes. The metal mixtures were then transferred to 96-well plates, and a pre-screened combined chemical probe 1 was added. The samples were then optically characterized using a microplate reader with a wide detection range of 230-780 nm (covering the UV-Vis characteristic absorption region), and high-resolution spectral acquisition was performed at 2 nm intervals. During the scanning process, absorbance data for each sample well was calculated using the absorbance calculation formula for 276 wavelength nodes, constructing a two-dimensional absorbance matrix (number of samples × number of wavelengths). The absorbance calculation formula is as follows:
[0077]
[0078] in, The summation symbol indicates a linear superposition of the five terms (i is the index of the component, from 1 to 5), and A(λ) is the absorbance at wavelength λ, ∈ i (λ) is the molar absorptivity of the i-th heavy metal, c i Let be the concentration of the i-th metal, and l be the optical path length.
[0079] (3) High-dimensional characterization of UV-Vis digital fingerprint spectrum based on stoichiometric transformation: Based on the obtained UV-VIS absorption curve dataset of multi-component heavy metal dispersion samples, one-dimensional spectral data is transformed into two-dimensional digital fingerprint spectrum through a specific algorithm to realize high-dimensional characterization of spectral signals and provide rich feature input for subsequent deep learning modeling.
[0080] The obtained absorbance data underwent spectral transformation, resulting in the transformation of 2700 Gram difference field (GADF), 2700 Gram sum field (GASF), and 2700 wavelet transform (CWT) spectra. Examples of the transformations can be found in [link to example]. Figure 4 The final result is three types of spectral data: Gram angle difference field (GADF), Gram angle sum field (GASF), and wavelet transform (CWT) formulas, as follows:
[0081]
[0082] Among them, GADF i,j and GASFi,j The element in the i-th row and j-th column of the matrix represents the collaborative relationship between time points i and j. and The normalized time series elements (the values at the i and j-th time points, ranging from [0,1]);
[0083]
[0084] Among them, W f (a, b) are wavelet transform coefficients, a is the scaling parameter, b is the translation parameter, f(t) is the continuous signal to be analyzed, φ is the conjugate function, and t is the integration variable.
[0085] (4) Data standardization and dataset preparation: In order to improve the generalization performance of the model, systematic preprocessing and enhancement operations were performed on the original spectral data. By introducing dynamic Gaussian noise, spectral mixing and concentration gradient perturbation, a variety of samples were generated, thereby effectively expanding the scale of the dataset. At the same time, the spectral features were standardized to eliminate the influence of different dimensions. In view of the skewed distribution of heavy metal concentration, the natural logarithmic transformation was used to narrow the numerical range of high concentration samples, further improving the analytical accuracy and sensitivity of the model in the low concentration region.
[0086] The original spectral dataset and concentration data were first augmented by adding Gaussian noise and combining it with spectral mixing techniques. Then, the absorbance data was standardized, and subsequently, a natural logarithmic transformation was used to improve the skewed distribution of the data. The formulas for Gaussian noise, spectral mixing, absorbance data standardization, and natural logarithmic transformation are as follows:
[0087] A noisy =A orig +N(b, a) 2 (8)
[0088] Among them, A noisy For the data after adding noise, A orig The original spectral data are given, and N is Gaussian noise with mean b and standard deviation a.
[0089] A mixed =αA1+(1-α)A2(α~(0.2-0.8)) (9)
[0090] Among them, A mixed The mixed spectral data are represented by A1 and A2, which are two randomly selected spectral samples, and α is the mixing ratio (between 0.2 and 0.8).
[0091]
[0092] Among them, X scaledFor the standardized data, X is the original input data (features or labels), μ is the mean absorbance, and σ is the standard deviation;
[0093] y log =In(1+y)(11)
[0094] Where y is the original concentration value, y log This is the value after logarithmic transformation.
[0095] (5) Development of a deep learning model for multi-component heavy metal dispersions: Through a deep learning model and an end-to-end design, the entire process from input of raw spectral data to output of multi-component concentrations is automated. Figure 2 The model input is a standardized digital fingerprint spectral feature vector. This vector is processed through multiple 2D convolutional layers to capture local spectral correlation patterns. Channel and spatial attention are then inserted for computation. After global average pooling to compress the feature dimension, a fully connected layer outputs the true concentration values. The deep learning model is a CNN-CBAM model, which includes an input layer receiving the 2D digital fingerprint spectrum, sequentially passing through a 3×3 convolutional layer, a batch normalization layer, and a ReLU activation layer to extract local features. A CBAM attention module is inserted, which first calculates inter-channel dependencies through channel attention and then generates spatial position weights through spatial attention. After global average pooling to compress the feature dimension, a fully connected layer outputs the concentration values of 5-10 heavy metals.
[0096] In terms of training strategy, the data was divided into training and testing sets. The training set was used for model validation and dynamic hyperparameter tuning. The Adam optimizer was employed, with mean absolute error (MAE) as the loss function. Stratified random sampling was used to divide the original dataset into training, validation, and testing sets in an 8:1:1 ratio. The optimization process used the Adam optimizer (learning rate 1×10⁻⁶). -4 The model uses the mean absolute error (MAE) as the loss function and improves generalization ability through a triple mechanism: dynamic learning rate decay, early stopping mechanism, and Dropout regularization (ratio 0.1). The training batch size is set to 16, and the maximum number of iterations is limited to 200. Finally, the trained CNN-CBAM model can directly predict the concentration of each metal in the solution based on the absorbance of the metal mixture solution.
[0097] (6) Model evaluation: Root mean square error (RMSE) and coefficient of determination (R²) were used. 2 The model performance is evaluated using the mean absolute error (MAE). When the evaluation index meets the preset threshold, the model is used for intelligent detection of multi-component heavy metal concentrations. Figure 5 , Figure 6 and Figure 7The CNN-CBAM model demonstrates excellent linear fitting between predicted and true values on the Gram angle difference field, Gram angle sum field, and wavelet transform map datasets, indicating that the model achieves high-precision prediction across the entire concentration gradient space. This superior performance is achieved, on the one hand, through…
[0098] The CNN-CBAM model extracts local features in the spectral space through convolutional neural networks and focuses on key information and suppresses noise using the CBAM attention mechanism, thus improving the ability to resolve complex features. On the other hand, it benefits from the rich information support of the spectral representation. For example, the Gram angle field converts time-series data into a two-dimensional image through trigonometric function sum and difference operations, preserving the dynamic correlation of the concentration sequence. The wavelet transform spectrum reveals the frequency domain features of the signal through multi-scale decomposition. The fusion of the two constructs a composite feature space containing spatiotemporal characteristics and frequency features. This synergistic effect of "efficient model resolution" and "multi-dimensional spectral representation" not only verifies the effectiveness of the CNN-CBAM model, but also illustrates the feasibility of the metal concentration prediction method based on spectral analysis from the dual dimensions of data representation and model design.
[0099] Coefficient of determination (R) 2 The formula is as follows:
[0100]
[0101] Among them, R 2 As the coefficient of determination, The summation symbol represents summing over samples 1 to N, where N is the total number of samples, and x... i This represents the true value of the independent variable for the i-th sample. Let y be the sample mean of the independent variable x. i This represents the true value of the dependent variable for the i-th sample. Let y be the sample mean of the independent variable y;
[0102] The formula for the root mean square error (RMSE) is as follows:
[0103]
[0104] Where RMSE is the root mean square error. n is the sample size, and y i This represents the true value of the dependent variable for the i-th sample. This represents the model prediction value for the i-th sample. The symbol represents the summation over samples 1 to N, where N is the total number of samples.
[0105] The formula for Mean Absolute Error (MAE) is as follows:
[0106]
[0107] Where MAE is the mean absolute error, n is the sample size, and y i This represents the true value of the dependent variable for the i-th sample. This represents the model prediction value for the i-th sample. Let be the absolute value of the residual of the i-th sample. The summation symbol represents summing over samples 1 to N, where N is the total number of samples.
[0108] In the task of predicting metal concentration, the study used R 2 The performance of the model is evaluated using the MAE and RMSE metrics, with preset thresholds for the evaluation metrics: R 2 ≥0.9, RMSE≤1.5, MAE≤1.0, when the model linearly fits the scatter plot of the predicted and actual values on the test set with R0.9. 2 The model is deemed qualified if the above thresholds are met.
[0109] (7) Development of an intelligent detection platform for the concentration of multi-component heavy metal dispersions: To ensure the operability and flexibility of the method, this invention develops an interactive prediction platform based on a trained deep learning model. This platform can directly predict the concentration of metals contained in a sample end-to-end by directly inputting the digital fingerprint spectrum of the heavy metal mixed solution, such as... Figure 8 As shown, the intelligent heavy metal concentration prediction system adopts a modular PyQt5 interface design, featuring a Windows 11-style light gray background with blue interactive elements, constructing a fully visualized operation platform from spectral data import to prediction result output. The top function control area integrates four buttons: "Select Data," "Fingerprint Spectrum Generation," "Predictive Analysis," and "Save Results." It supports loading Excel files containing UV-VIS spectral data, generating two-dimensional spectra based on the Gram Angular Difference Field (GADF) algorithm, calling pre-trained CNN-CBAM model inference, and exporting results in Excel / CSV format. The middle section is divided into left and right areas by a horizontal divider; the left side displays sample GADF spectra with ID labels in a scrolling layout, while the right side...
[0110] The QTableWidget dynamically generates a results table containing the "sample ID" and the concentrations of five metals: Sb, Fe, Ni, Cd, and Cu. The bottom status feedback area displays the processing progress and operation prompts in real time via a progress bar and dynamic labels.
[0111] In summary, the multi-component heavy metal intelligent detection system constructed in this invention is used to implement the aforementioned method, and its main components include:
[0112] Spectral acquisition module: includes a UV-VIS spectrometer and a 96-well plate. The spectrometer is used to acquire the UV-Vis absorption curve of the heavy metal mixed sample after the addition of combined chemical probes. The 96-well plate is used to hold the sample and carry out the colorimetric reaction.
[0113] Data processing module: used to convert one-dimensional absorption curves into two-dimensional digital fingerprint maps and perform data augmentation and standardization processing;
[0114] Model inference module: Embedded with a trained CNN-CBAM deep learning model, it receives digital fingerprint data and outputs the concentration of multiple heavy metal components;
[0115] Visualization and Interaction Module: A graphical user interface built with PyQt5 supports spectral data import, fingerprint spectrum generation, concentration prediction result display, and export.
[0116] The system is compatible with portable devices, supports offline inference and local data storage, and can be applied to the simultaneous detection of multi-component heavy metals (antimony, iron, nickel, cadmium, and copper) at wastewater treatment plants, river monitoring stations, and industrial pollution sources. The detection cycle is ≤10 minutes. It is compatible with low-cost portable devices, enabling minute-level, high-precision, multi-component heavy metal simultaneous detection. This system not only significantly shortens detection time and reduces equipment costs but also improves the convenience of on-site operation, making it suitable for various scenarios such as wastewater treatment plants, river monitoring stations, and industrial pollution sources. This achievement provides strong support for the development of environmental monitoring technology towards intelligence and portability, and also lays a solid foundation for the expansion of applications in related fields.
[0117] By designing multi-ligand chemical probes, highly specific heavy metal spectral fingerprint signals can be generated, significantly improving the sensitivity and selectivity of multi-component detection. Simultaneously, by introducing deep learning models into the field of UV-Vis digital fingerprint spectroscopy analysis, their self-attention mechanism is utilized to capture global dependencies, overcoming the limitation of traditional models that can only extract local features.
[0118] This paper combines deep learning models with digital fingerprint spectroscopy technology for the detection of heavy metals. By fusing full-band spectral features and constructing a nonlinear model, a novel technical approach is provided for the simultaneous analysis of multi-component heavy metals.
[0119] By organically combining ultraviolet-visible digital fingerprinting spectroscopy with deep learning, this approach comprehensively addresses the bottlenecks of existing technologies in areas such as multi-component detection efficiency, spectral resolution accuracy, model training costs, and device portability. Specific objectives include the following:
[0120] In terms of analytical detection methods, by introducing combined chemical probes and full-spectrum ultraviolet-visible digital fingerprinting technology, differentiated characteristic expression of multi-component heavy metal spectral fingerprint signals is achieved, effectively suppressing cross-interference and significantly improving detection sensitivity under low concentration conditions. This improvement successfully overcomes the detection limitations of traditional probes caused by signal overlap and stringent reaction conditions.
[0121] At the model construction level, the powerful analytical capabilities of deep learning models for global features are fully utilized to overcome the limitations of traditional linear models in complex spectral analysis, establishing an end-to-end nonlinear modeling framework. This framework can accurately identify high-dimensional feature correlations in digital fingerprint spectra, thereby achieving simultaneous prediction of multi-component heavy metal concentrations and significantly improving the accuracy and reliability of detection.
[0122] In terms of data and algorithm optimization, a data augmentation strategy is adopted to generate large-scale virtual samples, which effectively alleviates the challenges brought about by small sample training, while enhancing the model's adaptability to instrument fluctuations and environmental noise, and ensuring the stability and consistency of detection results across the entire concentration gradient range.
[0123] In terms of application system design, a lightweight end-to-end detection platform was developed, which seamlessly integrates deep learning models with low-cost digital fingerprint spectrometers, realizing fully automated operation from "spectral acquisition-model prediction-result output". This platform not only significantly reduces equipment costs, but also supports one-click operation on site and wireless data transmission, providing strong support for the development of environmental monitoring technology towards intelligence and portability.
[0124] By adopting a closed-loop technical path of "multi-probe fingerprinting - full-spectrum feature modeling - lightweight integration", chemical mechanisms and deep learning are deeply integrated. This not only improves the separation and characterization of spectral signals, but also breaks through the data dependence and analysis bottlenecks of traditional models. Ultimately, a fast, high-precision, and low-cost multi-component heavy metal detection system is formed, providing real-time and reliable technical support for actual environmental monitoring.
[0125] Atomic absorption spectrometry (AAS) is an analytical method that determines the content of an element by measuring the degree to which its atomic vapor absorbs light of a specific wavelength. For example, it can be used to detect heavy metal content in water. The principle is that heavy metal atoms absorb light of specific wavelengths, and the concentration is calculated by measuring the amount of absorption.
[0126] Inductively coupled plasma mass spectrometry (ICP-MS) is a highly sensitive analytical technique that uses inductively coupled plasma to ionize a sample, and then analyzes the mass-to-charge ratio of the ions using a mass spectrometer to determine the elemental composition and content. It is commonly used for trace heavy metal detection, but the equipment is expensive and the operation is complex.
[0127] Ultraviolet-Visible (UV-VIS): The collective term for the ultraviolet (UV) and visible (Vis) regions of the electromagnetic spectrum. Many substances have specific absorption or reflection characteristics in this wavelength range, which can be used to detect the composition and concentration of substances, such as the detection of heavy metal liquid dispersions mentioned in this article.
[0128] Molar absorptivity: In spectral analysis, it is a physical quantity that characterizes a substance's ability to absorb light of a specific wavelength. Together with the substance's concentration and optical path length, it determines the magnitude of absorbance and reflects the substance's efficiency in absorbing light.
[0129] Combinatorial chemical probes: Probes composed of multiple ligands can bind characteristically to different heavy metal ions, forming unique "spectral fingerprint clusters" in digital fingerprint spectra. These probes are used for the detection of multi-component heavy metals and solve problems such as cross-interference of traditional probes.
[0130] Digital fingerprint spectroscopy: A wide-band spectrum covering the entire spectrum, capturing the full-spectrum characteristic fingerprint signal after the interaction of heavy metal ions with probes using a high-resolution spectrometer, and converting it into a high-dimensional feature vector map, providing rich input features for deep learning.
[0131] Spectral fingerprint clusters: A unique set of spectral signals formed in digital fingerprint spectra after a combination of chemical probes binds to different heavy metal ions, which can be used to distinguish different heavy metals.
[0132] High-throughput colorimetric reaction experiment: This experimental method utilizes equipment such as 96-well plates to simultaneously perform colorimetric reactions on a large number of samples. It can process and detect samples in batches in a short time, improving experimental efficiency. In this paper, it is used for UV-VIS superimposed spectral acquisition of mixed heavy metals.
[0133] Deep learning model: A machine learning model based on multi-layer neural networks that can automatically learn complex features and patterns from data. In this paper, it is used to predict heavy metal concentration from digital fingerprint spectra.
[0134] Self-attention mechanism: A mechanism in deep learning models that can capture long-distance dependencies between elements in data, such as global dependencies across wavelengths in digital fingerprint spectra, thereby improving the model's ability to analyze global features.
[0135] Data augmentation: New training data is generated by adding noise, spectral mixing and other techniques to increase the diversity and quantity of data, meet the needs of deep learning models for large-scale data and improve the generalization ability of the models.
[0136] Root Mean Square Error (RMSE): A metric that measures the deviation between model predictions and actual values. It is calculated by taking the square root of the average of the squared prediction errors. The smaller the value, the higher the prediction accuracy.
[0137] Coefficient of determination (R) 2 ): An index that evaluates how well a model fits the data. It ranges from 0 to 1. The closer it is to 1, the higher the proportion of data variance that the model can explain and the better the prediction effect.
[0138] Mean Absolute Error (MAE): Calculates the average absolute error between the predicted value and the actual value. It measures the average magnitude of the prediction error. The smaller the value, the more accurate the prediction.
[0139] Gram angular difference field (GADF): a technique that maps one-dimensional data to polar coordinate space to generate two-dimensional spectra. It preserves the periodicity and trend of data through trigonometric function difference operations and is used in this paper for feature mining of spectral data.
[0140] Gram angle sum and field (GASF): Similar to GADF, it uses trigonometric functions and operations to map data to polar coordinate space to generate a spectrum and capture data features.
[0141] Wavelet transform: a multi-scale analysis method that decomposes a signal into components of different frequencies and scales to extract local features of the data. In this paper, it is used for spectral data processing.
[0142] Multiwell plate: An experimental plate with multiple wells, such as a 96-well plate, which can process multiple samples at the same time and is used for high-throughput experiments. In this paper, it is used to store mixed solutions of heavy metals and to carry out colorimetric reactions.
[0143] Microplate reader: an instrument used to detect the optical properties of samples, which can measure parameters such as absorbance. In this article, it is used to acquire spectral signals from samples in a 96-well plate.
[0144] Lightweight system integration: Optimize complex detection systems to make them small in size, low in power consumption, and low in cost, adaptable to portable devices, and enable rapid on-site detection. This is achieved through technologies such as low-power spectral sensors.
[0145] Embedded deployment: This involves embedding software and models into small hardware devices, enabling them to operate independently without relying on large servers. In this paper, it is used in a portable detection system that supports offline inference and local data storage.
[0146] Biotoxicity: The property of a substance to cause harm to living organisms (such as humans, animals and plants). In this article, it refers to the harm of heavy metals such as mercury and cadmium to human health and ecosystems.
[0147] Bioaccumulation in the food chain: The phenomenon in which harmful substances such as heavy metals gradually accumulate as the trophic level increases in the food chain. For example, if small fish ingest heavy metals, the heavy metals will accumulate in the bodies of large fish after they eat the small fish, which may eventually harm humans.
[0148] Orthogonal experimental design: a scientific method for designing experiments. By rationally arranging experimental factors and levels, the number of experiments can be reduced while obtaining comprehensive experimental information. In this paper, it is used to screen the optimal reaction conditions for combined chemical probes.
[0149] Gaussian noise: Random noise that follows a Gaussian distribution (normal distribution) can be used in data processing to simulate interference in the real environment. In this paper, it is used for data augmentation to improve the model's noise resistance.
[0150] Spectral mixing technique: A technique that mixes multiple spectral samples in a certain proportion to generate new samples, used for data augmentation and to increase the diversity of training data.
[0151] Skewed distribution: The data distribution is asymmetrical, with a longer tail on one side. This paper uses natural logarithmic transformation to process the skewed distribution of heavy metal concentrations, thereby improving the model's analytical accuracy for low-concentration samples.
[0152] Natural logarithm transformation: A mathematical transformation method that takes the natural logarithm of data, which can transform skewed distributed data into one that is closer to a normal distribution and narrow the range of high-value data. In this paper, it is used to process heavy metal concentration data.
[0153] PyQt5: A Python library for creating graphical user interfaces (GUIs). In this paper, it is used to develop a visual interface for an intelligent prediction system of heavy metal concentrations.
[0154] The preferred embodiments of the present invention disclosed above are merely illustrative of the invention. These preferred embodiments do not exhaustively describe all details, nor do they limit the invention to the specific implementations described. Clearly, many modifications and variations can be made based on the content of this specification. This specification selects and specifically describes these embodiments to better explain the principles and practical applications of the invention, thereby enabling those skilled in the art to better understand and utilize the invention. The invention is limited only by the claims and their full scope and equivalents.
Claims
1. A digital fingerprint spectroscopy-driven intelligent detection method for multi-component heavy metals, characterized in that, Includes the following steps: (1) Preparation of multi-component heavy metal dispersion system combined chemical probes: Based on the chemical characteristics of the target heavy metal ions, colorimetric reagents containing different coordination groups are selected and combined into combined chemical probes through orthogonal experimental design. Then, these combined chemical probes are mixed with single heavy metal solutions respectively. By comparing and analyzing the reaction effects of different probe combinations, their response sensitivity, selectivity and stability to the target metal ions are examined. Finally, probe combinations that can produce obvious differential signals are screened out, thereby ensuring effective signal separation and low interference analysis in the detection of multi-metal mixed samples. (2) Ultraviolet-visible spectral acquisition of multi-component heavy metal dispersion system: Different types of heavy metals are mixed in a random combination manner and a large number of multi-component heavy metal mixed simulation samples are prepared. Then, a chemical probe is added to the prepared multi-component heavy metal mixed simulation sample to induce a specific color reaction. Next, the UV-VIS absorption curve data of the multi-component heavy metal dispersion system sample is collected in batches using a UV-VIS spectrometer. (3) High-dimensional characterization of UV-Vis digital fingerprint spectrum based on stoichiometric transformation: Based on the acquired UV-VIS absorption curve dataset of multi-component heavy metal dispersion samples, one-dimensional spectral data is transformed into two-dimensional digital fingerprint spectrum through a specific algorithm to realize high-dimensional characterization of spectral signals and provide rich feature inputs for subsequent deep learning modeling. (4) Data standardization and dataset preparation: In order to improve the generalization performance of the model, systematic preprocessing and enhancement operations were performed on the original spectral data. Dynamic Gaussian noise, spectral mixing and concentration gradient perturbation were introduced to generate diverse samples and expand the scale of the dataset. At the same time, the spectral features were standardized to eliminate the influence of different dimensions. In view of the skewed distribution of heavy metal concentration, the natural logarithmic transformation was used to narrow the numerical range of high concentration samples and improve the analytical accuracy and sensitivity of the model in the low concentration region. (5) Development of deep learning model for multi-component heavy metal dispersion system. Through deep learning model, end-to-end design is adopted to realize full-process automation from input of raw spectral data to output of multi-component concentration. The model input is a standardized digital fingerprint spectral feature vector. After capturing the local correlation pattern of the spectrum through multiple two-dimensional convolutional layers, dual attention calculation of channel and space is inserted. After compressing the feature dimension through global average pooling, the fully connected layer outputs the real concentration value. In terms of training strategy, the data is divided into training set and test set. The training set is used for model validation and dynamic hyperparameter tuning. The Adam optimizer is used with mean absolute error (MAE) as the loss function. (6) Model evaluation: Root mean square error (RMSE) and coefficient of determination (R²) were used. 2 The model performance is evaluated using the mean absolute error (MAE). When the evaluation index meets the preset threshold, the model is used for intelligent detection of multi-component heavy metal concentration. Coefficient of determination (R) 2 The formula is as follows: Among them, R 2 As the coefficient of determination, The summation symbol represents summing over samples 1 to N, where N is the total number of samples, and x... i This represents the true value of the independent variable for the i-th sample. Let y be the sample mean of the independent variable x. i This represents the true value of the dependent variable for the i-th sample. Let y be the sample mean of the independent variable y; The formula for the root mean square error (RMSE) is as follows: Where RMSE is the root mean square error. n is the sample size, and y i This represents the true value of the dependent variable for the i-th sample. This represents the model prediction value for the i-th sample. The symbol represents the summation over samples 1 to N, where N is the total number of samples. The formula for Mean Absolute Error (MAE) is as follows: Where MAE is the mean absolute error, n is the sample size, and y i This represents the true value of the dependent variable for the i-th sample. This represents the model prediction value for the i-th sample. Let be the absolute value of the residual of the i-th sample. The symbol represents the summation over samples 1 to N, where N is the total number of samples. (7) Development of intelligent detection platform for concentration of multi-component heavy metal dispersion: To ensure the operability and flexibility of the method, an interactive prediction platform is developed based on the trained deep learning model. This platform can directly predict the concentration of metals contained in the sample end-to-end by directly inputting the digital fingerprint spectrum of the heavy metal mixed solution.
2. The digital fingerprint spectral-driven intelligent detection method for multi-component heavy metals according to claim 1, characterized in that, First, antimony, iron, nickel, cadmium, and copper were selected as target analytes, and multiple combined chemical probes were screened based on their chemical properties. Then, the prepared combined chemical probes were mixed with single metal solutions, and their colorimetric reaction effects were observed. By comparing and analyzing the color difference values corresponding to each combined chemical probe, the combined chemical probe 1 with the best color difference performance was selected.
3. The digital fingerprint spectral-driven intelligent detection method for multi-component heavy metals according to claim 1, characterized in that, In step (2), Sb, Fe, Ni, Cd, and Cu metals were randomly mixed in volume to achieve different concentration gradients, and the mixed solutions were stored in test tubes. The metal mixtures in the test tubes were then transferred to 96-well plates, and pre-screened combined chemical probe 1 was added. The samples were then optically characterized using a microplate reader, with a wide detection range of 230-780 nm (covering the UV-Vis characteristic absorption region), and high-resolution spectral acquisition was performed at 2 nm intervals. During the scanning process, absorbance data for each sample well was calculated using the absorbance calculation formula for 276 wavelength nodes, constructing a two-dimensional absorbance matrix (number of samples × number of wavelengths). The absorbance calculation formula is as follows: in, The summation symbol indicates a linear superposition of the five terms (i is the index of the component, from 1 to 5), and A(λ) is the absorbance at wavelength λ, ∈ i (λ) is the molar absorptivity of the i-th heavy metal, c i Let be the concentration of the i-th metal, and l be the optical path length.
4. The digital fingerprint spectral-driven intelligent detection method for multi-component heavy metals according to claim 1, characterized in that, In step (3), the obtained absorbance data is converted into spectral data, resulting in 2700 Gram difference field (GADF), 2700 Gram sum field (GASF), and 2700 wavelet transform (CWT) maps, ultimately forming three types of spectral data. The formulas for Gram difference field (GADF), Gram sum field (GASF), and wavelet transform (CWT) are as follows: Among them, GADF i,j and GASF i,j The element in the i-th row and j-th column of the matrix represents the collaborative relationship between time points i and j. and The normalized time series elements (the values at the i and j-th time points, ranging from [0,1]); Among them, W f (a, b) are wavelet transform coefficients, a is the scaling parameter, b is the translation parameter, f(t) is the continuous signal to be analyzed, φ is the conjugate function, and t is the integration variable.
5. The digital fingerprint spectral-driven intelligent detection method for multi-component heavy metals according to claim 1, characterized in that, In step (4), the original spectral dataset and concentration data are first augmented by adding Gaussian noise and combining it with spectral mixing techniques. Then, the absorbance data is standardized, and the natural logarithmic transformation is used to improve the skewed distribution of the data. The formulas for Gaussian noise, spectral mixing, absorbance data standardization, and natural logarithmic transformation are as follows: A noisy =A orig +N(b,a 2 ) (8) Among them, A noisy For the data after adding noise, A orig The original spectral data are given, and N is Gaussian noise with mean b and standard deviation a. A mixed =αA1+(1-α)A2(α~(0.2-0.8)) (9) Among them, A mixed The mixed spectral data are represented by A1 and A2, which are two randomly selected spectral samples, and α is the mixing ratio (between 0.2 and 0.8). Among them, X scaled For the standardized data, X is the original input data (features or labels), μ is the mean absorbance, and σ is the standard deviation; y log =ln(1+y) (11) Where y is the original concentration value, y log This is the value after logarithmic transformation.
6. The digital fingerprint spectral-driven intelligent detection method for multi-component heavy metals according to claim 1, characterized in that, The deep learning model is a CNN-CBAM model, which includes: an input layer that receives a two-dimensional digital fingerprint map, which is then processed sequentially through a 3×3 convolutional layer, a batch normalization layer, and a ReLU activation layer to extract local features; Insert a CBAM attention module, which first calculates the inter-channel dependencies through channel attention, and then generates spatial position weights through spatial attention; After compressing the feature dimensions through a global average pooling layer, the concentration values of 5-10 heavy metals are output through a fully connected layer. In terms of training strategy, stratified random sampling is used to divide the original dataset into training set, validation set and test set in an 8:1:1 ratio; The optimization process uses the Adam optimizer (learning rate 1×10⁻⁶). -4 The model uses mean absolute error as the loss function and improves generalization ability through a triple mechanism: dynamic learning rate decay, early stopping mechanism, and Dropout regularization (ratio 0.1). The training batch size is set to 16, and the maximum number of iterations is limited to 200. Finally, the trained CNN-CBAM model can directly predict the concentration of each metal in the solution based on the absorbance of the metal mixture solution.
7. The digital fingerprint spectral-driven intelligent detection method for multi-component heavy metals according to claim 1, characterized in that, In step (6), in the metal concentration prediction task, the study used R... 2 The performance of the model is evaluated using the MAE and RMSE metrics, with preset thresholds for the evaluation metrics: R 2 ≥0.9, RMSE≤1.5, MAE≤1.0, when the model linearly fits the scatter plot of the predicted and actual values on the test set with R0.
9. 2 The model is deemed qualified if the above thresholds are met.
8. A multi-component heavy metal intelligent detection system, used to implement the digital fingerprint spectral-driven multi-component heavy metal intelligent detection method as described in any one of claims 1-7, characterized in that, include: Spectral acquisition module: includes a UV-VIS spectrometer and a 96-well plate. The spectrometer is used to acquire the UV-Vis absorption curve of the heavy metal mixed sample after the addition of combined chemical probes. The 96-well plate is used to hold the sample and carry out the colorimetric reaction. Data processing module: used to convert one-dimensional absorption curves into two-dimensional digital fingerprint maps and perform data augmentation and standardization processing; Model inference module: Embedded with a trained CNN-CBAM deep learning model, it receives digital fingerprint data and outputs the concentration of multiple heavy metal components; Visualization and Interaction Module: A graphical user interface built with PyQt5 supports spectral data import, fingerprint spectrum generation, concentration prediction result display, and export.
9. The multi-component heavy metal intelligent detection system according to claim 8, characterized in that, The system is compatible with portable devices, supports offline inference and local data storage, and can be applied to the simultaneous detection of multiple heavy metals (antimony, iron, nickel, cadmium and copper) at sewage treatment plants, river monitoring stations and industrial pollution sources, with a detection cycle of ≤10 minutes.