Water quality heavy metal pollution detection method and system based on Raman spectrum

Through the detection method of water quality heavy metal pollution based on Raman spectroscopy, wavelet transformation, principal component analysis and partial least squares regression algorithm are used to solve the problem that the heavy metal content in water sources cannot be monitored in real time in the existing technology, and fast and accurate detection and detailed pollution assessment are achieved.

CN119985429AInactive Publication Date: 2025-05-13CSSC HAISHEN MEDICAL TECH CO LTD

Patent Information

Application Number
CN202411871017.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-18
Publication Date
2025-05-13
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The prior art cannot meet the need to monitor the heavy metal content in water sources in real time, and there are problems such as complex operation, time-consuming and costly.

Method used

The water quality heavy metal pollution detection method based on Raman spectroscopy is adopted to quickly scan water samples through a high-sensitivity Raman spectrometer, and combined with wavelet transformation algorithm, principal component analysis technology and partial least squares regression algorithm, denoising processing, feature extraction and concentration prediction are performed.

Benefits of technology

It realizes rapid and accurate detection of heavy metal concentration in water sources, improves detection speed and data quality, reduces operating costs, and provides detailed water quality heavy metal pollution detection reports.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119985429A_ABST
    Figure CN119985429A_ABST
Patent Text Reader

Abstract

The invention provides a water quality heavy metal pollution detection method and system based on Raman spectrum. The method comprises the following steps: collecting a water sample through a water area to be detected, and scanning the water sample by adopting a Raman spectrometer to generate an original Raman spectrogram; based on the original Raman spectrogram, using a wavelet transform algorithm to perform de-noising processing, reducing background interference, using a principal component analysis technology to extract characteristic peaks, and generating a heavy metal preliminary prediction result; based on the preliminary heavy metal prediction result, constructing a heavy metal concentration prediction model by applying a partial least squares regression algorithm, performing quantitative analysis, evaluating prediction accuracy by adopting a cross validation technology, and generating a heavy metal actual concentration result; and based on the actual heavy metal concentration result, comparing with a water quality safety standard, evaluating a pollution condition, and generating a water quality heavy metal pollution detection report. The technical scheme provided by the invention has high sensitivity, can efficiently denoise and accurately predict, and provides a scientific means for monitoring and treating heavy metal pollution of water.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the technical field of water quality detection, and in particular to a method and system for detecting heavy metal pollution in water based on Raman spectroscopy. Background Art

[0002] With the acceleration of industrialization and the continuous advancement of urbanization, water pollution is becoming increasingly serious. Heavy metal pollutants such as lead and mercury can enter the human body through drinking water or the food chain, causing chronic poisoning or even acute poisoning. Therefore, it is necessary to quickly and accurately detect the content of heavy metals in water sources. Many industrial activities will produce wastewater containing heavy metals. Direct discharge of these wastewaters into the environment without treatment will cause serious ecological damage. Therefore, an effective monitoring mechanism needs to be established to ensure that the discharge meets the standards. Natural disasters or man-made accidents may cause sudden water pollution incidents. In this case, the ability to respond quickly and assess the degree of pollution is crucial.

[0003] There are many methods for detecting heavy metals in water on the market, including atomic absorption spectroscopy, inductively coupled plasma mass spectrometry, and traditional chemical analysis methods. These technologies have their own characteristics, but they generally face problems such as complex operation, long time consumption, and high cost.

[0004] Traditional chemical analysis methods usually take a long time to complete sample pretreatment steps and cannot meet the needs of real-time monitoring; certain heavy metals at low concentration levels are difficult to be effectively detected by existing technologies, limiting their scope of application; advanced instruments and equipment not only have high purchase costs, but also require regular calibration and maintenance during daily use, increasing overall operating costs; some technologies have specific requirements for sample status (such as pH value, temperature), which increases the difficulty of actual operation. Summary of the invention

[0005] The embodiments of the present application provide a method and system for detecting heavy metal pollution in water based on Raman spectroscopy, so as to solve the problem that the existing technology cannot meet the requirements of real-time monitoring.

[0006] In a first aspect, the present application provides a method for detecting heavy metal pollution in water based on Raman spectroscopy, comprising:

[0007] Collect water samples from the water area to be tested, and quickly scan the water samples using a high-sensitivity Raman spectrometer to generate an original Raman spectrum;

[0008] Based on the original Raman spectrum, a wavelet transform algorithm is used to perform denoising, reduce background interference to improve spectrum quality, and a principal component analysis technique is used to extract characteristic peaks related to heavy metal ions to generate preliminary prediction results of heavy metals;

[0009] Based on the preliminary prediction results of heavy metals, a heavy metal concentration prediction model is constructed using a partial least squares regression algorithm, the actual concentration of heavy metals in the water samples is quantitatively analyzed, and cross-validation technology is used to evaluate the model stability and prediction accuracy to generate actual heavy metal concentration results;

[0010] Based on the actual heavy metal concentration results, a comparative analysis is conducted with national or regional water quality safety standards, a comprehensive assessment of the heavy metal pollution in water bodies is conducted, and a water quality heavy metal pollution detection report is generated.

[0011] Optionally, based on the original Raman spectrum, a wavelet transform algorithm is used to perform denoising processing to reduce background interference to improve spectrum quality, and a principal component analysis technique is used to extract characteristic peaks related to heavy metal ions to generate a preliminary prediction result of heavy metals, including:

[0012] Based on the original Raman spectrum, normalization processing is performed to eliminate intensity changes caused by concentration differences of different samples, ensure that subsequent analysis is consistent and accurate, and generate normalized spectrum data;

[0013] Based on the normalized spectral data, a wavelet transform algorithm is performed, a suitable mother wavelet function is selected, the number of decomposition layers is set, and the signal noise component and the useful signal are separated by multi-scale analysis to reduce background interference and generate a denoised Raman spectrum;

[0014] Based on the denoised Raman spectrum, principal component analysis technology is used to perform dimensionality reduction processing, extract the principal components representing the characteristics of heavy metal ions, obtain characteristic peaks related to heavy metal ions, and generate characteristic peak data;

[0015] Based on the characteristic peak data, combined with the known heavy metal ion standard spectrum database, a comparative analysis is performed to preliminarily determine the possible types and relative contents of heavy metal ions in the water sample and generate a preliminary prediction result of heavy metals.

[0016] Optionally, the normalized spectral data is processed by wavelet transform algorithm, a suitable mother wavelet function is selected, the number of decomposition layers is set, and the signal noise component and the useful signal are separated by multi-scale analysis to reduce background interference and generate a denoised Raman spectrum, including:

[0017] Based on the normalized spectral data, preprocessing is performed to determine a data baseline and a range, thereby generating preprocessed spectral data;

[0018] Based on the pre-processed spectral data, a mother wavelet function matching the spectral data characteristics is selected, an appropriate number of decomposition layers is set, and a wavelet transform algorithm is used to perform wavelet transform processing to generate wavelet transform coefficients;

[0019] Based on the wavelet transform coefficients, the useful signal components in the spectral data are identified and enhanced by multi-scale analysis technology, the noise components are removed, the background interference is reduced, and the optimized wavelet transform coefficients are generated;

[0020] Based on the optimized wavelet transform coefficients, an inverse wavelet transform process is performed to restore the time domain signal and generate a denoised Raman spectrum.

[0021] Optionally, based on the pre-processed spectral data, a mother wavelet function matching the spectral data characteristics is selected, a suitable number of decomposition layers is set, and a wavelet transform algorithm is used to perform wavelet transform processing to generate wavelet transform coefficients, including:

[0022] Based on the pre-processed spectral data, a Gaussian filtering method is used to reduce noise, and a background signal is removed by polynomial fitting to ensure the effectiveness of nonlinear conversion, so as to generate a nonlinear conversion output;

[0023] The nonlinear conversion output is calculated using the following formula:

[0024]

[0025] Among them, Y(t) is the nonlinear conversion output; X(t) is the original spectral data; γ is the conversion rate, which determines the steepness of the nonlinear conversion; θ is the conversion threshold, which determines the center point of the nonlinear conversion; λ1 is the first amplitude, which controls the intensity of the first periodic change; f1 is the first frequency, which defines the speed of the first periodic change; is the first initial phase, which defines the starting position of the first periodic change; λ2 is the second amplitude, which controls the intensity of the second periodic change; f2 is the second frequency, which defines the speed of the second periodic change; is the second initial phase, defining the starting position of the second periodic change; t is the time variable, indicating the time point of the spectral data;

[0026] Based on the nonlinear conversion output, a Gaussian kernel function is introduced to perform local weighting to highlight local features, and sine and cosine terms are introduced to capture periodic components in spectral data to generate spectral features;

[0027] The extracted spectral features are calculated using the following formula:

[0028]

[0029] Z(k) is the extracted spectral feature; Y(t) is the nonlinear conversion output; Y(iΔt) is the value of the nonlinear conversion output at time point iΔt; i is the index of the data point, from 1 to N; N is the total number of data points; Δt is the sampling interval; H(x) is the Gaussian kernel function, used for local weighting; τ is the time constant, which determines the influence range of the Gaussian kernel function; λ3 is the influence factor, which determines the rate of exponential decay; λ4 is the third amplitude, which controls the intensity of the third periodic change; f3 is the third frequency, which defines the speed of the third periodic change; is the third initial phase, defining the starting position of the third periodic change; λ5 is the fourth amplitude, controlling the intensity of the fourth periodic change; f4 is the fourth frequency, defining the speed of the fourth periodic change; is the fourth initial phase, defining the starting position of the fourth periodic change; k is a discrete position, indicating the position of feature extraction;

[0030] Based on the extracted spectral features, combined with the spectral data features and frequency domain characteristics, a suitable mother wavelet function is selected, the number of wavelet transform decomposition layers is determined through experimental results, and wavelet transform coefficients are generated.

[0031] Optionally, the denoised Raman spectrum is subjected to dimensionality reduction processing using principal component analysis technology to extract principal components representing the characteristics of heavy metal ions, obtain characteristic peaks related to heavy metal ions, and generate characteristic peak data, including:

[0032] Based on the denoised Raman spectrum, each spectrum sample is sorted, and a spectrum data matrix is ​​constructed to integrate each spectrum sample and wave number point to generate a spectrum data matrix;

[0033] Based on the spectral data matrix, statistical analysis is performed on the spectral information of all samples to reveal the correlation of spectral signals between different wave number points and generate a covariance matrix;

[0034] Based on the covariance matrix, the eigenvalues ​​and eigenvectors are solved by mathematical methods, and the eigenvectors with the largest eigenvalues ​​are selected as principal components, so as to retain the main variation information of the original Raman spectrum to the maximum extent and generate a principal component set;

[0035] Based on the principal component set, the high-dimensional spectral data is projected into the low-dimensional space constituted by the principal components, the principal components reflecting the characteristics of heavy metal ions are extracted, the positions and intensities of characteristic peaks related to heavy metal ions are analyzed and determined, and characteristic peak data are generated.

[0036] Optionally, based on the preliminary prediction results of heavy metals, a partial least squares regression algorithm is used to construct a heavy metal concentration prediction model, quantitatively analyze the actual concentration of heavy metals in the water sample, and use cross-validation technology to evaluate the model stability and prediction accuracy to generate the actual concentration results of heavy metals, including:

[0037] Based on the preliminary prediction results of heavy metals, Raman spectral data of standard heavy metal solutions with known concentrations are collected to generate a training data set;

[0038] Based on the training data set, a partial least squares regression algorithm is used to maximize the covariance between the input variables and the output variables to find the optimal linear combination and generate a heavy metal concentration prediction model;

[0039] Based on the heavy metal concentration prediction model, a cross-validation technique is used to evaluate the model, and the divided subsets are used as independent test sets in turn, and the test process is repeated to summarize the test results and generate a model evaluation result;

[0040] Based on the model evaluation results, the actual concentration of heavy metals in the water sample is quantitatively analyzed to generate actual concentration results of heavy metals.

[0041] Optionally, based on the training data set, a partial least squares regression algorithm is used to maximize the covariance between the input variables and the output variables to find the optimal linear combination and generate a heavy metal concentration prediction model, including:

[0042] Based on the training data set, standardization processing is performed to eliminate the differences in dimensions and numerical ranges between different spectral data, and each variable in the model training process is fairly compared to generate a preprocessed training data set;

[0043] Based on the preprocessed training data set, a partial least squares regression algorithm is used to select spectral data and known heavy metal concentration values ​​as input and output variables, and the optimal linear combination is found by maximizing the covariance between the input variables and the output variables to generate a preliminary linear combination model;

[0044] Based on the preliminary linear combination model, the principal components of the input variables are gradually extracted to control the cumulative explained variance to reach a predetermined threshold value, thereby generating an optimized principal component set;

[0045] Based on the optimized principal component set, potential multicollinearity problems among multivariate data are handled to generate a heavy metal concentration prediction model.

[0046] Optionally, based on the preprocessed training data set, a partial least squares regression algorithm is used to select spectral data and known heavy metal concentration values ​​as input and output variables, and the optimal linear combination is found by maximizing the covariance between the input variables and the output variables to generate a preliminary linear combination model, including:

[0047] Based on the preprocessed training data set, using a box plot to detect and remove outliers;

[0048] Applying a smoothing filter to further reduce noise and improve the data signal-to-noise ratio to generate nonlinear correlations;

[0049] The nonlinear correlation is calculated using the following formula:

[0050]

[0051] Among them, R xy is the nonlinear correlation between the input variable and the output variable; i is the sample index of the training data set, from 1 to N; N is the number of samples in the training data set; X i is the spectral data of the i-th sample; is the average value of all sample spectral data; σ X is the standard deviation of all sample spectral data; Y i is the known heavy metal concentration value of the i-th sample; is the average value of known heavy metal concentrations of all samples; σ Y is the standard deviation of the known heavy metal concentration values ​​of all samples; β is the nonlinear adjustment parameter used to control the sensitivity of nonlinear correlation;

[0052] Based on the nonlinear correlation, an initialization method based on prior knowledge is used to obtain a reasonable initial weight vector, a gradient descent method is used to solve the maximization objective function, and the weight vector is adjusted through multiple iterations to achieve convergence conditions to generate a weight vector of the optimal linear combination;

[0053] The weight vector of the optimal linear combination is calculated using the following formula:

[0054]

[0055] Where W is the weight vector of the optimal linear combination; w is the candidate weight vector; R xy is the nonlinear correlation matrix between input variables and output variables; α is the balance parameter used to control the influence of the diagonal matrix and the unit matrix; S x is the covariance matrix of the input variables; diag(w T S x w) is a diagonal matrix, and the diagonal elements are w T S x The elements of w; I is the identity matrix; γ is an additional adjustment parameter used to control the influence of the covariance matrix of the output variable; S y is the covariance matrix of the output variables; diag(w T S y w) is a diagonal matrix with diagonal elements s T S y The elements of w; δ is another adjustment parameter used to control the influence of the cross covariance matrix between the input variables and the output variables; S xy is the cross covariance matrix of the input variables and the output variables; is a diagonal matrix with diagonal elements elements; λ is a nonlinear adjustment parameter used to control the nonlinear influence of the cross-covariance matrix;

[0056] Based on the weight vector of the optimal linear combination, feature selection is performed by recursive feature elimination method, and a regularization term is introduced to control the complexity of the model to prevent overfitting. The model is fitted with the known heavy metal concentration values ​​to generate a preliminary linear combination model.

[0057] Optionally, the actual heavy metal concentration result is compared and analyzed with national or regional water quality safety standards to comprehensively evaluate the heavy metal pollution in the water body and generate a water quality heavy metal pollution detection report, including:

[0058] Based on the actual heavy metal concentration results, collect and sort out the current national or regional water quality safety standards, obtain the heavy metal concentration limit values, and generate a water quality safety standard comparison table;

[0059] Based on the water quality safety standard comparison table, the actual heavy metal concentration results are compared one by one, the exceeded items and the exceeded multiples are marked, and a heavy metal concentration exceeded situation table is generated;

[0060] Based on the table of excessive heavy metal concentrations, comprehensively assess the heavy metal pollution in water bodies, conduct hazard analysis, and generate a conclusion on the assessment of heavy metal pollution in water bodies;

[0061] Based on the conclusions of the heavy metal pollution assessment of the water body, combined with the specific sampling time and location, a water quality treatment plan and monitoring plan are formulated to generate a water quality heavy metal pollution detection report.

[0062] In a second aspect, the present application provides a water quality heavy metal pollution detection system based on Raman spectroscopy, comprising:

[0063] A collection module is used to collect water samples from the water area to be tested, and quickly scan the water samples using a high-sensitivity Raman spectrometer to generate an original Raman spectrum;

[0064] A processing module is used to perform denoising based on the original Raman spectrum using a wavelet transform algorithm to reduce background interference to improve spectrum quality, extract characteristic peaks related to heavy metal ions using a principal component analysis technique, and generate a preliminary prediction result of heavy metals;

[0065] A construction module is used to construct a heavy metal concentration prediction model based on the preliminary prediction results of heavy metals by using a partial least squares regression algorithm, quantitatively analyze the actual concentration of heavy metals in the water sample, evaluate the model stability and prediction accuracy by using a cross-validation technique, and generate the actual concentration results of heavy metals;

[0066] The analysis module is used to compare and analyze the actual heavy metal concentration results with national or regional water quality safety standards, comprehensively evaluate the heavy metal pollution in the water body, and generate a water quality heavy metal pollution detection report.

[0067] In an embodiment of the present application, water samples are collected from the water area to be tested, and a high-sensitivity Raman spectrometer is used to quickly scan the water samples to generate an original Raman spectrum; based on the original Raman spectrum, a wavelet transform algorithm is used to perform denoising processing to reduce background interference to improve the spectral quality, and a principal component analysis technique is used to extract characteristic peaks related to heavy metal ions to generate a preliminary prediction result of heavy metals; based on the preliminary prediction result of heavy metals, a partial least squares regression algorithm is used to construct a heavy metal concentration prediction model, and the actual concentration of heavy metals in the water sample is quantitatively analyzed, and the cross-validation technique is used to evaluate the model stability and prediction accuracy, and the actual concentration result of heavy metals is generated; based on the actual concentration result of heavy metals, a comparative analysis is performed with national or regional water quality safety standards, the heavy metal pollution situation of the water body is comprehensively evaluated, and a water quality heavy metal pollution detection report is generated.

[0068] The technical solution of this application has the following beneficial effects:

[0069] By using a high-sensitivity Raman spectrometer, this method can quickly and accurately scan water samples and generate original Raman spectra, which improves the detection speed, ensures high resolution and high sensitivity of the data, and can obtain a large amount of spectral information in a short time; the wavelet transform algorithm is used for denoising, which effectively reduces background interference and improves the spectral quality. Combined with the principal component analysis technology, it can extract the characteristic peaks related to heavy metal ions, thereby enhancing the identifiability and accuracy of the signal; the partial least squares regression algorithm is used to construct a heavy metal concentration prediction model, and the actual concentration of heavy metals in water samples is quantitatively analyzed. It can effectively process multivariate data, solve the problem of multicollinearity, and evaluate the stability and prediction of the model through cross-validation technology. The accuracy of the results ensures the reliability and robustness of the results; based on the actual heavy metal concentration results, a comparative analysis is conducted with national or regional water quality safety standards to comprehensively evaluate the heavy metal pollution in water bodies and generate a detailed water quality heavy metal pollution detection report to provide a scientific basis for relevant departments and support decision-making and the implementation of pollution control measures; this method combines advanced spectral technology and data analysis methods to achieve full process automation and intelligence from sample collection to result generation, reducing human errors and improving detection efficiency and accuracy; through regular testing and real-time monitoring, the changing trend of heavy metal pollution in water quality can be discovered in a timely manner, providing important data support for environmental monitoring and early warning systems, and helping to prevent and control the occurrence of water pollution incidents.

[0070] Furthermore, through normalization processing, the problem of spectral intensity variation caused by differences in sample concentrations can be eliminated, ensuring the consistency and accuracy of subsequent analysis, and helping to obtain more reliable results under different conditions; denoising using wavelet transform algorithm can effectively separate background noise and useful signals, significantly improving the quality of Raman spectra and making subsequent feature extraction more accurate; the application of principal component analysis technology not only reduces data dimensions, but also helps identify characteristic peaks related to heavy metal ions, thereby simplifying the process of extracting key information from complex spectra and speeding up detection; combined with known heavy metal ion standard spectral databases for comparative analysis, the types of heavy metals in water samples and their relative contents can be more accurately determined, enhancing the sensitivity and specificity of the method.

[0071] Furthermore, based on the training data set constructed based on the preliminary prediction results, the partial least squares regression algorithm can effectively find the optimal linear relationship between the input variables (such as characteristic peak data) and the output variables (heavy metal concentration), thereby improving the accuracy of the model in predicting the heavy metal concentration of actual samples; cross-validation technology is used to evaluate the model performance, and through multiple iterative tests and summary of the results, the stability and generalization ability of the model can be comprehensively tested, reducing the risk of overfitting; through the effective use of existing data sets and efficient mathematical modeling methods, the scheme can achieve rapid quantitative analysis of unknown samples without increasing additional experimental costs, thereby improving overall work efficiency; once the established heavy metal concentration prediction model has been fully verified, it can be applied to on-site or online monitoring systems to provide timely and effective monitoring means for water quality safety, which is conducive to rapid response to potential pollution incidents.

[0072] These and other aspects of the present application will become more clearly understood in the description of the following embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0073] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, a brief introduction will be given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0074] Figure 1 A flow chart of a method for detecting heavy metal pollution in water based on Raman spectroscopy provided in an embodiment of the present application;

[0075] Figure 2 A schematic diagram of the structure of a water quality heavy metal pollution detection system based on Raman spectroscopy provided in an embodiment of the present application;

[0076] Figure 3A schematic diagram of the structure of a computing device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0077] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application.

[0078] In some of the processes described in the specification and claims of this application and the above-mentioned figures, multiple operations that appear in a specific order are included, but it should be clearly understood that these operations may not be executed in the order in which they appear in this article or executed in parallel. The serial numbers of the operations, such as 101, 102, etc., are only used to distinguish between different operations, and the serial numbers themselves do not represent any execution order. In addition, these processes may include more or fewer operations, and these operations may be executed in sequence or in parallel. It should be noted that the descriptions of "first", "second", etc. in this article are used to distinguish different messages, devices, modules, etc., do not represent the order of precedence, and do not limit the "first" and "second" to be different types.

[0079] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of this application.

[0080] Figure 1 A flowchart of a method for detecting heavy metal pollution in water based on Raman spectroscopy is provided for the embodiment of the present application. Figure 1 As shown, the method includes:

[0081] 101. Collect water samples from the water area to be tested, and quickly scan the water samples using a high-sensitivity Raman spectrometer to generate an original Raman spectrum;

[0082] The water area to be tested refers to a specific area where water quality testing is required, such as a river, lake or ocean.

[0083] A water sample is a certain amount of water sample collected from the water area to be tested for subsequent analysis.

[0084] A high-sensitivity Raman spectrometer is an instrument that can measure the intensity distribution of scattered light caused by changes in the vibration and rotation states of a substance's molecules. It has high sensitivity and can detect very weak signals, making it suitable for the analysis of low-concentration substances.

[0085] The original Raman spectrum is a data image obtained after scanning the water sample with a Raman spectrometer. It shows the intensity of scattered light at different wavelengths and reflects the information of various chemical components that may exist in the water.

[0086] In the embodiment of the present application, first, determine the specific location of the water area to be tested and prepare the corresponding sampling tools; second, collect a sufficient amount of water samples at the selected location to ensure the representativeness and uniformity of the samples; third, send the collected water samples to a high-sensitivity Raman spectrometer set up in the laboratory or on-site, and adjust the instrument parameters to meet the current test requirements; finally, start the Raman spectrometer to quickly scan the water sample, record the data generated during the entire process, and form an original Raman spectrum.

[0087] Suppose you need to assess the heavy metal pollution in a river;

[0088] Firstly, a representative section of the river was selected as the water area to be tested, and a 1-liter water sample was collected from about 30 cm below the water surface using a professional sampling bottle; secondly, the 1-liter water sample was brought back to a laboratory equipped with a high-sensitivity Raman spectrometer, where the working parameters of the Raman spectrometer were pre-set, including laser power, integration time, and scanning range, in order to optimize the signal detection capability within the wavelength range sensitive to heavy metal ions; thirdly, a portion of the water sample was placed in a special sample pool and placed on the detection platform of the spectrometer, and the equipment was started to start the automatic scanning program. After about a few minutes, a detailed original Raman spectrum was obtained, which could show the characteristic information of all identifiable compounds (including potential heavy metals) in the water sample; finally, through further processing and analysis of this spectrum, it can be preliminarily determined whether there is heavy metal pollution in this river section and its type.

[0089] Through the above steps, high-sensitivity Raman spectroscopy technology can be used to quickly screen pollutants in environmental water bodies, providing a scientific basis and technical support for environmental protection work.

[0090] 102. Based on the original Raman spectrum, a wavelet transform algorithm is used to perform denoising, reduce background interference to improve spectrum quality, and a principal component analysis technique is used to extract characteristic peaks related to heavy metal ions to generate a preliminary prediction result of heavy metals;

[0091] The wavelet transform algorithm is a signal processing method used to analyze the changes of different frequency components over time. It is particularly suitable for denoising non-stationary signals.

[0092] Principal component analysis is a dimensionality reduction method in statistics. It converts the original data into a set of new variables (principal components) that are linear combinations of the original variables and are sorted by variance to extract the most important information.

[0093] Characteristic peaks are peaks that appear at certain specific wavelengths in the Raman spectrum, and these peaks are associated with specific chemical components in the sample.

[0094] The preliminary prediction results of heavy metals are estimates of the types of heavy metals that may exist in water samples and their approximate contents based on the denoised Raman spectral data and the characteristic information extracted by principal component analysis technology.

[0095] In this step, first, based on the original Raman spectrum obtained, the wavelet transform algorithm is used to denoise the data to remove background noise and other interference factors; secondly, the principal component analysis technology is used to identify and separate the significant characteristic peaks related to heavy metal ions using the cleaned data; thirdly, through the analysis of these characteristic peaks, it is preliminarily determined which heavy metal elements exist in the water sample and their relative intensities; finally, based on the above analysis results, a prediction report on the possibility of heavy metal presence is generated to provide a basis for further quantitative analysis.

[0096] Optionally, the method in step 102 is based on the original Raman spectrum, performs denoising processing using a wavelet transform algorithm, reduces background interference to improve spectral quality, uses principal component analysis technology to extract characteristic peaks related to heavy metal ions, and generates a preliminary prediction result of heavy metals, including: based on the original Raman spectrum, performs normalization processing to eliminate intensity changes caused by concentration differences in different samples, ensures that subsequent analysis is consistent and accurate, and generates normalized spectrum data; based on the normalized spectrum data, performs wavelet transform algorithm processing, selects a suitable mother wavelet function, sets the number of decomposition layers, separates signal noise components from useful signals through multi-scale analysis, reduces background interference, and generates a denoised Raman spectrum; based on the denoised Raman spectrum, performs dimensionality reduction processing using principal component analysis technology, extracts principal components representing the characteristics of heavy metal ions, obtains characteristic peaks related to heavy metal ions, and generates characteristic peak data; based on the characteristic peak data, performs comparative analysis in combination with a known heavy metal ion standard spectrum database, preliminarily determines the possible types and relative contents of heavy metal ions in the water sample, and generates a preliminary prediction result of heavy metals.

[0097] Among them, based on the normalized spectral data, wavelet transform algorithm processing is performed, a suitable mother wavelet function is selected, the number of decomposition layers is set, and the signal noise component and the useful signal are separated by multi-scale analysis, background interference is reduced, and a denoised Raman spectrum is generated, including: based on the normalized spectral data, preprocessing is performed to determine the data baseline and range, and preprocessed spectral data is generated; based on the preprocessed spectral data, a mother wavelet function matching the spectral data characteristics is selected, a suitable number of decomposition layers is set, and a wavelet transform algorithm is used to perform wavelet transform processing to generate wavelet transform coefficients; based on the wavelet transform coefficients, the useful signal components in the spectral data are identified and enhanced by multi-scale analysis technology, the noise components are removed, the background interference is reduced, and optimized wavelet transform coefficients are generated; based on the optimized wavelet transform coefficients, inverse wavelet transform processing is performed to restore the time domain signal to generate a denoised Raman spectrum.

[0098] Normalization is a data preprocessing method that converts data to the same magnitude or range to eliminate intensity variations caused by concentration differences in different samples, ensuring consistency and accuracy of subsequent analysis.

[0099] Wavelet transform is a signal processing technology used to decompose signals into different frequency components and effectively separate noise and useful signals. It is particularly suitable for denoising non-stationary signals.

[0100] The mother wavelet function is the basis function used in wavelet transform. Different mother wavelet functions are suitable for different types of data features.

[0101] The number of decomposition levels is a parameter in wavelet transform, which indicates the number of levels at which the signal is decomposed and affects the resolution and denoising effect of the signal.

[0102] Multiscale analysis uses wavelet transform to analyze signals at different scales, identify and enhance useful signal components, and remove noise components.

[0103] The principal component analysis technique is a statistical method used to identify major patterns in data and reduce dimensionality by transforming the original variables into a set of new variables (i.e., principal components) that are linear combinations of the original variables and are ordered by variance to capture the maximum variability in the data.

[0104] Characteristic peaks are peaks in the Raman spectrum that represent the presence of specific chemical substances. These peaks are associated with the molecular structure present in the sample and can be used to identify the type of substance and its concentration.

[0105] The preliminary prediction results of heavy metals are based on the denoised Raman spectral data and the characteristic information extracted after principal component analysis, which can estimate the types of heavy metals that may exist in water samples and their approximate contents.

[0106] In the embodiment of the present application, first, normalization processing is performed based on the original Raman spectrum to eliminate the intensity changes caused by concentration differences in different samples, and normalized spectrum data is generated; secondly, the normalized spectrum data is preprocessed to determine the data baseline and range, and preprocessed spectrum data is generated; thirdly, a mother wavelet function matching the characteristics of the spectrum data is selected, an appropriate number of decomposition layers is set, and a wavelet transform algorithm is used to perform wavelet transform processing to generate wavelet transform coefficients; then, a multi-scale analysis technique is used to identify and enhance useful signal components in the spectrum data, remove noise components, reduce background interference, and generate optimized wavelet transform coefficients; based on the optimized wavelet transform coefficients, an inverse wavelet transform is performed to restore the time domain signal to generate a denoised Raman spectrum; then, the principal component analysis technique is used to perform dimensionality reduction processing on the denoised Raman spectrum, extract the principal components representing the characteristics of heavy metal ions, obtain characteristic peaks related to heavy metal ions, and generate characteristic peak data; finally, combined with a known heavy metal ion standard spectrum database, a comparative analysis is performed to preliminarily determine the possible types and relative contents of heavy metal ions in the water sample, and generate a preliminary prediction result of heavy metals.

[0107] Suppose you need to test the water quality at the outlet of an industrial wastewater treatment plant;

[0108] First, collect water samples from the discharge port and use a high-sensitivity Raman spectrometer to quickly scan and generate original Raman spectra; second, normalize the original Raman spectra to eliminate the intensity changes caused by the concentration differences of different samples, ensure the consistency and accuracy of subsequent analysis, and generate normalized spectral data; third, based on the normalized spectral data, preprocessing is performed to determine the data baseline and range to generate preprocessed spectral data; then, select a mother wavelet function (such as Daubechies wavelet) that matches the characteristics of the spectral data, set an appropriate number of decomposition layers (for example, 5 layers), use a wavelet transform algorithm to perform wavelet transform processing, and generate wavelet transform coefficients; identify and enhance useful signal components in spectral data through multi-scale analysis technology , remove noise components, reduce background interference, and generate optimized wavelet transform coefficients; then, based on the optimized wavelet transform coefficients, perform inverse wavelet transform processing to restore the time domain signal and generate a denoised Raman spectrum; next, use principal component analysis technology to perform dimensionality reduction processing, extract the principal components representing the characteristics of heavy metal ions, obtain the characteristic peaks related to heavy metal ions, and generate characteristic peak data; finally, combine with the known heavy metal ion standard spectral database for comparative analysis, preliminarily judge the types of heavy metal ions (such as lead, cadmium, etc.) that may exist in the water sample and their relative contents, and generate preliminary prediction results for heavy metals; through the above steps, it is possible to effectively detect and evaluate the heavy metal pollution of water quality at the discharge outlet of industrial wastewater treatment plants, and provide a scientific basis for subsequent governance.

[0109] The present application takes into account that, when processing Raman spectral data, in order to improve signal quality and extract characteristic peaks associated with heavy metal ions, it is necessary to reduce noise through Gaussian filtering method, remove background signals through polynomial fitting, use nonlinear transformation formula, introduce Gaussian kernel function for local weighting, and determine the number of wavelet transform decomposition layers to generate wavelet transform coefficients.

[0110] Optionally, based on the pre-processed spectral data, a mother wavelet function matching the spectral data characteristics is selected, a suitable number of decomposition layers is set, and a wavelet transform algorithm is used to perform wavelet transform processing to generate wavelet transform coefficients, including:

[0111] Based on the pre-processed spectral data, a Gaussian filtering method is used to reduce noise, and a background signal is removed by polynomial fitting to ensure the effectiveness of nonlinear conversion, so as to generate a nonlinear conversion output;

[0112] The nonlinear conversion output is calculated using the following formula:

[0113]

[0114] Among them, Y(t) is the nonlinear conversion output; X(t) is the original spectral data; γ is the conversion rate, which determines the steepness of the nonlinear conversion; θ is the conversion threshold, which determines the center point of the nonlinear conversion; λ1 is the first amplitude, which controls the intensity of the first periodic change; f1 is the first frequency, which defines the speed of the first periodic change; is the first initial phase, which defines the starting position of the first periodic change; λ2 is the second amplitude, which controls the intensity of the second periodic change; f2 is the second frequency, which defines the speed of the second periodic change; is the second initial phase, defining the starting position of the second periodic change; t is the time variable, indicating the time point of the spectral data;

[0115] Based on the nonlinear conversion output, a Gaussian kernel function is introduced to perform local weighting to highlight local features, and sine and cosine terms are introduced to capture periodic components in spectral data to generate spectral features;

[0116] The extracted spectral features are calculated using the following formula:

[0117]

[0118] Z(k) is the extracted spectral feature; Y(t) is the nonlinear conversion output; Y(iΔt) is the value of the nonlinear conversion output at time point iΔt; i is the index of the data point, from 1 to N; N is the total number of data points; Δt is the sampling interval; H(x) is the Gaussian kernel function, used for local weighting; τ is the time constant, which determines the influence range of the Gaussian kernel function; λ3 is the influence factor, which determines the rate of exponential decay; λ4 is the third amplitude, which controls the intensity of the third periodic change; f3 is the third frequency, which defines the speed of the third periodic change; is the third initial phase, defining the starting position of the third periodic change; λ5 is the fourth amplitude, controlling the intensity of the fourth periodic change; f4 is the fourth frequency, defining the speed of the fourth periodic change; is the fourth initial phase, defining the starting position of the fourth periodic change; k is a discrete position, indicating the position of feature extraction;

[0119] Based on the extracted spectral features, combined with the spectral data features and frequency domain characteristics, a suitable mother wavelet function is selected, the number of wavelet transform decomposition layers is determined through experimental results, and wavelet transform coefficients are generated.

[0120] This method aims to improve the quality of Raman spectral data and accurately identify heavy metal ions in water samples through mathematical processing and feature extraction technology, enhance signal discrimination through Sigmoid function, and capture key information in combination with periodic variation terms. Gaussian kernel function is used to emphasize local features, and sine and cosine terms are introduced to capture periodic components, generate high-quality spectral features, and provide a reliable data basis for subsequent analysis.

[0121] In the nonlinear transformation output, the Sigmoid function term The original data is transformed nonlinearly through the Sigmoid function to make it more distinguishable; γ controls the steepness of the transformation, and θ controls the center point of the transformation; the first periodic change term λ1·cos(2πf1t+φ1): used to capture the periodic characteristics in the spectral data, λ1 controls the amplitude, f1 controls the frequency, and φ1 controls the phase; the second periodic change term λ2·sin(2πf2t+φ2): further enhances the ability to capture periodic characteristics, λ2 controls the amplitude, f2 controls the frequency, and φ2 controls the phase;

[0122] Among them, γ is determined by experimental optimization, usually ranging from 0.1 to 10; θ: determined according to experimental data, usually taken as the average or median of the original data; λ1, λ2, f1, f2, φ1, φ2 are determined by frequency domain analysis and experimental optimization to capture the periodic components of specific frequencies;

[0123] In extracting spectral features, the local weighted term The Gaussian kernel function is used for local weighting to highlight local features. λ3 controls the exponential decay rate and τ controls the impact range. The third periodic variation term λ4·sin(2πf3(iΔt)+φ3): further captures the periodic components in the spectral data. λ4 controls the amplitude, f3 controls the frequency and φ3 controls the phase. The fourth periodic variation term λ5·cos(2πf4(iΔt)+φ4): increases the diversity of periodic components. λ5 controls the amplitude, f4 controls the frequency and φ4 controls the phase.

[0124] Among them, Y(iΔt) is calculated by the nonlinear transformation output formula; H(x) is the standard Gaussian kernel function, usually in the form of H(x) = exp(-x 2 / 2); τ is determined based on experimental data and is usually several times the sampling interval; λ3 is determined through experimental optimization and is usually in the range of 0.1 to 10; λ4, λ5, f3, f4, φ3, φ4 are determined through frequency domain analysis and experimental optimization to capture the periodic components of specific frequencies;

[0125] Suppose you need to process the Raman spectrum data of a water sample;

[0126] The original spectral data X(t) is known, the sampling interval Δt = 0.01, the total number of data points N = 1000; the parameters are assumed to be γ = 5; θ = 100; λ1 = 2; f1 = 5; φ1 = 0; λ2 = 1; f2 = 10; τ=0.1; λ3=1; λ4=0.5; f3=7; λ5=0.3;f4=15; Assume X(t) = 105;

[0127]

[0128] Assuming that the threshold is set to 10, since the calculation result of the extracted spectral features 12.3 is greater than the set threshold, it means that the spectral features at this time point are very significant. This may be due to the presence of strong local features and periodic components in the spectral data at this position, which are consistent with the Raman spectral features of specific heavy metal ions. Therefore, it can be preliminarily determined that the water sample contains the corresponding heavy metal ions. Through the above steps, the spectral features can be successfully extracted and the types of heavy metals in the water sample can be identified.

[0129] Optionally, based on the denoised Raman spectrum, principal component analysis technology is used to perform dimensionality reduction processing, principal components representing the characteristics of heavy metal ions are extracted, characteristic peaks related to heavy metal ions are obtained, and characteristic peak data are generated, including: based on the denoised Raman spectrum, each spectral sample is sorted, and a spectral data matrix is ​​constructed to integrate each spectral sample and wave number point to generate a spectral data matrix; based on the spectral data matrix, all sample spectral information is statistically analyzed to reveal the correlation of spectral signals between different wave number points to generate a covariance matrix; based on the covariance matrix, eigenvalues ​​and eigenvectors are solved by mathematical methods, and eigenvectors with maximized eigenvalues ​​are selected as principal components to retain the main variation information of the original Raman spectrum to the maximum extent and generate a principal component set; based on the principal component set, high-dimensional spectral data is projected into a low-dimensional space composed of principal components to extract principal components reflecting the characteristics of heavy metal ions, and the positions and intensities of characteristic peaks related to heavy metal ions are analyzed and determined to generate characteristic peak data.

[0130] The spectral data matrix consists of the Raman spectral data of all samples, with each sample corresponding to a row and each column representing the spectral intensity value at a specific wave number point.

[0131] The covariance matrix is ​​used to describe the degree of linear relationship between different variables (here, different wavenumber points), reflecting the correlation of spectral signals between different wavenumber points.

[0132] Eigenvalues ​​and eigenvectors are a set of values ​​(eigenvalues) and directions (eigenvectors) obtained by eigendecomposing the covariance matrix in principal component analysis, where the eigenvalue represents the amount of information in the direction represented by the corresponding eigenvector.

[0133] The principal component set is composed by selecting eigenvectors with the largest eigenvalues, which can retain the variation information in the original data to the greatest extent.

[0134] The characteristic peak data is the peak information related to heavy metal ions extracted from the principal components after dimensionality reduction, including position and intensity.

[0135] In the embodiments of the present application, firstly, based on the denoised Raman spectrum, each spectral sample is sorted and a spectral data matrix is ​​constructed; secondly, based on the constructed spectral data matrix, statistical analysis is performed to reveal the correlation between spectral signals at different wavenumber points, thereby generating a covariance matrix; thirdly, mathematical methods are used to solve the eigenvalues ​​and corresponding eigenvectors of the covariance matrix, and those eigenvectors with the largest eigenvalues ​​are selected as the main components; finally, the high-dimensional original spectral data is projected into a low-dimensional space composed of the selected main components, and the main components reflecting the characteristics of heavy metal ions are extracted therefrom, the positions and intensities of the characteristic peaks related to heavy metal ions are determined, and the characteristic peak data are generated.

[0136] Suppose you want to detect heavy metal pollution in a city’s drinking water source;

[0137] First, denoised Raman spectra of multiple water samples from drinking water sources were collected and organized into a spectral data matrix, where each row represents a water sample and each column represents the spectral intensity at a specific wavenumber point. Subsequently, a comprehensive statistical analysis was performed based on this spectral data matrix to calculate the correlation between spectral signals at different wavenumber points to form a covariance matrix. Next, a mathematical algorithm was used to solve the eigenvalues ​​of the covariance matrix and its corresponding eigenvectors, from which several eigenvectors with the highest eigenvalues ​​were selected as principal components to ensure that these principal components can retain the main change information of the original spectral data to the greatest extent. Finally, all the original spectral data were reprojected into a new low-dimensional space according to this set of selected principal components. During this process, we focused on identifying and locating characteristic peaks that are closely related to the characteristics of heavy metal ions.

[0138] Through the above steps, not only the data dimension was effectively reduced, but also the key heavy metal ion characteristic peaks were successfully identified, providing an important basis for subsequent heavy metal pollution assessment.

[0139] 103. Based on the preliminary prediction results of heavy metals, a partial least squares regression algorithm is used to construct a heavy metal concentration prediction model, quantitatively analyze the actual concentration of heavy metals in the water samples, and use cross-validation technology to evaluate the model stability and prediction accuracy to generate the actual concentration results of heavy metals;

[0140] The partial least squares regression algorithm is a multivariate statistical modeling method that combines the advantages of principal component analysis and multivariate linear regression, aiming to find the optimal linear relationship between multiple independent variables and dependent variables.

[0141] Cross-validation technology is a method for evaluating model performance, which usually includes K-fold cross-validation, etc. It trains the model by splitting the data set and tests its generalization ability.

[0142] The heavy metal concentration prediction model is a mathematical model constructed based on the partial least squares regression algorithm, which is used to predict the actual concentration of heavy metals in water samples.

[0143] The actual heavy metal concentration results are the specific concentration values ​​of various heavy metals in the water samples calculated by the partial least squares regression algorithm model.

[0144] In this step, first, based on the preliminary prediction results of heavy metals, appropriate parameter settings are selected to build a heavy metal concentration prediction model; second, the model is used to accurately estimate the actual concentration of heavy metals in water samples; third, through cross-validation techniques such as K-fold cross-validation, the training and testing process is repeated to ensure the stability and accuracy of the model; finally, the model parameters are adjusted according to the results of cross-validation until the optimal state is reached, thereby obtaining the final prediction value of the actual concentration of heavy metals.

[0145] Optionally, the method in step 103 is based on the preliminary prediction results of heavy metals, uses a partial least squares regression algorithm to construct a heavy metal concentration prediction model, quantitatively analyzes the actual concentration of heavy metals in the water sample, and uses a cross-validation technique to evaluate the model stability and prediction accuracy to generate the actual concentration result of heavy metals, including: based on the preliminary prediction results of heavy metals, collects Raman spectrum data of standard heavy metal solutions of known concentrations to generate a training data set; based on the training data set, uses a partial least squares regression algorithm to maximize the covariance between input variables and output variables to find the optimal linear combination to generate a heavy metal concentration prediction model; based on the heavy metal concentration prediction model, uses a cross-validation technique to evaluate the model, takes turns to use the divided subsets as independent test sets, repeats the test process to summarize the test results, and generates a model evaluation result; based on the model evaluation result, quantitatively analyzes the actual concentration of heavy metals in the water sample to generate the actual concentration result of heavy metals.

[0146] Among them, based on the training data set, the partial least squares regression algorithm is used to maximize the covariance between the input variables and the output variables to find the optimal linear combination and generate a heavy metal concentration prediction model, including: based on the training data set, standardization processing is performed to eliminate the dimension and numerical range differences between different spectral data, and each variable in the model training process is fairly compared to generate a pre-processed training data set; based on the pre-processed training data set, the partial least squares regression algorithm is used to select spectral data and known heavy metal concentration values ​​as input and output variables, and by maximizing the covariance between the input variables and the output variables, the optimal linear combination is found to generate a preliminary linear combination model; based on the preliminary linear combination model, the principal components of the input variables are gradually extracted to control the cumulative explained variance to reach a predetermined threshold, and an optimized principal component set is generated; based on the optimized principal component set, potential multicollinearity problems between multivariate data are processed to generate a heavy metal concentration prediction model.

[0147] The training data set contains Raman spectral data of standard heavy metal solutions with known concentrations, which is used to train and build prediction models.

[0148] The partial least squares regression algorithm is a multivariate statistical method that finds the optimal linear combination by maximizing the covariance between input variables and output variables. It is suitable for dealing with multicollinearity problems in multivariate data.

[0149] Standardization is the preprocessing of data to make it have the same dimension and numerical range to ensure fair comparison of variables during model training.

[0150] The principal components are the main components obtained by dimensionality reduction of the input variables, which can retain the information of the original data to the greatest extent.

[0151] The cumulative explained variance measures the cumulative contribution of the principal component set to the variability of the original data, and a threshold is usually set to control the number of principal components.

[0152] Multicollinearity refers to the high correlation between multiple input variables, which may lead to unstable estimation of model parameters.

[0153] In the embodiment of the present application, first, based on the preliminary prediction results of heavy metals, Raman spectral data of standard heavy metal solutions of known concentrations are collected to generate a training data set; secondly, the training data set is standardized to eliminate the differences in dimensions and numerical ranges between different spectral data, and a preprocessed training data set is generated; thirdly, based on the preprocessed training data set, a partial least squares regression algorithm is used to select spectral data and known heavy metal concentration values ​​as input and output variables, and the optimal linear combination is found by maximizing the covariance between the input variables and the output variables to generate a preliminary linear combination model; then, the principal components of the input variables are gradually extracted, the cumulative explained variance is controlled to reach a predetermined threshold, and an optimized set of principal components is generated; then, based on the optimized set of principal components, the potential multicollinearity problem between multivariate data is processed to generate a heavy metal concentration prediction model; finally, a cross-validation technique is used to evaluate the model, and the divided subsets are taken as independent test sets in turn, and the test process is repeated to summarize the test results, generate a model evaluation result, and based on the result, the actual concentration of heavy metals in the water sample is quantitatively analyzed to generate the actual concentration result of heavy metals.

[0154] Suppose you need to detect heavy metal pollution in wastewater discharged from an industrial area;

[0155] Firstly, based on the preliminary prediction results of heavy metals, Raman spectral data of multiple standard heavy metal solutions with known concentrations were collected to generate a training data set; then, the training data set was standardized to eliminate the differences in dimensions and numerical ranges between different spectral data, and a preprocessed training data set was generated; then, based on the preprocessed training data set, the partial least squares regression algorithm was used to select spectral data and known heavy metal concentration values ​​as input and output variables, and the optimal linear combination was found by maximizing the covariance between the input variables and the output variables, and a preliminary linear combination model was generated; the principal components of the input variables were gradually extracted, and the cumulative explained variance was controlled to reach a predetermined threshold of 85%, and an optimized set of principal components was generated; again, based on the optimized set of principal components, the potential multicollinearity problem between multivariate data was handled, and a heavy metal concentration prediction model was generated; finally, cross-validation technology was used to evaluate the model, and the divided subsets were taken as independent test sets in turn. The test process was repeated to summarize the test results, and the model evaluation results were generated. Based on the results, the actual concentrations of heavy metals in the wastewater discharged from the industrial area were quantitatively analyzed, and the actual concentration results of heavy metals were generated.

[0156] Through the above steps, not only a stable and accurate heavy metal concentration prediction model was constructed, but also the heavy metal content in the wastewater discharged from the industrial area was successfully and accurately quantitatively analyzed, providing a reliable scientific basis for environmental monitoring and pollution control.

[0157] This application takes into account the complex correlation between spectral data and heavy metal concentrations. It adopts a partial least squares regression algorithm, combined with nonlinear correlation calculation and weight optimization technology, solves the optimal linear combination through the gradient descent method, and introduces regularization and feature selection strategies to generate an accurate preliminary linear combination model, which effectively improves the prediction performance.

[0158] Optionally, based on the preprocessed training data set, a partial least squares regression algorithm is used to select spectral data and known heavy metal concentration values ​​as input and output variables, and the optimal linear combination is found by maximizing the covariance between the input variables and the output variables to generate a preliminary linear combination model, including:

[0159] Based on the preprocessed training data set, using a box plot to detect and remove outliers;

[0160] Applying a smoothing filter to further reduce noise and improve the data signal-to-noise ratio to generate nonlinear correlations;

[0161] The nonlinear correlation is calculated using the following formula:

[0162]

[0163] Among them, R xyis the nonlinear correlation between the input variable and the output variable; i is the sample index of the training data set, from 1 to N; N is the number of samples in the training data set; X i is the spectral data of the i-th sample; is the average value of all sample spectral data; σ X is the standard deviation of all sample spectral data; Y i is the known heavy metal concentration value of the i-th sample; is the average value of known heavy metal concentrations of all samples; σ Y is the standard deviation of the known heavy metal concentration values ​​of all samples; β is the nonlinear adjustment parameter used to control the sensitivity of nonlinear correlation;

[0164] Based on the nonlinear correlation, an initialization method based on prior knowledge is used to obtain a reasonable initial weight vector, a gradient descent method is used to solve the maximization objective function, and the weight vector is adjusted through multiple iterations to achieve convergence conditions to generate a weight vector of the optimal linear combination;

[0165] The weight vector of the optimal linear combination is calculated using the following formula:

[0166]

[0167] Where W is the weight vector of the optimal linear combination; w is the candidate weight vector; R xy is the nonlinear correlation matrix between input variables and output variables; α is the balance parameter used to control the influence of the diagonal matrix and the unit matrix; S x is the covariance matrix of the input variables; diag(w T S x w) is a diagonal matrix, and the diagonal elements are w T S x The elements of w; I is the identity matrix; γ is an additional adjustment parameter used to control the influence of the covariance matrix of the output variable; S y is the covariance matrix of the output variables; diag(w T S y w) is a diagonal matrix, and the diagonal elements are w T S y The elements of w; δ is another adjustment parameter used to control the influence of the cross covariance matrix between the input variables and the output variables; S xy is the cross covariance matrix of the input variables and the output variables; is a diagonal matrix with diagonal elements elements; λ is a nonlinear adjustment parameter used to control the nonlinear influence of the cross-covariance matrix;

[0168] Based on the weight vector of the optimal linear combination, feature selection is performed by recursive feature elimination method, and a regularization term is introduced to control the complexity of the model to prevent overfitting. The model is fitted with the known heavy metal concentration values ​​to generate a preliminary linear combination model.

[0169] This method aims to improve the accuracy and reliability of Raman spectroscopy data in predicting heavy metal concentrations by combining statistical and machine learning techniques and utilizing advanced data processing and analysis methods, thereby more effectively identifying and quantifying heavy metal pollutants in water samples.

[0170] In nonlinear correlation, the standardized input variables Standardize the input variables so that the spectral data of each sample is compared with a mean of 0 and a standard deviation of 1, thereby eliminating the influence of different scales; the standardized output variables Standardize the output variables so that the known heavy metal concentration values ​​of each sample are compared with a mean of 0 and a standard deviation of 1, thereby eliminating the influence of different scales; Gaussian kernel function The Gaussian kernel function is introduced to capture the nonlinear relationship between the input variables and the output variables. β controls the width of the Gaussian kernel function, thereby affecting the sensitivity of nonlinear correlation.

[0171] Among them, X i Obtained from the preprocessed training dataset; It is obtained by calculating the average value of all sample spectral data; σ X It is obtained by calculating the standard deviation of all sample spectral data; Y i Obtained from the preprocessed training dataset; It is obtained by calculating the average value of known heavy metal concentrations of all samples; σ Y It is obtained by calculating the standard deviation of the known heavy metal concentration values ​​of all samples; β is determined by experimental optimization, and usually ranges from 0.1 to 10;

[0172] In the weight vector of the optimal linear combination, the product of the weight vector and the nonlinear correlation matrix w T R xy :The correlation between the input variable and the output variable under the current weight vector is measured by multiplying the weight vector and the nonlinear correlation matrix; the diagonal matrix α·diag(w T S x w): diagonal elements are w T S x The elements of w are used to control the importance of the input variables. α is a balancing parameter that adjusts the influence of the diagonal matrix and the identity matrix. The identity matrix (1-α)·I: maintains the stability of the weight vector to avoid overfitting. The influence of the covariance matrix of the output variable γ·diag(w T Sy w): diagonal elements are w T S y The elements of w are used to control the importance of the output variables, γ is an additional adjustment parameter; the nonlinear effects of the cross-covariance matrix The diagonal elements are The elements of are used to control the nonlinear effects of the cross-covariance matrix, δ is another adjustment parameter, and λ is a nonlinear adjustment parameter;

[0173] Among them, w is obtained through prior knowledge or random initialization; R xy Calculated by nonlinear correlation formula; α is determined by experimental optimization and usually takes a value between 0 and 1; S x Obtained by calculating the covariance of the input variables; I standard unit matrix; γ is determined by experimental optimization, usually in the range of 0.1 to 10; S y It is obtained by calculating the covariance of the output variables; δ is determined by experimental optimization, usually ranging from 0.1 to 10; S xy It is obtained by calculating the cross covariance between the input variable and the output variable; λ is determined by experimental optimization, and usually ranges from 0.1 to 10;

[0174] Suppose the experimenter needs to deal with one;

[0175] Assume that the original spectrum data X(t) is known, the sampling interval Δt = 0.01 seconds, the total number of data points N = 1000; Assume that the parameters are set to β = 5; α = 0.5; γ = 1; δ = 0.8; λ = 2; Assume that X i and Y i The specific values ​​are X1=100, Y1=1.2; X2=105, Y2=1.5; X3=95, Y3=1.0; ...; Assume that S x , S y , S xy , w is known;

[0176]

[0177] Assume that the threshold of nonlinear correlation is set to 0.5. xy=0.6 is greater than the set threshold, indicating that the nonlinear relationship between the input variable and the output variable is strong, which helps the model to capture the complex relationship between the two more accurately; assuming that the threshold of the weight vector is set to 0.1, since all components in W = [0.3, 0.4, 0.2, …] are greater than the set threshold, it means that each input variable has a significant contribution in the model, which further supports the effectiveness and predictive ability of the model; through the above steps, the weight vector of nonlinear correlation and optimal linear combination was successfully calculated based on the Raman spectral data of lake water samples, providing a reliable data basis for subsequent heavy metal concentration prediction.

[0178] 104. Based on the actual heavy metal concentration results, a comparative analysis is conducted with national or regional water quality safety standards to comprehensively assess the heavy metal pollution in water bodies and generate a water quality heavy metal pollution detection report.

[0179] National or regional water safety standards are a series of water quality index limits set by government agencies to protect human health and the environment from pollution.

[0180] The water quality heavy metal pollution detection report is a document formed by integrating all test results, data analysis and comparison with water quality standards, which comprehensively reflects the heavy metal pollution status in the water body.

[0181] In this step, first, the actual concentration results of the heavy metals are collected and sorted; second, the measured heavy metal concentrations are compared with the latest national or regional water quality safety standards to determine whether they exceed the allowable range; third, a comprehensive pollution status assessment is conducted based on the comparison results; finally, a detailed water quality heavy metal pollution detection report is compiled, covering the detection methods, heavy metal types and concentrations, exceeding standards, potential impacts, and recommended control measures, etc., which will help to formulate effective pollution prevention and control strategies.

[0182] Optionally, the actual heavy metal concentration results in step 104 are compared and analyzed with national or regional water quality safety standards to comprehensively assess the heavy metal pollution in the water body and generate a water quality heavy metal pollution detection report, including: based on the actual heavy metal concentration results, collecting and collating the current national or regional water quality safety standards, obtaining heavy metal concentration limit values, and generating a water quality safety standard comparison table; based on the water quality safety standard comparison table, the actual heavy metal concentration results are compared one by one, and the exceeded items and the exceeded multiples are marked to generate a heavy metal concentration exceeded situation table; based on the heavy metal concentration exceeded situation table, the heavy metal pollution in the water body is comprehensively assessed, a hazard analysis is performed, and a water body heavy metal pollution assessment conclusion is generated; based on the water body heavy metal pollution assessment conclusion, in combination with the specific sampling time and location, a water quality treatment plan and monitoring plan are compiled to generate a water quality heavy metal pollution detection report.

[0183] The water quality safety standard comparison table contains the heavy metal concentration limit values ​​specified in the current national or regional water quality safety standards, which is used to compare with the actual test results.

[0184] The table of excessive heavy metal concentrations records the names of the excessive items in the actual test results and their excess multiples, which is used to identify the degree of pollution.

[0185] The conclusion of the heavy metal pollution assessment in water bodies is a comprehensive assessment of the heavy metal pollution situation in water bodies based on the table of exceeding standards, and an analysis of possible hazards.

[0186] The water quality control plan and monitoring plan are to formulate specific control measures and future monitoring plans based on the conclusions of the pollution assessment to ensure water quality improvement and continuous monitoring.

[0187] In the embodiments of the present application, first, based on the actual heavy metal concentration results, the heavy metal concentration limit values ​​in the current national or regional water quality safety standards are collected and sorted out to generate a water quality safety standard comparison table; secondly, based on the water quality safety standard comparison table, the actual heavy metal concentration results are compared one by one, the exceeded items and the exceeded multiples are marked, and a heavy metal concentration exceeded situation table is generated; thirdly, based on the heavy metal concentration exceeded situation table, the heavy metal pollution situation in the water body is comprehensively evaluated, a hazard analysis is performed, and an assessment conclusion of the heavy metal pollution in the water body is generated; finally, based on the assessment conclusion of the heavy metal pollution in the water body, combined with the specific sampling time and location, a water quality treatment plan and monitoring plan are compiled to generate a water quality heavy metal pollution detection report.

[0188] Suppose you need to detect heavy metal pollution in a city’s lake;

[0189] First, based on the actual results of heavy metal concentrations, the heavy metal concentration limit values ​​in the current national water quality safety standards were collected and organized to generate a water quality safety standard comparison table; then, based on the comparison table, the actual heavy metal concentration results of lake water samples were compared one by one, and the exceeded items (such as lead and cadmium) and their exceedance multiples were marked to generate a heavy metal concentration exceedance situation table; then, based on the exceedance situation table, a comprehensive assessment of heavy metal pollution in lakes was conducted, and the possible hazards of these heavy metals to the environment and human health were analyzed, and an assessment conclusion on heavy metal pollution in water bodies was generated; finally, based on the assessment conclusions, combined with the specific sampling time and location, a detailed water quality treatment plan and a monitoring plan for the next year were compiled, including recommended treatment measures (such as sediment cleaning, water source protection, etc.) and regular monitoring time points, and a complete water quality heavy metal pollution detection report was generated.

[0190] Through the above steps, not only the heavy metal pollution status of the lake was comprehensively assessed, but also a scientific basis and action guide were provided to relevant departments, which helped to take timely measures to ensure water quality safety.

[0191] Figure 2 The present invention provides a schematic diagram of a water heavy metal pollution detection system based on Raman spectroscopy. Figure 2 As shown, the device comprises:

[0192] The collection module 21 is used to collect water samples from the water area to be tested, and quickly scan the water samples using a high-sensitivity Raman spectrometer to generate an original Raman spectrum;

[0193] The processing module 22 is used to perform denoising based on the original Raman spectrum using a wavelet transform algorithm to reduce background interference to improve spectrum quality, extract characteristic peaks related to heavy metal ions using a principal component analysis technique, and generate a preliminary prediction result of heavy metals;

[0194] A construction module 23 is used to construct a heavy metal concentration prediction model based on the preliminary prediction results of heavy metals by using a partial least squares regression algorithm, quantitatively analyze the actual concentration of heavy metals in the water sample, evaluate the model stability and prediction accuracy by using a cross-validation technique, and generate the actual concentration results of heavy metals;

[0195] The analysis module 24 is used to compare and analyze the actual heavy metal concentration results with national or regional water quality safety standards, comprehensively evaluate the heavy metal pollution in the water body, and generate a water quality heavy metal pollution detection report.

[0196] Figure 2 The water quality heavy metal pollution detection system based on Raman spectroscopy can be performed Figure 1 The implementation principle and technical effect of the method for detecting heavy metal pollution in water based on Raman spectroscopy described in the embodiment shown are not repeated here. The specific way in which each module and unit performs operations in the system for detecting heavy metal pollution in water based on Raman spectroscopy in the above embodiment has been described in detail in the embodiment of the method, and will not be elaborated here.

[0197] In one possible design, Figure 2 The water quality heavy metal pollution detection system based on Raman spectroscopy in the embodiment shown can be implemented as a computing device, such as Figure 3 As shown, the computing device may include a storage component 31 and a processing component 32;

[0198] The storage component 31 stores one or more computer instructions, wherein the one or more computer instructions are called and executed by the processing component 32 .

[0199] The processing component 32 is used to: collect water samples from the water area to be tested, quickly scan the water samples using a high-sensitivity Raman spectrometer, and generate an original Raman spectrum; based on the original Raman spectrum, use a wavelet transform algorithm to perform denoising processing to reduce background interference to improve spectral quality, use principal component analysis technology to extract characteristic peaks related to heavy metal ions, and generate preliminary prediction results of heavy metals; based on the preliminary prediction results of heavy metals, use a partial least squares regression algorithm to construct a heavy metal concentration prediction model, quantitatively analyze the actual concentration of heavy metals in the water samples, use cross-validation technology to evaluate the model stability and prediction accuracy, and generate actual heavy metal concentration results; based on the actual heavy metal concentration results, compare and analyze with national or regional water quality safety standards, comprehensively evaluate the heavy metal pollution of water bodies, and generate a water quality heavy metal pollution detection report.

[0200] The processing component 32 may include one or more processors to execute computer instructions to complete all or part of the steps in the above method. Of course, the processing component may also be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors or other electronic components to perform the above method.

[0201] The storage component 31 is configured to store various types of data to support operations at the terminal. The storage component can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.

[0202] Of course, the computing device may also include other components, such as input / output interfaces, display components, communication components, etc.

[0203] The input / output interface provides an interface between the processing component and the peripheral interface module, which may be an output device, an input device, etc.

[0204] The communication component is configured to facilitate, among other things, wired or wireless communications between the computing device and other devices.

[0205] Among them, the computing device can be a physical device or an elastic computing host provided by a cloud computing platform, etc. In this case, the computing device can refer to a cloud server, and the above-mentioned processing components, storage components, etc. can be basic server resources rented or purchased from the cloud computing platform.

[0206] The present application also provides a computer storage medium storing a computer program, wherein the computer program can achieve the above-mentioned Figure 1 The illustrated embodiment is a method for detecting heavy metal pollution in water based on Raman spectroscopy.

[0207] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0208] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Those of ordinary skill in the art may understand and implement it without creative effort.

[0209] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0210] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit it. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A method for detecting heavy metal pollution in water based on Raman spectroscopy, characterized in that: include: Collect water samples from the water area to be tested, and quickly scan the water samples using a high-sensitivity Raman spectrometer to generate an original Raman spectrum; Based on the original Raman spectrum, a wavelet transform algorithm is used to perform denoising, reduce background interference to improve spectrum quality, and a principal component analysis technique is used to extract characteristic peaks related to heavy metal ions to generate preliminary prediction results of heavy metals; Based on the preliminary prediction results of heavy metals, a heavy metal concentration prediction model is constructed using a partial least squares regression algorithm, the actual concentration of heavy metals in the water samples is quantitatively analyzed, and cross-validation technology is used to evaluate the model stability and prediction accuracy to generate actual heavy metal concentration results; Based on the actual heavy metal concentration results, a comparative analysis is conducted with national or regional water quality safety standards, a comprehensive assessment of the heavy metal pollution in water bodies is conducted, and a water quality heavy metal pollution detection report is generated.

2. The method according to claim 1, characterized in that Based on the original Raman spectrum, the wavelet transform algorithm is used for denoising, background interference is reduced to improve the spectrum quality, and the principal component analysis technology is used to extract the characteristic peaks related to heavy metal ions to generate the preliminary prediction results of heavy metals, including: Based on the original Raman spectrum, normalization processing is performed to eliminate intensity changes caused by concentration differences of different samples, ensure that subsequent analysis is consistent and accurate, and generate normalized spectrum data; Based on the normalized spectral data, a wavelet transform algorithm is performed, a suitable mother wavelet function is selected, the number of decomposition layers is set, and the signal noise component and the useful signal are separated by multi-scale analysis to reduce background interference and generate a denoised Raman spectrum; Based on the denoised Raman spectrum, principal component analysis technology is used to perform dimensionality reduction processing, extract the principal components representing the characteristics of heavy metal ions, obtain characteristic peaks related to heavy metal ions, and generate characteristic peak data; Based on the characteristic peak data, combined with the known heavy metal ion standard spectrum database, a comparative analysis is performed to preliminarily determine the possible types and relative contents of heavy metal ions in the water sample and generate a preliminary prediction result of heavy metals.

3. The method according to claim 2, characterized in that The method of performing wavelet transform algorithm processing based on the normalized spectral data, selecting a suitable mother wavelet function, setting the number of decomposition layers, separating the signal noise component from the useful signal through multi-scale analysis, reducing background interference, and generating a denoised Raman spectrum includes: Based on the normalized spectral data, preprocessing is performed to determine a data baseline and a range, thereby generating preprocessed spectral data; Based on the pre-processed spectral data, a mother wavelet function matching the spectral data characteristics is selected, an appropriate number of decomposition layers is set, and a wavelet transform algorithm is used to perform wavelet transform processing to generate wavelet transform coefficients; Based on the wavelet transform coefficients, the useful signal components in the spectral data are identified and enhanced by multi-scale analysis technology, the noise components are removed, the background interference is reduced, and the optimized wavelet transform coefficients are generated; Based on the optimized wavelet transform coefficients, an inverse wavelet transform process is performed to restore the time domain signal and generate a denoised Raman spectrum.

4. The method according to claim 3, characterized in that: Based on the pre-processed spectral data, a mother wavelet function matching the spectral data characteristics is selected, an appropriate number of decomposition layers is set, and a wavelet transform algorithm is used to perform wavelet transform processing to generate wavelet transform coefficients, including: Based on the pre-processed spectral data, a Gaussian filtering method is used to reduce noise, and a background signal is removed by polynomial fitting to ensure the effectiveness of nonlinear conversion, so as to generate a nonlinear conversion output; The nonlinear conversion output is calculated using the following formula: Among them, Y(t) is the nonlinear conversion output; X(t) is the original spectral data; γ is the conversion rate, which determines the steepness of the nonlinear conversion; θ is the conversion threshold, which determines the center point of the nonlinear conversion; λ1 is the first amplitude, which controls the intensity of the first periodic change; f1 is the first frequency, which defines the speed of the first periodic change; is the first initial phase, which defines the starting position of the first periodic change; λ2 is the second amplitude, which controls the intensity of the second periodic change; f2 is the second frequency, which defines the speed of the second periodic change; is the second initial phase, defining the starting position of the second periodic change; t is the time variable, indicating the time point of the spectral data; Based on the nonlinear conversion output, a Gaussian kernel function is introduced to perform local weighting to highlight local features, and sine and cosine terms are introduced to capture periodic components in spectral data to generate spectral features; The extracted spectral features are calculated using the following formula: Z(k) is the extracted spectral feature; Y(t) is the nonlinear conversion output; Y(iΔt) is the value of the nonlinear conversion output at time point iΔt; i is the index of the data point, from 1 to N; N is the total number of data points; Δt is the sampling interval; H(x) is the Gaussian kernel function, used for local weighting; τ is the time constant, which determines the influence range of the Gaussian kernel function; λ3 is the influence factor, which determines the rate of exponential decay; λ4 is the third amplitude, which controls the intensity of the third periodic change; f3 is the third frequency, which defines the speed of the third periodic change; is the third initial phase, defining the starting position of the third periodic change; λ5 is the fourth amplitude, controlling the intensity of the fourth periodic change; f4 is the fourth frequency, defining the speed of the fourth periodic change; is the fourth initial phase, defining the starting position of the fourth periodic change; k is a discrete position, indicating the position of feature extraction; Based on the extracted spectral features, combined with the spectral data features and frequency domain characteristics, a suitable mother wavelet function is selected, the number of wavelet transform decomposition layers is determined through experimental results, and wavelet transform coefficients are generated.

5. The method according to claim 2, characterized in that: Based on the denoised Raman spectrum, the principal component analysis technology is used to perform dimensionality reduction processing, extract the principal components representing the characteristics of heavy metal ions, obtain the characteristic peaks related to heavy metal ions, and generate characteristic peak data, including: Based on the denoised Raman spectrum, each spectrum sample is sorted, and a spectrum data matrix is ​​constructed to integrate each spectrum sample and wave number point to generate a spectrum data matrix; Based on the spectral data matrix, statistical analysis is performed on the spectral information of all samples to reveal the correlation of spectral signals between different wave number points and generate a covariance matrix; Based on the covariance matrix, the eigenvalues ​​and eigenvectors are solved by mathematical methods, and the eigenvectors with the largest eigenvalues ​​are selected as principal components, so as to retain the main variation information of the original Raman spectrum to the maximum extent and generate a principal component set; Based on the principal component set, the high-dimensional spectral data is projected into the low-dimensional space constituted by the principal components, the principal components reflecting the characteristics of heavy metal ions are extracted, the positions and intensities of characteristic peaks related to heavy metal ions are analyzed and determined, and characteristic peak data are generated.

6. The method according to claim 1, characterized in that Based on the preliminary prediction results of heavy metals, a partial least squares regression algorithm is used to construct a heavy metal concentration prediction model, quantitatively analyze the actual concentration of heavy metals in the water samples, and use cross-validation technology to evaluate the model stability and prediction accuracy to generate actual heavy metal concentration results, including: Based on the preliminary prediction results of heavy metals, Raman spectral data of standard heavy metal solutions with known concentrations are collected to generate a training data set; Based on the training data set, a partial least squares regression algorithm is used to maximize the covariance between the input variables and the output variables to find the optimal linear combination and generate a heavy metal concentration prediction model; Based on the heavy metal concentration prediction model, a cross-validation technique is used to evaluate the model, and the divided subsets are used as independent test sets in turn, and the test process is repeated to summarize the test results and generate a model evaluation result; Based on the model evaluation results, the actual concentration of heavy metals in the water sample is quantitatively analyzed to generate actual concentration results of heavy metals.

7. The method according to claim 6, characterized in that Based on the training data set, the partial least squares regression algorithm is used to maximize the covariance between the input variables and the output variables to find the optimal linear combination and generate a heavy metal concentration prediction model, including: Based on the training data set, standardization processing is performed to eliminate the differences in dimensions and numerical ranges between different spectral data, and each variable in the model training process is fairly compared to generate a preprocessed training data set; Based on the preprocessed training data set, a partial least squares regression algorithm is used to select spectral data and known heavy metal concentration values ​​as input and output variables, and the optimal linear combination is found by maximizing the covariance between the input variables and the output variables to generate a preliminary linear combination model; Based on the preliminary linear combination model, the principal components of the input variables are gradually extracted to control the cumulative explained variance to reach a predetermined threshold value, thereby generating an optimized principal component set; Based on the optimized principal component set, potential multicollinearity problems among multivariate data are handled to generate a heavy metal concentration prediction model.

8. The method according to claim 7, characterized in that Based on the preprocessed training data set, the partial least squares regression algorithm is used to select spectral data and known heavy metal concentration values ​​as input and output variables, and the optimal linear combination is found by maximizing the covariance between the input variables and the output variables to generate a preliminary linear combination model, including: Based on the preprocessed training data set, using a box plot to detect and remove outliers; Applying a smoothing filter to further reduce noise and improve the data signal-to-noise ratio to generate nonlinear correlations; The nonlinear correlation is calculated using the following formula: Among them, R xy is the nonlinear correlation between the input variable and the output variable; i is the sample index of the training data set, from 1 to N; N is the number of samples in the training data set; X i is the spectral data of the i-th sample; is the average value of all sample spectral data; σ X is the standard deviation of all sample spectral data; Y i is the known heavy metal concentration value of the i-th sample; is the average value of known heavy metal concentrations of all samples; σ Y is the standard deviation of the known heavy metal concentration values ​​of all samples; β is the nonlinear adjustment parameter used to control the sensitivity of nonlinear correlation; Based on the nonlinear correlation, an initialization method based on prior knowledge is used to obtain a reasonable initial weight vector, a gradient descent method is used to solve the maximization objective function, and the weight vector is adjusted through multiple iterations to achieve convergence conditions to generate a weight vector of the optimal linear combination; The weight vector of the optimal linear combination is calculated using the following formula: Where W is the weight vector of the optimal linear combination; w is the candidate weight vector; R xy is the nonlinear correlation matrix between input variables and output variables; α is the balance parameter used to control the influence of the diagonal matrix and the unit matrix; S x is the covariance matrix of the input variables; diag(w T S x w) is a diagonal matrix, and the diagonal elements are w T S x The elements of w; I is the identity matrix; γ is an additional adjustment parameter used to control the influence of the covariance matrix of the output variable; S y is the covariance matrix of the output variables; diag(w T S y w) is a diagonal matrix, and the diagonal elements are w T S y The elements of w; δ is another adjustment parameter used to control the influence of the cross covariance matrix between the input variables and the output variables; S xy is the cross covariance matrix of the input variables and the output variables; is a diagonal matrix with diagonal elements elements; λ is a nonlinear adjustment parameter used to control the nonlinear influence of the cross-covariance matrix; Based on the weight vector of the optimal linear combination, feature selection is performed by recursive feature elimination method, and a regularization term is introduced to control the complexity of the model to prevent overfitting. The model is fitted with the known heavy metal concentration values ​​to generate a preliminary linear combination model.

9. The method according to claim 1, characterized in that: Based on the actual heavy metal concentration results, a comparative analysis is conducted with national or regional water quality safety standards to comprehensively assess the heavy metal pollution in water bodies and generate a water quality heavy metal pollution detection report, including: Based on the actual heavy metal concentration results, collect and organize the current national or regional water quality safety standards, obtain the heavy metal concentration limit values, and generate a water quality safety standard comparison table; Based on the water quality safety standard comparison table, the actual heavy metal concentration results are compared one by one, the exceeding items and the exceeding multiples are marked, and a heavy metal concentration exceeding condition table is generated; Based on the table of excessive heavy metal concentrations, comprehensively assess the heavy metal pollution in water bodies, conduct hazard analysis, and generate a conclusion on the assessment of heavy metal pollution in water bodies; Based on the conclusions of the heavy metal pollution assessment of the water body, combined with the specific sampling time and location, a water quality treatment plan and monitoring plan are formulated to generate a water quality heavy metal pollution detection report.

10. A water quality heavy metal pollution detection system based on Raman spectroscopy, characterized in that: include: A collection module is used to collect water samples from the water area to be tested, and quickly scan the water samples using a high-sensitivity Raman spectrometer to generate an original Raman spectrum; A processing module is used to perform denoising based on the original Raman spectrum using a wavelet transform algorithm to reduce background interference to improve spectrum quality, extract characteristic peaks related to heavy metal ions using a principal component analysis technique, and generate a preliminary prediction result of heavy metals; A construction module is used to construct a heavy metal concentration prediction model based on the preliminary prediction results of heavy metals by using a partial least squares regression algorithm, quantitatively analyze the actual concentration of heavy metals in the water sample, evaluate the model stability and prediction accuracy by using a cross-validation technique, and generate the actual concentration results of heavy metals; The analysis module is used to compare and analyze the actual heavy metal concentration results with national or regional water quality safety standards, comprehensively evaluate the heavy metal pollution in the water body, and generate a water quality heavy metal pollution detection report.

Citation Information

Patent Citations

  • Raman spectral preprocessing method

    CN103217409A

  • Raman analysis method for trace impurities in parachlorotoluene

    CN108982468A

  • Method for detecting multiple substances in real time in microbial culture process

    CN117216724A

  • Soil heavy metal detection system and method

    CN118671303A

  • Aquatic heavy metal ion detecting system based on surface reinforcing raman spectroscopy

    CN204832041U

Cited By

  • Low-temperature crystallization process early warning method and system based on digital twin polypeptide drugs

    CN120473014A

  • Pesticide residue detection method based on Raman spectrum and deep learning and electronic equipment

    CN120490045A

  • Pesticide residue detection method based on raman spectrum and deep learning and electronic device

    CN120490045B

  • Mass spectrum detection and analysis system and method based on glucan

    CN121141552A

  • Dextran quality spectrum detection analysis system and method

    CN121141552B