A method for evaluating heavy metal pollution in marine environment

CN122835983APending Publication Date: 2026-09-29HAINAN ACAD OF ENVIRONMENTAL SCI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611302098.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-26
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

[0008]针对现有技术的缺陷或改进需求,本发明提供了一种海洋环境重金属污染程度的评价方法,其目的在于通过构建基于动态本底替代统一背景值的评价基准、基于生物有效性替代总量的风险表征、以及基于光谱智能预测的高效监测手段,实现三维耦合的综合评价,从而解决现有方法中背景值时空代表性差、忽略生物有效性、以及监测成本高技术缺陷的技术问题

Benefits of technology

本发明通过正定矩阵因子分解模型为每个站位解析出专属的动态本底值,替代了传统方法中区域统一的固定背景值,能够有效区分自然地质高背景与人为输入,避免天然高背景区域被误判为污染,也避免人为输入被低估,显著提高了污染程度评价的准确性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122835983A_ABST
    Figure CN122835983A_ABST
Patent Text Reader

Abstract

This invention discloses a method for assessing heavy metal pollution in marine environments, belonging to the field of marine environmental monitoring and assessment technology. It involves collecting sediment samples and visible-near-infrared reflectance spectral data; measuring the total concentration and active form concentration of heavy metals in the sediments; using a positive definite matrix factorization model to analyze the dynamic background values ​​corresponding to natural parent material source factors at each station, and calculating anthropogenic enrichment factors accordingly; determining the ecological risk level based on the active form concentration using an active form risk index; training a dual-scale long short-term memory network model with an attention mechanism using spectral data as input and measured total and active form concentrations as labels, outputting predicted concentrations and confidence levels; integrating the three dimensions to generate a comprehensive risk rating and dynamically adjusting monitoring strategies accordingly. This invention achieves accurate pollution assessment with a single station and a complete background value, overcoming the shortcomings of traditional methods such as inaccurate background values, neglect of morphological differences, and high monitoring costs.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of marine environmental monitoring and assessment technology, and in particular to an assessment method for heavy metal pollution in the marine environment. Background Technology

[0002] With the rapid development of coastal industries and marine aquaculture, heavy metal pollution in the marine environment has become increasingly prominent. Heavy metals are characterized by their difficulty in degradation, easy accumulation in sediments, and amplification through the food chain, posing a potential threat to marine ecosystems and human health. Therefore, scientifically and accurately assessing the degree of heavy metal pollution in the marine environment is of great significance for pollution control and ecological risk management.

[0003] Currently, the assessment methods for heavy metal pollution in marine environments mainly include the following categories: The first category is single-factor or multi-factor index methods based on total amount. Examples include geoaccumulation index methods, single-factor pollution index methods, and comprehensive pollution index methods. These methods determine the degree of pollution by comparing the measured concentration of heavy metals in sediments with background values ​​or standard limits. However, traditional methods typically use a uniform background value for the region, ignoring the differences in natural geological conditions at different stations. This leads to naturally high background areas being misjudged as polluted, while areas with anthropogenic input are underestimated. Secondly, these methods only focus on the total amount of heavy metals, ignoring the occurrence forms and bioavailability of heavy metals. In fact, the toxicity of different forms of heavy metals to organisms varies greatly, and total amount alone cannot accurately assess ecological risk.

[0004] The second category is the potential ecological risk index method. This method combines heavy metal concentration and toxicity coefficient to assess potential harm to ecosystems. However, this method suffers from numerous misuses and abuses in practical applications: on the one hand, it was originally designed for aquatic sediments but has been widely used for other media such as soil and dust; on the other hand, its risk classification standards need to be adjusted according to the type and quantity of pollutants being evaluated, but many studies directly adopt the original standards, leading to biased evaluation results.

[0005] The third category is the sediment quality benchmark method. This method sets thresholds based on biotoxicity experimental data. However, most existing benchmarks are derived from international species toxicity data and lack chronic toxicity data based on representative native species. In addition, this method fails to fully consider the significant impact of environmental conditions such as salinity, temperature, pH, and dissolved organic carbon on heavy metal toxicity, which may underestimate or overestimate the actual risk.

[0006] The fourth category is traditional chemical detection methods. Although these methods are highly accurate, they are time-consuming, costly, and require complex sample pretreatment, making it difficult to achieve large-scale, high-frequency monitoring.

[0007] In summary, existing evaluation methods have shortcomings in areas such as background value determination, bioavailability characterization, monitoring efficiency, and data-driven prediction. There is an urgent need for a comprehensive evaluation scheme that can overcome these deficiencies. Therefore, this paper proposes an evaluation method for heavy metal pollution in marine environments. Summary of the Invention

[0008] To address the shortcomings or improvement needs of existing technologies, this invention provides a method for evaluating the degree of heavy metal pollution in the marine environment. The purpose is to achieve a comprehensive evaluation through three-dimensional coupling by constructing an evaluation benchmark based on dynamic background values ​​replacing unified background values, risk characterization based on bioavailability replacing total amount, and efficient monitoring methods based on spectral intelligent prediction. This solves the technical problems of poor spatiotemporal representativeness of background values, neglect of bioavailability, and high monitoring costs in existing methods.

[0009] To achieve the above objectives, this invention provides a method for evaluating the degree of heavy metal pollution in the marine environment, comprising the following steps: Step S1: Site setup and sample collection.

[0010] Several sampling stations were set up in the study area, arranged in a grid pattern to prioritize coverage of pollution source areas, control areas, and ecologically sensitive areas. Surface sediment samples were collected from each station using a grab-type sediment sampler at a depth of 0–5 cm. The collected sediment samples were placed in sealed polyethylene bags, stored at 4°C, and transported back to the laboratory as soon as possible.

[0011] Meanwhile, a portable ground object spectrometer was used to collect visible-near-infrared reflectance spectral data of sediment surfaces at each station. The preferred spectral wavelength range was 350~2500nm. The average value was taken after repeated measurements at each station. Environmental parameters at each station, including temperature, salinity and pH value, were recorded simultaneously.

[0012] Step S2: Determination of total heavy metal content in sediments.

[0013] The collected sediment samples were freeze-dried, ground, and sieved through a 200-mesh sieve. Sample pretreatment was performed using microwave digestion, followed by determination of the total concentration of multiple heavy metal elements using inductively coupled plasma mass spectrometry (ICP-MS) or inductively coupled plasma atomic emission spectrometry (ICP-AES). The heavy metal elements determined were selected from at least five of the following: arsenic (As), cadmium (Cd), chromium (Cr), copper (Cu), mercury (Hg), lead (Pb), zinc (Zn), and nickel (Ni).

[0014] Simultaneously, auxiliary parameters of the sediment samples were determined, including particle size distribution, organic carbon content, and iron and manganese oxide content. Particle size distribution was determined using a laser particle size analyzer, organic carbon content was determined using the potassium dichromate oxidation method, and iron and manganese oxide content was determined using a chemical extraction method.

[0015] Step S3: Determination of the concentration of active heavy metals.

[0016] The concentration of active heavy metals in sediment samples from each station was determined using the Chelex-100 resin extraction method. Specifically, a certain amount of fresh sediment sample was weighed, and artificial seawater pretreated with Chelex-100 resin was added. Extraction was carried out for a predetermined time under isothermal oscillation conditions. After centrifugation, the supernatant was collected, and the concentration of active metals in the supernatant was determined by inductively coupled plasma mass spectrometry.

[0017] As a preferred option, parameters such as extraction time and solid-liquid ratio are optimized and determined through preliminary experiments based on the characteristics of the research sea area.

[0018] Step S4: Dynamic background calculation based on the positive definite matrix factorization model.

[0019] Input the total heavy metal concentration data of each station obtained in step S2 into the positive definite matrix factorization model. Run the positive definite matrix factorization model to resolve the source factors of heavy metals at each station. The number of resolved factors is preferably 3 to 5. Identify the factor type by the combination of characteristic elements of each factor: If Fe, Al and Cr, Ni and other elements are strongly correlated in a certain factor, then the factor is identified as a factor representing the source of natural parent material. If elements such as Cu, Pb, Zn, and As are strongly correlated in a certain factor, then that factor is identified as a factor originating from human activities.

[0020] The contribution concentration of each element corresponding to the factor representing the natural source of the parent material is used as the dynamic background value of the site.

[0021] Based on this, the human enrichment factor for each site is calculated using the following formula: EF PMF =C 实测 / C 本底 , where C 实测 C represents the measured total concentration of a certain element. 本底 This is the dynamic baseline value of the element on this site.

[0022] The level of anthropogenic pollution is determined based on the numerical range of the anthropogenic enrichment factor: When EF PMF When the value is less than 1, it is determined to be uncontaminated; When 1≤EF PMF When the concentration is less than 1.5, it is considered lightly polluted. When 1.5≤EF PMF When the level is less than 3, it is considered moderate pollution. When EF PMF When the concentration is ≥3, it is considered to be severely polluted.

[0023] Step S5: Calculation of the active state risk index.

[0024] Based on the active state concentration measured in step S3, the active state risk index (BRI) for each station is calculated using the formula: BRI = Σ(C 活性,i / TEL i )×W i , where C 活性,i Let TEL be the active state concentration of the i-th heavy metal element. i Let W be the critical effect concentration of the i-th heavy metal element. i Let be the toxicity weight coefficient of the i-th heavy metal element derived from the species sensitivity distribution.

[0025] Critical effect concentration TEL i The value of is preferably determined by referring to the critical effect concentration value in the sediment quality benchmark. Toxicity weighting coefficient W i The method for determining the toxicity coefficients is as follows: collect toxicity data of representative near-shore species, construct species sensitivity distribution curves, and derive the toxicity weight coefficients of each heavy metal element from these curves. When toxicity data of local species is lacking, toxicity response coefficients can be used temporarily, but these coefficients need to be gradually updated to values ​​based on local species in subsequent evaluations.

[0026] The ecological risk level is determined based on the numerical range of the active state risk index: When BRI < 1, it is judged as low ecological risk; When 1≤BRI<2, it is judged as medium ecological risk; When 2 ≤ BRI < 3, it is judged as high ecological risk; When BRI ≥ 3, it is judged as extremely high ecological risk.

[0027] Furthermore, it also includes an in-situ bioaccumulation verification step: pre-labeled indicator organisms with stable isotopes are released to a selected site, exposed for a predetermined period, and then recovered. The changes in the isotope ratios in the organisms are measured to obtain bioaccumulation kinetic parameters under field conditions, which serve as supplementary verification indicators for the active state risk index.

[0028] Step S6: Construct a prediction model based on spectral data and long short-term memory networks.

[0029] The visible-near-infrared reflectance spectral data collected in step S1 are preprocessed, including denoising, smoothing, baseline correction, and first-order or second-order differential transformation, in order to highlight the characteristic absorption peaks related to heavy metals and form a standardized spectral dataset.

[0030] Using preprocessed spectral data as input features, and the total heavy metal concentration measured in step S2 and the active state concentration measured in step S3 as labels, the training set and validation set are divided according to a predetermined ratio to construct a dual-scale long short-term memory network model with attention mechanism.

[0031] The dual-scale long short-term memory network model with attention mechanism includes a dual-scale feature extraction module and an attention mechanism module. The dual-scale feature extraction module extracts local and global features from the spectral data, respectively: local feature extraction uses a first-scale long short-term memory network unit to capture short-range correlations and local absorption features in the spectral data; global feature extraction uses a second-scale long short-term memory network unit to capture long-range dependencies and overall trend features in the spectral data. The attention mechanism module weights and focuses the features output by the dual-scale feature extraction module, automatically identifying the spectral feature bands that contribute most to heavy metal concentration prediction and assigning higher weight coefficients to the corresponding features.

[0032] The training set data is input into the model for training, with the objective function being to minimize the root mean square error between the predicted and measured values. The backpropagation algorithm is used to update the model parameters. After the model training is complete, the validation set data is input into the model for validation.

[0033] After the model is validated, the spectral data of each station is input into the trained prediction model, which outputs the predicted values ​​of total heavy metal concentration and active state concentration of each station, as well as the confidence level of the predicted values ​​of each station.

[0034] The confidence level is determined as follows: based on the width of the prediction interval output by the prediction model, a high confidence level is determined when the prediction interval width is less than a first preset threshold, a low confidence level is determined when the prediction interval width is greater than a second preset threshold, and a medium confidence level is determined when the width falls between the two. The first and second preset thresholds are determined by those skilled in the art based on specific model accuracy requirements and monitoring objectives.

[0035] As a preferred option, the coefficient of determination R is used. 2 The root mean square error is used as a model evaluation metric, when the coefficient of determination R on the validation set... 2 If the value is less than 0.7, return to adjust the model structure or add training data and retrain.

[0036] Step S7: Construction of a three-dimensional risk matrix and comprehensive risk rating.

[0037] For each station, a three-dimensional risk matrix is ​​constructed based on the anthropogenic pollution level determined by the anthropogenic enrichment factor calculated in step S4, the ecological risk level determined by the active state risk index calculated in step S5, and the confidence level output in step S6.

[0038] In the three-dimensional risk matrix, the level of anthropogenic pollution, the level of ecological risk, and the confidence level are each considered as a dimension. The final comprehensive risk rating for each station is generated through a preset decision matrix or a weighted summation method.

[0039] When using a weighted summation method, the weights of each dimension are determined by the analytic hierarchy process or the entropy weight method. Preferably, the weights of the anthropogenic pollution level are 0.35 to 0.45, the ecological risk level are 0.35 to 0.45, and the confidence level is 0.10 to 0.30, with the sum of the three being 1.

[0040] The comprehensive risk rating is divided into four levels: Level I, Level II, Level III, and Level IV. Level I represents low risk, Level II represents medium risk, Level III represents high risk, and Level IV represents extremely high risk.

[0041] Step S8: Determine the dynamic monitoring strategy.

[0042] Based on the comprehensive risk rating results from step S7, a spatial distribution map of the risk in the study area is drawn and visualized using a geographic information system.

[0043] The subsequent monitoring strategy will be determined based on the comprehensive risk rating results. For high-risk sites with comprehensive risk ratings of Level III and IV, maintain high-frequency chemical field measurements; For locations with low confidence levels, regardless of the risk rating, actual data must be supplemented to correct the model. For stations with a comprehensive risk rating of Level I and Level II and a high confidence level, the monitoring method will be changed to primarily using rapid spectral scanning, supplemented by quarterly spot checks.

[0044] Furthermore, it also includes a dynamic model update step: each newly acquired chemical measurement data is added to the training set, and the prediction model in step S6 is retrained periodically to achieve continuous optimization of the model.

[0045] Preferably, before step S4, a data quality control step is also included: outlier removal and missing value imputation are performed on the total heavy metal concentration data of sediments at each station obtained in step S2, outlier data other than the mean ± 3 times the standard deviation are removed, and missing values ​​are filled in using multiple imputation method.

[0046] Preferably, an independent model validation step is also included: 10% to 30% of the sampling stations are randomly selected as an independent validation set; the spectral data of the independent validation set is input into the prediction model trained in step S6; the predicted heavy metal concentration output by the model is compared with the measured values ​​in steps S2 and S3, and the coefficient of determination R is calculated. 2 and root mean square error, when R2 If the value is less than 0.7, return to step S6 to retrain the model or adjust the model structure. Beneficial effects

[0047] Compared with the prior art, the present invention has the following beneficial effects: This invention uses a positive definite matrix factorization model to analyze a unique dynamic background value for each station, replacing the fixed background value uniform across the region in traditional methods. This effectively distinguishes between high natural geological background and human input, avoiding misjudging naturally high background areas as polluted and preventing human input from being underestimated, thus significantly improving the accuracy of pollution level assessment.

[0048] This invention uses the concentration of active form extracted by Chelex-100 instead of the total amount to characterize ecological risk. The risk index calculated based on this can more accurately reflect the actual harm of heavy metals to the ecosystem.

[0049] This invention establishes a deep learning prediction model between visible-near-infrared spectral characteristics of sediments and heavy metal concentration, enabling the use of rapid and inexpensive spectral scanning to replace time-consuming and expensive traditional chemical analysis, thus making large-scale, high-frequency, and low-cost marine heavy metal monitoring possible.

[0050] This invention organically couples three dimensions: dynamic background calculation, bioavailability evaluation, and intelligent spectral prediction. It incorporates prediction reliability into the evaluation system through confidence levels, achieving a complete closed loop from data acquisition, model building, comprehensive evaluation to monitoring strategy optimization. Subsequent monitoring strategies are dynamically adjusted based on risk ratings and confidence levels, ensuring monitoring quality in high-risk and model-weak areas while optimizing monitoring costs in low-risk areas. Attached Figure Description

[0051] Figure 1 This is a schematic diagram of the overall process of the method for evaluating the degree of heavy metal pollution in the marine environment provided in an embodiment of the present invention.

[0052] Figure 2 This is a schematic diagram of the structure of a dual-scale long short-term memory network model with attention mechanism provided in an embodiment of the present invention.

[0053] Figure 3 This is a schematic diagram of the distribution of sampling stations in a research sea area, provided as an embodiment of the present invention. Detailed Implementation

[0054] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0055] Example: This embodiment uses a nearshore sea area as the research object and applies the method of this invention to comprehensively evaluate the degree of heavy metal pollution in the sediments of this sea area. The specific steps are as follows: Step 1, Site setup and sample collection Several sampling stations were deployed in a grid pattern in the research sea area, such as Figure 3 As shown, the spacing between stations is determined based on the area of ​​the study region and the distribution characteristics of pollution sources, and priority is given to covering the following three types of areas: 1) potential pollution source areas, including areas near sewage outlets, port and wharf areas, river estuaries, etc.; 2) ecologically sensitive areas, including aquaculture areas, protected areas, spawning grounds, etc.; 3) control areas, located in the open sea or clean areas far away from known pollution sources.

[0056] Spatial coordinates of each station were recorded using the Global Positioning System (GPS). At each station, surface sediment samples were collected using a grab sampler or box sampler at a depth of 0–5 cm. At least three parallel samples were collected from each station: one for total heavy metal determination, one for active form extraction determination, and one as a backup. The collected sediment samples were placed in acid-washed polyethylene sealed bags, air was removed, and the bags were sealed and stored at 4°C in the dark. They were then transported back to the laboratory for processing within 48 hours.

[0057] Visible-near-infrared reflectance spectral data of the sediment surface were collected simultaneously. Portable or laboratory ground-based spectrometers were used to collect reflectance spectra within the wavelength range of 350–2500 nm. For field measurements, clear weather and calm sea conditions were selected. The spectrometer probe was pointed vertically downwards at the sediment surface at a constant distance, and measurements were repeated at least 10 times at each station. The average value was taken as the raw spectral data for that station. Calibration was performed using a standard white board. Environmental parameters such as water depth, water temperature, salinity, and pH were recorded simultaneously at each station.

[0058] Step 2, determination of total heavy metal content in sediments Sediment samples used for total sediment determination were freeze-dried, and after drying, gravel and biodebris were removed. The samples were then ground in an agate mortar until they all passed through a 200-mesh nylon sieve. After mixing, the samples were placed in sealed bags for later use.

[0059] Sample pretreatment was performed using microwave digestion. A predetermined amount of sediment powder was accurately weighed and placed in a microwave digestion vessel. An optimized ratio of digestion acid, selected from at least two combinations of nitric acid, hydrochloric acid, hydrofluoric acid, and hydrogen peroxide, was added. The digestion vessel was tightly sealed and placed in a microwave digester, and digestion was carried out according to the set temperature program. After digestion, the mixture was cooled to room temperature, transferred to a polytetrafluoroethylene (PTFE) beaker, and the acid was removed to near dryness on a hot plate. The solution was then diluted to a predetermined volume with dilute nitric acid solution and shaken well before use. A blank sample was prepared simultaneously, using the same digestion and volume adjustment steps but without sediment, to subtract reagent blanks.

[0060] The concentration of the heavy metal element to be measured in the digestion solution was determined by inductively coupled plasma mass spectrometry (ICP-MS) or inductively coupled plasma atomic emission spectrometry (ICP-AES). The heavy metal elements to be measured were selected from at least five of the following: arsenic, cadmium, chromium, copper, mercury, lead, zinc, and nickel. The detection limits, precision, and accuracy of each element were controlled using standard reference materials, and the measured values ​​of the standard reference materials must be within the allowable range of the standard values.

[0061] Simultaneously, auxiliary parameters of the sediment samples were determined. Particle size distribution was measured using a laser particle size analyzer. The specific steps were as follows: an appropriate amount of sediment sample was added to a dispersant, ultrasonically vibrated, and then measured using the analyzer. The percentage content of each component (clay, silt, and sand) was output. Organic carbon content was determined using the potassium dichromate oxidation-external heating method: the sediment sample was accurately weighed, potassium dichromate standard solution and concentrated sulfuric acid were added, and the mixture was heated and digested in an oil bath. The remaining potassium dichromate was titrated with ferrous sulfate standard solution, and the organic carbon content was calculated based on the consumption. Iron and manganese oxide content was determined using a chemical extraction method, selectively extracting heavy metal components bound to iron and manganese oxides using a specific extractant.

[0062] Step 3, Determination of active heavy metal concentration The active concentrations of heavy metal elements in sediment samples from each station were determined using the Chelex-100 resin extraction method. The specific operational steps are as follows: 3.1) Pretreatment of Chelex-100 resin: Weigh an appropriate amount of Chelex-100 resin, wash it repeatedly with deionized water, then soak it in dilute hydrochloric acid solution and sodium hydroxide solution in turn, and finally wash it with artificial seawater until it is neutral.

[0063] 3.2) Active Form Extraction: Accurately weigh a predetermined amount of fresh sediment sample (dry weight), place it in a centrifuge tube, and add artificial seawater pretreated with Chelex-100 resin. The salinity of the artificial seawater should be consistent with the salinity at the sampling site. Tightly cap the tube and place it in a constant-temperature shaker. Shake at a constant speed for a predetermined time at the set temperature. After shaking, centrifuge the tube in a high-speed centrifuge and collect the supernatant.

[0064] 3.3) Determination: After acidifying the supernatant to the predetermined pH value with dilute nitric acid, the concentration of active metal in the supernatant was determined by inductively coupled plasma mass spectrometry.

[0065] Parameters such as extraction time, solid-liquid ratio, oscillation temperature, and rotation speed were optimized and determined through preliminary experiments based on the characteristics of sediments in the study area. The preliminary experimental method was as follows: sediments from representative stations were selected, and different extraction time gradients and solid-liquid ratio gradients were set. The metal concentration extracted under each condition was measured, and the parameters corresponding to the plateau period of the extracted concentration were taken as the optimized values.

[0066] Step 4: Dynamic background calculation based on the positive definite matrix factorization model The total heavy metal concentration measurements obtained in Step 2 were subjected to quality control and preprocessing. First, outliers were removed. Data from stations where the concentration of any element exceeded the average ± 3 standard deviations were marked as suspicious. After verifying the sampling and analysis records, a decision was made on whether to remove these values ​​or re-measure them. For a small number of missing values, multiple interpolation was used to fill them in. Specifically, a multiple interpolation model was established based on the concentration distribution of the same element at other stations and auxiliary parameters. After multiple interpolations, the average value was taken as the filled value.

[0067] The preprocessed total heavy metal concentration data and uncertainty data of each element from each station were input into a positive definite matrix factorization model. The model input parameters included the concentration values ​​of each element at each station and their corresponding uncertainties. During model operation, a range for the number of factors was set, and multiple iterative calculations were performed. The model's built-in residual analysis, Q-value test, and explained variance were used to evaluate the model's goodness of fit under different numbers of factors, and the optimal number of factors was selected.

[0068] After the model converges, it outputs the source component spectrum of each factor (i.e., the concentration contribution rate of each element in each factor) and the contribution concentration of each factor at each station. Factor types are identified by combining the characteristic elements of each factor. The identification rules are as follows: If aluminum, iron, and elements with geochemical behavior consistent with aluminum and iron (including chromium and nickel) have a high contribution rate in a certain factor, and the contribution concentration of the factor is evenly distributed in space and has no obvious correlation with the intensity of human activities, then the factor is identified as a factor representing the source of natural parent material. If one or more elements such as copper, lead, zinc, and arsenic have a high contribution rate in a certain factor, and the concentration of the factor's contribution shows a gradient distribution characteristic that decreases from the pollution source outwards in space, then the factor is identified as a factor from human activities. Different types such as industrial sources, agricultural sources, and transportation sources can be further distinguished based on the differences in characteristic elements. If the contribution rates of each element in a factor are low and there are no obvious characteristic elements, it is identified as a mixed or other source factor.

[0069] The contribution concentration values ​​of each element corresponding to the factor representing the natural source of the parent material were used as the dynamic background values ​​for each site. For sites identified as having no individual natural source factor, the spatial interpolation results of the contributions from neighboring natural source factors were used as the dynamic background values ​​for that site.

[0070] Based on this, the artificial enrichment factor for each element at each site is calculated using the following formula: EF PMF =C 实测 / C 本底 , where C 实测 C represents the measured total concentration of a certain element at this station. 本底 This represents the dynamic background concentration of a certain element at this station.

[0071] Based on the numerical range of the anthropogenic enrichment factor, the degree of anthropogenic pollution of each element at each station is divided into several levels. The specific classification criteria are as follows: when EF... PMF When the value is less than the first threshold, it is considered uncontaminated; when EF PMF When the value is greater than or equal to the first threshold and less than the second threshold, it is judged as slightly contaminated; when EF PMF When the level is greater than or equal to the second threshold and less than the third threshold, it is judged as moderate pollution; when EF PMF When the pollution level is greater than or equal to the third threshold, it is considered severely polluted. The first threshold is preferably 1, the second threshold is preferably 1.5, and the third threshold is preferably 3.

[0072] The overall pollution level of each station is determined by the highest pollution level among all elements at that station, or by the EF levels of each element. PMF The comprehensive value is determined by weighted summation of the values, and the specific method is selected according to the purpose of the evaluation.

[0073] Step 5, Calculation of the active state risk index Based on the active state concentrations of each element measured in step 3, the active state risk index (BRI) for each station is calculated using the formula: BRI = Σ(C 活性,i / TEL i )×W i , where C 活性,i This represents the measured active concentration of the i-th heavy metal element at this station; TEL i W represents the critical effect concentration of the i-th heavy metal element, i.e., the critical effect concentration value set in the sediment quality standard; i is the toxicity weighting coefficient of the i-th heavy metal element.

[0074] Toxicity weighting coefficient W i The method for determining it is as follows: Collect acute and chronic toxicity data of representative benthic species in the target sea area, including at least one species each of crustaceans, mollusks and polychaetes; The collected toxicity data were screened, and data that did not meet the quality control standards were removed. The toxicity data of each element after screening were logarithmically transformed to construct the species sensitivity distribution curve, and a preset distribution model was used for fitting. Extract the preset quantile values ​​corresponding to each element from the fitted species sensitivity distribution curve, normalize the quantile values ​​of each element so that the sum of the weights of each element is 1, and obtain the toxicity weight coefficient of each element.

[0075] It should be noted that when local species toxicity data are lacking, the toxicity response coefficient is temporarily used to replace the weighting coefficient, but this needs to be updated after data accumulation.

[0076] The ecological risk level is determined based on the numerical range of the Biotic Activity Index (BRI). The classification criteria are as follows: when the BRI is less than the first threshold, it is considered low ecological risk; when the BRI is greater than or equal to the first threshold and less than the second threshold, it is considered medium ecological risk; when the BRI is greater than or equal to the second threshold and less than the third threshold, it is considered high ecological risk; and when the BRI is greater than or equal to the third threshold, it is considered extremely high ecological risk. The first threshold is preferably 1, the second threshold is preferably 2, and the third threshold is preferably 3.

[0077] Optionally, the calculation of the active state risk index can also adjust the weight allocation method according to the evaluation purpose and regional characteristics, or use other equivalent mathematical forms to express the relationship between active state concentration and ecological risk.

[0078] Furthermore, the calculated active state risk index can be validated using in-situ bioaccumulation experiments. The specific validation method is as follows: Indicator organisms that are common in the target sea area and sensitive to environmental changes are selected. The indicator organisms are placed in water containing heavy metals labeled with stable isotopes for pre-exposure, so that a certain amount of labeled isotopes accumulate in the organisms. The pre-exposed organisms are placed in a net cage and suspended at a selected station for a predetermined underwater depth and exposure period. Biological samples were recovered at different time points during the exposure period, and the changes in the content of labeled isotopes in the organisms were measured to calculate the bioaccumulation rate and excretion rate. Correlation analysis was performed between bioaccumulation kinetic parameters and the bioactive state risk index at the same site. If the two showed a significant positive correlation, the effectiveness of the bioactive state risk index was verified.

[0079] Step 6: Constructing a prediction model based on spectral data and long short-term memory networks. The visible-near-infrared reflectance spectral data acquired in step 1 are preprocessed. The preprocessing steps include: 6.1) Average the spectral data from multiple repeated measurements to obtain the average spectral curve; 6.2) The Savitzky-Golay convolution smoothing method is used to smooth and denoise the average spectral curve, and an appropriate window width and polynomial order are set. 6.3) Use standard normal variable transformation or multivariate scattering correction for spectral correction to eliminate the influence of particle size and surface scattering effects; 6.4) First-order or second-order differential transforms are used to process the corrected spectral data to highlight spectral characteristic peaks and eliminate the effects of baseline drift. The preprocessed spectral data forms a standardized spectral dataset.

[0080] The preprocessed spectral data was used as the model input features, and the total heavy metal concentrations of each element measured in step 2 and the active state concentrations of each element measured in step 3 were used as training labels. The dataset was divided into training, validation, and test sets according to a preset ratio. Stratified random sampling was used for the partitioning to ensure that samples of different concentration gradients were distributed in each dataset.

[0081] A dual-scale long short-term memory network model with an attention mechanism is constructed. The model structure is as follows: Figure 2 As shown, the specific design is as follows: The input layer receives preprocessed spectral data, with the data dimension being the number of spectral bands.

[0082] The dual-scale feature extraction module includes parallel first-scale extraction units and second-scale extraction units. The first-scale extraction unit consists of a first long short-term memory (LSM) network layer containing a predetermined number of hidden units and a small time step, used to extract local features from the spectral data. These local features reflect spectral absorption characteristics within a specific narrow spectral band. The second-scale extraction unit consists of a second LSM network layer containing a predetermined number of hidden units and a larger time step, used to extract global features from the spectral data. These global features reflect spectral variation trends over a wide spectral band. The outputs of the first and second LSM network layers are regularized using dropout layers to prevent overfitting.

[0083] The attention mechanism module receives the feature matrix output by the dual-scale feature extraction module and calculates the attention weight for each feature vector using an attention scoring function based on a multilayer perceptron. The calculated attention weights are then normalized using a softmax function to obtain a normalized attention weight distribution. Finally, the normalized attention weights are weighted and summed with the original feature matrix to obtain a weighted feature vector. This attention mechanism enables the model to automatically identify the spectral feature bands that contribute most to heavy metal concentration prediction.

[0084] The fully connected layer receives the weighted feature vector output by the attention mechanism module, performs feature transformation through at least one fully connected hidden layer, and finally outputs the predicted concentration values ​​of each heavy metal element using a linear activation function.

[0085] The model training employs a pre-defined optimization algorithm, using root mean square error or mean absolute error as the loss function, and utilizes a learning rate decay strategy and an early stopping mechanism to optimize training performance. During training, a validation set is used to monitor the model's generalization performance, and training is stopped when the validation set loss no longer decreases for several consecutive rounds.

[0086] After the model is trained, it is independently evaluated using a test set. Evaluation metrics include the coefficient of determination (R²). 2 Root mean square error and mean absolute error. When the coefficient of determination R... 2 If the value falls below a preset threshold, the system will return to the previous state and retrain after adjusting the model structure or supplementing training data.

[0087] After the model training is completed and validated, the spectral data of each station is input into the model, and the predicted values ​​of the total concentration and active state concentration of each element of heavy metals at each station are output.

[0088] Meanwhile, the confidence level of the predicted values ​​for each station was determined using the following method: Based on the prediction error distribution of the test set samples, for each predicted value, a prediction interval is calculated using either the Monte Carlo dropout method or the Bayesian approximation method. The prediction interval characterizes the range of uncertainty of the predicted value. The confidence level is divided into several levels according to the relative width of the prediction interval, i.e., the ratio of the prediction interval width to the prediction mean. Specifically, a prediction interval relative width less than a first preset threshold is considered high confidence, a prediction interval relative width greater than a second preset threshold is considered low confidence, and a width in between is considered medium confidence. The first and second preset thresholds are determined based on model accuracy requirements and monitoring objectives.

[0089] Step 7: Construction of a three-dimensional risk matrix and comprehensive risk rating For each station, the following three dimensions of evaluation information are integrated: 1) Level of human-caused pollution: derived from the level of human-caused pollution determined in step 4; 2) Ecological risk level: derived from the ecological risk level determined in step 5; 3) Confidence level: The confidence level determined in step 6.

[0090] A three-dimensional risk matrix is ​​constructed, with the three dimensions being the level of anthropogenic pollution, the level of ecological risk, and the confidence level. The final comprehensive risk rating for each station is generated through table lookup or calculation using the three-dimensional risk matrix.

[0091] The specific method for generating the comprehensive risk rating is selected from one of the following: Method 1, Pre-set Decision Matrix Method: A comprehensive risk rating rule is pre-defined based on different combinations of anthropogenic pollution level, ecological risk level, and confidence level. Each station is then rated according to this rule. Specifically, when both the anthropogenic pollution level and ecological risk level are low and the confidence level is high, the rating is low risk; when at least one of the anthropogenic pollution level or ecological risk level is high and the confidence level is high, the rating is high risk; when the confidence level is low, the comprehensive risk rating is increased by one level from the original risk level.

[0092] Method 2, Weighted Summation Method: This method assigns values ​​to the degree of anthropogenic pollution, the ecological risk level, and the confidence level. These values ​​are then weighted and summed. The overall risk rating is determined based on the numerical range corresponding to the weighted sum. The weights for each dimension are determined using the analytic hierarchy process (AHP) or the entropy weight method, and are set by those skilled in the art based on the specific regional characteristics.

[0093] The comprehensive risk rating is divided into several levels, including at least four levels: low risk, medium risk, high risk, and extremely high risk.

[0094] Step 8: Determine the dynamic monitoring strategy Based on the comprehensive risk rating results from step 7, a spatial distribution map of the risk in the study area is drawn using a geographic information system. Specific steps include: using the geographic coordinates of each station as location information and the comprehensive risk rating of each station as attribute information, a continuous risk distribution surface for the entire study area is generated using spatial interpolation methods, presenting the spatial distribution pattern of risk levels in different colors or contour lines.

[0095] Based on the comprehensive risk rating results, the subsequent monitoring strategies for each station are determined, including: 8.1) For sites with a comprehensive risk rating of high risk and very high risk, maintain high-frequency chemical field monitoring, that is, recollect sediment samples and perform field analysis in steps 2 to 5 according to the predetermined cycle. The cycle is preferably once a month or once a quarter. 8.2) For stations with low confidence levels, regardless of the overall risk rating, actual measurement data must be added in the next monitoring cycle to verify and correct the output of the prediction model. 8.3) For stations with a comprehensive risk rating of low risk and medium risk and a high confidence level, the monitoring method will be changed to a method that mainly uses rapid spectral scanning and is supplemented by regular sampling. That is, spectral data will be collected and the concentration prediction value will be output through the prediction model in step 6. Chemical sampling will be carried out only at predetermined time intervals to verify the stability of the model.

[0096] Furthermore, the model includes a dynamic update step: each newly acquired chemical measurement data and corresponding spectral data are added to the training set in step 6, and the prediction model is retrained according to a preset time period to achieve continuous optimization of model parameters. As data accumulates, the prediction accuracy of the prediction model at each station and concentration range gradually improves, and correspondingly, the coverage of high-confidence stations gradually expands, further reducing the workload of chemical measurements.

[0097] Furthermore, it also includes an independent, periodic model validation step: after each model update, all data are randomly sampled into a test set at a predetermined ratio, and the updated model is independently evaluated to calculate the coefficient of determination R. 2 And root mean square error. When the coefficient of determination R 2 If the accuracy falls below a preset threshold, the model is restructured and retrained until the required accuracy is met. The preset threshold is preferably 0.7.

[0098] The method of this invention can be applied to the investigation, monitoring, and assessment of heavy metal pollution in sediments in various nearshore sea areas, bays, estuaries, and other marine environments. It can also provide technical support for marine environmental management departments in pollution classification, risk warning, and monitoring optimization decision-making. This method has a clear process, standardized operation, and is easy to promote, exhibiting good industrial applicability.

[0099] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely illustrative of the principles of the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the present invention as claimed. The scope of protection of this invention is defined by the appended claims and their equivalents.

Claims

1. A method for evaluating heavy metal pollution in marine environments, characterized in that, include: By collecting surface sediment samples and their visible-near-infrared reflectance spectral data at several sampling stations set up in the study area, the total concentration of heavy metal elements, auxiliary parameters, active state concentration, anthropogenic enrichment factor and active state risk index in the samples were determined. Using preprocessed spectral data as input features and measured total heavy metal concentration and active state concentration as labels, a dual-scale long short-term memory network model with attention mechanism is trained to establish a prediction model between spectral features and heavy metal concentration, and the confidence level of the predicted value of each station is output. The anthropogenic pollution level is determined based on the anthropogenic enrichment factor, the ecological risk level is determined based on the active state risk index, and a three-dimensional risk matrix is ​​constructed by combining the confidence level to generate a comprehensive risk rating for each station.

2. The method for evaluating heavy metal pollution in marine environments according to claim 1, characterized in that, The heavy metal element is selected from at least five of the following: arsenic, cadmium, chromium, copper, mercury, lead, zinc, and nickel; the auxiliary parameters include particle size distribution, organic carbon content, and iron and manganese oxide content.

3. The method for evaluating heavy metal pollution in marine environments according to claim 1, characterized in that, The procedure for determining the artificial enrichment factor includes: Input the total heavy metal concentration data of each station into the positive definite matrix factor decomposition model to analyze the source factors of heavy metals at each station; Factors representing the natural source of the parent material were identified, and the contribution concentrations of each element corresponding to these factors were used as the dynamic background values ​​for each site. Based on this, the anthropogenic enrichment factor EF for each site was calculated. PMF The calculation formula is: EF PMF =Measured concentration / Dynamic background concentration; Where, when EF PMF When <1, it is considered uncontaminated; when 1≤EF PMF A value less than 1.5 is considered slightly polluted; when 1.5 ≤ EF... PMF A value <3 indicates moderate pollution. PMF A value of ≥3 indicates severe pollution.

4. The method for evaluating heavy metal pollution in marine environments according to claim 1, characterized in that, The formula for calculating the Active State Risk Index (BRI) is: BRI = Σ(C 活性,i / TEL i )×W i C 活性,i Let TEL be the active state concentration of the i-th heavy metal element. i Let W be the critical effect concentration of the i-th heavy metal element. i Let be the toxicity weight coefficient of the i-th heavy metal element derived from the species sensitivity distribution.

5. The method for evaluating heavy metal pollution in marine environments according to claim 1, characterized in that, The dual-scale long short-term memory network model with attention mechanism includes a dual-scale feature extraction module and an attention mechanism module. The dual-scale feature extraction module is used to extract local features and global features from the spectral data respectively. The attention mechanism module is used to weight and focus the extracted features and automatically identify the spectral feature bands that contribute the most to the prediction of heavy metal concentration.

6. The method for evaluating heavy metal pollution in marine environments according to claim 1, characterized in that, The preprocessing includes denoising, smoothing, baseline correction, and first-order or second-order differential transformation of the visible-near-infrared reflectance spectral data.

7. The method for evaluating heavy metal pollution in marine environments according to claim 1, characterized in that, The confidence level is divided based on the width of the prediction interval output by the prediction model: when the prediction interval width is less than the first preset threshold, it is determined to be high confidence; when the prediction interval width is greater than the second preset threshold, it is determined to be low confidence; and when it is in between, it is determined to be medium confidence.

8. The method for evaluating heavy metal pollution in marine environments according to claim 1, characterized in that, In the three-dimensional risk matrix, the anthropogenic pollution level, ecological risk level, and confidence level are each used as three dimensions. The final comprehensive risk rating is generated through a preset decision matrix or a weighted summation method. The comprehensive risk rating is divided into Level I, Level II, Level III, and Level IV.

9. The method for evaluating heavy metal pollution in marine environments according to claim 1, characterized in that, The method also includes: drawing a risk spatial distribution map of the study sea area based on the comprehensive risk rating results, and determining the subsequent monitoring strategy: maintaining high-frequency chemical measurements for high-risk stations and stations with confidence levels below a preset threshold; and switching to a monitoring method that mainly uses rapid spectral scanning supplemented by random sampling for low-risk stations.

10. The method for evaluating heavy metal pollution in marine environments according to claim 1, characterized in that, The method further includes: randomly selecting 10% to 30% of the sampling stations from all sampling stations as an independent validation set; inputting the spectral data of the independent validation set into the trained prediction model; comparing the predicted heavy metal concentration values ​​output by the model with the measured values; and calculating the coefficient of determination R. 2 and root mean square error, when R 2 If the value is less than 0.7, retrain the model or adjust the model structure.