Toxicity risk evaluation method for low-concentration composite pollutants

By exposing fish gill cells to sewage samples and establishing a water quality toxicity risk model using multi-parameter phenotype spectrum and machine learning model, the complexity problem of water quality toxicity risk assessment is solved, and a more comprehensive and accurate toxicity risk assessment is achieved.

CN120220892APending Publication Date: 2025-06-27NANJING UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510338696.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-21
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The prior art is difficult to effectively evaluate the toxicity risk of low-concentration composite pollutants on water quality, especially under the complex combined toxicity effect, and traditional methods are difficult to fully reflect the toxicity risk of water quality.

Method used

By placing fish gill cells in filtered sewage samples and exposing the staining, combining mixed fluorescent stain labeling and high-connotation automatic imaging analysis system to obtain multi-parameter phenotype spectrum, unsupervised learning clustering and machine learning models (such as gradient lifting trees) to establish a water quality toxicity risk model, and then evaluate the toxicity risk level of sewage.

Benefits of technology

This method can more comprehensively and accurately reflect the toxicity risk of water quality, have the ability to learn dynamically and update data, and can provide timely scientific basis for water quality management and decision-making, and solve the problems of limited application scope, incomplete evaluation and inability to evaluate toxicity risk levels.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120220892A_ABST
    Figure CN120220892A_ABST
Patent Text Reader

Abstract

The invention discloses a toxicity risk evaluation method for low-concentration composite pollutants. The method comprises the following steps: placing gill cells in a filtered sewage sample for exposure contamination; the method comprises the following steps: marking cell structures of gill cells by using a mixed fluorescent staining agent, and carrying out high-throughput shooting on stained and marked cell images by using a high-content automatic imaging analysis system to obtain a multi-parameter phenotype spectrum; unsupervised learning clustering is carried out on the multi-parameter phenotype spectrum to form different clusters to complete dimension reduction, the sewage toxicity risk level is judged through machine learning model training, and a gradient boosting tree is finally selected through a hyper-parameter optimization model to establish a water quality toxicity risk model; inputting the multi-parameter phenotype spectrum into the water quality toxicity risk model, and screening by clustering categories to obtain a toxicity risk grade; the method can effectively and comprehensively evaluate the toxicity risk of the low-concentration composite pollutants in the effluent of the sewage treatment plant, and has practical significance for guiding the water quality risk control of the sewage treatment plant.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for evaluating toxicity risk, and particularly to a method for evaluating the toxicity risk of low-concentration complex pollutants. Background Art

[0002] With the acceleration of industrialization and urbanization, the toxic effects of low-concentration complex pollutants existing in water bodies on water quality have attracted increasing attention. Research shows that various low-concentration pollutants generally exhibit complex combined toxic effects such as addition, synergy, or antagonism in water bodies, and the toxicity is significantly higher than that of single pollutants. For example, when certain heavy metal ions coexist with other pollutants, they may exhibit completely different toxicity characteristics from when they exist alone. Sewage containing low-concentration complex pollutants not only has a complex composition but also requires high monitoring accuracy. Currently, commonly used physical and chemical indicators or single-pollutant concentration assessment methods are difficult to comprehensively reflect the toxicity risk of sewage, resulting in a relatively high uncertainty in water quality toxicity risk assessment.

[0003] Aquatic organism toxicity testing is the most intuitive and comprehensive method for evaluating the toxicity risk of sewage. The US Environmental Protection Agency uses zebrafish as an important biological model for water quality toxicity determination. However, with the increasing attention to animal experiment ethics and the promotion of scientific progress, regulatory authorities have clearly emphasized reducing the dependence on vertebrate testing and actively advocated using NAMs methods to replace traditional animal experiments. In the field of aquatic toxicology, models such as cell-based cytotoxicity tests and fish embryo acute toxicity (FET) tests have been proven to be effective alternatives and improvements to fish toxicity tests. In the international standard ISO21115:2019 "Water quality - Determination of the acute toxicity of water samples and chemicals to the rainbow trout gill cell line (RTgill-W1)", the permanent cell line of rainbow trout gill (RTgill-W1) is used as the test organism and cultured in a monolayer in a 24-well tissue culture plate. Then, these cells are exposed to water samples or chemicals for 24 hours. After the exposure ends, the cell viability is evaluated through a fluorescent cell viability indicator dye, and the toxicity is quantified based on the response curve between the cell viability and the concentration of the water sample or chemical. However, this method has certain limitations and is difficult to achieve the toxicity assessment of low-concentration complex pollutants; moreover, the biological characteristics based on which the biological response is calculated are relatively single and cannot comprehensively reflect the complexity of water quality toxicity risk; in addition, the existing method can only determine whether a water sample has potential toxicity and cannot give a clear grade evaluation of the toxicity risk of the water sample. Summary of the Invention

[0004] Object of the Invention: The object of the present invention is to provide a method for evaluating the toxicity risk of low-concentration complex pollutants to comprehensively evaluate the characteristics and evaluate the water quality toxicity risk level of low-concentration complex pollutants.

[0005] Technical solution: A method for evaluating the toxicity risk of low-concentration composite pollutants, comprising the following steps:

[0006] S1. Expose and poison fish gill cells in the filtered sewage sample.

[0007] S2. Use a mixed fluorescent stain to label the cell structure of fish gill cells, and obtain a multi-parameter phenotypic profile by high-throughput photographing of the stained and labeled cell images through a high-content automated imaging analysis system.

[0008] S3. Perform unsupervised learning clustering on the multi-parameter phenotypic profile to form different clusters for dimensionality reduction, determine the sewage toxicity risk level through machine learning model training, optimize the model through hyperparameters, and finally select a gradient boosting tree to establish a water quality toxicity risk model.

[0009] S4. Input the multi-parameter phenotypic profile described in S2 into the water quality toxicity risk model described in S3, and obtain the toxicity risk level by screening the clustering categories.

[0010] Preferably, the sewage sample in S1 is filtered through one of a 0.22 μm cellulose acetate filter membrane, a nitrocellulose filter membrane, and a mixed cellulose ester filter membrane.

[0011] Preferably, the fish gill cells in S1 are exposed and poisoned for 24 to 48 h.

[0012] Preferably, the mixed fluorescent stain in S2 is composed of 3 to 10 μg / mL Hoechst 33342 dye, 6 to 12 μM SYTO14 dye, 20 nM to 500 nM Mito Tracker Deep Red dye, 5 to 20 μL / mL concanavalin A / AlexaFluor 488 dye, 1 to 3 μg / mL Wheat-germ Agglutinin AlexaFluor 555 conjugate dye, and 3 to 12 μL / mL Phalloidin Alexa Fluor 568 conjugate dye.

[0013] Preferably, the shooting conditions of the high-content automated imaging analysis system in S2 are: synchronously shoot in 5 fluorescence channels of DNA, RNA, ER, AGP, and Chl using a 20x water immersion objective lens, and obtain the stained and labeled cell images in the form of 4 to 16 imaging fields / well, z-axis layer stacking, optimal focus projection, and confocal imaging.

[0014] Preferably, the process of obtaining the multi-parameter phenotypic spectrum in S2 is as follows: Select the grayscale image taken in bright field for calibration; Segment the cell image by one of the algorithms of image threshold segmentation, edge segmentation, morphological watershed segmentation or region-based segmentation algorithm, delimit the ROI region to separate the cells from the background, and then extract the biological parameters of each cell. The biological parameters include fluorescence intensity parameters, cell morphological feature parameters, and cell texture feature parameters, so as to obtain the multi-parameter phenotypic spectrum.

[0015] Preferably, after the multi-parameter phenotypic spectrum data in S3 is subjected to standardization processing and PCA dimensionality reduction, unsupervised clustering is performed. The Agglomerative Clustering algorithm is used for hierarchical clustering, with the Euclidean distance as the distance metric and the Ward criterion as the linkage criterion.

[0016] Preferably, after obtaining the multi-parameter phenotypic spectrum data in S3, it is divided into a training set, a validation set, and a test set. Define the model and parameter grid for training, select the optimal parameters through grid search to retrain the model and generate a classification report, and the report results are presented in terms of accuracy, recall rate, and F1 score.

[0017] Preferably, the machine learning models in S3 include logistic regression, gradient boosting tree, random forest, support vector machine, K-nearest neighbor, decision tree, XGBoost, and naive Bayes.

[0018] Preferably, the toxicity risk levels in S4 include low risk, medium risk, and high risk.

[0019] Beneficial effects: Compared with the prior art, the present invention has the following remarkable advantages: 1. By establishing a water quality risk model, it can better capture the complex interactions between water quality parameters, introduce machine learning to automatically identify patterns and relationships in the data, and has the ability of dynamic learning and data update, which can provide timely scientific basis for water quality management and decision-making, thus solving the problems of limited application scope, incomplete evaluation, and inability to evaluate the water quality toxicity risk level in the existing methods; 2. Utilize the biological reactions of fish gill cells to sewage at the subcellular structure level, obtain a large number of cell biological parameter index characteristics, and screen high-dimensional feature data through machine learning algorithms to obtain effective information, which can more comprehensively and accurately reflect the toxicity risk of water quality; 3. Obtain the biotoxicity effect characteristics of fish gill cells after exposure and poisoning through a high-throughput automated platform. The rich biological information data expands the application scope and enables the toxicity risk assessment of low-concentration complex pollutants to be realized. Description of the Drawings

[0020] Figure 1 It is a schematic flow chart of the present invention;

[0021] Figure 2 It is a schematic diagram of the hierarchical clustering screening results in Embodiment 1 of the present invention;

[0022] Figure 3 Schematic diagram of the subcellular and organelle structures of the fish gill cells in Example 1 of the present invention;

[0023] Figure 4 Schematic diagram of the subcellular and organelle structures of the fish gill cells in Example 2 of the present invention;

[0024] Figure 5 Schematic diagram of the subcellular and organelle structures of the fish gill cells in Example 3 of the present invention. Detailed implementation manners

[0025] The technical solution of the present invention will be described in detail below with reference to the accompanying drawings.

[0026] Example 1

[0027] The application object of this example is the effluent of a municipal sewage treatment plant located in Jilin Province. The daily treatment capacity of this plant is 10 cubic meters per day, and the effluent contains 44.93 mg / L of COD, 12.21 mg / L of total nitrogen, and 1.12 mg / L of total phosphorus. The toxicity risk assessment method for low-concentration composite pollutants has the following specific steps:

[0028] S1. Select fish gill cells of the RTgill-W1 strain, and expose the fish gill cells to the effluent sewage sample filtered through a 0.22 μm filter membrane for 24 h.

[0029] S2. Prepare a mixed fluorescent staining agent using 3 μg / mL Hoechst 33342 dye, 6 μM SYTO14 dye, 100 nM Mito Tracker Deep Red dye, 5 μL / mL concanavalin A / Alexa Fluor 488 dye, 3 μL / mL Phalloidin Alexa Fluor 568 conjugate dye, and 1.5 μg / mL Wheat-germ Agglutinin Alexa Fluor 555 conjugate dye, and use the prepared mixed fluorescent staining agent to label the cell structure of the fish gill cells; use a 20-fold water immersion objective lens to synchronously capture images in 5 fluorescence channels of DNA, RNA, ER, AGP, and Chl, and obtain the cell images after staining and labeling in the manner of 9 imaging fields / well, z-axis layer stacking, optimal focus projection, and confocal imaging. The subcellular and organelle structure images are as Figure 3 shown;

[0030] Select the grayscale cell image taken under bright field conditions for image correction; segment the cell image using one of the algorithms of image threshold segmentation, edge segmentation, morphological watershed segmentation, or region-based segmentation algorithm, delimit the ROI region to separate the cells from the background, and then extract the biological parameters of each cell, including fluorescence intensity parameters, cell morphological feature parameters, and cell texture feature parameters.

[0031] S3. Establish a water quality toxicity risk model:

[0032] Through preliminary experiments, obtain the multi-parameter phenotypic spectra of 128 effluent samples from municipal wastewater treatment plants. After standardizing the data and performing PCA dimensionality reduction, perform unsupervised clustering. Use the Agglomerative Clustering algorithm for hierarchical clustering. Use the Euclidean distance as the distance metric and the Ward linkage as the linkage criterion to maintain the compactness and consistency of the clustering; the clustering results form 10 different clusters as Figure 1 shown, that is, each wastewater sample corresponds to ten-dimensional eigenvalue as the input value for subsequent machine learning;

[0033] Select eight machine learning models of logistic regression, gradient boosting tree, random forest, support vector machine, K-nearest neighbor, decision tree, XGBoost, and naive Bayes to train and determine the sewage toxicity risk level; after loading the eigenvalues of the sewage samples, stratify and divide the training set, validation set, and test set, define the model and parameter grid for training; select the optimal parameters through grid search to retrain the model and generate a classification report. The classification report results of each model are presented in terms of accuracy, recall rate, and F1 score. According to the results, select the gradient boosting tree as the optimal model. After inputting the ten-dimensional eigenvalue, the output risk level result is low risk, medium risk, or high risk.

[0034] S4. The eigenvalues obtained after unsupervised learning clustering screening of the multi-parameter phenotypic spectra of the effluent samples from a municipal wastewater treatment plant in Jilin Province are: -2.0103262, -3.229465671, -3.49524468, -6.616062575, -13.25572898, 0.510507696, -0.692938068, -4.391582289, -19.06649167, -0.404621519.

[0035] Input the eigenvalues into the toxicity risk model to obtain the toxicity risk level result of the effluent as: low risk.

[0036] Example 2

[0037] The application object of this embodiment is the effluent of a municipal sewage treatment plant located in Hubei Province. The daily treatment capacity of this plant is 1 million cubic meters per day, and the effluent contains 45.09 mg / L of COD, 9.93 mg / L of total nitrogen, and 1.18 mg / L of total phosphorus. The toxicity risk assessment method for low-concentration composite pollutants has the following specific steps:

[0038] S1. Select RTgill-W1 strain of gill cells, and expose the gill cells to the effluent sewage sample filtered through a 0.22 μm filter membrane for 24 hours.

[0039] S2. Prepare a mixed fluorescent staining agent using 3 μg / mL Hoechst 33342 dye, 6 μM SYTO14 dye, 100 nM Mito Tracker Deep Red dye, 5 μL / mL concanavalin A / Alexa Fluor 488 dye, 3 μL / mL Phalloidin Alexa Fluor 568 conjugate dye, and 1.5 μg / mL Wheat-germ Agglutinin Alexa Fluor 555 conjugate dye, and use the prepared mixed fluorescent staining agent to label the cell structure of the gill cells; use a 20x water immersion objective lens to synchronously capture images in 5 fluorescence channels of DNA, RNA, ER, AGP, and Chl, and obtain the cell images after staining and labeling in the confocal imaging mode with 9 imaging fields / well, z-axis layer stacking, optimal focus projection. The subcellular and organelle structure images are as Figure 4 shown;

[0040] Select the grayscale cell image captured under bright field conditions for image correction; segment the cell image by one of the algorithms of image threshold segmentation, edge segmentation, morphological watershed segmentation, or region-based segmentation algorithm, delimit the ROI area to separate the cells from the background, and then extract the biological parameters of each cell, including fluorescence intensity parameters, cell morphological feature parameters, and cell texture feature parameters.

[0041] S3. Establish a water quality toxicity risk model:

[0042] For the multi-parameter phenotypic spectrum obtained through preliminary experiments, standardize the data, perform PCA dimensionality reduction, and then perform unsupervised clustering. Use the Agglomerative Clustering algorithm for hierarchical clustering. Use the Euclidean distance as the distance metric and the Ward linkage as the linkage criterion to maintain the compactness and consistency of the clustering; the clustering results form 10 different clusters as Figure 1 shown, that is, each sewage sample corresponds to ten-dimensional eigenvalue as the input value for subsequent machine learning;

[0043] Eight machine learning models, namely logistic regression, gradient boosting tree, random forest, support vector machine, K-nearest neighbor, decision tree, XGBoost and naive Bayes, are selected to train and determine the sewage toxicity risk level; after loading the eigenvalue of the sewage sample, the training set, validation set and test set are stratified and divided, and the model and parameter grid are defined for training; the optimal parameters are selected through grid search to retrain the model and generate a classification report. The classification report results of each model are presented by accuracy, recall rate and F1 score. According to the results, the gradient boosting tree is selected as the optimal model. After inputting the ten-dimensional eigenvalue, the output risk level result is low risk, medium risk or high risk.

[0044] S4. Conduct experiments on the effluent samples of a municipal sewage treatment plant in Hubei Province. The eigenvalues after unsupervised learning clustering screening of the multi-parameter phenotypic spectrum are: 6.285298055, -2.162563324, -3.491549387, -7.714040249, -17.27115125, -6.747862588, -4.368943918, -3.907625065, -19.06477471, -9.004467244.

[0045] Input the eigenvalue into the toxicity risk model, and the obtained toxicity risk level result of the effluent is: medium risk.

[0046] Example 3

[0047] The application object of this example is the effluent of a municipal sewage treatment plant located in Hebei. The daily treatment capacity of a municipal sewage treatment plant located in Hebei is 300,000 cubic meters per day, and the effluent contains 225.18 mg / L of COD, 25.79 mg / L of total nitrogen, 2.86 mg / L of total phosphorus. The specific steps of the toxicity risk assessment method for low-concentration complex pollutants are as follows:

[0048] S1. Select the fish gill cells as the RTgill-W1 strain, and expose the fish gill cells to the effluent sewage sample filtered through a 0.22 μm filter membrane for 24 h.

[0049] S2. Prepare a mixed fluorescent staining agent using 3 μg / mL Hoechst 33342 dye, 6 μM SYTO14 dye, 100 nM Mito Tracker Deep Red dye, 5 μL / mL concanavalin A / Alexa Fluor 488 dye, 3 μL / mL Phalloidin Alexa Fluor 568 conjugate dye, and 1.5 μg / mL Wheat-germ Agglutinin Alexa Fluor 555 conjugate dye, and use the prepared mixed fluorescent staining agent to label the cell structures of fish gill cells; synchronously capture images in 5 fluorescence channels of DNA, RNA, ER, AGP, and Chl using a 20x water immersion objective lens, and obtain cell images after staining and labeling in the manner of 9 imaging fields / well, z-axis layer stacking, optimal focus projection, and confocal imaging. Subcellular and organelle structure images are as Figure 5 shown;

[0050] Select the grayscale cell image captured under bright field conditions for image correction; segment the cell image using one of the algorithms of image threshold segmentation, edge segmentation, morphological watershed segmentation, or region-based segmentation algorithm, delimit the ROI region to separate the cells from the background, and then extract the biological parameters of each cell, including fluorescence intensity parameters, cell morphological feature parameters, and cell texture feature parameters.

[0051] S3. Establish a water quality toxicity risk model:

[0052] For the multi-parameter phenotypic spectra obtained from previous experiments, standardize the data, perform PCA dimensionality reduction, and then perform unsupervised clustering. Use the Agglomerative Clustering algorithm for hierarchical clustering. Use the Euclidean distance as the distance metric and the Ward linkage as the linkage criterion to maintain the compactness and consistency of the clustering; the clustering results form 10 different clusters as Figure 1 shown, that is, each sewage sample corresponds to ten-dimensional eigenvalue as the input value for subsequent machine learning;

[0053] Eight machine learning models, namely logistic regression, gradient boosting tree, random forest, support vector machine, K-nearest neighbor, decision tree, XGBoost, and naive Bayes, are selected to train and determine the sewage toxicity risk level. After loading the eigenvalue of the sewage sample, the training set, validation set, and test set are stratified and divided. The model and parameter grid are defined for training. The optimal parameters are selected through grid search to retrain the model and generate a classification report. The classification report results of each model are presented in terms of accuracy, recall, and F1 score. According to the results, the gradient boosting tree is selected as the optimal model. After inputting the ten-dimensional eigenvalue, the output risk level result is low risk, medium risk, or high risk.

[0054] S4. The eigenvalues obtained after unsupervised learning clustering screening of the multi-parameter phenotypic spectra from the effluent samples of a municipal sewage treatment plant in Hebei Province are: 7.697818738, -11.73713852, -4.454121971, -7.44308021, -12.29608172, -9.598931099, -3.014325632, -4.803087544, -19.00803381, 0.422383606.

[0055] The eigenvalue is input into the toxicity risk model, and the obtained toxicity risk level result of the effluent is: high risk.

Claims

1. A method for toxicity risk assessment of low-concentration composite pollutants, characterized in that: The following steps are involved: S1. Expose fish gill cells to toxic substances in filtered sewage samples; S2, using mixed fluorescent dyes to mark the cell structure of fish gill cells, and using a high-content automatic imaging analysis system to take high-throughput images of the stained cells to obtain a multi-parameter phenotypic spectrum; S3, clustering the multi-parameter phenotype spectrum by unsupervised learning to form different clusters to complete dimensionality reduction, determining the sewage toxicity risk level through machine learning model training, optimizing the model through hyperparameters, and finally selecting the gradient boosting tree to establish a water quality toxicity risk model; S4. Input the multi-parameter phenotype spectrum described in S2 into the water quality toxicity risk model described in S3, and obtain the toxicity risk level through cluster category screening.

2. The toxicity risk assessment method according to claim 1, characterized in that: The sewage sample described in S1 is filtered through a 0.22 μm acetate fiber filter membrane, a nitrocellulose filter membrane, or a mixed cellulose ester filter membrane.

3. The toxicity risk assessment method according to claim 1, characterized in that: The fish gill cells described in S1 were exposed to the poison for 24 to 48 hours.

4. The toxicity risk assessment method according to claim 1, characterized in that: The mixed fluorescent dye S2 consists of 3-10 μg / mL Hoechst 33342 dye, 6-12 μM SYTO14 dye, 20 nM-500 nM Mito Tracker DeepRed dye, 5-20 μL / mL concanavalinA / Alexa Fluor 488 dye, 1-3 μg / mL Wheat-germAgglutinin Alexa Fluor 555 conjugate dye and 3-12 μL / mL Phalloidin Alexa Fluor568 conjugate dye.

5. The toxicity risk assessment method according to claim 1, characterized in that: The shooting conditions of the high-content automatic imaging analysis system described in S2 are: using a 20x water immersion objective to simultaneously shoot in the five fluorescent channels of DNA, RNA, ER, AGP and Chl, and obtaining the stained and labeled cell images with 4 to 16 imaging fields / wells, z-axis layer stacking, optimal focus projection, and confocal imaging.

6. The toxicity risk assessment method according to claim 1, characterized in that: The process of obtaining the multi-parameter phenotypic spectrum described in S2 is: selecting a grayscale image taken in the bright field as correction; segmenting the cell image through an algorithm selected from image threshold segmentation, edge segmentation, morphological watershed segmentation or a region-based segmentation algorithm, defining the ROI area to separate the cells from the background, and then extracting the biological parameters of each cell, wherein the biological parameters include fluorescence intensity parameters, cell morphological characteristic parameters, and cell texture characteristic parameters, to obtain a multi-parameter phenotypic spectrum.

7. The toxicity risk assessment method according to claim 1, characterized in that: The multi-parameter phenotypic spectrum data described in S3 were standardized and PCA dimension reduced before unsupervised clustering, and hierarchical clustering was performed using the Agglomerative Clustering algorithm, with Euclidean distance as the distance metric and Ward's criterion as the linkage criterion.

8. The toxicity risk assessment method according to claim 1, characterized in that: After S3 obtains the multi-parameter phenotypic spectrum data, it divides it into training set, validation set and test set, defines the model and parameter grid for training, selects the optimal parameters through grid search, retrains the model and generates a classification report. The report results are presented in terms of accuracy, recall rate and F1 score.

9. The toxicity risk assessment method according to claim 1, characterized in that: The machine learning models described in S3 include logistic regression, gradient boosting tree, random forest, support vector machine, K nearest neighbor, decision tree, XGBoost and naive Bayes.

10. The toxicity risk assessment method according to claim 1, characterized in that: The toxicity risk levels described in S4 include low risk, medium risk and high risk.