Pollutant mixture toxicity prediction method and device, storage medium and electronic equipment

By constructing an AI-based method for predicting the mixed toxicity of pollutants, automatically collecting and structuring pollutant data from surface water bodies, and combining high-throughput toxicity testing and machine learning models, the accuracy problem of assessing the mixed toxicity of pollutants in surface water bodies was solved, enabling the identification and risk prediction of key pollutant components.

CN122494014APending Publication Date: 2026-07-31CHINESE RES ACAD OF ENVIRONMENTAL SCI
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHINESE RES ACAD OF ENVIRONMENTAL SCI
Filing Date
2026-03-31
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

In existing technologies, the assessment of mixed toxicity of pollutants in surface water bodies suffers from insufficient representativeness of pollutant combinations, making it difficult to accurately identify key components, resulting in low accuracy of ecological risk assessment. Furthermore, existing methods are insufficient to cover real mixed exposure combinations at the national or regional scale.

Method used

We automatically collect pollutant exposure concentration data from multiple sources using web crawling and text recognition technologies, construct a structured database, combine frequent itemset mining and high-throughput microplate toxicity analysis, use machine learning algorithms to build a mixture toxicity prediction model, and identify high-risk pollutants through global sensitivity analysis.

Benefits of technology

It enables automated data acquisition and processing of mixed toxicity data of pollutants in surface water bodies, improves the efficiency of identifying mixed exposure combinations and the accuracy of risk assessment, can identify key pollutant components on a large scale, and supports efficient prediction and risk screening of mixed toxicity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122494014A_ABST
    Figure CN122494014A_ABST
Patent Text Reader

Abstract

This invention relates to a method, apparatus, storage medium, and electronic device for predicting the toxicity of mixed pollutants. The method includes: acquiring structured data using web crawling and text recognition technology; determining target pollutant sequences based on the pollutant names contained in the structured data; constructing a pollutant presence / absence matrix based on the contained pollutant names and the target pollutant sequences; constructing typical pollutant combinations based on a frequent itemset mining algorithm and the pollutant presence / absence matrix; performing high-throughput microplate toxicity analysis based on various concentration compositions of the typical pollutant combinations to obtain mixture toxicity data; constructing a mixture toxicity prediction model based on machine learning algorithms, various concentration compositions, and mixture toxicity data; obtaining mixture toxicity prediction data for untested pollutant combinations based on the mixture toxicity prediction model; and performing sensitivity analysis on the mixture toxicity prediction data to identify high-risk pollutants. This can improve the accuracy of risk assessment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of environmental monitoring technology, and in particular to a method, apparatus, storage medium, and electronic device for predicting the mixed toxicity of pollutants in surface water bodies based on artificial intelligence. Background Technology

[0002] With the acceleration of industrialization and urbanization, a large number of pollutants, such as pharmaceuticals and personal care products, pesticides, and industrial organic matter, are continuously discharged into surface water bodies through urban sewage, industrial wastewater, and agricultural runoff. Although the concentration of individual pollutants is mostly at a low level, the diverse structures of various pollutants and the complex environment of surface water bodies make pollutants exhibit the characteristics of widespread coexistence and long-term exposure in space and time, thus posing a cumulative risk to aquatic ecosystems and human health.

[0003] In related technologies, ecological risk assessments of pollutants in surface water bodies mostly focus on single pollutants or, based on experience, select a limited number of pollutants with high concentrations to conduct mixed toxicity experiments, and then conduct risk assessments based on the mixed toxicity data obtained from these experiments. However, because this method relies on empirically selected pollutant combinations, the resulting combinations may lack representativeness. For example, the selected pollutants may not effectively characterize the key components in surface water bodies. Especially in mixed exposure studies of mixed toxicity experiments, the limited pollutant combinations are insufficient to reflect the true mixed exposure spectrum of pollutants in surface water bodies at the national or regional scale. Consequently, it is difficult to identify and determine the key pollutant components affecting mixed toxicity, leading to low accuracy in the risk assessment of surface water bodies. Summary of the Invention

[0004] In view of this, the present invention provides a method, apparatus, storage medium and electronic device for predicting the mixed toxicity of pollutants.

[0005] Specifically, the present invention is achieved through the following technical solution: According to a first aspect of the present invention, a method for predicting the mixed toxicity of pollutants is provided, the method comprising: Using web crawling and text recognition technologies, target data files corresponding to pollutants in surface water bodies are obtained from literature databases, environmental monitoring reports, and public data platforms, and structured data is constructed based on the target data files; Based on the pollutant names contained in the structured data, the target pollutant sequence is determined. The structured data is preprocessed, and a pollutant presence / absence matrix for each sampling point is constructed based on the pollutant names contained in the preprocessed structured data and the target pollutant sequence. Based on a pre-set frequent itemset mining algorithm, typical pollutant combinations are constructed based on the pollutant presence / absence matrix. High-throughput microplate toxicity analysis was performed based on various concentration compositions of typical pollutant combinations to obtain toxicity data for mixtures. Based on pre-set machine learning algorithms, various concentration compositions of typical pollutant combinations, and mixture toxicity data, a mixture toxicity prediction model is constructed. Based on the constructed mixture toxicity prediction model, mixture toxicity prediction data of untested pollutant combinations are obtained. Using a pre-set sensitivity analysis method, sensitivity analysis was performed on the toxicity prediction data of mixtures of untested pollutant combinations. Based on the sensitivity analysis results, high-risk pollutants affecting the mixed toxicity of pollutant combinations in the untested pollutant combinations were identified.

[0006] According to a second aspect of the present invention, a pollutant mixed toxicity prediction device is provided, the pollutant mixed toxicity prediction device comprising: The data acquisition module is used to acquire target data files corresponding to pollutants in surface water bodies from literature databases, environmental monitoring reports and public data platforms using web crawling and text recognition technology, and to construct structured data based on the target data files; The matrix construction module is used to determine the target pollutant sequence based on the pollutant names contained in the structured data, perform data preprocessing on the structured data, and construct a pollutant presence / absence matrix for each sampling point based on the pollutant names contained in the preprocessed structured data and the target pollutant sequence. The itemset mining module is used to construct typical pollutant combinations based on the pollutant presence / absence matrix, using a pre-set frequent itemset mining algorithm. The toxicity testing module is used for high-throughput microplate toxicity analysis based on various concentration compositions of typical pollutant combinations to obtain toxicity data for mixtures. The toxicity prediction module is used to build a mixture toxicity prediction model based on pre-set machine learning algorithms, various concentration compositions of typical pollutant combinations, and mixture toxicity data. Based on the built mixture toxicity prediction model, it obtains mixture toxicity prediction data for untested pollutant combinations. The toxicity analysis module is used to perform sensitivity analysis on the toxicity prediction data of mixtures of untested pollutant combinations using pre-set sensitivity analysis methods, and to identify high-risk pollutants in the untested pollutant combinations that affect the toxicity of the mixtures based on the sensitivity analysis results.

[0007] According to a third aspect of the invention, a storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the steps of the pollutant mixed toxicity prediction method in any possible implementation of the first aspect.

[0008] According to a fourth aspect of the present invention, an electronic device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps of the pollutant mixed toxicity prediction method in any possible implementation of the first aspect. Attached Figure Description

[0009] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention.

[0010] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, those skilled in the art can obtain other drawings based on these drawings without creative effort.

[0011] Figure 1 A schematic flowchart of a pollutant mixed toxicity prediction method provided in an embodiment of the present invention; Figure 2 This is a schematic diagram of a pollutant mixed toxicity prediction device provided in an embodiment of the present invention; Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation

[0012] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0013] In related technologies, ecological risk assessment of surface water bodies is generally based on a single pollutant or by selecting a limited number of pollutants based on experience to conduct mixed toxicity experiments. The ecological risk of pollutants in the surface water body is assessed based on the obtained mixed toxicity data. However, the key components of the pollutants in the pollutant combination are not representative enough, and the results of the mixed toxicity experiments are difficult to effectively characterize the ecological risk of the target surface water body, resulting in low accuracy of ecological risk assessment.

[0014] Furthermore, current data on the exposure concentrations of pollutants (pollutant components) in surface water bodies nationwide still primarily rely on researchers manually reviewing literature and monitoring reports, and manually entering and organizing data. This approach struggles to address the challenges posed by the numerous, heterogeneous, and frequently updated data sources from different research institutions or organizations. This not only results in a massive workload and a high risk of errors but also makes it difficult to construct a unified structured database of "sampling point – time – pollutant name – concentration" at the national or watershed scale. Consequently, it hinders the identification of key pollutant components in multi-pollutant mixed exposure combinations, leading to insufficient breadth and depth in the risk assessment of surface water bodies, and resulting in low confidence levels.

[0015] Regarding the ecological risk assessment of pollutants in surface water bodies, there is currently no method that organically integrates artificial intelligence technologies such as web crawling, layout analysis, pre-trained language models, named entity recognition, and rule matching to automatically acquire and standardize pollutant exposure concentration data in surface water bodies. Furthermore, based on this acquired data, a method is needed to identify typical mixed exposure combinations from various pollutant exposure combinations by combining 0-1 concentration processing and frequent itemset mining. Simultaneously, in terms of mixed toxicity experiments, existing studies mostly rely on pollutant exposure combinations composed of a limited number of pollutants to obtain mixed toxicity data. This fails to cover the scenario where a large number of pollutant exposure combinations actually exist in surface water bodies. There is also a lack of machine learning toxicity prediction models built based on real pollutant exposure combinations and high-throughput toxicity data, making it impossible to predict the mixed toxicity and screen for risks of untested pollutant exposure combinations. Therefore, there is an urgent need to develop an end-to-end technical solution for predicting the toxicity of pollutant exposure combinations by automatically constructing a structured database of pollutant exposure concentrations in surface water bodies based on multi-source literature and monitoring data, coupled with a machine learning toxicity prediction model.

[0016] In this embodiment, advancements in artificial intelligence, data mining, and high-throughput toxicity testing technologies have enabled the automatic collection and structured organization of large-scale environmental monitoring data from multi-source publicly available data. Furthermore, based on association analysis methods such as frequent itemset mining, this embodiment can identify co-occurrence patterns from a large number of "present / absent" records. Combined with a high-throughput micro-plate platform, it can rapidly acquire mixed toxicity data for multiple pollutant exposure combinations. Utilizing global sensitivity analysis methods, it provides an effective analytical tool for quantifying the response of each pollutant component in a mixed exposure combination to mixed toxicity experiments under multi-parameter conditions.

[0017] In this embodiment, based on the constructed structured database of pollutant exposure concentrations and high-throughput toxicity data of representative pollutant mixed exposure combinations, a machine learning toxicity prediction model for mixed exposure of pollutants in surface water is constructed to achieve mixed toxicity prediction of complex pollutant mixed exposure combinations. Combined with global sensitivity analysis, the key pollutant components driving mixed toxicity are identified, forming an integrated technical route of "automated data acquisition and processing + mixed exposure combination identification + high-throughput toxicity testing + machine learning prediction + sensitivity analysis".

[0018] See Figure 1 This invention provides a method for predicting the mixed toxicity of pollutants, which may include the following steps: S101. Using web crawler and text recognition technology, obtain target data files corresponding to pollutants in surface water bodies from literature databases, environmental monitoring reports and public data platforms, and construct structured data based on the target data files; In this embodiment, as an optional embodiment, the environmental monitoring report includes, but is not limited to, annual water environment monitoring reports and special monitoring reports. The public data platform includes, but is not limited to, the water quality monitoring data open platform. The target data files include: surface water pollutant monitoring literature published in the literature database, surface water pollutant data in the annual water environment monitoring reports issued by the ecological and environmental authorities, surface water pollutant data in the special monitoring reports, and surface water pollutant monitoring datasets issued by the water quality monitoring data open platform.

[0019] In this embodiment, artificial intelligence technology is used to automatically collect target data files containing pollutant exposure concentration data. As an optional embodiment, web crawling and text recognition technology are used to obtain target data files corresponding to pollutants in surface water bodies from literature databases, environmental monitoring reports, and public data platforms, including: Set up a search query for collecting pollutant exposure concentration data, and use web crawling and text recognition technology to collect target data files containing pollutant exposure concentration data from pre-acquired literature databases, environmental monitoring reports, and public data platforms based on the search query.

[0020] In this embodiment, a multi-source database is constructed from literature databases, environmental monitoring reports, and public data platforms. The search query includes search conditions related to pollutants in surface water bodies. Target data files are automatically downloaded and / or retrieved from the multi-source database according to the preset search query.

[0021] In this embodiment, the formats of literature databases and monitoring reports compiled or published by different research institutions or organizations vary, requiring corresponding processing. As an optional embodiment, structured data is constructed based on the target data file, including: A11. Using a pre-built layout analysis model, the target data file is identified to obtain the table area and the text area. Using a natural language processing model, information is extracted from the table area and the text area to obtain the initial structure data. The information extraction includes: sampling point information extraction, sampling time information extraction, pollutant name information extraction, and pollutant exposure concentration value information extraction. In this embodiment, a page layout analysis model is used to identify the page structure of the target data file, distinguishing between table areas and text areas. Then, a natural language processing model is used to extract sampling point information, sampling time information, pollutant name information, and pollutant exposure concentration information from the table areas and / or text areas. For each sampling point, an initial structured data of sampling point-sampling time-pollutant name-pollutant exposure concentration value is formed.

[0022] In this embodiment, a layout analysis model can be used to identify the table areas and text areas contained in each page structure of the target data file, and to determine the table boundaries, header rows, and data cell positions within the table areas, thereby facilitating subsequent information extraction. Within the text areas, paragraphs or sentences related to pollutant information can be located.

[0023] In this embodiment, as an optional embodiment, the natural language processing model includes, but is not limited to, a pre-trained language model and other natural language processing models. Using the natural language processing model, sampling point information, sampling time information, pollutant name information (pollutant name) and pollutant exposure concentration value information are automatically extracted from the table area and text area obtained by the layout analysis model. Based on the information extraction, the corresponding initial extraction results are generated, that is, the initial structure data containing sampling point-time-pollutant name-concentration is constructed.

[0024] A12, based on the pre-acquired pollutant dictionary, maps the pollutant names in the initial structure data to the standardized pollutant names corresponding to the pollutant dictionary; In this embodiment, pollutant names are standardized. As an optional embodiment, a named entity recognition model is used to identify chemical substance entities in the initial structured data to obtain pollutant names. A pollutant dictionary containing the pollutant's Chinese and English names, abbreviations, common aliases, and standardized pollutant names is then queried to obtain the standardized pollutant name matching the chemical substance entity. This unifies chemical substance entities with different names to the same standardized pollutant name. As an optional embodiment, the chemical substance entities identified by the named entity recognition model can also be verified and corrected using a preset rule base before matching is performed based on the pollutant dictionary.

[0025] A13 standardizes the concentration units corresponding to the pollutant exposure concentration values ​​in the initial structured data and updates the pollutant exposure concentration values ​​accordingly, constructing structured data based on the sampling point-sampling time-pollutant name-pollutant exposure concentration value structure.

[0026] In this embodiment, automatic identification and unified conversion of concentration units are performed. As an optional embodiment, the physical unit corresponding to the concentration value can be identified based on text information such as table headers, annotations, and method descriptions. Then, according to preset unit conversion rules, concentration values ​​under different units are uniformly converted to concentration values ​​under the target concentration unit (standard concentration unit). For initialization structure data where the concentration unit is not explicitly given, the unit type is determined and the corresponding unit conversion is performed through context inference and unit conversion rules corresponding to the initialization structure data.

[0027] In this embodiment, after constructing structured data based on the target data file, a database can be constructed based on the structured data.

[0028] In this embodiment, as an optional implementation, structured data of "sampling point – sampling time – pollutant name – pollutant exposure concentration value" is constructed based on the standardized sampling point information, sampling time, standard name of pollutant, and concentration value under standardized units (pollutant exposure concentration value). As an optional implementation, each sampling point corresponds to one set of structured data, which includes, but is not limited to, being represented by tables or documents.

[0029] In this embodiment, as an optional embodiment, the method further includes: All the structured data constructed will be preprocessed and stored in the database to form a structured database of surface water pollutant exposure concentrations covering multiple watersheds, multiple water body types and multiple types of pollutants.

[0030] S102. Based on the pollutant names contained in the structured data, determine the target pollutant sequence, perform data preprocessing on the structured data, and construct a pollutant presence / absence matrix for each sampling point based on the pollutant names contained in the preprocessed structured data and the target pollutant sequence. In this embodiment, as an optional embodiment, data preprocessing includes quality control and data cleaning of pollutant exposure concentration data in structured data. The quality control and data cleaning include, but are not limited to: identification and removal of concentration outliers, handling of missing concentration values, merging of duplicate sampling records, and standardization detection of concentration units.

[0031] In this embodiment, as an optional implementation, various pollutants in the statistically acquired structured data are used as the target pollutant sequence. Alternatively, pollutants with negligible concentrations can be removed from the acquired structured data to obtain the target pollutant sequence. Subsequently, the target pollutant sequence can be adjusted based on incrementally acquired structured data, or updated based on changes in surface water monitoring. For each target pollutant, a pre-set presence threshold is established. For each sampling point, when the target pollutant is present at the sampling point and its exposure concentration is greater than or equal to the set presence threshold, the row and column values ​​of the pollutant presence / absence matrix corresponding to that sampling point are recorded as 1, indicating that the target pollutant is detected at that sampling point. When the target pollutant is not present at the sampling point, or its exposure concentration is less than the presence threshold, the row and column values ​​of the pollutant presence / absence matrix corresponding to that sampling point are recorded as 0. In this embodiment, each target pollutant is used as a column of the matrix, and each sampling point corresponds to a row of the matrix, thus obtaining a 0-1 matrix (pollutant presence / absence matrix) characterizing the presence / absence of the target pollutant at the sampling point.

[0032] In this embodiment, for each target pollutant in the database (structured database of surface water pollutant exposure concentration), a corresponding presence threshold is set, thereby performing 0-1 processing on the structured data of each sampling point. When the exposure concentration value of a pollutant at a sampling point is greater than or equal to the corresponding presence threshold, the pollutant at that sampling point is recorded as 1, otherwise it is recorded as 0, thus obtaining the 0-1 presence / absence matrix of sampling point – target pollutant.

[0033] In this embodiment, as an optional embodiment, the threshold for the presence of the target pollutant is set to be at least one of the following indicators: the limit of quantitation for the monitoring method of the target pollutant, the environmental quality benchmark value corresponding to the target pollutant, and the preset quantile based on the statistical distribution of the exposure concentration of the target pollutant.

[0034] S103, Based on the pre-set frequent itemset mining algorithm, construct typical pollutant combinations based on the pollutant presence / absence matrix; In this embodiment, as an optional implementation, based on a pre-set frequent itemset mining algorithm and the pollutant presence / absence matrix, a typical pollutant combination is constructed, including: Traverse the pollutant presence / absence matrix, count the number of supporting rows with the same target pollutant set, and determine the support of the target pollutant set based on the number of supporting rows and the total number of rows in the pollutant presence / absence matrix. From each set of target pollutants, obtain the set of target pollutants with a support greater than a preset support threshold, obtain the frequent pollutant combination, and use the frequent pollutant combination as the typical pollutant combination.

[0035] In this embodiment, as an optional implementation, each sampling point is regarded as a transaction, and the target pollutant with a value of 1 in the pollutant presence / absence matrix corresponding to the sampling point is regarded as an item in the transaction. A support threshold is set, and the Apriori algorithm and / or FP-Growth algorithm are used to perform frequent itemset mining on the items of each transaction in the pollutant presence / absence matrix to obtain frequent pollutant combinations that meet the support threshold.

[0036] In this embodiment, each sampling point is treated as a transaction, and each target pollutant (pollutant with a value of 1) in that sampling point is considered as an item in the transaction. Combinations of these items form an itemset, and each itemset corresponds to a pollutant combination; that is, an itemset corresponds to a chemical combination composed of the target pollutants. Support is defined as the percentage of sampling points containing the same item out of the total number of sampling points. A support threshold is set based on the sampling point size, support, and research objective; for example, the support threshold is preferably in the range of 0.01 to 0.5. For a pollutant combination, the support is defined as the percentage of sampling points containing the same item in that combination out of the total number of sampling points in that combination. Itemets with support greater than the support threshold are constructed as frequent pollutant combinations. Thus, frequent itemset mining algorithms, such as the Apriori algorithm and the FP-Growth algorithm, are used to mine the pollutant presence / absence matrix to obtain frequent pollutant combinations that meet the support threshold. For example, assuming the pollutant presence / absence matrix includes 100 sampling points corresponding to 10 pollutants, the resulting pollutant presence / absence matrix is ​​100x10. Assuming the number of sampling points containing only pollutants A and B is 5, the number containing only pollutants A and C is 36, the number containing only pollutants A, B, and C is 15, ..., and the most common pollutant types are A, B, C, D, E, F, and G, with a corresponding sampling point count of 25, then the support of pollutant combination 1 (containing pollutants A and B) is 0.05, the support of pollutant combination 2 (containing pollutants A and C) is 0.36, the support of pollutant combination 3 (containing pollutants A, B, and C) is 0.15, ..., and the support of pollutant combination 4 (containing pollutants A, B, C, D, E, F, and G) is 0.25. If the support threshold is set to 0.2, then the frequent pollutant combinations obtained based on the frequent itemset mining algorithm include: pollutant combination 2 and pollutant combination 4.

[0037] In this embodiment, as an optional embodiment, the method further includes: Based on the support of frequent pollutant combinations, typical mixture combinations containing representative pollutants are selected from each frequent pollutant combination.

[0038] In this embodiment, as an optional implementation, the frequent pollutant combinations are sorted according to their support, and one or more frequent pollutant combinations with the highest support are selected as representative surface water pollutant mixed exposure combinations, i.e., typical mixture combinations.

[0039] S104, based on the various concentration compositions of typical pollutant combinations, high-throughput microplate toxicity analysis was performed to obtain mixture toxicity data; In this embodiment, as an optional implementation, high-throughput microplate toxicity analysis is performed based on various concentration compositions of typical pollutant combinations to obtain toxicity data for the mixture, including: Based on the various concentration compositions of typical pollutant combinations, multiple sets of typical pollutant concentration combinations are constructed. For each typical pollutant concentration combination, a pollutant solution of corresponding concentration is prepared to obtain the pollutant concentration test group corresponding to that typical pollutant concentration combination. High-throughput microplate toxicity analysis was performed for each pollutant concentration test group to obtain the mixture toxicity data for that pollutant concentration test group.

[0040] In this embodiment, based on the pre-selected toxicity endpoint of the test organism, and based on the actual concentration range of each target pollutant in the surface water environment in the representative surface water pollutant mixed exposure combination (typical pollutant combination), the concentration composition of each target pollutant in the typical pollutant combination is designed, and toxicity test is carried out using a 96-well high-throughput microplate to obtain the toxicity data of the mixture.

[0041] In this embodiment, a mixture sample solution of corresponding concentration is prepared for each target pollutant in the typical pollutant combination to obtain multiple typical pollutant concentration combinations. The mixture toxicity test is then performed on the test organism (e.g., luminescent bacteria) to obtain mixture toxicity data.

[0042] S105. Based on a pre-set machine learning algorithm, various concentration compositions of typical pollutant combinations and mixture toxicity data, a mixture toxicity prediction model is constructed. Based on the constructed mixture toxicity prediction model, mixture toxicity prediction data of untested pollutant combinations is obtained. In this embodiment, as an optional embodiment, a mixture toxicity prediction model is constructed based on a pre-set machine learning algorithm, various concentration compositions of typical pollutant combinations, and mixture toxicity data, including: The concentration information of each pollutant in the pollutant concentration test group was used as the input to the machine learning algorithm model. Using the toxicity data of the mixture corresponding to the pollutant concentration test group as the output of the machine learning algorithm model, the machine learning algorithm model is trained by using propagation neural network, random forest, graph neural network or gradient algorithm in the backpropagation machine learning algorithm to obtain the mixture toxicity prediction model.

[0043] In this embodiment, a mixture toxicity prediction model is constructed based on the obtained structured exposure database and experimental data (mixture toxicity data) from high-throughput microplate toxicity analysis. The input features of the initialized machine learning algorithm model include, but are not limited to, the environmental exposure concentration of each target pollutant in the pollutant concentration test group.

[0044] In this embodiment, as an optional embodiment, the mixture toxicity data can also be the mixture toxicity data corresponding to each sampling point that has undergone toxicity testing collected during the data collection process. In this way, the concentration data of each pollutant contained in the sampling point and the mixture toxicity data corresponding to the sampling point can be used as sample training data, and combined with the mixture toxicity data obtained from high-throughput microplate toxicity experiments for training, so as to effectively increase the amount of sample data required for training the mixture toxicity prediction model.

[0045] In this embodiment, as an optional implementation, the machine learning algorithm includes, but is not limited to, Back Propagation Neural Network (BPNN), Convolutional Neural Network (CNN), Graph Neural Network (GNN), Random Forest, XGBoost, or other gradient boosting algorithms. Supervised learning is performed based on sample data (typical pollutant combinations) labeled with toxic endpoints, and the model structure and training parameters are optimized through cross-validation and hyperparameter search, so that the model can fit the toxic response of the mixture (such as EC50, inhibition rate, etc.) better.

[0046] In this embodiment, as an optional embodiment, the untested contaminant combinations include, but are not limited to, frequent contaminant combinations and other contaminant combinations present at each sampling point.

[0047] S106. Using a pre-set sensitivity analysis method, perform sensitivity analysis on the toxicity prediction data of mixtures of untested pollutant combinations, and identify high-risk pollutants in the untested pollutant combinations that affect the toxicity of the mixtures based on the sensitivity analysis results.

[0048] In this embodiment, as an optional embodiment, a pre-set sensitivity analysis method is used to perform sensitivity analysis on the toxicity prediction data of the mixture of untested pollutant combinations. Based on the sensitivity analysis results, high-risk pollutants affecting the mixed toxicity of pollutant combinations in the untested pollutant combinations are identified, including: Using the Sobol global sensitivity analysis method and / or the Morris sensitivity analysis method, based on the mixture toxicity prediction data of untested pollutant combinations, the first-order sensitivity index of each pollutant is calculated. Pollutants whose first-order sensitivity index exceeds the pre-set sensitivity threshold are identified as key pollutants that induce mixture toxicity. A key pollutant group is constructed based on each key pollutant. Based on the constructed mixture toxicity prediction model, key prediction data of mixture toxicity for key pollutant groups are obtained; If the predicted toxicity data of the mixture of untested pollutant combinations is not significantly different from the key predicted toxicity data of the mixture of key pollutant groups, the key pollutant group is identified as a high-risk pollutant group that affects the mixed toxicity of pollutants in the untested pollutant combination. If there is a significant difference between the predicted toxicity data of the mixture of untested pollutant combinations and the key predicted toxicity data of the mixture of key pollutant groups, calculate the higher-order sensitivity index of each pollutant corresponding to the untested pollutant combinations, reconstruct the key pollutant groups based on the higher-order sensitivity index and the pre-set higher-order sensitivity threshold, and execute the step of obtaining the key predicted toxicity data of the mixture of key pollutant groups based on the constructed mixture toxicity prediction model.

[0049] In this embodiment, based on the mixture toxicity prediction model and global sensitivity analysis method, the toxicity endpoint is predicted for each frequent pollutant combination to obtain high-risk pollutant combinations.

[0050] In this embodiment, as an optional implementation, the Sobol global sensitivity analysis method and / or the Morris sensitivity analysis method are used to calculate the first-order sensitivity index of each pollutant based on the toxicity prediction data of the mixture of untested pollutant combinations, including: By changing the concentration of a candidate pollutant in the untested pollutant combination, the corresponding mixture toxicity prediction data is obtained based on the constructed mixture toxicity prediction model. Based on the change in the mixture toxicity prediction data before and after the concentration change, the basic effect value of the candidate pollutant is obtained. The first-order sensitivity index of the candidate pollutant is determined based on the basic effect value of the candidate pollutant, the arithmetic mean of the basic effect values ​​of each candidate pollutant, the arithmetic mean of the absolute values ​​of the basic effect values ​​of each candidate pollutant, and the standard deviation of the basic effect values ​​of each candidate pollutant.

[0051] In this embodiment, the arithmetic mean of the basic effect values ​​and the arithmetic mean of the absolute values ​​of the basic effect values ​​of candidate pollutants can be used to determine the degree of influence of candidate pollutants on the toxicity of the mixture. The standard deviation can characterize the interaction between candidate pollutants. In this way, by using the basic effect values ​​of the candidate pollutants, the key components affecting the toxicity of the mixture can be accurately identified, thereby providing a scientific basis for prioritizing the control of which key pollutants in the risk assessment of complex mixtures.

[0052] In this embodiment, a trained machine learning model (mixture toxicity prediction model) is used to predict the toxicity endpoints of frequent pollutant combinations for which no actual toxicity tests have been conducted, thereby achieving efficient risk screening. Furthermore, model interpretation methods such as SHAP values ​​and ensemble gradients can be combined to quantify the contribution of each pollutant component in the frequent pollutant combination to the model output, assisting in the identification of potential critical pollutants from a model perspective.

[0053] In this embodiment, a machine learning algorithm model is constructed with the exposure concentration of each pollutant in a typical pollutant combination as the input variable and the corresponding mixture toxicity data as the output variable. Within the actual environmental concentration range of pollutants, input parameter combinations (concentration combinations) are generated through Monte Carlo sampling or quasi-random sampling methods. The Sobol global sensitivity analysis method and / or Morris method are used to calculate the first-order sensitivity index and the total sensitivity index of each pollutant corresponding to the input parameter combination. Target pollutants whose sensitivity indices exceed the preset sensitivity threshold are identified as key pollutant components that induce mixture toxicity. The identification results are verified through a toxicity comparison experiment between key component combinations containing only key pollutant components and full component combinations (input parameter combinations) containing all pollutant components.

[0054] In this embodiment, if the toxicity of the key component combination is not significantly different from or highly consistent with the toxicity of the whole component combination, it indicates that the key components identified by the sensitivity analysis can represent the main toxicity drivers of the mixture, thus verifying the effectiveness of the identification method. If the toxicity of the key component combination is significantly lower than that of the whole component combination, it indicates that there is an important synergistic effect in the toxicity of the mixture or that it is significantly affected by components with low sensitivity indices. Further correction of the key component combination is needed, for example, checking whether any potentially important variables have been missed, and, if necessary, using higher-order sensitivity analysis methods or introducing an investigation of inter-component interaction effects to improve the accuracy of key component identification.

[0055] The method of this embodiment has at least the following beneficial effects: 1. Automated acquisition and structured processing of new pollutant exposure concentrations in surface water for mixed exposure studies: This embodiment introduces artificial intelligence technologies such as web crawling, layout analysis, pre-trained language models, and named entity recognition to automatically extract "sampling point-time-new pollutant-concentration" information from multi-source heterogeneous monitoring literature and reports. This can significantly reduce the workload of manual review and data entry, improve data collection efficiency and accuracy, and provide standardized exposure data for mixed exposure studies.

[0056] 2. Supports large-scale, multi-dimensional mixed exposure feature recognition: In this embodiment, by unifying pollutant names and concentration units, a structured database of new pollutant exposure concentrations in surface water that can be integrated at the watershed, regional, and even national scales is constructed, providing complete and computable data support for the analysis of mixed exposure characteristics of new pollutants at different spatial scales and water body types.

[0057] 3. Efficient coupling with data mining methods such as frequent itemset mining: In this embodiment, based on a structured database, by performing 0-1 conversion and frequent itemset mining, it is possible to objectively identify new pollutant combinations that are real and commonly co-occurring in the environment, overcoming the limitations of previous subjective group selection based on researchers' experience, and improving the screening efficiency and rationality of representative mixture combinations.

[0058] 4. The data acquisition process is easy to expand and update: This embodiment features AI-powered data acquisition and structured processing, which offers excellent scalability. It can iterate based on updates to the new pollutant list, data source expansion, and algorithm upgrades. By re-executing the automatic extraction process, the exposure database can be quickly updated, reducing long-term operation and maintenance costs.

[0059] 5. Machine learning-based ability to predict the toxicity of mixtures: Based on high-throughput toxicity experimental data of representative mixtures, a machine learning prediction model can be constructed to rapidly predict the toxicity endpoints of new pollutant combinations that have not been tested in actual experiments. This enables efficient risk screening of large-scale potential combinations and significantly expands the coverage of traditional experimental studies in terms of the number of combinations and concentration scenarios.

[0060] 6. Organic integration with global sensitivity analysis and model interpretation methods: This embodiment utilizes global sensitivity analysis to quantitatively assess the impact of each pollutant on the mixed toxicity output from a parameter space perspective. On the other hand, it combines interpretability indicators from a machine learning model to identify key components from the perspective of the internal mechanism of the machine learning model. The two types of information complement each other, providing a more reliable and interpretable decision-making basis for the priority control of new pollutants and water environment management.

[0061] 7. Establish an integrated technical approach encompassing data acquisition, combination identification, toxicity testing, intelligent prediction, and key component identification: This embodiment, within a pre-set unified framework, organically links the automated collection and structured processing of literature and monitoring data, the identification of mixed exposure combinations, high-throughput toxicity testing, machine learning prediction, and global sensitivity analysis. This avoids information loss caused by the fragmentation of each link and improves the systematicness, accuracy, and practicality of the risk assessment of mixed exposure to new pollutants in surface water.

[0062] The following specific examples will be described in detail.

[0063] Example: AI-based data collection and structured database construction of new surface water pollutant exposure concentrations, using antibiotics as an example, and analysis of mixed exposures and toxicity prediction. This embodiment selects antibiotic-type new pollutants as the target new pollutants and takes surface water bodies in several typical watersheds as the research objects. The specific implementation steps are as follows: AI-driven acquisition and structured database construction of S1 antibiotic exposure concentration data: S1.1 Data Source Determination and Automatic Acquisition: Target data sources include: publicly available literature on surface water antibiotic monitoring in domestic and international literature databases; annual water environment monitoring reports and special monitoring reports issued by some provincial ecological and environmental departments; and dedicated water quality monitoring data open platforms. As an optional embodiment, a search query is constructed based on keywords such as "surface water," "antibiotics," and "antibiotics," and their combinations. A web crawler is used to search and access the aforementioned literature databases, automatically downloading full-text PDFs, HTML pages, and report files of literature that meet the search query conditions. For online open water quality monitoring datasets, the original data files are automatically obtained through interface requests or web crawling.

[0064] S1.2 Text Layout Analysis and Region Recognition: For unparseable formats such as PDFs and reports, layout analysis algorithms are used to identify table boundaries, header rows, and data cell areas on the page, and to distinguish between main text paragraphs, chart descriptions, and appendix content. For HTML or other parsable formats, table nodes and text nodes are located based on document structure tags, and table and text areas that may contain monitoring information are marked as candidate areas for subsequent information extraction.

[0065] S1.3 Extract monitoring information using natural language processing models: In the table area, the pre-trained language model and rule template are called to identify the column headers and meanings of the table. The columns corresponding to fields such as "sampling point name", "latitude and longitude", "administrative region", "sampling time", "antibiotic name", and "concentration" are automatically labeled, and specific values ​​or text are extracted from each cell. In the main text area, a pre-trained language model is used to perform sequence labeling and relation extraction on sentences in the main text area, identify sentences containing sampling point information, sampling time and monitoring results, and extract the corresponding fields from them; then, the extraction results from the table and the main text area are merged and matched to form an initial record set of "sampling point - time - candidate antibiotic name - concentration value".

[0066] S1.4 Antibiotic Name Standardization: Candidate antibiotic names are standardized using a pre-built antibiotic dictionary containing the English and Chinese names, abbreviations, and other aliases of common antibiotics. Specifically, a named entity recognition model is used to identify candidate antibiotic names in the extraction results. The identified candidate entities are matched with entries in the antibiotic dictionary based on similarity. Spelling variations, capitalization differences, and abbreviations are standardized. For example, different spellings such as "sulfamethoxazole," "SMX," and "sulfamethoxazole" are uniformly mapped to the standard name "sulfamethoxazole." Furthermore, candidate entities with low confidence or that do not match entries in the antibiotic dictionary are incorrectly removed or corrected through a combination of rule base and manual sampling.

[0067] S1.5 Automatic Identification and Unified Conversion of Concentration Units: In the table area, the concentration unit is automatically identified based on the table header and column descriptions, such as "ng / L", "μg / L", "mg / L", etc.; in the text area, the concentration unit is supplemented by combining the concentration units appearing in the method description and results description; as an optional embodiment, a target concentration unit is set, such as ng / L, and the concentration values ​​in all records are numerically converted according to the conversion relationship between the concentration unit and the target concentration unit ng / L; for some records where the concentration unit is not explicitly given but is consistent with the concentration unit of other records in the same table or document, the concentration unit is supplemented by inference from the context; for records where the concentration unit information cannot be reliably inferred, they are marked as uncertain and removed or stored separately.

[0068] S1.6 Structured Record Generation and Database Construction: Monitoring information, after name standardization and unified concentration unit conversion, is organized into structured records according to the field order of "Sampling Point ID – Sampling Point Spatial Location – Sampling Time – Standard Antibiotic Name – Concentration (ng / L)". The sampling point ID is uniquely identified by the sampling point name and latitude / longitude information, and the sampling time is uniformly formatted as a standard date. All structured records are imported into the database management system, thereby establishing a structured database of surface water antibiotic exposure concentrations covering multiple watersheds and water body types nationwide.

[0069] After constructing a structured database of antibiotic exposure concentrations in surface water, the antibiotic exposure concentration data were processed by 0-1 conversion, frequent combination mining, and representative combination screening. This resulted in obtaining multiple antibiotic mixed exposure combinations with high support and spatial coverage in different watersheds and water body types. Based on the antibiotic mixed exposure combinations, a high-throughput toxicity test scheme for representative mixtures was designed to obtain multi-endpoint mixed toxicity data.

[0070] In this embodiment, based on the constructed antibiotic exposure and toxicity database, machine learning models such as BPNN are used to build an input-output mapping relationship. The model is trained and validated using the component concentrations, 0-1 existence matrix, and frequent combination indexes of the antibiotic mixed exposure combination as inputs, and the experimental toxicity endpoints (multi-endpoint mixed toxicity data) as outputs. The trained model can be used to predict the toxic effects of the experimental combination (different concentration combinations). Combined with global sensitivity analysis, key antibiotic components that contribute significantly to mixed toxicity are identified, thereby achieving quantitative assessment of antibiotic mixed exposure risk and screening of priority control substances.

[0071] Based on the same inventive concept, such as Figure 2 As shown, this embodiment of the invention also provides a pollutant mixed toxicity prediction device, the device comprising: Data acquisition module 201 is used to acquire target data files corresponding to pollutants in surface water bodies from literature databases, environmental monitoring reports and public data platforms using web crawler and text recognition technology, and to construct structured data based on the target data files; In this embodiment, as an optional embodiment, the data acquisition module 201 is specifically used for: Set up a search query for collecting pollutant exposure concentration data, and use web crawling and text recognition technology to collect target data files containing pollutant exposure concentration data from pre-acquired literature databases, environmental monitoring reports, and public data platforms based on the search query.

[0072] In this embodiment, as another optional embodiment, the data acquisition module 201 is further used for: Using a pre-built layout analysis model, the target data file is identified to obtain the table area and the main text area. Using a natural language processing model, information is extracted from the table area and the main text area to obtain the initial structure data. The information extraction includes: sampling point information extraction, sampling time information extraction, pollutant name information extraction, and pollutant exposure concentration value information extraction. Based on the pre-acquired pollutant dictionary, the pollutant names in the initial structure data are mapped to the standardized pollutant names corresponding to the pollutant dictionary; The concentration units corresponding to the pollutant exposure concentration values ​​in the initial structured data are standardized, and the pollutant exposure concentration values ​​are updated accordingly to construct structured data based on the structure of sampling point-sampling time-pollutant name-pollutant exposure concentration value.

[0073] In this embodiment, as another optional embodiment, the data acquisition module 201 is further configured to: All the structured data constructed will be preprocessed and stored in the database to form a structured database of surface water pollutant exposure concentrations covering multiple watersheds, multiple water body types and multiple types of pollutants.

[0074] The matrix construction module 202 is used to determine the target pollutant sequence based on the pollutant names contained in the structured data, perform data preprocessing on the structured data, and construct a pollutant presence / absence matrix for each sampling point based on the pollutant names contained in the preprocessed structured data and the target pollutant sequence. Itemset mining module 203 is used to construct typical pollutant combinations based on the pollutant presence / absence matrix, according to a pre-set frequent itemset mining algorithm. In this embodiment, as an optional embodiment, the itemset mining module 203 is specifically used for: Traverse the pollutant presence / absence matrix, count the number of supporting rows with the same target pollutant set, and determine the support of the target pollutant set based on the number of supporting rows and the total number of rows in the pollutant presence / absence matrix. From each set of target pollutants, obtain the set of target pollutants with a support greater than a preset support threshold, obtain the frequent pollutant combination, and use the frequent pollutant combination as the typical pollutant combination.

[0075] In this embodiment, as another optional embodiment, the itemset mining module 203 is specifically used for: Traverse the pollutant presence / absence matrix, count the number of supporting rows with the same target pollutant set, and determine the support of the target pollutant set based on the number of supporting rows and the total number of rows in the pollutant presence / absence matrix. From each set of target pollutants, obtain the set of target pollutants with a support greater than a pre-set support threshold to obtain frequent pollutant combinations; Based on the support of frequent pollutant combinations, typical mixture combinations containing representative pollutants are selected from each frequent pollutant combination.

[0076] In this embodiment, as an optional embodiment, the threshold for the presence of the target pollutant is set to be at least one of the following indicators: the limit of quantitation for the monitoring method of the target pollutant, the environmental quality benchmark value corresponding to the target pollutant, and the preset quantile based on the statistical distribution of the exposure concentration of the target pollutant.

[0077] The toxicity testing module 204 is used for high-throughput microplate toxicity analysis based on various concentration compositions of typical pollutant combinations to obtain toxicity data of mixtures. In this embodiment, as an optional embodiment, the toxicity testing module 204 is specifically used for: Based on the various concentration compositions of typical pollutant combinations, multiple sets of typical pollutant concentration combinations are constructed. For each typical pollutant concentration combination, a pollutant solution of corresponding concentration is prepared to obtain the pollutant concentration test group corresponding to that typical pollutant concentration combination. High-throughput microplate toxicity analysis was performed for each pollutant concentration test group to obtain the mixture toxicity data for that pollutant concentration test group.

[0078] The toxicity prediction module 205 is used to construct a mixture toxicity prediction model based on a pre-set machine learning algorithm, various concentration compositions of typical pollutant combinations and mixture toxicity data, and to obtain mixture toxicity prediction data of untested pollutant combinations based on the constructed mixture toxicity prediction model. In this embodiment, as an optional embodiment, the toxicity prediction module 205 is specifically used for: The concentration information of each pollutant in the pollutant concentration test group was used as the input to the machine learning algorithm model. Using the toxicity data of the mixture corresponding to the pollutant concentration test group as the output of the machine learning algorithm model, the machine learning algorithm model is trained by using propagation neural network, random forest, graph neural network or gradient algorithm in the backpropagation machine learning algorithm to obtain the mixture toxicity prediction model.

[0079] The toxicity analysis module 206 is used to perform sensitivity analysis on the toxicity prediction data of the mixture of untested pollutant combinations using a pre-set sensitivity analysis method, and to identify high-risk pollutants in the untested pollutant combinations that affect the toxicity of the pollutant mixture based on the sensitivity analysis results.

[0080] In this embodiment, as an optional embodiment, the toxicity analysis module 206 is specifically used for: Using the Sobol global sensitivity analysis method and / or the Morris sensitivity analysis method, based on the mixture toxicity prediction data of untested pollutant combinations, the first-order sensitivity index of each pollutant is calculated. Pollutants whose first-order sensitivity index exceeds the pre-set sensitivity threshold are identified as key pollutants that induce mixture toxicity. A key pollutant group is constructed based on each key pollutant. Based on the constructed mixture toxicity prediction model, key prediction data of mixture toxicity for key pollutant groups are obtained; If the predicted toxicity data of the mixture of untested pollutant combinations is not significantly different from the key predicted toxicity data of the mixture of key pollutant groups, the key pollutant group is identified as a high-risk pollutant group that affects the mixed toxicity of pollutants in the untested pollutant combination. If there is a significant difference between the predicted toxicity data of the mixture of untested pollutant combinations and the key predicted toxicity data of the mixture of key pollutant groups, calculate the higher-order sensitivity index of each pollutant corresponding to the untested pollutant combinations, reconstruct the key pollutant groups based on the higher-order sensitivity index and the pre-set higher-order sensitivity threshold, and execute the step of obtaining the key predicted toxicity data of the mixture of key pollutant groups based on the constructed mixture toxicity prediction model.

[0081] In this embodiment, as an optional implementation, the Sobol global sensitivity analysis method and / or the Morris sensitivity analysis method are used to calculate the first-order sensitivity index of each pollutant based on the toxicity prediction data of the mixture of untested pollutant combinations, including: By changing the concentration of a candidate pollutant in the untested pollutant combination, the corresponding mixture toxicity prediction data is obtained based on the constructed mixture toxicity prediction model. Based on the change in the mixture toxicity prediction data before and after the concentration change, the basic effect value of the candidate pollutant is obtained. The first-order sensitivity index of the candidate pollutant is determined based on the basic effect value of the candidate pollutant, the arithmetic mean of the basic effect values ​​of each candidate pollutant, the arithmetic mean of the absolute values ​​of the basic effect values ​​of each candidate pollutant, and the standard deviation of the basic effect values ​​of each candidate pollutant.

[0082] Based on the same inventive concept, embodiments of the present invention also provide a storage medium storing a computer program thereon, which, when executed by a processor, implements the steps of the pollutant mixed toxicity prediction method in any of the above possible implementations.

[0083] Optionally, the storage medium may be a non-transitory computer-readable storage medium, such as a ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device.

[0084] Based on the same inventive concept, see [link to inventive concept] Figure 3 This invention also provides an electronic device, including a memory 101 (e.g., non-volatile memory), a processor 102, and a computer program stored on the memory 101 and executable on the processor 102. When the processor 102 executes the program, it implements the steps of the pollutant mixed toxicity prediction method in any of the above possible implementations, which can be equivalent to the pollutant mixed toxicity prediction device described above. Of course, the processor can also be used to process other data or perform calculations. This electronic device can be a PC, server, terminal, or other similar device.

[0085] like Figure 3 As shown, the electronic device may also include: memory 103, network interface 104, and internal bus 105. In addition to these components, other hardware may also be included, which will not be described in detail here.

[0086] It should be noted that the above-mentioned pollutant mixed toxicity prediction device can be implemented by software. As a device in a logical sense, it is formed by the processor 102 of the electronic device in which it is located reading the computer program instructions stored in the non-volatile memory into the memory 103 for execution.

[0087] The embodiments of the subject matter and functional operation described in this specification can be implemented in the following ways: digital electronic circuits, tangibly embodied computer software or firmware, computer hardware including the structures disclosed in this specification and their structural equivalents, or combinations thereof. Embodiments of the subject matter described in this specification can be implemented as one or more computer programs, i.e., one or more modules of computer program instructions encoded on a tangible, non-transitory program carrier for execution by a data processing apparatus or for controlling the operation of a data processing apparatus. Alternatively or additionally, the program instructions may be encoded on artificially generated propagation signals, such as machine-generated electrical, optical, or electromagnetic signals, which are generated to encode information and transmit it to a suitable receiving device for execution by the data processing apparatus. The computer storage medium may be a machine-readable storage device, a machine-readable storage substrate, a random or serial access memory device, or combinations thereof.

[0088] The processing and logic flow described in this specification can be executed by one or more programmable computers that execute one or more computer programs to perform corresponding functions by operating on input data and generating output. The processing and logic flow can also be executed by special-purpose logic circuitry—such as FPGA (Field Programmable Gate Array) or ASIC (Application-Specific Integrated Circuit), and the device can also be implemented as special-purpose logic circuitry.

[0089] Suitable computers for executing computer programs include, for example, general-purpose and / or special-purpose microprocessors, or any other type of central processing unit. Typically, the central processing unit receives instructions and data from read-only memory and / or random access memory. The basic components of a computer include a central processing unit for implementing or executing instructions and one or more memory devices for storing instructions and data. Typically, a computer will also include one or more mass storage devices for storing data, such as disks, magneto-optical disks, or optical disks, or the computer will be operatively coupled to such mass storage devices to receive data from or transfer data to them, or both. However, a computer is not required to have such devices. Furthermore, a computer can be embedded in another device, such as a mobile phone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a global positioning system (GPS) receiver, or a portable storage device such as a universal serial bus (USB) flash drive, to name a few.

[0090] Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, such as semiconductor memory devices (e.g., EPROM, EEPROM, and flash memory devices), magnetic disks (e.g., internal hard disks or removable disks), magneto-optical disks, and CD-ROM and DVD-ROM disks. Processors and memory may be supplemented by or incorporated into dedicated logic circuitry.

[0091] While this specification contains numerous specific implementation details, these should not be construed as limiting the scope of any invention or the scope of the claims, but rather are primarily used to describe features of specific embodiments of a particular invention. Certain features described in the various embodiments herein may also be implemented in combination in a single embodiment. Conversely, various features described in a single embodiment may also be implemented separately in various embodiments or in any suitable sub-combination. Furthermore, while features may function in certain combinations as described above and even initially claimed in this way, one or more features from a claimed combination may be removed from that combination in some cases, and a claimed combination may refer to a sub-combination or a variation thereof.

[0092] Similarly, although the operations are depicted in a specific order in the accompanying drawings, this should not be construed as requiring these operations to be performed in the specific order shown or sequentially, or requiring all illustrated operations to be performed to achieve the desired result. In some cases, multitasking and parallel processing may be advantageous. Furthermore, the separation of various system modules and components in the above embodiments should not be construed as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.

[0093] Thus, specific embodiments of the subject matter have been described. Other embodiments are within the scope of the appended claims. In some cases, the actions recited in the claims may be performed in a different order and still achieve the desired result. Furthermore, the processes depicted in the drawings are not necessarily shown in a specific order or sequence to achieve the desired result. In some implementations, multitasking and parallel processing may be advantageous.

[0094] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes the element.

[0095] The above are merely specific embodiments of the present invention, enabling those skilled in the art to understand or implement the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the present invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features claimed herein.

Claims

1. A method for predicting the mixed toxicity of pollutants, characterized in that, include: Using web crawling and text recognition technologies, target data files corresponding to pollutants in surface water bodies are obtained from literature databases, environmental monitoring reports, and public data platforms, and structured data is constructed based on the target data files; Based on the pollutant names contained in the structured data, the target pollutant sequence is determined. The structured data is preprocessed, and a pollutant presence / absence matrix for each sampling point is constructed based on the pollutant names contained in the preprocessed structured data and the target pollutant sequence. Based on a pre-set frequent itemset mining algorithm, typical pollutant combinations are constructed based on the pollutant presence / absence matrix. High-throughput microplate toxicity analysis was performed based on various concentration compositions of typical pollutant combinations to obtain toxicity data for mixtures. Based on pre-set machine learning algorithms, various concentration compositions of typical pollutant combinations, and mixture toxicity data, a mixture toxicity prediction model is constructed. Based on the constructed mixture toxicity prediction model, mixture toxicity prediction data of untested pollutant combinations are obtained. Using a pre-set sensitivity analysis method, sensitivity analysis was performed on the toxicity prediction data of mixtures of untested pollutant combinations. Based on the sensitivity analysis results, high-risk pollutants affecting the mixed toxicity of pollutant combinations in the untested pollutant combinations were identified.

2. The method for predicting the mixed toxicity of pollutants according to claim 1, characterized in that, The construction of structured data based on the target data file includes: Using a pre-built layout analysis model, the target data file is identified to obtain the table area and the main text area. Using a natural language processing model, information is extracted from the table area and the main text area to obtain the initial structure data. The information extraction includes: sampling point information extraction, sampling time information extraction, pollutant name information extraction, and pollutant exposure concentration value information extraction. Based on the pre-acquired pollutant dictionary, the pollutant names in the initial structure data are mapped to the standardized pollutant names corresponding to the pollutant dictionary; The concentration units corresponding to the pollutant exposure concentration values ​​in the initial structured data are standardized, and the pollutant exposure concentration values ​​are updated accordingly to construct structured data based on the structure of sampling point-sampling time-pollutant name-pollutant exposure concentration value.

3. The method for predicting the mixed toxicity of pollutants according to claim 1, characterized in that, The pre-set frequent itemset mining algorithm constructs typical pollutant combinations based on the pollutant presence / absence matrix, including: Traverse the pollutant presence / absence matrix, count the number of supporting rows with the same target pollutant set, and determine the support of the target pollutant set based on the number of supporting rows and the total number of rows in the pollutant presence / absence matrix. From each set of target pollutants, obtain the set of target pollutants with a support greater than a preset support threshold, obtain the frequent pollutant combination, and use the frequent pollutant combination as the typical pollutant combination.

4. The method for predicting the mixed toxicity of pollutants according to claim 1, characterized in that, The pre-set frequent itemset mining algorithm constructs typical pollutant combinations based on the pollutant presence / absence matrix, including: Traverse the pollutant presence / absence matrix, count the number of supporting rows with the same target pollutant set, and determine the support of the target pollutant set based on the number of supporting rows and the total number of rows in the pollutant presence / absence matrix. From each set of target pollutants, obtain the set of target pollutants with a support greater than a pre-set support threshold to obtain frequent pollutant combinations; Based on the support of frequent pollutant combinations, typical mixture combinations containing representative pollutants are selected from each frequent pollutant combination.

5. The method for predicting the mixed toxicity of pollutants according to claim 1, characterized in that, The high-throughput microplate toxicity analysis based on various concentration compositions of typical pollutant combinations yields toxicity data for the mixtures, including: Based on the various concentration compositions of typical pollutant combinations, multiple sets of typical pollutant concentration combinations are constructed. For each typical pollutant concentration combination, a pollutant solution of corresponding concentration is prepared to obtain the pollutant concentration test group corresponding to that typical pollutant concentration combination. High-throughput microplate toxicity analysis was performed for each pollutant concentration test group to obtain the mixture toxicity data for that pollutant concentration test group.

6. The method for predicting the mixed toxicity of pollutants according to any one of claims 1 to 5, characterized in that, The mixture toxicity prediction model is constructed based on a pre-set machine learning algorithm, various concentration compositions of typical pollutant combinations, and mixture toxicity data, including: The concentration information of each pollutant in the pollutant concentration test group was used as the input to the machine learning algorithm model. Using the toxicity data of the mixture corresponding to the pollutant concentration test group as the output of the machine learning algorithm model, the machine learning algorithm model is trained by using propagation neural network, random forest, graph neural network or gradient algorithm in the backpropagation machine learning algorithm to obtain the mixture toxicity prediction model.

7. The method for predicting the mixed toxicity of pollutants according to any one of claims 1 to 5, characterized in that, The method utilizes a pre-set sensitivity analysis approach to perform sensitivity analysis on the toxicity prediction data of mixtures of untested contaminant combinations. Based on the sensitivity analysis results, high-risk contaminants affecting the mixed toxicity of contaminant combinations in the untested contaminant combinations are identified, including: Using the Sobol global sensitivity analysis method and / or the Morris sensitivity analysis method, based on the mixture toxicity prediction data of untested pollutant combinations, the first-order sensitivity index of each pollutant is calculated. Pollutants whose first-order sensitivity index exceeds the pre-set sensitivity threshold are identified as key pollutants that induce mixture toxicity. A key pollutant group is constructed based on each key pollutant. Based on the constructed mixture toxicity prediction model, key prediction data of mixture toxicity for key pollutant groups are obtained; If the predicted toxicity data of the mixture of untested pollutant combinations is not significantly different from the key predicted toxicity data of the mixture of key pollutant groups, the key pollutant group is identified as a high-risk pollutant group that affects the mixed toxicity of pollutants in the untested pollutant combination. If there is a significant difference between the predicted toxicity data of the mixture of untested pollutant combinations and the key predicted toxicity data of the mixture of key pollutant groups, calculate the higher-order sensitivity index of each pollutant corresponding to the untested pollutant combinations, reconstruct the key pollutant groups based on the higher-order sensitivity index and the pre-set higher-order sensitivity threshold, and execute the step of obtaining the key predicted toxicity data of the mixture of key pollutant groups based on the constructed mixture toxicity prediction model.

8. A device for predicting the toxicity of mixed pollutants, characterized in that, The pollutant mixed toxicity prediction device includes: The data acquisition module is used to acquire target data files corresponding to pollutants in surface water bodies from literature databases, environmental monitoring reports and public data platforms using web crawling and text recognition technology, and to construct structured data based on the target data files; The matrix construction module is used to determine the target pollutant sequence based on the pollutant names contained in the structured data, perform data preprocessing on the structured data, and construct a pollutant presence / absence matrix for each sampling point based on the pollutant names contained in the preprocessed structured data and the target pollutant sequence. The itemset mining module is used to construct typical pollutant combinations based on the pollutant presence / absence matrix, using a pre-set frequent itemset mining algorithm. The toxicity testing module is used for high-throughput microplate toxicity analysis based on various concentration compositions of typical pollutant combinations to obtain toxicity data for mixtures. The toxicity prediction module is used to build a mixture toxicity prediction model based on pre-set machine learning algorithms, various concentration compositions of typical pollutant combinations, and mixture toxicity data. Based on the built mixture toxicity prediction model, it obtains mixture toxicity prediction data for untested pollutant combinations. The toxicity analysis module is used to perform sensitivity analysis on the toxicity prediction data of mixtures of untested pollutant combinations using pre-set sensitivity analysis methods, and to identify high-risk pollutants in the untested pollutant combinations that affect the toxicity of the mixtures based on the sensitivity analysis results.

9. A storage medium, characterized in that, A program or instruction is stored on a storage medium, and the program or instruction is executed by a processor to implement the steps of the pollutant mixed toxicity prediction method as described in any one of claims 1 to 7.

10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the pollutant mixed toxicity prediction method according to any one of claims 1 to 7.