Method for predicting harm of new pollutant PFAS to human health based on large language model

By combining graph neural networks and large language models, and integrating PFAS molecular structure and toxicological semantic information, the problem of insufficient accuracy and interpretability in multi-organ toxicity prediction in existing technologies is solved, and joint prediction of multi-organ toxicity and risk level classification are realized.

CN120998484APending Publication Date: 2025-11-21HUANGHUAI LABORATORY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511001860.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-21
Publication Date
2025-11-21

AI Technical Summary

Technical Problem

Existing technologies, when predicting the health hazards of PFAS to humans, rely on a single structural feature expression and lack the fusion of toxicological semantic information, making it difficult to achieve joint prediction of multi-organ toxicity. Furthermore, the prediction results lack interpretability and risk level classification.

Method used

A graph neural network is used to extract the molecular structural features of PFAS, and a large language model is used to analyze the toxicology corpus. The structural vectors and semantic evidence are fused through the Transformer multi-task framework to output the multi-organ toxicity probability and five-level risk level, generating a structured interpretive report.

Benefits of technology

It achieves high accuracy and stability in the joint prediction of multi-organ toxicity, provides interpretable prediction results and clear risk levels, and supports health risk assessment in complex exposure scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120998484A_ABST
    Figure CN120998484A_ABST
Patent Text Reader

Abstract

The invention discloses a method for predicting harm of a new pollutant PFAS to human health based on a large language model, and aims to solve the problems of single structural feature expression, lack of toxicological semantic fusion, insufficient multi-organ toxicity prediction capability and the like in the prior art. According to the method, PFAS multi-dimensional structure features are extracted through combination of a graph neural network and molecular fingerprints, toxicology literature semantic evidence is mined based on a large language model, structures and semantic evidence vectors are fused by adopting a Transform multi-task framework, a multi-label classification model covering six organ systems including the liver, the kidney and the nerves is constructed, the toxicity probability of each organ is output, and the toxicity probability of each organ is calculated. Five risk levels are divided, and interpretable results are provided in combination with structure fragments and literature evidence. According to the method, the accuracy and stability of toxicity prediction are improved, and PFAS substitute screening and health risk management and control are assisted.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of environmental science and machine learning, in particular to a method for predicting the harm of new pollutants PFAS to human health based on a large language model. BACKGROUND

[0002] Perfluoroalkyl and polyfluoroalkyl substances (PFAS) are widely used in industry, consumer products and the environment due to their excellent chemical stability and surface activity. However, PFAS have received increasing attention in recent years due to their persistence, bioaccumulation and potential toxicity, especially to multiple organ systems of the human body. Exposure of the human body to PFAS can cause liver damage, kidney dysfunction, nervous system diseases, reproductive toxicity, endocrine disruption and abnormal immune function, and other health risks.

[0003] Given the wide variety of PFAS and their complex structures, traditional in vitro and in vivo toxicity test-based risk assessment methods have the drawbacks of high cost, long time consumption and ethical limitations, making it urgent to use computational methods to accurately and explainably predict the toxicity of PFAS.

[0004] In recent years, machine learning techniques, especially graph neural networks (GNN), have shown great power in molecular structure representation and property prediction. At the same time, large language models (LLM) have rapidly developed in the field of natural language processing and professional semantic understanding, providing a new technical path for toxicology literature mining and semantic evidence extraction. Combining molecular structure and toxicity semantic information to achieve joint prediction of multi-organ toxicity has become a research hotspot in the field of PFAS risk assessment.

[0005] Currently, research on PFAS toxicity prediction mainly focuses on the following aspects:

[0006] 1. Single toxicity indicator prediction based on molecular fingerprint and traditional machine learning

[0007] Some studies use traditional machine learning methods based on molecular fingerprint encoding such as ECFP (e.g. random forest, support vector machine) to predict single toxicity indicators of PFAS such as acute toxicity and carcinogenicity. Although this method has high computational efficiency, it has limited prediction accuracy due to the inability to fully capture molecular topological structure and key functional group information, and often ignores the synergistic toxicity risk of multiple organ systems.

[0008] 2. Application of graph neural networks in PFAS molecular structure feature extraction

[0009] Some frontier research introduces GNN, directly learns the complex relationship between atoms and bonds based on molecular graph structure, and improves the structure expression ability and toxicity prediction performance. However, these models mostly rely only on structure data, without combining external toxicology semantic information, making it difficult to fully utilize rich literature knowledge support.

[0010] 3. Toxicology literature retrieval and correlation analysis based on keyword matching

[0011] Traditional toxicology literature mining mostly uses keyword search, text matching and other methods to obtain information about the association between PFAS and organ toxicity. Such methods have a rough understanding of semantics and are difficult to capture implicit toxicity mechanisms and complex contextual causal relationships, and there are problems of missed detection and false detection.

[0012] Although existing technologies have made some progress in the field of PFAS toxicity prediction, there are still the following main deficiencies:

[0013] 1. Single structure feature expression, difficult to fully capture molecular complexity

[0014] Most methods rely only on traditional molecular fingerprints or simple molecular descriptors, and do not fully utilize molecular topological structure information, resulting in limited recognition ability of key structural fragments and functional groups, affecting the accuracy of toxicity prediction.

[0015] 2. Lack of deep integration of toxicology semantic information

[0016] Existing toxicity prediction models usually use structure data or text data independently, and lack the integration of molecular structure features and the implicit toxicity mechanisms and organ toxicity information in a large number of professional toxicology literature, resulting in insufficient understanding of complex toxicity mechanisms by the model.

[0017] 3. Literature mining technology is rough, and the semantic understanding ability is insufficient

[0018] Most literature retrieval relies on keyword matching, ignoring contextual semantics and implicit relationships, which can easily cause information to be missed or misdetected, making it difficult to effectively extract high-quality evidence to support toxicity prediction.

[0019] 4. Insufficient joint prediction ability of multi-organ toxicity

[0020] Existing models mostly focus on a single toxicity indicator, lack simultaneous assessment and correlation analysis of toxicity risks to multiple human organ systems, and are difficult to meet the health risk assessment needs under complex exposure conditions.

[0021] 5. Lack of effective explanation and risk classification of prediction results

[0022] The prior art often cannot provide intuitive correlation between structural features and literature evidence, lacks mechanistic explanation of toxicity prediction results and scientifically reasonable risk level division, and limits practical application value and regulatory reference significance. SUMMARY

[0023] In view of the problems of long job completion time, low scheduling resource utilization, and uneven task load in solving the above multi-tractor cooperative scheduling process, the application provides a multi-tractor cooperative scheduling method, constructs a scheduling model with dynamic task cooperation and multi-objective optimization capability, faces the typical scene of multiple tractors working on multiple farmland tasks, and considers two core optimization objectives of minimizing the maximum task completion time and minimizing the total scheduling path cost, thereby improving the overall scheduling efficiency.

[0024] The technical scheme adopted by the application is: a method for predicting the harm of new pollutants PFAS to human health based on a large language model, comprising the following steps:

[0025] Step 1, input the SMILES structure expression of the PFAS molecule, extract atomic, bond and topological relationship features to generate a structure vector through GNN, and extract toxicity-related functional group substructures in combination with molecular fingerprint encoding;

[0026] Step 2, analyze the toxicology corpus through the fine-tuned large language model to establish a semantic mapping relationship between the structure features of PFAS and the organ-level toxicity labels;

[0027] Step 3, taking the PFAS structure name or feature fragment as a search input, extracting a structured semantic evidence vector from a toxicology database to generate four-tuple information including the target organ of toxicity, the description of the toxicity mechanism, the evidence strength and the literature source;

[0028] Step 4, fuse the structure vector of step 1 and the semantic evidence vector of step 3, input into the Transformer multi-task framework, and output the probability values of liver toxicity, kidney toxicity, neurotoxicity, reproductive toxicity, endocrine disruption and immunotoxicity;

[0029] Step 5, according to the toxicity probability value, divide the five-level risk level, and generate an explanatory report supporting the structure feature description and literature abstract.

[0030] Further, the step 1 specifically comprises:

[0031] Step 1-1, collect the SMILES format structure information of the target PFAS and perform standardization processing;

[0032] Step 1-2, convert the standardized structure into a molecular graph through a graph neural network, extract atomic, bond and topological relationship features, and generate a first structure feature vector;

[0033] Step 1-3, extract functional groups and substructure features by molecular fingerprinting method, generate the second structure feature vector;

[0034] Step 1-4, automatically annotate key structural units related to organ toxicity, generate toxicity correlation feature labels;

[0035] Step 1-5, fuse the first structure feature vector, the second structure feature vector and the toxicity correlation feature label to form a multi-dimensional structure feature expression, which provides input data for the PFAS toxicity prediction model.

[0036] Further, the key structural units in step 1-4 include -CF3, -SO3H or long-chain fluorinated alkyl groups.

[0037] Further, the toxicology corpus in step 2 includes ToxRefDB, CTD, PubChem toxicity data, EPA evaluation report and PubMed screened literature.

[0038] Further, the structured semantic evidence in step 3 is generated by the following methods:

[0039] (1) Based on semantic understanding, open corpus of PubMed, EPA report and toxicology database is retrieved;

[0040] (2) Automatically identify toxicity relationship expressions across contexts, implicit expressions or indirect reasoning;

[0041] (3) Output structured quadruples containing target organs, toxicity mechanisms, evidence types and literature sentences.

[0042] Further, the Transformer multi-task learning framework in step 4 includes:

[0043] Input layer, concatenate GNN structure vector and semantic evidence vector to form multi-modal input;

[0044] Transformer encoding layer, capture cross-modal feature dependency through Transformer self-attention mechanism;

[0045] Multi-label output layer, use Sigmoid activation function to independently calculate the probability of six organ toxicities.

[0046] Further, the five-level risk grade classification standard in step 5 is:

[0047] Extremely high risk: toxicity probability ≥ 0.85;

[0048] High risk: toxicity probability ∈ [0.70, 0.85);

[0049] Medium risk: toxicity probability ∈ [0.50, 0.70);

[0050] Low risk: toxicity probability ∈ [0.30, 0.50);

[0051] Very low risk: toxicity probability < 0.30.

[0052] Further, the explanatory report generated in step 5 includes:

[0053] (1) Key toxic structure fragment annotation;

[0054] (2) Support literature abstract and PMID number;

[0055] (3) Organ toxicity probability distribution graph and risk level mark.

[0056] Further, the training of the Transformer multi-task framework uses a weighted binary cross-entropy loss function, and the label imbalance processing strategy is used to alleviate the imbalance problem of organ toxicity data.

[0057] The present application has the beneficial effects of:

[0058] 1. The structure-semantics fusion prediction capability is better:

[0059] Existing technologies mostly use QSAR models based only on molecular structure, without fusing literature semantic evidence, and are difficult to capture cross-modal information. The present application innovatively fuses the molecular structure features extracted by the graph neural network with the toxicology semantic features analyzed by the large language model, realizes the joint modeling of structure and literature evidence through the Transformer multi-task learning framework, and significantly improves the accuracy and stability of multi-organ toxicity prediction.

[0060] 2. It has the ability of joint prediction of multi-organ toxicity:

[0061] Most existing technologies can only output a single toxicity label (such as LD50, carcinogenicity), and lack of specific organ evaluation. The present application constructs a multi-label classification model covering six organ systems of liver, kidney, nerve, reproduction, endocrine and immune, which can output the toxicity probability of each organ at the same time, realize the fine evaluation of joint toxicity risk of multi-organ system, and meet the health risk assessment needs under complex exposure scenarios.

[0062] 3. Stronger semantic causal reasoning and explainability:

[0063] The prediction results of the prior art are mostly "black box" outputs, lacking tracking and explanation of the source of toxicity. The present application can automatically identify key structural fragments related to toxicity (such as -CF3, -SO3H, long-chain fluoralkyl) through the semantic causal reasoning ability of large language models, and match corresponding literature evidence (including PMID numbers, core sentence segments, etc.), generating a complete causal chain of "structure-toxicity-literature", greatly enhancing the credibility and explainability of the prediction results.

[0064] 4. Higher risk level classification and regulatory adaptability:

[0065] The output results of the prior art are not uniform, making it difficult to be directly used for risk classification or policy decision. The present application designs a five-level risk classification system based on multi-dimensional toxicity probability values, outputs structured results containing structure diagrams, toxicity labels, risk levels, mechanism explanations, and cited literature, which can be directly adapted to research analysis, risk supervision, and industrial screening scenarios, providing clear references for safety substitute screening and health risk management.

[0066] 5. Automation, high throughput, and strong universality:

[0067] The prior art often relies on manual rules or single-point modeling, making it difficult to handle a large number of compounds or respond quickly to new pollutants. The present application realizes the full-process automation from structure input to toxicity prediction and explanation output, supports batch analysis of hundreds of thousands of PFAS structures, and the model can be extended to polychlorinated biphenyls (PCBs), polycyclic aromatic hydrocarbons (PAHs), drugs, and other structural pollutants, with wide applicability. BRIEF DESCRIPTION OF DRAWINGS

[0068] Figure 1 The method flowchart of the present application;

[0069] Figure 2 The Transformer multi-task framework model schematic diagram in the present application. DETAILED DESCRIPTION

[0070] The present application will be further described below in conjunction with the accompanying drawings.

[0071] As shown in Figure 1 , the present application is a method for predicting the harm of new pollutants PFAS to human health based on a large language model, comprising the following steps:

[0072] Step 1, PFAS structure data collection and coding; comprising the following steps:

[0073] Step 1-1, collect the structure information of PFAS, support input in SMILES standard format, and perform standardization processing on the structure to ensure the consistency and universality of the molecular expression;

[0074] Step 1-2, convert PFAS structures into molecular graph form through GNN, extract atomic, bond, and topological relationship features, and generate the first structure feature vector with rich structural information;

[0075] Step 1-3, extract functional group and substructure features through molecular fingerprint (e.g. ECFP) encoding method, and generate the second structure feature vector;

[0076] Step 1-4, automatically label key structural units related to organ toxicity, including -CF3, -SO3H or long-chain fluorinated alkyl groups, and generate toxicity-related feature labels;

[0077] Step 1-5, fuse the first structure feature vector, the second structure feature vector, and the toxicity-related feature labels to form a multi-dimensional structure feature representation, providing input data for the PFAS toxicity prediction model.

[0078] Step 2, large language model construction and training; based on existing general large language model (such as GPT-4), fine-tune it to understand and analyze professional corpus in the field of toxicology. Training corpus sources include ToxRefDB, Comparative Toxicogenomics Database (CTD), PubChem toxicity data, U.S. Environmental Protection Agency (EPA) toxicological evaluation reports, and a large number of manually screened toxicity research literature from PubMed. Through supervised and instruction fine-tuning training on the above corpus, the model can automatically identify the toxicity relationship between PFAS compounds and specific organs of the human body, extract organ-level toxicity labels such as "liver damage", "neurotoxicity", "endocrine disruption", etc. At the same time, the model supports the association between molecular structure features (such as functional groups, fluorocarbon chain length, functional group combinations) and semantic information, achieving semantic mapping and reasoning ability between structure and toxicity mechanism, providing basic support for subsequent high-throughput toxicity prediction.

[0079] Step 3, literature semantic evidence retrieval; use standardized PFAS structure names, structure fragments or fingerprint features as retrieval inputs, combined with the semantic reasoning ability of the large language model, efficiently retrieve literature fragments related to specific organ toxicity from PubMed, EPA reports, toxicology databases and other open corpus. Compared with traditional keyword matching methods, this module has semantic understanding ability, and can identify toxicity relationship expressions expressed in cross-context, implicit or indirect reasoning, such as "PFHxA disrupts estrogen homeostasis in exposed populations" or "long-chain PFAS are associated with liver enzyme elevation in cohort studies".

[0080] The system can automatically extract the following core information and structure the output:

[0081] (1) Toxic target organs (such as liver, kidney, nerve, reproduction, endocrine, immune);

[0082] (2) Toxicity mechanism description (such as endocrine disruption, immune suppression, apoptosis, etc.);

[0083] (3) Evidence strength (such as whether it is human data, animal experiments, in vitro studies, etc.);

[0084] (4) Literature sources and sentences (including PMID number, publication year, core sentence, etc.);

[0085] Through the structured toxicity evidence output, this module not only improves the interpretability and traceability of the model prediction results, but also can be used to construct a knowledge four-tuple relationship graph of "PFAS-organ-toxicity-literature", supporting the construction of toxicity correlation network and mechanism analysis.

[0086] Step 4, fuse the structure vector of step 1 and the semantic evidence vector of step 3, input into the Transformer multi-task framework, and output the probability values of liver toxicity, kidney toxicity, nerve toxicity, reproductive toxicity, endocrine disruption and immune toxicity.

[0087] As shown in Figure 2 , the Transformer multi-task framework includes:

[0088] (1) Input layer:

[0089] Structure feature vector: learned by GNN from the graph structure of PFAS molecule, which can capture chemical information such as molecular bonding relationship, atomic environment and spatial configuration.

[0090] Semantic evidence vector: based on large language model or natural language processing method, the knowledge related to organ toxicity is extracted from relevant literature and database, reflecting the context information in the field of toxicology.

[0091] The above two vectors are merged through the fusion layer to form a unified multi-modal feature representation.

[0092] (2) Transformer encoding layer:

[0093] Use the Transformer self-attention mechanism to capture the complex dependency relationship between the dimensions of the input features, and enhance the model's ability to recognize key structural fragments and semantic clues.

[0094] (3) Multi-label output layer:

[0095] The Transformer output is mapped to six organ toxicity dimensions through a fully connected layer, each corresponding to an independent output unit. A Sigmoid activation function is used to independently calculate the probability value of each toxicity dimension, ranging from 0 to 1, representing the probability of toxicity of that organ system. This design ensures independent prediction between multiple labels, while the model can capture potential correlations between labels by sharing parameters.

[0096] Binary cross-entropy is used as the loss function for multi-label classification, and the loss is calculated independently for each label and summed for all labels. The model parameters are optimized through overall backpropagation. Label imbalance handling strategies such as weighted BCE, sample resampling, etc. are used to alleviate the bias caused by the difference in data distribution of different organ toxicities.

[0097] Step 5: Divide the five-level risk grade according to the toxicity probability value:

[0098]

[0099] The system automatically outputs the corresponding structural feature explanation and literature fragment summary for each prediction, such as:

[0100] Liver toxicity risk: high (0.82);

[0101] Supporting structural features: C8 chain length CF3 group exists;

[0102] Supporting literature: [PMID:31504583] "Long-chain PFAS significantly increased liver enzymes in exposed workers..."

[0103] The final system generates structured prediction results: PFAS molecular structure diagram; multi-organ toxicity probability value; risk level label; supporting mechanism feature fragment; literature citation summary and link, output results to adapt to the call of scientific research, risk supervision or industrial screening platform.

[0104] The application has multiple dimensional significant advantages: through the architecture of GNN and LLM fusion, the joint input of PFAS structure vector and semantic vector is realized, providing a strong foundation for cross-modal fusion of multi-organ toxicity joint prediction; relying on the semantic evidence automatic retrieval and reasoning module, the organ toxicity evidence related to the PFAS structure can be accurately extracted from a large amount of toxicological literature, effectively assisting model reasoning and enhancing prediction basis; the multi-label Transformer prediction mechanism is adopted, the probability modeling of multiple organ toxicities can be realized at the same time, and the feature explanation is realized combined with the attention mechanism, the accuracy and traceability of multi-target prediction are improved; with the help of toxicity mechanism explanation and risk level automatic generation mechanism, through structure fragment labeling, semantic evidence matching and clear risk grading standard, the explainability of the system is greatly improved; in combination with the design of a structured output interface, API or batch processing calling is supported, and rich information such as structure diagram, toxicity label, risk level, mechanism explanation and reference literature can be output, which is suitable for multi-scene application requirements.

Claims

1. A method for predicting the health hazards of emerging pollutants PFAS based on large language models, characterized in that, Includes the following steps: Step 1: Input the SMILES structural expression of the PFAS molecule, extract the atomic, bond and topological relationship features through GNN to generate a structural vector, and combine the molecular fingerprint encoding to extract the substructure of toxicity-related functional groups. Step 2: Analyze the toxicology corpus using a fine-tuned large language model to establish a semantic mapping relationship between PFAS structural features and organ-level toxicity labels; Step 3: Using the PFAS structure name or feature fragment as the search input, extract the structured semantic evidence vector from the toxicology database and generate a structured semantic quadruple containing the target organ of toxicity, description of the toxic mechanism, strength of evidence and literature source. Step 4: Combine the structural vector from Step 1 with the semantic evidence vector from Step 3, input them into the Transformer multi-task framework, and output the probability values ​​of hepatotoxicity, nephrotoxicity, neurotoxicity, reproductive toxicity, endocrine disruption, and immunotoxicity. Step 5: Divide the risk into five levels based on the toxicity probability value, and generate an explanatory report that supports the structural feature description and literature abstract.

2. The method for predicting the health hazards of novel pollutant PFAS based on a large language model according to claim 1, characterized in that, Step 1 specifically includes: Step 1-1: Collect the SMILES format structure information of the target PFAS and perform standardization processing; Steps 1-2: The standardized structure is transformed into a molecular graph using a graph neural network, and features of atoms, bonds and topological relationships are extracted to generate the first structural feature vector. Steps 1-3: Extract functional groups and substructure features using molecular fingerprinting to generate a second structure feature vector; Steps 1-4: Automatically label key structural units related to organ toxicity and generate toxicity-related feature identifiers; Steps 1-5: The first structural feature vector, the second structural feature vector, and the toxicity-related feature identifier are fused to form a multidimensional structural feature expression, providing input data for the PFAS toxicity prediction model.

3. The method for predicting the health hazards of novel pollutant PFAS based on a large language model according to claim 2, characterized in that, The key structural units in steps 1-4 include -CF3, -SO3H, or long-chain fluoroalkyl groups.

4. The method for predicting the health hazards of novel pollutant PFAS based on a large language model according to claim 1, characterized in that, The toxicological corpus sources mentioned in step 2 include ToxRefDB, CTD, PubChem toxicity data, EPA assessment reports, and PubMed screening literature.

5. The method for predicting the health hazards of novel pollutant PFAS based on a large language model according to claim 1, characterized in that, The structured semantic evidence mentioned in step 3 is generated in the following way: (1) Retrieve open corpora from PubMed, EPA reports and toxicology databases based on semantic understanding; (2) Automatically identify toxic relationship expressions that cross context, implicit expression, or indirect inference; (3) Output a structured quadruple containing the target organ, toxic mechanism, evidence type and literature segment.

6. The method for predicting the health hazards of the novel pollutant PFAS based on a large language model according to claim 1, characterized in that, Its features are, The Transformer multi-task learning framework described in step 4 includes: The input layer concatenates the GNN structure vector with the semantic evidence vector to form a multimodal input; The Transformer encoding layer captures cross-modal feature dependencies through the Transformer self-attention mechanism; The multi-label output layer uses the Sigmoid activation function to independently calculate the toxicity probability of the six major organs.

7. The method for predicting the health hazards of novel pollutant PFAS based on a large language model according to claim 1, characterized in that, The five-level risk classification criteria for step 5 are as follows: Extremely high risk: probability of toxicity ≥ 0.85; High risk: toxicity probability ∈ [0.70, 0.85); Medium risk: toxicity probability ∈ [0.50, 0.70); Low risk: toxicity probability ∈ [0.30, 0.50); Extremely low risk: probability of toxicity < 0.

30.

8. The method for predicting the health hazards of novel pollutant PFAS based on a large language model according to claim 1, characterized in that, The explanatory report generated in step 5 includes: (1) Labeling of key toxic structural fragments; (2) Supporting document abstracts and PMID numbers; (3) Probability distribution of toxicity in each organ and risk level label.

9. The method for predicting the health hazards of the novel pollutant PFAS based on a large language model according to claim 6, characterized in that, The Transformer multi-task framework is trained using a weighted binary cross-entropy loss function, and the imbalance of organ toxicity data is mitigated through a label imbalance handling strategy.